Job Description
This role is for one of the Weekday's clients
Min Experience: 3+ years
Location: Bengaluru
JobType: full-time
We are looking for a highly skilled Big Data Engineer with 3+ years of experience in building scalable data pipelines and distributed systems. The ideal candidate will have strong expertise in Apache Spark (Scala), experience working across on-premise and AWS environments, and a solid understanding of large-scale data processing in AdTech ecosystems.
This role involves working on high-volume datasets (billions of records), optimizing distributed jobs, and contributing to the design of robust data infrastructure powering analytics and identity-driven use cases.
Requirements
Key Responsibilities:
- Design, develop, and optimize large-scale batch data pipelines using Apache Spark (Scala)
- Process and transform high-volume datasets (TBs of data) in distributed environments
- Build and maintain data pipelines across hybrid infrastructure (On-Prem + AWS)
- Work with object storage systems such as S3 and MinIO for efficient data access and storage
- Develop reusable and scalable data processing frameworks
- Optimize Spark jobs for performance (memory tuning, partitioning, shuffling, etc.)
- Manage and orchestrate workloads using HashiCorp Nomad
- Integrate data pipelines with PostgreSQL and other downstream systems
- Ensure data quality, consistency, and reliability across pipelines
- Troubleshoot production issues and perform root cause analysis
- Contribute to system design discussions, especially for high-scale AdTech use cases (identity resolution, user profiling, etc.)
Required Skills:
- Strong programming experience in Scala
- Good working knowledge of Python (for auxiliary tasks, scripting, or ML integration)
- Deep expertise in Apache Spark (Core and SQL)
- Strong understanding of distributed data processing and large-scale systems
- Experience working with AWS (S3, EMR or equivalent ecosystem) and on-prem clusters
- Hands-on experience with object storage systems (S3 / MinIO)
- Experience with HashiCorp Nomad or similar orchestration tools
- Solid understanding of data modeling and ETL pipeline design
- Experience working with PostgreSQL or similar relational databases
- Strong debugging and performance tuning skills for Spark jobs
- Familiarity with Unix/Linux environments and shell scripting
Good to Have:
- Experience in AdTech, Identity Graph, or User Profiling systems
- Exposure to machine learning pipelines or feature engineering workflows
- Experience with data lake architectures
- Understanding of cost optimization and resource management in AWS
Tech Stack Summary:
- Languages: Scala (Primary), Python (Secondary)
- Processing: Apache Spark (Core and SQL)
- Infrastructure: AWS and On-Prem
- Storage: S3, MinIO
- Orchestration: HashiCorp Nomad
- Database: PostgreSQL
Must-have skills
Spark, SQL, Scala
Good-to-have skills
Python, AWS, Big Data