Job Title: Data Engineer - PySpark, Python, SQL, Git, AWS Services – Glue, Lambda, Step Functions, S3, Athena.
Job Description:
We are seeking a talented Data Engineer with expertise in PySpark, Python, SQL, Git, and AWS to join our dynamic team. The ideal candidate will have a strong background in data engineering, data processing, and cloud technologies. You will play a crucial role in designing, developing, and maintaining our data infrastructure to support our analytics.
Preferred Skills:
1. Knowledge of data warehousing concepts and data modeling.
2. Familiarity with big data technologies like Hadoop and Spark.
3. AWS certifications related to data engineering.
Join our team and contribute to our mission of turning data into actionable insights. If you're a motivated data engineer with expertise in PySpark, Python, SQL, Git, and AWS, we want to hear from you. Apply now to be part of our innovative and dynamic data engineering team.
Responsibilities:
1. Develop and maintain ETL pipelines using PySpark and AWS Glue to process and transform large volumes of data efficiently.
2. Collaborate with analysts to understand data requirements and ensure data availability and quality.
3. Write and optimize SQL queries for data extraction, transformation, and loading.
4. Utilize Git for version control, ensuring proper documentation and tracking of code changes.
5. Design, implement, and manage scalable data lakes on AWS, including S3, or other relevant services for efficient data storage and retrieval.
6. Develop and optimize high-performance, scalable databases using Amazon DynamoDB.
7. Proficiency in Amazon QuickSight for creating interactive dashboards and data visualizations.
8. Automate workflows using AWS Cloud services like event bridge, step functions.
9. Monitor and optimize data processing workflows for performance and scalability.
10. Troubleshoot data-related issues and provide timely resolution.
11. Stay up-to-date with industry best practices and emerging technologies in data engineering.
Qualifications:
1. Bachelor's degree in Computer Science, Data Science, or a related field. Master's degree is a plus.
2. Strong proficiency in PySpark and Python for data processing and analysis.
3. Proficiency in SQL for data manipulation and querying.
4. Experience with version control systems, preferably Git.
5. Familiarity with AWS services, including S3, Redshift, Glue, Step Functions, Event Bridge, CloudWatch, Lambda, Quicksight, DynamoDB, Athena, CodeCommit etc.
6. Excellent problem-solving skills and attention to detail.
7. Strong communication and collaboration skills to work effectively within a team.
8. Ability to manage multiple tasks and prioritize effectively in a fast-paced environment.
Responsibilities:
1. Develop and maintain ETL pipelines using PySpark and AWS Glue to process and transform large volumes of data efficiently.
2. Collaborate with analysts to understand data requirements and ensure data availability and quality.
3. Write and optimize SQL queries for data extraction, transformation, and loading.
4. Utilize Git for version control, ensuring proper documentation and tracking of code changes.
5. Design, implement, and manage scalable data lakes on AWS, including S3, or other relevant services for efficient data storage and retrieval.
6. Develop and optimize high-performance, scalable databases using Amazon DynamoDB.
7. Proficiency in Amazon QuickSight for creating interactive dashboards and data visualizations.
8. Automate workflows using AWS Cloud services like event bridge, step functions.
9. Monitor and optimize data processing workflows for performance and scalability.
10. Troubleshoot data-related issues and provide timely resolution.
11. Stay up-to-date with industry best practices and emerging technologies in data engineering.
Qualifications:
1. Bachelor's degree in Computer Science, Data Science, or a related field. Master's degree is a plus.
2. Strong proficiency in PySpark and Python for data processing and analysis.
3. Proficiency in SQL for data manipulation and querying.
4. Experience with version control systems, preferably Git.
5. Familiarity with AWS services, including S3, Redshift, Glue, Step Functions, Event Bridge, CloudWatch, Lambda, Quicksight, DynamoDB, Athena, CodeCommit etc.
6. Excellent problem-solving skills and attention to detail.
7. Strong communication and collaboration skills to work effectively within a team.
8. Ability to manage multiple tasks and prioritize effectively in a fast-paced environment.

Choosing a digital partner is about more than capabilities — it’s about collaboration and character.
Unrealistic overhauls and off-the-shelf products ignore what matters most — your unique needs, culture, goals, and your legacy data and technology environments.
At EXL, our collaboration is built on ongoing listening and learning to adapt our methodologies. We’re your business evolution partner—tailoring solutions that make the most of data to make better business decisions and drive more intelligence into your increasingly digital operations.
Whether your goals are scaling the use of AI and digital, redesign operating models, or driving better and faster decisions, we’re here to partner with you to help you gain—and maintain—competitive advantage with efficient, sustainable models at scale.
Our expertise in transformation, data science, and change management helps make your business more efficient and effective, improve customer relationships and enhance revenue growth. Instead of focusing on multi-year, resource- and time-intensive platform designs or migrations, we look deeper at your entire value chain to integrate strategies with impact.
We use our specialization in analytics, digital interventions, and operations management—alongside deep industry expertise — to deliver solutions that help you outperform the competition.
At EXL, it’s all about outcomes—your outcomes—and delivering success on your terms. Share your goals with us and together, we’ll optimize how you leverage data to drive your business forward.
For more information, visit www.exlservice.com.