We are looking for an experienced Senior Data Engineer with 8+ years of hands-on experience in designing, developing, and maintaining scalable data platforms, data pipelines, and analytics solutions. The ideal candidate will have strong expertise in Python/SQL, ETL/ELT, cloud data platforms, data warehousing, distributed data processing, orchestration, and data architecture
The candidate will work closely with Data Scientists, BI Developers, Software Engineers, Product Managers, and business stakeholders to build reliable, secure, high-performance data solutions that support business-critical analytics and AI/ML initiatives.
Design, develop, and maintain scalable and reliable batch and real-time data pipelines
Build robust ETL/ELT workflows to ingest, transform, validate, and distribute data from multiple sources.
Develop highly optimized and complex SQL queries, stored procedures, and data transformations
Design and implement data warehouses, data lakes, lakehouses, and dimensional data models
Work with large datasets using distributed processing technologies such as Apache Spark/PySpark
Develop data pipelines using orchestration tools such as Apache Airflow, Azure Data Factory, AWS Glue, or similar platforms
Implement data solutions on major cloud platforms such as AWS, Azure, or GCP
Design and optimize cloud data platforms and services such as Amazon Redshift, Snowflake, Databricks, Azure Synapse, BigQuery, or equivalent technologies
Implement data quality, data validation, reconciliation, monitoring, and observability frameworks.
Develop solutions for incremental data processing, CDC, slowly changing dimensions, partitioning, and performance optimization
Build and maintain real-time/streaming data pipelines using technologies such as Kafka, Kinesis, or equivalent tools.
Implement appropriate data security, governance, access control, encryption, and compliance practices.
Collaborate with data architects to translate business requirements into scalable technical solutions.
Perform performance tuning of data pipelines, databases, Spark jobs, and cloud data workloads.
Establish and maintain CI/CD practices for data engineering workflows.
Write unit, integration, and data-quality tests to ensure reliability of production pipelines.
Troubleshoot production data issues and participate in incident resolution and root-cause analysis.
Conduct code reviews and promote engineering best practices across the data engineering team.
Mentor junior and mid-level data engineers and provide technical leadership.
Document data architecture, pipeline designs, data models, operational procedures, and technical decisions.
Stay current with emerging technologies in cloud, big data, data engineering, data platforms, and AI/ML
Strong proficiency in Python
Advanced SQL skills.
Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle
Experience with NoSQL databases such as MongoDB, DynamoDB, Cassandra, or similar is advantageous.
Strong understanding of database design, indexing, query optimization, and transaction management.
Strong experience with Apache Spark / PySpark
Experience with Hadoop ecosystem technologies is desirable.
Understanding of distributed computing, partitioning, parallel processing, and performance optimization.
Extensive experience building ETL/ELT pipelines
Experience with tools such as:
Apache Airflow
Azure Data Factory
AWS Glue
dbt
Informatica
Talend
SSIS
Experience handling structured, semi-structured, and unstructured data.
Strong experience with at least one major cloud platform:
AWS
S3
Glue
EMR
Redshift
Lambda
Kinesis
Athena
IAM
Azure
Azure Data Factory
Azure Data Lake Storage
Azure Databricks
Azure Synapse Analytics
Azure Functions
Event Hubs
Key Vault
GCP
BigQuery
Cloud Storage
Dataflow
Dataproc
Pub/Sub
Cloud Composer
Strong understanding of data warehouse architecture
Experience with Snowflake, Databricks, Redshift, Synapse, BigQuery, or equivalent.
Expertise in:
Star and Snowflake schemas
Fact and dimension tables
Slowly Changing Dimensions (SCD)
Data marts
Data lakes
Lakehouse architecture
Partitioning and clustering
Data modeling
Experience with Apache Kafka or equivalent streaming platforms.
Understanding of producers, consumers, topics, partitions, offsets, consumer groups, and schema management.
Experience developing real-time or near-real-time data processing pipelines.
Experience with Git/GitHub/GitLab/Bitbucket
Experience with CI/CD pipelines
Knowledge of Docker and Kubernetes is desirable.
Experience with Infrastructure as Code tools such as Terraform is advantageous.
Familiarity with automated testing, deployment, monitoring, and observability.
Understanding of data governance, metadata management, lineage, data cataloging, and data quality
Experience implementing role-based access control and secure data access.
Knowledge of privacy and compliance requirements such as GDPR, CCPA, HIPAA, or equivalent regulations, depending on business requirements.
Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
8+ years of professional experience in Data Engineering, Big Data, or related disciplines
Experience leading data engineering projects from requirements through production deployment.
Experience working in Agile/Scrum environments.
Experience with BI and analytics platforms such as Power BI, Tableau, Looker, or similar
Understanding of Machine Learning data pipelines and MLOps is a plus.
Experience with modern data stack technologies such as dbt, Databricks, Snowflake, Kafka, and cloud-native services is highly desirable.
Strong analytical and problem-solving skills.
Excellent understanding of data architecture and engineering principles.
Ability to translate complex business requirements into scalable technical solutions.
Strong communication and stakeholder-management skills.
Ability to work independently and collaboratively in a distributed team.
Strong ownership and accountability for production data systems.
Ability to mentor engineers and provide technical leadership.
Focus on performance, reliability, scalability, security, and maintainability.
The ideal candidate should demonstrate experience with:
Enterprise-scale data platforms
High-volume data processing
Batch and streaming architectures
Cloud migration and modernization
Data warehouse and lakehouse implementations
ETL/ELT modernization
Data quality and observability
Performance and cost optimization
API and database integrations
Real-time analytics
Data governance and security
Production support and incident management
Technical leadership and mentoring

The Hudson Group comprises of:
HudsonIT Consultancy Ltd – a dynamic Software Solutions and IT Staffing firm in United States and
Hudson Manpower Inc – specializing in offering comprehensive recruitment services for technical industries worldwide, ensuring quality hires for various sectors.
We began our journey in early 2019 and take pride in our ability to transform businesses through the seamless integration of cutting-edge digital solutions and outstanding staffing services. With a rich history of staffing experience, we have not only adapted but thrived in the ever-changing landscape of technological innovation.