o Python , Scala , Apache Spark (DataFrames, Spark SQL, performance tuning)
o SQL (advanced joins, window functions, query tuning)
o ADO adherence
o Basics of Java – Good to have
o Data warehousing concepts: star/snowflake schemas, facts & dimensions
o Data modelling & mapping understanding
o Good understanding on ETL & data processing
o Designing batch and streaming pipelines
o Data integration - files, message queues etc
o Hadoop ecosystem (HDFS, Hive) ;Distributed computing concepts (partitioning, shuffling etc)
o Data validation, profiling, and monitoring
o DQ Controls and framework alignment
o Basic knowledge of data governance, security, and compliance controls
o Version control and branching strategies
o Automated builds, tests and deployments; Pipeline-as-code (e.g. YAML-based pipelines)
o Managing artefacts, versioning and rollbacks
o Release Planning & Coordination; Code Validation & Post Deployment Checks
o Rollback & Incident Handling
o Continuous Improvement of Release Process
null
null

HCLTech is a global technology company, home to more than 226,600 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending September 2025 totaled $14.2 billion. To learn how we can supercharge progress for you, visit hcltech.com.