We are looking for experienced SRE leader to drive enterprise-wide reliability transformation across cloud, digital, SaaS, industrial, AI-enabled, and mission-critical platforms. This role combines technical expertise with strategic transformation leadership to modernize operations through SRE, AI-driven observability, autonomous operations, platform engineering, and intelligent automation practices.
As a trusted advisor and hands-on transformation leader, you will partner with engineering, operations, cloud, platform, architecture, security and AI-engineering teams to institutionalize modern Site Reliability Engineering practices at scale.
Lead and drive enterprise-wide SRE transformation initiatives across business applications and platforms.
Assess current operational maturity and define SRE adoption roadmaps, operating models, governance frameworks, reliability standards, and implementation strategies.
Coach and mentor engineering, operations, cloud, platform, and application teams on SRE principles, SLI/SLO frameworks, Error Budgets, reliability engineering, and operational excellence.
Establish and institutionalize reliability metrics, service health management practices, reliability governance models, and continuous service improvement frameworks.
Define and govern enterprise observability strategies, telemetry standards, monitoring frameworks, and best practices leveraging Splunk Observability Cloud, Splunk ITSI, and Splunk Enterprise.
Lead and guide the implementation of enterprise observability capabilities, including monitoring standards, OpenTelemetry adoption, Business Observability, service health monitoring, KPI frameworks, alert quality management, and operational analytics.
Guide teams in implementing effective Incident Management, Problem Management, Root Cause Analysis (RCA), and Blameless Postmortem practices.
Drive continuous improvement initiatives focused on reducing MTTD and MTTR, improving MTBF, and enhancing overall service reliability and operational efficiency.
Promote automation-first operations through toil reduction, intelligent automation, self-healing, and operational excellence initiatives.
Advise and guide teams on implementing AIOps capabilities, including event correlation, anomaly detection, intelligent alerting, operational intelligence, and predictive operations.
Collaborate with business, engineering, architecture, cloud, platform, security, and operations stakeholders to embed reliability into technology delivery and operational processes.
Drive adoption of Agile and Scrum practices through reliability reviews, retrospectives, and data-driven continuous improvement initiatives.
Measure and report SRE adoption, reliability KPIs, operational maturity, and business outcomes to leadership teams."
"Must Have Skills :
Good to Have Skills :

HCLTech is a global technology company, home to more than 226,600 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending September 2025 totaled $14.2 billion. To learn how we can supercharge progress for you, visit hcltech.com.