The TJX Companies, Inc.

SDM SRE Senior Engineer

The TJX Companies, Inc.  •  Republic of India (Onsite)  •  3 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

TJX Companies

At TJX Companies, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our four global Home Offices, Distribution Centers or Retail Stores—TJ Maxx, Marshalls, Homegoods, Homesense, Sierra, Winners, and TK Maxx, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer.

Location:India IT Office, Hyderabad / India

Whatyou’lldiscover

  • Inclusive culture and career growth opportunities

  • Global IT Organization which collaborates across U.S., Canada, Europe,Indiaand Australia

  • Challenging, collaborative, and team-based environment

Whatyou’lldo

The Infrastructure and Operations organization embodies the hub of lifecycle engineering at TJX, delivering,maintaining, andoptimizingour technology portfolio at cloud scale. We are a service-oriented team aimed at providing extraordinary experiences to thousands of TJX associates, business partners, and application delivery teams across the portfolio.

We are seeking a Site Reliability Engineer – AI Operations to implement Site Reliability Engineering practices with a strong focus on automation, observability, AIOps, Gen AI, and self-healing operations. This role will help improve production reliability, reduce manual toil, accelerate incident response, and enable intelligent operational automation across enterprise applications and platforms.

The ideal candidate combines strong production operations experience with engineering, automation, observability, and AI-enabled operations capabilities. This role will work closely with application, infrastructure, cloud, security, operations, and leadership teams to strengthen reliability, improve operational efficiency, and modernize how production services are supported.

Key Responsibilities

Site Reliability Engineering&Operations

  • Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes.

  • Support production reliability by improving availability, performance, scalability, and operational readiness of applications and platforms.

  • Participate in incident response, troubleshooting, root cause analysis, postmortems, and corrective/preventive action planning.

  • Drive shift-left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines.

  • Support reliable application and platform operations in Azure Cloud.

Automation&Self-Healing

  • Build andmaintainautomation scripts, runbooks, self-healing workflows, and operational tools to reduce manual effort and improve MTTR.

  • Identifyrepetitive operational activities and automate them whereappropriate

  • Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows.

  • Create andmaintainreusableautomationplaybooks, operational procedures, and knowledge articles.

  • Ensure automation activities follow change management, governance, security, and compliance processes.

Observability, Monitoring&Reliability Metrics

  • Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring.

  • Work with monitoring and APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.

  • Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards.

  • Analyze service health trends, alert patterns, and operational data toidentifyimprovement opportunities.

  • Improve alert quality, event correlation, and operational visibility across enterprise applications and platforms.

AIOps, Gen AI&Intelligent Operations

  • Support AI-driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations.

  • Contribute to the design and implementation of Gen AI and AIOpsusecases that improve incident response and operational productivity.

  • Explore opportunities to use AI-enabled operations to reduce toil, improve service reliability, and enhance support experiences.

  • Ensure AI-enabled operations follow enterprise security, compliance, data privacy, and responsible AI guidelines.

  • Collaborate with engineering, operations, and platform teams to evaluate emerging AI and automation capabilities.

Collaboration&Continuous Improvement

  • Work closely with application, infrastructure, cloud, operations, security, and leadership teams.

  • Support cross-functional initiatives focused on reliability, automation, service health, and operational excellence.

  • Communicate technical issues, risks, and improvementopportunities clearlyto stakeholders.

  • Contribute to continuous improvement byidentifyinggaps in process, tooling, monitoring, and documentation.

  • Support knowledge sharing and operational readiness across global teams.

Whatyou’llneed

Weseekcreative, customer-focused individuals with strong SRE, DevOps, production operations, automation, observability, and AI-enabled operations experience. This role requires a continuous improvement mindset and the ability to balance innovation with reliability, stability, security, and compliance.

Education

  • Bachelor’s degree in Information Technology, Computer Science, Engineering, or equivalent practical experience.

Experience

  • 6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles.

  • Hands-on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation.

  • Experience supporting production applications and platforms in large-scale enterprise environments.

  • Experience working acrossapplication, infrastructure, cloud, operations, security, and leadership teams.

  • Experience working within ITIL-based operational environments.

Required Technical Skills

  • Strong experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools.

  • Experience with SLIs, SLOs, SLAs, operational KPIs, service health dashboards, and reliability metrics.

  • Strong coding/scripting experience in one or more languages such as Python, Java, PowerShell, Go, or Shell scripting.

  • Hands-on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate,Rundeck, or similar tools.

  • Knowledge of AIOps, Gen AI, or AI-enabled operations use cases such as anomaly detection, log analysis, ticket classification, event correlation, and incident summarization.

  • Experience integrating enterprise tools using APIs, webhooks, scripts, and automation workflows.

  • Hands-on experience supporting or implementing solutions in Azure Cloud.

  • Understanding ofdistributed systems, APIs, microservices, cloud-native applications, and reliability engineering principles.

  • Strong troubleshooting, analytical, communication, and stakeholder management skills.

Preferred Qualifications

  • Experience with Agentic AI, AI agents, or Gen AI-based automation for IT operations.

  • Experience with Azure OpenAI, Microsoft Copilot Studio,LangChain, Semantic Kernel, vector databases, or RAG-based solutions.

  • Experience with ITSM and incident response tools such as ServiceNow, Jira Service Management, PagerDuty,xMatters,Opsgenie, or similar platforms.

  • Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code,GitOps, orDevSecOps

  • Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time-series analysis.

  • Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls.

  • Experience in the retail domain.

  • Relevant certifications in Azure, DevOps, SRE, AI, or ITIL.

Key Competencies

  • Strong analytical and problem-solving mindset.

  • Automation-first and continuous improvement mindset.

  • Strong communicationand stakeholder management skills.

  • Ability to influence without direct authority.

  • Ability to work across product, engineering, operations, cloud, security, and leadership teams.

  • Curiosity and willingness to explore AIOps, Gen AI, and Agentic AI capabilities.

  • Ability to balance innovation with reliability, stability, security, and compliance.

  • Customer-focused approach with strong ownership and accountability.

Working Expectations

This role is based in India and requires collaboration with global teams across U.S., Canada, Europe, India, and Australia. The candidate should be comfortable supporting flexible working hours as needed to engage with global stakeholders,participatein critical meetings, support escalations, and ensure seamless delivery across regions.

Join us

Join us and Discover Different at TJX.

In addition to our open door policy and supportive work environment, we also strive to provide a competitive salary and benefits package. TJX considers all applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, gender identity and expression, marital or military status, or based on any individual's status in any group or class protected by applicable federal, state, or local law. TJX also provides reasonable accommodations to qualified individuals with disabilities in accordance with the Americans with Disabilities Act and applicable state and local law.

Address:

Salarpuria Sattva Knowledge City, Inorbit Road

Location:

APAC Home Office Hyderabad IN

The TJX Companies, Inc.

About The TJX Companies, Inc.

TJX is the leading off-price apparel and home fashions retailer in the U.S. and worldwide, with four global home offices, seven brands, nearly 4,700 stores in nine countries, and five distinctive branded e-commerce sites. As Associates, we make a difference with our contributions—collaborating in delighting shoppers with hidden treasures.

Industry
Retail & Ecommerce
Company Size
10,000+ employees
Headquarters
Framingham, MA
Year Founded
Unknown
Website
tjx.com
Social Media