LTIMindtree

Specialist - Software Engineering

LTIMindtree  •  Mexico (Onsite)  •  1 month ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Job Title

Site Reliability Engineer (SRE)

We are seeking an experienced Site Reliability Engineer (SRE) with strong DevOps and automation expertise to ensure the reliability, scalability, and performance of distributed systems. This role focuses on CI/CD automation, monitoring, observability, and system troubleshooting across cloud-native and Kubernetes-based environments.

You will play a critical role in building and maintaining monitoring platforms, automating operational processes, and improving system reliability across multiple application domains.

Key Responsibilities

  • Apply Site Reliability Engineering (SRE) and DevOps best practices to improve system availability, performance, and scalability.
  • Design, build, and maintain CI/CD pipelines with a strong focus on automation.
  • Implement and manage metrics collection, monitoring, and alerting across platforms.
  • Perform system troubleshooting and problem-solving across infrastructure and application layers.
  • Create, operate, and maintain Prometheus and Grafana clusters for monitoring Kubernetes environments.
  • Implement and support observability standards, including OpenTelemetry
  • Develop and maintain automation tools and scripts using Python, Groovy, and Shell
  • Collaborate with engineering and platform teams to improve reliability, deployment processes, and operational efficiency.

Required Skills & Qualifications

  • Hands-on experience in Site Reliability Engineering (SRE) and DevOps roles.
  • Strong expertise in CI/CD pipelines, automation, and deployment strategies.
  • Experience with metrics collection, monitoring, and alerting systems
  • Proven ability in system troubleshooting and root cause analysis across platforms and applications.
  • Hands-on experience managing Prometheus and Grafana for Kubernetes cluster monitoring.
  • Strong automation and scripting skills using:
    • Python
    • Shell scripting
    • Groovy
  • Experience working with OpenTelemetry for distributed tracing and observability.

Key Skills

  • SRE experience managing Google Cloud services and accounts
  • Strong Prometheus and Grafana querying and dashboarding skills.
  • Observability and monitoring best practices.
  • Automation-first mindset with strong scripting capabilities.
  • Kubernetes monitoring and cloud-native operations experience.
LTIMindtree

About LTIMindtree

LTIMindtree is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies. As a digital transformation partner to more than 700 clients, LTIMindtree brings extensive domain and technology expertise to help drive superior competitive differentiation, customer experiences, and business outcomes in a converging world. Powered by 86,000+ talented and entrepreneurial professionals across more than 30 countries, LTIMindtree — a Larsen & Toubro Group company — combines the industry-acclaimed strengths of erstwhile Larsen and Toubro Infotech and Mindtree in solving the most complex business challenges and delivering transformation at scale.

For more info, please visit www.ltimindtree.com.

Industry
IT & Software
Company Size
10,000+ employees
Headquarters
Mumbai, IN
Year Founded
Unknown
Social Media