Job Description
Proficient Mission-Critical Application Monitoring & Support Engineer
Responsibilities
Support and maintain mission-critical applications in a production environment. The role focuses on proactively monitoring application health, managing alerts and incidents, performing advanced troubleshooting, analyzing logs, conducting root cause analysis (RCA), and ensuring compliance with strict SLAs. The candidate will work closely with application teams, infrastructure teams, vendors, and business stakeholders to maintain service availability, performance, and reliability.
Responsibilities include monitoring dashboards and business-critical transactions, investigating and resolving incidents, performing log and trend analysis, tuning alerts to reduce noise and improve monitoring effectiveness, documenting activities in ServiceNow, leading or contributing to RCAs, maintaining operational procedures and SOPs, and driving continuous improvement initiatives. The ideal candidate should demonstrate strong analytical and troubleshooting skills, a high sense of urgency, operational ownership, and the ability to communicate effectively during critical situations.
Preferred Qualifications
- 3 to 5+ years of experience in Application Support, Production Support, Monitoring Operations, SRE (Site Reliability Engineering) or IT Operations.
- Experience supporting mission-critical applications in production environments.
- Hands-on experience with monitoring and observability platforms such as New Relic, Dynatrace, Datadog, Splunk, Grafana, AppDynamics, or similar tools.
- Strong experience in application and infrastructure log analysis and troubleshooting.
- Experience managing incidents in SLA-driven environments.
- Knowledge of SQL and database troubleshooting.
- Understanding of APIs, integrations, and distributed application architectures.
- Experience with cloud environments such as Azure, AWS, or Google Cloud Platform.
- Familiarity with ServiceNow or similar ITSM platforms.
- Knowledge of ITIL Incident Management, Problem Management, and Change Management processes.
- Experience participating in or leading Root Cause Analysis (RCA) activities.
- Experience supporting high-availability, 24x7 operational environments.
- Excellent verbal and written communication skills in English.
Required Technical Skills
- Monitoring and observability tools (New Relic, Dynatrace, Datadog, Splunk, Grafana, AppDynamics, or equivalent).
- Log analysis and troubleshooting.
- SQL and database querying.
- Incident management and escalation processes.
- Application performance monitoring (APM).
- Basic cloud troubleshooting (Azure, AWS, GCP).
- ServiceNow or ITSM platforms.
- Application infrastructure fundamentals (web servers, APIs, integrations, middleware).
- Dashboard creation and operational reporting.
- Alert tuning and threshold optimization.
Preferred Certifications
- ITIL Foundation Certification.
- New Relic Certified Professional.
- Splunk Core Certified User/Power User.
- Dynatrace Associate Certification.
- Microsoft Azure Fundamentals (AZ-900).
- AWS Cloud Practitioner.
- Google Associate Cloud Engineer.
- Site Reliability Engineering (SRE) training or certification.
- ServiceNow Fundamentals Certification.
Key Competencies
- Strong analytical and problem-solving skills.
- Incident ownership and accountability.
- Sense of urgency and ability to work under pressure.
- Root cause analysis and critical thinking.
- Effective stakeholder communication.
- Customer-focused mindset.
- Operational leadership during critical incidents.
- Risk identification and proactive issue prevention.
- Continuous improvement mindset.
- Collaboration across technical and business teams.
- Attention to detail and execution discipline.
Success Profile
A successful candidate is someone who proactively identifies risks before they impact the business, effectively manages incidents through resolution, performs detailed log and trend analysis, communicates clearly with stakeholders, drives RCA and corrective actions, and continuously improves monitoring and operational processes. They take full ownership of production issues, demonstrate strong technical judgment, and contribute to maintaining the stability, availability, and performance of mission-critical applications.