Omnissa

Site Reliability Engineer (SRE) IT Platform Engineering

Omnissa  •  Republic of India (Onsite)  •  2 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Site Reliability Engineer (SRE) – Observability, Tracing & Platform Operations

AboutOmnissa— Our Company, Mission & Vision

Omnissais an independent global leader in digital work platforms, originating from the former VMware End User Computing business and nowoperatingunder KKR ownership.Omnissaempowers organizations to deliver seamless, flexible, and secure digital work experiences to employees everywhere. Our platform supports over 26,000 customers worldwide, including 7 of the top 10 Fortune 500 enterprises

Our purpose is clear:

Guided by decades of innovation and strengthened by significant investment in AI, open APIs, and nextgeneration digital workspace technologies,Omnissais building the industry’s first autonomous workspace experience—a platform that simplifies operations for IT while unlocking higher productivity for employees.

Platform Engineering atOmnissa(IT Organization)

The Platform Engineeringteam withinOmnissa’sinternal IT organizationis responsible forarchitecting,operating, and continuously improving our enterprisegrade infrastructure platforms. Our mission is to deliver highly resilient, scalable, and secure systems that powerOmnissa’sinternal operations andcustomerfacingservices.

Our environment includes:

Core Platforms

  • On-premisecloud environmentsbuilt on:

  • VMware Cloud Foundation (vCF)

  • ApacheCloudStack

  • ProxmoxVirtualization Stack

  • Kubernetesbasedorchestrationfor containerized workloads

  • Opensource S3compatible object storage systems

Observability Infrastructure

We have developedextensive Observability Infrastructure thatmonitorsallOmnissainternal services that is managed automaticallyleveraging

  • Prometheus

  • Grafana

  • Loki

  • Ansible

AIDrivenAutomation & Incident Response

We have developed an internal AIpowered incident diagnosis and firstresponse platform,leveragingcuttingedgeopen technologies including:

  • Ollama

  • n8n

  • Various Model Context Protocol (MCP) servers

These systems help us reducemeantimeto detect(MTTD) andmeantime to-resolve (MTTR) through automated analysis, enrichment, and intelligent triage.

— Site Reliability Engineer (SRE)

We areseekinga highly skilled SRE with deepexpertisein Observability, particularly in:

  • Automation

  • Grafana

  • Loki

  • Prometheus

  • Development/Scripting

This role is critical tomaintainingthe reliability, performance, and operational integrity of our platforms. You will support both plannedand unplannedworkstreams, collaborating closely with engineering, incident management, and service owners.

The role includes participation in an on-call rotation, including nights and weekends, to ensure continuous coverage for missioncriticalsystems.

Key Responsibilities

Observability Engineering

  • Design, deploy, andmaintainLoki, Grafana, Prometheus, and integrated observability pipelines.

  • Contribute to new monitoring initiatives by developingnew monitoring checks for services.

  • Maintain and improve our automation workflows that manage the infrastructure.

  • Develop and refine AI workflowsfor incident analysis and auto-remediation.

  • Continuously enhance logging, metrics, and tracingcoverage across services.

Reliability, Resilience & Performance

  • Ensure high availability, capacity planning, and performance optimization across platforms.

  • Drive reliability improvements through automation, SLIs/SLOs, and rootcauseanalysis.

  • Partner with development and platform teams to embed reliability best practices.

Incident Management & On-Call

  • Participate inthe globalon-callrotation, including weekends.

  • Leverage our AIdrivenincident diagnosis tools(Ollama,n8n, MCP) to accelerate response.

  • Manage unplanned work such as production incidents, outages, highurgency escalationsandparticipatein post – mortem reviews

  • Coordinate postincidentreviews and continuous improvement initiatives.

Planned & Unplanned Work Management

  • Utilize the Atlassian toolset(Jira, Confluence,Opsgenie, etc.) for structured task, change, and incident management.

  • Manage planned maintenance, releases, and platform improvements.

  • Collaborate with crossfunctionalteams to prioritize backlog and operational tasks.

Platform Operations

  • Support and enhance internal clouds based on vCF, CloudStack, and Proxmox

  • Operate Kubernetes clusters and improve the reliability of containerized workloads.

  • Maintain S3compatible storage platforms used across the enterprise.

Required Skills & Experience

  • Familiarity with at least one scripting/programming language.

  • Strong hands-onexpertisewith:

  • Grafana, Loki, Tempo(or similar tracing systems), Prometheus

  • Experience with Configuration Management tools (e.g., Ansible/Saltstack)

  • Proficiencyin operating modern Linux-baseddistributed systems

  • Experience supporting largescale,highly availablearchitectures.

  • Familiarity with Kubernetes, CI/CD pipelines, and Infrastructure as Code.

  • Comfortable with on-callparticipation and incident leadership.

  • Experience with Atlassian tools (Jira, Confluence,Opsgenie).

  • Proficiencyin Linux& Windows

Nice to Have

  • Exposure to Ollama, N8N, or similar AI orchestration/automation tooling.

  • Experience with S3 storage internals oropensourceobject stores (e.g.,SeaweedFS, Ceph).

  • Understanding ofvirtualization stacks such as Proxmox, vSphere/vCF, orCloudStack

  • Background in SREdrivenculture, including SLIs/SLOs and error budgeting.

Omnissa

About Omnissa

Omnissa® is the digital work platform leader, trusted by thousands of organizations worldwide as the former VMware End-User Computing business.

We make digital work, work – for businesses and their people. No painful IT processes or productivity trade-offs. Instead, a seamlessly delivered digital employee experience that simplifies work. Our comprehensive digital work platform enables IT teams to provide secure, personalized experiences for every employee, on any device. Omnissa unifies, automates, and efficiently scales the digital workspace. By empowering employees to do their best work, anywhere, we help workforces everywhere unlock exponential business value.

All is made possible with the Omnissa platform, the first AI-driven digital work platform for smart, seamless, and secure work experiences from anywhere. It integrates multiple industry-leading solutions across unified endpoint management (UEM), virtual desktops and apps, digital employee experience (DEX), and security and compliance. By continuously adapting to users’ work styles, Omnissa optimizes user experience, security, IT operations and costs.

Industry
IT & Software
Company Size
1,001-5,000 employees
Headquarters
Mountain View, California
Year Founded
Unknown
Social Media