
Site Reliability Engineer (SRE) – Observability, Tracing & Platform Operations
AboutOmnissa— Our Company, Mission & Vision
Omnissais an independent global leader in digital work platforms, originating from the former VMware End User Computing business and nowoperatingunder KKR ownership.Omnissaempowers organizations to deliver seamless, flexible, and secure digital work experiences to employees everywhere. Our platform supports over 26,000 customers worldwide, including 7 of the top 10 Fortune 500 enterprises
Our purpose is clear:
Guided by decades of innovation and strengthened by significant investment in AI, open APIs, and nextgeneration digital workspace technologies,Omnissais building the industry’s first autonomous workspace experience—a platform that simplifies operations for IT while unlocking higher productivity for employees.
Platform Engineering atOmnissa(IT Organization)
The Platform Engineeringteam withinOmnissa’sinternal IT organizationis responsible forarchitecting,operating, and continuously improving our enterprisegrade infrastructure platforms. Our mission is to deliver highly resilient, scalable, and secure systems that powerOmnissa’sinternal operations andcustomerfacingservices.
Our environment includes:
Core Platforms
On-premisecloud environmentsbuilt on:
VMware Cloud Foundation (vCF)
ApacheCloudStack
ProxmoxVirtualization Stack
Kubernetesbasedorchestrationfor containerized workloads
Opensource S3compatible object storage systems
Observability Infrastructure
We have developedextensive Observability Infrastructure thatmonitorsallOmnissainternal services that is managed automaticallyleveraging
Prometheus
Grafana
Loki
Ansible
AIDrivenAutomation & Incident Response
We have developed an internal AIpowered incident diagnosis and firstresponse platform,leveragingcuttingedgeopen technologies including:
Ollama
n8n
Various Model Context Protocol (MCP) servers
These systems help us reducemeantimeto detect(MTTD) andmeantime to-resolve (MTTR) through automated analysis, enrichment, and intelligent triage.
— Site Reliability Engineer (SRE)
We areseekinga highly skilled SRE with deepexpertisein Observability, particularly in:
Automation
Grafana
Loki
Prometheus
Development/Scripting
This role is critical tomaintainingthe reliability, performance, and operational integrity of our platforms. You will support both plannedand unplannedworkstreams, collaborating closely with engineering, incident management, and service owners.
The role includes participation in an on-call rotation, including nights and weekends, to ensure continuous coverage for missioncriticalsystems.
Key Responsibilities
Observability Engineering
Design, deploy, andmaintainLoki, Grafana, Prometheus, and integrated observability pipelines.
Contribute to new monitoring initiatives by developingnew monitoring checks for services.
Maintain and improve our automation workflows that manage the infrastructure.
Develop and refine AI workflowsfor incident analysis and auto-remediation.
Continuously enhance logging, metrics, and tracingcoverage across services.
Reliability, Resilience & Performance
Ensure high availability, capacity planning, and performance optimization across platforms.
Drive reliability improvements through automation, SLIs/SLOs, and rootcauseanalysis.
Partner with development and platform teams to embed reliability best practices.
Incident Management & On-Call
Participate inthe globalon-callrotation, including weekends.
Leverage our AIdrivenincident diagnosis tools(Ollama,n8n, MCP) to accelerate response.
Manage unplanned work such as production incidents, outages, highurgency escalationsandparticipatein post – mortem reviews
Coordinate postincidentreviews and continuous improvement initiatives.
Planned & Unplanned Work Management
Utilize the Atlassian toolset(Jira, Confluence,Opsgenie, etc.) for structured task, change, and incident management.
Manage planned maintenance, releases, and platform improvements.
Collaborate with crossfunctionalteams to prioritize backlog and operational tasks.
Platform Operations
Support and enhance internal clouds based on vCF, CloudStack, and Proxmox
Operate Kubernetes clusters and improve the reliability of containerized workloads.
Maintain S3compatible storage platforms used across the enterprise.
Required Skills & Experience
Familiarity with at least one scripting/programming language.
Strong hands-onexpertisewith:
Grafana, Loki, Tempo(or similar tracing systems), Prometheus
Experience with Configuration Management tools (e.g., Ansible/Saltstack)
Proficiencyin operating modern Linux-baseddistributed systems
Experience supporting largescale,highly availablearchitectures.
Familiarity with Kubernetes, CI/CD pipelines, and Infrastructure as Code.
Comfortable with on-callparticipation and incident leadership.
Experience with Atlassian tools (Jira, Confluence,Opsgenie).
Proficiencyin Linux& Windows
Nice to Have
Exposure to Ollama, N8N, or similar AI orchestration/automation tooling.
Experience with S3 storage internals oropensourceobject stores (e.g.,SeaweedFS, Ceph).
Understanding ofvirtualization stacks such as Proxmox, vSphere/vCF, orCloudStack
Background in SREdrivenculture, including SLIs/SLOs and error budgeting.

Omnissa® is the digital work platform leader, trusted by thousands of organizations worldwide as the former VMware End-User Computing business.
We make digital work, work – for businesses and their people. No painful IT processes or productivity trade-offs. Instead, a seamlessly delivered digital employee experience that simplifies work. Our comprehensive digital work platform enables IT teams to provide secure, personalized experiences for every employee, on any device. Omnissa unifies, automates, and efficiently scales the digital workspace. By empowering employees to do their best work, anywhere, we help workforces everywhere unlock exponential business value.
All is made possible with the Omnissa platform, the first AI-driven digital work platform for smart, seamless, and secure work experiences from anywhere. It integrates multiple industry-leading solutions across unified endpoint management (UEM), virtual desktops and apps, digital employee experience (DEX), and security and compliance. By continuously adapting to users’ work styles, Omnissa optimizes user experience, security, IT operations and costs.