Omnicell

Sr. Manager, Site Reliability

Omnicell  •  Austin, TX (Hybrid)  •  7 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

About this opportunity

Omnicell is building a Global Cloud Operations organization from the ground up as our business shifts from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform that hospitals depend on 24/7. The Site Reliability Engineering function is the reliability engine of that organization, and this role is the first senior SRE hire — the person who will design the practice, set the standards, and then run the plays themselves until the team is large enough to delegate.

This is not a role where reliability practices already exist and you tune them. It is a role where you define what good looks like for Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud Operations and an Engineer III SRE you will coach and grow.

The environment is hybrid. Some of our products are still hardware in hospitals communicating with cloud services; others are fully SaaS. Some customers access us over private circuits, others over the public internet. We operate in a regulated environment — HIPAA, SOC 2, and in some engagements FedRAMP — which means reliability, security, and auditability are not separable concerns. The person we hire will be comfortable with that complexity and will help the organization design for it rather than around it.

This role also anchors Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner of how that gets introduced, prioritized against foundational reliability work, and validated in a regulated environment.

What you will own

Reliability practice (the coach half)

  • Define and publish SLOs and SLIs for the top 5–10 Tier-1 customer-facing services, in partnership with Product and Engineering. Establish error budget policy and the enforcement mechanism when budgets burn.

  • Design the incident command structure: severity rubric, declaration criteria, war-room protocol, stakeholder communication cadence, and the postmortem template. Train the first cohort of incident commanders across Engineering and Support.

  • Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet.

  • Partner with the VP to migrate the interim incident response RACI — currently held by matrixed individuals across IT, Engineering, Support, and Enterprise Security — into a durable SRE-owned model.

  • Establish the on-call rotation model, including fair distribution, compensation approach, paging discipline, and the handoff protocol with our existing managed services partners (IBM, HCL) who provide L1/L2 coverage.

  • Develop and track operational KPIs — MTTR, SLO attainment, change-failure rate, recurrence, cost per workload — and present reliability metrics and improvement roadmaps to senior leadership in the monthly Cloud Ops executive review.

Hands-on engineering (the player half)

  • Instrument Tier-1 services yourself. Write the dashboards. Write the alerts. Write the runbooks. Do not wait for the team to grow before the work starts.

  • Take the pager. Commander Sev-1 and Sev-2 incidents until a broader on-call rotation is staffed. Lead blameless postmortems and drive follow-up work to resolution.

  • Contribute code and infrastructure-as-code (Terraform preferred; Chef/Puppet acceptable) to the platform. Oversee the design and evolution of CI/CD pipelines — our current stack includes CodeFresh, TeamCity, GitHub Actions, and Octopus Deploy, and we are consolidating over time.

  • Administer and scale our Kubernetes platform, including secure and compliant cluster configurations. Working knowledge of Docker, Helm, and Service Mesh (Istio or Linkerd) expected.

  • Run chaos and failover exercises (Chaos Monkey, LitmusChaos, or equivalent). Validate that what we think is resilient actually is.

AI-driven operations

  • Architect Omnicell's AIOps direction: evaluate and introduce ML-based anomaly detection, predictive alerting, automated root cause analysis, and LLM-assisted runbook or triage pipelines.

  • Make informed build-versus-buy calls across the AIOps landscape. Integrate AI-assisted tooling into the observability and incident response stack where it adds measurable value; resist the hype where it does not.

  • Ensure AI-assisted operations meet the auditability and explainability bar required in a HIPAA and SOC 2 environment.

Coaching and team-building

  • Coach one Engineer III SRE who joins shortly after you do. Pair on incidents. Review their design proposals. Help them grow toward senior. This is a formal, named relationship, not a side duty.

  • Design the next 2–4 SRE hires. Write the requisitions, run the interview loops, make the calls. Your operating assumption is that the team grows under your direction over the next 12–18 months.

  • Represent SRE in architecture reviews, product launch readiness reviews, and the monthly executive Cloud Ops metric review. Be the person in the room who knows what reliability costs and what it is worth.

  • Partner with Enterprise Security, Compliance, and Architecture to ensure platform services meet regulatory and security requirements in a healthcare environment.

What success looks like in the first six months

Concrete outcomes this role will be evaluated against in the first half-year. These are drawn from the Cloud Ops 90-day plan and its extension into the following quarter.

  • Month 1: SLOs drafted for the top 5 Tier-1 services with Product sign-off. Severity rubric published. First live tabletop Sev-1 run against the interim RACI.

  • Month 2: Observability platform selection finalized. Instrumentation standard published. Engineer III SRE hired and onboarded.

  • Month 3: On-call rotation live. First real Sev-1 commanded under the new structure with a blameless postmortem completed and follow-ups tracked.

  • Month 4–6: Error budget policy in effect for the first 3 services. First incident review at executive level. Interview loop running for the next SRE hires. Initial AIOps evaluation and pilot scope defined.

Required knowledge and skills

  • Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment — with demonstrated experience building a practice from zero or near-zero: you have set SLOs, defined incident command, and introduced error budget thinking to an organization that did not have it.

  • Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.

  • Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, GitHub Actions, Jenkins, TeamCity, or equivalent).

  • Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.

  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.

  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations. Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd).

  • Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like.

  • Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be applied in a regulated environment.

  • Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice and documented follow-up discipline.

  • Ability to coach and mentor, with direct evidence of growing junior and mid-level engineers. You are not a manager in this role, but you are a formal coach.

  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.

  • Excellent communication and stakeholder management skills; ability to translate complex technical concepts for non-technical audiences.

Basic requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience

  • 8+ years of experience in software or platform engineering, with at least 4 of those in an SRE, DevOps, or platform reliability role.

  • Proven Experience advising and influencing senior technical or operations leaders using data driven recommendations.

  • At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.

Preferred knowledge and skills

  • Masters Degree

  • Prior experience in healthcare, clinical workflows, or another regulated vertical.

  • Experience transitioning from MSP-heavy operations to internal-first, or integrating managed service providers (IBM, HCL, or similar) into an SRE operating model.

  • Exposure to hybrid hardware-plus-cloud products, where device reliability and cloud reliability are jointly owned.

  • Experience building or integrating AIOps platforms for automated incident triage and remediation.

  • Familiarity with large language model APIs or agentic AI frameworks applied to on-call automation or runbook generation.

  • Experience deploying and managing stateful distributed services in Kubernetes.

  • Hands-on experience with security scanning and intrusion detection systems in regulated environments (HIPAA, SOC 2, or equivalent).

  • Experience with messaging systems such as Kafka or RabbitMQ.

  • Familiarity with chaos engineering principles and tooling (Chaos Monkey, LitmusChaos, or similar).

  • Working knowledge of Databricks, Team Foundation Server, Octopus Deploy, or similar tools in Omnicell's current stack.

  • Experience with FinOps practices and cloud cost optimization strategies.

Who you will work with

You will report to the VP, Global Cloud Operations, who is joining Omnicell in parallel with this role. You will be the VP's first senior technical hire and their primary partner in standing up the SRE function.

You will work daily with Platform Engineering, the future NOC function (initially staffed through our existing IBM and HCL partnerships), Product Engineering leads, Enterprise Security, and the Customer Support organization. You will also partner with Enterprise Security's SOC during security-relevant incidents per the established CloudOps–Security partnership model.

You will coach one Engineer III SRE directly and help design the interview loop for subsequent SRE hires.

Why this role, why now

Most senior SRE roles have you tuning a practice that already exists. This one has you building it. At Omnicell you will make foundational calls — what SLOs Tier-1 services carry, what the incident command model looks like, what our observability stack is, how we integrate AI-driven operations, how we partner with our managed service providers — that will shape how the company operates for the next several years.

You will also be building in a domain where reliability genuinely matters. The products Omnicell builds dispense medications in hospitals. When our cloud services degrade, pharmacy workflows are affected and patient care can be too. That raises the stakes of the work, and it raises the quality of the conversations you will have with Product and Engineering about reliability trade-offs.

On the player-coach dynamic

This role is explicitly a player-coach, not a manager. You will not have direct reports in the first six months. You will have a formal coaching relationship with an Engineer III SRE who joins shortly after you do.

The split in practice is roughly 60 percent hands-on engineering and incident response, 25 percent practice design and coaching, 15 percent cross-functional partnership work. As the team grows, the ratio will shift toward coaching and practice leadership, but hands-on work never goes to zero in this role. If you want to stop being hands-on, this is not the right fit.

Growth and career path

The natural next step for someone successful in this role is Manager or Director of SRE as the team grows past the first few hires, or a Principal / Distinguished Engineer path if you want to stay individual contributor and technical. Both paths are supported and neither is forced. Internally this role is mapped to the Manager, Site Reliability Engineering compensation band, but the expected day-to-day operating mode is senior individual contributor with formal coaching responsibilities.

Work conditions

  • Corporate office or lab environment; remote or hybrid arrangement supported.

  • Ability to travel up to 10% of the time.

  • On-call participation expected as part of the SRE rotation.


Since 1992, Omnicell has been committed to transforming pharmacy care through outcomes-centric innovation designed to optimize clinical and business outcomes across all settings of care. We strive to be the healthcare provider’s most trusted partner by our guiding promise of “Outcomes. Defined and Delivered.”  

Our comprehensive portfolio of robotics, smart devices, intelligent software, and expert services is helping healthcare facilities worldwide to improve business and clinical outcomes as they move closer to the industry vision of the Autonomous Pharmacy. 

Our guiding principles inform everything we do: 
  • As Passionate Transformers, we find a better way to innovate relentlessly. 
  • Being Mission Driven, we consistently deliver on our promises. 
  • Our Entrepreneurial spirit makes the most of EVERY opportunity for innovation. 
  • Understanding that Relationships Matter creates synergies that yield the greatest benefits for all.
  • Intellectually Curious, eager to think deeper to learn and improve.
  • In Doing the Right Thing, we lead by example in ALL we do. 

We are deeply committed to Environmental, Social, and Governance (ESG) initiatives. Our ESG efforts focus on creating an inclusive culture and a healthier world. This includes our Employee Impact Groups, which foster inclusion and belonging, as well as our learning and well-being programs that support personal and professional growth. We also prioritize sustainability in our operations, aiming to reduce our environmental footprint and promote responsible business practices. Join us in transforming the pharmacy care delivery model, making patient care safer and smarter for all.
Omnicell

About Omnicell

Omnicell is transforming pharmacy and nursing care through outcomes-centric solutions designed to optimize clinical and business outcomes across all settings of care. Our comprehensive portfolio of robotics and smart devices, intelligent software workflows, and data and analytics, all optimized by expert services are helping healthcare facilities worldwide to reduce costs, improve labor efficiency, establish new revenue streams, enhance supply chain control, support compliance, and move closer to the industry vision of the Autonomous Pharmacy. To learn more, visit omnicell.com.

Industry
Healthcare & Social Services
Company Size
1,001-5,000 employees
Headquarters
Fort Worth, Texas
Year Founded
1992
Social Media