Meta

Production Engineering Manager

Meta  •  London, GB (Onsite)  •  1 day ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Meta is seeking a Production Engineering Manager to lead a team responsible for the reliability, scalability, and operational excellence of Meta's production infrastructure and services. In this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning and performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure Meta's services operate at global scale with high availability and efficiency.

Responsibilities
Manage a team of production engineers delivering on reliability, scalability, and operational efficiency across multiple interdependent production systems
* Drive roadmap creation for infrastructure reliability initiatives, capacity planning, and automation efforts, increasing team scope as AI-driven productivity improves throughput
* Lead adoption of AI-augmented engineering workflows across the team, sharing learnings and best practices with the broader production engineering organization
* Contribute hands-on to technical work including code, system design reviews, and incident response, using AI tooling to expand personal and team reach across disciplines
* Partner cross-functionally with software engineering, data science, and product teams to unblock dependencies and ensure smooth execution of infrastructure and reliability projects
* Proactively identify and resolve sources of operational toil — including on-call load, alerting gaps, and technical debt — and implement automation to increase team efficiency and scope
* Set clear goals and expectations for individual team members, provide timely and actionable feedback, and actively develop engineers' skills including proficiency with AI-augmented workflows
* Establish and monitor service-level objectives, reliability metrics, and engineering efficiency indicators to maintain high engineering craft and product quality
* Communicate production system health, incident learnings, and infrastructure strategy effectively to engineering leadership and cross-functional stakeholders
* Recruit, onboard, and retain production engineering talent, ensuring the team structure minimizes single points of failure and supports sustainable growth

Qualifications
4+ years of experience in production engineering, site reliability engineering, or systems software engineering
* 2+ years of experience managing production engineering or infrastructure engineering teams
* Experience driving reliability and scalability improvements for large-scale distributed systems, including incident management, capacity planning, and performance optimization
* Experience coding and debugging in at least one systems or scripting language (such as Python, C++, Go, or Bash) and contributing technically alongside a team
* Experience setting team goals, managing execution against roadmaps, and communicating infrastructure strategy to technical and non-technical stakeholders Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
* Track record of cross-functional collaboration with software engineering and data science teams to co-own reliability outcomes for consumer-facing or infrastructure services
* Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
* Experience leading adoption of AI-assisted tooling or automation frameworks within an engineering team to expand operational scope and reduce toil
* Experience managing on-call rotations, defining service-level objectives, and implementing observability and alerting improvements at scale
* Familiarity with container orchestration, service mesh architectures, or large-scale deployment pipelines in a production environment
* Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Meta

About Meta

Meta's mission is to build the future of human connection and the technology that makes it possible.

Our technologies help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology.

To help create a safe and respectful online space, we encourage constructive conversations on this page. Please note the following:

• Start with an open mind. Whether you agree or disagree, engage with empathy.

• Comments violating our Community Standards will be removed or hidden. Please treat everybody with respect.

• Keep it constructive. Use your interactions here to learn about and grow your understanding of others.

• Our moderators are here to uphold these guidelines for the benefit of everyone, every day.

• If you are seeking support for issues related to your Facebook account, please reference our Help Center (https://www.facebook.com/help) or Help Community (https://www.facebook.com/help/community).

For a full listing of our jobs, visit https://www.metacareers.com

Industry
IT & Software
Company Size
10,000+ employees
Headquarters
Menlo Park, CA
Year Founded
2004
Website
meta.com
Social Media