Job Description
The Meta Technical Program Management (TPM) community is pioneering technologies to bring people and businesses closer together at global scale. TPMs work at the cross-section of technical execution and business strategy and partner closely with Engineering, Data Science, and Product teams. Being a TPM at Meta means driving impact by delivering measurable results across a wide range of areas — defining and guiding high-level goals and roadmaps, monitoring and communicating progress, and shaping the decisions that determine where and how we invest. It also means having a strong technical background, understanding system architecture, and the experience to collaborate effectively across functions and organizations to deliver impact.
Central Products builds the shared platforms, infrastructure, and cross-cutting capabilities that power Meta's family of apps. Within Central Products, the Capacity team ensures that compute and infrastructure resources — increasingly, the AI and GPU capacity behind Meta's most important product bets — are forecast, allocated, and optimized so that the highest-priority products can scale without being constrained by infrastructure. This is a role for someone who is equally comfortable building rigorous quantitative models and translating them into clear, executive-ready narratives that drive decisions.
One of the fastest-growing areas we support is Meta Business AI — the AI agents that help businesses connect with their customers. As this product scales rapidly, its demand for large-scale AI inference capacity grows with it. As the Technical Program Manager for Capacity, you will own the capacity strategy for this space end-to-end and, above all, build the analytical and communication backbone that enables Engineering, Data Science, Product, and leadership to make the right capacity, cost, and efficiency tradeoffs at speed.
Responsibilities
Own the capacity demand model, forecasting, and supply plan for a large-scale AI product area, partnering with Engineering, Data Science, Infrastructure, and Product to align roadmaps, instrumentation, and tradeoffs
* Lead leadership communications for capacity: build the narratives, reviews, scorecards, and decision frameworks that give senior leadership a clear, data-grounded view of capacity health, risks, tradeoffs, and resource decisions
* Drive a deep, quantitative view of AI inference capacity — modeling compute and GPU demand, and driving infrastructure efficiency and cost optimization
* Build the analytical foundation — models, scorecards, and dashboards — that makes performance, efficiency, quality, cost, and capacity tradeoffs explicit and enables faster, better-grounded cross-functional decisions
* Partner with Engineering to protect reliability and capacity headroom so that product growth is never limited by infrastructure, including planning for demand spikes and graceful degradation
* Define and drive end-to-end program plans across multiple organizations; own risk, escalation, and change management
* Establish the operating cadence — capacity reviews, planning cycles, and ramp gates — and act as the connective tissue that keeps a fast-moving, multi-team effort aligned on the right tradeoffs
* Influence program and product direction across Engineering, Product, and Infrastructure through data-driven analysis and thought leadership, simplifying complexity and building alignment without direct authority
Qualifications
B.S. in Computer Science or a related technical discipline, or equivalent experience
* 15+ years of software engineering, systems engineering, or technical product/program management experience, including owning large-scale, cross-organization technical programs
* Experience delivering technical programs or products from inception to delivery jointly with Engineering, Data Science, and Product partners
* Experience with capacity planning, infrastructure, or ML-serving / performance / efficiency programs at scale
* Experience building quantitative models (demand, cost, capacity, or quality) and using them to inform senior-leadership decisions
* Experience communicating and translating technical topics for executive and non-technical audiences, including leadership-facing narratives, reviews, and written communication
* Experience influencing strategy and outcomes across multiple organizations and partners in different time zones, without direct authority
* Experience gathering requirements, defining scope, and performing risk and change management on large programs Experience in Business AI, ads, or AI-agent / agentic-commerce domains
* Experience owning executive communications and decision frameworks for large, cross-functional technical programs
* Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
* Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
* Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
* Experience with GenAI / LLM inference at scale — inference efficiency, model right-sizing, batching, quantization, or multi-model routing
* Experience with capacity and cost economics, including infrastructure cost modeling, unit economics, and ROI analysis