
How do bold AI research ideas become diligent training, evaluation, and production systems? NVIDIA’s Deep Learning Software team is looking for a Senior Technical Program Manager to lead programs across model pre-training, production RL runs, evaluation, and agentic AI infrastructure. We build the software foundations that help research and engineering teams train, evaluate, and deliver sophisticated AI models. We partner across research, platform engineering, distributed computing, evaluation, and open-source development to turn sophisticated technical goals into clear software plans. Come help us improve how sophisticated AI systems are built and validated!
What you'll be doing:
Lead multi-functional programs across training frameworks, evaluation environments, agent and model runtimes, datasets, verifiers, and distributed training infrastructure.
Partner with AI researchers, engineering leaders, product, infrastructure, and QA teams to define roadmaps, achievements, release plans, and measurable success criteria.
Coordinate large-scale RL training and evaluation experiments, including handling GPU resources, dependency tracking, run scheduling, results reporting, release readiness, technical decisions, integration plans, and program updates.
What we need to see:
Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent experience.
10+ years of technical program management, engineering program management, or related experience delivering sophisticated software platforms.
Experience leading global, matrixed programs across research, software engineering, infrastructure, QA, release teams, and partner groups.
Strong understanding of the AI model lifecycle, including training, post-training, evaluation, experimentation, production readiness, reinforcement learning concepts, and GPU-accelerated distributed systems.
Experience running software releases across repositories, dependencies, test configurations, quality gates, collaborator approvals, open-source workflows, CI/CD systems, and tools such as GitHub, Git, Jira, Linear, Aha!, or Confluence.
Ways to stand out from the crowd:
Experience supporting reinforcement learning, post-training, agentic AI, or large-scale model evaluation programs.
Familiarity with PPO, GRPO, asynchronous RL, distributed inference, rollout generation, policy optimization, evaluation harnesses, verifiers, benchmark development, or reproducible experimentation.
Knowledge of GPU infrastructure, distributed training, Kubernetes, workload schedulers, cluster capacity management, performance analysis, open-source contributor workflows, release readiness, or operational metrics.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 258,750 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.
You will also be eligible for equity and benefits
Applications for this job will be accepted at least until August 1, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.#deeplearning

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.