Recruiting from Scratch

Team Lead, GPU Sandboxes

Recruiting from Scratch  •  $250k - $300k/yr  •  San Francisco, CA (Onsite)  •  3 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Who is Recruiting from Scratch:
Recruiting from Scratch is a specialized talent firm dedicated to helping companies build exceptional teams. We partner closely with our clients to deeply understand their needs, then connect them with top-tier candidates who are not only highly skilled but also the right fit for the company’s culture and vision. Our mission is simple: place the best people in the right roles to drive long-term success for both clients and candidates.

Team Lead, GPU Sandboxes

Location: San Francisco, CA
Company Stage of Funding: Series A
Office Type: In Person
Salary: $250,000-$300,000 + Equity

We're representing a rapidly growing infrastructure company building the compute layer for AI agents. Its platform provides secure, isolated, instantly available sandboxes where AI agents can execute code, interact with computers, and run training workloads.

Following its Series A, the company has grown 10x and is now expanding its GPU infrastructure to support increasingly demanding agentic AI, reinforcement learning, training, and inference workloads. The team is building serverless GPU environments that combine full GPU passthrough with fast startup, checkpointing, snapshotting, and workload forking.

This is a highly technical player-coach opportunity for an experienced infrastructure engineer to lead a small team while remaining deeply hands-on with the systems powering production GPU workloads.

What You Will Do

  • Lead the architecture and delivery of the company's serverless GPU sandbox runtime.
  • Drive the transition from MIG-based GPU environments to full VFIO GPU passthrough.
  • Build and operate virtualization infrastructure using KVM/QEMU, VFIO/PCIe passthrough, IOMMU isolation, and guest image management.
  • Take existing GPU checkpoint and restore capabilities from a working implementation into reliable production infrastructure.
  • Design systems that enable GPU-attached workloads to pause, snapshot, restore, and fork efficiently.
  • Optimize GPU fleet utilization through intelligent scheduling, bin-packing, and capacity management.
  • Reduce cold-start latency and idle GPU costs while maintaining a serverless developer experience.
  • Own reliability and security for multi-tenant GPU infrastructure, including strong workload isolation guarantees.
  • Manage GPU driver lifecycle, hardware health, fleet maintenance, and production operational issues.
  • Partner with bare-metal and cloud infrastructure providers on GPU capacity, hardware qualification, and fleet expansion.
  • Lead a pod of four senior engineers while remaining a primary contributor on the team's most technically challenging systems.
  • Hire, mentor, and unblock engineers while setting a high bar for technical quality and execution.
  • Partner with Product on roadmap decisions, pricing inputs, SLAs, quotas, regions, and customer commitments.
  • Participate in on-call and own the production outcomes of the systems you and your team build.

Ideal Background

  • 8+ years of professional systems, infrastructure, or platform engineering experience.
  • 2+ years of experience leading engineers responsible for shipping production systems.
  • Deep hands-on Linux virtualization expertise, including KVM, QEMU, libvirt, VFIO/PCIe passthrough, and IOMMU.
  • Production experience operating GPU infrastructure at fleet scale.
  • Strong understanding of GPU scheduling, utilization, health monitoring, driver lifecycle, and hardware failure management.
  • Strong systems programming experience with Go, Rust, C, or C++.
  • Experience designing and operating high-performance, multi-tenant infrastructure.
  • Strong understanding of reliability, isolation, security, and performance in production infrastructure.
  • Comfortable acting as a player-coach—leading engineers while spending the majority of your time shipping production code.
  • Proven ability to own complex infrastructure from architecture through deployment, operations, and on-call.
  • Comfortable operating with significant autonomy in a fast-moving startup environment.

Preferred

  • Deep experience with NVIDIA data center GPUs such as H100s or A100s.
  • Experience with NVML, DCGM, and GPU observability or health-management tooling.
  • Strong understanding of NVIDIA MIG architecture and the tradeoffs between MIG and full GPU passthrough.
  • Experience building GPU clouds, AI compute platforms, serverless infrastructure, or hyperscaler compute systems.
  • Background at a GPU cloud provider, hyperscaler infrastructure organization, or high-scale compute platform.
  • Experience with GPU checkpoint/restore, workload snapshotting, or live workload migration.
  • Experience optimizing scheduling and bin-packing for expensive heterogeneous compute resources.
  • Familiarity with reinforcement learning, model training, inference, or agent-driven ML workloads.
  • Experience building infrastructure specifically for AI agents.
  • Strong understanding of the economics of GPU infrastructure, including utilization, capacity planning, and idle compute costs.

Compensation and Benefits

  • Competitive base salary plus equity.
  • Full-time position reporting directly to the CTO.
  • Location options in San Francisco or Croatia.
  • Lead a focused pod of four senior engineers alongside a dedicated Product Manager.
  • Player-coach structure where technical contribution remains the majority of the role.
  • Opportunity to own a technically challenging serverless GPU platform supporting production AI agent and reinforcement learning workloads.
  • Join shortly after the company's Series A during a period of rapid growth, with the platform already having grown approximately 10x
  • Significant ownership over GPU architecture, fleet economics, reliability, security, team development, and long-term technical direction.
Salary Range: $250,000-$300,000 base.
Recruiting from Scratch

About Recruiting from Scratch

Recruiting from Scratch provides recruiting services for companies that need to hire the best talent in software engineering, hardware engineering, product design, product management, marketing, GTM, and accounting & finance.

Industry
HR & Recruiting
Company Size
51-200 employees
Headquarters
New York, NY
Year Founded
2021
Social Media