Genentech

Orchestration Workload Engineer - ACE - AI Factory

Genentech  •  Switzerland (Onsite)  •  10 days ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

As a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be recognised internally as an expert in workload orchestration, owning and advancing our scheduler tech stack across our High-Performance Computing (HPC) platforms. With the rapid expansion of our compute infrastructure, your broad expertise will drive the efficient scheduling, policy management, and resource optimization of our multi-node CPU and GPU environments.

In this role, you will use your expertise to bridge traditional scientific computing with modern AI paradigms, while acting as a coach and mentor to help colleagues develop technical expertise. You will solve unique, unprecedented scheduling and infrastructure challenges that directly impact Roche’s compute architecture, ensuring our researchers, data scientists, and engineers can execute compute workloads reliably, efficiently, and successfully.

Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.

The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute—helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value.

The Opportunity

SLURM Architecture & Ecosystem Leadership

  • Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC (and AI environments) to ensure high availability and dynamic resource distribution.
  • Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.
  • Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms (e.g., SLURM Slinky or Run:ai) to streamline job submission workflows across architectures.

Hybrid Workload & Kubernetes Integration

  • Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration to support hybrid AI/HPC workloads.
  • Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.

Technical Leadership, Mentorship & Governance

  • Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns across Roche compute environments.
  • Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning across the chapter.
  • Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization, utilizing configuration-as-code to deploy policies uniformly.

Who You Are

  • Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.
  • Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.
  • Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.
  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.
  • SLURM Architecture & Optimization: Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources.
  • SLURM Operations, Accounting & Observability: Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions.
  • Kubernetes & Container Knowledge: Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context.
  • AI Infrastructure & Interconnects: Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).
  • Automation: Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines.
  • Broad Platform Expertise: Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management to solve orchestration problems spanning multiple technology domains.
  • Domain Expertise & Problem Solving: Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads.
  • Coaching & Collaboration: Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams.
  • Cross-Organizational Coordination: Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders.
  • Strategic Vision: Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows.

#RDT2026

Genentech is an equal opportunity employer. It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence. The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.

If you have a disability and need an accommodation in relation to the online application process, please contact us by completing this form Accommodations for Applicants.

Genentech

About Genentech

About Genentech

We're passionate about finding solutions for people facing the world's most difficult-to-treat conditions. That is why we use cutting-edge science to create and deliver innovative medicines around the globe. To us, science is personal.

Making a difference in the lives of millions starts when you make a change in yours. If you’d like to join our team, view our openings at gene.com/careers.

Our patient resource center is dedicated to getting patients and caregivers to the right resources. You can reach them at 1 (877) GENENTECH (436-3683)

Monday-Friday, 6am-5pm PST or patientinfo@gene.com.

Community Guidelines:

1. We want to foster positive conversation around the issues we are passionate about. To that end, we remove profanity, content that contains threatening language, content that is aimed at private individuals, personal information, and repeated unwanted messages.

2. Don’t mention any medicines by name — ours or anyone else’s.

Because of the fair balance rules governing our industry, we cannot post any comments that reference any pharmaceutical brand, product, or service. Please do not mention any specific medicines by name, or include any links to third party sites in your comments.

3. This isn’t the place to report or discuss side effects.

This site is not intended as a forum for reporting side effects experienced while taking a Genentech product. Instead, you should report any side effects to Genentech Drug Safety at 1-888-835-2555. You can also report side effects of any prescription product directly to the FDA at 1-800-FDA-1088 or by visiting www.FDA.gov/medwatch.

4. Don’t pitch your product or service.

Please don't use our page as a place to promote your product or pitch your services. Please also avoid posting links to external sites. We reserve the right to remove any posts that are deemed promotional.

Industry
Biotech & Life Sciences
Company Size
10,000+ employees
Headquarters
South San Francisco, California
Year Founded
1976
Website
gene.com
Social Media