Rivian

Staff Software Engineer, Compute & Storage, Autonomy

Rivian  •  $207k - $258k/yr  •  United States (Onsite)  •  5 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Location(s): Palo Alto, California

Category: Autonomous Driving

About Rivian

Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract.

As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.

Role Summary

Rivian's Autonomy org needs a Staff Software Engineer, Compute & Storage to own the distributed compute and storage platform that every autonomy workload runs on. This sits in the Platform Services team in the AI Platform organization in the Autonomy team. Autonomy training, simulation base evaluation / validation, and autonomy visualization all depend on the same two things: available compute and fast access to data. The role requires deep expertise in Kubernetes-based distributed compute, large-scale object storage, and the performance and cost tradeoffs of running both at petabyte scale.

You'll work with the AI Platform, Perception, Planning, Simulation, and Vehicle Integration, Product Management, and other technology partners to operate a platform serving a fleet of over 100,000 vehicles, hundreds of petabytes of drive and simulation data, and training clusters of thousands of GPUs across multiple clouds. This is a platform ownership role, and it's measured by what it lets other engineers do: how fast someone goes from idea to trained model or data pipeline, how many scenarios simulation runs per day, and what each of those costs.

Responsibilities

  • Own the architecture and roadmap for Autonomy's distributed compute platform: job scheduling, quota and fair-share across teams, autoscaling, and spot and preemption strategy across multiple clouds.
  • Build scalable tools and APIs that turn high-level job requests into executed work, making large-scale computation accessible to engineers who aren't infrastructure specialists.
  • Read the jobs other teams run, profile them, and find inefficiencies and bottlenecks. Fix them directly, or give the team the tooling to see them.
  • Treat cluster efficiency as a primary metric: eliminate GPU fragmentation, right-size quota, and reclaim capacity stranded on partially filled nodes.
  • Own the storage architecture for autonomy data at petabyte scale, including layout, partitioning, tiering across hot, warm and cold, lifecycle policy, and the caching and prefetch layers that keep training and simulation jobs from starving on I/O.
  • Own the multi-cloud compute and data path as training extends beyond a single provider, including replication strategy, consistency, and cross-provider egress economics.
  • Drive throughput and turnaround time for the heaviest workloads: training data loading, large-scale log replay, and batch resimulation running tens of thousands of concurrent jobs.
  • Own the queue-based and event-driven infrastructure behind job submission and autoscaling, and keep it stable as queue depth moves by orders of magnitude.
  • Own platform observability: define the metrics and build the dashboards and alerting that make job throughput, queue health, cluster utilization, and per-team cost visible to you and the teams you serve.
  • Define SLAs for job admission, completion, and data availability. Set the on-call strategy and take part in the rotation.
  • Own compute and storage cost as a first-class engineering metric, instrumented per team and per workload, across multiple AWS accounts and services.
  • Work with the security & privacy team on data governance, access control, retention, and audit for vehicle-collected data.
  • Set technical standards for how distributed workloads are built and run, and raise the bar through design review, code review, and mentorship of senior engineers.

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 8+ years of software engineering experience building and operating distributed systems in production.
  • 5+ years owning a distributed compute, batch execution, or scheduling platform used by other engineering teams: building the scheduler, orchestration layer, or execution engine itself, not just jobs that ran on one.
  • 5+ years with Kubernetes as an execution substrate: scheduling, resource management, custom controllers or operators, autoscaling, and failure modes at scale.
  • 5+ years with large-scale object storage: data layout, partitioning, lifecycle and tiering, caching, and the performance and cost tradeoffs among them.
  • 3+ years with a distributed processing or training framework (Ray, Spark, Flink, or Dask).
  • 3+ years hands-on with Go, C++ or Rust, plus strong Python.
  • 3+ years with infrastructure as code and configuration management (Terraform, AWS CDK, or CloudFormation), and container build and supply chain.
  • 3+ years with queueing and event-driven systems (SQS, Kafka, Kinesis, or equivalent), including autoscaling workloads on queue depth.
  • 3+ years with production monitoring and alerting (Prometheus, Grafana, Datadog, or CloudWatch): defining the metrics and building the dashboards, not just consuming them.
  • 3+ years debugging production distributed systems, including resource contention, stragglers, data skew, network saturation, and cascading failure, and running root-cause analysis on incidents.
  • 2+ years carrying an infrastructure cost target and meeting it.
  • Ability to turn ambiguous, high-level requirements into a detailed system design and drive it to completion unprompted.
  • Technical influence beyond your own commits: designs others build on, standards others adopt, engineers who improved from working with you.
  • Nice to have: GPU cluster management and utilization, including topology-aware placement, gang scheduling, multi-tenant sharing, and fragmentation.
  • Nice to have: autonomous vehicle, robotics, or another domain with continuous high-volume sensor data.
  • Nice to have: multi-cloud or hybrid cloud and on-premise operation, including data movement economics between providers.
  • Nice to have: Linux internals, including the CPU scheduler, memory management, file systems, and networking.

Pay Disclosure

The salary range for this role is $206,500-$258,100 for San Francisco Bay Area based applicants. This is the lowest to highest salary we in good faith believe we would pay for this role at the time of this posting. An employee’s position within the salary range will be based on several factors including, but not limited to, specific competencies, relevant education, qualifications, certifications, experience, skills, geographic location, shift, and organizational needs.

We offer a comprehensive package of benefits for full-time and part-time employees, their spouse or domestic partner, and children up to age 26, including but not limited to paid vacation, paid sick leave, and a competitive portfolio of insurance benefits including life, medical, dental, vision, short-term disability insurance, and long-term disability insurance to eligible employees. You may also have the opportunity to participate in Rivian’s 401(k) Plan and Employee Stock Purchase Program if you meet certain eligibility requirements. Full-time employee coverage is effective on their first day of employment. Part-time employee coverage is effective the first of the month following 90 days of employment. More information about benefits is available at rivianbenefits.com.

Equal Opportunity

Rivian is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law.

Rivian is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at candidateaccommodations@rivian.com.

Candidate Data Privacy and Technology

Rivian may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes (“Candidate Personal Data”). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) recordkeeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.

Rivian may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian affiliates; and (iii) Rivian’s service providers, including providers of background checks, staffing services, and cloud services.

Rivian may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, the United Kingdom, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions.

How We Use AI in Our Hiring Process: To ensure transparency, we want candidates to know that Rivian uses iCIMS Talent Cloud Artificial Intelligence (TCAI) and AI-enabled tools to assist with screening, reviewing, organizing and highlighting profiles and applications that match the key requirements for each role.

AI does not make hiring decisions: Qualified candidate applications are reviewed by a member of our team, and all decisions throughout the process are made by humans. We use AI to support efficiency and consistency, not to replace human judgment. We are committed to a fair, thoughtful, and equitable experience for every candidate.

Participation in AI profile matching is entirely voluntary. If you prefer that your profile not be used in this process, you can opt out at any time. Opting out means your profile will be excluded from automated matching and will not be surfaced for additional roles through this system. Your current application remains active and will not be affected in any way.

Please note that we are currently not accepting applications from third party application services.

Rivian

About Rivian

Doing something different is never easy. It requires courage, optimism and grit. Core to our mission is building a team of adventurous individuals determined to make a positive impact on the world. This means challenging ourselves constantly. Stretching beyond the bounds of conventional thinking. Reframing old problems. Seeking new solutions. And operating comfortably in a space of uncertainty. While our backgrounds are diverse, our team shares a love of the outdoors and a desire to protect it for future generations. Do you like doing the impossible? We’d love to hear from you.

Industry
Automotive & Mobility
Company Size
10,000+ employees
Headquarters
Irvine, California
Year Founded
2009
Social Media