Apple

Site Reliability Engineer — Data Platforms, IS&T Ai & Data Platforms

Apple  •  Shanghai, CN (Onsite)  •  17 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

The AiDP Data Platforms team builds and operates data platforms at scale on the Cloud, helping Apple process, store, and access petabytes of data. We’re seeking an SRE to own the reliability, performance, and operability of our data and ML platforms — someone who thinks in SLOs, failure modes, and blast radius, and is passionate about keeping large-scale distributed systems fast, available, and cost-efficient.

You’ll operate and harden our big data platform — built on open source and other technologies — that powers critical applications like analytics, reporting, and AI/ML. This means driving down MTTR, automating operational toil, tuning performance and cost, running capacity planning, and root-causing production incidents before and after they happen. You’re an independent, self-directed problem-solver who communicates clearly with both engineers and non-technical partners, and you’ll work across many teams to keep the platform running at Apple’s standard.

Preferred Qualifications

Experience contributing to open source projects or operating across multiple public cloud providers.
In-depth operational knowledge of specific distributed frameworks — Spark, Flink, or Kafka Streams, Trino, Iceberg — including debugging via component-specific logs, and experience managing multi-tenant Kubernetes clusters at scale.
Experience with workflow/pipeline orchestration tools (Airflow, dbt) and understanding of data modeling and warehousing concepts.
Experience debugging Kubernetes/Spark production issues via logs and metrics, with a continuous-improvement mindset for self, team, and org.
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.

Minimum Qualifications

3+ years of experience operating and supporting critical, large-scale distributed systems in production, with scripting/programming ability in Python, Go, Java, Scala, or Bash for automation and tooling.
Deep understanding of reliability principles — fault tolerance, high availability, low latency, graceful degradation — and how to instrument and enforce them (SLIs/SLOs, alerting, on-call practices, incident response, postmortems).
Hands-on experience operating data processing ecosystems and distributed computing frameworks (Spark, Flink) and MPP query engines (Trino, StarRocks), including performance tuning and capacity management.
Proficiency operating Kubernetes/Helm at scale, building and maintaining CI/CD pipelines (GitHub Actions, Jenkins), managing infrastructure as code (Terraform, Pulumi), and running service-oriented architectures across multi-cloud environments.
Strong troubleshooting and performance analysis skills in complex production environments; fluency in Unix/Linux and command-line diagnostics, with excellent problem-solving and communication skills.
Apple

About Apple

We’re a diverse collective of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. And the same innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it. This is where your work can make a difference in people’s lives. Including your own.

Apple is an equal opportunity employer that is committed to inclusion and diversity. Visit apple.com/careers to learn more.

Industry
Hardware & Semiconductors
Company Size
10,000+ employees
Headquarters
Cupertino, California
Year Founded
1976
Website
apple.com
Social Media