We are hiring a Site Reliability Engineer to Build, support and improve Keystone, an ETL application/platform operating in the China region. Keystone supports data ingestion, transformation, and loading workflows across Kubernetes-based environments, including Airflow-based loader jobs that load data into Snowflake and Spark-on-EKS loader jobs that load data into a Datalake or Lakehouse environment.
The SRE will own reliability, availability, performance, data freshness, and operational readiness for production ETL pipelines. This role requires strong hands-on experience with Linux, Kubernetes, Spark, Airflow, Kafka or streaming systems, Snowflake, object storage, observability, and incident response.
The ideal candidate can troubleshoot distributed data platform issues from logs, metrics, timestamps, pipeline state, and infrastructure signals, and can drive permanent fixes through configuration improvements, automation, tuning, and engineering partnership.
This position is based in Shanghai, China.
Key Responsibility
Experience supporting metadata-driven ETL platforms or internal ETL frameworks similar to Keystone.
Experience with Iceberg, Hive Metastore, Parquet, ORC, or similar Lakehouse technologies.
Experience with GitOps or source-of-truth configuration management.
Experience with Snowflake performance tuning, warehouse sizing, query history analysis, and load optimization.
Experience with Spark performance tuning at scale.
Experience operating multi-region or region-specific data platforms.
Experience with certificate, PKI, TLS, JKS/truststore, Kerberos, or key rotation processes.
Experience with infrastructure as code tools such as Terraform, Helm, Argo CD, Ansible, or similar.
Familiarity with API gateways, RabbitMQ, Redis, Cassandra, or other platform dependencies.
4+ years of experience in SRE, DevOps, platform engineering, data infrastructure, or production operations.
Strong Linux troubleshooting skills.
Strong Kubernetes/EKS operations experience, including kubectl, deployments, statefulsets, pods, resource limits, service accounts, secrets, logs, events, and workload debugging.
Hands-on experience supporting Apache Spark in production, preferably Spark on Kubernetes.
Experience tuning Spark jobs, reading driver/executor logs, troubleshooting memory issues, shuffle failures, failed stages, and performance bottlenecks.
Production experience with Apache Airflow, including DAG operations, task failures, retries, scheduling, SLA misses, and dependency troubleshooting.
Practical experience with Kafka or similar messaging/streaming platforms, including consumer groups, lag, offsets, partitions, and secure client connectivity.
Familiarity with Snowflake loading patterns, connectivity, roles, warehouses, stages, query/load history, and load monitoring.
Familiarity with Datalake or Lakehouse architectures using object storage such as S3 or equivalent.
Solid SQL skills and ability to analyze large operational or data pipeline datasets.
Experience with monitoring and logging tools such as Splunk, Prometheus, Grafana, CloudWatch, ELK, Datadog, or equivalent.
Scripting experience with Python and Bash.
Strong incident management, RCA, change management, and operational documentation skills.
Ability to troubleshoot distributed systems using evidence from logs, metrics, events, timestamps, and configuration.
Professional working proficiency in Mandarin and English.

We’re a diverse collective of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. And the same innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it. This is where your work can make a difference in people’s lives. Including your own.
Apple is an equal opportunity employer that is committed to inclusion and diversity. Visit apple.com/careers to learn more.