Job Description
Imagine what you could do here. At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. We are looking for great Systems Reliability Engineers to design, build and operate tools and automated systems to help run the edge infrastructure that delivers Apple's services to customers across China. If you are passionate and curious in how internet and DNS works at a fundamental level and like operating systems at scale, this is the role for you.
You will be in charge of designing and building the infrastructure to support services at Apple scale. Day to day, you will design and tune resolution paths, traffic-steering and failover policies, and capacity; automate operations; harden the platform against load spikes and attacks; and lead the response when incidents occur. You have a passion for automation and making resilient systems that scale. The ideal candidate has a strong knowledge of internet protocols foundation, solid Linux skills, cloud and compute experience, working familiarity with proxies and load balancers, along with strong coding ability in Go and scripting languages .
Preferred Qualifications
Hands-on experience with reverse proxies or load balancers such as NGINX or Envoy
Configuration management systems such as Saltstack, Puppet, or Chef
Experience with continuous / rapid release engineering
Strong tooling and automations development experience
Experience working in a 24/7/365 service environment
Deep Linux and networking experience: kernel and network-stack performance tuning, packet-level troubleshooting, and operating networks at scale
Deep experience running DNS and GSLB at large scale: authoritative and recursive DNS platforms (e.g., BIND, NSD, Knot, PowerDNS) and commercial GSLB systems, with anycast and BGP-based routing for global traffic distribution
DNS security at scale: DNSSEC, plus DDoS mitigation and resilience at the DNS layer
Experience designing for high availability and disaster recovery across regions
Depth in cloud platforms and compute, with infrastructure-as-code and CI/CD
A track record building SRE practices: SLOs, error budgets, and blameless incident response
Minimum Qualifications
BS in Computer Science or a related field, or equivalent job-related experience
Linux systems administration experience
5+ years operating production infrastructure at scale at a SRE level
Hands-on experience operating DNS in production (authoritative and/or recursive)
Proficiency in a programming or scripting language such as Go or Python
Experience operating systems in cloud platforms
Fluent English and Mandarin