Job Description
Infrastructure Services is part of IS&T and the foundation of Apple's global network operations — managing data center equipment and systems to deliver compute, storage, and networking services for teams across Apple, including its internal developer community. From individual facilities to a worldwide network, Infrastructure Services ensures the technology underneath everything works without question
Apple's services reach hundreds of millions of people every day, and Edge Services builds and runs the authoritative DNS and global server load balancing infrastructure that sits at the front door of all of them. Every request to an Apple service begins here, in the first few milliseconds where our systems decide how and where to answer. The reliability, correctness, and performance of that front door directly shape the experience people have with Apple, and keeping it fast and dependable at global scale is the work of our team.
We're looking for a thoughtful, experienced engineer to bring technical leadership to this infrastructure. You'll operate and extend bare-metal and cloud systems distributed across dozens of data centers and network points of presence worldwide, build deep traffic and resolution visibility, drive performance and capacity planning, and create the automation that keeps our 24x7 operations calm and reliable. You'll also mentor engineers across the team and partner closely with our software, network, and product engineering groups. We believe great systems are built by great teams and care as much about how we work together as about what we ship. If you're energized by owning the reliability of critical infrastructure at global scale, we'd like to hear from you.
This is a senior individual-contributor role and a technical anchor for the Edge Services DNS and GSLB systems. The essential functions of the job are to operate and extend distributed Linux infrastructure as code across geographically dispersed sites; to define and maintain the observability, alerting, and service-level objectives that keep the systems healthy; to build automation that removes operational toil; to plan fleet capacity and hardware lifecycle across globally distributed sites; and to mentor engineers while partnering across software, network, and product teams. The role includes incident and process management and participation in a shared 24x7 production on-call rotation.
Preferred Qualifications
Extensive experience operating large-scale infrastructure across multiple data centers or network points of presence.
Expertise with anycast routing and BGP, and a strong understanding of how DNS, load balancing, and the network layer interact.
Depth in Linux internals, including kernel networking, and in package management and software deployment at fleet scale.
Experience defining service-level objectives and indicators (SLOs/SLIs) and driving reliability programs across distributed systems.
Familiarity with secure-by-default operations, such as DNSSEC key management, change safety, and progressive configuration rollout.
Experience with scale and performance testing, disaster recovery, and capacity planning.
Experience managing hardware lifecycle across edge and regional network environments.
Track record of leading end-to-end projects, defining technical roadmaps, and driving cross-functional alignment on architecture and best practices.
Experience mentoring engineers and leading code reviews.
Minimum Qualifications
Demonstrated experience operating Linux systems in production, including software deployment and CI/CD workflows.
Working knowledge of networking fundamentals, with hands-on experience troubleshooting TCP/UDP and common layer 2–3 issues.
Experience with configuration management or infrastructure-as-code tooling (for example, Salt, Ansible, Puppet, or Terraform) and with observability tooling (for example, Prometheus and Grafana, or equivalents).
Proficiency in at least one programming language used for automation and tooling (for example, Python, Go, Rust, or Swift).
Willingness to participate in a shared 24x7 on-call rotation.
Bachelor's degree in Computer Science or a related field, or equivalent practical experience.