Job Description
Imagine what we could do together. At Apple, new ideas have a way of becoming excellent products, services, and customer experiences very quickly. Bring passion and dedication to your job and there’s no telling what you could accomplish! The people here at Apple don’t just build products - they craft the kind of wonder that’s revolutionized entire industries. It’s the diversity of those people and their ideas that encourages the innovation that runs through everything we do, from amazing technology to industry-leading environmental efforts.
At ETS Team, we take pride in developing groundbreaking, world-changing platforms and services. Our ETS applications play a crucial role in supporting the Apple ecosystem by offering identity management, factory and device support, infrastructure support, platform support, and collaboration tools. Whether you're logging into Apple, making a purchase, or enabling Apple devices, our applications are there every step of the way, ensuring a flawless and secure experience.
As Site Reliability and Operations Engineer (SRE), you’ll be part of the action—working closely with application teams to automate operations, optimize infrastructure, and solve issues in an exciting, fast-paced environment. You’ll play a vital role in ensuring that our systems are reliable, scalable, and high-performing.
This role is designed for driven individuals who:
- Have a passion for designing and building reliable systems with delivering quality in a dynamic, high-energy workplace
- Strong sense of ownership and integrity demonstrated through clear communication and collaboration and can stay calm under pressure.
- Love learning new technologies and thrive in solving sophisticated challenges.
- Are independent, motivated, and excited to take on ambitious projects.
We are seeking dedicated Site Reliability Engineers (SREs) to join our ETS Platform team.
Preferred Qualifications
Experience managing and scaling distributed systems in public, private, or hybrid cloud environments.
Hands-on experience operating large fleets of diverse systems using automation/configuration platforms (Puppet, Ansible, Spinnaker)
Strong proficiency with Linux, core networking concepts (TLS/SSL, DNS, load balancers), and troubleshooting at scale.
Understanding of security standards, cryptography, and enterprise policies.
Experience with Incident / Problem management and RCA
Strong Network, Load Balancing (Nginx, Envoy, NetScaler) experience is a huge plus
Good solid understanding using Kubernetes concepts such as networking, Storage, Secrets, Deployments & Containerization.
Hands-on experience with AliCloud, AWS, or GCP is preferred.
Strong analytical skills, Demonstrated ability to deliver results on time with high quality
Familiarity with micro services architecture and container orchestration with Docker & Kubernetes
Minimum Qualifications
Bachelor’s or Master’s degree in Computer Science or a related field (or equivalent practical experience).
5–7 years in a Reliability Engineering, DevOps, or infrastructure-focused role.
Advanced experience with programming languages or scripting (Python / Bash / LUA / GoLang )
Hands-on experience with relational and/or NoSQL databases (Oracle, MongoDB )
Deep understanding of systems and infrastructure fundamentals.
Advanced knowledge and hands-on experience with CI/CD pipelines, Git-based source control, release engineering, and DevOps practices.
Experience deploying, supporting, and monitoring new and existing services, platforms, and application stacks.
Strong advocate for automating manual operational work through software leveraging AI/GenAI technologies.
Solid understanding of Linux internals, standard networking protocols, and distributed system components.