Job Description
Team Introduction:
Our Site Reliability Engineering (SRE) team blends software and systems engineering to build and operate large-scale data infrastructure with high reliability and efficiency. We provide a dependable cloud environment that powers our global business. In this role, you will leverage your expertise in data center architecture, data infrastructure services, and systems and tools development to solve complex scaling and reliability challenges.
We’re looking for a Technical Lead (SRE) who can provide deep technical leadership, drive architectural improvements, and collaborate effectively across multiple organizations. You’ll partner with engineering, product, data, and infrastructure teams to deliver resilient, scalable platforms. This is a highly technical, hands-on role that requires strong problem-solving ability, clear communication, and the ability to influence without formal authority.
Responsibilities
- Strong hands-on skills in the design, development, and operation of large-scale cloud infrastructure and distributed systems.
- Collaborate with cross-functional teams (e.g., Advertising, Machine Learning, E-commerce, and Core Infra) to drive system reliability, performance, and scalability.
- Lead initiatives to automate operations, eliminate toil, and improve overall system efficiency.
- Troubleshoot complex production issues, perform root-cause analysis, and drive long-term reliability improvements.
- Promote best practices in system design, observability, performance optimization, and cost efficiency.
- Communicate complex technical concepts effectively to both technical and non-technical stakeholders.
The base salary range for this position in the selected city is $208800 - $438000 annually.