Job Description
Team Introduction
Our team is responsible for building ByteDance's cloud-native infrastructure, which powers all of the company's products. We specialize in cutting-edge technologies in areas such as Kubernetes cluster management, runtime resource optimization, multi-cloud and multi-cluster solutions, and cloud-native infrastructure stability. The team has released open-source projects such as Kubebrain and Katalyst, and we are focused on groundbreaking research and development of core technologies around resource pooling and elasticity.
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Key Responsibilities
- Build and Optimize Large-Scale Kubernetes Clusters;
- Design and evolve system architectures for ultra-large-scale Kubernetes clusters,
- Continuously improve the performance and stability of control systems in scenarios like big data and machine learning.
- Define and Enhance Kubernetes SLOs;
- Establish and optimize Service Level Objectives (SLOs) for Kubernetes clusters,
- Analyze end-to-end latency, identify performance bottlenecks, propose solutions, and implement them in production environments.
- Develop Observability for Kubernetes Clusters;
- Build and enhance observability systems to improve issue diagnosis efficiency,
- Create observability data warehouses and use data-driven approaches to optimize cluster performance.