Job Description
The Data-AML-Engine Orchestration team builds large-scale machine learning infrastructure that powers online model serving across ByteDance products, including TikTok. We develop the orchestration, scheduling, and resource management systems that connect heterogeneous compute infrastructure with production ML workloads.
You will work on systems that directly affect GPU utilization, serving latency and availability, infrastructure reliability, and MLE productivity. Depending on your background and interests, you may focus on one or more of the following areas.
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Responsibilities:
- Design and build foundational orchestration capabilities for machine learning platforms, including Kubernetes Operators, container runtimes, and lifecycle management for jobs, services, and stateful workloads.
- Build multi-tenant resource and quota systems that support priorities, preemption, fair sharing, elasticity, and cross-cluster scheduling. Improve GPU utilization and cost efficiency through resource pooling and FinOps.
- Build lifecycle orchestration for online model serving, including model and image distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operation, and disaster recovery.
- Build serving orchestration and traffic management capabilities for disaggregated serving clusters, including topology-aware scheduling, KV Cache affinity, intelligent request routing, and QoS/SLA management.
annually.