Job Description
Our team develops the core training and serving infrastructure that powers one of the world's largest recommendation systems, enabling billions of personalized recommendations every day. We are also advancing the next generation of AI infrastructure for foundation models and LLMs, driving innovation in large-scale model training, online inference, and GPU optimization. As part of the team, you will work on distributed training and inference systems, high-performance GPU computing, and scalable LLM infrastructure. You'll collaborate closely with experienced engineers and researchers to transform cutting-edge AI technologies into production systems that directly impact the experience of hundreds of millions of TikTok users. This role is ideal for candidates who are passionate about LLM systems, distributed computing, GPU programming, and building AI systems at massive scale. We are looking for passionate students to join our Model Infrastructure team, building the next generation of infrastructure for TikTok's For You recommendation system and Large Language Models (LLMs).
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Successful candidates must be able to commit to at least 3 months long internship period.
Responsibilities
- Design, build and optimize distributed training infrastructure and low-latency online inference systems for large-scale recommendation models and Large Language Models
- Develop high-performance GPU kernel implementations and efficient inter-node communication primitives to improve training and inference efficiency
- Build compiler optimization passes and operator fusion technologies for deep learning frameworks to accelerate model execution
Minimum Qualification(s):
- Currently pursuing a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Information Systems, or a related technical discipline
- Strong programming skills in C++ or Python, with solid understanding of data structures and algorithms
- Familiarity with PyTorch or TensorFlow, and basic knowledge of Transformer architectures and Large Language Models (LLMs)
Preferred Qualification(s):
- Hands-on experience with LLM training or inference through academic projects, internships, or open-source contributions
- Familiarity with distributed training concepts (DP, TP, PP, FSDP, ZeRO) or GPU programming technologies (CUDA, Triton)
- Strong problem-solving skills and passion for building large-scale AI systems