
Recommendation Architecture AI/ML Infrastructure Engineer Intern (Data-Arch-TikTok Live) - 2027 Summer
TikTokSan Jose, California, United StatesInternshipEntry-levelPosted Today
Our team develops the core training and serving infrastructure that powers one of the world's largest recommendation systems, enabling billions of personalized recommendations every day. We are also advancing the next generation of AI infrastructure for foundation models and LLMs, driving innovation in large-scale model training, online inference, and GPU optimization.
As part of the team, you will work on distributed training and inference systems, high-performance GPU computing, and scalable LLM infrastructure. You'll collaborate closely with experienced engineers and researchers to transform cutting-edge AI technologies into production systems that directly impact the experience of hundreds of millions of TikTok users. This role is ideal for candidates who are passionate about LLM systems, distributed computing, GPU programming, and building AI systems at massive scale.We are looking for talented individuals to join our Model Infrastructure team, building the next generation of infrastructure for TikTok's For You recommendation system and Large Language Models (LLMs).
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Online Assessment
Candidates who pass resume screening will be invited to participate in Our Company's technical online assessment.
Responsibilities
• Build and optimize infrastructure for large-scale model training and online inference.
• Develop distributed systems supporting large recommendation models and LLMs.
• Improve training and inference performance through GPU optimization and efficient communication.
• Collaborate with researchers to develop and deploy LLM training and serving solutions.
• Analyze system bottlenecks and implement performance optimizations.
Minimum Qualifications:
- Currently pursuing a Bachelor's or Master's degree in Artificial Intelligence, Software Development, Computer Science, Computer Engineering or a related technical discipline.
- Strong programming skills in C++ or Python. •Good understanding of data structures, algorithms, and computer systems.
- Familiarity with PyTorch or TensorFlow.
- Knowledge of Transformer architectures and Large Language Models (LLMs).
- Strong problem-solving skills and a passion for building large-scale AI systems.
Preferred Qualifications:
- Hands-on experience with LLM training or inference through research, internships, or open-source projects. •Familiarity with distributed training concepts (e.g., DP, TP, PP, FSDP, ZeRO).
- Experience with GPU programming using CUDA, Triton, or similar technologies.
- Understanding of LLM serving techniques such as KV Cache, Continuous Batching, or FlashAttention. •Contributions to open-source projects or research in machine learning systems, distributed systems, or LLM infrastructure.
By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://careers.tiktok.com/legal/privacy
As part of the team, you will work on distributed training and inference systems, high-performance GPU computing, and scalable LLM infrastructure. You'll collaborate closely with experienced engineers and researchers to transform cutting-edge AI technologies into production systems that directly impact the experience of hundreds of millions of TikTok users. This role is ideal for candidates who are passionate about LLM systems, distributed computing, GPU programming, and building AI systems at massive scale.We are looking for talented individuals to join our Model Infrastructure team, building the next generation of infrastructure for TikTok's For You recommendation system and Large Language Models (LLMs).
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Online Assessment
Candidates who pass resume screening will be invited to participate in Our Company's technical online assessment.
Responsibilities
• Build and optimize infrastructure for large-scale model training and online inference.
• Develop distributed systems supporting large recommendation models and LLMs.
• Improve training and inference performance through GPU optimization and efficient communication.
• Collaborate with researchers to develop and deploy LLM training and serving solutions.
• Analyze system bottlenecks and implement performance optimizations.
Minimum Qualifications:
- Currently pursuing a Bachelor's or Master's degree in Artificial Intelligence, Software Development, Computer Science, Computer Engineering or a related technical discipline.
- Strong programming skills in C++ or Python. •Good understanding of data structures, algorithms, and computer systems.
- Familiarity with PyTorch or TensorFlow.
- Knowledge of Transformer architectures and Large Language Models (LLMs).
- Strong problem-solving skills and a passion for building large-scale AI systems.
Preferred Qualifications:
- Hands-on experience with LLM training or inference through research, internships, or open-source projects. •Familiarity with distributed training concepts (e.g., DP, TP, PP, FSDP, ZeRO).
- Experience with GPU programming using CUDA, Triton, or similar technologies.
- Understanding of LLM serving techniques such as KV Cache, Continuous Batching, or FlashAttention. •Contributions to open-source projects or research in machine learning systems, distributed systems, or LLM infrastructure.
By submitting an application for this role, you accept and agree to our global applicant privacy policy, which may be accessed here: https://careers.tiktok.com/legal/privacy