San Jose, California, United StatesFull TimeEntry-levelPosted Today
The TikTok Agentic Arch team is building AI-native engineering infrastructure that transforms how software is designed, built, and operated at TikTok scale. We research and deploy reliable long-horizon agents that can reason over large codebases, interact with complex tools and environments, and complete consequential engineering work across distributed production systems.
Our environment provides a unique opportunity to advance agent research through real-world execution feedback. You will develop frontier methods, evaluate them rigorously, and bring them into production through platforms used at large scale—creating measurable improvements in engineering productivity, software quality, and operational effectiveness.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
- Conduct frontier research on reliable long-horizon agents for complex, multi-step software engineering and operations tasks.
- Advance agent capabilities in areas such as hierarchical planning and reasoning, memory and context management, tool use, environment interaction, and multi-agent coordination.
- Co-design agents and models through post-training, reinforcement learning, learning from execution feedback, search, and test-time scaling to improve performance on real-world tasks.
- Build scalable agent infrastructure for orchestration, evaluation, observability, and reliable execution across large codebases, engineering toolchains, and distributed production environments.
- Apply and validate new methods in representative workflows such as cross-repository software changes, testing and verification, large-scale migrations, deployment, and incident diagnosis and remediation.
- Translate research into production systems, define rigorous evaluation methodologies, and measure impact through task success, software quality, engineering efficiency, and system performance.
- Collaborate with researchers, infrastructure teams, developer-platform teams, and product engineers to deploy solutions at scale and produce publishable research and broader scientific insights where appropriate.
Minimum Qualifications
- Individuals who are completing or have recently completed a PhD degree in Computer Science or a related discipline.
- Demonstrated research or engineering experience in AI agents or closely related areas such as large language model reasoning, reinforcement learning, program synthesis, or AI for code.
- Strong understanding of one or more relevant areas, including agent planning and reasoning, model post-training, reinforcement learning, memory and context systems, tool learning, multi-agent systems, or agent evaluation.
- Strong programming and systems-building ability in at least one language such as Python, C++, Go, or Java, with the ability to turn research ideas into robust implementations.
- Ability to formulate ambiguous real-world problems, design rigorous experiments and evaluations, analyze results, and iterate from evidence.
- Strong communication and collaboration skills, with the ability to work across research, infrastructure, platform, and product teams.
Preferred Qualifications
- Evidence of research excellence through influential publications, open-source work, deployed systems, or other significant contributions. Relevant venues include NeurIPS, ICML, ICLR, ACL, MLSys, OSDI, SOSP, NSDI, ICSE, and FSE.
- Experience building or deploying LLM agents, AI developer tools, or agent platforms in complex or large-scale environments.•Hands-on experience with model post-training, reinforcement learning, execution-feedback loops, tool-using agents, distributed agent runtimes, or scalable evaluation systems.
- Experience working with large codebases, distributed systems, CI/CD and testing infrastructure, developer platforms, or production operations
- A track record of translating research into measurable production impact.
Our environment provides a unique opportunity to advance agent research through real-world execution feedback. You will develop frontier methods, evaluate them rigorously, and bring them into production through platforms used at large scale—creating measurable improvements in engineering productivity, software quality, and operational effectiveness.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Responsibilities
- Conduct frontier research on reliable long-horizon agents for complex, multi-step software engineering and operations tasks.
- Advance agent capabilities in areas such as hierarchical planning and reasoning, memory and context management, tool use, environment interaction, and multi-agent coordination.
- Co-design agents and models through post-training, reinforcement learning, learning from execution feedback, search, and test-time scaling to improve performance on real-world tasks.
- Build scalable agent infrastructure for orchestration, evaluation, observability, and reliable execution across large codebases, engineering toolchains, and distributed production environments.
- Apply and validate new methods in representative workflows such as cross-repository software changes, testing and verification, large-scale migrations, deployment, and incident diagnosis and remediation.
- Translate research into production systems, define rigorous evaluation methodologies, and measure impact through task success, software quality, engineering efficiency, and system performance.
- Collaborate with researchers, infrastructure teams, developer-platform teams, and product engineers to deploy solutions at scale and produce publishable research and broader scientific insights where appropriate.
Minimum Qualifications
- Individuals who are completing or have recently completed a PhD degree in Computer Science or a related discipline.
- Demonstrated research or engineering experience in AI agents or closely related areas such as large language model reasoning, reinforcement learning, program synthesis, or AI for code.
- Strong understanding of one or more relevant areas, including agent planning and reasoning, model post-training, reinforcement learning, memory and context systems, tool learning, multi-agent systems, or agent evaluation.
- Strong programming and systems-building ability in at least one language such as Python, C++, Go, or Java, with the ability to turn research ideas into robust implementations.
- Ability to formulate ambiguous real-world problems, design rigorous experiments and evaluations, analyze results, and iterate from evidence.
- Strong communication and collaboration skills, with the ability to work across research, infrastructure, platform, and product teams.
Preferred Qualifications
- Evidence of research excellence through influential publications, open-source work, deployed systems, or other significant contributions. Relevant venues include NeurIPS, ICML, ICLR, ACL, MLSys, OSDI, SOSP, NSDI, ICSE, and FSE.
- Experience building or deploying LLM agents, AI developer tools, or agent platforms in complex or large-scale environments.•Hands-on experience with model post-training, reinforcement learning, execution-feedback loops, tool-using agents, distributed agent runtimes, or scalable evaluation systems.
- Experience working with large codebases, distributed systems, CI/CD and testing infrastructure, developer platforms, or production operations
- A track record of translating research into measurable production impact.
