Singapore, SingaporeFull TimeEntry-levelPosted Today
Team Introduction
Volcano Ark is Volcano Engine's one-stop foundation model and Agent platform for enterprises and developers. Our team is building the next-generation infrastructure for general-purpose Agents capable of handling complex tasks—from Agent runtimes and execution engines, to platform-level resource models, lifecycle management, and public APIs, as well as evaluation and observability systems.
We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.
Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to ByteDance and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities
-For training track, develop the Volcano Ark training platform, enabling both internal and external users to perform serverless post-training (e.g., Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL)) on the Ark platform.
-Design elastic training solutions for complex multi-tenant workloads, supporting mixed-tenant training across multiple data centers and heterogeneous hardware while optimizing training throughput, resource utilization, and system stability.
-Build next-generation reinforcement learning infrastructure to improve training efficiency, while designing intuitive and developer-friendly APIs for RL training workflows.
-For inference track, develop the Volcano Ark MaaS (Model-as-a-Service) inference platform, optimizing the performance, cost efficiency, and reliability of large language model inference across hyperscale heterogeneous GPU clusters.
Reduce inference costs through advanced system optimizations, including disaggregated inference architectures, distributed KV cache systems, heterogeneous inference, elastic compute scheduling, and multi-tenant co-located inference.
Volcano Ark is Volcano Engine's one-stop foundation model and Agent platform for enterprises and developers. Our team is building the next-generation infrastructure for general-purpose Agents capable of handling complex tasks—from Agent runtimes and execution engines, to platform-level resource models, lifecycle management, and public APIs, as well as evaluation and observability systems.
We are looking for talented individuals to join our team in 2027. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Launch your career where inspiration is infinite at ByteDance.
Successful candidates must be able to commit to an onboarding date by end of year 2027. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to ByteDance and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities
-For training track, develop the Volcano Ark training platform, enabling both internal and external users to perform serverless post-training (e.g., Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL)) on the Ark platform.
-Design elastic training solutions for complex multi-tenant workloads, supporting mixed-tenant training across multiple data centers and heterogeneous hardware while optimizing training throughput, resource utilization, and system stability.
-Build next-generation reinforcement learning infrastructure to improve training efficiency, while designing intuitive and developer-friendly APIs for RL training workflows.
-For inference track, develop the Volcano Ark MaaS (Model-as-a-Service) inference platform, optimizing the performance, cost efficiency, and reliability of large language model inference across hyperscale heterogeneous GPU clusters.
Reduce inference costs through advanced system optimizations, including disaggregated inference architectures, distributed KV cache systems, heterogeneous inference, elastic compute scheduling, and multi-tenant co-located inference.
