Bellevue, WAFull TimeSeniorPosted Today
Meta is seeking a principal-level Software Engineer to drive technical strategy and execution across our Systems ML Engineering organization. In this role, you will define the architectural foundations that power large-scale machine learning infrastructure, spanning training systems, inference pipelines, ML compilers, high-performance computing frameworks, and on-device optimization. You will identify and solve the hardest cross-system ML infrastructure challenges, shape multi-year technical roadmaps, and amplify the impact of engineering teams through AI-native workflows and deep systems expertise. This is a role for engineers who identify problems others miss and drive them to resolution at organizational scale.
Responsibilities
Identify and solve the most complex cross-system ML infrastructure challenges spanning training, inference, compiler optimization, and hardware-software co-design, including problems that have resisted prior solution attempts
* Define extensible architectural standards and technical foundations for ML systems that enable consistency and reliability across multiple engineering organizations
* Develop and own the multi-year technical roadmap for ML systems infrastructure, balancing short-term delivery with long-term platform health and competitive positioning
* Leverage AI-native tooling and workflows as a force multiplier to eliminate entire categories of engineering toil and accelerate cross-disciplinary work across the ML systems stack
* Drive performance improvements across large-scale ML training and inference systems by identifying bottlenecks that span multiple subsystems, ownership boundaries, and abstraction layers
* Establish invariants, correctness proofs, and systemic reliability practices that prevent whole classes of failures across ML infrastructure pipelines
* Partner with research, hardware, and product engineering teams to translate theoretical ML systems advances into production infrastructure that delivers measurable efficiency and capability gains
* Assess emerging AI and computing technologies, evaluate competitive ML infrastructure trends, and influence organizational strategy to ensure technical competitiveness
* Mentor engineers across the organization by providing customized coaching, leading engineering programs, and establishing a culture of thoroughness and high craft in ML systems development
* Communicate complex ML systems architecture and strategy clearly to technical and non-technical audiences, producing reference-quality design documents and roadmap artifacts
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
* 12+ years of experience in software engineering with deep specialization in one or more ML systems domains including AI infrastructure, ML compilers, high-performance computing, GPU architecture, ML frameworks, or on-device optimization
* Experience architecting and delivering large-scale ML training or inference infrastructure that has had measurable impact across multiple engineering organizations
* Experience leading multi-year cross-functional technical initiatives, including defining metrics, managing dependencies, and driving execution across organizational boundaries
* Experience developing high-performance ML systems infrastructure in C++, Python, or CUDA, including work at the intersection of hardware and software
* Experience influencing technical direction and engineering practices across multiple teams through written proposals, design reviews, and stakeholder alignment Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
* Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
* Experience contributing to industry-wide ML systems efforts through publications, open-source projects, or standards bodies
* Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
* Track record of applying AI tools and automation to redesign engineering workflows, with demonstrated efficiency or quality improvements at organizational scale
* Experience with ML compiler stacks such as MLIR, XLA, or TVM, or with hardware-software co-design for custom ML accelerators
* Experience defining and operationalizing reliability, performance, and correctness standards for distributed ML training or large-scale inference systems
Responsibilities
Identify and solve the most complex cross-system ML infrastructure challenges spanning training, inference, compiler optimization, and hardware-software co-design, including problems that have resisted prior solution attempts
* Define extensible architectural standards and technical foundations for ML systems that enable consistency and reliability across multiple engineering organizations
* Develop and own the multi-year technical roadmap for ML systems infrastructure, balancing short-term delivery with long-term platform health and competitive positioning
* Leverage AI-native tooling and workflows as a force multiplier to eliminate entire categories of engineering toil and accelerate cross-disciplinary work across the ML systems stack
* Drive performance improvements across large-scale ML training and inference systems by identifying bottlenecks that span multiple subsystems, ownership boundaries, and abstraction layers
* Establish invariants, correctness proofs, and systemic reliability practices that prevent whole classes of failures across ML infrastructure pipelines
* Partner with research, hardware, and product engineering teams to translate theoretical ML systems advances into production infrastructure that delivers measurable efficiency and capability gains
* Assess emerging AI and computing technologies, evaluate competitive ML infrastructure trends, and influence organizational strategy to ensure technical competitiveness
* Mentor engineers across the organization by providing customized coaching, leading engineering programs, and establishing a culture of thoroughness and high craft in ML systems development
* Communicate complex ML systems architecture and strategy clearly to technical and non-technical audiences, producing reference-quality design documents and roadmap artifacts
Qualifications
Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
* 12+ years of experience in software engineering with deep specialization in one or more ML systems domains including AI infrastructure, ML compilers, high-performance computing, GPU architecture, ML frameworks, or on-device optimization
* Experience architecting and delivering large-scale ML training or inference infrastructure that has had measurable impact across multiple engineering organizations
* Experience leading multi-year cross-functional technical initiatives, including defining metrics, managing dependencies, and driving execution across organizational boundaries
* Experience developing high-performance ML systems infrastructure in C++, Python, or CUDA, including work at the intersection of hardware and software
* Experience influencing technical direction and engineering practices across multiple teams through written proposals, design reviews, and stakeholder alignment Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
* Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
* Experience contributing to industry-wide ML systems efforts through publications, open-source projects, or standards bodies
* Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
* Track record of applying AI tools and automation to redesign engineering workflows, with demonstrated efficiency or quality improvements at organizational scale
* Experience with ML compiler stacks such as MLIR, XLA, or TVM, or with hardware-software co-design for custom ML accelerators
* Experience defining and operationalizing reliability, performance, and correctness standards for distributed ML training or large-scale inference systems
