Minimum qualifications:
- Bachelor's degree in Engineering or equivalent practical experience
- 15 years of experience in ML infrastructure planning or deployment - in NPI or production
Preferred qualifications:
- Master's degree in Engineering
- Solid understanding of infrastructure deployment processes, dependencies, constraints and acceleration approaches
About the job
Google’s AI & Infrastructure (AI2) team is responsible for the global infrastructure that powers Google. We do everything from designing and building gigawatts of data centers and petabits of networks around the world, to designing and manufacturing the servers, AI accelerators, and optical switches which comprise our fleet, to developing the infrastructure software which runs on that fleet and underpins all Google’s multi-billion-user services and Cloud platforms.
Our organization delivers sustainable and resource-efficient infrastructure that supports Google and our customers. We provide a foundation that operates safely, securely, and sustainably, and are chartered with ensuring Google’s data center landscape can support rapidly changing demand while delivering on Google’s commitment to sustainability.
We seek a leader that will be responsible for driving execution of current ML infrastructure deployments and working closely with all partner teams involved.
In this key leadership role, you will be a key lead responsible for managing the execution and orchestration of ML infrastructure deployment for Google Product Areas and external customers. In addition, you’ll be a key stakeholder in ensuring that our ML infrastructure planning is robust and comprehensive.Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.US: $281000 - $392000 (USD) + 30% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities
- Work with planning leads to develop executable and deployment plans for Google’s ML infrastructure. Ensure ML infrastructure plans make efficient use of power and cooling at a data center level.
- Drive execution of ML infrastructure deployments, ensuring timely delivery to meet Google's demand. Oversee execution escalations to ensure timely closure of issues and on-time delivery of capacity.
- Lead, mentor, and develop a high-performing team of execution leads for TPU/GPU platforms. Manage and allocate staff, budget, and resources effectively to achieve key objectives.
- Lead and influence cross-functional ML workstreams in a matrix environment to deliver outstanding acceleration outcomes.
- Work closely with leads in functional areas and data centers to drive scalable execution and acceleration at a multi-GW scale. Establish key performance indicators (KPIs) and metrics to monitor and improve the effectiveness of planning and execution.
