The Director of Cloud Engineering is a senior engineering leader accountable for the strategy, architecture, and operational excellence of the company's multi-cloud platform. Reporting to the Sr. Director of IT Operations, this role owns the end-to-end cloud estate across Google Cloud Platform (GCP) and Microsoft Azure, setting the technical direction for cloud-native infrastructure, platform engineering, and developer experience.
You will build and lead a high-performing distributed team of cloud engineers, SREs, and platform engineers while partnering closely with product, security, data, and application engineering to deliver a reliable, secure, and cost-efficient cloud foundation. This is a hands-on leadership role that balances deep technical credibility with the strategic thinking and organizational influence expected at the Director level.
Key Responsibilities:
Cloud Platform Strategy & Architecture
• Define and own the multi-cloud architecture strategy across GCP and Azure, aligning platform capabilities with business objectives, scalability requirements, and total cost of ownership targets.
• Lead architectural decisions for cloud-native workloads, including microservices, containerization (Kubernetes/GKE/AKS), serverless functions, and event-driven patterns.
• Drive the evolution of the Internal Developer Platform (IDP), enabling self-service infrastructure provisioning, golden-path templates, and standardized deployment pipelines.
• Establish and enforce cloud governance frameworks, including resource hierarchy, landing zones, policy-as-code, and guardrails across both cloud providers.
• Evaluate emerging technologies—AI/ML infrastructure, LLMOps tooling, edge compute—and make build-vs-buy recommendations grounded in cost, risk, and strategic fit.
Infrastructure Engineering & Operations
• Own Infrastructure as Code (IaC) standards and practices; ensure all infrastructure changes flow through version-controlled, peer-reviewed pipelines.
• Oversee GitOps workflows and CI/CD pipeline infrastructure, partnering with application teams to reduce deployment lead time and increase release frequency.
• Direct SRE practices across the cloud estate: define and track SLIs/SLOs, drive blameless post-mortems, and lead reliability improvements through systematic error-budget management.
• Own observability strategy—logging, metrics, tracing, and alerting—across cloud environments, ensuring engineers have actionable insight into system health and performance.
• Lead disaster recovery design, runbook development, and regular DR/BCP testing exercises; drive remediation of identified gaps to closure.
• Administer and maintain all Microsoft licensing through NCE and MPSA agreements
• Maintain network architecture and connectivity across cloud environments, data centers, and supported locations, including VPN, private connectivity (Interconnect/ExpressRoute), and DNS.
Security, Compliance & Risk
• Partner with the Security team to implement Zero Trust network architecture, enforce least-privilege IAM, and operationalize Cloud Security Posture Management (CSPM) tooling.
• Ensure cloud environments meet regulatory and compliance requirements; support internal and external audits by providing architecture documentation, access logs, and control evidence.
• Define and maintain cloud security baselines, secrets management practices, and data encryption standards across the shared responsibility model.
• Drive incident response processes for cloud infrastructure events, ensuring timely escalation, containment, and root-cause resolution.
FinOps & Cloud Cost Management
• Own the cloud infrastructure budget; implement FinOps practices including cost allocation tagging, chargeback/showback reporting, reserved capacity planning, and rightsizing recommendations.
• Establish cloud cost visibility and governance mechanisms so engineering teams can make cost-aware architecture decisions in real time.
• Negotiate vendor agreements and manage strategic relationships with GCP, Azure, and key third-party tooling providers.
AI-Enabled Cloud Engineering
Evaluate, integrate, and champion the responsible use of AI tools across cloud engineering and platform teams, setting standards for AI-driven productivity and innovation.
Treat AI as a productivity multiplier, not a replacement — ensuring engineers maintain sound engineering judgment and that AI-generated code and configuration meet the same quality, security, and reliability standards as hand-written work.
Leadership & People Management
• Lead, mentor, and grow a distributed team of cloud engineers, platform engineers, and SREs (onshore and offshore); establish clear career ladders and development paths.
• Build a team culture grounded in psychological safety, continuous learning, and engineering excellence; champion internal tech talks, documentation habits, and knowledge sharing.
• Manage headcount planning, recruiting, and onboarding; partner with HR and technical leads to define roles and evaluate candidates.
• Deliver timely, constructive performance reviews; proactively manage performance issues and recognize high-impact contributions.
• Serve as a technical escalation path and executive sponsor for major infrastructure initiatives; represent the cloud engineering team in IT governance and steering forums.
QualificationsExperience
• 15+ years in IT/infrastructure engineering, with at least 8 years focused on public cloud platforms (GCP and/or Azure).
• 5+ years in a senior engineering leadership role (Director, Principal, or Staff Engineer equivalent) managing teams of 8 or more engineers across onshore and offshore locations.
• Demonstrated hands-on proficiency with both GCP and Azure services, including but not limited to: GKE, Cloud Run, Cloud Armor, BigQuery, VPC Service Controls (GCP) and AKS, Azure Policy, Defender for Cloud, ExpressRoute, and Entra ID (Azure).
• Proven track record delivering large-scale cloud migrations, re-architecture projects, or greenfield platform builds in a regulated or enterprise environment.
• Deep expertise with Infrastructure as Code (Terraform required; Pulumi or CDK a plus) and GitOps tooling (ArgoCD, Flux, or equivalent).
• Experience designing and operating CI/CD pipelines at enterprise scale (GitHub Actions, Cloud Build, Azure DevOps, or equivalent).
• Direct experience with FinOps practices: cloud cost allocation, tagging governance, reserved instance/committed-use optimization, and budget reporting.
• Familiarity with SRE principles: SLI/SLO frameworks, error budgets, chaos engineering, and reliability-focused incident management.
Education
• Bachelor’s degree in Computer Science, Information Systems, or a closely related engineering discipline, or equivalent professional experience.
Certifications
• At least one of: Google Cloud Professional Cloud Architect, Google Cloud Professional DevOps Engineer, Microsoft Azure Solutions Architect Expert (AZ-305), or Microsoft Azure DevOps Engineer Expert (AZ-400).
Preferred Qualifications
• Experience building or maturing an Internal Developer Platform (IDP) using tools such as Backstage, Crossplane, or similar platform-engineering frameworks.
• Exposure to AI/ML infrastructure on GCP (Vertex AI) or Azure (Azure AI Studio, Azure Machine Learning) and an understanding of LLMOps patterns such as model serving, inference optimization, and vector database infrastructure.
• Working knowledge of service mesh technologies (Istio, Linkerd) and zero-trust network segmentation at the workload level.
• Experience with cloud-native observability stacks (Prometheus/Grafana, Datadog, Dynatrace, or Google Cloud Operations Suite).
• Familiarity with enterprise IT governance frameworks such as ITIL v4, COBIT 2019, or TOGAF; direct experience supporting SOC 2, ISO 27001, or HIPAA/PCI-DSS compliance programs.
• Experience in financial services, healthcare, or another regulated industry where cloud security posture and audit readiness are operationally critical.
• Additional certifications: Certified Kubernetes Administrator (CKA), HashiCorp Terraform Associate, Google Cloud Professional Security Engineer, or AWS Solutions Architect (for multi-cloud context).
