LondonFull TimeSeniorPosted Today
Summary
At Apple, we believe that innovation flourishes in an environment where ideas are challenged, collaboration is encouraged, and technology is pushed to its limits. This environment is only possible when diverse minds come together, bringing unique perspectives and experiences. Our people and their ideas inspire innovation in everything we do. Imagine what you could accomplish here! Join Apple and help us make the world a better place.As an SRE on our team, you'll own the reliability, performance, and scale of the distributed storage and data platform systems that power Apple's services. You'll debug replication and consensus failures, tune systems at petabyte scale, and write production code that operates the platform. We firmly believe in ownership, with software engineers accountable for the code they write.
Description
The Apple Services Engineering (ASE) organisation builds and provides systems and infrastructure that fuel Apple’s services — iCloud, iTunes, Siri, and Maps. Our team builds and operates the data platform infrastructure behind them, keeping petabyte-scale workloads fast, resilient, and reliable.The platform runs on large-scale distributed systems, including object stores, databases, and data pipelines, on Linux across private and hybrid cloud. You'll work on storage engines, distributed consensus, and data-flow internals, partnering with development teams on system-wide architecture rather than individual components.
Preferred Qualifications
Contributions to distributed-systems internals, open-source data infrastructure, or storage/database engines.Experience defining SLIs/SLOs, building observability, and using error budgets to drive reliability decisions.
A good grasp of Unix internals and networking fundamentals.
Experience with data migration, disaster recovery, or capacity planning at scale.
Minimum Qualifications
Experience in managing and scaling large-scale distributed systems in a private or hybrid cloud environment.Comfortable designing, writing, and releasing production code in languages such as Go or Python.
Able to debug and reason about how distributed systems fail and perform at scale.
Willingness to take part in on-call rotations and incident response to keep critical systems healthy.
