Back to Jobs
BT Group

Site Reliability Engineering Specialist

BT Group
Ipswich, GB, IP5 3REFull TimePosted Today

Working Style: 3 days a week in office, 2 days from home

About the role

Professional Services was formed as a progressive development towards the convergence of multiple domains across BT. We pride ourselves on providing expert third line support to an extensive range of services; ensuring the required levels of availability are maintained. The team is widely recognised for getting things done, while making transformational improvements along the way. We do this by ensuring we have the right people to achieve our high ambitions.

The purpose of this role is to apply SRE principles to BT's DWDM optical transport infrastructure, delivering third line operational support, major technology introductions, and complex fault resolution across our DWDM platforms. You will take the lead on high impact incident resolution and drive automation of change activities to deliver flawless, repeatable deployments across the optical network. You will also build and own automated processes that reduce operational toil and improve the reliability of our DWDM services.

What you’ll be doing
  • Provides expert third line support for BT's DWDM platforms, acting as a final technical escalation point for complex optical network faults spanning multiple vendor equipment and service layers.
  • Leads on major incident resolution within the DWDM domain, coordinating across teams to restore service and minimise customer impact.
  • Builds and owns automated change processes for DWDM network activities, utilising CI/CD pipelines to deliver consistent, repeatable, and auditable deployments.
  • Leads blameless post-incident reviews to uncover systemic root causes and convert learnings into concrete reliability, automation, and process improvements.
  • Champions a reliability-first change culture, promoting safe deployment patterns, blameless learning, and continuous improvement across the DWDM engineering team.
  • Collaborates with design & platform teams to support the implementation of flawless change into the live optical network, including new DWDM circuit provisioning, capacity upgrades, and hardware refreshes.
  • Acts as a subject matter expert within the DWDM and optical transport domain, applying this expertise to troubleshoot faults across Nokia and Ciena platforms.
  • Configures and maintains monitoring and observability solutions for DWDM platforms, ensuring comprehensive visibility of optical network health, performance metrics, and alerting.
  • Embeds secure by design principles when building new change processes and solutions.
  • Will champion and build effective working relationships, both internally and externally to deliver business outcomes.
  • Champions the adoption of Site Reliability Engineering practices within Professional Services, driving cultural change towards automation, observability, and reduced operational toil on the optical network.
  • Be prepared to support and drive resolution of Critical Incidents out of hours as part of an on-call rota.
Essential Skills / Experience
  • Strong understanding of DWDM and optical transport technologies, including wavelength provisioning, optical amplification, multiplexing/demultiplexing, and coherent optics. Broader networking experience may be considered where direct DWDM expertise is limited.
  • Hands-on experience with at least one programming or scripting language, preferably Python or Bash, to support automation and operational efficiency.
  • Experience working within an SRE, Network Operations, or Infrastructure environment, with strong knowledge of monitoring and observability tools such as Grafana, Prometheus, or equivalent, and a solid understanding of UNIX/Linux systems for service health and performance management.
  • Practical experience with Nokia and/or Ciena optical platforms.
  • Confidence and professionalism in communicating with all stakeholders, both locally and with members of the Senior Management Team.
  • Resilience and adaptability in managing operational challenges and changing priorities.
Desirable Skills / Experience
  • Strong Python programming proficiency for building automation tooling.
  • A good understanding of Linux operating systems or similar
  • A strong understanding of containerisation using Docker, Podman or a similar container engine.
  • Experience in building CI/CD pipelines.
  • A good knowledge of coding best practices including code structure, peer review & testing.
  • Experience with optical network planning tools and capacity management.
  • A strong understanding of optical network management systems and their role in fault detection, performance monitoring, and circuit provisioning.
  • A strong understanding of change & incident management best practice within a live network environment.
Our Package

Tailored benefits make a real difference. That’s why we offer a comprehensive range to support your growth, wellbeing, and everyday life.
You can design the package to suit you and your lifestyle. Your core benefits include:

  • 10% on target annual bonus
  • Access to an online private GP 24/7 for you and your immediate family
  • Market-leading paid carers leave with up to 2 weeks off
  • Equalized maternity, paternity, and adoption leave – 18 weeks’ full pay and 8 weeks’ half pay
  • Discounted EE and BT products, including mobile and broadband
  • Market leading Pension scheme – 5% from you and 10% from us
  • Holiday purchase scheme

You can select additional benefits, including healthcare, dental, gym memberships and more when you’re ready.
Ready to connect for good and help shape the future? Apply now

BT Group is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet. Our purpose is as simple as it is ambitious: we connect for good.  Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers. 

BT Group’s role is about setting direction, unlocking value and creating the conditions for our brands and businesses to thrive.

Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably. Group teams shape strategy, policy, brand, capital allocation and transformation, helping the whole organisation perform at its best.

We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected. These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country.   Joining BT Group means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies.

Ready to apply? You'll be taken to BT Group's application page.
Site Reliability Engineering Specialist at BT Group