Site Reliability Engineer

BT Group · London Area, United Kingdom
LinkedIn

Posted

Aug 18, 2026 (11d ago)

Seniority

Lead

Work Model

Not Specified

Type

Not Specified

Category

DevOps & SRE

Salary

Not specified

Skills

Ansible AWS Azure CI/CD Datadog GCP Generative AI Grafana Helm Java LLM Microservices Observability Prometheus Python Slack SRE Terraform

Description

Why BT? BT has a key role in British society, fostering change and leading technology innovation. From delivering the Olympics, to supporting the emergency services, to investing more into research than any other UK technology company, we take pride in everything we do – and in the people who work here. We’re now a global company operating at the forefront of the information age, employing 90,000 people in 180 countries. And we’re on a mission. Guided by our core values of Personal, Simple and Brilliant our goal is to help customers, communities and businesses overcome barriers and release their potential. So, if you’re interested in the power of potential, why not join us today and release yours? Read more about what it’s like to work at BT. Why Business Unit/Function The Site Reliability Engineering Specialist independently executes activities that help ensures BT is in the best position to deliver the service performance, reliability and availability that internal and external customers expect, through enabling cross-team engineering discussions to achieve scalable, measurable, fault-tolerant, and cost-effective cloud services. Why this job matters The Site Reliability Engineering Specialist plays a critical role in safeguarding BT’s ability to deliver exceptional service performance, reliability, and availability across its digital platforms. In today’s fast-paced, cloud-driven AI environment, customers expect seamless experiences, and this position ensures those expectations are met by driving scalable, fault-tolerant, and cost-effective solutions. By enabling cross-team collaboration and implementing automation, monitoring, and resilience strategies, the specialist not only minimizes downtime and operational risk but also accelerates innovation and system evolution. This role is pivotal in maintaining BT’s reputation for reliability while empowering the business to adapt quickly to emerging technologies and deliver consistent value to customers worldwide. What I’ll be doing – your accountabilities Executes the implementation of new software development life cycle automation tools, frameworks, and code pipelines (continuous integration/continuous delivery pipelines whilst executing best practices with a focus on the re-use of application code, demonstrates consistent software delivery practices and produces continuous integration/continuous delivery platform solutions using Amazon Web Services cloud, infrastructure as code (IaC), GitOps, and container technologies Coordinates a diverse team and creates the initial test schedule to deliver all aspects of testing to time, budget and quality targets, ensuring producing outlines of solutions and defining depth of testing required Executes the implementation of automation technologies to ensure repeatability, eliminating toil, reducing mean time to detection and resolution and repair services Proactively identifies and manages risk through regular assessment and diligent execution of controls and mitigations, proactively raising any concerns Leads scale testing to measure, tune and optimise system performance Executes metric/monitoring analysis that creates stability, security, and performance improvements Designs, analyses, develops and troubleshoots highly distributed large-scale production systems spanning on-prem and cloud-based hosting Executes approaches that scale systems sustainably through mechanisms like automation and evolves systems by pushing for changes that improve reliability and velocity Writes and delivers infrastructure as code software to improve the availability, scalability, latency, and efficiency of services Implements robust monitoring and alerting systems and performs root cause analysis and post-mortems with an eye towards future prevention Inspects queue and support processing to ensure early warning of support issues Executes retrospective and preventive actions after each high severity production incident Analyses complex systems from a reliability and resilience perspective and identifies sources of instability in distributed systems Champions, continuously develops and shares with team knowledge on emerging trends and changes in site reliability engineering best practices and industry standards Mentors other site reliability engineers, helping to improve the team’s abilities by acting as a technical resource Uses the network of site reliability engineers, removing BTs organisational boundaries to deliver improvements that are in synergy with initiatives being driven by other SREs. Skills required for the job A degree in IT, Maths or Science A deep understanding of full stack monitoring solutions such as Dynatrace to ensure current end to end performance and trends of owned CDO Applications Strong proficiency in one or more programming languages (e.g. Java, Python). Experience with cloud platforms (AWS, Azure, or GCP). Solid understanding of software architecture, design patterns, and microservices. Familiarity with CI/CD tools and DevOps practices. High levels of quality presentation and reporting capabilities to collate output from Managed Service Partners. Resilience to ensure support teams are engaged 24x7x365 to support priority incident resolution. Ability to adapt to latest industry trends CI/CD/CT Pipeline management Micro-Service functionality Business Process Improvement Growth mindset AI driven‑ Observability & AIOps AIOps fundamentals (cross domain‑ telemetry ingestion, event correlation, topology/context building, and remediation augmentation). Agentic/autonomous observability skills (using intelligent agents to detect anomalies, correlate signals, and trigger guarded remediations to cut MTTR). AI assisted alerting & noise reduction (designing contextual, business‑ impact‑ ‑aware alerts; prioritization via ML). Connected leaders behaviours Solution Focussed Achiever: deliver ambitious goals, outcomes and timelines. Cutting through complexity and obstacles, working end to end with other SRE leads to get to the right solution at the right time. Inspiring Communicator: Understanding the current condition of the problem area that you are looking to improve so that you are readily able to articulate the benefit that YOU are delivering. Lead Collaboration: Understanding the agendas and needs of others, alongside the needs of the business. Breaking down silos, working brilliantly with partners both within and outside of the organisation to deliver business results Experience you would be expected to have Incident Response with AI LLM assisted‑ incident workflows (AI summaries, timeline drafting, suggested fixes, and post‑mortems integrated with Slack/Teams). Runbook automation with AI (building AI assisted, context‑ aware runbooks and approval gates for high‑ risk‑ actions). Generative AI for coordination & RCA (using LLMs to accelerate investigation and communications; understanding current accuracy limits and human in‑ the‑ ‑loop needs). ML Ops for Reliability SRE principles applied to ML systems (SLOs/SLIs/error budgets for ML services; capacity planning and model freshness). Production ML observability (data/concept/label drift detection, automated retraining triggers, explainability traces). Telemetry & visualization for model health (instrumentation with Prometheus/Grafana for drift and degradation). AI enhanced‑ Automation & CI/CD AI ‑augmented IaC and pipelines (LLM generated‑ Terraform/Helm/Ansible, policy enforcement, drift detection in infra). AIOps in delivery (change ‑impact hints, automated triage, and GitOps based auto‑ remediation‑). AI pair programming‑ ergonomics (using Copilot responsibly; measuring impact on quality/velocity and guardrails). AI + Chaos Engineering (Resilience) Designing AI guided chaos experiments (intelligent fault selection‑, anomaly detection during experiments, learning from outcomes). Reinforcement learning‑ ‑driven fault injection (automated scenario generation to expose latent weaknesses and improve recovery times). Operationalizing lessons from chaos + ML (predictive failure analysis and proactive controls). Platform & Tool Literacy (AI r‑eady) Hands‑on with AIOps/observability platforms (event correlation and unified incident views at scale). Familiarity with AI enabled‑ incident tooling (e.g., incident.io/Rootly/PagerDuty/Datadog for AI triage and summaries). Governance, Safety & Measurement Human in‑ the‑ ‑loop guardrails (approval policies, rollback safety, and compliance in autonomous actions). Trustworthy AI practices (explainability, data/model/process trust; aligning metrics with business outcomes). Outcome measurement for AI adoption (MTTR, alert noise, developer experience/velocity with AI tools).