Job Title
Senior Platform SRE
IG is a FTSE 100 fintech operating across five continents, serving over 1.3m customers and handling billions of dollars in transactions – built on scale, trust, and proof. We didn't pivot to innovation; it's how we've always operated What that means for the people who work here is real: genuinely complex problems to solve, the technology and resources to tackle them properly, and the kind of scope that's rare in established businesses.
The bar is high – bring a curious and forward-thinking mindset and we'll give you the platform to define what comes next. Join us at IG – the future gets built here.
The Platform SRE team is the engine of IG’s reliability programme. We sit within Infrastructure & Operations, working across IG’s hybrid estate of on-premisesHashiCorpNomad and AWS.
We are not a reactive ops team. We build the platform, standards, and tooling that make reliability the default for every engineering team at IG. Through the SRE Guild, we connect with Domain SREs and Reliability Champions across the organisation, setting the bar and lifting it together.
You will be a hands-on technical contributor at the heart of the Platform SRE team, owning pieces of the reliability platform that hundreds of engineers depend on. You will work at the intersection of software engineering, observability, and systems reliability, turning reliability from a reactive concern into a proactive engineering discipline.
You will partner with Platform Engineering, product teams, and Reliability Champions to define what good looks like in production and then make it the default. You will contribute to the SRE Guild, mentor engineers across the organisation, and when things go wrong, you will be on the call helping to mitigate, understand, and prevent a repeat.
Build and own the reliability platform
Implement comprehensive monitoring and observability usingOpenTelemetryand distributed tracingMaintain SLO, error budgetsand burn-rate tracking
Establish andmaintain24/7 operational readiness including automated deployments, blue/green releases, and zero-downtime patching strategies
Engineer self-healing capabilities: auto-remediation, error-budget-gated rollback, and automated traffic rerouting
Design and run chaos experiments across the AWS estate, turning severe-but-plausible failure scenarios into engineering improvements
Build automation tools and CI/CD pipelines that embed reliability practices, whileapplyingsoftware engineering discipline including version control, code reviews, and testing
Contribute to the SRE AI agent, IG’s agentic tooling for incident investigation and reliability review, built on AWS frontier models
Mentor junior SREs and Reliability Champions on reliability patterns and production engineering discipline
Set and uphold standards
Author and evolve the SRE standards that underpin the Guild: SLOmethodology, error budget policy, observability instrumentation guide, andProduction Readiness Review(PRR)checklist
Mentor developers on reliability patterns including circuit breakers, retry logic, and fault tolerance
Work with development teams and Reliability Champions to design SLOs on customer journeys rather than per-service
Assistand guide teamsin system design, capacity planning,architecturalreviewsandclosingobservability gaps
Own incident response and learning
Facilitate blameless post-incident reviews (PIRs) within five working days using contributing-factormethodology
Maintain the Lessons Register, track remediation actions to closure, and surface patterns across incidents quarterly
What you'll need for this role:
Extensive working experience in all the below mentioned areas.
Essential Technical Skills
Observability and instrumentation:hands-onOpenTelemetryexperience (spans, metrics, traces, context propagation) and production use of Honeycomb, Datadog, Dynatrace, or Grafana; able to instrument Java or Python services directly
SLOs and error budgets:proventrack recorddesigning customer-meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and working with development teams on reliability measurement
CI/CD and release engineering:experience building pipelines with safety mechanisms: blue/green and canary releases, automated rollback, and DORA metrics integration
Container orchestration:Kubernetes (EKS, AKS, or GKE)required;HashiCorpNomad is a strong advantage on IG’s hybrid estate; solid understanding of cloud networking andIaC(Terraform preferred)
Software engineering:production-quality coding in Java and/or Python; comfortable contributing to application codebases to implement reliability patterns, not just configuring infrastructure around them
Distributed systems:strong understanding of how large-scale systems fail and how to make them fail safely; circuit breakers, bulkheads, idempotency, graceful degradation, and load-shedding; high-throughput, low-latency environments preferred
Incident management:on-call experience on production systems, blameless PIR facilitation, contributing-factor analysis, and driving action items to closure; PagerDuty and ServiceNow familiarity helpful
Chaos engineering:experience designing and executing hypothesis-driven experiments with blast-radius controls and gap-to-impact-tolerance analysis; AWS FIS, Gremlin, or equivalent
Community and standards:at ease in a guild or community-of-practice model; comfortable writing RFCs, presenting at engineering forums, and building standards that others will adopt
Experience Requirements
Track recordin high-throughput, production environments (financial services, trading platforms, or similar mission-critical systems preferred)
Demonstrated ability to improve system reliability and performance at scale
Experience working collaboratively with development teams to implement observability and reliability improvements
Strong troubleshooting skills in distributed systems environments
Core Competencies
Systems thinking approach to problem-solving
Excellent communication skills for cross-functional collaboration and technical enablement
Ability to balance hands-on development work with operational responsibilities
Strong bias toward automation andeliminatingmanual toil
We try to take a thoughtful approach to our ways of working as a company. We follow a hybrid working model with 3 days in the office -- which we think balances the need to collaborate effectively and connect with each other. When it comes to how we deliver, there are 5 things we want everyone to do to drive high performance, better learning and career satisfaction:
We believe that diversity is vital to success, it fuels creativity, drives innovation and sets us up for global success. We're committed to building teams with a variety of perspectives and skills to help us realise our vision and strategy, that's why we encourage applications from people with diverse backgrounds and experiences to join us on this journey. Learn more about our D&I approach here
Your growth fuels our success! Thrive with tailored development programs, mentoring opportunities with leaders, and clear career progression. Expand your network through committees, sports and social clubs. Enjoy extra time off for volunteering and community work.
Competitive salary
Flexible Benefits Package on top of your salary (12%)
Private medical cover for you and your family
Life insurance
Contribution to gym memberships
25 Days holiday, with 1 additional day off to celebrate your Birthday & 2 additional days off a year for voluntary work (28 in total
The option to buy or sell holiday days.
Unlimited access to the LinkedIn Learning Platform
A comprehensive global and local onboarding process
Employee-led LGBTQ+, Women’s, Black and Parents & Carers networks with an annual budget for organising events & projects that foster an open, diverse and inclusive culture
Enhanced primary (maternity), secondary (paternity), and shared parental pay and leave, as well as a range of support and benefits for parents
Option to participate and create ESG initiatives based on IG Brighter Future Fund
Learn more about the Perks here!
Join us for this exciting journey. Apply now!
Number of openings
1

We’ve been at the forefront of trading innovation since 1974, taking on the challenge to deliver an unmatched experience for our clients and raise the bar for tomorrow’s opportunities.
Today, we’re a global fintech company incorporating the IG, tasty, IG Prime, Spectrum and DailyFX brands, with a presence in 18 countries across five continents – Europe, North America, Africa, Asia-Pacific and the Middle East.
We’re an organisation of positive problem-solvers, united and inspired by our purpose, which is to power the pursuit of financial freedom for the ambitious. Our award-winning products and platforms empower go-getters the world over to unlock opportunities around the clock, giving them access to over 19,000 financial markets.
Today, more than 400,000 clients call IG Group home.
IG Group Holdings plc is an established member of the FTSE 250 and holds a long-term investment grade credit rating of BBB- with a stable outlook from Fitch Ratings.