Site Reliability Engineer / Production Support
London, United KingdomPosted Aug 5, 2026
Site Reliability Engineer / Production Support
hackajob London, England, United KingdomSite Reliability Engineer / Production Support
hackajob
London, England, United Kingdom
1 week ago
Be among the first 25 applicants
See who hackajob has hired for this role
Save
hackajob is collaborating with Monument to connect them with exceptional professionals for this role.Site Reliability Engineer/ Production Support
Location London (Oxford Circus) | Hybrid: 2 days per week | Reports to Head of Cloud Operations
About Monument
We're building something genuinely rare: a financial brand designed for the mass affluent, the professionals, entrepreneurs and ambitious savers that traditional banks have systematically underserved for decades.
We exist to make managing wealth simpler, smarter and more human, treating every client's wealth with the same care as if it were our own.
We hold over £7 billion in client savings, serve more than 100,000 clients, and were named the UK's fastest growing fintech in 2025. The momentum is real.
THE OPPORTUNITY
Monument’s production environment is the heartbeat of a licensed bank, and the SRE role is the single point of ownership when incidents occur. You will directly oversee the offshore Production Support team, run on-call and incident response, and ensure fast detection, triage and restoration of services.
This is not a passive monitoring role. You are expected to understand at a working level all of Monument’s key system flows from the services, partners and teams in play and to actively debug incidents, escalate effectively and drive permanent fixes. You will also be a builder: using AI tools for automated alert correlation, root cause analysis and runbook generation.
For the right person, this is a rare opportunity to own production reliability at a pre-IPO challenger bank, operating at the intersection of deep engineering and real commercial consequence.
What You'll Do
- Directly oversee the offshore Production Support team and be the single point person when incidents occur, escalating only to Head Of when required.
- Run on-call and incident response; ensure fast detection, triage, and restoration.
- Maintain observability standards (logs, metrics, traces) and alert quality (low noise, high signal).
- Understand at a working level all key system flows, the services, partners, and teams in play, and how to actively debug an incident.
- Lead reliability engineering: resilience patterns, performance tuning, capacity planning.
- Facilitate post-incident reviews and track actions to completion.
- Use AI tools for automated alert correlation, root cause analysis, and runbook generation.
- Hunt for routine/common tasks and formulate plans on how to automate and then execute them.
You’ll thrive here if you live by the same principles that define all Monument builders:
- Ownership under pressure - when things go wrong, you are calm, decisive, and effective. You own the incident until it's resolved.
- Builder - you don't just respond to incidents; you build the automation that prevents them or resolves them faster next time.
- Deeply curious - you understand the full system landscape and how services interact. You can debug across layers.
- Automation-first - every manual task is a candidate for automation. You actively hunt for toil and eliminate it.
- Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source.
- Strong SRE or production support experience with accountability for incident response in a production environment.
- Deep understanding of observability tools, alerting, logging, and distributed systems debugging.
- Experience managing and working with offshore support teams.
- Hands-on experience with reliability engineering: resilience patterns, performance tuning, capacity planning.
- Active use of AI tools for incident triage, automation, and runbook generation.
- Ability to understand complex system flows across multiple services and third-party integrations.
- Experience in financial services or similarly regulated environments is a strong advantage.
Be the person who keeps Monument running - your work directly protects clients and the business. Build automation that genuinely matters: every runbook you automate and every alert you tune makes the system more resilient. Work with modern AI tools as a core part of your daily workflow, not as a novelty. Own production reliability at a critical stage of Monument’s growth, with real responsibility and real impact.
OUR VALUES
At Monument, our values shape how we make decisions, how we treat each other when things get hard, and how we show up for clients who expect more than standard banking.
We set ambitious goals and hold ourselves to them, not because it looks good, but because our clients' outcomes depend on it. When something isn't working, we say so early, learn from it, and move. We don't wait for perfect conditions, and we don't protect egos over progress.
We work as a genuine team, which means real collaboration, honest conversations when we disagree, and shared accountability when things go wrong. We know better decisions come from different perspectives, so we actively value the range of experiences and backgrounds our people bring.
We're always asking whether there's a smarter way to do what we do, not for the sake of change, but because standing still isn't an option in the market we're in.
If that sounds like how you like to work, we'd like to hear from you.
-
Seniority level
Mid-Senior level -
Employment type
Full-time -
Job function
Engineering and Information Technology -
Industries
IT Services and IT Consulting
Referrals increase your chances of interviewing at hackajob by 2x
See who you know Get notified when a new job is posted.Similar jobs
-
Site Reliability Engineer
Site Reliability Engineer
Alpaca
London, England, United Kingdom 2 months ago -
Production Support Engineer
Production Support Engineer
intro
London Area, United Kingdom 1 day ago -
Site Reliability Engineer
Site Reliability Engineer
Ensono
London, England, United Kingdom 2 weeks ago -
Production Support Engineer
Production Support Engineer
Stanford Black Limited
London Area, United Kingdom 3 weeks ago -
Production Engineer
Production Engineer
beyond-tabs.com
London, England, United Kingdom 1 day ago -
Site Reliability Engineer
Site Reliability Engineer
Allwyn UK
Watford, England, United Kingdom 1 week ago -
Production Support Engineer
Production Support Engineer
Selby Jennings
London, England, United Kingdom 1 week ago -
Cloud Support Engineer
Cloud Support Engineer
Thought Machine
London, England, United Kingdom 1 week ago -
Production Support Engineer | Trading Technology
Production Support Engineer | Trading Technology
Selby Jennings
London Area, United Kingdom 1 week ago -
Platform Engineer
Platform Engineer
END.
Washington, England, United Kingdom 1 month ago -
Production Support Engineer, Senior Associate
Production Support Engineer, Senior Associate
State Street
London, England, United Kingdom 2 days ago -
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)
Signify Technology
London, England, United Kingdom 6 days ago -
Site Reliability Engineer
Site Reliability Engineer
Thought Machine
Greater London, England, United Kingdom 5 days ago -
Engineer - Site Reliability
Engineer - Site Reliability
Cboe Global Markets
London, England, United Kingdom 2 weeks ago -
DevOps Support Engineer
DevOps Support Engineer
TEKsystems
London, England, United Kingdom £450.00 - £650.00 2 days ago -
Site Reliability Engineer
Site Reliability Engineer
Trainline
London, England, United Kingdom 2 weeks ago -
Systems Engineer
Systems Engineer
mthree
London Area, United Kingdom 1 week ago -
Operations Engineer
Operations Engineer
GradBay
Wallingford, England, United Kingdom 3 days ago -
SaaS Operations Engineer
SaaS Operations Engineer
Bottomline
Theale, England, United Kingdom 2 weeks ago -
Site Reliability Engineer
Site Reliability Engineer
Joybuy
London Area, United Kingdom 22 hours ago -
Site Reliability Engineer
Site Reliability Engineer
Mphasis
London Area, United Kingdom 1 day ago -
Production Engineer
Production Engineer
Liquidnet
London, England, United Kingdom 1 week ago -
Site Reliability Engineer
Site Reliability Engineer
Intapp
London, England, United Kingdom 1 week ago -
Production Support - Senior Engineer
Production Support - Senior Engineer
Hyperlayer
London, England, United Kingdom 3 weeks ago -
Reliability Engineer
Reliability Engineer
Two Sigma
London, England, United Kingdom 3 days ago -
SRE, London
SRE, London
Apple
London, England, United Kingdom 1 month ago -
Production Engineer - Database Operations
Production Engineer - Database Operations
Palantir Technologies
London, England, United Kingdom 2 months ago
People also viewed
-
Observability Engineer
Observability Engineer
London, England, United Kingdom 6 days ago -
Trading Support Engineer
Trading Support Engineer
London Area, United Kingdom £100,000.00 - £150,000.00 3 weeks ago -
Observability Engineer
Observability Engineer
London, England, United Kingdom 6 days ago -
Site Reliability Engineer
Site Reliability Engineer
City Of London, England, United Kingdom 2 weeks ago -
DevOps Engineer
DevOps Engineer
London Area, United Kingdom £80,000.00 - £95,000.00 1 day ago -
NOC Engineer (AWS)
NOC Engineer (AWS)
Basingstoke, England, United Kingdom £60,000.00 - £60,000.00 1 week ago -
Systems Engineer (AWS)
Systems Engineer (AWS)
London, England, United Kingdom 6 days ago -
Site Reliability Engineer
Site Reliability Engineer
London, England, United Kingdom 1 week ago -
Infrastructure Engineer
Infrastructure Engineer
Greater London, England, United Kingdom 2 days ago -
Site Reliability Engineer
Site Reliability Engineer
London Area, United Kingdom £400,000.00 - £600,000.00 16 hours ago
Similar Searches
-
Site Reliability Engineer jobs
14,094 open jobs -
Platform Engineer jobs
7,837 open jobs -
Linux System Engineer jobs
2,599 open jobs -
Senior Infrastructure Engineer jobs
3,526 open jobs -
Build And Release Engineer jobs
2,631 open jobs -
Linux Engineer jobs
3,388 open jobs -
Production Support Engineer jobs
21,767 open jobs -
System Development Engineer jobs
22,005 open jobs -
Principal Infrastructure Engineer jobs
861 open jobs -
Network Project Engineer jobs
1,039 open jobs -
Linux System Administrator jobs
659 open jobs -
Infrastructure Engineer jobs
8,209 open jobs -
Unix Analyst jobs
27 open jobs -
Production Support Manager jobs
1,274 open jobs -
Reliability Engineer jobs
8,873 open jobs -
Domain Administrator jobs
1,393 open jobs -
Linux Administrator jobs
575 open jobs -
Application Support Engineer jobs
1,854 open jobs -
Senior Reliability Engineer jobs
347 open jobs -
Senior Linux Administrator jobs
396 open jobs -
Middleware Administrator jobs
174 open jobs -
Cloud Architect jobs
5,755 open jobs -
Performance Architect jobs
3,385 open jobs -
System Operations Engineer jobs
21,886 open jobs -
Engineer jobs
79,088 open jobs