Staff Site Reliability Engineer
Obsidian Security, City Centre, Manchester
Staff Site Reliability Engineer
Salary not available. View on company website.
Obsidian Security, City Centre, Manchester
- Full time
- Permanent
- Onsite working
Posted today, 31 Aug | Get your application in now to be one of the first to apply.
Closing date: Closing date not specified
Job ref: 2166e2cf31144ea98f1d6af49b3ae4c0
Location ref: City Centre, Manchester
Full Job Description
Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted-consistently and predictably. This is a hands-on technical role that involves architecting and leading the implementation of systems that handle real-world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission-critical enterprise workloads.,
- Reliability Strategy & Architecture - Define and lead long-term reliability strategy across services. Establish end-to-end system visibility frameworks and guide architecture for observability, detection, and resilience.
- Cross-Org Leadership - Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert.
- Detection & Observability - Build intelligent detection systems (anomaly detection, connector health models) and enable self-service observability.
- Incident Management - Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust.
- Execution - Contribute hands-on to system design, monitoring, and debugging across distributed systems and data pipelines.
5+ years in SRE, Production Engineering, or related roles - 3+ years operating at a senior or technical leadership level (Staff or equivalent scope)
- Deep expertise in:
- AWS and/or GCP
- Kubernetes and Helm
- Observability stacks (Prometheus, Grafana, or equivalent)
- CI/CD systems (GitLab CI/CD, ArgoCD, etc.) Proven experience designing and scaling reliability systems for multi-tenant SaaS platforms Strong debugging and systems thinking across distributed microservices and legacy systems Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience Hands-on engineering approach with a track record of building-not just configuring-reliability systems,
- Experience in B2B SaaS serving enterprise or financial customers
- Familiarity with third-party SaaS connector architectures and ingestion patterns
- Experience building anomaly detection or intelligent alerting systems
- Experience designing customer-facing status pages and incident communication frameworks
Our competitive benefits packages are designed to support our employees' well-being, both at work and at home. Our US-based employees enjoy competitive compensation with equity and 401k, comprehensive healthcare with dental and vision coverage, flexible paid time off and paid holiday time off, 12 weeks of new parent or family leave, and personal and professional development resources. For more details on our US benefits, or for information on our international benefits, please see here., Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company. Base Salary Range: £124,000 GBP - £141,000 GBP