SRE Service Model Overview
The SRE Service Model defines how Managed Service SRE teams engage with delivery teams, transition services into operational ownership, and provide ongoing reliability management after go-live.
The model recognises that successful managed service delivery requires SRE involvement before a service reaches production. Reliability cannot be established through operational handover alone. SRE must work alongside delivery teams during design, build and testing phases to understand the service, influence reliability decisions and establish the operational capabilities required to support the service.
The SRE team becomes accountable for the operational reliability of the service after a structured transition period, including a formal handover and hypercare phase. This transition typically requires approximately three months of collaboration between the delivery team and the SRE organisation before production ownership transfers fully to SRE.
Purpose
The SRE Service Model provides a consistent approach for introducing new services into a managed SRE operating model.
It defines:
- How SRE engages during service delivery.
- The activities required before operational ownership transfer.
- The expectations for handover and hypercare.
- The responsibilities of delivery teams and SRE teams.
- How reliability is maintained after go-live.
- How continuous improvement is managed throughout the service lifecycle.
The goal is to ensure that services entering managed support are operationally ready, observable, supportable and aligned with SRE reliability standards.
Service Model Principles
SRE involvement starts during delivery
Managed Service SRE engagement must begin before production readiness activities.
Waiting until deployment or go-live to involve SRE creates operational risk because critical reliability requirements may not have been considered during design and implementation.
SRE engagement during delivery helps establish:
- Appropriate service architecture considerations.
- Observability requirements.
- Monitoring and alerting standards.
- Operational procedures.
- Incident response expectations.
- Capacity and performance requirements.
- Support ownership boundaries.
The delivery team remains responsible for building the service, but SRE provides operational input to make sure the service can be effectively supported once live.
Operational ownership requires a planned transition
A production deployment does not represent an immediate transfer of operational ownership.
A controlled transition period is required to allow SRE to:
- Understand the service behaviour.
- Validate operational processes.
- Confirm monitoring and alerting coverage.
- Learn service dependencies.
- Build confidence responding to operational events.
- Identify gaps before full ownership transfer.
The transition period typically includes a three-month handover and hypercare phase involving the delivery team and wider SRE organisation.
Reliability is a shared responsibility during transition
During delivery and transition, reliability responsibilities are shared between the delivery team and SRE.
Delivery team provides:
- Service knowledge.
- Application expertise.
- Architectural context.
- Support for defect resolution.
- Engineering changes required to meet operational standards.
SRE provides:
- Operational readiness assessment.
- Reliability standards.
- Observability validation.
- Operational process development.
- Monitoring and alerting implementation.
- Incident response preparation.
The responsibility gradually moves from shared ownership during transition to SRE operational ownership after acceptance.
SRE owns service reliability after handover
Once the service completes transition and enters steady-state operations, SRE becomes responsible for operational management.
This includes:
- Monitoring service health.
- Responding to operational alerts.
- Managing incidents.
- Maintaining operational documentation.
- Reporting reliability performance.
- Identifying reliability improvements.
- Managing operational risks.
Application teams continue to own application development, defects and product changes. SRE owns the reliability of the production service and coordinates with application teams when remediation is required.
Service Lifecycle
The managed SRE service lifecycle consists of five stages:
┌──────────────────────┐
│ Delivery Engagement │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Operational Readiness│
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Handover & Hypercare │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Managed Operations │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Continuous Improvement│
└──────────────────────┘
Delivery Engagement
SRE engagement begins during the delivery lifecycle rather than immediately before go-live.
The delivery team works with SRE to establish the operational requirements needed for successful support.
Key activities include:
- Reviewing service architecture and dependencies.
- Defining service ownership boundaries.
- Establishing monitoring and observability requirements.
- Reviewing logging, metrics and tracing requirements.
- Identifying operational procedures.
- Understanding expected service behaviour.
- Reviewing resilience and recovery expectations.
The earlier SRE is involved, the more opportunity exists to address reliability concerns before they become operational issues.
Operational Readiness
Before a service can transition into managed support, it must meet agreed operational readiness expectations.
Operational readiness activities include the following areas:
| Area | Expectations | Typical owner |
|---|---|---|
| Service ownership | Escalation paths and responsibilities are documented. | Service owner / delivery team |
| Observability | Metrics, logging and tracing provide sufficient visibility into service health. | Delivery team / SRE |
| Monitoring | Monitoring identifies customer-impacting issues and supports effective response. | SRE |
| Alerting | Alerts are actionable, owned and aligned with service impact. | SRE |
| Documentation | Operational documentation and support procedures are available. | Delivery team |
| Access | SRE has appropriate access to required platforms, tooling and environments. | Delivery team / Platform teams |
| Recovery | Backup, recovery and failure handling processes are understood and tested where required. | Delivery team / Platform teams |
| Dependencies | Service dependencies and operational impacts are documented. | Delivery team |
A service should not enter production support without sufficient operational visibility and support capability.
Handover and Hypercare
The handover and hypercare period provides a controlled transition between delivery ownership and SRE operational ownership.
This period typically lasts around three months, although the required duration may vary depending on service complexity, risk and maturity.
During hypercare:
Delivery teams provide
- Detailed service knowledge transfer.
- Support for early production issues.
- Assistance understanding unexpected behaviour.
- Resolution of outstanding defects.
- Context around design decisions.
SRE provides
- Operational support ownership preparation.
- Validation of monitoring and alerting.
- Incident response practice.
- Operational documentation updates.
- Reliability assessment.
- Identification of improvement opportunities.
The objective is not simply knowledge transfer. It is to build operational confidence within SRE before ownership transfer.
A service should only move into steady-state operation when:
- SRE understands the service architecture and behaviour.
- Operational procedures have been validated.
- Monitoring and alerting are effective.
- Support processes are established.
- Outstanding operational risks have owners and agreed actions.
Managed Operations
After successful handover, SRE assumes operational ownership of the service.
Managed operations include:
- Continuous service monitoring.
- Incident response and coordination.
- Problem management.
- Reliability reporting.
- Operational maintenance.
- Observability improvement.
- Capacity and performance monitoring.
- Reliability engineering activities.
SRE operates the service using agreed standards for monitoring, logging, metrics, alerting and incident management.
Continuous Reliability Improvement
Managed SRE ownership includes ongoing improvement after service transition.
Reliability improvements are identified through:
- Incident reviews.
- Alert quality reviews.
- Service performance trends.
- SLO performance.
- Error budget consumption.
- Operational workload analysis.
- Recurring operational issues.
Improvement activities may include:
- Reducing operational toil.
- Improving automation.
- Improving observability coverage.
- Increasing resilience.
- Reducing incident frequency.
- Improving recovery processes.
Roles and Responsibilities
| Role | Responsibilities |
|---|---|
| Delivery team | Builds the service, provides service knowledge, resolves application issues and supports transition activities. |
| Service owner | Defines business expectations, priorities and acceptable reliability outcomes. |
| Managed Service SRE team | Provides operational ownership, monitoring, incident response and reliability improvement after transition. |
| SRE leadership | Ensures appropriate capability, standards and governance are applied. |
| Platform teams | Provide supporting infrastructure, shared services and platform capabilities. |
Service Transition Exit Criteria
A service is ready for full SRE ownership when:
- Service documentation is complete.
- Ownership boundaries are agreed.
- Monitoring and alerting meet required standards.
- SRE has appropriate access and operational knowledge.
- Incident processes have been validated.
- Known risks have owners and remediation plans.
- Delivery team dependencies have been reduced to normal operational engagement.
- SRE has confidence operating the service independently.
Reliability Outcomes
A successful Managed Service SRE engagement results in:
- Clear operational ownership.
- Reliable service operation.
- Reduced operational risk.
- Faster detection and recovery from incidents.
- Improved observability.
- Reduced manual operational effort.
- Continuous reliability improvements driven by operational insight.
The SRE Service Model provides the structure required to move services from delivery into stable, measurable and continuously improving managed operations.