Guide

Where to start with Google’s free SRE books

In brief

A chapter-level starting point for alerting, on-call coverage and postmortems, with examples of adapting the practices to local staffing and authority.

4 min read

Sources
An open reference book links to a rehearsal board with two role markers and a route that stops before an external receiver.
Conceptual reading-to-practice trial: a chapter can improve coordination while a mitigation still depends on reaching the authorized owner. The unfinished route identifies an assumption to test locally; no practice adoption or recovery has been proved.
On this page3 sections

The official Google SRE books page makes Site Reliability Engineering, The Site Reliability Workbook and Building Secure & Reliable Systems available to read online for free. Their breadth is an advantage when used as a reference: a team can bring a current service question to the relevant chapter without first reading every book in order.

The first book develops the discipline and its foundations. The workbook offers practical approaches and examples. Building Secure & Reliable Systems connects security and reliability in system design and operation. The useful choice is where to begin given the decision your team is struggling to make.

Chapters for alerting, coverage and postmortems

Difficulty agreeing on acceptable reliability points toward objectives and error budgets. Frequent pages with no useful response suggest monitoring and on-call practices. Reviews that produce little follow-through can be read alongside the postmortem material and the team's actual action lists.

For alerting, keep a noisy alert beside the workbook chapter on alerting from SLOs. For coverage, use the on-call chapter to follow the primary responder’s route to the expertise and authority needed for recovery. Trying that route locally reveals which parts of the practice already fit and which need staffing, access or fallback work. That gives the team a specific question to investigate before adopting the practice.

Some established practices will fit directly. Use them with their original attribution when local conditions support them, and adapt the parts whose assumptions do not hold. There is no need to invent a local replacement simply to make adoption look original.

The staffing and authority a practice assumes

A practice may depend on clear ownership, enough staffing, dependable telemetry or permission to change the release plan. Copying the visible steps does not create those conditions. Naming the dependency explains why a sound method can remain difficult to use in a different organization.

Consider a hypothetical team adopting an incident-command model while a consequential mitigation still requires an unavailable external owner. The new roles improve coordination, but the same action remains blocked. The local adaptation must address authority and fallback, rather than treating compliance with the role diagram as completed adoption.

An error-budget policy raises a related question. Its release guidance can change behavior only when the people involved are able to act on it. Reading with that assumption in mind separates a practice ready for immediate use from one requiring a staffing, access or decision-making change first.

Turn a chapter into a small local experiment
Turn a chapter into a small local experiment
Suggested reading workflow. Use a current service question to select material, identify its assumptions and test one adapted practice before expanding it.
Read diagram description

Suggested reading workflow. Use a current service question to select material, identify its assumptions and test one adapted practice before expanding it. Diagram labels: Current service question: What does the team need to improve?; Relevant chapter: Read the mechanism and its assumptions; Local adaptation: Check roles, access and available capacity; Small experiment: Record what transferred and what did not.

A local trial connects the chapter to the service

For the team adopting incident command, a local trial can follow a mitigation request from the incident lead to the person authorized to approve it. The present obstacle is already known: that owner may be unavailable. Observing whether the proposed fallback makes the decision possible gives this trial a result to examine before the arrangement is extended to other services.

A monitoring review might examine the last month of pages to identify which required immediate action. A postmortem review might inspect whether selected improvements have owners and acceptance evidence. Neither exercise requires assuming that another organization's targets or staffing arrangements are universal.

After the trial, return to the question that brought you to the chapter. Could the responder reach someone able to authorize the mitigation? Did the alert review preserve necessary coverage while identifying unnecessary interruption? Record what transferred directly, what needed adaptation and what remains unresolved.

If the mitigation remains blocked by unavailable authority, the team now knows which support arrangement needs work. If the new response path works, it has a reason to extend the practice. In either case, the chapter has helped explain a local problem and supplied something concrete to try, making the books easier to return to when the next question arises.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

More on Google

All Google coverage

Related reading

Explore a related question