NotebookLM for SRE: build a source-backed incident dossier
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
In-depth explanations, practical guides and independent analysis. Choose a topic and format to find your next read.
Guides show how. Explainers unpack a concept. Commentary makes an argument. Explore Production notes.
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
Test evidence retrieval, long-session consistency, tool boundaries, and cost before letting a model influence production changes.
Define inputs, permissions, idempotency, verification, and stop conditions for one operational capability before expanding its authority.
Connect on-call interruptions, specialist demand, and recovery work to a concrete capacity decision using an on-call workload worksheet.
Turn reliability policy into a release decision with fresh evidence, limited exceptions, and a clear path to resuming normal changes.
A design guide for an incident assistant that handles duplicate events, partial failures, and reviewed AI summaries across tools.
Build a view that answers who is affected, what changed, and where to investigate, then test its queries and missing-data behavior.
A reusable prompt pattern for incident analysis, with checks for missing facts, unsafe recommendations, and unsupported certainty.
Adapt a complete SLO record, worked budget examples, and an exception log without confusing bad requests with minutes of downtime.
Use directory purpose, mount boundaries, and read-only checks to investigate missing files, full disks, and unexpected runtime state.
Measure a representative workload, identify the limiting resource, and test one reversible change instead of applying a universal sysctl recipe.
Design incident channels, bot actions, evidence links, and a fallback path that still works when the chat system is unavailable.