NotebookLM for SRE: build a source-backed incident dossier
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
Practical AI operations. Reliable systems.
Walkthroughs for observability, incident response, and AI-assisted operations.
Start with what AI can do for operations, and where judgment still matters.
Read AIOps fundamentals ↗Make metrics, logs, and traces work together.
Read Observability for SRE ↗Build an incident response practice that learns.
Read Incident management with AI ↗Instructions, examples, and implementation notes.
Use a curated notebook to compare runbooks and incident evidence, while keeping citations, freshness, and operational authority in view.
A design guide for an incident assistant that handles duplicate events, partial failures, and reviewed AI summaries across tools.
A reusable prompt pattern for incident analysis, with checks for missing facts, unsafe recommendations, and unsupported certainty.
Use Python logging to capture useful context without swallowing errors, duplicating exceptions, or exposing sensitive data.
Validate the source data, preserve useful denominators, and distinguish a failed collection from a real zero before aggregating metrics.
Build a ticket-routing evaluation around real ownership, uncertain cases, and correction cost. Start with an honest baseline before adding a model.