Company guide · AIOps & SRE

Google tools and cloud operations for SRE

From incident sources to running services

Google coverage on AIOpsSRE follows two kinds of work: understanding the material an engineer already has, and operating the services an AI application depends on. A notebook can help compare conflicting procedures. A gateway or telemetry agent changes how requests and observations move through a running system. The articles below explain those mechanisms so you can choose the part of the workflow that needs attention.

When the documents disagree

Start with comparing incident sources with NotebookLM when a runbook and an incident review give different instructions. Google's current documentation calls the product Gemini Notebook; the article retains NotebookLM in its title for readers familiar with that name. Its source-linked chat lets you follow an answer back to the relevant passage. See Google's chat documentation.

The guide's hypothetical restart example shows why the dates and conditions attached to each instruction matter. Finding both passages helps the runbook owner resolve the disagreement. A notebook built from historical documents cannot establish the service's current state, so the result belongs in a reviewed procedure rather than becoming an automatic restart instruction.

For a question about the team's operating practice, the guide to Google's free SRE books takes a different route. It connects alerting, on-call coverage and postmortems to specific chapters, then asks which staffing, access and authority assumptions need local adaptation.

Follow a slow or incomplete agent response

For traffic already passing through Apigee, the model and MCP dashboard explainer shows how to separate model latency from tool latency. The useful detail is the filter: pooled tool percentiles can change when the mix of calls changes, even if an individual tool has not slowed down. Select the affected tool and time window before following the investigation into backend traces. The Apigee dashboard reference describes the available views.

API Gateway streaming addresses another delay: waiting for a complete response before any output reaches the client. The September 29 preview report explains creation-time gateway settings and deadlines. Earlier visible text and a successfully completed answer are separate results. Streaming remains Public Preview; consult the current configuration documentation for its restrictions when deciding whether to migrate.

Retries and telemetry need their own checks

Cloud Tasks per-task retries explains how jobs on one queue can use different retry settings. An incident-summary request may expire sooner than a historical export. If a handler can change state, losing the response after a successful change also makes duplicate delivery consequential. The article follows both cases, including why an application deadline and the previous write's outcome still matter.

Use the legacy-agent migration report for existing Monitoring and Logging agents. Its disk-label example follows a measurement into the query responders actually use: incoming data can survive a migration while an old filter stops matching it. This is a different investigation from model latency, even when both appear on an operational dashboard.