Practical guidance for SRE and platform teams operating reliable services and production AI. Independent analysis, field guides, and working resources.

Selected reading

Prompt caching: measure the cost of useful work

Build a prompt-caching pilot that counts retries, failed work and unknown charges, with a local calculator and sample ledger.

A practical pilot for separating cache savings from the cost of completed work.

7 min read

Read the article
Reusable metal printing plate beside fresh geometric prints and a copper paperclip on a charcoal workbench.

Original writing / Newest first

More from AIOpsSRE

All articles
AIOps

What MTTD misses

Separate detection from response and find delays that an improving average can conceal.

Releases, research & operational impact

Latest AI news

All news

Agent evaluationNews analysis

SWE-Serve tests coding agents through live inference

NVIDIA’s new SGLang benchmark finds patches that pass local checks but fail a running server. Its result is a reason to examine what a green test actually covers.

3 min read

Local AI agentsNews analysis

Antigravity brings local models to its agent SDK

Google adds local Gemma and LiteRT workflows. Running the model on a workstation still leaves a separate decision about what its tools may read, change or send.

3 min read

From reading to doing

Work through it, step by step.

All guides