What MTTD misses
Separate detection from response and find delays that an improving average can conceal.
AI in production. Reliability in practice.
Start hereFeatured guide
A lost response can leave a worker running even when its caller sees a timeout. Follow the recovery from operation identity to destination evidence.
Selected reading
Understand your detection gaps
Separate detection from response and find delays that an improving average can conceal.
Follow through after an incident
Separate a completed task from evidence that the underlying risk was reduced.
Clarify who owns reliability
Clarify ownership across shared platforms, service reliability, and incident response.
Build a latency SLI from Prometheus histogram buckets, include failed requests, and compare classic and native queries with a worked search-service example.
A synthetic order-service incident follows a shared handoff through page edits, new findings and shift change, including the review and access work involved.
ninfer-ext trades a smaller EXL3 file for more decoding work in its published tests. The saved memory matters when it lets a request fit that would otherwise fail.
Follow three incident files through Pi’s agent loop to a report another engineer can verify, then decide which customization is worth maintaining.
Companies & products
AI Tools Guide
Coding assistants, code review and agent supervision.