Tracing an AI request through Cloudflare and its origin
Follow a slow AI request from Cloudflare to an instrumented origin, configure sampling, and verify that the trace preserves the evidence needed to investigate.
AI in production. Reliability in practice.
Latest edition ·The latest analysis
A million tokens can hold an incident’s history. The harder question is whether the model can find the evidence that matters.
Read the storyFrom the publication
Follow a slow AI request from Cloudflare to an instrumented origin, configure sampling, and verify that the trace preserves the evidence needed to investigate.
Build a latency SLI from Prometheus histogram buckets, include failed requests, and compare classic and native queries with a worked search-service example.
A synthetic order-service incident follows a shared handoff through page edits, new findings and shift change, including the review and access work involved.
ninfer-ext trades a smaller EXL3 file for more decoding work in its published tests. The saved memory matters when it lets a request fit that would otherwise fail.
Field guides
Follow a request. Review the code. Find out what actually happened.
Explore the guidesKeep the traces that explain a failure
Tail sampling selects recorded traces by error, latency and other evidence. A checkout timeout shows how routing, waiting time and memory affect the trace a responder can retrieve.
Give AI code review a useful job
CodeRabbit brings AI feedback into pull requests, editors and the CLI. A retry loop that exceeds its caller’s deadline shows the kind of finding that can repay review time.
Recover when a tool call times out
A missing response can leave a worker already created. Durable operation IDs and the destination’s idempotency contract let a replacement executor recover the original request.
Put it into practice
Free reliability tools. No account needed.
Open the workbenchThe AI tools guide
Coding assistants, code review and agent supervision.
Browse the guideCompanies & products