AI in production. Reliability in practice.

Latest edition ·

The latest analysis

A benchmark cannot separate the model from its harness

Read coding-agent benchmarks as system comparisons when harnesses differ, then use controlled contrasts and independent tests before attributing a result to the model.

5 min read

Read the story
Inside the story

From the publication

The latest

All articles

Field guides

Get closer
to the evidence.

Follow a request. Review the code. Find out what actually happened.

Explore the guides

Keep the traces that explain a failure

Give AI code review a useful job

Recover when a tool call times out

Put it into practice

Less blank page.
More working plan.

Free reliability tools. No account needed.

Open the workbench

The AI tools guide

Find the right AI tool for the job.

Coding assistants, code review and agent supervision.

Browse the guide

Companies & products

Know what you’re running

All coverage