Guide

Auditing AI-generated tests and documentation before cleanup

In brief

Review the claim each test or document protects, challenge the evidence, and record a keep, rewrite or retire decision before deleting repository history.

6 min read

Sources
An archivist sorts code and documentation sheets into keep, review, and retire boxes using an evidence ledger.
Original AI-generated conceptual illustration; not documentary photography or a product image.
On this page5 sections

Repository cleanup becomes risky when tests and documentation were produced faster than anyone could review their purpose. Some files are redundant. Others are awkward but preserve the only executable record of a failure, compatibility promise, or operating decision. Deleting by age, naming pattern, or authorship can make the tree look cleaner while removing the evidence that explains why the system behaves as it does.

The practical unit of review is not the file. It is the claim the file protects. A test may claim that a retry stops after three attempts. A runbook may claim that an operator can identify an unknown deployment outcome without rerunning the change. Before keeping, rewriting, or retiring either artifact, make that claim inspectable.

ChatGPT is most useful in an SRE workflow when the output has a factual boundary the responder can inspect. For repository cleanup, that boundary is the artifact, the claim it protects, the evidence that challenges the claim, and the recorded disposition. Authorship can guide review priority, but it cannot substitute for the evidence.

Inventory consumers, not just files

Start with a generated inventory of test files, documentation pages, ownership metadata, and references. Add each artifact's last meaningful change, not merely its latest formatting commit. Git log can follow a path and show the commits that introduced or revised it, but history is context rather than a keep rule. An old regression test may still protect a current path. A new guide may already duplicate the canonical runbook.

Search for consumers. Test names may appear in CI filters, dashboards, quarantine lists, or release procedures. Documentation URLs may be linked from alerts, onboarding checklists, support replies, or code comments. A page with little web traffic can still sit on a critical incident path. Record those relationships before moving anything.

Group the inventory by protected decision. Authentication expiry tests, token-refresh documentation, and an on-call note about stale sessions may all describe one operating boundary. Reviewing them together exposes contradiction and duplication more reliably than a file-by-file pass.

Do not label every AI-assisted artifact as suspect and every human-authored artifact as trusted. Instead, use generation provenance as one input to triage. Large batches created in one session, files without a named consumer, and repeated examples with only noun changes deserve early review because they can conceal superficial coverage. The disposition still depends on what they prove.

A repository audit moves from inventory to an explicit protected claim, challenges the claim, and records a keep, rewrite, or retire decision.
Evidence-preserving cleanup. Deletion is one possible disposition after the protected claim has been identified and challenged.

Translate coverage into protected behavior

Line coverage answers whether a statement ran. It does not show whether an assertion would fail when the behavior changes. Coverage.py's branch coverage adds evidence about alternative control-flow paths, which is useful when generated tests exercise only the success branch. Even branch coverage cannot tell you whether the asserted outcome is meaningful.

For each candidate test, write one sentence: “This test fails if...” Then verify that the sentence names an externally meaningful behavior, invariant, or failure boundary. “This test fails if the mock returns another value” is usually weak. “This test fails if a timed-out deployment is reported as failed before the destination state is read back” protects an operational decision.

Challenge the sentence. Temporarily change the protected behavior, use a mutation-testing tool, or replace the implementation with a plausible wrong answer. A surviving mutation does not automatically make the test useless; the mutant may be equivalent or outside the intended contract. It does reveal that the current evidence did not distinguish the change. Review that result before deciding.

Consider a hypothetical client with three generated retry tests. All three mock a timeout, call the same function, and assert that an exception is raised. The production contract is more specific: idempotent reads may retry, writes with an unknown outcome must stop and request reconciliation, and authentication failures must not retry. The three files produce activity but protect only one vague claim. Rewrite them as a small decision table with distinct assertions. Retire the originals only after the replacement fails against each deliberately wrong policy.

Make documentation prove a reader action

Documentation needs a similar challenge. A link checker can find missing targets, and MkDocs strict mode can turn warnings into build failures. Those checks establish that the page can be built and followed. They do not prove that the reader can make the promised decision.

Name the page's reader, trigger, and output. A runbook should say when to enter it, what evidence to collect, which actions are authorized, how to stop, and what state confirms recovery. A design note should preserve the rejected alternatives and the assumption that would reopen the choice. A tutorial should end in an observable result rather than a sequence of commands alone.

Then walk the page from a clean state. Resolve every internal link, run commands in the supported environment, and compare stated defaults with current configuration. For incident material, use a safe scenario or fixture. Do not test destructive steps against production merely to validate prose.

If two pages make the same promise, choose one canonical destination and redirect or replace the other. Keep a short compatibility note when external links or old alerts still use the retired URL. A silent deletion turns a documentation cleanup into an incident for the next reader.

Choose keep, rewrite, or retire

A keep decision means the claim is current, the evidence can still challenge it, and ownership is clear. A rewrite decision means the claim matters but the artifact expresses or verifies it poorly. A retire decision means the claim is obsolete, duplicated by a verified replacement, or no longer belongs to the supported product boundary.

Record the decision beside the evidence. Include the original path, protected claim, consumer search, challenge performed, replacement if any, reviewer, date, and rollback note. For large repositories, merge cleanup in small groups by decision boundary. That makes a regression easier to trace and a mistaken retirement easier to reverse.

The site's runbook template is useful for documentation that governs operational action. The incident-response agent evaluation applies the same principle to AI systems: a plausible output is not acceptance evidence unless an independent observation can disagree with it.

Stop the cleanup when a protected claim cannot yet be challenged. That boundary changes the recommendation from retire to keep or quarantine until the missing evidence exists. “No one remembers why this file exists” is a reason to investigate, not proof that the file is safe to delete.

Leave the repository more explainable

Measure the cleanup by what becomes easier to understand and verify. Useful outcomes include fewer duplicate claims, a smaller set of authoritative pages, tests that fail against meaningful counterexamples, explicit owners, and working redirects. Raw file reduction is secondary. It can reward deletion even when the remaining system is harder to operate.

AI-assisted tests and docs should pass the same evidence gate as any other contribution. The extra concern is scale: generation can produce many plausible artifacts before the first weak assumption is noticed. Inventory the claims, challenge them, and retain the disposition. A clean repository is not the one with the fewest files. It is the one whose important promises can still be found, tested, and explained.

Source context

Source links appear within this article. Read them alongside the author’s analysis and evaluate the guidance against your environment.

Report an error or outdated detail

Related reading

Explore a related question