Explainer

Pi.dev: a coding assistant you can adapt to your work

In brief

See how Pi works with your files and tools, then follow an incident-review example to decide whether it could make your team’s work easier.

9 min read

Sources
Conceptual steel chassis with swappable modules, one removed module being measured by a separate dial gauge.
Original AI-generated conceptual artwork for AIOpsSRE; not Pi hardware or a product screenshot.
On this page7 sections

Pi.dev: a coding assistant you can adapt to your work

An incident review can leave you with plenty of information and no clear starting point. The deployment history tells you when the software changed, application logs show what happened around that time, and a runbook suggests what to inspect. Before you can explain the failure, you have to bring those pieces together.

Pi can help with that first pass. It is a coding assistant that runs in a terminal and works with files on your machine. You give it a task and choose the model it will use; Pi can then read files, search for relevant details, run commands and edit code. For an SRE, that makes it worth exploring for work that combines written instructions with technical evidence.

What makes Pi particularly interesting is how much you can adapt. You can give it a smaller set of tools for a review, save instructions for a recurring job or add a connection to one of your own services. To see why that matters, let’s follow one hypothetical incident review from the files you supply to the report another engineer would receive.

Start with a question someone can answer

In this hypothetical example, a checkout service becomes slower after a deployment. You want to know what the available evidence supports and what to examine next. Use three files with credentials and private customer details removed: deployments.csv, containing revisions and rollout times; checkout.log, containing selected timestamped events; and runbook.md, containing the investigation procedure.

The folder deliberately has no information about the downstream services checkout depends on. That gives the exercise a useful gap. If Pi cannot see how those services behaved, its report should leave that question open. You are looking for help organizing an investigation, so a clear account of what is missing can be as useful as a promising explanation.

The request can be straightforward: “Review these files, separate observations from hypotheses, cite the filename and line numbers behind each factual claim, and suggest the next checks.” A report that connects the rollout time to the first slow requests could help the reviewer decide which change to examine. Its references should also make it easy to check whether those events actually line up.

This gives Pi a useful job before you connect it to anything live. It also gives you a way to judge the result: can another engineer understand the findings, check them against the files and choose a sensible next step with less effort?

A bounded Pi investigation
Sanitized read-only evidence flows through Pi to a candidate report, then an engineer checks cited files before accepting a next step.
Illustrative pilot: Pi produces a report for review. Production access and any resulting change remain outside this read-only workflow.
Read diagram description

A sanitized evidence bundle is mounted read-only. Pi reads and searches it without production credentials. Its candidate report separates cited observations, hypotheses and missing evidence. An engineer checks the claims against the original files before accepting a useful next check. Any later production action requires separate authorization and independent outcome verification.

How Pi works through the files

Pi sends your request to the selected model along with the conversation so far, relevant instructions and descriptions of the available tools. The model might ask to read deployments.csv, inspect part of checkout.log and then search for a particular error. Pi carries out those requests and returns the results to the model, which can continue investigating or write its answer. The documentation on how Pi works describes this repeating process.

Those file operations are what Pi calls tools. Its default selection includes read, bash, edit and write, so an ordinary coding session can inspect files, run shell commands and change code. For our investigation, file-reading and search tools are enough. Pi also provides grep, find and ls, and its --tools option lets you replace the selected set for that run. You can give the assistant what it needs for the question without offering code-editing tools it has no reason to use. The exact options are in the CLI reference.

The software managing this back-and-forth is often called an agent harness. In Pi, it handles model requests, tool execution and the conversation record. That record can branch when you want to explore another line of inquiry. During a long session, Pi can summarize older material for subsequent model requests while retaining the original entries. This can help you return to an earlier investigative path, although the saved conversation still needs the same care as the logs it contains. Pi’s session guide explains how history and summaries are stored.

Make a useful investigation easier to repeat

If the first exercise helps, the next improvement may be quite small. A prompt template can save the request so that reviewers do not have to reconstruct it each time. A skill can add instructions and supporting material for the task, such as how your team distinguishes an observation from a hypothesis. Both let you shape the work without starting by writing a custom program.

An extension becomes useful when you need code to do something new. Pi extensions are TypeScript modules that can add tools, respond to events or change the information sent to the model. For our example, an extension could provide a tool that retrieves an approved incident export. Pi’s extension guide explains the available building blocks. Teams can also distribute reviewed prompts, skills and extensions together as a Pi package, using a pinned version so that a shared setup changes deliberately.

There is now another way to connect services. Pi 0.99.0, released September 29, 2026, added built-in support for the Model Context Protocol, or MCP, and Codemode. MCP gives an assistant a common way to discover and call tools supplied by another program. If your evidence service offers an appropriate MCP connection, Pi could retrieve the bundle through it instead of relying on a manual export.

Codemode lets the model compose tool operations using JavaScript. For example, a connected tool could retrieve the incident export and another could list the related deployments, giving the model both results to examine. The MCP documentation covers configuration, while Earendil’s explanation of the release describes the design choice. These are options to explore after the basic investigation is useful; the file-based trial does not depend on them.

Read the report against the original evidence

Suppose the report says the slowdown began soon after the rollout and recommends inspecting the changed code. Open its cited lines and check that the timing supports that next step. The deployment may be a good lead, but its timing alone does not establish the cause. Because our bundle leaves out downstream-service information, a claim that those services were healthy would need evidence the assistant was never given.

That comparison is the important part of the trial. Keep the report beside its input files and note which findings helped, which needed correction and whether the suggested checks were practical. Include the time spent doing that review. A report produced quickly can still cost more time if its claims are difficult to verify.

Decide whether the extra control helps

Pi’s flexibility is valuable when it lets you make a recurring job easier to do well. In this example, that might mean a consistent request, the right files available from the start and a report whose references save the reviewer time. If an existing tool already does that adequately, maintaining a custom Pi setup may add work without improving the investigation.

If the main difficulty is finding which coding session needs attention, the Herdr guide covers that supervision problem. Pi is a closer fit when you want to change how the assistant does the work itself.

Keep the first trial small enough to learn from. If checking the reports takes more effort than doing the work directly, adjust the request, inputs or model before connecting additional systems. If the reports consistently make the next decision easier, you have a reason to develop the setup further. Any later ability to change a deployment needs its own allowed target, action and current-state checks, as described in the guide to operational agent skills.

If you want to try the three-file exercise, the setup details below explain how to limit access and run it. The automation notes are there for later, once you have a report worth producing again.

Optional setup details

For this exercise, run the whole Pi process in an isolated environment, such as a suitable container or sandbox, with only the sanitized folder mounted read-only. Leave out production credentials and allow only the model connection approved for the trial. This keeps the experiment focused on whether Pi helps with the report, without giving it access to make production changes.

These restrictions need to come from the environment because Pi uses the permissions of the account that runs it. Choosing a working folder does not confine access to that folder, and Pi does not ask before every tool call. Project trust controls which project settings and resources load; it does not restrict all later actions. Pi’s security guide explains the distinction.

Extensions make it worth checking where the complete process runs. If Pi and its extensions are inside the isolated environment, the same limits apply to both. If only selected tools are sent into a sandbox while Pi stays on the host, other extensions may still run on the host. Files, credentials and network access that you expose remain available inside the environment. The isolation guide compares the arrangements and provides setup instructions.

Once that environment is ready, follow the official quickstart. The npm installation uses @earendil-works/pi-coding-agent and requires Node.js 22.19 or newer. Choose an approved model provider and record the output of pi --version so you know what you tested. Model access is separate from installing Pi, so check both its cost and whether your files may be processed by that provider.

After the environment and provider are configured, this illustrative command selects file-reading and search tools and asks for the report described above:

pi --no-approve --no-extensions --no-skills \
  --no-prompt-templates --no-context-files \
  --tools read,grep,find,ls --print \
  "Review deployments.csv, checkout.log and runbook.md.
   Separate observations from hypotheses.
   Cite filenames and line numbers for factual claims.
   Identify missing evidence and propose the next checks."

Here, --no-approve skips trust-gated project resources; it does not turn on approval prompts. The other resource flags reduce loaded customization, and --tools selects the listed built-ins. The CLI reference documents these choices. The environment’s restrictions still determine what Pi can actually reach.

Optional automation details

Once the manual trial helps, you may want a script to start Pi whenever a new bundle arrives. The first detail to settle is how that script knows the report is ready. Print mode signals an error to the script when the final assistant response is marked error or aborted.

JSON mode writes a stream of records describing what Pi does, with one JSON object on each line. A failed or aborted assistant response does not by itself make that process exit with an error status, so the script needs to inspect the records. Their JSON format also does not require the model’s final answer to follow a particular structure.

For a long-running integration, Pi’s RPC interface lets another program send commands and receive responses. A successful prompt response can mean the request was accepted or queued. To wait for automatic work to finish, keep listening until agent_settled arrives; agent_end can still be followed by recovery or queued work. The CLI integration documentation explains these differences. Your program can then hand the report to a reviewer, who checks whether its conclusions follow from the evidence. Our guide to evaluating incident-response agents develops that approach further.

If the next version retrieves evidence through MCP, Pi can connect to a local server using stdio or a remote server over streamable HTTP, with OAuth available for remote sign-in. Check the server’s error behavior as part of that integration. Pi may retry resource reads after a temporary failure, but it does not automatically retry tool calls because a server may already have carried out the action. The distinction becomes important if the service eventually offers tools that can change something, so preserve the unresolved result for inspection before attempting it again. These details are covered in the MCP documentation.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

A useful next step

Continue the work