On this page6 sections
When an application shells out to codex exec, the agent becomes a worker process with a lifecycle the host must own. The interesting part is not spawning it. The interesting part is knowing which task is running, which events belong to it, whether a timeout ended the process or only the wait, and what state is safe to resume.
Consider a hypothetical internal release tool that starts a Codex worker to update one repository, expects a tested patch, and must recover after the host restarts. Its supervisor needs a stronger result than “the process stopped.”
A recent r/codex project makes that problem concrete. CAD Studio launches codex exec --json from a Mac application, maps JSONL events into chat steps, and resumes a thread for follow-up work. Its author reported that resume can fail while an earlier writer for the same thread is still alive, so the application waits, retries, and terminates child processes when it exits. That is a field report from one implementation, not a universal Codex guarantee, but it exposes the host responsibility that a successful demo can hide.
Choose the integration layer before writing a supervisor
Official OpenAI documentation separates three integration levels. codex exec fits scripts, CI jobs, and bounded background work. The Codex SDK fits application code that needs to start, resume, or stream tasks programmatically. App-server fits products that need persistent conversations, event streaming, interruption, tool exposure, and approval handling.
Use the lightest layer that matches the lifecycle. A one-shot repository check can run as a child process and return a structured result. An interactive product with multiple concurrent users, approvals, and durable conversations is already asking for a protocol and session manager. Rebuilding those behaviors around shell processes may be more work than choosing the SDK or app-server.
The Reddit example remains useful because many internal tools begin with a single worker. The host can make that worker reliable if it treats standard input, standard output, standard error, the process ID, working directory, thread ID, and generated artifacts as one execution record.

Treat JSONL as an event stream, not the result
OpenAI's Codex evaluation guidance documents that codex exec --json emits newline-delimited JSON events. Command executions and file changes appear as item events, and the trace can support deterministic checks. The official example saves the trace, parses each line, then inspects whether required commands and files appeared.
That stream explains progress, but a progress event is not the same as accepted completion. A command may start and later fail. A file may change before a build rejects it. A final agent message can sound complete while the required artifact is absent. Define completion in the host: expected terminal event, zero process exit when appropriate, required artifact checks, and the application-specific verifier.
Persist the raw event stream before transforming it for a UI. A renderer may collapse repeated command output or hide reasoning items. The raw trace lets a later verifier reconstruct the order, correlate a failed item, and distinguish an absent terminal event from a display bug. Redact secrets at the source and keep access to traces narrower than access to the finished artifact.
One task needs one process owner
Create an execution ID before spawning the child. Store the process ID, thread ID when present, working directory, requested sandbox and approval policy, start time, deadline, and expected outputs. Update a heartbeat from observed events. The parent process, not the agent, owns cancellation and cleanup.
Do not allow two writers to resume the same thread concurrently. A per-thread lease can be as simple as an atomic database row with an owner and expiry. Renew it while the child is alive. Release it only after the child exits and output streams are drained. On startup, reconcile an expired lease against the actual process and execution record before spawning a replacement.
An operational agent skill needs a contract for incomplete execution, not just instructions for the successful path. Specify what was authorized, which state was last established, and how the executor will reconcile a write whose result is unknown. The agent skill execution guide applies the same boundary to tool authorization and recovery.
That matters when the parent crashes after a file change but before receiving the terminal event. Starting the same task again may overwrite accepted work or repeat an external side effect. First inspect the workspace, Git status, expected artifacts, and any destination the agent was allowed to change. Classify the outcome as completed, failed, still running, or unknown. Unknown is a real state that requires reconciliation. The unknown-write reconciliation guide explains why a timed-out request must be read back before another write is issued.
Cancellation has more than one edge
A deadline should trigger an ordered shutdown. Stop accepting new input, send the supported interruption signal, wait a bounded grace period, terminate the child if it remains alive, and finally escalate to a hard kill only when necessary. Keep reading output during shutdown so the trace includes late events and errors.
Killing only the parent shell may leave descendants running. Spawn the worker in a process group or platform equivalent and terminate the owned group. Do not match broad process names, because another user's Codex session may be legitimate. Record every signal and the observed exit status.
After cancellation, never assume the workspace is clean. Run the same artifact checks used after normal completion, then decide whether to keep, revert, or quarantine changes. For a code task, that can include a diff, build, tests, and a check for untracked files. For a CAD workflow, it can include the source script, generated model, and a geometry verifier.
Sandbox and approval settings are part of the record
The Reddit command uses a workspace-write sandbox and an approval policy chosen by the host. Official guidance recommends the least permissions needed, especially in automation. Store those settings with the execution because they determine which actions could have occurred.
A host should not silently retry a task with broader authority because the first attempt stalled. If the worker requires network access, a larger write scope, or a consequential tool, move the execution into a state that the product can explain and approve. A retry with different permissions is a new decision, even if the prompt is unchanged.
Separate agent progress from business writes. Let the worker create or inspect artifacts inside its sandbox. Make the host validate those artifacts, then perform any irreversible publish, send, or delete through an explicit application action. This keeps a model retry from repeating an external side effect.
Verify the artifact and the supervisor
Test the supervisor with deliberate failures. Start a worker that exceeds its deadline, crashes after writing a file, emits malformed JSONL, exits without the expected terminal event, and leaves a descendant process. Attempt two resumes for one thread. Restart the host while a task is running. Each case should produce a bounded state and enough evidence to decide what happens next.
Then evaluate the agent's task separately. OpenAI's guidance recommends a small set of deterministic checks over the trace and artifacts, followed by qualitative checks where rules cannot decide. Keep those signals distinct. A perfect supervisor can reliably deliver bad work; an excellent agent can be undermined by a host that loses its exit state.
The worker pattern is complete when one execution ID connects the request, child process, event stream, thread, workspace changes, verifier, and accepted result. At that point codex exec is no longer an opaque subprocess. It is a bounded component the application can observe, stop, reconcile, and trust only as far as its evidence allows.
Source context
Source links appear within this article. Read them alongside the author’s analysis and evaluate the guidance against your environment.
Report an error or outdated detail