Guide

Python logging and exit status in a scheduled job

In brief

A complete Python example sends the result to stdout, records an exception once and returns the failure status its scheduler needs.

4 min read

Sources
A torn input sheet branches to a context scroll and a coral status pennant carrying the same failure mark.
Conceptual view of the example job: exception context helps the responder, and a nonzero exit status tells the scheduler that counting failed. The illustration is not an executed test.
On this page2 sections

Consider a hypothetical aggregation job that encounters an exception, writes an error to its log and exits with status zero. A person reading the log would call the job a failure. A scheduler using the exit status would call it a success and never take its failure route. Adding more detail to the error message would not resolve that disagreement.

A small command-line program therefore has two audiences. Its caller needs a dependable completion signal; the person investigating a failure needs context. Python's standard logging module supplies levels, handlers and exception information for the second audience. The program must also make sure its terminal status answers the first audience's question: did the requested work complete?

One exception record, one failing process status

The example counts records in a CSV file that must contain a service column. Save it as count_records.py and run python3 count_records.py input.csv. It uses only Python's standard library, prints the successful result to standard output and sends operational messages to standard error. That separation lets the caller consume the count without mistaking a log message for data.

import argparse
import csv
import logging

log = logging.getLogger("record_counter")

def count_records(path):
    with open(path, newline="", encoding="utf-8") as stream:
        rows = csv.DictReader(stream)
        if not rows.fieldnames or "service" not in rows.fieldnames:
            raise ValueError("CSV must contain a service column")
        count = 0
        for row in rows:
            if not (row.get("service") or "").strip():
                raise ValueError("CSV contains an empty service value")
            count += 1
        return count

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("path")
    args = parser.parse_args()
    logging.basicConfig(
        level=logging.INFO,
        format="%(asctime)s %(levelname)s %(name)s %(message)s",
    )
    try:
        count = count_records(args.path)
    except (OSError, UnicodeError, ValueError, csv.Error):
        log.exception("record_count_failed")
        return 1
    print(count)
    log.info("record_count_completed records=%d", count)
    return 0

if __name__ == "__main__":
    raise SystemExit(main())

Follow an invalid file through the program. The counting function checks the header before counting rows and rejects an empty service value as it reads. If validation fails, count_records raises an exception instead of returning a plausible count for incomplete work. The exception reaches main, where log.exception records the traceback and the operation returns a nonzero result.

That boundary is also a sensible place to log the terminal failure. Logging the same exception inside the counting function and again in the caller would create two records for one failed operation. Here, the lower-level function explains failure by raising it; the operation-level handler records what ended the job.

The distinction between invalid input and an empty result is deliberate. A completely empty file has no required header and fails. A valid header with no data rows represents a valid count of zero and succeeds. Both contain no records, but only one satisfies the input contract. Treating them identically would make a missing input structure look like completed work.

The last line passes the return value from main into SystemExit. The scheduler receives status 1 when counting fails, while the responder gets a traceback explaining that failure. Keeping both outputs aligned is what prevents this job from looking successful to its caller after it has logged an error.

Logs and process status serve different readers
Logs and process status serve different readers. The CSV example logs a terminal exception once at the operation boundary and exits with status 1. A successful run prints the count and exits with status 0. Confirm that the scheduler preserves the process status.
The CSV example logs a terminal exception once at the operation boundary and exits with status 1. A successful run prints the count and exits with status 0. Confirm that the scheduler preserves the process status.
Read diagram description

The CSV example logs a terminal exception once at the operation boundary and exits with status 1. A successful run prints the count and exits with status 0. Confirm that the scheduler preserves the process status. Diagram labels: Read + validate CSV: Required service field; invalid input raises failure; Success: Count to stdout; INFO to stderr; Failure: One exception log to stderr; no result; Exit 0: Caller sees success; Exit 1: Caller sees failure.

A recoverable error can have different semantics. If a job is designed to finish useful partial work, define what its result and status mean for that case rather than failing every error or treating every zero exit as complete success. The caller needs to understand the same contract as the program.

From a local traceback to the collected record

When this pattern grows into a service, stable event names and bounded fields such as component, deployment version and outcome make records easier to interpret. Correlation IDs can connect related events without copying the entire request into each one. Credentials, raw support messages and customer payloads should stay out by default. Even a traceback can contain paths or sensitive exception detail, so access and retention deserve attention.

Local output is only the start of a production logging path. A collector must preserve timestamps and fields, expose delivery failures and handle backpressure when events arrive faster than it can forward them. A helpful local traceback is of little use to a responder who cannot find it at the destination. A searchable observability log strategy develops that broader event-to-investigation path.

The program provides a small, understandable check of the whole arrangement. Run it with a valid file, an empty file and a file missing the required column, then compare the count, log and exit status. After adding a scheduler and collector, repeat those cases at their actual destinations. Agreement across those observations tells you that a failed count remains a failed job all the way to the system and person responsible for responding.

Sources & context

Sources linked in this article. Read alongside the author’s analysis; a citation does not independently verify a publisher’s claims.

Report an error or outdated detail

Related reading

Explore a related question