Guide

Collecting and aggregating operational data with Python

In brief

A complete Python example combines request and error counts, while a failed regional collection explains why correct arithmetic can still produce an incomplete report.

5 min read

Sources
Two count bundles on a three-position manifest feed one open bowl while the third source position is empty.
Conceptual collection manifest: received counts can be combined, but an absent expected source must remain visible as missing. The image gives no measured fleet rate or actual collection result.
On this page2 sections

In the illustrative input below, two services report errors over the same interval. One has two failures in 100 requests; the other has nine in 900. Adding the counts gives 11 failures in 1,000 requests, or 1.1%. Averaging their individual percentages would give 1.5%, because it would give the small service as much influence as the large one.

The Python program below performs the count-based calculation and validates its input. That solves an important part of operational aggregation, but it leaves a second question for the collector: did all the expected data arrive? Separating collection, validation and calculation makes both responsibilities easier to understand and test. Starting with a file lets us examine the arithmetic without depending on a live API.

Why the combined error rate is 1.1%

The program reads a JSON array in which each row represents a distinct service over the same time window. The requests and errors values must be integers. Save the file as aggregate.py; supplying [{"requests": 100, "errors": 2}, {"requests": 900, "errors": 9}] produces 1,000 requests, 11 errors and an error fraction of 0.011.

import json
import sys

def aggregate(rows):
    if not isinstance(rows, list) or not rows:
        raise ValueError("Expected a nonempty array of count records")
    requests = errors = 0
    for row in rows:
        if not isinstance(row, dict):
            raise ValueError("Each record must be an object")
        n, e = row.get("requests"), row.get("errors")
        if type(n) is not int or type(e) is not int:
            raise ValueError("Counts must be integers")
        if n < 0 or e < 0 or e > n:
            raise ValueError("Require 0 <= errors <= requests")
        requests += n
        errors += e
    return {"requests": requests, "errors": errors,
            "error_fraction": errors / requests if requests else None}

def main():
    if len(sys.argv) != 2:
        print("Usage: python3 aggregate.py input.json", file=sys.stderr)
        return 2
    try:
        with open(sys.argv[1], encoding="utf-8") as stream:
            result = aggregate(json.load(stream))
    except (OSError, UnicodeError, ValueError) as exc:
        print(f"Aggregation failed: {exc}", file=sys.stderr)
        return 1
    print(json.dumps(result, allow_nan=False))
    return 0

if __name__ == "__main__":
    raise SystemExit(main())

The input checks happen before a row contributes to the totals. The function requires a nonempty list of objects, integer counts and a valid relationship between the counts: neither can be negative, and errors cannot exceed requests. It then sums errors and requests separately before dividing. This preserves the relative size of the populations rather than averaging their rates.

A total of zero requests has no meaningful error fraction in this calculation. The function returns None for that field instead of claiming a zero error rate. The command-line wrapper also distinguishes an input or validation failure from a successful result by reporting the failure separately and returning a nonzero status.

Now consider a hypothetical fleet collector whose request to one region fails. If it replaces the failure with an empty list, the remaining rows can all pass validation while the affected region disappears from the report. The arithmetic may be correct for the received population even as the apparent fleet error rate falls for the wrong reason.

The function cannot discover a source that was never supplied. A collection manifest has to provide that context: expected sources, received sources, their shared observation window and any failed pages. With that record, the caller can hold an incomplete fleet report or publish an explicitly narrower result. Validation establishes the shape and values of received rows; the manifest establishes what collection did and did not cover.

The surrounding collector needs to distinguish a failed region, a successful collection with no observations and a partial collection. This function deliberately rejects an empty list, so the caller has to represent a valid no-observations result separately. That preserves the difference between having nothing to count and failing to obtain the counts.

The rows from reachable regions may still help an investigation. Reporting them as a partial result keeps that information available without presenting it as the fleet total. Which regions replied is therefore part of the output’s meaning, just as requests and errors are.

Aggregate counts before dividing
Aggregate counts before dividing. Disjoint populations over the same window: 2 + 9 = 11 errors and 100 + 900 = 1,000 requests, so the combined error fraction is 1.1%. Averaging the two percentages would incorrectly produce 1.5%.
Disjoint populations over the same window: 2 + 9 = 11 errors and 100 + 900 = 1,000 requests, so the combined error fraction is 1.1%. Averaging the two percentages would incorrectly produce 1.5%.
Read diagram description

Disjoint populations over the same window: 2 + 9 = 11 errors and 100 + 900 = 1,000 requests, so the combined error fraction is 1.1%. Averaging the two percentages would incorrectly produce 1.5%. Diagram labels: Service A: 2 errors / 100 requests = 2%; Service B: 9 errors / 900 requests = 1%; Sum counts: 11 errors / 1,000 requests; Combined error fraction: 11 ÷ 1,000 = 0.011 = 1.1%.

An API client brings another set of failure states

Replacing the file reader adds authentication, connection and read timeouts, response-size limits, pagination and response-schema checks. Unexpected status codes and malformed responses need visible failure states. Retries with bounds and backoff give temporary trouble a chance to resolve without letting a scheduled collection continue indefinitely.

As the client progresses, record received sources and pages, then carry the collection window and freshness into the output. A failure on one page should remain identifiable when the publication step decides whether the result is complete enough for its intended use. Otherwise the final calculation loses the context the collector had when it encountered the problem.

The example also assumes that populations are disjoint and windows match. It does not deduplicate records, align time zones or make different service objectives comparable. Those are data-contract decisions to settle before reusing the arithmetic for a fleet. The guide to metrics and denominators explains why they change the meaning of a rate.

Test the resulting pipeline with a complete collection, a failed source and a valid source with no observations, following each case through to the report. The arithmetic test tells you whether 11 failures in 1,000 requests becomes 1.1%. The collection test tells you whether the reader can distinguish that measured result from a claim about a region the system never reached. A dependable report needs both answers.

Source context

This article does not include external reference links. Read it as the author’s perspective and evaluate the guidance against your environment.

Report an error or outdated detail

Related reading

Explore a related question