# Inference maintenance pilot

This is a blank planning record, not authorization to drain a node.

## Scope and authority
- Service, namespace, owner:
- Approved staging node and node selector:
- Deployment and model versions:
- PDB selector, maxUnavailable or minAvailable, observedGeneration:
- Other maintenance or application rollouts excluded:
- Authorized executor and stop/recovery owner:

## Capacity and acceptance
- Offered request rate and request/output size distribution:
- Baseline completions, latency, rejections, queue age:
- Ready replicas and node placement:
- Replacement GPU availability and scheduling constraints:
- Measured scheduling, loading, warmup and first useful response times:
- Stop thresholds and maximum observation window:
- Temporary capacity, traffic and spending ceilings:

## Execution and recovery
- Approved command and start/end timestamps:
- Evicted pods, replacement placements and unresolved states:
- After-change completions, latency, rejections, queue age:
- Host safe to return? Evidence and approving owner:
- Uncordon or replacement action, with verification:
- Decision: proceed within tested scope / hold / change capacity plan:
- Untested workload, topology or demand conditions:
