Opus 5.5 adds conversation-replay checks
New model documentation makes retained instructions and thinking-block compatibility part of the compaction migration review.
Original report: Claude adds on-demand compaction for long-running agentsAIOps / Observability / Infrastructure
Developments in AI and the infrastructure behind it. What changed, how it works, and what it means for the people running it.
Our editorial standardsDevelopments since our original reporting, with dated updates and supporting sources.
New model documentation makes retained instructions and thinking-block compatibility part of the compaction migration review.
Original report: Claude adds on-demand compaction for long-running agentsReleases, research & operational impact
AI infrastructureNews analysis
NVIDIA introduces workload-driven GPU cluster validation. Choose the acceptance claim and thresholds before interpreting a completed run as readiness.
Read the analysisKubernetes operationsNews analysis
NodeWright brings Kubernetes-aware host changes into a declarative workflow. Check node selection, interruption limits and recovery before a fleet rollout.
AI modelsNews analysis
Anthropic releases Opus 5.5 with lower token prices. Evaluate operational exceptions and integration behavior before changing an agent workflow.
Inference infrastructureNews analysis
NVIDIA’s September 21 demonstration accelerates one video-generation workload across eight GPUs. A service decision still needs the same offered load and completed-work count.
Agent evaluationNews analysis
NVIDIA’s new SGLang benchmark finds patches that pass local checks but fail a running server. Its result is a reason to examine what a green test actually covers.
Local AI agentsNews analysis
Google adds local Gemma and LiteRT workflows. Running the model on a workstation still leaves a separate decision about what its tools may read, change or send.
AI inferenceNews analysis
NVIDIA reports low confidential-computing overhead on an eight-GPU test. Separate that measured performance result from your own security and recovery claims.
AI infrastructureNews analysis
New Topograph guidance connects physical network topology to workload schedulers. The operational check is whether the scheduler’s view survives cluster change.
AI model operationsNews analysis
Lower model prices and new caching controls make agent pilots cheaper to run. Measure accepted work before expanding the workload.
AI infrastructureNews analysis
Google, NVIDIA and Emerald AI are backing flexible data centers. Operators will need to define which AI work can yield power, and how it recovers.
Agent infrastructureNews analysis
The Messages API can now summarize conversation history when an application chooses. SRE teams should test which operational constraints survive.
AI agent governanceNews analysis
Anthropic has added Chrome-session transcripts to its Compliance API beta. Responders gain another evidence source, with important gaps to understand.
AI operationsNews analysis
Splunk’s September release connects agent evaluation and cost monitoring. The useful test is whether a responder can trace spending back to an outcome.
AI observabilityNews analysis
The stable Kubernetes attributes processor helps connect AI-agent traces to their workloads. Check the metadata joins before upgrading.
AI infrastructureNews analysis
Memory QoS reaches beta in Kubernetes 1.37. For AI inference teams, the rollout question is how host-memory policy affects serving latency and neighboring workers.