Recovering an agent task after OpenShell containment
OpenShell can stop an agent after Kubernetes has accepted its scaling request. Recovery then depends on the Deployment state, available replicas and customer-facing result.
Guides, news and explanations about AI and reliable systems.
OpenShell can stop an agent after Kubernetes has accepted its scaling request. Recovery then depends on the Deployment state, available replicas and customer-facing result.
NVIDIA’s Cluster Readiness Engine runs diagnostic, communication and training workloads on selected nodes. Workload-specific thresholds turn those results into an admission decision.
NodeWright coordinates package, configuration and tuning changes across Kubernetes hosts. Progressive rollout and package verification organize the update one selected node group at a time.
NVIDIA reports generation falling from 156.6 to 34.2 seconds on its example workload. The eight-device layout also changes how much capacity remains for other requests.
SWE-Serve checks coding-agent patches through a running SGLang server. In NVIDIA’s experiment, adding live-serving checks reduced the pass rate for the same patches.
An eight-B200 test retained 96.1% to 98.2% of baseline output-token throughput with confidential computing enabled. Its software and workload define what that comparison covers.
Topograph publishes discovered topology as Kubernetes labels or Slurm configuration. Placement policies can use that map to trade a shorter communication path against a longer queue.
The alliance supports storage, generation and flexible AI workloads as ways to change grid demand. For a scheduled evaluation, a safe pause can still miss the deadline the job serves.
Model loading and checkpoint recovery stress storage differently. Reference configurations help explain which workload and failure conditions their results cover.