SWE-Serve checks coding-agent patches through a running SGLang server. In NVIDIA’s experiment, adding live-serving checks reduced the pass rate for the same patches.
An eight-B200 test retained 96.1% to 98.2% of baseline output-token throughput with confidential computing enabled. Its software and workload define what that comparison covers.
Topograph publishes discovered topology as Kubernetes labels or Slurm configuration. Placement policies can use that map to trade a shorter communication path against a longer queue.
The alliance supports storage, generation and flexible AI workloads as ways to change grid demand. For a scheduled evaluation, a safe pause can still miss the deadline the job serves.
Kubernetes attributes connect agent calls to pods, workloads and paging routes. A processor migration can break those connections while telemetry still flows.
Host-memory throttling can keep an inference process alive while its requests slow down. Effective kubelet settings and request deadlines explain the tradeoff.