NVIDIA measures the overhead of confidential inference
An eight-B200 test retained 96.1% to 98.2% of baseline output-token throughput with confidential computing enabled. Its software and workload define what that comparison covers.
Guides, news and explanations about AI and reliable systems.
An eight-B200 test retained 96.1% to 98.2% of baseline output-token throughput with confidential computing enabled. Its software and workload define what that comparison covers.