Serving World Models: Latency vs Throughput with NVIDIA Cosmos 3 Super on Nebius | Cosmos Labs
Serving world models for video generation requires a different balance of latency and throughput than serving large language models. In this livestream, Nebius Physical AI team members Stewart Tong and Timothy Le examine NVIDIA Cosmos 3 Super benchmarks through vLLM-Omni on NVIDIA HGX B200 and HGX H200 systems. Their work compares four serving topologies, ranging from one eight-GPU replica to eight single-GPU replicas. See how replica count, parallelism, and concurrency affect request latency a