The key points
- GPU count does not describe the communication paths between devices.
- Scale-up and scale-out networks solve related but different problems.
- Measure end-to-end workload performance at the intended cluster size.
Eight GPUs is a quantity, not an architecture
Imagine two offers with the same accelerator model and the same device count. One suits a workload that runs independently on each device. The other is needed for a tightly coupled job that repeatedly exchanges intermediate results. Our starting point is that these are different requirements, even before comparing a price.
NVIDIA distinguishes scale-up networking, which connects accelerators inside a closely integrated domain, from scale-out networking between systems. NVLink is an example of a scale-up interconnect; InfiniBand and Ethernet are used for scale-out designs. A scale-up domain can span a rack, so “inside one server” is not a universal definition. [1]
The practical implication is to describe the actual boundary. Which GPUs share a high-bandwidth fabric? Which transfers cross network adapters and switches? A product label or a port speed is a starting point for that conversation, not the answer.
Why communication becomes part of execution time
Distributed programs use collective operations to coordinate data across devices. NCCL documents AllReduce as combining values across ranks and returning the result to each rank. AllGather distributes gathered inputs, while AlltoAll exchanges different data between ranks. These are different communication patterns, not one generic network transfer. [2]
When a step depends on exchanged information, communication can sit on the path to completing that step. Some communication can overlap computation, but that is an implementation property to measure. Buying more accelerators does not prove that the exchange will be hidden. [1]
For illustration, suppose a job takes 10 hours on four GPUs and 7 hours on eight, with identical hourly prices per device. Elapsed time improves, but usage rises from 40 to 56 GPU-hours. That can still be worthwhile for a deadline. It simply is not a cost saving. These numbers are hypothetical and do not describe a measured cluster.
Keep the different bandwidth numbers separate
Memory bandwidth, interconnect bandwidth and an external network connection describe different data paths. NVIDIA’s GPU architecture documentation separates memory movement from arithmetic execution. Its networking discussion adds domain topology and communication behaviour to the system-level picture. A large number on one path cannot establish the performance of another. [3][1]
A comparison should identify whether a rate is per port, per GPU, per server or an aggregate, and whether it counts one direction or both. Similarly, an internet-facing connection should not be mistaken for the fabric used between the GPUs assigned to the workload. These are the specifications we would ask a supplier to clarify before comparing two cluster offers.
Topology matters because the program’s communication pattern must operate over the connections actually provided. A configuration useful for many independent inference workers should not automatically be recommended for a single distributed job. The reverse is also true: a sophisticated interconnect may offer little benefit to a workload that rarely uses it.
A useful test begins with the intended job
Ask for the exact node layout, devices per node, fabric technology, sharing assumptions and software versions. Then test at the intended GPU count, with the model and parallelism configuration you plan to run. Record successful job duration, communication time where measurable, and the total billable resource usage.
A collective microbenchmark can help explain one part of the system. It is not a substitute for an application benchmark. NCCL’s collective definitions describe the communication operation; they do not promise the application’s final performance. [2]
Our view is that a cluster should be purchased as a working configuration. The question is not whether the component list looks impressive, but whether the connections let those components complete useful work together. No provider ranking or measured speedup is claimed here.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · NVLink: The Scale-Up Network for AI Factories ↗ (opens in a new tab)Manufacturer technical article · Checked 27 September 2026
- NVIDIA · NCCL collective operations ↗ (opens in a new tab)Manufacturer documentation · Checked 27 September 2026
- NVIDIA · GPU Performance Background ↗ (opens in a new tab)Manufacturer documentation · Checked 27 September 2026
Prepared with AI assistance. Publication approved by Tommaso Luci. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.