NVLink: the definition
NVLink is NVIDIA's interconnect technology for supported processor and accelerator configurations. Depending on the architecture, direct connections and switches provide high-bandwidth paths for communication, but do not automatically turn multiple GPUs into one universally transparent device.
The key points
- A GPU model name alone does not establish the deployed NVLink topology.
- Per-link, per-GPU and aggregate switch bandwidth are different measurement boundaries.
- Communication support does not remove software placement or per-device memory constraints.
- Scale-up links and a cluster's scale-out network solve related but different problems.
Why GPUs need to communicate
A workload can use more than one GPU by splitting data, model state or operations across devices. Those choices create communication: gradients may be combined, activations transferred or partial results exchanged. The time spent communicating can limit the benefit from extra arithmetic resources. NVLink addresses supported connections between processors and accelerators, while libraries and frameworks decide how the application uses them. [1][2]
For a hypothetical workload, adding a second GPU cuts arithmetic time from 100 milliseconds to 50 milliseconds but adds 40 milliseconds of non-overlapped communication. The resulting step takes 90 milliseconds, not 50. A faster link might reduce the added component, while a different parallelization strategy might reduce the amount of data exchanged. Both algorithm and interconnect matter.
Direct links and switched systems
NVLink capabilities depend on the supported GPU generation, package and system design. Some configurations use direct links between devices; others use switching to create a larger communication domain. NVIDIA describes NVLink Switch as part of its scale-up systems. A product supporting NVLink at the hardware level is not evidence that a rented node actually exposes the needed cables, bridges, switches or complete topology. [1]
A switch changes how endpoints can reach one another but does not make capacity infinite. Port counts, link rates, routing, shared resources and the supported system arrangement define the available paths. It is useful to ask for the actual topology and measured collective performance rather than accept an unqualified label such as NVLink enabled. [5]
Imagine four hypothetical GPUs arranged in two well-connected pairs, with a slower path between pairs. A two-GPU job placed inside one pair can behave differently from a four-GPU job spanning both. The aggregate GPU count is identical to that of a fully connected four-device configuration, but the communication opportunities are not. Placement is part of the supplied resource.
Reading bandwidth figures without mixing boundaries
An interconnect specification can describe one physical link, all links on one GPU, one switch or a whole system. It can report one-way throughput or add both directions. These figures are not interchangeable. NVLink generations have different published capabilities, so a numerical claim should include its architecture and measurement boundary rather than be treated as a permanent property of the brand name. [1]
Suppose a hypothetical GPU has four links, each providing 25 GB/s in one direction under an idealized specification. Summing the forward paths gives 100 GB/s; summing forward and reverse gives 200 GB/s. A one-way transfer from that GPU does not gain a 200 GB/s ceiling merely because a bidirectional total was quoted. Protocol behavior and the workload can reduce the useful rate further.
For another hypothetical calculation, transferring 8 GB over a path sustaining 200 GB/s takes at least 40 milliseconds. This excludes scheduling, synchronization and other computation. If the communication overlaps arithmetic, its contribution to step time can be smaller than its standalone duration. The useful measurement is the critical path of the application, not the sum of independently timed components when they run concurrently.
Does NVLink combine GPU memory?
Supported peer access lets devices communicate with memory on other devices under the platform's rules. That does not mean every program can treat all installed memory as one automatic, equally fast pool. Local and remote accesses have different paths. Frameworks still need an appropriate placement or parallelization strategy, and some state may remain replicated. Capacity must be evaluated at the level where an allocation or operation actually executes. [3][2]
Two hypothetical 80 GB devices offer 160 GB of installed memory in aggregate. A program allocating one 100 GB tensor on a single device will not necessarily succeed. A supported sharded representation can change the placement, but extra buffers, replication and communication must be included. The link enables useful cooperation; the software determines how the aggregate resources become usable.
Different parallel strategies need different traffic
Data-parallel training assigns different data to model replicas and synchronizes updates. Tensor parallelism partitions operations or their parameters across devices, often introducing communication within a layer. Pipeline approaches place different stages on different devices. These patterns have different communication sizes, frequencies and opportunities for overlap. A network that is sufficient for one pattern may limit another. [2][6]
Collective libraries such as NCCL implement operations including all-reduce, all-gather and reduce-scatter across devices. They select communication approaches based on topology and the operation. Their performance depends on more than the peak link rate: message size, device count, scheduling and surrounding computation matter. A tested collective is more informative than a topology label, but still does not fully replace the application benchmark. [4]
NVLink versus PCIe and scale-out networking
PCIe commonly connects endpoints to a host and can also carry supported peer traffic. NVLink supplies different supported high-bandwidth processor connections. Scale-out networks such as InfiniBand or Ethernet with RDMA connect nodes across a cluster. A complete system can use all of these. They should be understood as distinct paths in a hierarchy rather than competing names for one identical connection. [1][7]
A hypothetical eight-GPU node can have a strong internal fabric while connecting to the rest of the cluster through a comparatively constrained network. A job contained in one node may run well; a job spanning sixteen nodes may spend more time communicating. Internal connectivity does not establish inter-node bandwidth, and the number on a network adapter does not establish the internal GPU fabric.
What to verify in a compute offering
Ask for the exact accelerator form factor, number of devices, internal connections, switch topology, visible peer relationships and the network path between nodes. Establish whether the allocation is a whole device, a supported partition or a shared service, and whether the intended software can use the connections. Do not substitute a manufacturer's reference-system diagram for the provider's actual deployed configuration.
Measure the relevant device count and communication patterns with the intended framework. Include failure and restart behavior for long jobs, not just a short peak-throughput test. The central question is whether the complete configuration performs the workload reliably enough to justify its cost. NVLink can be a critical ingredient in that answer, but it is not the answer by itself.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Does an NVLink-capable model guarantee an NVLink-connected rental?
No. The deployed form factor, bridges, switches, topology and allocation must be verified separately.
2. Why can a good single-node benchmark mislead a multi-node buyer?
The larger job introduces scale-out communication, different collective behavior and more synchronization. The internal fabric does not establish the inter-node network.
3. Can aggregate GPU memory be treated as one allocation automatically?
Not generally. Software placement, supported peer access, local limits and replicated state determine how aggregate memory can be used.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · NVLink and NVLink Switch ↗ (opens in a new tab)Manufacturer documentation · Checked 28 September 2026
- PyTorch · Distributed Data Parallel ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
- NVIDIA · CUDA programming model ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · NCCL collective operations ↗ (opens in a new tab)Library documentation · Checked 28 September 2026
- NVIDIA · NCCL troubleshooting and topology ↗ (opens in a new tab)Library documentation · Checked 28 September 2026
- Hugging Face · Parallelism methods ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- NVIDIA · GPUDirect RDMA ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.