PCIe: the definition
PCI Express is a serial interconnect standard used to connect components such as GPUs, network adapters and storage controllers within a computer. Its usable performance depends on the link generation, active lane width, protocol overhead and complete path.
The key points
- A physical slot's size does not guarantee its electrical lane allocation or operating generation.
- GT/s describes transfers, not directly application bytes per second.
- One-way bandwidth and summed bidirectional bandwidth are different figures.
- Several fast endpoints can still share a narrower upstream path.
What PCI Express connects
PCI Express, abbreviated PCIe, is a connection standard for components inside a computing system. A discrete GPU, an NVMe storage device and a high-speed network adapter may all use it to communicate with the host. It defines more than a connector: physical signaling, link behavior and transactions are part of the standard. Device form factor and the protocol carried through a connection are therefore separate concepts. [1][2]
A server's PCIe topology usually has a host-side root complex, links, and endpoints, with switches where needed. A switch can connect multiple downstream devices to an upstream path. The device sees a transaction path rather than one universal shared memory bucket. Topology matters when several devices transfer data simultaneously or when peer-to-peer traffic must cross particular boundaries. [3][4]
Lanes and negotiated width
A PCIe lane provides signaling in both directions. Multiple lanes form a link, commonly described by widths such as x4, x8 or x16. Devices and the host negotiate an operating configuration subject to their capabilities and platform wiring. A mechanically long slot can have fewer electrical lanes than its shape suggests. A card's advertised maximum is not proof of the link it receives in a specific server. [5]
As a hypothetical example, an x16-capable GPU installed in a slot wired for eight lanes cannot use sixteen host lanes through that slot. If the negotiated generation is also lower than the device's maximum, both factors change the link's data-rate ceiling. The configuration should be inspected in the running system, not inferred solely from the card's product sheet.
Bifurcation divides a supported host lane group into several smaller links. It requires appropriate platform and firmware support, as well as compatible physical routing. It is not the same operation as placing a switch downstream. A passive adapter cannot create additional lanes or make an unsupported lane split appear in the host. [6]
Generations and transfer-rate units
PCIe generations define signaling and protocol capabilities. A transfer rate in gigatransfers per second is not automatically an equal number of gigabytes per second. The relationship depends on encoding, link width and protocol overhead. PCIe 5.0, for example, specifies 32 GT/s per lane. Later-generation signaling changes mean that a formula valid for one generation should not be blindly reused for every generation. [7][8]
For a PCIe 5.0 x16 link, using 128b/130b encoding, the simplified one-way data-rate ceiling before packet overhead is 32 billion transfers per second × 128/130 × 16 lanes ÷ 8 bits per byte. That is approximately 63.0 GB/s in decimal units. Summing both directions gives approximately twice that value, but it does not double the rate of a single one-way transfer. The application payload rate will be lower.
Suppose a hypothetical workload transfers 30 GB over a path sustaining 20 GB/s. The transfer alone needs at least 1.5 seconds. An advertised bidirectional total of 80 GB/s elsewhere in the system does not change that lower bound. Measurements must name direction, payload and the actual path to be useful.
Why the narrowest shared path matters
A PCIe switch can expose multiple device-facing ports while sharing an upstream connection. Aggregate demand can exceed the upstream path's capacity. Peer-to-peer operations may avoid the CPU memory path in supported configurations, but support depends on topology, access control and platform behavior. NVIDIA's GPUDirect documentation describes these constraints rather than promising that every pair of devices can communicate equally well. [4][3]
Imagine four hypothetical storage devices, each able to deliver 7 GB/s, attached behind an upstream path sustaining 14 GB/s. Reading one device may achieve its individual rate. Reading all four simultaneously cannot yield 28 GB/s through that same 14 GB/s bottleneck. The per-device specification remains true while the aggregate expectation is wrong.
On a NUMA system, a device can be closer to one CPU or memory node than another. Traffic reaching remote memory may cross another interconnect. Placing data-loader threads and buffers near the relevant device can change the observed path. Core count and total RAM do not reveal these relationships; the server topology and workload placement do. [9]
PCIe versus GPU interconnects
PCIe is widely used for host and peripheral connectivity. NVLink is a separate family of supported NVIDIA interconnects used for particular processor and accelerator relationships. A system can contain both. Their existence does not make them interchangeable, and a GPU family name does not establish whether a specific deployed configuration has NVLink connections or switches. [10][4]
Data remaining inside GPU memory may not continuously traverse PCIe during a kernel. In that situation, GPU memory bandwidth can matter much more than host-link bandwidth. Repeatedly transferring large inputs and outputs can produce the opposite result. The first step is to identify which transfers occur and whether they overlap useful work, not assume that the fastest host link always changes the application most. [11]
PCIe and NVMe are not synonyms
NVMe defines a storage command architecture and associated transports. Many local NVMe devices use PCIe, but PCIe also serves GPUs and network adapters, and NVMe can operate over fabrics. An M.2 form factor does not itself establish the storage protocol, endurance or sustained performance. A storage comparison needs controller, media, interface and workload information. [2]
A hypothetical SSD can use a fast PCIe link while its flash media sustains much less than that link's maximum. Likewise, small random requests can be limited by latency or queueing rather than sequential bandwidth. The transport ceiling is one constraint in a complete storage system, not a promise that every access reaches it.
A practical inspection checklist
Record the negotiated speed and width, endpoint placement, upstream sharing, CPU and memory locality, and the device's actual transfer pattern. Then benchmark directionally meaningful workloads, including concurrent transfers when the application uses them. Keep units consistent and distinguish physical connectors from electrical capabilities. These checks make a specification actionable without assuming that every x16 slot behaves identically.
The key idea is a path rather than a badge. PCIe connects devices through a topology whose links and shared resources have limits. Understanding that path explains many otherwise surprising differences between servers carrying the same GPU, network card or SSD model.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Can an x16-sized slot operate with only eight lanes?
Yes. Mechanical size, electrical wiring and negotiated width are separate. Inspect the platform and running link.
2. Why is bidirectional bandwidth not the speed of a one-way copy?
It adds the capacities of two directions. A copy going in only one direction cannot consume the reverse direction's capacity as extra forward bandwidth.
3. Why can several individually fast devices run more slowly together?
They may share a narrower upstream link, memory path or other resource. Aggregate demand must be evaluated against the complete topology.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- PCI-SIG · PCI Express specifications ↗ (opens in a new tab)Standards-body documentation · Checked 28 September 2026
- NVM Express · PCIe and NVMe terminology ↗ (opens in a new tab)Standards-body explanation · Checked 28 September 2026
- PCI-SIG · PCIe 5.0 switch listing ↗ (opens in a new tab)Compliance product record · Checked 28 September 2026
- NVIDIA · GPUDirect RDMA ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- Intel · What is PCIe 4.0 and 5.0? ↗ (opens in a new tab)Manufacturer explanation · Checked 28 September 2026
- ASUS · PCIe bifurcation support ↗ (opens in a new tab)Platform support documentation · Checked 28 September 2026
- PCI-SIG · PCIe compliance updates ↗ (opens in a new tab)Standards-body explanation · Checked 28 September 2026
- PCI-SIG · PCI Express 6.0 specification ↗ (opens in a new tab)Standards-body documentation · Checked 28 September 2026
- Linux Kernel · NUMA ↗ (opens in a new tab)Operating-system documentation · Checked 28 September 2026
- NVIDIA · NVLink and NVLink Switch ↗ (opens in a new tab)Manufacturer documentation · Checked 28 September 2026
- NVIDIA · CUDA Best Practices Guide ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.