NVMe: the definition
Non-Volatile Memory Express is a storage architecture and command family for communicating with nonvolatile storage through supported transports. It describes how storage is accessed, not the physical flash technology or a guaranteed application speed.
The key points
- NVMe is not the same thing as PCIe, an SSD form factor or NAND flash.
- IOPS, bandwidth and latency describe different aspects of a workload.
- Queue depth can raise throughput while also increasing waiting time.
- A namespace is a logical address space; it is not automatically a separate physical device.
Four layers that are often confused
Storage discussions commonly mix media, controller, protocol and physical connection. Flash cells store data. A controller manages access to that media. NVMe supplies a command architecture. PCIe is a common transport for locally attached NVMe devices. A physical shape such as M.2 is another property. Knowing one of these labels does not determine all the others, and a connector's appearance does not establish sustained performance. [1][2]
A hypothetical SSD can have a PCIe interface capable of more traffic than its media sustains. Another device can share the same interface generation while offering different endurance or latency. Calling both NVMe SSDs correctly identifies a broad class, but does not make them equivalent. Compare the complete configuration and workload rather than assume the protocol name is a speed rating.
Commands and asynchronous queues
NVMe uses submission and completion mechanisms that allow software to have multiple operations in flight. The host submits commands and later observes their completion. Implementations can associate queues with processing resources to reduce coordination overhead. SPDK's NVMe driver illustrates an asynchronous, user-space implementation; operating-system drivers use their own integration approach. The same protocol can therefore participate in different software paths. [3]
Submission is not completion. An application that measures only the time to enqueue writes has not measured how long the storage takes to finish them. It must also define what completion means for its durability requirement. Buffers must remain valid for the operation's lifetime, and returned error conditions must be handled. A fast benchmark that ignores completion or failures does not establish usable throughput. [3]
Imagine a hypothetical warehouse receiving many delivery orders. Sending orders faster can keep workers busy, but it does not make the loading bay infinitely wide. Queues absorb outstanding work and help expose parallelism. Once the service capacity is saturated, adding more orders mainly increases waiting. This analogy explains why queue depth is an experimental setting, not an unconditional performance improvement.
IOPS, bandwidth and latency
IOPS counts completed input/output operations per second. Bandwidth measures bytes transferred per second. Latency measures time per operation, ideally with a distribution rather than only an average. Request size connects IOPS to bandwidth arithmetically, but performance at one size does not determine another. Sequential and random access also exercise a storage system differently. [2]
At a hypothetical 100,000 completed operations per second with 4 KiB per operation, payload bandwidth is 409.6 million bytes per second, about 0.410 GB/s. At the same operation rate with 1 MiB requests, the multiplication produces over 100 GB/s, which may exceed the actual device or link. The operation rate cannot be assumed constant while request size changes dramatically.
For an illustrative steady system, average outstanding operations equal throughput multiplied by average response time. A workload completing 100,000 operations per second at 100 microseconds average latency has about ten operations in flight on average. This relationship helps sanity-check measurements, but it does not specify the tail latency, queue policy or achievable rate of another workload.
Controllers and namespaces
An NVMe namespace is a collection of logical block addresses exposed to host software. A namespace identifier distinguishes it within the relevant controller context. Splitting a device into namespaces does not necessarily partition its flash chips into physically independent devices. The NVM Express organization explicitly distinguishes logical address isolation from physical block isolation. [4]
This matters when interpreting multi-tenant claims. Two logical namespaces can still share controller resources, media or failure domains. Separate names alone do not prove independent bandwidth or independent failure behavior. Isolation, encryption and quality-of-service capabilities must be checked for the actual implementation rather than inferred from the existence of multiple logical devices.
Persistence is a separate requirement
Nonvolatile media retains stored information without ordinary operating power, but an application write can pass through volatile buffers and several software layers before reaching the intended persistence boundary. A database's durability contract involves flushes, ordering and recovery behavior in addition to the storage protocol. The existence of an SSD is not proof that every acknowledged application operation survives every failure. [3]
Suppose a hypothetical training job writes a checkpoint into local storage that is erased when the rented machine is terminated. The medium can be nonvolatile while the service is still ephemeral from the customer's perspective. Hardware persistence, cloud resource lifecycle and disaster recovery are three different questions. Place important recovery state where the documented service behavior matches the requirement.
We recommend defining the failure being protected against: process crash, operating-system failure, host loss, accidental deletion or site loss. Then test recovery for that case. Replication and backups can serve different purposes. A storage benchmark measures performance; a restore test provides evidence about recoverability.
NVMe over Fabrics
NVMe is not confined to a local PCIe device. NVMe over Fabrics extends its use across supported network transports, allowing access to remote storage targets. SPDK supplies implementations for such configurations. The remote path adds network and target-system considerations, so a local-device benchmark should not be treated as a guarantee for remote storage bearing the same command-family name. [2]
A hypothetical remote storage service may contain fast SSDs but share a constrained network uplink among many clients. Alternatively, a well-designed shared service may provide better aggregate performance and durability than one local disk. The correct comparison includes client concurrency, network path, service-level limits and workload latency—not just the media installed behind the service.
What AI workloads ask from storage
An input pipeline can read many small files, large packed datasets or sequential shards. Checkpointing can create large bursts of writes. Metadata and directory operations may dominate some access patterns even when advertised sequential bandwidth is high. Start by measuring the application's actual requests and identifying which layer is limiting delivery to the GPU.
If a hypothetical dataset is 2 TB and a complete input path sustains 4 GB/s, reading it once requires at least 500 seconds in decimal units. That estimate excludes decoding, preprocessing, repeated epochs and contention. Caching can change which storage tier supplies later reads, so report whether a benchmark measures cold storage or a warmed cache.
A meaningful storage specification
Record capacity units, local versus remote attachment, request size, read/write mix, queue depth, concurrency and latency percentiles. Include endurance and service lifecycle where relevant, and state whether tests ran long enough to move beyond short-lived caching behavior. A single peak IOPS or bandwidth number is not an all-purpose description of storage quality.
The lasting concept is separation of layers. NVMe defines a way to request storage operations; media, controllers, transports, filesystems and applications determine the completed workflow. Understanding those boundaries makes both performance claims and failure assumptions easier to challenge constructively.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Is NVMe another name for PCIe?
No. NVMe describes a storage architecture and commands; PCIe is one transport used by local devices. NVMe can also operate across supported fabrics.
2. Why can increasing queue depth worsen latency?
Once service capacity is saturated, more outstanding work can mainly increase waiting instead of increasing completed throughput.
3. Does a namespace guarantee a physically separate SSD?
No. It is a logical address space. Resource sharing and physical failure isolation depend on the implementation.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVM Express · PCIe and NVMe terminology ↗ (opens in a new tab)Standards-body explanation · Checked 28 September 2026
- SPDK · Storage architecture overview ↗ (opens in a new tab)Project documentation · Checked 28 September 2026
- SPDK · NVMe driver ↗ (opens in a new tab)Driver documentation · Checked 28 September 2026
- NVM Express · NVMe namespaces ↗ (opens in a new tab)Standards-body explanation · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.