NVIDIA L40S: the definition
The NVIDIA L40S is a PCIe data center GPU based on the Ada Lovelace architecture, with 48 GB of GDDR6 memory, FP8 Tensor Cores and ray-tracing cores. NVIDIA positions it as a general-purpose GPU for AI inference, smaller-scale training and fine-tuning, 3D graphics and video.
The key points
- The L40S is an Ada Lovelace PCIe card with 48 GB of GDDR6, 864 GB/s of bandwidth and a 350 W power limit.
- Its fourth-generation Tensor Cores support FP8, and it keeps RT cores and video engines for graphics and media work.
- Compared with the A100 and H100 it offers competitive peak compute but much lower memory bandwidth and no NVLink.
- It fits best for inference of models that fit in 48 GB, fine-tuning of smaller models, rendering and video, rather than large-scale training.
What the L40S is
NVIDIA introduced the L40S on August 8, 2023, at the same time as a new generation of its OVX servers, which hold up to eight L40S GPUs. NVIDIA described it as a universal data center GPU for generative AI, graphics-intensive applications and video, and said systems would be available from the fall of 2023 from makers including ASUS, Dell Technologies, GIGABYTE, HPE, Lenovo, QCT and Supermicro. [2]
The L40S uses NVIDIA's Ada Lovelace architecture rather than the Hopper architecture used in the H100. That shows in its design: it keeps display outputs, ray-tracing cores and video encoders from NVIDIA's visualization line, and uses GDDR6 memory rather than the more expensive HBM used on NVIDIA's top data center GPUs. [1][4]
Key specifications
The L40S has 18,176 CUDA cores, 568 fourth-generation Tensor Cores and 142 third-generation RT cores. It carries 48 GB of GDDR6 with ECC and 864 GB/s of memory bandwidth, and connects to the host over PCIe Gen4 x16 at 64 GB/s bidirectional. It is a dual-slot card, 4.4 inches high and 10.5 inches long, with a passive heatsink that relies on server airflow, a 16-pin power connector and a maximum power of 350 W. [1][3]
For media and graphics it has three NVENC encoders and three NVDEC decoders with AV1 support, four DisplayPort 1.4a outputs and support for NVIDIA virtual GPU software. It supports secure boot and is NEBS Level 3 ready for telecom environments. It does not support NVLink or Multi-Instance GPU. [1][3]
The L40S is closely related to the earlier L40. Both use Ada Lovelace with 48 GB of GDDR6 and 864 GB/s, but the L40S raises the power limit from 300 W to 350 W, which lets it sustain higher clocks. ServeTheHome described the power increase as the main difference between the two cards. [4]
AI performance and FP8
NVIDIA rates the L40S at 91.6 teraFLOPS of FP32, 366 teraFLOPS of TF32, 733 teraFLOPS of FP16 or BF16 and 1,466 teraFLOPS of FP8 on its Tensor Cores. The Tensor Core figures include structured sparsity, so dense throughput is half: roughly 183 teraFLOPS TF32, 366 teraFLOPS FP16 or BF16 and 733 teraFLOPS FP8. Like Hopper, its Tensor Cores support FP8 through NVIDIA's Transformer Engine. [1][2]
At launch NVIDIA claimed the L40S delivers up to 1.2 times the generative AI inference performance of the A100, up to 1.7 times its training performance and nearly 5 times its FP32 performance. Lenovo's product guide makes the training comparison explicit: eight L40S GPUs in a mainstream server against an eight-GPU HGX A100 system. NVIDIA also claims up to 5 times the inference performance of the previous-generation A40. [2][3][1]
Peak Tensor Core figures overstate what the L40S achieves on many language model workloads. Generating tokens from a large model is usually limited by memory bandwidth, because each step reads the model weights from memory, and the L40S has less than half the bandwidth of an A100 and about a quarter of an H100 SXM. Its compute advantage shows most in compute-bound work such as image generation, prompt processing and batch inference of smaller models.
L40S versus A100 and H100
Memory is the largest difference. The A100 80GB has 80 GB of HBM2e with 1,935 GB/s on the PCIe card and 2,039 GB/s on SXM. The H100 SXM has 80 GB of HBM3 at 3.35 TB/s. The L40S has 48 GB of GDDR6 at 864 GB/s. [5][6][1]
On dense 16-bit Tensor throughput the L40S, at about 366 teraFLOPS, is ahead of the A100 at 312 teraFLOPS but well behind the H100 SXM at about 990 teraFLOPS (1,979 teraFLOPS with sparsity). The L40S and H100 both support FP8; the A100 does not. On power, the L40S is rated at 350 W, the A100 at 300 W for PCIe and 400 W for SXM, and the H100 SXM at up to 700 W. [1][5][6]
Multi-GPU scaling is where the L40S falls furthest behind. The A100 offers 600 GB/s of NVLink and the H100 SXM 900 GB/s, while L40S GPUs exchange data over PCIe Gen4 at 64 GB/s. The A100 and H100 also support MIG partitioning into as many as seven instances; the L40S does not. [5][6][1]
A model with 30 billion parameters in FP8 needs about 30 GB for its weights, which fits on one L40S with room for a moderate key-value cache. A 70-billion-parameter model in FP8 needs about 70 GB, so it would have to be split across two L40S cards over PCIe, while it fits on a single 80 GB H100 only with very little room left, or comfortably on a larger GPU such as the H200.
Where the L40S fits
NVIDIA's own positioning covers generative AI inference and training, 3D graphics and rendering, NVIDIA Omniverse workloads and video processing. The combination of RT cores, AV1 video engines, display outputs and Tensor Cores in one card is what makes it suitable for mixed graphics and AI deployments. [2][1]
In our view the L40S is a sensible choice for serving small and mid-sized language models, image and video generation, parameter-efficient fine-tuning such as LoRA on models that fit in 48 GB, and mixed pipelines that combine rendering with AI. It is a poor fit for pre-training or full fine-tuning of large models, where memory bandwidth, capacity and NVLink matter more than peak FLOPS.
The L40S in 2026 and its successor
NVIDIA's Blackwell-generation successor in this segment is the RTX PRO 6000 Blackwell Server Edition. It doubles memory to 96 GB of GDDR7 with 1,597 GB/s of bandwidth, moves to PCIe Gen5, adds MIG with up to four instances and FP4 support, and has a configurable power limit of up to 600 W. NVIDIA describes it as a large performance step over the L40S for generative and multimodal AI. [7]
As of September 2026 the L40S is a previous-generation part, but that does not make it irrelevant for renters. Older GPUs tend to become cheaper as newer ones arrive, and for workloads that fit in 48 GB and are not bandwidth-bound, a lower rate can outweigh the newer card's speed.
Renting an L40S today
L40S capacity can typically be rented as single GPUs or small multi-GPU servers, which suits inference endpoints and fine-tuning jobs that do not need a full eight-GPU HGX node. When comparing offers, check the number of GPUs per instance, the CPU and system memory that come with each GPU, and whether the workload is limited by memory bandwidth. On Kovara, the L40S 48GB page shows current listings, the GPU prices view compares providers, and Compare lets you set the L40S against an A100, H100 or newer option for the model you plan to run.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1Is the L40S faster than the A100?
It depends on the workload. The L40S has higher peak dense FP16 and BF16 Tensor throughput than the A100 (about 366 versus 312 teraFLOPS) and adds FP8, which the A100 lacks. The A100 has far more memory bandwidth (about 2 TB/s versus 864 GB/s), more memory in its 80 GB version, and NVLink, so memory-bound and multi-GPU workloads often run better on the A100.
2Can the L40S be used for training?
Yes, for smaller models and fine-tuning. NVIDIA claimed an eight-GPU L40S server delivers up to 1.7 times the training performance of an eight-GPU HGX A100. The lack of NVLink and the 48 GB memory limit make it a weak choice for large distributed training.
3Does the L40S support NVLink or MIG?
No. The L40S connects to the host and to other GPUs only over PCIe Gen4 x16, and it does not support Multi-Instance GPU partitioning. It does support NVIDIA virtual GPU software.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · L40S GPU for AI and Graphics Performance ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA Newsroom · NVIDIA, Global Data Center System Manufacturers to Supercharge Generative AI and Industrial Digitalization ↗ (opens in a new tab)Press release · Checked 29 September 2026
- Lenovo Press · ThinkSystem NVIDIA L40S 48GB PCIe Gen4 Passive GPU Product Guide ↗ (opens in a new tab)Product guide · Checked 29 September 2026
- ServeTheHome · NVIDIA L40S GPU for Data Center Visualization Launched ↗ (opens in a new tab)News report · Checked 29 September 2026
- NVIDIA · A100 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · H100 Tensor Core GPU ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
- NVIDIA · RTX PRO 6000 Blackwell Server Edition ↗ (opens in a new tab)Manufacturer documentation · Checked 29 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.
