Tensor cores: the definition
Tensor cores are NVIDIA execution resources specialized for supported matrix operations. They accelerate particular calculations inside a GPU; they are not complete processors, memory devices or a guarantee that every model operation runs faster.
The key points
- Matrix arithmetic combines many multiply-and-accumulate operations in a regular structure.
- Supported numerical formats and operation shapes differ between architectures.
- Input precision, accumulation precision and application quality are different questions.
- A sparse or low-precision theoretical peak must not be compared directly with an unrelated dense or high-precision result.
Why matrix multiplication matters
A matrix is a two-dimensional array of values. Multiplying an M-by-K matrix by a K-by-N matrix produces an M-by-N result. Each output combines corresponding inputs through multiplication and addition. Large matrices expose many operations that can be organized in parallel, with repeated reuse of inputs. Neural networks use this structure in linear layers and other operations, making efficient matrix arithmetic an important part of accelerator design. [1]
For an illustrative two-by-two multiplication, output row one, column one combines the first input row with the first column of the other input. Other output positions use different combinations. The full problem has both independence between output elements and reuse of input values. Real implementations divide much larger matrices into tiles that match memory and execution resources rather than treating every multiplication as a completely isolated instruction.
What the specialized hardware does
Tensor cores execute supported matrix operations using specialized hardware inside a GPU. They coexist with other execution resources that handle general arithmetic and control. A matrix operation can perform much more arithmetic per instruction than a scalar operation, but the hardware still needs its inputs delivered and results stored. A tensor-core count is consequently not comparable to a count of full CPU cores. [1][2]
The name does not mean that tensor cores execute every possible operation on a tensor. A tensor is a software data structure with shape and representation; not every transformation of that structure maps to a specialized matrix instruction. Indexing, reductions, transfers, synchronization and irregular operations can use other resources or become bottlenecks. Profiling should establish which parts of a model actually benefit. [3]
Support changes with architecture. Numerical formats, instruction shapes and acceleration features must be read from the relevant generation's documentation. A statement about a newer device should not be projected backward onto every NVIDIA GPU, and another manufacturer's matrix units have their own terminology and execution contracts. [4][5]
Precision is a computational choice
Floating-point formats divide a fixed number of bits among sign, exponent and significand information. Different formats trade representable range against precision. Using fewer bits can reduce storage and transfer requirements and enable faster supported arithmetic. It can also change numerical error. A machine-learning deployment therefore needs both a speed measurement and an evaluation of whether the result remains suitable for its task. [6][7]
Mixed precision uses different representations for different operations or stages instead of blindly converting every value to one small format. Some operations can use lower-precision inputs while accumulating in a wider format. Training may use gradient scaling where appropriate to reduce underflow problems. These techniques are implementation choices, not proof that an arbitrary model preserves its quality under every low-precision mode. [6]
Suppose a hypothetical inference service meets its accuracy requirement in one representation but fails that requirement after a more aggressive conversion. Even if the second implementation reports twice as many tokens per second, it is not twice as productive at the original task. The meaningful unit is acceptable output. An evaluation should define that standard before choosing the fastest numerical mode.
Shapes, tiling and memory reuse
Matrix dimensions influence how well work maps to hardware tiles. Small or poorly aligned shapes can leave resources underused or require different kernels. Increasing batch size may create larger operations with better reuse, but can also increase memory consumption and the waiting time for individual requests. The best batch for offline throughput need not be the best batch for an interactive service. [1][8]
Arithmetic intensity describes arithmetic performed per byte moved across a specified memory boundary. Matrix multiplication can reuse a tile's inputs many times, improving this ratio. If data movement is the bottleneck, improving reuse can be more valuable than increasing theoretical arithmetic capacity. Conversely, a well-fed large matrix operation can be compute-limited. The workload determines which ceiling applies. [2]
As an illustrative calculation, multiplying square matrices of dimension 1,000 involves roughly two billion floating-point operations under the common multiply-plus-add counting convention. At a hypothetical sustained 20 trillion relevant operations per second, arithmetic alone would take about 0.1 milliseconds. Launch overhead, transfers and other operations can dominate at that scale. This is a lower-bound exercise, not a prediction for a real application.
Dense and sparse peaks are different
A dense calculation treats all values in the specified operation as present. Sparsity-based acceleration relies on a supported pattern or representation that lets hardware avoid some arithmetic. The required structure, metadata and software path matter. A specification quoting sparse throughput is not evidence that a dense application with arbitrary inputs obtains the same rate. [9]
Consider a hypothetical comparison where device A advertises a sparse low-precision peak and device B advertises a dense FP32 peak. Dividing those figures creates a number but not a meaningful application speed ratio. To compare them fairly, align representation, operation type, sparsity assumptions and the workload, then measure the complete execution path.
How applications reach tensor cores
Frameworks and libraries can dispatch supported operations to kernels that use tensor cores. The chosen data type, dimensions, device capability and library implementation influence the path. A developer does not usually allocate a tensor core as though it were a separate cloud server. High-level software expresses an operation; the compiled or selected implementation maps it onto available hardware. [4][6]
We recommend measuring before writing custom matrix kernels. Established libraries include substantial tuning, and an application can lose more time in redundant copies or unfused surrounding operations than in its matrix multiplication itself. A custom implementation is valuable when it addresses a demonstrated limitation and has correctness tests across the intended input shapes.
Interpreting an AI performance claim
Ask which representation is used for inputs and accumulation, whether sparsity is assumed, whether the number is theoretical or measured, and which problem sizes were used. Also check memory fit, quality criteria and total runtime. Tokens per second, time per training step and peak FLOPS answer different questions. None should be silently substituted for another.
The enduring lesson is specialization with conditions. Tensor cores can make suitable matrix arithmetic extremely efficient. Their value is realized only when software maps useful work to them and the rest of the system supplies data at the required rate. A faster matrix engine is an opportunity, not a universal speedup applied to every line of a program.
Check your understanding
Try answering before opening the explanation. Your answers are not collected or scored.
1. Does a tensor stored on a GPU necessarily use tensor cores?
No. Its operations must map to supported matrix instructions through the chosen software implementation. Many tensor operations use other resources.
2. Why is sparse throughput not a universal dense throughput figure?
Sparse acceleration depends on supported data structure and execution conditions. Dense inputs do not automatically satisfy those conditions.
3. Why evaluate output quality alongside speed?
Changing numerical representation can change errors or model behavior. More output is not more acceptable work if the original quality requirement is no longer met.
Sources & editorial note
Reference documentation is listed below with its recorded check date. Technical statements are attributed; passages framed as our view or recommendation are editorial interpretation. Examples are hypothetical unless explicitly identified otherwise. No independent Kovara hardware testing is claimed.
- NVIDIA · Matrix Multiplication Background ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · GPU Performance Background ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · CUDA Best Practices Guide ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- NVIDIA · CUDA programming model ↗ (opens in a new tab)Developer documentation · Checked 28 September 2026
- AMD · GPU microarchitecture ↗ (opens in a new tab)Architecture documentation · Checked 28 September 2026
- PyTorch · Automatic Mixed Precision ↗ (opens in a new tab)Framework tutorial · Checked 28 September 2026
- Hugging Face · Quantization ↗ (opens in a new tab)Framework documentation · Checked 28 September 2026
- NVIDIA · Inference optimization ↗ (opens in a new tab)Manufacturer technical article · Checked 28 September 2026
- NVIDIA · Structured sparsity and sparse Tensor Cores ↗ (opens in a new tab)Manufacturer technical article · Checked 28 September 2026
Prepared with AI assistance. Publication authorized by Tommaso Luci; this does not claim independent technical peer review. Kovara Research is the publication label, not a claim of an independent laboratory or a named analyst team.