Kovara Research
Articles & analysis
29 publications
Why GPU memory matters as much as compute
A model must fit before it can run well. Capacity, bandwidth and the KV cache explain why peak compute is only part of the story.27 September 2026A GPU cluster is only as useful as its connections
NVLink, scale-out networks and collective communication turn a list of accelerators into a system—or a bottleneck.27 September 2026AI capacity starts before the first GPU is installed
Power, rack density, cooling and commissioning explain why a data-center announcement is not the same as deployable compute.27 September 2026Liquid cooling changes more than the temperature
Why chip-level heat removal is also a conversation about facility design, water boundaries and operational responsibility.27 September 2026The cheapest GPU-hour is not always the cheapest job
A practical way to compare published rates without losing sight of configuration, commitment and completed work.27 September 2026A published price is not a capacity reservation
Listing, availability, quote and reservation are different claims. Keeping them separate makes sourcing decisions clearer.27 September 2026How to read evidence on Kovara
What sources, dates, confidence grades and unknown fields tell you—and why none should be read as an unconditional guarantee.27 September 2026ASIC vs GPU for AI: Custom Accelerators Compared with General-Purpose GPUs
How custom AI chips such as TPU, Trainium, Maia, MTIA, Gaudi, Cerebras and Groq compare with GPUs on flexibility, software, efficiency and who can actually buy them.29 September 2026China's Semiconductor Industry: SMIC, YMTC, CXMT and Huawei Ascend
A neutral overview of China's chip industry as of September 2026: foundries, memory makers, Huawei's Ascend AI chips, domestic GPU firms, state funding and US export controls.29 September 2026Europe's Semiconductor Landscape: ASML, Infineon, imec and New Fabs
Europe leads in lithography, chip equipment and research but makes under a tenth of the world's chips. A September 2026 look at ASML, Zeiss, ASM, Infineon, ST, NXP, imec and ESMC.29 September 2026GPU Cloud in Europe: Hubs, Power, Sovereignty and AI Factories
Where GPU cloud capacity sits in Europe, why grid limits are pushing it from FLAP-D to the Nordics, and how GDPR, the AI Act and EU AI Factories shape it.29 September 2026Types of GPU Cloud Providers: Hyperscalers, Neoclouds and Marketplaces
A guide to the kinds of GPU compute suppliers, from hyperscalers and neoclouds to marketplaces, sovereign clouds, bare metal and colocation, and what each suits.29 September 2026HBM generations explained: HBM2, HBM2E, HBM3, HBM3E and HBM4
How each HBM generation raised stack height, per-pin speed and bandwidth per stack, which JEDEC standards define them, and which GPUs use which memory.29 September 2026HBM supply: SK hynix, Samsung and Micron and the AI memory bottleneck
Who makes HBM, how the three suppliers' positions have shifted, how they are expanding capacity, and why HBM and its packaging became a constraint on AI accelerators.29 September 2026History of NVIDIA: From NV1 and GeForce to CUDA and AI Data Centers
NVIDIA’s path from its 1993 founding through NV1, RIVA 128, GeForce 256 and CUDA to the deep-learning boom and a data-center business worth tens of billions per quarter.29 September 2026History of the Integrated Circuit: Kilby, Noyce, Fairchild and Intel
How Jack Kilby and Robert Noyce invented the integrated circuit, why Jean Hoerni’s planar process mattered, and how Fairchild gave rise to Silicon Valley and Intel.29 September 2026History of the Transistor: From Bell Labs to FinFET and Gate-All-Around
How the transistor evolved from Bell Labs' 1947 point-contact device through the MOSFET, high-k metal gates and FinFETs to today's gate-all-around nanosheets.29 September 2026The History of TSMC and the Foundry Model That Powers AI Chips
How Morris Chang founded TSMC in 1987, why the pure-play foundry model let fabless firms like NVIDIA, AMD and Apple thrive, and TSMC’s central role in AI chips today.29 September 2026How Semiconductors Are Made: From Sand to Chip, Step by Step
A step-by-step guide to chipmaking: silicon ingots and wafers, lithography, deposition, etching, doping, metal wiring, testing and packaging, and why fabs cost so much.29 September 2026India Semiconductor Mission: Fabs, Chip Plants and AI Compute in 2026
What India's Semiconductor Mission has approved and built as of September 2026, from Tata-PSMC Dholera and Micron Sanand to ISM 2.0 and IndiaAI's subsidised GPU capacity.29 September 2026Japan's Semiconductor Revival: Rapidus, TSMC Kumamoto and Kioxia
From 1980s DRAM dominance to decline and a state-backed comeback: Rapidus 2nm, TSMC's Kumamoto fabs, Kioxia, Micron Hiroshima and Japan's equipment strengths.29 September 2026NVIDIA GPU Architectures Timeline: Tesla to Blackwell and Rubin
A generation-by-generation guide to NVIDIA’s data-center GPU architectures, from G80 and Fermi to Volta’s first Tensor Cores, Hopper, Blackwell and Rubin.29 September 2026On-Demand vs Reserved vs Spot GPUs: Cloud Pricing Models Explained
How on-demand, spot, reserved, capacity-block and private-contract GPU rentals work, and the trade-offs each makes between price, availability and risk.29 September 2026The Semiconductor Supply Chain Explained: From Design to Packaging
A step-by-step guide to how chips are made: EDA and IP, fabless design, equipment, materials, fabrication, memory and packaging, and where each step is concentrated.29 September 2026South Korea's Memory Chip Industry: Samsung, SK hynix and HBM
How South Korea came to lead DRAM, NAND and HBM through Samsung and SK hynix, where its fabs are, and how government policy supports the industry as of September 2026.29 September 2026Why Taiwan Is Central to the Semiconductor Industry
How Taiwan became the hub of leading-edge chipmaking: ITRI, UMC and TSMC, the foundry model, the island's manufacturing share, and the risks analysts discuss.29 September 2026The memory wall in AI: why bandwidth, not FLOPS, limits LLM inference
Why memory bandwidth and capacity often limit AI performance more than raw compute, how the KV cache and arithmetic intensity explain it, and how hardware and software are responding.29 September 2026Where AI Data Centers Are Built and Why: Power, Grids, Water and Policy
Why AI data centers cluster where they do: power and grid access, interconnection queues, water and cooling, latency, land and incentives, and the hubs as of September 2026.29 September 2026The learning library
Computing8 minadvanced packagingWhat chiplets, interposers and 2.5D and 3D packaging are, how CoWoS, InFO, EMIB, Foveros, SoIC and UCIe fit together, and why packaging capacity became a limit on AI GPU supply.GPUs9 minAI chip export controlsHow US export controls on advanced AI chips work, what performance thresholds like TPP mean, how the rules changed from 2022 to 2026, and how they affect GPU availability.GPUs9 minAMD Instinct MI300XA guide to AMD's Instinct MI300X, MI325X and MI350X/MI355X GPUs: CDNA 3 and CDNA 4 chiplets, 192 to 288 GB of HBM, ROCm, and how they compete with NVIDIA.Cloud8 minAWS TrainiumA guide to AWS custom AI silicon: Inferentia, Inferentia2, Trainium, Trainium2 and Trainium3, the Neuron SDK, UltraServers, and where they fit next to GPU instances.Computing8 minCHIPS ActWhat the 2022 US CHIPS and Science Act funds, how manufacturing incentives were awarded, the largest projects, and how the program has changed as of September 2026.Cloud7 minContainerLearn what a container packages, what it shares with the host, and how persistent storage, GPU access and build identity affect reliable deployment.Computing7 minCPUUnderstand the processor that runs an operating system, coordinates applications and feeds accelerator workloads—and why clock speed alone is not enough.Computing8 minCUDAUnderstand host and device code, kernels, threads, memory, streams and compatibility before confusing a toolkit version with an entire working environment.Data centers7 minData centerUnderstand the physical systems behind cloud computing, distinguish a facility from a region or campus, and learn why announced megawatts are not the same as deployable compute.Computing7 minEuropean Chips ActHow the EU Chips Act of 2023 works, what its three pillars fund, which projects it supported, what auditors found, and where Chips Act 2.0 stands as of September 2026.Computing8 minEUV lithographyExtreme ultraviolet lithography prints the smallest features on modern chips. How 13.5 nm light is made, why ASML is the only supplier, and what High-NA EUV changes.AI & models8 minFP8How FP32, TF32, BF16, FP16, FP8, MXFP4, NVFP4 and INT8 differ, why lower precision speeds up AI training and inference, and what it costs in accuracy.GPUs8 minGDDR vs HBMHow DDR5, LPDDR5X, GDDR6/GDDR7 and HBM differ in bandwidth, capacity, power and cost, and why CPUs, gaming GPUs, Grace and data-center GPUs each use a different one.GPUs17 minGPUA detailed introduction to graphics processing units, from parallel execution and memory hierarchies to AI workloads, interconnects and meaningful performance comparisons.Computing7 minGPU clusterUnderstand nodes, parallelism, network topology, scheduling, storage and recovery before assuming that twice as many GPUs means a job finishes twice as fast.Buyer guides8 minGPU Spot pricingLearn how interruptible GPU pricing works, why quota is not available stock, and how checkpointing, retries and deadlines change a price comparison.GPUs7 minHBMExplore stacked DRAM, wide interfaces, packaging, bandwidth calculations and the distinction between a memory stack and a complete accelerator.Networking7 minInfiniBandExplore adapters, switches, fabric management, RDMA and collective communication—and learn why a port speed is not a complete cluster-performance specification.Data centers7 minLiquid coolingUnderstand cold plates, immersion, coolant distribution units and facility heat rejection, including why a closed loop does not automatically mean zero water use.GPUs8 minMIGHow NVIDIA Multi-Instance GPU splits an A100, H100, H200 or B200 into up to seven isolated GPUs, how MIG profiles work, and how MIG differs from vGPU and time-slicing.Computing9 minMoore’s lawWhat Gordon Moore actually predicted in 1965 and 1975, how the trend held for decades, what the ‘end of Moore’s law’ debate means, and why AI shifted focus to packaging and memory.Cloud9 minNeocloudWhat neoclouds such as CoreWeave, Lambda, Nebius and Crusoe are, how they differ from AWS, Azure and Google Cloud, and how they sell and finance GPU capacity.GPUs10 minNVIDIA BlackwellHow NVIDIA Blackwell works: the dual-die B200 and B300 GPUs, GB200 and GB300 superchips, FP4, NVLink 5, memory configurations and what changed from Hopper.GPUs8 minNVIDIA H100What the NVIDIA H100 is: the Hopper GPU's SXM, PCIe and NVL versions, Transformer Engine and FP8, 80 and 94 GB memory, NVLink 4, and why it became the AI reference GPU.GPUs8 minNVIDIA L40SWhat the NVIDIA L40S is: an Ada Lovelace PCIe GPU with 48 GB of GDDR6, FP8 support and RT cores, and how it compares with the A100 and H100 for AI and graphics.Data centers9 minNVL72GB200 NVL72 and GB300 NVL72 explained: 72 Blackwell GPUs in one NVLink domain, rack layout, liquid cooling, power per rack and why rack-scale design matters.Networking6 minNVLinkUnderstand GPU-to-GPU communication, switches, topology and bandwidth boundaries without assuming that every GPU offering includes the same connections.Storage7 minNVMeUnderstand the difference between NVMe, PCIe and flash media, then learn how request size, queue depth, latency and durability shape storage behavior.Storage7 minObject storageLearn how object storage differs from block devices and filesystems, and how identity, request behavior, lifecycle policies and permissions affect AI data workflows.Networking7 minPCIeLearn how PCI Express connects devices, how to interpret transfer rates, and why slot size, negotiated width and upstream topology must be checked separately.Computing8 minprocess nodeWhy chip process names like 3nm, 2nm and Intel 18A no longer measure anything physical, and how transistor density, gate-all-around and backside power define progress.Data centers7 minPUELearn the power usage effectiveness formula, its measurement boundary, common averaging mistakes and why a lower ratio is not automatically lower total energy or carbon impact.Networking7 minRDMAUnderstand how remote direct memory access changes the network data path, why registered buffers matter, and what a completed transfer does—and does not—prove.Networking7 minRoCELearn how RoCE transports remote-memory operations, why routing and congestion matter, and how to distinguish adapter capability from an engineered end-to-end fabric.GPUs7 minSXMUnderstand what an SXM GPU specifies, how it fits into a server, and why module format, host connectivity and GPU-to-GPU links are separate questions.GPUs7 minTensor coresUnderstand the specialized arithmetic behind many AI speed claims, including matrix multiplication, mixed precision, sparsity and the limits of headline FLOPS.Computing10 minTPUWhat Google's Tensor Processing Units are, how systolic arrays and TPU pods work, how each generation from TPU v1 to Ironwood and TPU 8 differs, and how to rent them.AI & models7 minTraining & inferenceUnderstand optimization, gradients, inference execution, fine-tuning, memory state and the difference between fast tensor operations and a useful AI service.Cloud7 minVirtual machineUnderstand hypervisors, guest operating systems, virtual CPUs, device access and persistence before treating a cloud instance as an identical copy of a physical server.GPUs7 minVRAMLearn what GPU memory holds, how it differs from system RAM, and how to estimate a workload's memory requirement without confusing capacity with speed.Computing8 minwafer yieldWhat a 300 mm wafer is, how many dies fit on one, how defect density drives yield, what binning does, and why huge AI GPU dies cost far more to make.
