the short version

A consumer card is designed to render frames in a desktop. A datacenter card is designed to do arithmetic in a rack, for years, next to seven of its siblings. They are often built on the same architecture and share an instruction set, so the differences are not mainly about the die. They are about memory, connectivity, reliability and the conditions the card is expected to survive.

The price gap is large enough that it is worth knowing exactly what you are buying, because for a lot of work the cheap one is the correct answer.

what actually differs

ConsumerDatacenter
ExamplesGeForce, RadeonH100, MI300 and successors
Memory typeGDDR on the boardHBM on the package
Memory capacityRoughly 8 to 32 GBWell past 100 GB, and climbing each generation
ECCGenerally absentAlways on
Peer interconnectNone to speak ofNVLink with NVSwitch, or Infinity Fabric / xGMI
Display outputsYesUsually none at all
CoolingIts own fans, designed for a desktop casePassive, relying on rack airflow. It will overheat on a bench.
FP64Deliberately weakStrong, and the reason scientific computing still cares
PartitioningNoHardware partitioning into isolated instances
Form factorLarge multi-slot cardsRack modules, often a socketed board carrying several devices

memory is usually the whole decision

In practice the first question is not how fast the card is but whether the job fits on it, because it is a hard wall rather than a soft one. A model that does not fit does not run slowly, it does not run. Capacity is the single line item that most often forces the expensive choice, and the gap is roughly an order of magnitude.

That is also why the workarounds people reach for are all capacity tricks rather than speed tricks: quantizing weights to fewer bits, offloading layers to host memory, sharding a model across several cards, or recomputing activations rather than storing them. Every one of those buys capacity by spending time or accuracy.

ECC, and why the rack always runs it

Memory occasionally flips a bit on its own. At a desktop scale, over a few hours of gaming, the chance is small enough to ignore and the consequence is a stray pixel. At the scale of a training run, thousands of devices and weeks of wall clock, it stops being unlikely and the consequence is a silently corrupted model or a crash a long way from the cause.

Error correcting memory detects and fixes single-bit errors and reports the count, which is the part that matters operationally: a card quietly accumulating corrected errors is a card to replace before it takes a run down with it. The cost is a small slice of capacity and bandwidth, which is an easy trade at that scale and a pointless one on a desktop.

the interconnect gap

This is the difference that is hardest to work around. Datacenter parts carry a dedicated high-speed network to their neighbours, so that splitting a model across eight devices does not mean routing every exchange through the host over PCIe. Consumer cards generally have nothing of the kind, so multi-GPU work on them is bounded by the host link.

For a single-device workload this does not matter at all. For distributed training it is decisive, and it is the reason a multi-card consumer build rarely scales the way the raw compute suggests it should.

the third tier

Between the two sits a workstation class, the Nvidia RTX professional cards and AMD's Radeon Pro line. These take the consumer die and add the parts that matter for professional work: more memory, ECC, certified drivers, and a form factor that fits under a desk rather than in a rack. They keep display outputs and their own cooling.

If you need ECC and capacity but not a datacenter's interconnect or its power and cooling, this tier is often the honest answer, and it is routinely skipped in comparisons that present the choice as a straight two-way split.

when the cheap card is the right call

A consumer card is genuinely the better tool for learning GPU programming, for developing and debugging kernels before they run anywhere expensive, for inference on models that fit its memory, and for fine-tuning small models with the capacity tricks above. The instruction set is close enough that code written on one runs on the other, which is the entire point of developing locally.

Two caveats. Sustained full-load work on a desktop card in a warm room will thermally throttle in a way a rack module will not, so your local timings are not the timings you will get in production. And Nvidia's licence terms restrict deploying GeForce cards in datacenters, which does not affect your desk but very much affects anyone thinking of building a cheap cluster out of them.

see it on your own machine

nvidia-smi --query-gpu=name,memory.total,ecc.mode.current,display_active --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | capacity, whether ECC is on, and whether the card is driving a display, which together identify the tier |'cd_smi'
nvidia-smi --query-gpu=ecc.errors.corrected.volatile.total --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | corrected error count, the number worth watching on a card that is about to become a problem |'cd_ecc'
rocm-smi --showproductname --showmeminfo vram | https://rocm.docs.amd.com/projects/rocm_smi_lib/en/latest/ | the same identification on AMD hardware |'cd_rocmsmi'

related topics

Anatomy of a GPU — where these two words were first defined, and the HBM against GDDR split.
The GPU Memory Hierarchy — what the capacity at the bottom of the ladder is actually holding.
GPU Memory Planning for LLMs — turning the capacity wall into a number before you buy anything.

reference

NVIDIA GPU architecture whitepapers
AMD Instinct accelerators
AMD HIP documentation
NVIDIA CUDA C++ Programming Guide