self-hosted/ai
§01·tools · /compare

Compare GPUs

Side-by-side benchmarks for 2–4 GPUs at a time. Each cell shows whether the model fits, how fast, and how much VRAM it peaked at.

Every column answers for all 97 models in the catalogue — including the ones that will not fit, which is the answer you most need when choosing between two cards. Memory across the catalogue spans 8 GB — where 4 cards sit — up to the Apple M2 Max at 64GB, and Apple's share of that is unified memory rather than dedicated VRAM.

loading compare…

§02·reading a cell

What a cell is telling you

verified
a published recipe targets this exact model and this exact card.
fits
no recipe for the pair, but one on the same vendor's hardware puts the memory floor at or below this card.
won't fit
a floor measured on any platform is above this card's memory.
untested
it fits in memory as far as we know, and nothing establishes a runtime path.

A comparison is only worth as much as the rows both columns can speak to, which is why every model gets a labelled cell rather than a blank one. A blank is ambiguous in the worst possible way: it reads the same whether a model is far too large for both cards or simply has not been tried, and those two facts point a buyer in opposite directions.

Where the two columns differ, the difference worth trusting is the verdict, not the speed. Verdicts on both sides come out of the same evidence through the same code — a test walks a comparison row by row against the advisor's own answer for each card. Speed figures come from whatever a contributor published, so two cells can differ because of quantisation, context length or a newer runtime rather than the hardware. Open the recipe behind a cell before reading a gap as a hardware gap.

§03·what the rows cover

97 models across 8 modalities

Modality changes what a card needs to be good at. Text generation is bound by memory bandwidth once the weights fit; image and video generation are bound by raw compute and by peak VRAM during the decode step, which is where a card with the same nominal memory can still fall over.

cards you can put in a column

§04·questions

Common questions

How many GPUs can I compare at once?

Two to four. Fewer than two is not a comparison, and past four the row labels stop fitting next to the columns on a phone. The selection lives in the URL, so a comparison you have set up can be sent to someone.

Why does every model appear, even ones that clearly will not run?

Because an omission is unreadable. A table that silently dropped whatever it could not confirm would look identical whether the answer was "too big for both cards" or "nobody has tried it" — and those are opposite pieces of advice. Every model in the catalogue gets a labelled cell in every column instead.

Are the numbers comparable between columns?

The verdicts are, because both columns are derived from the same evidence by the same code. Raw speed figures need more care: they come from whatever run a contributor published, at that run's quantisation, context length and software version, so a difference between two cells can be the software rather than the silicon. The unit strings are not normalised either — one card may report tok/s and another tokens/s.

Can I compare a Mac against a graphics card?

Yes, and the memory column is the interesting part: Apple silicon shares one pool between CPU and GPU, so a 64 GB machine can hold a model no consumer card can — while a discrete card of the same nominal size will usually finish a given run sooner. The verdicts account for the platform difference; they never infer that something runs on Metal from a CUDA recipe.

Only have one card in mind? Ask the advisor what it runs.