chonk.siUnofficial

Can my hardware run it?

Mistral Large 4 has 1.05 trillion parameters, of which 52 billion are used for each token, according to its model card. Even compressed, the weights alone need hundreds of gigabytes of memory.

Source: Mistral's model card

Memory

Memory for all weights
Memory needed for Mistral Large 4's weights at each precision
PrecisionBits per weightAll weightsRead per token
BF16162,100 GB104 GB
FP881,050 GB52 GB
4-bit4525 GB26 GB

"Read per token" is the size of the active weights, which every generated token reads once. All sizes are lower bounds: parameters × bits per weight ÷ 8, in decimal gigabytes. Quantized files are a little larger because they also store scaling factors. Running the model needs more memory on top, for the context (the KV cache) and for the inference software. Mistral has not published the architecture details needed to estimate the KV cache yet.

Your setup

Rent hardware

  • Vast.aiA marketplace for renting GPUs by the hour.
  • RunpodA GPU cloud for renting pods by the hour or running serverless GPUs.

These are affiliate links: chonk.si may earn a commission if you sign up. It does not change any fact or number on this site.

Speed estimate

Generating one token means reading every active weight from memory once. Dividing your memory bandwidth by the size of the active weights gives the most tokens per second a single request can reach. Compute, the KV cache and the links between GPUs all push the real number lower, so treat it as a ceiling, not a prediction.

Sources