# Mistral Large 4: inference memory estimates

Prepared 8 October 2026. All memory figures below are calculations, not measured hardware usage.

## Findings and source

Using the published rounded total of 1 trillion parameters, ideal packed weight payloads are about 2,000 GB at 16-bit, 1,000 GB at 8-bit, and 500 GB at 4-bit. The illustrative runtime scenario raises these to about 2,286, 1,220, and 670 GB respectively.

Source: [Mistral, Mistral Large 4 announcement, 6 October 2026](https://mistral.ai/news/mistral-large-4/). Evidence used: user-supplied transcribed source notes, recorded as retrieved on 8 October 2026. This session did not independently retrieve the announcement. No weights were downloaded and no inference was run.

Published facts as reported in those notes: 1 trillion total parameters; 52 billion active parameters; native multimodality; public API preview announced on 6 October; weights planned for release by the end of October, with further architecture details alongside them. The supplied announcement notes do not specify exact layer count, KV heads, head dimension, quantization layout, or deployment memory measurements.

The active fraction is approximately 52 billion / 1 trillion = 5.2%. Active parameters describe selected computation. In a fully resident mixture-of-experts deployment, unselected expert weights remain available because subsequent tokens can select different experts. Sharding distributes storage across devices; offloading moves some weights to RAM or storage and adds transfer costs. Neither implies that the remaining parameters disappear.

## Calculation table

GB = 10^9 bytes; GiB = 2^30 bytes. Each table cell below is **GB / GiB**. Decimal places document conversions, not accuracy of the rounded parameter count.

| Precision scenario | Weight payload | Quantization metadata | KV cache | Runtime allowance | Total |
|---|---:|---:|---:|---:|---:|
| 16-bit | 2000.00 / 1862.65 | 0.00 / 0.00 | 68.72 / 64.00 | 217.18 / 202.26 | 2285.90 / 2128.91 |
| 8-bit | 1000.00 / 931.32 | 31.25 / 29.10 | 68.72 / 64.00 | 120.30 / 112.04 | 1220.27 / 1136.47 |
| 4-bit | 500.00 / 465.66 | 31.25 / 29.10 | 68.72 / 64.00 | 70.30 / 65.48 | 670.27 / 624.24 |

The accompanying CSV preserves the component byte counts and their GB/GiB conversions. Its baseline is four simultaneous sequences, each with 32,768 cached tokens. Totals are computed before rounding.

## Formulas and assumptions

Let P = 10^12 total parameters and b be bits per weight.

- Weight payload: W = P × b / 8 bytes. Assumes every parameter uses that precision and is ideally packed. These scenarios do not establish official ML4 precision support.
- Quantization metadata: M = (P / 128) × 4 bytes at 8-bit and 4-bit; M = 0 in the 16-bit scenario. Assumes one 2-byte scale and one 2-byte zero point for each 128 weights. This adds 0.25 bits per parameter, or 31.25 GB / 29.10 GiB. Actual formats may use different groups, omit zero points, introduce padding, or retain some tensors at higher precision.
- Conventional attention KV cache: K = 2 × L × B × T × H_KV × d × s bytes. The leading 2 represents keys and values. Hypothetical architecture: L = 128 layers, H_KV = 8 KV heads, d = 128 dimensions per head, s = 2 bytes per element (16-bit/BF16 cache). B is simultaneous sequences; T is cached tokens per sequence, including cached prompt and generated tokens. The baseline uses B = 4 and T = 32,768, producing 64 GiB.
- Other runtime allowance: R = 0.10 × (W + M) + 16 × 2^30 bytes.
- Estimated total: W + M + K + R.

All architecture and overhead settings above are illustrative assumptions, not official ML4 specifications. The runtime allowance is a planning placeholder for activations, workspaces, communication buffers, and allocator overhead; it is not an empirically validated reserve.

## KV-cache sensitivity

Same hypothetical architecture and 16-bit cache in every row; only B and T change.

| Simultaneous sequences B | Cached tokens per sequence T | KV cache, GB | KV cache, GiB |
|---:|---:|---:|---:|
| 1 | 8,192 | 4.29 | 4.00 |
| 1 | 32,768 | 17.18 | 16.00 |
| 4 | 8,192 | 17.18 | 16.00 |
| 4 | 32,768 | 68.72 | 64.00 |

Cache consumption is 524,288 bytes (0.5 MiB) per cached token per sequence under these assumptions. It scales linearly with B and T: increasing either fourfold increases cache memory fourfold. The largest table scenario uses 16 times the cache of the smallest. This cache table applies to all three weight-precision scenarios because cache precision is held fixed; quantizing weights does not automatically quantize the KV cache.

## Interpretation and limits

These totals estimate aggregate inference memory for one resident model copy, not per-GPU capacity or training memory. They exclude replicated model copies, loading peaks, host-side duplication, and deployment-specific imbalance. Actual activations and workspace peaks vary with batching, prefill implementation, and parallelism. Sliding-window attention, compressed caches, cache quantization, or different multimodal processing can change memory needs. No model-specific multimodal memory requirement is inferred here.

The published rounded parameter count supports approximate weight arithmetic. Exact runtime requirements require the released architecture/checkpoint details and measurements of the intended serving configuration.
