Skip to content

Free tool

Fine-Tuning VRAM Calculator

How much GPU memory does it take to fine-tune an LLM? Pick the model and see full fine-tuning, LoRA, and QLoRA side by side, with the minimum GPU each one needs. Training memory is a different world from inference, and the method you pick is the whole story.

Fine-tuning Qwen2.5 7B. QLoRA fits on a single consumer GPU, about 6 GB.. Full fine-tuning needs 141 GB GPU (H200).

Fine-tuning Qwen2.5 7B

QLoRA fits on a single consumer GPU (~6 GB). Full fine-tuning needs 141 GB GPU (H200).

  • Full fine-tune
    126 GB

    Every weight trained (Adam, mixed-precision). Best quality, brutally heavy, and usually multi-GPU.

    Minimum: 141 GB GPU (H200)

  • LoRA
    16.8 GB

    Base frozen in 16-bit; train small adapters. ~10–20× lighter than full, near-full quality for most tasks.

    Minimum: 24 GB GPU (RTX 4090 / 3090)

  • QLoRA
    6 GB

    Base quantized to 4-bit + adapters. The consumer-GPU sweet spot: fine-tune big models on one card.

    Minimum: 8 GB GPU (RTX 4060 / 3070)

Just want to run a model, not train it?

Estimates in bytes-per-parameter, mixed-precision: full ≈ 18 (weights + grads + Adam m/v + fp32 master + overhead), LoRA ≈ 2.4 (frozen 16-bit base dominates), QLoRA ≈ 0.85 (4-bit base + paged adapter optimizer). These include a modest single-batch activation allowance, and real VRAM moves with batch size, sequence length and gradient checkpointing (which can cut activation memory a lot at some speed cost). Dense models; not a benchmark. Re-verify against your framework before committing hardware.

Frequently asked

How much VRAM do I need to fine-tune a 7B model?

It depends entirely on the method. A full fine-tune of a 7B model (mixed-precision Adam) needs roughly 18 GB per billion params for weights, gradients, optimizer states and a master copy (about 126 GB), which wants a 141 GB GPU (H200), or two 80 GB A100/H100s. LoRA drops that to ~16.8 GB because the base model stays frozen in 16-bit and you only train small adapters; it fits a 24 GB GPU (RTX 4090 / 3090). QLoRA quantizes the frozen base to 4-bit and needs only ~6 GB, so a 7B fine-tune runs on an 8 GB GPU (RTX 4060 / 3070).

Why does fine-tuning need so much more VRAM than inference?

Inference only stores the weights plus a small KV cache. Training adds three big costs on top: gradients (same size as the weights), optimizer state (Adam keeps two fp32 moments per parameter, which is 8 bytes/param alone), and activations saved for backpropagation. For a full fine-tune that's roughly 18 bytes per parameter versus ~2 for 16-bit inference, about 9× more.

What's the difference between LoRA and QLoRA for VRAM?

Both freeze the base model and train small low-rank adapters, so the optimizer cost is tiny. The difference is how the frozen base is stored: LoRA keeps it in 16-bit, QLoRA quantizes it to 4-bit, and across a whole model that lands at ~2.4 against ~0.85 bytes per parameter all in. That 2.8× cut on the largest memory component is why QLoRA fine-tunes a 70B on an 80 GB GPU (A100 / H100), where LoRA asks for 168 GB and needs 3× 80 GB (A100 / H100).

Can I fine-tune a 70B model on a single GPU?

Not with full fine-tuning, which needs about 1260 GB — 18× 80 GB (A100 / H100). But QLoRA makes it possible: a 70B model in 4-bit plus adapters lands around 59.5 GB, which fits an 80 GB GPU (A100 / H100). That is the headline result from the original QLoRA paper. On consumer hardware you'd step down to a 24B or smaller model: 24B at 4-bit is 20.4 GB, which fits a 24 GB GPU (RTX 4090 / 3090).

Does gradient checkpointing reduce the VRAM needed?

Yes, gradient (activation) checkpointing recomputes activations during the backward pass instead of storing them, which can cut activation memory substantially at the cost of ~20–30% more compute time. The estimates here assume a modest single-batch activation allowance without aggressive checkpointing; turning it on, and lowering batch size or sequence length, all reduce the real requirement.

Why fine-tuning needs more VRAM than inference

Running a model needs enough memory for its weights. Training or fine-tuning one needs several times more, because the process must store gradients and optimizer state alongside the weights, plus the activations from each step.

Full fine-tuning is memory-hungry

A full fine-tune keeps the weights, their gradients, and optimizer state (often two extra copies for an optimizer like Adam) in memory at once. That can be many times the memory of inference, often eight times or more, which is why full fine-tuning of even mid-size models needs data-center GPUs.

LoRA and QLoRA change the math

Parameter-efficient methods freeze the base weights and train only a small set of added parameters, so gradients and optimizer state cover a tiny fraction of the model. QLoRA also loads the frozen base in 4-bit. Together they bring fine-tuning of a 7B-class model within reach of a single consumer GPU.

Batch size and sequence length

Activation memory scales with batch size and sequence length, so long training samples or large batches raise the requirement quickly. If you run out of memory, cutting the batch size or turning on gradient checkpointing is usually the first fix.

Where this answer comes from

Last checked on Sep 10, 2026, and reviewed on a 60-day schedule. Figures come from the company that publishes them, never from another site’s summary. How we check this.

Related reading

Go deeper on how this works and what to pick.

Newsletter

Liked the tool? Get the signal.

One weekly email on the AI + hardware that actually matters, from the people who build these calculators.

Free · unsubscribe anytime · no spam.