Racks & Watts
We may earn a commission when you buy through links on this site, at no extra cost to you. Our recommendations are based on genuine research, not commission. Learn more.

A builder's guide to budget NVIDIA cards for Ollama

Best budget GPU for local LLMs in 2026: 6 GB, 8 GB or 12 GB, measured

September 2026 prices and our bench numbers: 12 GB is the line that matters for local LLMs. A refurbished RTX 3060 12 GB beats a new 8 GB card for Ollama on a $250–$550 budget.

NVIDIAGPU

VRAM decides what you can run. 12 GB is the line worth paying for in 2026.

Pick your model size first7B/8B fits in 8 GB; 14B Q4 needs about 9 GB to stay on GPU

Price 12 GB refurbished, not newRTX 3060 12 GB refurbs listed $359.99–$429.97 vs $457.58+ new

Cap power, keep the PSUOur 3060 Ti held 76.3 tok/s at a 150 W cap vs 200 W rated

Illustrated summary · not a benchmark chart

7B on RTX 3060 Ti76.3 tok/s at 150 W cap
14B Q4 default tagabout 9 GB
14B spilled to CPU~38 → 24.9 tok/s
Research checkedSeptember 25, 2026

The short answer

VRAM decides what you can run. 12 GB is the line worth paying for in 2026.

If you want 14B models fully on GPU, buy a refurbished RTX 3060 12 GB, seen from $359.99. If you'll only run 7B/8B, an 8 GB card suffices, but don't pay $500+ for one. Treat 6 GB as a stopgap.

A stronger fit

Builders with $360–$460 who want 14B-class models (qwen2.5:14b) fully in VRAM, or 7B/8B models with room for longer context, on a 550 W-class PSU.

Think twice if

Anyone set on 7B/8B only and paying $500+ for a new 8 GB card, or buying a 6 GB card expecting it to grow into bigger models.

Our assessment is based on vendor documentation, not a scored hands-on test. How we evaluate products.

Plans at a glance

NVIDIA's tiers, side by side.

RTX 2060 6 GB

$185.99+ used
6 GB GDDR6, 192-bit

A cheap stopgap for small models; new listings at $259+ are poor value.

Check current terms ↗

RTX 3050 6 GB

$249.99 new
6 GB variant

Only if you need a new card under $260 and accept small models.

Check current terms ↗

RTX 3060 12 GB

$359.99 refurb
12 GB GDDR6, 192-bit, 170 W

Our pick: the cheapest NVIDIA route to 14B Q4 fully on GPU.

Check current terms ↗

RTX 4060 8 GB

$499.99+ new
8 GB GDDR6, 128-bit

Hard to justify for LLMs: 8 GB caps you at 7B/8B at this price.

Check current terms ↗

Prices as listed on the retailer or manufacturer page on September 25, 2026. Listings and stock change daily; verify before you buy.

Buy for the model size you want, because VRAM is the real limit

The card you want is the one that holds your model entirely in VRAM, and that is decided by gigabytes, not CUDA cores. On our test bench, a 7B/8B Q4 model ran fully on an 8 GB RTX 3060 Ti at 70+ tokens/s (bench notes). A 14B Q4 model needs about 9 GB to stay fully on GPU and lost more than a third of its speed when it spilled to CPU.

Key takeaway: 8 GB runs 7B/8B well. 14B needs about 9 GB, so a 12 GB card is the cheapest way to get it. In September 2026 the refurbished RTX 3060 12 GB is the value pick.

Ollama publishes download size per tag, but runtime VRAM is weights plus KV cache and overhead, so you need headroom beyond the file (Ollama qwen2.5 tags).

Model (Q4_K_M default) Size on disk Fits 6 GB? Fits 8 GB? Fits 12 GB?
llama3.1:8b ~4.7 GB (tags) Tight, unmeasured Yes, measured Yes
qwen2.5:7b ~4.7 GB (tags) Tight, unmeasured Yes, measured Yes
qwen2.5:14b ~9 GB (tags) No No, spills Should fit, unmeasured on a 3060

Expect 70+ tokens/s from 8 GB on 7B/8B models

We measured this, and 8 GB is plenty for this class. With a 150 W software cap, the RTX 3060 Ti ran qwen2.5:7b at 76.3 tokens/s generation and 3,847 tokens/s prompt processing, at 149.6 W and 51 C. llama3.1:8b ran at 73.9 tokens/s, 149.9 W, 51 C (bench notes).

Method: Ollama 0.34.2, Q4_K_M, one /api/generate call with stream:false and num_predict 320 after a warm-up. Tokens/s is eval_count/eval_duration, single runs, so treat ±5% as noise. Idle board power was 16.8 W (driver-reported).

Tip: The 3060 Ti is rated at 200 W (NVIDIA); we capped it at 150 W and still got 76 tokens/s. Cap first before you shop for a bigger PSU.

Pay for 12 GB if you want 14B, because spilling costs a third of your speed

If 14B models are the goal, 8 GB is a trap. Our qwen2.5:14b run split 6.7 GB on the 3060 Ti, 1.7 GB on the RTX 2060 and about 0.5 GB on CPU: 24.9 tokens/s. With the whole 2060 free, the same model sat fully in VRAM across both cards at about 38 tokens/s (bench notes).

Only the RTX 3060 12 GB tops 8 GB below the 3070 in NVIDIA’s 30-series table (NVIDIA compare). The RTX 4060 Ti also comes in 16 GB (NVIDIA), but our sources have no 16 GB price. The 3060 12 GB uses a 192-bit bus vs 256-bit on the 3060 Ti (NVIDIA), so we’d expect lower tokens/s than our 3060 Ti numbers; that’s our expectation, not a measurement.

Example (illustrative): You want qwen2.5:14b for coding help. At ~9 GB it spills on any 8 GB card, landing near our 24.9 tokens/s split result or worse. A 12 GB card should keep it fully on GPU.

Skip new 8 GB cards at $500+, and buy 12 GB refurbished

The old “$280 RTX 3060 12 GB” advice is dead, but refurbished units still undercut new 8 GB cards. On September 25, 2026, Newegg listed a refurbished MSI Ventus 2X OC 12G at $359.99 and a refurbished EVGA XC at $394.99, versus new 3060 12 GB cards from $459.97 (Newegg). Amazon Renewed 3060 12 GB listings ran $399.97–$429.97, mostly 1–3 in stock, and new ones $457.58–$539.99 (Amazon).

The same day, a new RTX 4060 8 GB was $530.00+ on Amazon (Amazon) and $499.99–$543.36 on Newegg (Newegg). A new RTX 5060 8 GB was $539.99 (Amazon). That’s $140–$180 more for less VRAM.

Watch out: New RTX 2060 6 GB listings ran $259–$369.99 (Amazon). At $339–$370 you’re a few dollars from a refurbished 12 GB 3060. Only consider a 2060 on used offers, which started at $185.99.

The Intel Arc B580 12 GB listed at $329.99 new (Amazon), but it’s outside this NVIDIA guide and we haven’t tested it.

Run this worksheet before you click buy

  1. Write down the biggest model you’ll actually run and its default tag size (qwen2.5, llama3.1).
  2. Add headroom for KV cache and overhead. Our rule: a 14B Q4 needs about 9 GB to stay on GPU.
  3. If that number is at or under ~6 GB, an 8 GB card works. If it’s around 9 GB, you need 12 GB.
  4. Price that VRAM tier refurbished first, then used, then new. Note stock counts; renewed listings were 1–3 deep.
  5. Check your PSU: NVIDIA lists 550 W system power for the 3060 and 600 W for the 3060 Ti (NVIDIA).
  6. Plan a power cap. Ours kept 70+ tokens/s at 150 W on the 3060 Ti.
  7. Confirm the exact variant: the RTX 3060 also exists as an 8 GB, 128-bit card. The 12 GB model is the one you want.

What we couldn’t verify: we haven’t measured an RTX 3060 12 GB, RTX 4060 or a 6 GB card running a model alone, we have no RTX 3060 Ti price in our sources, and we can’t vouch for refurbished warranty terms or card condition from listings alone.

What we would do next

Three moves before you commit.

Decide 8B-class or 14B-class

If you'll live on llama3.1:8b or qwen2.5:7b, 8 GB is enough at 70+ tokens/s on our bench. If you want qwen2.5:14b, plan for about 9 GB of VRAM, which means a 12 GB card.

Price a refurbished RTX 3060 12 GB

Compare Newegg refurbished (seen from $359.99) against Amazon Renewed ($399.97–$429.97) and check the listing's warranty and return terms. Confirm it's the 12 GB, 192-bit model, not the 8 GB variant.

Proposed test on arrival: warm up, cap, measure

Proposed: set a power cap, run one warm-up, then time qwen2.5:14b with num_predict 320 and read eval_count/eval_duration. Confirm Ollama reports the model fully on GPU, not split to CPU.

Common questions

Quick answers.

Is the RTX 3060 12 GB still worth it in 2026?

For LLMs, yes, if you buy refurbished. It's the only 30-series card below the 3070 with more than 8 GB, and refurbished listings started at $359.99 on September 25, 2026. At $457+ new it's harder to justify, but it still undercuts new 8 GB cards.

Can a 6 GB card run a 7B model?

The default 7B/8B Q4 files are about 4.7 GB, and runtime needs headroom beyond that for KV cache and overhead. We didn't measure a 6 GB card running one alone, so treat it as tight and unverified.

How much faster is a model that fits fully in VRAM?

On our bench, qwen2.5:14b ran at about 38 tokens/s fully in VRAM across two cards, versus 24.9 tokens/s when about 0.5 GB spilled to CPU. That's a loss of more than a third.

Do I need a bigger power supply?

Maybe not. NVIDIA lists 550 W system power for the RTX 3060 and 600 W for the 3060 Ti. We capped our 3060 Ti at 150 W, below its 200 W rating, and still measured 76.3 tokens/s on qwen2.5:7b.

Is a new RTX 4060 8 GB a good LLM card?

Not at September 2026 prices. New listings ran $499.99–$669.99 for 8 GB on a 128-bit bus, which limits you to 7B/8B-class models fully on GPU. A refurbished 12 GB 3060 costs less and holds 14B Q4.

Sources & methodology

Where these facts come from.

We read the manufacturer's specifications and the retailer listings on the dates shown. We have not run a hands-on test unless the article says so.

Ready to look for yourself?

Open NVIDIA's product page with this guide beside it.

Use the worksheet above while you are on their site. If the VRAM or the watts do not fit your plan, you will know before you buy.

See NVIDIA's current pricing ↗

Prepared by Racks & Watts. Measurements are from our own hardware unless the article says otherwise; prices change, verify on the retailer's page before you buy.