Best Local AI Hardware: Builds from $600 to $3,499

Updated: September 18, 2026

Five complete build tiers from a CPU-only starter to a 32GB RTX 5090 workstation: the exact part lists, the memory-bandwidth math that decides how fast each one generates tokens, and which tier gives the best value for your budget.

Need software next? Explore the models directory for downloads, grab optimized picks from the 8GB model guide, and keep our troubleshooting playbook handy while you build. Also see our 2026 AI PC Build Guide with updated parts lists and pricing. Ready to master AI? Follow our AI Learning Path to learn what AI really is, not just how to use it.

135
Total Models
$600
Starting Price
5
Build Tiers
32GB
Top Build VRAM
💻

Budget Builds

$600-$900 • Runs 48 models (up to 7B)

Performance Builds

$1,200-$2,500 • Runs 96 models (up to 34B)

🚀

Enterprise Builds

$5,000 and up • Check model and multi-GPU requirements

Once your hardware is sorted

Bought the hardware? Now get the most out of it.

Every course on running local models, RAG, agents and fine-tuning, so the machine you just speced actually earns its price.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

Cheapest AI PC: The Quick Answer

The cheapest unified-memory AI PC is the base Mac mini with the M6 chip and 16GB of unified memory (check current price): it runs 7-8B models well out of the box with zero build effort. The cheapest discrete-GPU route is a used RTX 3060 12GB build, which handles everything up to the 14B class and leaves you an upgrade path.

One buying note: Apple moved the Mac mini line to the M6 and M5 Pro chips in September 2026, so older guides still quoting M4 configurations and prices are out of date. Apple lists 153GB/s of memory bandwidth for the entry M6 configuration, 170GB/s for the 24GB and 32GB M6 configurations, and 307GB/s for the M5 Pro. Check the configuration you are buying against Apple's official Mac mini tech specs before you order, since memory is not upgradeable after purchase. For the Mac route, see our Apple Silicon AI guide; for the RTX 3060 route, the budget local AI machine guide has the full parts list, and this page shows exactly what a 12GB card runs. Already own a card? We keep per-GPU model picks (RTX 3060/4060/4070/4090) and per-Mac picks, and a full local AI server comparison by budget.

Local AI PC: What to Buy

A local AI PC is an ordinary desktop with one job that other PCs do not have: holding an entire model in graphics memory. If the model's weights do not fit in VRAM, the card sits half idle while the CPU and system RAM do the work several times slower. So pick the graphics card first, by memory size, and spend what is left on everything else.

  • 16GB of VRAM: RTX 5060 Ti 16GB ($429 MSRP) or, if you want more speed, RTX 5070 Ti 16GB ($749 MSRP). AMD's RX 9070 XT 16GB ($599 SEP) sits between them on bandwidth; confirm it in Ollama's GPU list before buying. Runs the whole 7B-14B class at Q4 with room for context, plus 20B mixture-of-experts models such as gpt-oss:20b.
  • 24GB of VRAM: a used RTX 3090 24GB. No current GeForce card offers 24GB below the RTX 5090's 32GB, so the used 3090 is still the way into the 27B-32B class, including Qwen3.6-27B at Q4.
  • Everything else: a current 6-core CPU (Ryzen 5 9600X class), 32GB of DDR5, a 1TB NVMe drive and a power supply sized for the card. None of these change how fast a GPU-resident model generates tokens.

Best PC for Local AI: 16GB or 24GB?

The best PC for local AI is the one whose VRAM matches the largest model you actually want to run, not the one with the fastest CPU. Generating a token means reading every weight in the model once, so the card's memory size decides what fits and its memory bandwidth decides how fast; the CPU mostly waits. A 16-core CPU does nothing for a model that lives on the GPU, while doubling bandwidth doubles the speed ceiling.

  • Best at 16GB: RTX 5070 Ti 16GB (896GB/s) if speed matters, RTX 5060 Ti 16GB (448GB/s) if budget matters. Same models fit on both; the 5070 Ti has twice the ceiling.
  • Best at 24GB: used RTX 3090 (936GB/s). Unlocks the 27B-32B class, which is where the strongest single-GPU models live right now.
  • Best if money is no object: RTX 5090 32GB (1,792GB/s, $1,999 MSRP). Runs the 32B class with long context and roughly doubles the 3090's ceiling.
  • Why VRAM beats CPU: a Ryzen 5 9600X and a Ryzen 9 9950X give identical token speeds once the model sits on the card. Put the CPU money into the GPU's memory size, and 32GB of system RAM so large models can still load when they spill.

August 2026 Update: What the New Model Wave Means for Your Hardware

The summer 2026 wave of open models has kept splitting the hardware answer in two directions:

  • Small cards got better, not obsolete. Google's Gemma 4 family (Apache 2.0) runs offline on everything from a Raspberry Pi to a mid-range GPU — an 8-12GB card is still a genuinely useful AI machine.
  • 24GB is the new sweet spot for one-GPU builds. Qwen3.6-27B — the current best single-GPU model — fits an RTX 4090/5090 at ~17GB. See what a 24GB card runs best.
  • No GPU at all? System RAM sets the ceiling instead — see the picks for 32GB of RAM.
  • The new giants need server-class memory. GLM-5.2 (753B, MIT) is self-hostable — but quantized it wants a 256-512GB unified-memory Mac Studio or an EPYC-class rig, and the 2.8T Kimi K3 is out of home reach entirely. Don't buy a bigger GPU chasing them.
  • August did not move the hardware goalposts. Meta's Muse Glimmer and Alibaba's Qwen3.8-27B both land in the same 24GB class as the models above, and NVIDIA's Nemotron 3.5 Lightning is a 30B mixture-of-experts with only 3B active — so it runs far lighter than its parameter count suggests. None of them is a reason to buy a bigger card.

The builds below remain current — nothing in the new wave changes the price-to-capability picks. One caveat on timing: GPU street prices are running 50%+ over MSRP right now and the RTX 50 Super refresh is delayed — see why GPU prices are so high and what to buy right now.

5 AI PC Builds Compared

Five part lists from under $1,000 to about $3,500 at MSRP, the range most people actually shop in, re-checked in September 2026 against NVIDIA, AMD and Apple spec pages and against the models we recommend now: Qwen3.6-27B, the Qwen 14B class, and Gemma 4. Only vendor MSRPs appear below, because GPU street prices move week to week; check current prices on the day you buy. The speeds under each build are arithmetic ceilings: card memory bandwidth divided by the weight bytes a Q4_K_M model moves per token. Real output lands below those ceilings.

Budget Champion

under $1,000 at MSRP
check current price

CPU: Ryzen 5 9600X (6-core)

RAM: 32GB DDR5-5600

GPU: None (CPU only)

Storage: 1TB NVMe

📐 Bandwidth ceiling (Q4_K_M):

  • • Llama 3.1 8B (~4.9GB): ~18 tok/s ceiling
  • • Mistral 7B (~4.4GB): ~20 tok/s ceiling
  • • Gemma 4 E2B/E4B: sized for this class
  • Dual-channel DDR5-5600 ≈ 90GB/s; CPU-only decode lands furthest below its ceiling.

Verdict: Perfect starter. Handles all models in our 8GB guide smoothly.

BEST VALUE

Sweet Spot Build

about $1,100 to $1,300 at MSRP
check current price

CPU: Ryzen 5 9600X (6-core)

RAM: 32GB DDR5-5600

GPU: RTX 5060 Ti 16GB ($429 MSRP)

Storage: 2TB NVMe

📐 Bandwidth ceiling (Q4_K_M):

  • • Llama 3.1 8B (~4.9GB): ~91 tok/s ceiling
  • • Qwen 3 14B (~8.4GB): ~53 tok/s ceiling
  • • Gemma 3 12B (~7.2GB): ~62 tok/s ceiling
  • RTX 5060 Ti: 448GB/s, 180W (NVIDIA spec). 16GB leaves room for long context on the 14B class.

Verdict: Best bang-for-buck. The cheapest new 16GB card, so the whole 14B class fits with context to spare, and 20B MoE models such as gpt-oss:20b fit too. See what a 16GB card runs and full GPU comparisons. AMD alternative: the RX 9070 XT 16GB ($599 SEP) has more bandwidth at 640GB/s; confirm ROCm support in Ollama's GPU list first.

24GB on a Budget

about $1,500 with a used 3090
check current price

CPU: Ryzen 7 9700X (8-core)

RAM: 64GB DDR5-5600

GPU: RTX 3090 24GB (used, previous generation)

Storage: 1TB NVMe

📐 Bandwidth ceiling (Q4_K_M):

  • • Qwen3.6-27B (~16.2GB): ~58 tok/s ceiling
  • • Qwen 2.5 Coder 32B (~19.2GB): ~49 tok/s ceiling
  • • Card power: 350W board TGP (NVIDIA spec)
  • RTX 3090: 936GB/s — the cheapest high-bandwidth 24GB card.

Verdict: The one previous-generation card on this page, kept on purpose: no current GeForce card offers 24GB below the RTX 5090's 32GB (the RTX 5080 is 16GB), so the used 3090 is still the cheapest route to 24GB, and 24GB runs the best single-GPU models at full Q4. See what 24GB unlocks. Used prices move week to week, so check current sold listings before you commit. Want 70B? One 24GB card cannot hold it at Q4 (~42.5GB of weights); you need a second 3090 for 48GB pooled — the server comparison covers that tier.

Performance 16GB

about $1,700 to $1,900 at MSRP
check current price

CPU: Ryzen 7 9700X (8-core)

RAM: 32GB DDR5-5600

GPU: RTX 5070 Ti 16GB ($749 MSRP)

Storage: 2TB Gen4 NVMe

📐 Bandwidth ceiling (Q4_K_M):

  • • Llama 3.1 8B (~4.9GB): ~183 tok/s ceiling
  • • Qwen 3 14B (~8.4GB): ~107 tok/s ceiling
  • • gpt-oss:20b: MoE. OpenAI's model card lists ~3.6B active params per token, so it decodes far faster than its 20B size suggests
  • RTX 5070 Ti: 896GB/s, 300W (NVIDIA spec). Twice the Sweet Spot's bandwidth on the same 16GB.

Verdict: Same models as the Sweet Spot, twice the speed ceiling. Run a dev environment and an AI coding assistant side by side; the 16GB tier is where MoE models shine. The RTX 5080 16GB ($999 MSRP, 960GB/s) adds little for LLMs over this card.

Ultimate Workstation

about $3,500 at MSRP
check current price

CPU: Ryzen 9 9950X (16-core)

RAM: 64GB DDR5-5600

GPU: RTX 5090 32GB ($1,999 MSRP)

Storage: 4TB Gen5 NVMe

📐 Bandwidth ceiling (Q4_K_M):

  • • Llama 3.1 8B (~4.9GB): ~366 tok/s ceiling
  • • Qwen3.6-27B (~16.2GB): ~111 tok/s ceiling
  • • Qwen 2.5 Coder 32B (~19.2GB): ~93 tok/s ceiling
  • • Llama 3.3 70B (Q4): loads via RAM offload, but slow. ~10GB of the ~42.5GB spills past the 32GB card, and the spilled layers run at DDR5 speed instead of the card's 1792GB/s
  • RTX 5090: 1792GB/s, 575W (NVIDIA spec). Budget a 1000W or larger ATX 3.1 power supply.

Verdict: Runs Qwen3.6-27B, the current best single-GPU model, at full Q4 with long context, and the whole 32B class alongside it. Roughly double the used 3090's ceiling; only worth it if you are VRAM-bound or speed-bound today.

💡 Where These Numbers Come From

We do not own these machines. Every tokens-per-second figure on this page is an arithmetic upper bound derived from published memory bandwidth, not a benchmark run. Token generation on a local LLM is memory-bandwidth bound: every weight has to be read once per token. At Q4_K_M that is roughly 0.6GB per billion parameters, so the ceiling is simply memory bandwidth ÷ weight bytes: an RTX 3090 at 936GB/s against a 32B model's ~19.2GB of weights puts the bound near 49 tok/s. Card VRAM, bandwidth, power and MSRP figures come from NVIDIA's and AMD's published specs and launch announcements, Mac memory figures from Apple's tech specs, all re-checked September 2026; VRAM fit is cross-checked against our VRAM guides. Four things pull real output below the bound:

  • • Prompt processing is compute-bound, not bandwidth-bound — long prompts hit a different limit
  • • The KV cache grows with context, taking VRAM and slowing each step
  • • Sampling, framework overhead and quantization layout all cost throughput
  • • CPU-only builds fall furthest below the bound, since DDR5 never reaches its theoretical rate

New to local AI? Start with the Windows installation guide or check which models work on your current hardware in our 8GB RAM guide. Before you buy a card, confirm it is supported in Ollama's official GPU compatibility list.

Do You Need a Dedicated AI Server Instead of a Desktop?

Only if the machine will sit headless on your network and serve models to other devices around the clock. A server trades the desktop's flexibility for 24/7 uptime, low idle draw and more VRAM per dollar (usually via used cards), and it is a different shopping list from every build above — different case, different CPU priorities, different noise and power budget.

We keep that comparison on its own page, the best local AI server guide, rather than duplicating it here: every server tier from a ~$1,000 homelab box to dual-3090 and Mac Studio options, with the VRAM math that decides which models each one can actually hold.

Explore Hardware for 135 AI Models

$600$2,500$5,000$10,000

Your Recommended Build

Developer/Professional Build
$1,899

Ideal for software developers using AI coding assistants

Specifications:

  • • CPU: AMD Ryzen 7 7700X (8-core, 4.5GHz)
  • • RAM: 32GB DDR5-5600 (2x16GB)
  • • GPU: RTX 4070 12GB
  • • Storage: 1TB Samsung 980 PRO NVMe
Check the exact model and quantisation
Use the memory checker below to narrow your choices

Model Shortlist

62 / 135
Matching Model Candidates
Suggested system RAM:32GB

GPU Recommendation:

RTX 4070 12GB / RTX 4070 Ti 16GB
12-16GB VRAM • $600-$800

Measure Performance on Your Machine

This shortlist uses budget, model size and use case. Confirm memory fit and runtime support separately. Speed depends on the exact model, quantisation, context and hardware.

  1. Load your chosen model and run ollama ps to check GPU or CPU placement.
  2. Use the same prompt and context length for each comparison. Exclude the first warm-up run.
  3. With Ollama’s generation API, divide eval_count by eval_duration / 1e9 for generated tokens per second. Record the model, quantisation and hardware alongside the result.

Check model placement · Ollama generation metrics

Model Memory Checker for 135 AI Models

Compare the catalogue’s memory estimates with your hardware. Leave room for the operating system, runtime and context cache; a memory match does not guarantee that a model will run. Cloud services require their provider’s API. Check the exact model file and runtime hardware support before buying.

Select Your Hardware

NVIDIA GPUs

Apple Silicon

Cloud GPUs (Monthly)

Type

gpu

Memory

12GB

Reference price

$799

Within listed memory

65/135

Airoboros 70B

Exceeds listed memory
Required:40GB+
Category:general

Airoboros L2 70B

Exceeds listed memory
Required:40GB+
Category:experimental

Alpaca 7B

Within listed memory
Required:4-8GB
Category:general

Aquila 7B

Within listed memory
Required:4-8GB
Category:general

Baichuan2 13B

Exceeds listed memory
Required:8-16GB
Category:chat

ChatGLM3 6B

Within listed memory
Required:4-8GB
Category:chat

Chronos 70B

Exceeds listed memory
Required:40GB+
Category:experimental

Claude 3 Haiku

Cloud service
Required:Cloud
Category:chat

Claude 3 Opus

Cloud service
Required:Cloud
Category:general

Claude 3 Sonnet

Cloud service
Required:Cloud
Category:general

CodeGemma 7B

Within listed memory
Required:4-8GB
Category:coding

CodeLlama 7B

Within listed memory
Required:4-8GB
Category:coding

CodeLlama 13B

Exceeds listed memory
Required:8-16GB
Category:coding

CodeLlama 34B

Exceeds listed memory
Required:20GB+
Category:coding

CodeLlama 70B

Exceeds listed memory
Required:40GB+
Category:coding

CodeLlama Instruct 7B

Within listed memory
Required:4-8GB
Category:coding

CodeLlama Python 7B

Within listed memory
Required:4-8GB
Category:coding

CodeLlama Python 13B

Exceeds listed memory
Required:8-16GB
Category:coding

CodeLlama Python 34B

Exceeds listed memory
Required:20GB+
Category:coding

Codestral 22B

Exceeds listed memory
Required:16GB+
Category:coding

Coqui TTS

Within listed memory
Required:4-8GB
Category:voice

Whisper Large v3

Within listed memory
Required:10GB
Category:voice

Bark

Check memory headroom
Required:8-12GB
Category:voice

DeepSeek Coder V2 16B

Exceeds listed memory
Required:10-16GB
Category:coding

DeepSeek Coder V2 236B

Exceeds listed memory
Required:100GB+
Category:coding

DeepSeek LLM 7B

Within listed memory
Required:4-8GB
Category:general

Dolphin 2.6 Mistral 7B

Within listed memory
Required:4-8GB
Category:chat

Dolphin 2.6 Mixtral 8x7B

Exceeds listed memory
Required:24GB+
Category:chat

Dolphin Mixtral 8x7B

Exceeds listed memory
Required:24GB+
Category:general

Dragon 7B

Within listed memory
Required:4-8GB
Category:general

Showing 30 of 135 models

View All Models →

Need More Memory for a Downloadable Model?

Compare renting a GPU with buying hardware for your chosen model. Check the model’s licence, runtime support and memory needs first. Hosted API models stay with their provider.

Hardware Requirements by Model Category

Check Memory Before Comparing Speed

Budget for the model weights, context cache and runtime overhead. On a Mac, the operating system shares the same memory pool. For a GPU, verify how much VRAM is actually free. After loading a model, ollama ps shows whether it uses the GPU, CPU or both. Check Ollama’s memory-placement guide.

Tiny & Small (1-7B)

48 Models
  • RAM: 8GB minimum, 16GB recommended
  • CPU: 4+ cores, modern architecture
  • Storage: 50GB+ SSD space
  • Check: Quantisation and space for the context cache
Llama 3.1 8B, Qwen 2.5 7B, Mistral 7B, Gemma 4 E4B, Qwen 2.5 Coder 7B

Medium (8-34B)

48 Models
  • RAM: 32GB minimum, 64GB recommended
  • CPU: 8+ cores, high performance
  • Storage: 100GB+ NVMe SSD
  • Check: Exact model file size and available VRAM
Qwen 3 14B, Phi-4 14B, Gemma 4 12B Unified, Qwen3.6-27B, Qwen 2.5 Coder 32B

Large & Massive (70B+)

36 Models
  • RAM: 64GB minimum, 128GB+ ideal
  • CPU: 16+ cores, server-grade
  • Storage: 200GB+ enterprise SSD
  • Check: GPU residency versus system-RAM offload
Llama 3.3 70B, Qwen 2.5 72B, Llama 3.1 405B, GLM-5.2 (server-class)

Coding Models

26
16GB+ RAM, Fast SSD

Vision Models

7
12GB+ VRAM Required

Chat Models

20
8GB+ RAM, Fast Response

Math Models

5
32GB+ RAM for Precision

Best GPUs for Local AI Acceleration

⭐ Recommended

NVIDIA RTX 4060 Ti 16GB

Best budget GPU for local AI with ample VRAM

  • 16GB VRAM for large models
  • CUDA cores for AI acceleration
  • Runs 13B models smoothly
  • Low power consumption

NVIDIA RTX 4070 Ti

Excellent price/performance for serious AI work

  • 16GB VRAM
  • Superior CUDA performance
  • Handles 30B models
  • DLSS 3 support

NVIDIA RTX 4090 24GB

Professional-grade AI workstation GPU

  • 24GB VRAM for 70B models
  • Fastest inference speeds
  • Professional AI training
  • Future-proof investment

Recommended RAM Upgrades for Local AI

⭐ Recommended

Corsair Vengeance 32GB Kit

Sweet spot for most local AI workloads

  • 2x16GB DDR4-3600
  • Optimized for AMD & Intel
  • Run 13B models comfortably
  • Excellent heat spreaders

G.Skill Ripjaws DDR5 32GB

Latest DDR5 for newest systems

  • 2x16GB DDR5-5600
  • Intel XMP 3.0
  • On-die ECC
  • Future-ready performance

Crucial 64GB DDR5 Kit

Maximum capacity for large models

  • 2x32GB DDR5-6000
  • Run 70B models
  • Premium Samsung B-die
  • RGB lighting

Corsair Vengeance LPX 16GB DDR4

Affordable RAM upgrade for basic AI models

  • 2x8GB DDR4-3200
  • Low profile design
  • XMP 2.0 support
  • Lifetime warranty

Pre-Built Systems for Local AI

HP Victus Gaming Desktop

Ready-to-run AI desktop under $1000

  • AMD Ryzen 7 5700G
  • 16GB DDR4 RAM
  • RTX 3060 12GB
  • 1TB NVMe SSD

Dell Precision 3680 Tower

Professional AI development machine

  • Intel Xeon W-2400
  • 64GB ECC RAM
  • RTX 4000 Ada
  • ISV certified
⭐ Recommended

Mac Mini M2 Pro

Compact powerhouse for local AI

  • M2 Pro chip
  • 32GB unified memory
  • Run 30B models
  • Silent operation

Mac Studio M2 Max

Ultimate Mac for AI workloads

  • M2 Max chip
  • 64GB unified memory
  • Run 70B models
  • 32-core GPU

Can\'t Afford $1,000+ for Hardware? Try Cloud GPUs

Access the same powerful GPUs without the upfront cost. Perfect for testing models, occasional use, or when you need more power than your hardware provides.

Quick Cost Comparison Calculator

Cloud GPU Cost

$10-30/month
No upfront investment

Hardware Cost

$800-1,500 upfront
Plus electricity costs
💡 Recommendation: For 20 hours/month, try Paperspace Free Tier or Vast.ai
Most Popular

RunPod

Affordable cloud GPUs starting at $0.2/hour

  • RTX 4090 at $0.74/hour
  • RTX 3090 at $0.44/hour
  • No setup required
  • Pay per second billing
From $0.2/hour
Save $1,500+ vs buying
Try RunPod
Best Value

Vast.ai

Decentralized GPU marketplace with best prices

  • RTX 4090 from $0.40/hour
  • 50% cheaper than AWS
  • Global availability
  • Instant deployment
From $0.15/hour
Save $2,000+ vs buying
Try Vast.ai
Pro Choice

Lambda Labs

Professional GPU cloud for AI/ML teams

  • A100 80GB available
  • Persistent storage
  • Jupyter notebooks
  • Team collaboration
From $1.10/hour
Enterprise grade
Try Lambda Labs
Free Tier

Paperspace

User-friendly GPU cloud with free tier

  • Free GPU tier available
  • One-click templates
  • AutoML tools
  • Gradient notebooks
Free tier + $0.45/hour
Start free
Try Paperspace

Cloud vs Local: Quick Comparison

AspectCloud GPULocal Hardware
Initial Cost✓ $0 upfront✗ $800-15,000
Scalability✓ Instant scaling✗ Fixed capacity
Maintenance✓ Zero maintenance✗ Your responsibility
Privacy⚠ Data leaves premises✓ 100% local
Latency⚠ Network dependent✓ No network latency
24/7 Usage✗ Expensive✓ Fixed cost

Start with Cloud, Upgrade to Local Later

The smart approach: Test models and learn on cloud GPUs for $20-50/month. Once you know exactly what you need, invest in the right hardware.

🎓 Learn How to Use Cloud GPUs

Step-by-step tutorials showing exactly how to run AI models on cloud GPUs. Start in 5 minutes for just $10.

Reference Build Guides

Detailed component lists optimized for different model sizes and use cases. Parts re-checked September 2026. GPU street prices are volatile right now, so treat GPU line items as the number to verify on the day you buy.

Student Build

$799
48 Models
Supported (up to 7B)
  • • AMD Ryzen 5 5600 (6-core)
  • • 16GB DDR4-3200 RAM
  • • 500GB NVMe SSD
  • • Used RTX 3060 12GB
  • • 550W PSU, mATX case
Best for: Llama 3.1 8B, Phi-4 14B, Gemma 3 12B, Qwen 2.5 Coder 7B
Check GPU placement with ollama ps after loading.

Developer Build

$1,899
89 Models
Supported (up to 34B)
  • • AMD Ryzen 7 7700X (8-core)
  • • 32GB DDR5-5600 RAM
  • • 1TB Samsung 980 PRO
  • • RTX 4070 12GB
  • • 750W Gold PSU
Best for: Qwen 2.5 Coder 14B, Qwen 3 14B, DeepSeek-Coder-V2 Lite
Leave VRAM headroom for your coding context.

AI Researcher

$3,499
16GB VRAM
Up to ~24B dense, 20B MoE
  • • Intel i9-13900K (24-core)
  • • 64GB DDR5-6000 RAM
  • • 2TB Samsung 990 PRO
  • • RTX 4080 16GB
  • • 1000W Platinum PSU
Best for: gpt-oss:20b, Devstral 24B, Qwen 3 14B at Q5
Choose the quantised file before sizing VRAM.

Mac mini M5 Pro

Check current price
24GB Unified
Runs the 7B-14B class well; 48GB and 64GB options reach the 27-32B class
  • • M5 Pro chip (15-core CPU)
  • • 24GB unified memory (307GB/s, Apple spec)
  • • 512GB SSD
  • • 16-core GPU
  • • Silent, tiny, ~no maintenance
Best for: Llama 3.1 8B, Qwen 3 14B, Gemma 3 12B — see the Apple Silicon guide
⚡ Note: memory is fixed at purchase. Confirm the configuration on Apple's Mac mini tech specs before ordering.

Pro Workstation

$5,999
24GB VRAM
+128GB RAM for 70B offload
  • • AMD Threadripper PRO
  • • 128GB ECC RAM
  • • 4TB NVMe RAID
  • • RTX 4090 24GB
  • • 1600W Redundant PSU
Best for: Qwen3.6-27B, Qwen 2.5 Coder 32B, Llama 3.3 70B
RAM offload changes throughput; measure your workload.

Enterprise Server

$10K+
Multi-GPU Planning
Check per-device memory and runtime support
  • • Dual EPYC or Xeon
  • • 256GB+ ECC RAM
  • • 8TB Enterprise SSD
  • • Dual RTX 4090/5090 or A6000
  • • 4U Rackmount
Best for: GLM-5.2 (quantized), Llama 3.1 405B, production serving
⚡ Multiple models simultaneously

RTX 5090 Flagship

$5,500+
32GB VRAM
Every single-GPU model, fast
  • • AMD Ryzen 9 9950X (16-core)
  • • 64GB DDR5-6000
  • • 2TB Gen5 NVMe
  • • RTX 5090 32GB
  • • 1200W ATX 3.1 PSU
Best for: Qwen3.6-27B, Qwen 2.5 Coder 32B; check context-cache memory for long prompts
⚠ NVIDIA's MSRP is $1,999; street prices sit well above it right now, so check current price and only buy if you're VRAM-bound today

Throughput Ceilings by Build Tier

These are not benchmark results. Each figure is the arithmetic upper bound for Q4_K_M decoding on that machine — published memory bandwidth divided by the weight bytes moved per token. Treat them as the speed a tier cannot exceed, and expect real output somewhere below.

Arithmetic Ceiling — Llama 3.1 8B at Q4_K_M (~4.9GB of weights)

Ultimate Workstation (RTX 5090, 1792GB/s)366 tok/s upper bound
366
Performance 16GB (RTX 5070 Ti, 896GB/s)183 tok/s upper bound
183
Sweet Spot Build (RTX 5060 Ti, 448GB/s)91 tok/s upper bound
91
Budget Build (CPU only, ~90GB/s DDR5-5600)18 tok/s upper bound
18
MacBook Pro M3 Max (400GB/s, Apple spec)82 tok/s upper bound
82
Hardware ConfigurationModel (Q4_K_M)Memory BandwidthWeights Read Per TokenArithmetic Ceiling
Budget Build (Ryzen 5 9600X, 32GB DDR5, CPU only)Llama 3.1 8B~90GB/s (dual-channel DDR5-5600)~4.9GB~18 tok/s
Sweet Spot Build (Ryzen 5 9600X, 32GB, RTX 5060 Ti 16GB)Llama 3.1 8B448GB/s~4.9GB~91 tok/s
Sweet Spot Build (Ryzen 5 9600X, 32GB, RTX 5060 Ti 16GB)Qwen 2.5 Coder 14B448GB/s~8.4GB~53 tok/s
Performance 16GB (Ryzen 7 9700X, 32GB, RTX 5070 Ti 16GB)Qwen 3 14B896GB/s~8.4GB~107 tok/s
Ultimate Workstation (Ryzen 9 9950X, 64GB, RTX 5090 32GB)Qwen3.6-27B1792GB/s~16.2GB~111 tok/s
AI Server (Ryzen 7 7700, 64GB, dual used RTX 3090 — 48GB pooled)Llama 3.3 70B936GB/s per card (layers run serially, so one card's bandwidth sets the pace)~42.5GB~22 tok/s

* Ceilings, not benchmarks: bandwidth ÷ weight bytes, using NVIDIA's and AMD's published bandwidth figures and ~0.6GB of weights per billion parameters at Q4_K_M. Real throughput is lower once prefill, KV cache growth and sampling overhead are included. The 70B row is the dual-3090 tier, not a single card: a 70B at Q4_K_M is ~42.5GB of weights, so 48GB of pooled VRAM is the floor — see the local AI server comparison for the full breakdown.

GPU Memory Bandwidth Comparison

Bandwidth is the number that decides decoding speed, so it is the honest way to rank cards. The ceiling column is bandwidth ÷ ~8.4GB, the weights a 14B model moves per token at Q4_K_M.

GPUVRAMMemory Bandwidth14B CeilingLargest Dense Model at Q4
RTX 306012GB360GB/s~43 tok/s14B class
RTX 407012GB504GB/s~60 tok/s14B class
RTX 5060 Ti (16GB)16GB448GB/s~53 tok/s14B class, 20B MoE
RX 9070 XT (AMD)16GB640GB/s~76 tok/s14B class, 20B MoE
RTX 4080 Super16GB736GB/s~88 tok/s20-24B
RTX 5070 Ti16GB896GB/s~107 tok/s20-24B
RTX 508016GB960GB/s~114 tok/s20-24B
RTX 3090 (used)24GB936GB/s~111 tok/s27-32B
RTX 409024GB1008GB/s~120 tok/s27-32B
RTX 509032GB1792GB/s~213 tok/s32B with long context; ~49B is tight

Bandwidth and VRAM are NVIDIA's and AMD's published specs. Ceilings are arithmetic upper bounds, not measured throughput. Street prices are running far above MSRP right now — see why GPU prices are so high.

Hardware FAQ

Do I need a GPU for local AI?

Not necessarily. Modern CPUs can run smaller models (3B-8B) effectively. However, a GPU provides 2-5x speed improvements and enables running larger models more efficiently. If you plan to use AI regularly or work with larger models, a GPU is highly recommended.

How much RAM do I really need?

RAM is crucial for local AI. As a rule of thumb: model size + 4-8GB for the operating system. For an 8B model (~5GB), you need at least 12GB RAM, but 16GB+ is recommended for smooth operation. For 70B models, you need 64GB+ RAM.

Should I build a desktop or a dedicated AI server?

Build a desktop if you are the only user and you want the machine to do other things too — every build on this page fits that case. Go headless-server only when models need to be available on your network 24/7, which changes the CPU, case, noise and power calculus entirely. Tier-by-tier server options (homelab, dual-3090, Mac Studio) are compared on the best local ai server page, with the single-box parts list in the $1,500 AI server build guide.

Is Apple Silicon (M-series) good for AI?

Yes! Apple Silicon offers excellent AI performance with unified memory architecture. A base Mac Mini M4 (16GB) handles 7-8B models, M4 Pro/Max configurations run the 14B-70B range depending on memory, and unified memory means the model shares one big pool instead of fighting a VRAM ceiling. See the Apple M4 local AI guide for chip-by-chip picks.

Can I upgrade my existing computer?

Often yes! The most impactful upgrades are usually RAM (if your motherboard supports more) and adding a GPU. However, very old CPUs (pre-2018) may become bottlenecks. Check your motherboard specifications for RAM and GPU compatibility.

Which models can I run with my hardware?

Start by checking the Local AI Models directory to filter by parameters, modality, and context window that match your build. If you're on a lean system, jump into the 8GB optimization guide for hand-picked quantized models before upgrading to larger tiers.

Was this helpful?

Get Hardware Updates & Deals

Get the latest hardware recommendations, price and availability updates, and deals delivered weekly, written for people running models at home.

Reading now
Join the discussion

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📅 Published: 2025-10-28🔄 Last Updated: September 18, 2026✓ Manually Reviewed

Related Guides

Continue your local AI journey with these comprehensive guides

Free Tools & Calculators