★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own it all: Lifetime $149, pay once
Hardware

Strix Halo / AMD Ryzen AI Max+ 395 for Local AI (2026): 128GB Unified Memory in a Mini PC

May 1, 2026
22 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Got the hardware sorted? Now build on it. You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Start free
Or own it for life — Lifetime $149, pay once

The verdict by memory tier: buy the 128 GB configuration or do not buy Strix Halo. At 128 GB, with the 96 GB GPU allocation that Framework and HP state as the maximum, a Ryzen AI Max+ 395 box holds a 70B model at Q4_K_M (about 40 GB) with generous context, a 70B at Q8_0 (about 75 GB), 100B-class MoE models at 4-bit, and the 82.5 to 91 GB builds of DeepSeek V4-Flash. At 64 GB you get 32B-class models at Q4 or Q8 and a 70B at Q4 with little room to spare. At 32 GB you have an ordinary 14B-class machine that a 16 GB discrete GPU would beat on speed. The trade in every tier is the same: 256 GB/s of memory bandwidth, about a quarter of an RTX 4090, so anything that fits in a discrete card's VRAM runs faster there.

Updated 19 September 2026: prices replaced with Framework's current MSRP, speed tables replaced with bandwidth ceilings (we do not own this hardware and the old figures could not be sourced), the 70B BF16 claim removed (a 70B at BF16 is about 141 GB and does not fit in 128 GB), and the Ryzen AI Max+ PRO 495 / 192 GB tier added.

This guide covers the hardware spec, the systems that ship it (Framework Desktop, Asus ROG Flow Z13, HP ZBook Ultra G1a and Z2 Mini G1a, GMKtec EVO-X2), ROCm setup for gfx1151, BIOS memory allocation, what the bandwidth caps, how it compares with the current Mac Studio and with discrete NVIDIA cards, Ollama / vLLM / llama.cpp recipes, the XDNA 2 NPU situation, and the workloads where Strix Halo wins or loses. If you are still deciding between a unified-memory box and a discrete card, the hardware guide is the wider map.

Table of Contents

  1. What Strix Halo Is
  2. Why 128 GB Unified Memory Matters
  3. Hardware Specs
  4. The 2026 Additions: Max+ 392 / 388 and the PRO 400 series
  5. Available Systems (Framework, Asus, HP, GMKtec)
  6. What 256 GB/s Actually Caps
  7. vs Mac Studio
  8. vs Discrete NVIDIA GPUs
  9. BIOS: Allocating Memory to the iGPU
  10. ROCm Setup for gfx1151
  11. Ollama on Strix Halo
  12. llama.cpp Native Build
  13. vLLM-ROCm
  14. PyTorch + Hugging Face
  15. Image and Video Generation
  16. The XDNA 2 NPU
  17. Power, Thermals, Acoustics
  18. Use Cases Where Strix Halo Wins
  19. Where Discrete GPUs Still Win
  20. Buying Advice
  21. Troubleshooting
  22. FAQ

Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

What Strix Halo Is

Strix Halo is the codename for AMD's Ryzen AI Max 300 series, the high-end APU platform sold as Ryzen AI Max+ 395 with lower-tier siblings (Max 390, Max 385). Architecture of the 395:

  • CPU: 16 Zen 5 cores / 32 threads, up to 5.1 GHz boost
  • iGPU: Radeon 8060S, 40 RDNA 3.5 compute units, up to 2.9 GHz, ROCm target gfx1151
  • NPU: XDNA 2, 50 TOPS INT8 (separate from CPU and iGPU)
  • Memory: up to 128 GB LPDDR5X-8000, soldered, shared by CPU and iGPU
  • Memory bus: 256-bit, which at 8,000 MT/s is 256 GB/s of theoretical bandwidth
  • GPU-addressable memory: up to 96 GB (Framework and HP both state this figure)
  • TDP: configurable; the 45 to 120 W range is widely reported and Framework's Desktop runs a 120 W ceiling

The core counts, clocks, GPU model and memory type above are taken from Framework's and HP's product pages for the machines that use the chip; AMD's own spec page did not load when this page was refreshed, so the TDP row is the one figure here we could not re-verify at source.

Released early 2025. In January 2026 AMD added the Max+ 392 and Max+ 388, and on 20 May 2026 it announced the Ryzen AI Max PRO 400 series (codename Gorgon Halo), whose top part, the Max+ PRO 495, lifts the memory ceiling to 192 GB. See the 2026 additions section.


Why 128 GB Unified Memory Matters

For local LLMs the first constraint is whether the model fits in the memory the GPU can reach. Discrete consumer GPUs stop at 24 GB (RTX 3090, 4090, 7900 XTX) or 32 GB (RTX 5090). At 24 GB you can run:

  • 8B models in FP16
  • 14B models in Q4 to Q8
  • 32B models in Q4
  • 70B models only with CPU offload (slow)
  • 70B models in Q8 or FP16: no

A 128 GB Strix Halo box with 96 GB assigned to the GPU unlocks:

  • 70B models at Q4_K_M (about 40 GB) with long context, or at Q8_0 (about 75 GB) with room for context
  • 100B-class MoE models at 4-bit
  • The 82.5 to 91 GB DeepSeek V4-Flash builds (see DeepSeek V4 hardware requirements)
  • Long contexts (131K) on 70B-class models with a quantised KV cache

What it does not unlock, despite what earlier versions of this page said, is a 70B model at BF16. Llama 3.1 70B in BF16 is about 141 GB of weights, which is larger than the whole memory pool. The 128 GB tier is a Q8-and-below tier for 70B.

The trade-off: 256 GB/s of unified bandwidth is much lower than discrete VRAM (1,008 GB/s on an RTX 4090, 1,792 GB/s on an RTX 5090; NVIDIA's product pages print bus width and memory type, and these are the standard spec-sheet bandwidths derived from them), so per-token speed is lower on every model that would have fit the discrete card. For models that would not fit at all, the comparison is against CPU offload, and there Strix Halo wins.


Hardware Specs

SpecRyzen AI Max+ 395Ryzen AI Max 390Ryzen AI Max 385
Cores / threads16 / 3212 / 248 / 16
Boost clock (CPU)up to 5.1 GHzup to 5.0 GHzup to 5.0 GHz
iGPURadeon 8060SRadeon 8050SRadeon 8050S
iGPU CUs403232
iGPU clockup to 2.9 GHzup to 2.8 GHzup to 2.8 GHz
NPU TOPS (INT8)505050
MemoryLPDDR5X-8000LPDDR5X-8000LPDDR5X-8000
Max memory128 GB128 GB128 GB (Framework sells it at 32 GB)
Memory bandwidth (256-bit x 8,000 MT/s)256 GB/s256 GB/s256 GB/s

Sources: Framework Desktop configuration page (385 and 395 clocks, CU counts, memory), HP ZBook Ultra G1a page (PRO 390 and PRO 385 core and CU counts). For local AI the 395 is the one to buy among the original SKUs: fewer CUs on the 390 and 385 mean proportionally lower compute, and the bandwidth is the same on all three. Always pair it with 128 GB.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

The 2026 Additions: Max+ 392 / 388 and the PRO 400 Series

Short answer: two things changed in 2026. At CES in January AMD added the Ryzen AI Max+ 392 (12 cores) and Max+ 388 (8 cores), which keep the full 40-CU iGPU. On 20 May 2026 AMD announced the Ryzen AI Max PRO 400 series, codename Gorgon Halo, whose top part, the Ryzen AI Max+ PRO 495, supports 192 GB of LPDDR5X-8533 at 273 GB/s. Framework has a 192 GB Desktop with the PRO 495 listed as "coming soon" with no price or date. The 395 itself is unchanged.

SpecMax+ 395 (2025)Max+ 392 / 388 (Jan 2026)Max+ PRO 495 (May 2026)
CPU cores / threads16 / 3212 / 24 and 8 / 1616 / 32
Boost clockup to 5.1 GHzup to 5.0 GHzup to 5.2 GHz
iGPURadeon 8060S, 40 CURadeon 8060S, 40 CURadeon 8065S, 40 CU
NPU (INT8)50 TOPS50 TOPS55 TOPS
MemoryLPDDR5X-8000LPDDR5X-8000 (8533 validated on some SKUs)LPDDR5X-8533
Max unified memory128 GB128 GB192 GB
Bandwidth256 GB/s256 GB/s (273 GB/s at 8533)273 GB/s (Framework's stated figure)

The 392 and 388 matter because they give the full 40-CU Radeon 8060S to cheaper, lower-core-count parts. For LLM inference, which is iGPU- and bandwidth-bound, a 12-core 392 has the same ceiling as the 16-core 395.

The PRO 495 matters for one reason: 192 GB. That is the tier where a 70B model at Q8 fits with a long context, where the 128 GB DeepSeek V4-Flash UD-Q3_K_XL build fits with room to spare, and where the 155 GB native-precision V4-Flash build becomes possible on one box. The 7 percent bandwidth lift from 8533 memory is real but small; the capacity jump is the reason to wait if you can. Framework's 192 GB page states 273 GB/s and 131 TOPS of combined AI compute, and nothing else yet.

Ryzen AI 400 is a different, non-Halo line. Also announced at CES 2026, the Ryzen AI 400 mobile series (Gorgon Point) is a dual-channel thin-and-light laptop platform on Strix Point and Krackan Point silicon. Despite the higher number it is not a 128 GB quad-channel unified-memory part. If you want to run 70B locally you want a Max+ (395, 392, or PRO 495), not a Ryzen AI 400.


Available Systems (Framework, Asus, HP, GMKtec)

Systems shipping Strix Halo as of 19 September 2026. Only Framework publishes a fixed list price; for the others the vendor page is the price.

SystemForm factorMemory optionsPrice
Framework Desktop (DIY Edition)4.5-litre mini-ITX-style desktop32 GB (Max 385) / 64 GB / 128 GB (Max+ 395)$1,269 / $1,959 / $3,449 MSRP, excluding tax; listed out of stock on 19 Sep 2026
Framework Desktop, 192 GB (Max+ PRO 495)same chassis192 GB"Coming soon", no MSRP published
Asus ROG Flow Z13 (GZ302, 2025)13-inch 2-in-1 tablet128 GB LPDDR5X-8000 (Max+ 395 SKU)check current price
HP ZBook Ultra G1a 14-inchMobile workstationup to 128 GB (PRO 395 / 390 / 385)check current price
HP Z2 Mini G1aMini workstation, 300 W internal PSUup to 128 GB (up to PRO 395)check current price
GMKtec EVO-X2Mini-PC64 / 128 GB LPDDR5X-8000check current price

Best value at 128 GB: still the Framework Desktop, but the value story has changed. The 128 GB configuration launched at $1,999 in February 2025 and is $3,449 MSRP today because LPDDR5X pricing moved; that is the same memory squeeze that pushed Apple to raise its own memory upgrade prices. Framework's chassis is the closest to open hardware in this segment, has full Linux support, and the mainboard is sold separately for people who want to build their own case.


What 256 GB/s Actually Caps

We do not own a Strix Halo machine and this page no longer prints tokens-per-second figures we cannot source. What can be stated exactly is the ceiling. Token generation is memory-bandwidth bound: every generated token reads the active weights once, so the upper bound is bandwidth divided by bytes read per token. The table uses 256 GB/s and approximate file sizes for common quantisations; real throughput is always below the ceiling, on this platform and every other.

Model and quantisationApprox. weights read per tokenCeiling at 256 GB/sFits in 96 GB GPU allocation?
8B, Q4_K_M~4.9 GB~52 tok/sYes
14B, Q4_K_M~9 GB~28 tok/sYes
32B, Q4_K_M~20 GB~13 tok/sYes
70B, Q4_K_M~40 GB~6 tok/sYes, with room for context
70B, Q8_0~75 GB~3.4 tok/sYes, tight
70B, BF16~141 GBdoes not fitNo
DeepSeek V4-Flash UD-IQ1_S (82.5 GB, 13B of 284B active)~3.8 GB (active share of the file)~68 tok/s ceiling, but only if all experts are residentYes

Two things to read off that table. First, for dense 70B models the honest number is single digits, on Strix Halo and on any other 256 GB/s machine; a Mac Studio M5 Max at 614 GB/s has a ceiling of about 15 tok/s on the same file, an M5 Ultra at 1.2 TB/s about 30. Second, MoE models are where a 128 GB box shines: the file must be resident, but only the active experts are read per token, so the ceiling is that of a much smaller model. That is why the platform's best use is large MoE models, not dense 70B models.

For the file sizes behind these rows see the Ollama model RAM and VRAM table.


vs Mac Studio

Apple's current Mac Studio (announced 25 August 2026, shipping from 22 September) comes with M5 Max or M5 Ultra. Figures are Apple's published specs and US MSRPs.

AspectStrix Halo (Framework Desktop 128 GB)Mac Studio M5 MaxMac Studio M5 Ultra
Starting MSRP$3,449 (128 GB, DIY)$2,499 (36 GB base)$5,499 (96 GB base)
Memory options32 / 64 / 128 GB36 / 48 / 64 / 128 GB96 / 256 / 512 GB
Memory bandwidth256 GB/s460 GB/s (32-core GPU) or 614 GB/s (40-core GPU)1.2 TB/s
70B Q4 ceiling (40 GB file)~6 tok/s~15 tok/s~30 tok/s
GPU-addressable memoryup to 96 GBwhole pool minus macOSwhole pool minus macOS
OSLinux / WindowsmacOSmacOS
Software ecosystemROCm, vLLM, llama.cpp, OllamaMLX, llama.cpp, OllamaMLX, llama.cpp, Ollama
Form factor4.5-litre desktopStudioStudio

The Mac wins on bandwidth at every tier and on capacity at the top, where 256 GB and 512 GB configurations exist and Strix Halo has nothing to answer until the 192 GB PRO 495 machines ship. What Strix Halo offers is an x86 Linux box with ROCm and vLLM, and a 128 GB configuration whose MSRP is below what a 128 GB Mac Studio configures to (Apple does not publish the upgrade price on its specs page; check the configurator). For the Mac side of this decision see Best Mac for Local AI.


vs Discrete NVIDIA GPUs

The comparison that matters is bandwidth against capacity. The RTX 4090 is a 24 GB card on a 384-bit bus (1,008 GB/s at its GDDR6X speed) and the RTX 5090 a 32 GB card on a 512-bit bus (1,792 GB/s at its GDDR7 speed); NVIDIA's pages print the bus widths, the bandwidths are the standard spec-sheet figures.

  • Anything that fits in 24 or 32 GB has a ceiling four to seven times higher on the discrete card. An 8B Q4 model that caps at about 52 tok/s on Strix Halo caps at about 200 tok/s on a 4090 and 365 tok/s on a 5090.
  • A 70B at Q4 (about 40 GB) does not fit either card, so on the NVIDIA side it runs partly from system RAM at DDR5 speeds, and Strix Halo's fully resident 6 tok/s ceiling starts to look good.
  • Two 4090s (48 GB) hold a 70B at Q4 with room for context and keep the bandwidth advantage, at several times the power draw and a price that depends on the used-card market; check current prices. The dual 3090 vs single 5090 build works through that maths.
  • A 16 GB card is a different class of machine altogether; the best models for an RTX 5080 list is where that budget goes.

Strix Halo wins on 70B-at-Q8 capability, power, noise and form factor. Discrete cards win on speed for every model that fits them, on image generation, and on the CUDA ecosystem.


BIOS: Allocating Memory to the iGPU

By default, Strix Halo systems hand a small carve-out to the iGPU and leave the rest as system RAM. To run 70B models on the iGPU you raise that allocation. Framework and HP both state a 96 GB maximum for GPU-assigned memory on a 128 GB system.

Most BIOSes expose this under Advanced > AMD CBS > NBIO Common Options > GFX Configuration > UMA Frame Buffer Size. Framework exposes a simpler memory reservation menu. Reboot after changing it.

WorkloadSuggested iGPU allocation on 128 GB
LLMs only96 GB (the vendor maximum; leaves 32 GB for the OS and runtime)
LLM + image generation80 GB (more system RAM for ComfyUI buffers)
Mixed workstation64 GB

On Linux, llama.cpp can additionally use GTT memory beyond the BIOS carve-out; that is a driver detail, not a vendor-stated capacity, so plan around 96 GB.


ROCm Setup for gfx1151

ROCm 6.3 and newer include Strix Halo (gfx1151):

wget https://repo.radeon.com/amdgpu-install/6.3/ubuntu/jammy/amdgpu-install_6.3.60300-1_all.deb
sudo apt install ./amdgpu-install*.deb
sudo amdgpu-install --usecase=rocm,hiplibsdk -y
sudo usermod -aG render,video $USER
sudo reboot

Some workloads still need the gfx version override:

echo 'export HSA_OVERRIDE_GFX_VERSION=11.5.1' >> ~/.bashrc
echo 'export HCC_AMDGPU_TARGET=gfx1151' >> ~/.bashrc
source ~/.bashrc

Verify:

rocminfo | grep gfx
# Expected: gfx1151 (Strix Halo iGPU)
rocm-smi --showmeminfo vram
# Should show up to ~96 GB available depending on BIOS allocation

Newer ROCm releases exist; the 6.3 line is the first with gfx1151 and the commands above track whatever version you install. The AMD ROCm setup guide covers the general case.


Ollama on Strix Halo

curl -fsSL https://ollama.com/install.sh | sh

Edit systemd to set the override:

sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf <<EOF
[Service]
Environment="HSA_OVERRIDE_GFX_VERSION=11.5.1"
Environment="HCC_AMDGPU_TARGET=gfx1151"
EOF
sudo systemctl daemon-reload
sudo systemctl restart ollama

# Run a 70B model at Q4: about 40 GB, fits in the 96 GB GPU allocation
ollama run llama3.1:70b

For a 70B at Q8 (about 75 GB, fits with a 96 GB allocation and a modest context):

ollama run llama3.1:70b-instruct-q8_0

Do not pull the fp16 tag of a 70B on this machine; at about 141 GB it will not load.


llama.cpp Native Build

git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp

HIPCXX="$(hipconfig -l)/clang" \
HIP_PATH="$(hipconfig -R)" \
cmake -B build \
    -DGGML_HIP=ON \
    -DAMDGPU_TARGETS=gfx1151 \
    -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

./build/bin/llama-cli -m llama-3.1-70b-instruct-Q4_K_M.gguf -ngl 999 -fa

For long context: enable a quantised KV cache (--cache-type-k q8_0 or q4_0) to halve or quarter KV memory.


vLLM-ROCm

docker pull rocm/vllm:latest

docker run --device /dev/kfd --device /dev/dri \
    --group-add video --group-add render \
    --security-opt seccomp=unconfined \
    --shm-size 16G \
    -e HSA_OVERRIDE_GFX_VERSION=11.5.1 \
    -p 8000:8000 \
    rocm/vllm:latest \
    vllm serve casperhansen/llama-3.1-70b-instruct-awq \
    --quantization awq \
    --max-model-len 16384 \
    --gpu-memory-utilization 0.85

vLLM-ROCm on Strix Halo is functional but bandwidth-bound: single-stream decode has the same ceiling as llama.cpp or Ollama, and multi-user batching has far less headroom than a discrete GPU with four to seven times the bandwidth.


PyTorch + Hugging Face

pip install torch torchvision --index-url https://download.pytorch.org/whl/rocm6.3

Run any HF Transformers model that fits:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-32B-Instruct",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

A 32B model in BF16 is about 65 GB and loads inside the 96 GB GPU allocation. A 70B in BF16 does not fit; load it quantised or use llama.cpp. Training is impractical on this platform (compute-bound, not memory-bound).


Image and Video Generation

ComfyUI and Automatic1111 run on ROCm with gfx1151, and the memory pool means SDXL, Flux and the current video models (Wan, HunyuanVideo) load unquantised without offload tricks. That is the good news.

The bad news is that diffusion is compute-bound, and 40 RDNA 3.5 compute units are a fraction of a discrete card's shader throughput, with no FP8 hardware path. Expect image generation to be several times slower than on an RTX 4090 and video generation to be slow enough that routine use is impractical. Earlier versions of this page printed per-workflow timings; they were removed because we cannot source them. If image or video generation is your primary workload, buy a discrete GPU.


The XDNA 2 NPU

The 50-TOPS XDNA 2 NPU is largely unused for general LLM inference as of September 2026. AMD's Lemonade SDK and Ryzen AI Software target it for specific small models, but Ollama, llama.cpp and vLLM all run on the iGPU.

The NPU is for:

  • Always-on background AI (Windows Studio Effects, Recall)
  • Battery-sensitive laptop workloads
  • Specific quantised models AMD ships optimised kernels for

For mainstream LLM inference, ignore the NPU and use the iGPU.


Power, Thermals, Acoustics

The part is configurable across a wide TDP range and vendors expose profiles in BIOS or a utility. Because decode is bandwidth-bound, a lower power cap costs less generation speed than you would expect; prompt prefill and image generation, which are compute-bound, feel the cap more. For a 24/7 inference box a mid profile is the usual compromise between fan noise and prefill speed. We have no measured tok/s per profile to print, and the bandwidth ceiling is the same at every profile.


Use Cases Where Strix Halo Wins

  1. 70B at Q4 or Q8 on a single small box with no offload and no multi-GPU rig.
  2. Large MoE models: the 82.5 to 91 GB DeepSeek V4-Flash builds, 100B-class MoE at 4-bit. Resident file, small active set, high ceiling.
  3. Multi-tenant home server: 96 GB of GPU memory holds several medium models at once.
  4. Quiet, efficient desktop for usable 70B inference at low power.
  5. Privacy or air-gapped 70B: one mini-PC, easy to lock down. See Air-Gapped AI Deployment.
  6. Long-context RAG and agents on 70B with a quantised KV cache.
  7. A Linux/x86 alternative to the Mac Studio for people who need ROCm, vLLM or a PC software stack.

Where Discrete GPUs Still Win

  1. Speed on anything that fits in VRAM: an RTX 4090's 1,008 GB/s gives a ceiling four times higher than Strix Halo's 256 GB/s at the same model size; an RTX 5090's 1,792 GB/s, seven times.
  2. Image and video generation: compute-bound, and the discrete card has many times the shader throughput.
  3. Multi-user serving: vLLM's batching scales with bandwidth.
  4. Fine-tuning: compute throughput dominates.
  5. TensorRT-LLM and the CUDA ecosystem: NVIDIA only.
  6. FP8 hardware: RDNA 3.5 has none.
  7. Time to first token on long prompts.

Buying Advice

Buy Strix Halo (128 GB) if:

  • You want to run 70B models at Q4 or Q8 locally without offload.
  • Large MoE models are on your list.
  • You value a quiet, low-power mini-PC.
  • You are comfortable with Linux + ROCm or Windows + WSL2.
  • LLMs are the primary workload; image generation is occasional.

Wait for the 192 GB PRO 495 systems if:

  • You want the 128 GB DeepSeek V4-Flash Q3 builds, or the 155 GB native build, on one box.
  • You can wait; Framework lists it as coming soon with no date.

Buy a discrete GPU instead if:

  • You primarily run 7B to 32B models.
  • You do heavy image or video generation.
  • You need TensorRT-LLM or ExLlamaV2.
  • You serve many concurrent users.

Buy a Mac Studio instead if:

  • Bandwidth matters more than price: 614 GB/s (M5 Max) or 1.2 TB/s (M5 Ultra) against 256 GB/s.
  • You need 256 GB or 512 GB in one machine today.
  • You are already in the Apple ecosystem or want MLX.

Troubleshooting

SymptomCauseFix
iGPU only sees 16 GBBIOS UMA too lowRaise the allocation toward the 96 GB maximum in BIOS
Ollama uses CPU onlygfx version mismatchSet HSA_OVERRIDE_GFX_VERSION=11.5.1
ROCm "no agents found"Driver / groupsAdd render and video groups, reboot
70B fp16 tag fails to loadFile is about 141 GB, larger than the poolUse the Q8_0 or Q4_K_M tag
Throttling under sustained loadTDP / thermalRaise the cTDP profile or improve cooling
WSL2 GPU not visibleAMD WSL driver missingInstall AMD Software for WSL on the host
Long prompt prefill slowCompute-boundExpected; Strix Halo trades compute for memory

FAQ

See answers to common Strix Halo questions below.


Sources: Framework Desktop configuration page and Framework Desktop overview (clocks, CU counts, LPDDR5x-8000, 256-bit bus, 96 GB graphics-addressable memory, MSRPs, 192 GB / 273 GB/s tier) | Asus ROG Flow Z13 (2025) specifications | HP ZBook Ultra G1a and HP Z2 Mini G1a (PRO 395/390/385 core and CU counts, 96 GB GPU memory) | GMKtec EVO-X2 | AMD Ryzen AI Max+ PRO 495 product page | Apple Mac Studio tech specs and Apple Newsroom, 25 August 2026 | ROCm documentation. Bandwidth ceilings are arithmetic on vendor-stated bandwidth and approximate file sizes, not measurements.

Related guides:

🎯
AI Learning Path

Got the hardware sorted? Now build on it.

You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Once your hardware is sorted

Decide before you spend a thousand pounds

The AI Hardware course sizes your build properly — VRAM ladder, real bottlenecks, budget builds — and Pick the Right Model tells you what to run on it.

$149 once unlocks everything, forever — about $0.27/chapter for life. Prefer to spread it out? Pro is $79/year (saves 27%) or $8.99/month.
Secure checkout by Lemon Squeezy — your card never touches this siteInstant access the moment you payFirst chapter of every course is free — try before you buy

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
Tagsstrix haloryzen ai maxamdunified memorymini pcframework desktopgfx1151

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path
More on Local AI Hardware
See the full AI Hardware Guide 2026 guide.

Comments (0)

No comments yet. Be the first to share your thoughts!

📅 Published: May 1, 2026🔄 Last Updated: September 19, 2026✓ Manually Reviewed

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks. Includes Strix Halo-tuned Ollama config. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Was this helpful?

LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Got the hardware sorted? Now build on it.

You know what to buy — the courses show you what to actually run, fine-tune, and ship on it. First chapter free, no card.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators