★ Reading this for free? Get 25 structured AI courses + per-chapter AI tutor — the first chapter of every course free, no card.Start free in 30 secondsOr own every course: $149 once
Troubleshooting

Local Reasoning Model Never Finishes: Fix Think Loops

October 4, 2026
13 min read
LocalAimaster Research Team

Want to go deeper than this article?

Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.

📚AI Learning Path

Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.

Start free
Or own it for life — Lifetime $149, pay once

Four different things look identical from the outside — a model that "never finishes" — and each has a different fix. Runaway generation is a budget problem. Empty content with think: false is usually a version bug. Leaked <think> or <|im_end|> tags are a template/parser problem. And the same request behaving differently on /api/chat versus /v1/chat/completions is a field-name problem, because the two endpoints accept disjoint value sets for the same feature.

Work in that order and you will localise the fault in about two minutes instead of swapping quants for an afternoon. Everything below is dated and versioned, because several of these are live regressions rather than settled behaviour: Ollama's current release at the time of writing is v0.32.14, published 2026-08-15 (GitHub releases), and the llama.cpp flags quoted are from common/arg.cpp on master, read 2026-08-18.


Which Fault Do You Have?

Start here — match the symptom, then jump to the section. Do not change sampler settings before you have done this; temperature is almost never the cause.

What you seeMost likely faultSection
Generates for minutes, no answer, never returnsUnbounded reasoning + a prompt it cannot satisfyFault A
content: "", done_reason: "stop", low eval_countEngine/model bug on think: falseFault B
Long thinking, then stops with empty contentModel/quant fails the thinking→content transitionFault B
Literal <think> or `<im_end>` in the visible reply
Works on one endpoint, 400s or ignores you on the otherField name / value set divergenceFault D

The reason these get confused is that a chat UI shows you only the last box. Stop using the UI to diagnose. Every step below uses curl against the raw API so you can see done_reason, eval_count and the thinking field, which is where the actual evidence lives.


Reading articles is good. Building is better.

Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.

Step 1: Read the Template the Model Is Actually Using

Before you change anything, print the template and the stop parameters — a surprising share of "broken model" reports are a repackaged GGUF with a template that does not match its own EOS token.

Ollama ships this in the CLI. These flags are what ollama show --help prints:

ollama show --template   qwen3.5:8b   # the Jinja chat template
ollama show --parameters qwen3.5:8b   # PARAMETER lines, including stop sequences
ollama show --modelfile  qwen3.5:8b   # everything the blob carries

What you are looking for, in order:

  1. Does the template emit a reasoning block at all? If the template has no <think> handling but the model was trained to produce one, the tags have nowhere to go and end up in your content.
  2. Do the stop parameters match the tokens the template closes with? A ChatML-style template that closes turns with <|im_end|> needs that token registered as EOS. If ollama show --parameters is empty and the template uses <|im_end|>, you have found your leak.
  3. Is there a num_predict or num_ctx set? Usually not — and that matters for Fault A.

The llama.cpp equivalent is to let the model's own metadata drive the template rather than guessing:

llama-server -m model.gguf --jinja           # use the template from model metadata
llama-server -m model.gguf --jinja --chat-template-file ./fixed.jinja   # override, only if you have proven the packaged one wrong

Per common/arg.cpp, --chat-template and --chat-template-file default to "template taken from the model's metadata". Override last, not first: an override that disagrees with the tokenizer produces exactly the same symptoms you are trying to fix.

If none of this is familiar territory, our complete Ollama guide covers Modelfiles and parameters from the ground up.


Fault A: It Reasons Forever and Never Answers

The cause is almost never the model — it is that nothing bounds the reasoning, and your prompt contains something the model keeps failing to verify. Ollama's num_predict defaults to unlimited, and llama.cpp's --n-predict documents its default as "-1, -1 = infinity" (common/arg.cpp). Give a self-doubting model infinite rope and it will use it.

The reproducible trigger: exact word counts

This is the one that has been isolated cleanly. ollama/ollama issue #17512 (filed 2026-08-02) reports that with thinking on, a prompt specifying an exact word count sends the model into "an unbounded self-verification loop: it drafts, attempts to count words, decides the count is wrong, and rewrites — indefinitely, never emitting a stop token." The reporter's controlled table, same models and machine, varying only thinking mode:

PromptThinkinggemma4:12bqwen3.5:9b
"Write a 200-word story about a lighthouse."onnever terminatesnever terminates
"Write a 200-word story about a lighthouse."off229 tokens, 46s320 tokens, 45s
"Write a short story about a lighthouse."on1,423 tokens, terminates2,916 tokens, terminates

Source: ollama/ollama issue #17512, reproduced on Ollama 0.32.5 on an Intel Core Ultra 7 255H across CPU, Vulkan and SYCL backends. That last row is the tell: thinking on is fine on its own. Three conditions had to line up — thinking on, an exact count, and num_predict left at its default of -1.

The secondary damage is worth knowing about: with OLLAMA_NUM_PARALLEL=1, one stuck request holds the only slot forever and every later request queues behind it, so the whole server looks hung to clients that are perfectly healthy.

Fixes, in the order to try them

  1. Rewrite the instruction. "About 200 words" or "roughly 4 short paragraphs" instead of "exactly 200 words". Reasoning models cannot count tokens, so an exact count is an unsatisfiable verification target. This costs nothing and fixes the documented case outright.
  2. Bound the generation. Nothing should ever run unbounded in production:
curl -s http://localhost:11434/api/chat -d '{
  "model": "qwen3.5:8b",
  "messages": [{"role":"user","content":"Write a short story about a lighthouse."}],
  "stream": false,
  "options": { "num_predict": 1200 }
}'
  1. Bound the thinking specifically, if your runtime supports it. llama.cpp has a first-class control — from common/arg.cpp, --reasoning-budget N is documented as "token budget for thinking: -1 for unrestricted, 0 for immediate end, N>0 for token budget (default: -1)", with a companion --reasoning-budget-message that injects a message before the end-of-thinking tag when the budget runs out. That combination is the cleanest way to force a model out of a loop and into an answer:
llama-server -m qwen3.5-8b.gguf --jinja \
  --reasoning-budget 1024 \
  --reasoning-budget-message "Budget reached. Answer now with what you have."

On the Ollama side, the equivalent is coarser: you get effort levels rather than a token budget. think accepts a boolean or a level — the API reference describes it as "should the model think before responding? Can be a boolean or a thinking level ("low", "medium", "high", or "max")". Dropping from "high" to "low" shortens traces; it does not hard-cap them. A token-budget equivalent is an open feature request (ollama/ollama #17561, "Proposal: bound thinking with a token budget (think: N, effort levels, PARAMETER think_budget)"), so at the time of writing there is no think: 1024.

  1. Only then look at sampling. A repeat-penalty of 1.0 on a model tuned for one can make loops worse, but this is a distant fourth — see LLM sampling parameters explained for what each knob really does.

Fault B: Empty Content — done_reason: "stop" With Nothing In It

If you get content: "" back, look at eval_count before you touch anything else. A tiny eval_count means the model produced almost nothing (a bug). A large one with a full thinking field means it produced plenty and failed to hand off into the answer (a different bug).

B1 — think: false returns nothing

This is a known live regression class, not your config. ollama/ollama issue #17823 (open, filed 2026-08-17) reports that on 0.32.14, an /api/chat request with "think": false against a Gemma 4 MLX model returns an empty assistant message — content: "", eval_count: 1, done_reason: "stop" — while the identical request against the identical model on 0.32.5 returns the answer directly. Crucially, in the same report "omitting think, or sending think: true / "low" / "medium" / "high", all return content normally".

That gives you an immediate workaround: stop asking for think: false and strip the thinking yourself.

# instead of "think": false — let it think, ignore the trace
curl -s http://localhost:11434/api/chat -d '{
  "model": "gemma4:26b",
  "messages": [{"role":"user","content":"What is the capital of France? One word."}],
  "stream": false,
  "think": "low"
}' | jq -r '.message.content'

.message.thinking and .message.content are separate fields in the /api/chat response, so ignoring the trace costs you one jq selector, not a regex. Related open work in the same family is worth checking against your model: issue #17459 (Gemma 4 emitting repeated <unused49> tokens when think=false, filed 2026-07-30) and PR #17128, an open parser fix that suppresses thinking output when think=false for DeepSeek — worth watching if you are on R1-family models, because it means the leak is acknowledged upstream and not something you should work around permanently.

B2 — Long thinking, then stops with empty content

Different failure, same appearance. ollama/ollama issue #17777 (open, filed 2026-08-15) documents a community Qwen3.5 fine-tune where, with "think": true, the model "reliably fails to ever produce message.content once the user/system instruction goes beyond a short, single-sentence request" — it generates a large message.thinking, then stops with done_reason: "stop" and empty content, well before hitting the model's 8192-token num_ctx, so context exhaustion is ruled out.

The signal to check on your own machine, in this order:

  • Does the failure track prompt complexity? Short prompt works, longer prompt returns nothing → thinking-to-content transition failure, which points at the model artefact.
  • Does it happen on the official tag but not a community fine-tune, or vice versa? #17777 is specifically a fine-tune of a supported architecture. Fine-tunes ship their own template and frequently ship it wrong.
  • Is eval_count well under num_ctx? If yes, this is not a context problem and raising num_ctx will waste your time.

If a fine-tune fails and the official tag of the same architecture does not, the fine-tune is the fault. Pull the official one, confirm, and file against the fine-tune rather than the runtime.


Own it instead of renting it

Run this on your own machine and stop paying every month

Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.

Fault C: Raw <think> or <|im_end|> In Your Visible Output

Leaked tags are a parsing decision, not a model malfunction — and in llama.cpp it is a single flag. From common/arg.cpp, --reasoning-format "controls whether thought tags are allowed and/or extracted from the response, and in which format they're returned", with these exact options:

--reasoning-formatDocumented behaviour
none"leaves thoughts unparsed in message.content"
deepseek"puts thoughts in message.reasoning_content"
deepseek-legacy"keeps <think> tags in message.content while also populating message.reasoning_content"
(default)auto

Source: llama.cpp common/arg.cpp on master, read 2026-08-18. The related -rea, --reasoning [on|off|auto] flag switches reasoning itself on or off, default auto (detected from the template).

So: seeing <think> in your content means you are effectively on none — either explicitly, or because auto did not recognise the model's reasoning syntax. Set --reasoning-format deepseek and read message.reasoning_content instead.

A leaked <|im_end|> is a different animal and is worth separating carefully. That token is a turn terminator, not a reasoning tag. If it reaches your screen, the runtime did not treat it as an end-of-generation token — which means the EOS token recorded in the GGUF metadata disagrees with what the template emits. That is a repackaging defect in the specific quant you downloaded, and the productive move is to try a different quant or a different uploader of the same model before writing a stop-sequence workaround. If you must patch it in place, register it explicitly:

# Ollama: add the terminator the blob forgot
printf 'FROM broken-model:latest\nPARAMETER stop "<|im_end|>"\n' > Modelfile
ollama create fixed-model -f Modelfile
ollama show --parameters fixed-model   # confirm it stuck

Treat that as a patch on someone else's mistake, not a configuration you should need.


Fault D: The Same Request, Two Endpoints, Two Behaviours

This is the one nobody writes up, and it explains a large share of "it works in the terminal but not in my app" reports: /api/chat and /v1/chat/completions control thinking with different field names and disjoint value sets.

The clearest published statement of it is the table in ollama/ollama PR #17499 ("docs: note that reasoning_effort and think accept different values", open, filed 2026-08-01), tested by its author on Ollama 0.30.10 with qwen3.5:4b on Linux:

EndpointFieldAccepts"none"true/false
/api/chatthinkhigh, medium, low, max, true, false400accepted
/v1/chat/completionsreasoning_efforthigh, medium, low, max, noneok400

Source: ollama/ollama PR #17499. As that PR puts it: "Both endpoints control the same feature with different field names and disjoint value sets, and neither page references the other. Moving between them returns a 400 with no pointer to the correct spelling."

The practical consequences:

  • There is no boolean false on the OpenAI-compatible endpoint. To disable thinking there you send "reasoning_effort": "none". Sending "think": false to /v1/chat/completions does not disable anything — at best it is ignored.
  • Every OpenAI-compatible client inherits this. ollama/ollama issue #15288 (since closed) reported exactly this shape for Gemma 4: content always empty with all text in the reasoning field on /v1/chat/completions, while native /api/chat with "think": false worked — breaking any framework that reads choices[0].message.content. If your app is empty but curl to /api/chat is fine, this is your bug.
  • Some controls are silently ignored on both. ollama/ollama issue #17785 (open, filed 2026-08-15) reports NVIDIA Nemotron 3's reasoning-effort controls — chat_template_kwargs.enable_thinking, chat_template_kwargs.low_effort and reasoning_budget — producing "no measurable, directionally-consistent effect on generated reasoning volume" on either surface, tried both nested under extra_body on the OpenAI route and promoted to top-level keys on /api/chat.

Run this two-line comparison on your own model before you debug anything else in your stack:

# native
curl -s http://localhost:11434/api/chat -d '{"model":"qwen3.5:8b","stream":false,"think":false,
  "messages":[{"role":"user","content":"Capital of France? One word."}]}' | jq '.message'

# OpenAI-compatible — note the different field AND the different value
curl -s http://localhost:11434/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model":"qwen3.5:8b","reasoning_effort":"none",
  "messages":[{"role":"user","content":"Capital of France? One word."}]}' | jq '.choices[0].message'

If one returns text and the other returns an empty string or a 400, you have localised the fault to the endpoint layer and can stop suspecting the model.


Handling Thinking Correctly In Your App

Read the structured field; do not regex the text. Both runtimes hand you the reasoning separately, and the structured route survives model swaps that a regex will not.

  • Ollama /api/chat: the trace arrives in message.thinking, the answer in message.content. Render or discard the first, act on the second.
  • llama.cpp: run with --reasoning-format deepseek and read message.reasoning_content.

Three rules that save real debugging time:

  1. Never regex a streaming buffer for <think>. Opening and closing tags routinely arrive split across chunks, so a buffer-wide regex will match late, match twice, or never match. If you genuinely have to parse text, run a small state machine over the token stream: outside-block → saw opener → inside-block → saw closer.
  2. Do not feed reasoning traces back into history by default. Most templates expect the trace to belong to the last assistant turn only. llama.cpp exposes --reasoning-preserve for the templates that do want it retained — the server docs describe it as "preserve reasoning trace in the full history, not just the last assistant message", and it defaults to whatever the template asks for. That default tells you plainly which way to lean: leave it alone unless a template needs it.
  3. If you need machine-readable output, constrain it rather than hoping. Reasoning models are the worst case for "just ask for JSON", because the trace itself often contains JSON-looking text. Use a grammar or schema — see our JSON mode and grammars guide.

Is It A Config Error Or A Version Bug?

This distinction decides whether you keep debugging or pin a version and move on. Our read of the evidence as of 2026-08-18:

SymptomVerdictEvidence
Never terminates on exact word countsConfiguration/prompt — reproducible, avoidableollama/ollama #17512 (closed)
think: false returns empty on Gemma 4 MLXVersion bug — regression 0.32.5 → 0.32.14ollama/ollama #17823 (open)
Long thinking then empty content on a fine-tuneModel artefact — architecture-specificollama/ollama #17777 (open)
Nemotron 3 reasoning controls ignoredRuntime gap — both endpointsollama/ollama #17785 (open)
reasoning_effort vs think value mismatchBy design, badly documentedollama/ollama #17499 (open)
Literal <think> in contentConfiguration — one flagllama.cpp --reasoning-format

Issue states above were read from the GitHub API on 2026-08-18. Open items may be fixed by the time you read this — check the issue number before assuming you are hitting it, and check your own version first: ollama --version.


What We Could Not Test

Being straight about the boundaries of this page:

  • We did not reproduce the Gemma 4 MLX regression ourselves. That path needs Apple Silicon running Ollama's MLX engine at specific versions. #17823 is a careful report with both versions, the same model artefact and full response bodies, which is why we cite it — but it is a citation, not our measurement. What to check on your machine: run the identical request against your current build and one release earlier, and compare eval_count.
  • We did not benchmark the Nemotron 3 controls. That is a 120B-class model; #17785's own reproduction used 7 runs of a ~200-token prompt. If you have the hardware, the check is whether reasoning volume moves directionally when you change the field, not whether the request succeeds — it succeeds either way, which is precisely why it is easy to miss.
  • Effort levels are not token budgets. "low" shortens traces in the reports we read; we have no measurement letting us tell you how much, and it will differ per model. Treat num_predict as your only hard guarantee on Ollama today.
  • This moves fast. Thinking-by-default only became the local norm during 2026. Anything here tied to an open issue number is a snapshot, not a law.

For faults that turn out not to be reasoning-related at all, our Ollama troubleshooting guide covers the wider surface, and the best local models roundup lists which current models think by default in the first place.


Sources

  • ollama/ollama issues and PRs — issues #17512, #17823, #17777, #17785, #17561, #17459, #15288 and PRs #17499, #17128; titles, states and bodies read on 2026-08-18
  • ollama/ollama releases — v0.32.14, published 2026-08-15
  • Ollama API reference — the think parameter, the thinking response field, done_reason, num_predict, stop, and the OpenAI-compatibility page listing reasoning_effort
  • llama.cpp common/arg.cpp — verbatim help text for --reasoning-format, --reasoning, --reasoning-budget, --reasoning-budget-message, --reasoning-preserve, --chat-template-file, --n-predict
  • ollama show --help — flag list verified locally

FAQ

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once

Liked this? 25 full AI courses are waiting.

From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.

Reading now
Join the discussion
TagsOllamallama.cppReasoning ModelsChat TemplatesQwen3.5Gemma 4Troubleshooting

LocalAimaster Research Team

Local AI Master writes hands-on courses and hardware guides for running AI on machines you own. Content is checked against current releases and corrected when readers tell us it is wrong.

Build Real AI on Your Machine

RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.

Want the structured version?

Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.

AI Learning Path

Comments (0)

No comments yet. Be the first to share your thoughts!

Why does my Ollama model never stop thinking?

Almost always because nothing bounds the reasoning and something in your prompt makes the model unable to satisfy itself. Ollama's num_predict defaults to unlimited, so a model that keeps second-guessing itself will generate until the context fills or you kill it. The single most reproducible trigger reported so far is an exact word-count instruction: ollama/ollama issue #17512 documents "Write a 200-word story" never terminating with thinking on, and terminating in 229-320 tokens with thinking off, on the same models and machine. Fix order: remove the exact-count instruction, set num_predict, then set a reasoning budget.

Why do I get an empty response when I set think: false?

That is usually an engine or model-specific bug, not your configuration. ollama/ollama issue #17823 (open, filed 2026-08-17) reports that on Ollama 0.32.14 a /api/chat request with "think": false against a Gemma 4 MLX model returns content: "" with eval_count: 1 and done_reason: "stop", while the identical request on 0.32.5 returns the answer — a regression. In the same report, omitting think entirely or sending think: true / "low" / "medium" / "high" all return content normally. If think:false gives you nothing, first try omitting think and stripping the thinking field yourself, then check whether your Ollama version is in the affected range.

Why does the same request behave differently on /api/chat and /v1/chat/completions?

Because they are two different control surfaces with different field names and disjoint value sets. Per the table in ollama/ollama PR #17499 (tested on Ollama 0.30.10 with qwen3.5:4b): /api/chat takes think and accepts high, medium, low, max, true, false — and returns 400 for "none". /v1/chat/completions takes reasoning_effort and accepts high, medium, low, max, none — and returns 400 for true/false. So there is no way to spell think:false on the OpenAI-compatible endpoint using a boolean; you send reasoning_effort: "none" instead.

How do I see the chat template and stop tokens a local model is actually using?

Run ollama show --template MODEL for the Jinja template, ollama show --parameters MODEL for stop sequences and other PARAMETER lines, and ollama show --modelfile MODEL for everything the blob carries. Those four flags (--template, --parameters, --modelfile, --system, plus --license) are what ollama show --help lists. In llama.cpp, the Jinja engine is on by default on current builds (--jinja / --no-jinja), so the model's own template metadata drives the chat format; override with --chat-template-file only when you have confirmed the packaged one is wrong.

How do I stop raw <think> and <|im_end|> tags leaking into my output?

Those are parser or template failures, not model failures. In llama.cpp, --reasoning-format controls it directly: "none" leaves thoughts unparsed in message.content, "deepseek" puts them in message.reasoning_content, and "deepseek-legacy" keeps the tags in content while also populating reasoning_content (default: auto). If you see literal tags in content, you are effectively on "none" — either because the flag is set that way or because the runtime did not recognise the model's reasoning syntax. A leaked <|im_end|> instead means the EOS/stop token in the GGUF metadata does not match the template, which is a repackaging problem: try a different quant of the same model before you write a regex.

Should I strip <think> blocks with a regex in my app?

Only as a fallback. Prefer the structured field: /api/chat returns reasoning separately in message.thinking, and llama.cpp with --reasoning-format deepseek returns it in message.reasoning_content. Both survive model changes; a regex does not. If you must parse text while streaming, do it as a state machine over the token stream rather than a regex over the buffer, because the opening and closing tags routinely arrive split across chunks.

Ready to Go Beyond Tutorials?

25 structured courses with hands-on chapters - build RAG chatbots, AI agents, and ML pipelines on your own hardware.

Bonus kit

Ollama Docker Templates

10 one-command Docker stacks for local models — pinned versions, so a regression never surprises you mid-project. Included with paid plans, or free after subscribing to both Local AI Master and Little AI Master on YouTube.

See Plans →

Was this helpful?

📅 Published: October 4, 2026🔄 Last Updated: October 4, 2026✓ Manually Reviewed
LM

Written by the Local AI Master Team

The team behind Local AI Master

We build Local AI Master around practical, testable local AI workflows: model selection, hardware planning, RAG systems, agents, and MLOps. The goal is to turn scattered tutorials into a structured learning path you can follow on your own hardware.

✓ Local AI Curriculum✓ Hands-On Projects✓ Open Source Contributor
📚
Free · no account required

Grab the AI Starter Kit — career roadmap, cheat sheet, setup guide

No spam. Unsubscribe with one click.

🎯
AI Learning Path

Go from reading about AI to building with AI

25 structured courses. Hands-on projects. Runs on your machine. Start free.

Or own it for life — Lifetime $149 $599, pay once
Free Tools & Calculators