Local Reasoning Model Never Finishes: Fix Think Loops
Want to go deeper than this article?
Free account unlocks the first chapter of all 25 courses — RAG, agents, MCP, voice AI, MLOps, real GitHub repos.
Go from reading about AI to building with AI 25 structured courses. Hands-on projects. Runs on your machine. Start free.
Four different things look identical from the outside — a model that "never finishes" — and each has a different fix. Runaway generation is a budget problem. Empty content with think: false is usually a version bug. Leaked <think> or <|im_end|> tags are a template/parser problem. And the same request behaving differently on /api/chat versus /v1/chat/completions is a field-name problem, because the two endpoints accept disjoint value sets for the same feature.
Work in that order and you will localise the fault in about two minutes instead of swapping quants for an afternoon. Everything below is dated and versioned, because several of these are live regressions rather than settled behaviour: Ollama's current release at the time of writing is v0.32.14, published 2026-08-15 (GitHub releases), and the llama.cpp flags quoted are from common/arg.cpp on master, read 2026-08-18.
Which Fault Do You Have?
Start here — match the symptom, then jump to the section. Do not change sampler settings before you have done this; temperature is almost never the cause.
| What you see | Most likely fault | Section |
|---|---|---|
| Generates for minutes, no answer, never returns | Unbounded reasoning + a prompt it cannot satisfy | Fault A |
content: "", done_reason: "stop", low eval_count | Engine/model bug on think: false | Fault B |
Long thinking, then stops with empty content | Model/quant fails the thinking→content transition | Fault B |
Literal <think> or `< | im_end | >` in the visible reply |
| Works on one endpoint, 400s or ignores you on the other | Field name / value set divergence | Fault D |
The reason these get confused is that a chat UI shows you only the last box. Stop using the UI to diagnose. Every step below uses curl against the raw API so you can see done_reason, eval_count and the thinking field, which is where the actual evidence lives.
Reading articles is good. Building is better.
Free account = the first chapter of all 25 courses, with a per-chapter AI tutor. No card.
Step 1: Read the Template the Model Is Actually Using
Before you change anything, print the template and the stop parameters — a surprising share of "broken model" reports are a repackaged GGUF with a template that does not match its own EOS token.
Ollama ships this in the CLI. These flags are what ollama show --help prints:
ollama show --template qwen3.5:8b # the Jinja chat template
ollama show --parameters qwen3.5:8b # PARAMETER lines, including stop sequences
ollama show --modelfile qwen3.5:8b # everything the blob carries
What you are looking for, in order:
- Does the template emit a reasoning block at all? If the template has no
<think>handling but the model was trained to produce one, the tags have nowhere to go and end up in your content. - Do the
stopparameters match the tokens the template closes with? A ChatML-style template that closes turns with<|im_end|>needs that token registered as EOS. Ifollama show --parametersis empty and the template uses<|im_end|>, you have found your leak. - Is there a
num_predictornum_ctxset? Usually not — and that matters for Fault A.
The llama.cpp equivalent is to let the model's own metadata drive the template rather than guessing:
llama-server -m model.gguf --jinja # use the template from model metadata
llama-server -m model.gguf --jinja --chat-template-file ./fixed.jinja # override, only if you have proven the packaged one wrong
Per common/arg.cpp, --chat-template and --chat-template-file default to "template taken from the model's metadata". Override last, not first: an override that disagrees with the tokenizer produces exactly the same symptoms you are trying to fix.
If none of this is familiar territory, our complete Ollama guide covers Modelfiles and parameters from the ground up.
Fault A: It Reasons Forever and Never Answers
The cause is almost never the model — it is that nothing bounds the reasoning, and your prompt contains something the model keeps failing to verify. Ollama's num_predict defaults to unlimited, and llama.cpp's --n-predict documents its default as "-1, -1 = infinity" (common/arg.cpp). Give a self-doubting model infinite rope and it will use it.
The reproducible trigger: exact word counts
This is the one that has been isolated cleanly. ollama/ollama issue #17512 (filed 2026-08-02) reports that with thinking on, a prompt specifying an exact word count sends the model into "an unbounded self-verification loop: it drafts, attempts to count words, decides the count is wrong, and rewrites — indefinitely, never emitting a stop token." The reporter's controlled table, same models and machine, varying only thinking mode:
| Prompt | Thinking | gemma4:12b | qwen3.5:9b |
|---|---|---|---|
| "Write a 200-word story about a lighthouse." | on | never terminates | never terminates |
| "Write a 200-word story about a lighthouse." | off | 229 tokens, 46s | 320 tokens, 45s |
| "Write a short story about a lighthouse." | on | 1,423 tokens, terminates | 2,916 tokens, terminates |
Source: ollama/ollama issue #17512, reproduced on Ollama 0.32.5 on an Intel Core Ultra 7 255H across CPU, Vulkan and SYCL backends. That last row is the tell: thinking on is fine on its own. Three conditions had to line up — thinking on, an exact count, and num_predict left at its default of -1.
The secondary damage is worth knowing about: with OLLAMA_NUM_PARALLEL=1, one stuck request holds the only slot forever and every later request queues behind it, so the whole server looks hung to clients that are perfectly healthy.
Fixes, in the order to try them
- Rewrite the instruction. "About 200 words" or "roughly 4 short paragraphs" instead of "exactly 200 words". Reasoning models cannot count tokens, so an exact count is an unsatisfiable verification target. This costs nothing and fixes the documented case outright.
- Bound the generation. Nothing should ever run unbounded in production:
curl -s http://localhost:11434/api/chat -d '{
"model": "qwen3.5:8b",
"messages": [{"role":"user","content":"Write a short story about a lighthouse."}],
"stream": false,
"options": { "num_predict": 1200 }
}'
- Bound the thinking specifically, if your runtime supports it. llama.cpp has a first-class control — from
common/arg.cpp,--reasoning-budget Nis documented as "token budget for thinking: -1 for unrestricted, 0 for immediate end, N>0 for token budget (default: -1)", with a companion--reasoning-budget-messagethat injects a message before the end-of-thinking tag when the budget runs out. That combination is the cleanest way to force a model out of a loop and into an answer:
llama-server -m qwen3.5-8b.gguf --jinja \
--reasoning-budget 1024 \
--reasoning-budget-message "Budget reached. Answer now with what you have."
On the Ollama side, the equivalent is coarser: you get effort levels rather than a token budget. think accepts a boolean or a level — the API reference describes it as "should the model think before responding? Can be a boolean or a thinking level ("low", "medium", "high", or "max")". Dropping from "high" to "low" shortens traces; it does not hard-cap them. A token-budget equivalent is an open feature request (ollama/ollama #17561, "Proposal: bound thinking with a token budget (think: N, effort levels, PARAMETER think_budget)"), so at the time of writing there is no think: 1024.
- Only then look at sampling. A repeat-penalty of 1.0 on a model tuned for one can make loops worse, but this is a distant fourth — see LLM sampling parameters explained for what each knob really does.
Fault B: Empty Content — done_reason: "stop" With Nothing In It
If you get content: "" back, look at eval_count before you touch anything else. A tiny eval_count means the model produced almost nothing (a bug). A large one with a full thinking field means it produced plenty and failed to hand off into the answer (a different bug).
B1 — think: false returns nothing
This is a known live regression class, not your config. ollama/ollama issue #17823 (open, filed 2026-08-17) reports that on 0.32.14, an /api/chat request with "think": false against a Gemma 4 MLX model returns an empty assistant message — content: "", eval_count: 1, done_reason: "stop" — while the identical request against the identical model on 0.32.5 returns the answer directly. Crucially, in the same report "omitting think, or sending think: true / "low" / "medium" / "high", all return content normally".
That gives you an immediate workaround: stop asking for think: false and strip the thinking yourself.
# instead of "think": false — let it think, ignore the trace
curl -s http://localhost:11434/api/chat -d '{
"model": "gemma4:26b",
"messages": [{"role":"user","content":"What is the capital of France? One word."}],
"stream": false,
"think": "low"
}' | jq -r '.message.content'
.message.thinking and .message.content are separate fields in the /api/chat response, so ignoring the trace costs you one jq selector, not a regex. Related open work in the same family is worth checking against your model: issue #17459 (Gemma 4 emitting repeated <unused49> tokens when think=false, filed 2026-07-30) and PR #17128, an open parser fix that suppresses thinking output when think=false for DeepSeek — worth watching if you are on R1-family models, because it means the leak is acknowledged upstream and not something you should work around permanently.
B2 — Long thinking, then stops with empty content
Different failure, same appearance. ollama/ollama issue #17777 (open, filed 2026-08-15) documents a community Qwen3.5 fine-tune where, with "think": true, the model "reliably fails to ever produce message.content once the user/system instruction goes beyond a short, single-sentence request" — it generates a large message.thinking, then stops with done_reason: "stop" and empty content, well before hitting the model's 8192-token num_ctx, so context exhaustion is ruled out.
The signal to check on your own machine, in this order:
- Does the failure track prompt complexity? Short prompt works, longer prompt returns nothing → thinking-to-content transition failure, which points at the model artefact.
- Does it happen on the official tag but not a community fine-tune, or vice versa? #17777 is specifically a fine-tune of a supported architecture. Fine-tunes ship their own template and frequently ship it wrong.
- Is
eval_countwell undernum_ctx? If yes, this is not a context problem and raisingnum_ctxwill waste your time.
If a fine-tune fails and the official tag of the same architecture does not, the fine-tune is the fault. Pull the official one, confirm, and file against the fine-tune rather than the runtime.
Run this on your own machine and stop paying every month
Pay once and keep it. No renewal, no per-token bill, and nothing you feed it ever leaves your hardware.
Fault C: Raw <think> or <|im_end|> In Your Visible Output
Leaked tags are a parsing decision, not a model malfunction — and in llama.cpp it is a single flag. From common/arg.cpp, --reasoning-format "controls whether thought tags are allowed and/or extracted from the response, and in which format they're returned", with these exact options:
--reasoning-format | Documented behaviour |
|---|---|
none | "leaves thoughts unparsed in message.content" |
deepseek | "puts thoughts in message.reasoning_content" |
deepseek-legacy | "keeps <think> tags in message.content while also populating message.reasoning_content" |
| (default) | auto |
Source: llama.cpp common/arg.cpp on master, read 2026-08-18. The related -rea, --reasoning [on|off|auto] flag switches reasoning itself on or off, default auto (detected from the template).
So: seeing <think> in your content means you are effectively on none — either explicitly, or because auto did not recognise the model's reasoning syntax. Set --reasoning-format deepseek and read message.reasoning_content instead.
A leaked <|im_end|> is a different animal and is worth separating carefully. That token is a turn terminator, not a reasoning tag. If it reaches your screen, the runtime did not treat it as an end-of-generation token — which means the EOS token recorded in the GGUF metadata disagrees with what the template emits. That is a repackaging defect in the specific quant you downloaded, and the productive move is to try a different quant or a different uploader of the same model before writing a stop-sequence workaround. If you must patch it in place, register it explicitly:
# Ollama: add the terminator the blob forgot
printf 'FROM broken-model:latest\nPARAMETER stop "<|im_end|>"\n' > Modelfile
ollama create fixed-model -f Modelfile
ollama show --parameters fixed-model # confirm it stuck
Treat that as a patch on someone else's mistake, not a configuration you should need.
Fault D: The Same Request, Two Endpoints, Two Behaviours
This is the one nobody writes up, and it explains a large share of "it works in the terminal but not in my app" reports: /api/chat and /v1/chat/completions control thinking with different field names and disjoint value sets.
The clearest published statement of it is the table in ollama/ollama PR #17499 ("docs: note that reasoning_effort and think accept different values", open, filed 2026-08-01), tested by its author on Ollama 0.30.10 with qwen3.5:4b on Linux:
| Endpoint | Field | Accepts | "none" | true/false |
|---|---|---|---|---|
/api/chat | think | high, medium, low, max, true, false | 400 | accepted |
/v1/chat/completions | reasoning_effort | high, medium, low, max, none | ok | 400 |
Source: ollama/ollama PR #17499. As that PR puts it: "Both endpoints control the same feature with different field names and disjoint value sets, and neither page references the other. Moving between them returns a 400 with no pointer to the correct spelling."
The practical consequences:
- There is no boolean
falseon the OpenAI-compatible endpoint. To disable thinking there you send"reasoning_effort": "none". Sending"think": falseto/v1/chat/completionsdoes not disable anything — at best it is ignored. - Every OpenAI-compatible client inherits this. ollama/ollama issue #15288 (since closed) reported exactly this shape for Gemma 4:
contentalways empty with all text in thereasoningfield on/v1/chat/completions, while native/api/chatwith"think": falseworked — breaking any framework that readschoices[0].message.content. If your app is empty but curl to/api/chatis fine, this is your bug. - Some controls are silently ignored on both. ollama/ollama issue #17785 (open, filed 2026-08-15) reports NVIDIA Nemotron 3's reasoning-effort controls —
chat_template_kwargs.enable_thinking,chat_template_kwargs.low_effortandreasoning_budget— producing "no measurable, directionally-consistent effect on generated reasoning volume" on either surface, tried both nested underextra_bodyon the OpenAI route and promoted to top-level keys on/api/chat.
Run this two-line comparison on your own model before you debug anything else in your stack:
# native
curl -s http://localhost:11434/api/chat -d '{"model":"qwen3.5:8b","stream":false,"think":false,
"messages":[{"role":"user","content":"Capital of France? One word."}]}' | jq '.message'
# OpenAI-compatible — note the different field AND the different value
curl -s http://localhost:11434/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model":"qwen3.5:8b","reasoning_effort":"none",
"messages":[{"role":"user","content":"Capital of France? One word."}]}' | jq '.choices[0].message'
If one returns text and the other returns an empty string or a 400, you have localised the fault to the endpoint layer and can stop suspecting the model.
Handling Thinking Correctly In Your App
Read the structured field; do not regex the text. Both runtimes hand you the reasoning separately, and the structured route survives model swaps that a regex will not.
- Ollama
/api/chat: the trace arrives inmessage.thinking, the answer inmessage.content. Render or discard the first, act on the second. - llama.cpp: run with
--reasoning-format deepseekand readmessage.reasoning_content.
Three rules that save real debugging time:
- Never regex a streaming buffer for
<think>. Opening and closing tags routinely arrive split across chunks, so a buffer-wide regex will match late, match twice, or never match. If you genuinely have to parse text, run a small state machine over the token stream: outside-block → saw opener → inside-block → saw closer. - Do not feed reasoning traces back into history by default. Most templates expect the trace to belong to the last assistant turn only. llama.cpp exposes
--reasoning-preservefor the templates that do want it retained — the server docs describe it as "preserve reasoning trace in the full history, not just the last assistant message", and it defaults to whatever the template asks for. That default tells you plainly which way to lean: leave it alone unless a template needs it. - If you need machine-readable output, constrain it rather than hoping. Reasoning models are the worst case for "just ask for JSON", because the trace itself often contains JSON-looking text. Use a grammar or schema — see our JSON mode and grammars guide.
Is It A Config Error Or A Version Bug?
This distinction decides whether you keep debugging or pin a version and move on. Our read of the evidence as of 2026-08-18:
| Symptom | Verdict | Evidence |
|---|---|---|
| Never terminates on exact word counts | Configuration/prompt — reproducible, avoidable | ollama/ollama #17512 (closed) |
think: false returns empty on Gemma 4 MLX | Version bug — regression 0.32.5 → 0.32.14 | ollama/ollama #17823 (open) |
| Long thinking then empty content on a fine-tune | Model artefact — architecture-specific | ollama/ollama #17777 (open) |
| Nemotron 3 reasoning controls ignored | Runtime gap — both endpoints | ollama/ollama #17785 (open) |
reasoning_effort vs think value mismatch | By design, badly documented | ollama/ollama #17499 (open) |
Literal <think> in content | Configuration — one flag | llama.cpp --reasoning-format |
Issue states above were read from the GitHub API on 2026-08-18. Open items may be fixed by the time you read this — check the issue number before assuming you are hitting it, and check your own version first: ollama --version.
What We Could Not Test
Being straight about the boundaries of this page:
- We did not reproduce the Gemma 4 MLX regression ourselves. That path needs Apple Silicon running Ollama's MLX engine at specific versions. #17823 is a careful report with both versions, the same model artefact and full response bodies, which is why we cite it — but it is a citation, not our measurement. What to check on your machine: run the identical request against your current build and one release earlier, and compare
eval_count. - We did not benchmark the Nemotron 3 controls. That is a 120B-class model; #17785's own reproduction used 7 runs of a ~200-token prompt. If you have the hardware, the check is whether reasoning volume moves directionally when you change the field, not whether the request succeeds — it succeeds either way, which is precisely why it is easy to miss.
- Effort levels are not token budgets.
"low"shortens traces in the reports we read; we have no measurement letting us tell you how much, and it will differ per model. Treatnum_predictas your only hard guarantee on Ollama today. - This moves fast. Thinking-by-default only became the local norm during 2026. Anything here tied to an open issue number is a snapshot, not a law.
For faults that turn out not to be reasoning-related at all, our Ollama troubleshooting guide covers the wider surface, and the best local models roundup lists which current models think by default in the first place.
Sources
- ollama/ollama issues and PRs — issues #17512, #17823, #17777, #17785, #17561, #17459, #15288 and PRs #17499, #17128; titles, states and bodies read on 2026-08-18
- ollama/ollama releases — v0.32.14, published 2026-08-15
- Ollama API reference — the
thinkparameter, thethinkingresponse field,done_reason,num_predict,stop, and the OpenAI-compatibility page listingreasoning_effort - llama.cpp
common/arg.cpp— verbatim help text for--reasoning-format,--reasoning,--reasoning-budget,--reasoning-budget-message,--reasoning-preserve,--chat-template-file,--n-predict ollama show --help— flag list verified locally
FAQ
Go from reading about AI to building with AI
25 structured courses. Hands-on projects. Runs on your machine. Start free.
Liked this? 25 full AI courses are waiting.
From fundamentals to RAG, agents, MCP servers, voice AI, and production deployment with real GitHub repos. First chapter free, every course.
Build Real AI on Your Machine
RAG, agents, NLP, vision, and MLOps - chapters across 25 courses that take you from reading about AI to building AI.
Want the structured version?
Hands-on courses on local AI, from $8.99 a month. The first chapter of each is free.
Keep going
Comments (0)
No comments yet. Be the first to share your thoughts!