Reasoning Guard: Proxy-Level Loop Detection for Local LLM Inference

✍️ OpenClawRadar📅 Published: April 30, 2026🔗 Source
Reasoning Guard: Proxy-Level Loop Detection for Local LLM Inference
Ad

A developer running Qwen3.6 MoE behind a vLLM proxy hit a common reliability issue: runaway reasoning loops where the model repeats itself inside a reasoning block, burning tokens and stalling agents. At 180+ tokens/sec, even a 20–30 second loop wastes GPU time and blocks client requests. They built a lightweight guard that lives in the proxy layer and enforces deterministic checks on the streaming output before it reaches the client.

Architecture

Client → Proxy → vLLM → Model

The proxy intercepts the streaming response as it leaves vLLM. It does not modify model weights, call a second LLM, or use embeddings or semantic analysis. All checks are cheap and deterministic.

What It Checks

  • Reasoning token caps (configurable per effort level)
  • Repeated paragraph detection
  • Sliding-window n-gram repetition
  • Repeated sentence fingerprinting
  • Fuzzy opening-pattern detection (catches loops like “Actually, I think I’ve found it…”)
  • Cut-and-continue recovery path
Ad

Recovery Flow

When the guard triggers, it:

  • Stops the upstream stream
  • Captures the reasoning produced so far
  • Reissues the request with that reasoning baked in as prior assistant context
  • Disables thinking for the continuation
  • Merges phase 1 and phase 2 usage stats

Because vLLM prefix caching is already active, the continuation is effectively seamless. Phase 2 usually resumes with ~50–100ms TTFT, so the client sees reasoning flow directly into the final answer instead of hanging.

Observability

The proxy logs each trigger with:

  • Whether the guard fired
  • Trigger reason
  • Token cap used
  • Reasoning token count
  • Merged total usage
  • Stream-end metadata

Result

Before: occasional 2000+ token reasoning blocks that went nowhere. After: the model still reasons when useful, but runaway thinking gets cut and redirected into an answer. The author describes it as a “proxy-level seatbelt for local LLM inference.”

No model surgery, no extra LLM calls — just stream interception, token counting, loop detection, and a clean recovery path. The guard has been validated end-to-end through the live proxy against real trace logs.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Interact MCP: Faster Web Browsing for Claude Code with Persistent Chromium
Tools

Interact MCP: Faster Web Browsing for Claude Code with Persistent Chromium

Interact MCP is a Model Context Protocol tool that keeps a persistent Chromium browser in-process, reducing browser action times from 2-5 seconds to 5-50ms after the initial call. It features a ref system for element interaction without CSS selectors and includes 46 tools for web automation.

OpenClawRadar
🦀
Tools

Zillow-Full: An OpenClaw Skill That Turned Manual Property Research Into an Automated Deal Pipeline

A developer built 'zillow-full' on OpenClaw to pull Zestimates, tax history, price history, and comps per property. With a nightly cron scoring listings against deal criteria, wholesale deals went from 2 to 11 per month.

OpenClawRadar
Local Code Index for Coding Agents: Resolves Imports Without Language Server
Tools

Local Code Index for Coding Agents: Resolves Imports Without Language Server

Basemind is an MIT-licensed Rust tool that indexes local code and resolves imports without a language server, supporting real resolution for JS/TS, Python, and Java.

OpenClawRadar
mycrab.space introduces SKILL.md and Prompt Autocomposer for standardized app deployment
Tools

mycrab.space introduces SKILL.md and Prompt Autocomposer for standardized app deployment

mycrab.space has released SKILL.md, a Markdown blueprint for defining app dependencies and configuration, and a Prompt Autocomposer that generates ready-to-use deployment commands from these files. The system enables zero-config deployment of applications like VS Code in browser, personal music clouds, and AI agent interfaces.

OpenClawRadar