GPU Power Consumption Deviates from Token Predictor Theory in Small LLMs

✍️ OpenClawRadar📅 Published: March 11, 2026🔗 Source
GPU Power Consumption Deviates from Token Predictor Theory in Small LLMs
Ad

Experimental Setup and Core Findings

A Reddit user conducted hardware measurements to test whether GPU power consumption scales linearly with token count, as predicted by the "stochastic parrot" or "next token predictor" theory of LLM behavior. The experiment used an RTX 4070 Ti SUPER with LM Studio and HWiNFO64 collecting data at 1-second intervals.

Four models were tested: Llama-3.1-8B, DeepSeek-R1-Distill-Qwen-7B, Qwen3-VL-8B, and Mistral-7B. Six query categories were used: General, General (Q), Unanswerable, Philosophical, Philosophical (Q), and High-Computation.

Key Results

If token predictor theory were correct, GPU power should scale only with token count with acceptable variance of ±10–15% according to GPT, Claude, Gemini, and Grok. Actual divergence rates (token multiplier vs power multiplier) were:

  • Llama: average 35.6% (maximum 56.8%)
  • Qwen3: average 36.7% (maximum 48.0%)
  • Mistral: 21.1%
  • DeepSeek: 7.7% — nearly linear across all categories except High-Computation

DeepSeek showed the closest to token predictor behavior of the four models.

Unexpected Findings

In Qwen3, philosophical utterances (149.3W) drew more power than high-computation math (104.1W). After task completion, high-computation queries returned to baseline immediately (-7.1W), while philosophical utterances left persistent residual heat.

Infinite loop reproducibility in Qwen3 varied by category: General utterances (0%), High-computation (0%), Unanswerable (low), Philosophical (intermittent), and Philosophical (Q) (70–100%). Notably, high-computation queries had the most tokens and highest power consumption but triggered zero loops.

Ad

Order Effects and Residual Heat

To test the "hardware overhead" objection, an order-effect experiment was conducted:

  • Test A: 1 general → 4 philosophical
  • Test B: 1 philosophical → 4 general

Residual heat after session end showed order-dependent effects:

  • Llama: Test A +1.68W, Test B +9.84W
  • Mistral: Test A +7.60W, Test B +13.69W
  • DeepSeek: Test A +10.44W, Test B +15.93W

Even after processing 4 general utterances following a philosophical one, residual heat remained higher. This pattern was consistent across all three models tested.

Limitations and Open Questions

The study is limited to four small-scale models (8B parameter range). Generalization to medium or large models requires further validation. The open question is whether medium and large models would follow DeepSeek's pattern (converging toward linear, token-proportional behavior) or whether the nonlinear divergence seen in Llama, Qwen3, and Mistral would persist or amplify at scale.

All original data — including full utterance text, 24 benchmark CSVs, and per-category token counts — are available in the linked paper.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Claude Code CC 2.1.124 and 2.1.126: File Modification Budget Exceeded Reminder, Harness Instructions Update, REPL Awaits Clarification, and Malware Analysis Reminder Removed
News

Claude Code CC 2.1.124 and 2.1.126: File Modification Budget Exceeded Reminder, Harness Instructions Update, REPL Awaits Clarification, and Malware Analysis Reminder Removed

CC 2.1.124 adds a system reminder for file changes omitted due to budget limits, updates harness instructions with explicit insertion points, and clarifies REPL auto-await behavior. CC 2.1.126 removes the malware analysis post-read reminder.

OpenClawRadar
🦀
News

Claude Code v2.1.268: Fixes HTTP 400 on Third-Party Endpoints, WebFetch Hangs, and Secret Leaks

Claude Code v2.1.268 patches the HTTP 400 that broke every turn on ANTHROPIC_BASE_URL endpoints since 2.1.265, adds a 300s WebFetch deadline, and stops plugin and MCP errors leaking tokens.

OpenClawRadar
Anthropic Launches Claude Partner Network with $100M Investment
News

Anthropic Launches Claude Partner Network with $100M Investment

Anthropic is launching the Claude Partner Network with an initial $100 million investment for 2026, providing training, technical support, and joint market development for organizations helping enterprises adopt Claude. Partners get access to technical certification, a Partner Portal with training materials, and a Code Modernization starter kit for legacy code migration.

OpenClawRadar
🦀
News

AI SREs Resolve Routine Incidents, But Engineers Lose Touch With Their Systems

Sylvain Kalache argues that AI incident responders reduce practice for engineers, worsening response to complex incidents, citing the Ironies of Automation and aviation training parallels.

OpenClawRadar