Gemma 4 31B outperforms larger models on FoodTruck Bench

Benchmark results and analysis
Gemma 4 31B achieved 3rd place on the FoodTruck Bench benchmark, outperforming several larger and more established models. According to the Reddit discussion, the model beat GLM 5, Qwen 3.5 397B, and all Claude Sonnet variants.
The FoodTruck Bench is a benchmark that tests language models on complex, multi-step planning tasks. The original poster speculates that Gemma 4's performance suggests it handles long-horizon tasks better than previous models that failed to complete the benchmark. Specifically, the model appears to effectively listen to its own advice when planning for subsequent steps in the task sequence.
This result is notable because Gemma 4 31B is significantly smaller than some of the models it outperformed. Qwen 3.5 397B, for example, has approximately 12.8 times more parameters than Gemma 4 31B. The performance suggests that model architecture and training approaches may be as important as parameter count for certain types of reasoning tasks.
FoodTruck Bench tests models on practical planning scenarios that require maintaining context over extended sequences of actions. The benchmark's design makes it particularly relevant for developers working with AI agents that need to execute multi-step tasks in real-world applications.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Tencent Hosts Free OpenClaw Installation Event in Shenzhen Amid High Demand
Tencent organized 20 employees outside its Shenzhen office building to install OpenClaw for free on March 6, responding to reports of people paying over $70 for house-call installation services. The event used Tencent Cloud's Lighthouse platform, with most attendees being white-collar professionals facing workplace competition and AI adoption pressure.

Mercor Breach: 4TB of Voice Samples + IDs Stolen – What Attackers Can Do Now
4TB of voice recordings paired with government IDs stolen from 40,000 Mercor contractors. Attackers can clone voices from 15 seconds of clean audio and bypass bank voiceprint verification, deepfake calls, and insurance fraud.

SubQ: First Fully Subquadratic LLM with 12M-Token Context and 95% RULER Accuracy
Subquadratic launches SubQ 1M-Preview, a subquadratic LLM with linear compute scaling, 12M-token context, 52× faster sparse attention vs FlashAttention, and 95% on RULER 128K. Available via API, CLI code agent (SubQ Code), and search tool (SubQ Search).
Transformer Language Model Runs Locally on Stock Game Boy Color
Andrej Karpathy's TinyStories-260K model runs on a stock Game Boy Color via a custom ROM, using INT8 fixed-point math and bank-switched cartridge memory for weights and KV cache.