Local LLM Struggles with Unreal Engine Solitaire: Qwen 3.6-27B Burns 687k Tokens on One Card

A Reddit user's experiment with local LLMs for game development reveals severe practical limitations. Using Qwen 3.6-27B with access to unreal-mcpython, SearXNG, and GitHub, the task was to create a Solitaire game in Unreal Engine. After a few hours (much time waiting for user responses to prompts), the result was a single card with correct textures but no game logic, consuming ↑687k and ↓210k tokens.
Manual Interventions Required
- Downloading PNGs with card faces manually
- Creating a mesh with 3 materials (stock cube has only 1 side material)
- Constant prompting like "stop imagining things, use a bloody search"
- Repeated corrections: "the card has no texture" or "the card has ace of spades on both sides"
The two-sided card problem consumed the majority of time and tokens. The stock cube can only have one material on all sides; a custom mesh with 3 materials is required. Gemini Flash 3.5 generated the correct OBJ file in one attempt, but Qwen went in circles for hours despite finding concrete code examples. The model insisted on creating planes, compounding two planes with a cube, disabling substrate, or other non-working approaches. The user ultimately had to provide the mesh manually.
Gemma 4-31B was tested but couldn't make a meaningful MCP call and was disqualified early.
Practical takeaway: for Unreal Engine tasks involving custom geometry, local LLMs like Qwen 3.6-27B still require substantial hand-holding. Token budgets balloon quickly, and basic mesh operations remain a stumbling block.
📖 Read the full source: r/LocalLLaMA
👀 See Also

AI's PR Problem: Flat Wages, Soaring Capital, and Public Backlash
College wage premium flat for 25 years, S&P 500 up 380%. Workers see AI as theft enabler, leading to laws against data centers.

Qwen3.5-27B-FP8 performance benchmarks with OpenClaw agents
Testing shows Qwen3.5-27B-FP8 can run six OpenClaw agents simultaneously with throughput scaling to 120 tokens/second. The SGLang framework with prefix caching reduces 100K context prefill from 10 seconds to 200ms.

TabFM: Google's Zero-Shot Foundation Model for Tabular Data Classification and Regression
TabFM applies in-context learning to tabular data, eliminating hyperparameter tuning and feature engineering for classification and regression. Available on Hugging Face and GitHub.

Why Anthropic's Activation Steering Struggles with Generating Valid JSON
Activation steering, a technique used for AI safety, fails to generate valid JSON, achieving only 24.4% validity compared to 86.8% from the untrained base model.