SpruceChat Runs 0.5B LLM On-Device on Miyoo Handhelds via llama.cpp

What This Is
SpruceChat is a project that runs the Qwen2.5-0.5B language model entirely on-device on several handheld gaming consoles using llama.cpp. It requires no cloud connection or WiFi after the initial setup.
Key Details
The model lives in RAM after the first boot, and tokens stream in one by one during generation. It runs on the Miyoo A30, Miyoo Flip, Trimui Brick, and Trimui Smart Pro.
Performance on the Miyoo A30 (which has a Cortex-A7 quad-core processor):
- Model load: ~60 seconds on first boot
- Generation speed: ~1-2 tokens per second
- Prompt evaluation: ~3 tokens per second
The developer notes it's not fast, but it streams so you can watch it think. They mention 64-bit devices are quicker.
The AI is described as having "the personality of a spruce tree: patient, unhurried, quietly amazed by everything."
If the device is on WiFi, you can also hit the llama-server from a browser on a phone or laptop to chat with a real keyboard.
The repository is at https://github.com/RED-BASE/SpruceChat. The project was built with help from Claude, and there's already a collaborator working on expanding device support. The first release is up with both armhf and aarch64 binaries, and the model is included.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Google's HEIR Compiler: Practical Private AI with Homomorphic Encryption
Google open-sources HEIR, a compiler that lets developers run AI inference on encrypted data without decryption. Demos include fraud detection, recommendations, and more.

Audacity-MCP: Claude AI Integration for Local Audio Editing with 131 Tools
Audacity-MCP connects Claude to Audacity via pipe interface, enabling voice-controlled audio editing with 131 tools, 9 automated pipelines, and local Whisper transcription without cloud dependencies.
Tokenless API Gateway Routes AI Traffic Between Models to Cut Spend in Half
Tokenless is an API gateway that dynamically routes AI agent requests turn-by-turn between models. It matches Claude Fable 5 performance at half the cost by fanning out to multiple models and selecting the best one mid-generation.

Microsoft BitNet: 1-bit LLM inference framework for CPU and GPU
Microsoft released BitNet, an inference framework for 1-bit LLMs that achieves 1.37x to 6.17x speedups on CPUs and reduces energy consumption by 55.4% to 82.2%. It can run a 100B parameter model on a single CPU at 5-7 tokens per second.