Off Grid Mobile App Adds On-Device AI Tool Use with 3x Speed Improvement

Off Grid, an on-device AI mobile app, has been updated to add tool use capabilities and significant performance improvements. The app now allows AI models to call tools offline without requiring API keys, servers, or cloud functions.
Key Features and Performance
The update introduces automatic tool loops for web search, calculator, date/time functions, and device information access. According to the developer, this bridges the gap between "local toy" and "useful assistant" by enabling 3B parameter models to reason, call tools, and synthesize results directly on your phone.
Performance improvements come from configurable KV cache options. Users can now choose between three KV cache types:
f16q8_0q4_0
With q4_0 cache, models that previously generated 10 tokens/second now reach 30 tokens/second. The app includes a performance nudge feature that suggests faster settings after the first generation.
Model Support and Platform Availability
Off Grid supports GGUF format models, including:
- Qwen 3
- Llama 3.2
- Gemma 3
- Phi-4
- Other GGUF-compatible models
The app is now available on both major app stores without sideloading requirements. It can be installed directly from the App Store and Google Play.
Core Functionality and Philosophy
What hasn't changed in this update:
- MIT licensed and fully open source
- Zero data leaves the device (no analytics, telemetry, or anonymous usage data)
- Offline capabilities including text generation (15-30 tokens/second), image generation (5-10 seconds on NPU), vision AI, voice transcription, and document analysis
The developer states the project is motivated by the belief that "the phone in your pocket should be the most private computer you own — not the most surveilled."
📖 Read the full source: HN AI Agents
👀 See Also
Tokenless API Gateway Routes AI Traffic Between Models to Cut Spend in Half
Tokenless is an API gateway that dynamically routes AI agent requests turn-by-turn between models. It matches Claude Fable 5 performance at half the cost by fanning out to multiple models and selecting the best one mid-generation.

Practical Findings from 11 Multi-Agent Software Builds Without Programmatic Scaffolding
Analysis of 11 autonomous multi-agent builds shows scope enforcement works mechanically (20/20 success) not via prompts (0/20), orchestration costs are dominated by memory re-ingestion (~95% of input spend), and worker model capability creates 9.8x throughput gaps.

Autonomous coding workflow ships 163K lines overnight using Claude Code
A developer built an autonomous workflow that completed 72 tasks overnight, generating 163,643 lines of code and 6,400+ passing tests with an 85% first-attempt success rate.

Ninetails Memory Engine V4.5: Int8 Quantization + LRU Cache Cuts Local MCP Memory to 60MB
The Ninetails Memory Engine V4.5 uses Int8 scalar quantization and LRU cache eviction to reduce vector storage from 6KB to 1.5KB per embedding, keeping the entire engine at 40-60MB RAM. It combines 70% vector similarity with 30% BM25 search in a fully local SQLite implementation.