TinySearch v0.2.0: Lightweight Web Search for Local LLMs Now Backed by SearXNG

TinySearch v0.2.0 is a lightweight open-source MCP/FastAPI web-search tool for small local LLMs. It searches the web, crawls a few pages, chunks/retrieves/reranks useful parts, and outputs a compact context blob capped at 8k tokens — no more dumping 30k tokens of scraped nonsense into agent prompts.
Key Changes in v0.2.0
- SearXNG is now the default search backend (replaces DuckDuckGo)
- You can point TinySearch at your own SearXNG instance
- More flexible and less dependent on a single provider
- Output still capped at 8k tokens, optimized for LLM agents
DuckDuckGo started throwing limits and CAPTCHAs more often in recent weeks. For an MCP tool that agents depend on, that wasn't acceptable. SearXNG adds a bit of overhead — calls now take about 10-15 seconds — but the author says it's worth the convenience.
Who It's For
Developers using smaller local models with Cline, Roo, OpenCode, MCP agents, or any setup where context budgets are tight. The author uses TinySearch daily with Qwen3.5-9B for questions about library versions, function calls, and Azure/GCP API specifics.
Repository: github.com/MarcellM01/TinySearch
📖 Read the full source: r/LocalLLaMA
👀 See Also

Indie dev deploys full game studio site via Claude Code, including Steam API data layer
An indie game developer used Claude Code to build and deploy a game studio website without touching a terminal, including a data layer that pulls live info from the Steam API.

onWatch: Open-source local API quota tracker with SQLite storage
onWatch is a local-first API quota tracker that stores all data in a local SQLite database with no cloud service, telemetry, or account creation. It's a single binary (~13MB) that runs as a background daemon using <50MB RAM and serves a dashboard on localhost.

Developer Builds MCP Server for Claude WhatsApp Integration, Shares Challenges
A developer built an MCP server to give Claude access to real WhatsApp conversations, discovering that conversation context management was trickier than expected and required a database to track conversations.

Local Semantic Memory Search for OpenClaw Agents Using Harrier Embeddings
Run a local embedding server with Microsoft's Harrier model, expose an Ollama-compatible API, and wire OpenClaw's memorySearch config for local semantic memory retrieval without external services.