RouteLLM Setup for Cost-Effective AI Task Routing

Docker Compose Configuration for Hybrid AI Setup
A Reddit user posted a detailed Docker Compose setup that implements what they call "Poor Man's Superintelligence" - a hybrid AI system that routes tasks between local and cloud models based on complexity.
Core Components
The system uses four main services:
- vscode-openwire: Uses image
sendmeticket/vscode-openwire:1.0.0with ports 3000 and 3030 exposed. This provides access to GitHub Copilot through OpenWire, though the source notes this may violate TOS and suggests using an available API key instead. - ollama: Runs
ollama/ollama:latestwith port 11434 exposed. It automatically pulls and serves theqwen3.5:4bmodel as the local "weak" model. - openroutellm: Uses image
sendmeticket/openroutellm:1.0.0on port 6060. This is the routing service that decides which model handles each request. - openclaw: Runs
ghcr.io/openclaw/openclaw:latestwith ports 18789 and 18790 exposed, serving as the main interface.
RouteLLM Configuration
The openroutellm service is configured with specific parameters:
python -m routellm.openai_server --routers bert --default-router-threshold 0.75 --port 6060 --openwire-base-url http://vscode-openwire:3030/v1 --ollama-base-url http://ollama:11434/v1 --strong-model gpt-4o --weak-model qwen3.5:4bThis setup uses BERT-based routing with a 0.75 threshold to determine when to send tasks to the "strong" model (GPT-4o) versus the local "weak" model (Qwen3.5:4b).
How It Works
The system routes difficult tasks to the paid GPT-4o model through OpenWire/Copilot, while simpler tasks are handled by the local Qwen3.5:4b model running in Ollama. This creates what the author describes as a "fail-safe, local-first AI model with low base intelligence but really high max intelligence."
All services are connected through a custom Docker network (openclaw_net with subnet 172.10.10.0/24) and include health checks to ensure service availability.
📖 Read the full source: r/LocalLLaMA
👀 See Also

Benchmark Results: 15 LLMs Tested on 38 Real Workflow Tasks
A developer benchmarked 15 cloud and local LLMs on 38 tasks from their actual workflow, including CSV transforms, letter counting, modular arithmetic, and format compliance. Claude 3.5 Sonnet and Opus both scored 100%, but Sonnet costs 3.5x less per call.

How Claude Helped Reverse-Engineer Garmin’s BLE Protocols to Fake a Native Running Sensor
A developer used Claude to reverse-engineer Garmin’s undocumented BLE protocols, making an ESP32 look like a native HRM strap — dual identity switching and running dynamics RE.

I ripped out OpenClaw's default markdown memory and built a Node.js/Postgres API layer instead
A developer disabled OpenClaw's memory-core plugin and built a typed Node.js/Express + PostgreSQL backend. Context drift dropped to zero.

PowerShell Script Automates OpenClaw Docker Setup on Windows
A PowerShell script handles Windows-specific networking quirks and Docker configuration for OpenClaw, automating checks, image retrieval, setup guidance, and container deployment.