ClawRelay: macOS-native OpenAI-compatible LLM proxy with automatic failover

What ClawRelay does
ClawRelay is a native Swift application for macOS 15+ that runs an OpenAI-compatible HTTP server locally. You configure LLM providers in priority order (OpenAI, Groq, Nvidia NIMs, Ollama, or any service with a /v1/chat/completions endpoint). When a request comes in, it tries the first provider and automatically falls back to the next if there's a failure (rate limit, 5xx error, or timeout).
Setup and configuration
The app runs in the system tray with quick access and a full settings window. Provider API keys are stored in macOS Keychain. No Docker, Node.js, or config files are required.
To connect your tools:
- Base URL:
http://localhost:11434/v1 - API Key: optional for local use, can be generated in-app for LAN or tunnel setups
Works with Cursor, Continue.dev, LM Studio, the Python openai library, and any tool that accepts a custom base URL.
openClaw integration
For openClaw users, one command wires it up:
bash <(curl -fsSL https://www.desertstack.dev/clawrelay/enable-provider.sh ) \
--provider-id "clawrelay" \
--base-url "http://localhost:11434/v1" \
--api-key "claw_relay_key" \
--api "openai-completions" \
--model-id "clawrelay" \
--model-name "ClawRelay"Generate your key from the Servers tab in ClawRelay. Requires jq and the openclaw CLI.
Deployment options
Beyond localhost, you can bind ClawRelay to your LAN interface to reach it from any device on your network. You can also put Cloudflare Tunnel or ngrok in front to expose it to the internet. The same app and configuration work for all deployment scenarios.
Built-in features
- Request logs included
- System tray access
- Full settings window
- macOS Keychain storage for API keys
- Native Swift implementation
📖 Read the full source: r/clawdbot
👀 See Also

Why Your Claude Code UI Output Drifts and How a Structured Spec Fixes It
A developer explains that inconsistent UI output from Claude Code isn't a prompt problem — it's a format problem. Providing exact hex codes, font weights, spacing, screen states, and transitions eliminates drift. They also open-sourced an MCP server that converts screen recordings into structured specs.

Context Mode MCP Server Cuts Claude Code Context Usage by 98%
Context Mode is an MCP server that reduces Claude Code context consumption from 315 KB to 5.4 KB by sandboxing tool outputs. It supports 10 language runtimes and includes a knowledge base with full-text search.

Roost: A Single-Go-Binary Sidebar for Claude Code with Clickable Prompt History, File Tree, and Notifications
Roost is a single Go binary that adds a web-based sidebar to Claude Code: xterm.js terminal backed by tmux, file tree that follows your cd, clickable prompt history from ~/.claude/projects/*.jsonl, and push notifications via Claude Code's Stop hook. Run over SSH as single-user-per-instance; no build step on the frontend.

Cloken: A Chrome Extension That Shows Real-Time Claude Context Usage as a Percentage
Cloken is a free Chrome extension that displays your current Claude.ai chat context usage as a percentage — including messages, files, images, and system prompt.