free-claude-code adds GLM-5 support via NVIDIA NIM, expands to OpenRouter and Discord

free-claude-code, a lightweight proxy that converts Claude Code's Anthropic API requests into other provider formats, has been updated with GLM-5 support through NVIDIA NIM and several new features. The tool allows developers to use Claude Code's agentic coding interface without an Anthropic subscription by routing requests to alternative backends.
Key updates
NVIDIA added tool calling fixes for z-ai/glm5 to their NIM inventory, and free-claude-code now fully supports it. The NVIDIA NIM free tier provides 40 requests per minute with no credit card required.
- OpenRouter support: Use any model on OpenRouter's platform as your backend, including their free models
- Discord bot integration: Control Claude Code remotely via Discord in addition to the existing Telegram bot support
- LMStudio local provider support: Run models fully locally
- Claude Code VSCode extension support
Technical advantages
- Zero cost options: NVIDIA NIM free tier (40 reqs/min) and Open Router free models require no payment
- Interleaved thinking preservation: Native interleaved thinking tokens are preserved across turns, allowing models like GLM-5 and Kimi-K2.5 to leverage reasoning from previous turns
- 5 built-in optimizations: Fast prefix detection, title generation skip, suggestion mode skip, and other optimizations reduce unnecessary LLM calls
- Remote control: Telegram and Discord bots enable sending coding tasks from mobile devices with session forking and persistence
- Configurable rate limiter: Sliding window rate limiting for concurrent sessions
- Easy model support: New models launching on NVIDIA NIM can be used with no code changes
- Extensibility: Modular code structure makes it easy to add custom providers or messaging platforms
Supported models
Popular models include z-ai/glm5, moonshotai/kimi-k2.5, minimaxai/minimax-m2.5, qwen/qwen3.5-397b-a17b, and stepfun-ai/step-3.5-flash. The full list is available in nvidia_nim_models.json. With OpenRouter and LMStudio, virtually any model can be used as a backend.
The developer is currently working on automatic model selection based on availability and quality. The project is open source with issues and PRs welcome.
📖 Read the full source: r/ClaudeAI
👀 See Also

ClawNet: Peer-to-Peer AI Agent Network Without API Keys
ClawNet is a peer-to-peer network that allows AI agents to collaborate directly without API keys or platform fees. Installation is via a curl script, and features include a task bazaar, shell economy, and knowledge network.

Jentic Mini: Self-Hosted API and Action Execution Layer for OpenClaw
Jentic Mini is a self-hosted API and action execution layer that sits between AI agents and external APIs, storing credentials in an encrypted vault and providing scoped toolkits with individually revocable keys. It automatically imports 10,000+ OpenAPI specs and Arazzo workflow sources when credentials are added.

Orion: Bypassing CoreML to Run and Train LLMs Directly on Apple Neural Engine
Orion is an open-source Objective-C system that bypasses Apple's CoreML to run and train LLMs directly on the Apple Neural Engine (ANE), achieving 170+ tokens/s for GPT-2 124M decode and stable multi-step training on a 110M parameter transformer.

Memora v0.2.25 MCP Server: 5× Faster Writes on D1 Database
Memora v0.2.25, an MCP server for Claude persistent memory, achieves 5× faster writes on Cloudflare D1 with memory_create dropping from 10s+ to ~1.8s and memory_update from 10s+ to ~1.1s per call.