Bypassing NemoClaw Sandbox Isolation for Local Nemotron 9B Agent

✍️ OpenClawRadar📅 Published: March 18, 2026🔗 Source
Bypassing NemoClaw Sandbox Isolation for Local Nemotron 9B Agent
Ad

Local NemoClaw Workaround for Full Local Inference

A developer has documented a method to bypass NVIDIA's NemoClaw sandbox isolation to run a fully local AI agent. NemoClaw, launched at GTC, is an enterprise sandbox for AI agents built on OpenShell (k3s + Landlock + seccomp) that by default expects cloud API connections and heavily restricts local networking.

Ad

Technical Implementation Details

The developer wanted 100% local inference on WSL2 + RTX 5090 and punched through the sandbox to reach a vLLM instance. The solution involved multiple components:

  • Host iptables configuration: Allowed traffic from Docker bridge to vLLM on port 8000
  • Pod TCP Relay: Custom Python relay in the Pod's main namespace bridging sandbox veth → Docker bridge
  • Sandbox iptables injection: Used nsenter to inject ACCEPT rule into the sandbox's OUTPUT chain, bypassing the default REJECT
  • Tool Call Translation: Built a custom Gateway that intercepts the streaming SSE response from vLLM, buffers it, parses Nemotron 9B's <TOOLCALL>[...]</TOOLCALL> text output, and rewrites it into OpenAI-compatible tool_calls in real-time

This configuration allows opencode inside the sandbox to use Nemotron as a fully autonomous agent. Everything runs locally with no data leaving the machine. The setup is volatile (WSL2 reboots wipe the iptables hacks), but enables a 9B model to execute terminal commands inside a locked-down enterprise container.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Memento Vault: Local Tool for Persistent Context in Claude Code Sessions
Tools

Memento Vault: Local Tool for Persistent Context in Claude Code Sessions

Memento Vault is a set of hooks that automatically captures session transcripts, scores them, and stores atomic notes in a local git repo. It provides zero-cost retrieval via BM25 + vector search with 472ms average latency and injects relevant context at session start, on every prompt, and on file reads.

OpenClawRadar
Claw Compactor: 14-stage token compression engine for LLM pipelines
Tools

Claw Compactor: 14-stage token compression engine for LLM pipelines

Claw Compactor is an open-source LLM token compression engine using a 14-stage Fusion Pipeline to achieve 54% average compression with zero LLM inference cost. It includes specialized compressors for code, JSON, logs, diffs, and search results with reversible compression capabilities.

OpenClawRadar
New Structured Data API Provides Subscription Pricing for LLM Agents
Tools

New Structured Data API Provides Subscription Pricing for LLM Agents

A developer has released a structured data API that normalizes subscription pricing across streaming platforms, ride-share services, dating apps, and other subscription-based platforms. The API provides consistent JSON schemas, region-aware pricing where available, and MCP-compatible endpoints for LLM agents to consume without scraping.

OpenClawRadar
Files.md: Open-Source Local-First Markdown Note-Taking App with LLM-Friendly Design
Tools

Files.md: Open-Source Local-First Markdown Note-Taking App with LLM-Friendly Design

Files.md is an open-source, local-first markdown app for notes, tasks, and journals. 886 stars, built in Go, works offline, syncs via iCloud/Dropbox/self-hosted server or hosted beta app.files.md.

OpenClawRadar