Hybrid RAG for Local Agent Memory with OpenClaw, Ollama, and nomic-embed-text

✍️ OpenClawRadar📅 Published: March 10, 2026🔗 Source
Hybrid RAG for Local Agent Memory with OpenClaw, Ollama, and nomic-embed-text
Ad

Problem: Retrieval, Not Storage

The developer had months of daily memory logs stored in markdown files, which worked for saving information but not for finding it again. When the agent needed past context, it would fall back to running ls, opening files one by one, spending tokens, and sometimes missing relevant information. The issue was retrieval by meaning, not storage.

Solution: Hybrid RAG with Local Embeddings

The developer enabled memorySearch in OpenClaw using Ollama as the provider and nomic-embed-text for local embeddings, running in hybrid mode. Hybrid means 70% vector similarity (cosine via nomic-embed-text) combined with 30% BM25 keyword matching. Vector handles semantic proximity while BM25 handles exact names, versions, and IDs. MMR reduces redundant results, and temporal decay gives more weight to recent logs. Everything runs locally without external APIs.

Configuration

"memorySearch": {
  "provider": "ollama",
  "query": {
    "hybrid": {
      "enabled": true,
      "vectorWeight": 0.7,
      "textWeight": 0.3,
      "mmr": {
        "enabled": true,
        "lambda": 0.7
      },
      "temporalDecay": {
        "enabled": true,
        "halfLifeDays": 30
      }
    }
  }
}

Setup Instructions

  • OpenClaw detects Ollama automatically at localhost:11434
  • No need to specify baseUrl or model - it picks up nomic-embed-text if pulled
  • Run ollama pull nomic-embed-text first, then restart the gateway
  • Avoid setting provider: "openai" and pointing baseUrl to Ollama - use provider: "ollama" directly
Ad

Behavioral Change Required

Enabling the tool wasn't enough. Without explicit instructions to use memorySearch before reading files directly, the agent would skip it and take the slower, token-heavy route. The developer wrote a rule into both AGENTS.md and MEMORY.md in the workspace to make memory search part of the agent's normal workflow.

Before vs After Results

  • Before: Browse folders, open files blindly, hope wording matches, waste tokens, miss context
  • After: Run memory_search with semantic query, retrieve ranked results with similarity scores, open best match, answer from actual past notes
  • Similarity scores for relevant results typically range 0.45 to 0.48 for nomic-embed-text on prose logs

Practical Notes

  • nomic-embed-text has a 2048 token context limit by default, not 8192 - large files may get truncated at indexing
  • Memory files in Spanish work well - nomic-embed-text handles Spanish without issues
  • Retrieval quality depends on note quality - vague logs still cause semantic search struggles

Tech Stack

  • OpenClaw (local, self-hosted)
  • Ollama + nomic-embed-text:latest
  • SQLite with sqlite-vec and FTS5 (created automatically by OpenClaw on first use)
  • Mac mini M4, 16GB unified memory

📖 Read the full source: r/openclaw

Ad

👀 See Also

Practical AI Travel Planning Workflow: What Works and What Doesn't
Use Cases

Practical AI Travel Planning Workflow: What Works and What Doesn't

A developer shares their year-long experience using ChatGPT, Claude, and Perplexity to plan trips to six countries, detailing specific strengths like itinerary creation and budget accuracy, weaknesses including incorrect opening hours, and a five-step verification workflow.

OpenClawRadar
Developer Considers Switching from DeepSeek to Grok for Finance AI Agent
Use Cases

Developer Considers Switching from DeepSeek to Grok for Finance AI Agent

A developer building a finance AI web app in FastAPI/Python reports DeepSeek V3.2 Reasoning has 70s TTFT and ~25 t/s output speed, making streaming feel terrible. They're considering Grok 4.1 Fast Reasoning with ~15s TTFT and ~75 t/s output.

OpenClawRadar
OpenClaw user struggles with AI agent automation after successful Claude Code pipeline
Use Cases

OpenClaw user struggles with AI agent automation after successful Claude Code pipeline

A marketing agency owner successfully created an image recreation pipeline using Claude Code in one hour, but encountered problems when trying to teach the same process to an AI agent in OpenClaw running on Gemini 3.1 Pro, with issues including bad reasoning, slow responses, and incorrect outputs.

OpenClawRadar
Developer Implements AI-Ready Feedback Loop for Feature Shipping
Use Cases

Developer Implements AI-Ready Feedback Loop for Feature Shipping

A developer built a feedback system that captures app context and automatically generates structured GitHub issues, then uses Claude Code with a triage skill to turn those issues into scoped development tasks. Two features were shipped using this workflow from mobile devices.

OpenClawRadar