OmniRecall Beta: FAISS-Powered Memory Injection for Cloud LLM Chats

What OmniRecall Does
OmniRecall is a local mitmproxy bypass that intercepts traffic to cloud chat interfaces (tested on DeepSeek). It hacks into the proprietary SSE fragment stream and forces a long-term memory layer onto a system that was designed to be stateless.
Technical Mechanism
- Deep-Packet Parsing: Reconstructs the full assistant reply by tracking real-time patches
- Command Control: Detects [ADD], [UPDATE], [REMOVE], [CLEAR] from the AI's output
- Local Brain: Maintains memory.txt + FAISS index (sentence-transformers MiniLM-L6)
- Context Injection: Top recalled facts get force-fed into your next message as [RECALL: ...]
Current Status & Limitations
This is a beta/experimental release. The developer notes: "This is the closest I've gotten to the dream after weeks of debugging hell. It is buggy. It is experimental. [ADD] is mostly stable, but [SEARCH] is temperamental—if you want perfection, fix it yourself. I've hit my energy limit on this build."
Upstream UI changes will break it. The developer states: "If it breaks, that's on you now."
Requirements & Setup
Potato-PC Requirements:
- CPU only (faiss-cpu + all-MiniLM-L6-v2)
- No local LLM needed — augments the cloud models you already use
- Zero cost, zero API keys, 100% local data isolation
How to Deploy:
pip install mitmproxy faiss-cpu sentence-transformers numpyTrust the mitmproxy CA cert on your OS/browser (run mitmproxy once to generate it). Set system proxy to 127.0.0.1:8080. Then run:
mitmdump -s omnirecall.pyGo to chat.deepseek.com and start feeding it memories.
License Terms
The project uses an aggressively restrictive source-available license:
- No commercial use
- No private forks
- Mandatory public ALTERATIONS.md for any logic changes
- If you port to Claude/GPT-4o/whatever, keep it public per the license
The developer explains: "I've watched too many solo-dev projects get strip-mined, privatized, or turned into paid SaaS while the creator gets zero. This license isn't friendly—it's built to protect the work from exactly those people. If the terms scare you off, that's the point."
📖 Read the full source: r/LocalLLaMA
👀 See Also

Ink: A Deployment Platform Where Claude AI Agents Are the Primary Users
Ink (ml.ink) is a deployment platform designed for AI agents like Claude, featuring one tool call deployment, auto-detection of frameworks, and integrated services including compute, databases, DNS, secrets, domains, metrics, and logs.

Terminal-Based 3D Renderer Built with Multi-Agent Claude Code System
A developer created tortuise, a pure terminal-based 3D renderer that displays Gaussian splats using Unicode and ASCII symbols, built over 3 days using 70-80 AI agents coordinated through a Claude Code setup with subagents inside subagents.

htmLLM-124M v2 Released: Specialized HTML/Bootstrap Autocomplete Model
LH-Tech-AI released htmLLM-124M v2, a 124M parameter model specialized for HTML/Bootstrap autocompletion that achieves 0.91 validation loss and trains in ~8 hours on a single T4 GPU.

ACO System: Multi-Agent AI Pipeline from GitHub Issue to Merged PR
ACO System is an open-source multi-agent framework where six specialized AI agents autonomously run the entire dev pipeline from GitHub Issue to merged PR, with a deterministic Architect gate that rejects bad stories before they reach developers.