agentmemory V4 achieves 96.2% on LongMemEval benchmark, outperforms commercial AI memory systems

✍️ OpenClawRadar📅 Published: March 27, 2026🔗 Source
agentmemory V4 achieves 96.2% on LongMemEval benchmark, outperforms commercial AI memory systems
Ad

agentmemory V4 is an open-source memory system for AI agents that just achieved a world record score of 96.2% on LongMemEval, the standard benchmark for long-term AI agent memory.

Benchmark Performance

The system outperformed several funded AI memory companies:

  • PwC Chronos: 95.6%
  • Mastra: 94.87%
  • OMEGA: 93.2% (raw)
  • Supermemory: 85.86%
  • Emergence AI: 86%
  • Zep: 71.2%

Development Details

Built solo in 16 days on a mid-range gaming PC (i3-12100F) with a total cost of $1,000. The system uses Claude Opus as a generator and GPT-4o as a judge, but the retrieval architecture is the core innovation.

Technical Architecture

The system combines multiple retrieval techniques in a single SQLite-backed system:

  • HNSW (Hierarchical Navigable Small World) for approximate nearest neighbor search
  • BM25 for traditional text retrieval
  • Cross-encoder for relevance scoring
  • Knowledge graph integration
  • Temporal grounding for time-aware memory retrieval

Availability

The system is open source under the MIT license and available at: github.com/JordanMcCann/agentmemory

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

Leanstral: Open-Source Code Agent for Lean 4 and Formal Proof Engineering
Tools

Leanstral: Open-Source Code Agent for Lean 4 and Formal Proof Engineering

Mistral AI released Leanstral, the first open-source code agent designed for Lean 4, with 6B active parameters and Apache 2.0 licensing. Benchmarks show it outperforms larger open-source models and offers competitive performance to Claude at significantly lower cost.

OpenClawRadar
🦀
Tools

Back Office Simulator: A WebGL Tribute to Manual Data Entry

A browser-based game that simulates manual back-office data entry, built with WebGL 2 and JavaScript. It's a satirical take on the types of tasks AI is making obsolete.

OpenClawRadar
PromptFlow Voice: Speak Hindi, Get Structured Claude Code Prompts — Built with Claude Code
Tools

PromptFlow Voice: Speak Hindi, Get Structured Claude Code Prompts — Built with Claude Code

A developer used Claude Code and Codex to build PromptFlow Voice — a desktop app that converts natural Hindi speech into structured dev prompts for Claude Code, or ready-to-send emails for Gmail. The demo shows a Hindi voice request producing an English Claude Code prompt about socket reconnection logic with exponential backoff.

OpenClawRadar
MemAware Benchmark Tests AI Memory Beyond Keyword Search
Tools

MemAware Benchmark Tests AI Memory Beyond Keyword Search

MemAware is a benchmark with 900 questions across 3 difficulty levels that tests whether AI assistants with memory can surface relevant context when queries don't hint at it. Results show BM25 search scored 2.8% vs 0.8% with no memory, while vector search drops to 0.7% on cross-domain connections.

OpenClawRadar