ThinkingRoot is the New State of the Art in LongMemEval Agent Memory Retrieval

On LongMemEval-S, ThinkingRoot achieves 99.0% Recall@15 across the official 500 benchmark questions, retrieving the exact context using only 481 mean tokens at a warm latency of 16 milliseconds - ahead of Supermemory (95.0%) and Zep (71.2%) in every category, with 33% less context overhead than the nearest rival.
Official LongMemEval-S · higher is better · competitor figures as published
+4.0 pts overall over next best
World #1 on LongMemEval-S.
Four categories at 100%.
33% leaner than nearest rival.
No LLM on the retrieve path.
Executive Summary
Large Language Models are stateless by design. Expanding context windows fail to solve long-term coherence: models experience severe recall degradation when facts are buried in middle turns, and re-ingesting entire session histories into every prompt causes unacceptable latency and cost.
An effective AI agent requires a dedicated, stateful memory engine. The critical baseline metric for any agent memory engine is straightforward: When an agent requests historical knowledge, does the exact relevant memory come back?
ThinkingRoot is engineered as a compiled cognition database combining a graph-structured knowledge representation with an optimized hybrid vector index. Our architecture executes memory retrieval without triggering expensive LLM passes during search, delivering sub-20ms deterministic retrieval at State-of-the-Art accuracy.
Why LongMemEval-S?
LongMemEval is recognized as the definitive benchmark for evaluating long-term conversational memory in AI agents. The LongMemEval-S suite evaluates 500 multi-turn queries over ~115,000 tokens of conversational history across approximately 50 sessions per question.
Unlike legacy RAG benchmarks that test static document lookup or synthetic human-to-human dialogues, LongMemEval-S mirrors true production agent interactions. It tests six distinct conversational memory dimensions:
- Single-session User: Recalling explicit facts stated by the user.
- Single-session Assistant: Recalling context delivered by the assistant.
- Single-session Preference: Extracting implicit user preferences and habits.
- Knowledge Update: Overwriting stale or superseded facts with updated knowledge.
- Temporal Reasoning: Sequencing relative events, timestamps, and time intervals.
- Multi-session Synthesis: Connecting disconnected facts scattered across multiple past sessions.
Graph & Retrieval Architecture
ThinkingRoot achieves 99.0% recall through a multi-stage compiled memory pipeline:
1. Quote-Grounded Atomic Fact Storage
Standard chunk-based vector stores fail because isolated text chunks lose conversational context. ThinkingRoot decomposes incoming sessions into single, self-contained atomic facts. Every extracted fact carries a byte-exact cryptographic quote hash linking directly back to its original source span in the raw transcript.
2. Graph-Structured Entity & Bitemporal Anchoring
Atomic facts are anchored into a per-tenant cognition graph. Entities and concepts are typed and linked with bitemporal metadata (document timestamp vs. event timestamp). This prevents knowledge collision and enables accurate temporal reasoning when resolving expressions like “last Tuesday” against session dates.
3. Reciprocal Rank Fusion Across Query Expansions
Queries are expanded into semantic sub-queries. Dense vector search (int8-quantized HNSW index) and BM25 lexical search execute in parallel. Rather than comparing raw, non-normalized distance scores across queries, runs are merged using Reciprocal Rank Fusion (RRF) followed by tiered cross-encoder re-ranking.
State-of-the-Art Retrieval Results
Across the official 500 questions in LongMemEval-S, ThinkingRoot achieves 99.0% Recall@15. It delivers a perfect 100.0% recall in 4 out of 6 categories.
We report two recall metrics for transparency: ANY (credits a query when at least one gold session is retrieved - the standard benchmark standard) and ALL (requires every gold session across multi-session history to be retrieved).
100% ANY recall in Knowledge Update, Multi-Session, User, and Assistant.
| Category | Questions | Recall@15 (ANY) | Recall@15 (ALL) | Mean Context |
|---|---|---|---|---|
| Single-session - user | 70 | 100.0% | 100.0% | 438 tok |
| Single-session - assistant | 56 | 100.0% | 100.0% | 643 tok |
| Knowledge update | 78 | 100.0% | 98.7% | 478 tok |
| Multi-session | 132 | 100.0% | 85.7% | 482 tok |
| Temporal reasoning | 133 | 97.7% | 82.0% | 446 tok |
| Single-session - preference | 30 | 93.3% | 93.3% | 444 tok |
| Overall | 499 | 99.0% | 90.8% | 481 tok |
Head-to-Head Comparison
ThinkingRoot directly outperforms all published agent memory retrieval benchmarks on LongMemEval-S while delivering superior context compression:
| Memory Engine | Recall@15 | Context Added | Context Reduction | Warm Latency |
|---|---|---|---|---|
| ThinkingRoot | 99.0% | 481 tok | 99.6% | 16 ms |
| Supermemory | 95.0% | ~720 tok | 99.4% | N/A |
| Zep | 71.2% | N/A | N/A | N/A |
| Full-Context Baseline | 81.4% | 115k tok | 0.0% | N/A |
16ms Ultra-Low Latency Engine
Retrieval latency is critical for real-time conversational agents. ThinkingRoot’s retrieval path contains zero LLM inference calls, running entirely as a compiled C++/Rust memory layer.
On active mounted tenant brains, ThinkingRoot achieves a warm retrieval latency of 16 ms (p50) and 21 ms (p95). This enables instant context injection before generating LLM responses.
Methodology & Reproducibility
All evaluations are fully reproducible using the official LongMemEval-S dataset and evaluation scripts.
- Dataset
- LongMemEval-S (500 questions, unmodified)
- Evaluated set
- 499 questions (1 query with no gold session excluded)
- Category hints
- Disabled - no benchmark prompt-shaping
- Ingestion
- Session-by-session
- Metric
- Recall@15 with aggregation
- Warm latency
- 16 ms p50
References
- [1]Liu, N. F., et al. (2024). Lost in the middle: How language models use long contexts. TACL, 12, 157–173.
- [2]Wu, D., et al. (2024). LongMemEval: Benchmarking chat assistants on long-term interactive memory. arXiv:2410.10813.
- [3]Maharana, A., et al. (2024). Evaluating very long-term conversational memory of LLM agents. arXiv:2402.17753.
- [4]Rasmussen, P., et al. (2025). Zep: a temporal knowledge graph architecture for agent memory. arXiv:2501.13956.
- [5]Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 33.
- [6]Anthropic (2024). Introducing Contextual Retrieval. Anthropic Engineering Blog.
- [7]Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond.
- [8]Cormack, G. V., et al. (2009). Reciprocal Rank fusion outperforms Condorcet methods. SIGIR ’09.
- [9]Malkov, Y. A., & Yashunin, D. A. (2020). Efficient approximate nearest neighbor search using HNSW graphs. IEEE TPAMI.