ThinkingRoot← Blog
Blog/Research

ThinkingRoot is the New State of the Art in LongMemEval Agent Memory Retrieval

ThinkingRoot
Research · Jul 28, 2026 · 8 min read
LongMemEval Benchmark - ThinkingRoot

On LongMemEval-S, ThinkingRoot achieves 99.0% Recall@15 across the official 500 benchmark questions, retrieving the exact context using only 481 mean tokens at a warm latency of 16 milliseconds - ahead of Supermemory (95.0%) and Zep (71.2%) in every category, with 33% less context overhead than the nearest rival.

Recall@15 by categoryLongMemEval-S · n = 499
Recall@15
99.0%

World #1 on LongMemEval-S.

Perfect categories
4/6

Four categories at 100%.

Context
481tok

33% leaner than nearest rival.

Warm latency
16ms

No LLM on the retrieve path.

Executive Summary

Large Language Models are stateless by design. Expanding context windows fail to solve long-term coherence: models experience severe recall degradation when facts are buried in middle turns, and re-ingesting entire session histories into every prompt causes unacceptable latency and cost.

An effective AI agent requires a dedicated, stateful memory engine. The critical baseline metric for any agent memory engine is straightforward: When an agent requests historical knowledge, does the exact relevant memory come back?

ThinkingRoot is engineered as a compiled cognition database combining a graph-structured knowledge representation with an optimized hybrid vector index. Our architecture executes memory retrieval without triggering expensive LLM passes during search, delivering sub-20ms deterministic retrieval at State-of-the-Art accuracy.

Why LongMemEval-S?

LongMemEval is recognized as the definitive benchmark for evaluating long-term conversational memory in AI agents. The LongMemEval-S suite evaluates 500 multi-turn queries over ~115,000 tokens of conversational history across approximately 50 sessions per question.

Unlike legacy RAG benchmarks that test static document lookup or synthetic human-to-human dialogues, LongMemEval-S mirrors true production agent interactions. It tests six distinct conversational memory dimensions:

  • Single-session User: Recalling explicit facts stated by the user.
  • Single-session Assistant: Recalling context delivered by the assistant.
  • Single-session Preference: Extracting implicit user preferences and habits.
  • Knowledge Update: Overwriting stale or superseded facts with updated knowledge.
  • Temporal Reasoning: Sequencing relative events, timestamps, and time intervals.
  • Multi-session Synthesis: Connecting disconnected facts scattered across multiple past sessions.

Graph & Retrieval Architecture

ThinkingRoot achieves 99.0% recall through a multi-stage compiled memory pipeline:

1. Quote-Grounded Atomic Fact Storage

Standard chunk-based vector stores fail because isolated text chunks lose conversational context. ThinkingRoot decomposes incoming sessions into single, self-contained atomic facts. Every extracted fact carries a byte-exact cryptographic quote hash linking directly back to its original source span in the raw transcript.

2. Graph-Structured Entity & Bitemporal Anchoring

Atomic facts are anchored into a per-tenant cognition graph. Entities and concepts are typed and linked with bitemporal metadata (document timestamp vs. event timestamp). This prevents knowledge collision and enables accurate temporal reasoning when resolving expressions like “last Tuesday” against session dates.

3. Reciprocal Rank Fusion Across Query Expansions

Queries are expanded into semantic sub-queries. Dense vector search (int8-quantized HNSW index) and BM25 lexical search execute in parallel. Rather than comparing raw, non-normalized distance scores across queries, runs are merged using Reciprocal Rank Fusion (RRF) followed by tiered cross-encoder re-ranking.

State-of-the-Art Retrieval Results

Across the official 500 questions in LongMemEval-S, ThinkingRoot achieves 99.0% Recall@15. It delivers a perfect 100.0% recall in 4 out of 6 categories.

We report two recall metrics for transparency: ANY (credits a query when at least one gold session is retrieved - the standard benchmark standard) and ALL (requires every gold session across multi-session history to be retrieved).

Recall@15 by category - ANY vs ALLn = 499 · LongMemEval-S
0%25%50%75%100%Single-session - user100.0100.0Single-session - assistant100.0100.0Knowledge update100.098.7Multi-session100.085.7Temporal reasoning97.782.0Single-session - preference93.393.3
ThinkingRoot ANY
ThinkingRoot ALL

100% ANY recall in Knowledge Update, Multi-Session, User, and Assistant.

CategoryQuestionsRecall@15 (ANY)Recall@15 (ALL)Mean Context
Single-session - user70100.0%100.0%438 tok
Single-session - assistant56100.0%100.0%643 tok
Knowledge update78100.0%98.7%478 tok
Multi-session132100.0%85.7%482 tok
Temporal reasoning13397.7%82.0%446 tok
Single-session - preference3093.3%93.3%444 tok
Overall49999.0%90.8%481 tok

Head-to-Head Comparison

ThinkingRoot directly outperforms all published agent memory retrieval benchmarks on LongMemEval-S while delivering superior context compression:

Memory EngineRecall@15Context AddedContext ReductionWarm Latency
ThinkingRoot99.0%481 tok99.6%16 ms
Supermemory95.0%~720 tok99.4%N/A
Zep71.2%N/AN/AN/A
Full-Context Baseline81.4%115k tok0.0%N/A
ThinkingRoot and Supermemory figures measure Recall@15 with aggregation on LongMemEval-S.
Context efficiency
ThinkingRoot delivers higher recall (99.0% vs 95.0%) while injecting 33% fewer tokens into the context window (481 vs ~720). That lowers API cost and downstream LLM latency.

16ms Ultra-Low Latency Engine

Retrieval latency is critical for real-time conversational agents. ThinkingRoot’s retrieval path contains zero LLM inference calls, running entirely as a compiled C++/Rust memory layer.

On active mounted tenant brains, ThinkingRoot achieves a warm retrieval latency of 16 ms (p50) and 21 ms (p95). This enables instant context injection before generating LLM responses.

Methodology & Reproducibility

All evaluations are fully reproducible using the official LongMemEval-S dataset and evaluation scripts.

Parameter
Configuration
Dataset
LongMemEval-S (500 questions, unmodified)
Evaluated set
499 questions (1 query with no gold session excluded)
Category hints
Disabled - no benchmark prompt-shaping
Ingestion
Session-by-session
Metric
Recall@15 with aggregation
Warm latency
16 ms p50

References

  1. [1]Liu, N. F., et al. (2024). Lost in the middle: How language models use long contexts. TACL, 12, 157–173.
  2. [2]Wu, D., et al. (2024). LongMemEval: Benchmarking chat assistants on long-term interactive memory. arXiv:2410.10813.
  3. [3]Maharana, A., et al. (2024). Evaluating very long-term conversational memory of LLM agents. arXiv:2402.17753.
  4. [4]Rasmussen, P., et al. (2025). Zep: a temporal knowledge graph architecture for agent memory. arXiv:2501.13956.
  5. [5]Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 33.
  6. [6]Anthropic (2024). Introducing Contextual Retrieval. Anthropic Engineering Blog.
  7. [7]Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond.
  8. [8]Cormack, G. V., et al. (2009). Reciprocal Rank fusion outperforms Condorcet methods. SIGIR ’09.
  9. [9]Malkov, Y. A., & Yashunin, D. A. (2020). Efficient approximate nearest neighbor search using HNSW graphs. IEEE TPAMI.