REMORY: Learning Residual Memory for Context Compaction
Abstract
Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.
Community
Remory appends learned soft memory tokens to a compacted summary. On Qwen3.8-27B and GLM-5.3-Flash, it improves long-horizon benchmark performance while reducing repeated tool outputs and tool errors.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation (2026)
- CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory (2026)
- StateComp: Learning When to Compress History in Long Horizon Agents (2026)
- MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization (2026)
- AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents (2026)
- Persistent Context Graphs for Efficient Memory Compaction in LLM Agents (2026)
- Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Better source attribution at the same coverage is the result I'd care about most. When a summary drops where something came from, nobody can check it later. Do the memory tokens keep any link a person can trace back to the original step, or is that only recoverable through the model?
Thanks for the question!
- The soft tokens are not directly human-readable; their information is recovered through the LLM. This is also the training objective: REMORY learns to generate representations that the frozen LLM can interpret and use to supplement the summary, helping it approximate the actions it would take with the raw context.
- One hypothesis based on the experiments is that the soft tokens may serve as a form of associative memory, helping the LLM connect its working memory (context) with external long-term memory, such as saved files or memory.jsonl. They may preserve cues that help the model recall where relevant information resides and access it again.
Models citing this paper 2
mocoV3/GLM-5.3-Flash-REMORY-1.24B
Datasets citing this paper 0
No dataset linking this paper