REMORY: Learning Residual Memory for Context Compaction Paper • 2610.11287 • Published 4 days ago • 38
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling Paper • 2609.36529 • Published 13 days ago • 37
PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs Paper • 2609.34715 • Published 14 days ago • 62
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published about 1 month ago • 266
DepthBench: Measuring How Residual Connections Enable More Computational Depth Paper • 2609.32534 • Published 16 days ago • 32
Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models Paper • 2609.34972 • Published 14 days ago • 30
FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation Paper • 2609.27657 • Published 19 days ago • 9
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 18 days ago • 105
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 20 days ago • 37
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 25 days ago • 115
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 25 days ago • 228
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 25 days ago • 78