This repository curates papers and blogs on long-context language modeling, covering surveys; efficient attention; KV-cache optimization; recurrent transformers and state-space models; position encoding & length extrapolation; long-context training; long-term memory; retrieval-augmented generation; in-context learning; context and model compression; long reasoning (long CoT); long video & image; long-horizon agents; long-text generation; inference acceleration; benchmarks & evaluation; and technical reports.
π₯ Must-read papers for LLM-based Long Context Modeling.
π₯β‘π₯ Thanks for all the great contributors on GitHub!
ππ€π I have the privilege of joining [LCLM-Horizon] and collaborating with them on providing a very complete and comprehensive scholarly survey (A Comprehensive Survey on Long Context Language Modeling) and repository (A-Comprehensive-Survey-For-Long-Context-Language-Modeling) dedicated to Long Context Language Modeling. I look forward to collaborating with them to advance research and deepen understanding in this area!
Taxonomy at a glance
flowchart LR
LCLM["Long-Context Modeling"]
LCLM --> A["Attention & KV Cache"]
LCLM --> T["Training & Alignment"]
LCLM --> M["Memory & RAG"]
LCLM --> C["Compression"]
LCLM --> R["Reasoning & Generation"]
LCLM --> V["Multimodal / Video"]
LCLM --> E["Evaluation & Acceleration"]
A --> A1["Sparse / Linear / IO-aware Attention"]
A --> A2["Eviction / Quantization / Offloading"]
T --> T1["Continual Pretraining / Long-SFT"]
T --> T2["Adaptation & RL for Long Context"]
M --> M1["Long-Term Memory"]
M --> M2["RAG / Hybrid Long-Context"]
C --> C1["Context Compression"]
C --> C2["Model Compression"]
R --> R1["Long CoT"]
R --> R2["Long-Form Text Generation"]
If you find our repository and survey useful for your research, please consider citing the following paper:
@article{liu2025comprehensive,
title={A Comprehensive Survey on Long Context Language Modeling},
author={Liu, Jiaheng and Zhu, Dawei and Bai, Zhiqi and He, Yancheng and Liao, Huanxuan and Que, Haoran and Wang, Zekun and Zhang, Chenchen and Zhang, Ge and Zhang, Jiebin and others},
journal={arXiv preprint arXiv:2503.17407},
year={2025}
}- π’ News
- π Papers
- 1. Survey Papers
- 2. Efficient Attention
- 3. KV-Cache Optimization
- 4. Recurrent Transformers
- 5. State Space Models & Hybrids
- 6. Position Encoding & Length Extrapolation
- 7. Long-Context Training
- 8. Long-Term Memory
- 9. Retrieval-Augmented Generation
- 10. In-Context Learning (Many-shot / Long-ICL)
- 11. Context Compression
- 12. Model Compression for Long Context
- 13. Long Reasoning (Long CoT)
- 14. Long Video & Image
- 15. Long-Horizon Agents
- 16. Long-form Text Generation
- 17. Inference Acceleration & Serving
- 18. Benchmarks & Evaluation
- 19. Technical Reports (Long-Context Models)
- 20. Blogs & Tutorials
- Acknowledgements
-
[2026.08.14]
- Paper: SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
- Paper: KV Cache Compression Through the Lens of Transform Coding
- Paper: MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends
- Paper: AgentRewind: Recoverable Execution for Long-Horizon LLM Agents
- Paper: ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond
- Paper: MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
- Paper: Handover of In-Context Learning State Across Session Boundaries
-
[2026.08.13]
- Paper: The Query Knows What to Forget: A Second Erase Direction for Linear Attention
- Paper: When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation
- Paper: RippleMem: From Isolated Retrieval to Associative Recollection for Long-Term Agent Memory
- Paper: vToken: Token-Level Virtualization for Reclaimable KV Caches
- Paper: SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention
- Paper: LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
- Paper: AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
- Paper: PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
- Paper: Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- Paper: Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
- Paper: CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
- Paper: NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
- Paper: EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
-
[2026.08.12]
- Paper: EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
- Paper: Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
- Paper: MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
- Paper: Disentangling the Expressivity of RoPE
- Paper: LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
- Paper: Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
- Paper: The Sleeping Agent: What Gist-Based Context Compression Loses and Why
- Paper: Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
- Paper: Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem
- Paper: Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
- Paper: LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
- Paper: Hybrid Gated Attention
- Paper: Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
- Paper: Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
-
[2026.08.11]
- Paper: Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
- Paper: Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models
- Paper: ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover
- Paper: When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
- Paper: SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis
- Paper: Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory
- Paper: StreamFlow: Dynamic Memory Flows for Streaming Video Understanding
- Paper: InSight-doc: Agentic Visual Perception for Long-Document Understanding
- Paper: R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
- Paper: ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
-
[2026.08.10]
- Paper: Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension
- Paper: MixFormer: Linear Transformer with Mixture of Memory Experts
- Paper: KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
- Paper: MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory
- Paper: Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
- Paper: Evo-Bench: Can Language Models Improve Agent Harness?
- Paper: BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
- Paper: Motif 3: Technical Report
-
[2026.08.09]
- Paper: DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference
- Paper: RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation
- Paper: VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
- Paper: Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling
- Paper: VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
- Paper: Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
- Paper: Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
-
[2026.08.08]
- Paper: OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
- Paper: SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding
- Paper: CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents
- Paper: SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
-
[2026.08.07]
- Paper: CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
- Paper: HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management
- Paper: Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression
- Paper: Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
- Paper: StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear Recurrence
- Paper: The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents
- Paper: RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
- Paper: Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
- Paper: Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
- Paper: An AI4AI Framework for Visual Token Pruning
- Paper: DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
- Paper: MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
- Paper: Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
- Paper: MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
- Paper: HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
- Paper: Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
- Paper: CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
- Blog: Efficient Decode Context Parallelism with vLLM for Long Context Workloads
-
[2026.08.06]
- Paper: Retrofitting Linear Attention into Diffusion Language Models
- Paper: Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
- Paper: Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability
- Paper: StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
- Paper: One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
- Paper: TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
- Paper: Runtime Observability for Heterogeneous Attention Memory
- Paper: Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
- Paper: Retrofitting Linear Attention into Diffusion Language Models
-
[2026.08.05]
- Paper: Recursive Synthesis for Long-Horizon Terminal Tasks
- Paper: QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
- Paper: OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
- Paper: Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression
- Paper: Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning
- Paper: MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off
- Paper: Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems
- Paper: Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
- Paper: ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- Paper: EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
- Paper: Thinking with Anchors: Grounded and Efficient Document Reasoning
- Paper: Chained Recursive Language Models for Multi-Iteration Reasoning
- Paper: Training-Free Hashing-Based Attention via Binary Principal Components
- Paper: Recursive Synthesis for Long-Horizon Terminal Tasks
-
[2026.08.04]
- Paper: Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms
- Paper: TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series
- Paper: Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
- Paper: TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning
- Paper: PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory
- Paper: Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- Paper: Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention
- Paper: SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
- Paper: RUTA: Principled Visual Token Allocation via Rate-Utility Optimization
- Paper: Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
- Paper: GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
- Paper: StreamDAM: Presence-Aware Memory for Real-Time Streaming Video Object Segmentation
- Paper: When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
- Paper: Muon Meets Mamba: Spectral Optimization for State Space Models
- Paper: Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
- Paper: EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
- Paper: When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
- Paper: Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
- Paper: Interpretable Adaptive Sampling for LLM Test-Time Scaling
- Paper: SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval
- Paper: OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
Month Papers
-
[2026.08.03]
- Paper: ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference
- Paper: AnchorKV: Anchor-Residual KV Cache Compression
- Paper: Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
- Paper: DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
- Paper: Output-Aware Rotation for INT2 KV-Cache Quantization
- Paper: Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation
- Paper: LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
- Paper: Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression
- Paper: LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
- Paper: AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
- Paper: GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
- Paper: Learning What to Remember: Test-Time Training via Context Distillation
- Paper: Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
- Paper: HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses
- Paper: Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents
- Paper: Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
- Paper: IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
- Paper: DiffPrune: differentiable information throttling for token pruning in vision-language models
- Paper: ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs
- Paper: CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
- Paper: Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection
-
[2026.08.02]
- Paper: Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
- Paper: RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
- Paper: Practical Online KV Cache Compaction for LLM Agents: An Empirical Study
- Paper: Think in Sets for Streaming Video Token Compression
- Paper: PMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM Agents
- Paper: An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
- Paper: Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere
-
[2026.08.01]
-
[2026.07.31]
- Paper: SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering
- Paper: ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression
- Paper: Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models
- Paper: DarwinX: Evolving Agent Harnesses Through Natural Selection
- Paper: Cross-Benchmark Generalization in Long-Horizon Agents
-
[2026.07.30]
- Paper: SemPIC: Learning Semantic Position-Independent KV Caches
- Paper: Recall Before You Rank: Similarity-Guided Top-$K$ Reuse for Efficient Long-Context Attention
- Paper: MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
- Paper: ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
- Paper: Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
- Paper: VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
- Paper: ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
- Paper: LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
- Paper: Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
- Paper: RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
- Paper: Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
- Paper: SemPIC: Learning Semantic Position-Independent KV Caches
-
[2026.07.29]
- Paper: Metis: Memory Foundation Model
- Paper: MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
- Paper: ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
- Paper: Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
- Paper: Mergeable Model-Side Aggregation States for Long-Context Language Models
- Paper: FreqForcing: Autoregressive Long Video Generation via Spectral Self-Anchoring
- Paper: Metis: Memory Foundation Model
-
[2026.07.28]
- Paper: CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
- Paper: CHILL-Harness: Counterfactual Harness Learning for Efficient Reasoning in Long-Horizon Agents
- Paper: Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
- Paper: Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns
- Paper: CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
-
[2026.07.27]
- Paper: Kimi K3: Open Frontier Intelligence
- Paper: PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention
- Paper: Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- Paper: MemTX: Transactional Belief Commit for Stateful Agent Memory
- Paper: Addressable Recall Compaction for Long Context-Window Control in AI Agents
- Paper: DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation
- Paper: LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
- Paper: Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- Paper: DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- Paper: Kimi K3: Open Frontier Intelligence
-
[2026.07.26]
-
[2026.07.25]
-
[2026.07.24]
- Paper: HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding
- Paper: RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- Paper: StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
- Paper: Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
-
[2026.07.23]
- Paper: Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection
- Paper: Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets
- Paper: AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning
- Paper: Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
- Paper: Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
- Paper: SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
- Paper: Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
- Paper: MemTools: A Unified Research Framework for Interoperable Agent Memory
- Paper: Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
- Paper: Anti-Periodic Positional Encoding: MΓΆbius Boundary Conditions Make In-Context Retrieval Reliable
-
[2026.07.22]
- Paper: ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
- Paper: SLPO: Scaling Latent Reasoning via a Surrogate Policy
- Paper: Self Gradient Forcing: Native Long Video Extrapolation
- Paper: JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
- Paper: PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
- Paper: ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
- Paper: PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
- Paper: EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
- Paper: Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
- Paper: ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
-
[2026.07.21]
- Paper: ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- Paper: FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
- Paper: Supra Cognitive Modes: A Routed Architecture for Agent Memory
- Paper: ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning
-
[2026.07.20]
- Paper: C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
- Paper: Is Progressive Disclosure All You Need for Long-Context Agents?
- Paper: AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
- Paper: Surprise Forcing: What to Remember, When to Skip in Long Video Generation
- Paper: ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
- Paper: FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
- Paper: How Agent Skills Fail under Long Contexts: A White-Box Study in Code Auditing
- Paper: SALT: Salience-Aware Lexical Trie for Long-Context Compression
- Paper: Mechanistic Attention Guidance for Agent Memory Refinement
- Paper: Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
-
[2026.07.19]
-
[2026.07.18]
- Paper: From Memory to Skills: Evidence-Grounded Co-Evolution Governance for Long-Horizon LLM Agents
- Paper: SpecLA: Efficient Speculative Decoding for Linear-Attention Models
- Paper: RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
- Paper: Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty
- Paper: Technical Report: AI-Assisted Gated DeltaNet Optimization on NVIDIA Blackwell
-
[2026.07.17]
- Paper: SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation
- Paper: FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- Paper: Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- Paper: Recursive Harness Self-Improvement
- Paper: ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
- Paper: DSWorld: A Data Science World Model for Efficient Autonomous Agents
- Paper: LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
- Paper: Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
- Paper: Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
- Paper: Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding
- Paper: Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
- Paper: SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation
-
[2026.07.16]
- Paper: VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
- Paper: LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
- Paper: Long-Context Fine-Tuning with Limited VRAM
- Paper: VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
- Paper: Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers
-
[2026.07.15]
- Paper: Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity
- Paper: Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
- Paper: PReM: Learning What to Preserve and When to Refresh for Context Compression
- Paper: CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
- Paper: TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
-
[2026.07.14]
- Paper: ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
- Paper: Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- Paper: MemoHarness: Agent Harnesses That Learn from Experience
- Paper: Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
- Paper: MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
- Paper: VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
- Paper: FOLIO: Focused Semantic Memory for Streaming Video Understanding
- Paper: A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs
- Paper: Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
- Paper: Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents
-
[2026.07.13]
- Paper: Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
- Paper: LightMem-Ego: Your AI Memory for Everyday Life
- Paper: ToFu: A White-Box, Token-Efficient Agent Harness for Researchers
- Paper: StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
- Paper: SLVMBench: Skill Learning from Video Memory
-
[2026.07.12]
-
[2026.07.11]
-
[2026.07.10]
-
[2026.07.09]
- Paper: OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
- Paper: Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- Paper: What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents
- Paper: Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
-
[2026.07.08]
- Paper: Infinite Worlds with Versatile Interactions
- Paper: The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
- Paper: Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
- Paper: Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
- Paper: Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- Paper: AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
-
[2026.07.07]
- Paper: AlayaWorld: Long-Horizon and Playable Video World Generation
- Paper: Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure
- Paper: TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
- Paper: DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
- Paper: FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
- Paper: AlayaWorld: Long-Horizon and Playable Video World Generation
-
[2026.07.06]
- Paper: Multiplayer Interactive World Models with Representation Autoencoders
- Paper: Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
- Paper: CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents
- Paper: KVpop -- Key-Value Cache Compression with Predictive Online Pruning
- Paper: Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
-
[2026.07.04]
-
[2026.07.03]
-
[2026.07.02]
-
[2026.07.01]
- Paper: MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
- Paper: MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression
- Paper: Imprint: Online Memory Compression for Long-Horizon Egocentric QA
- Paper: Self-GC: Self-Governing Context for Long-Horizon LLM Agents
- Paper: HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching
- Paper: CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
- Paper: QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
- Paper: The risk of KV cache compression
- Paper: Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
- Paper: MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Paper entries live under papers/ so this README stays under GitHub's homepage size limit.
For an interactive chapter reader (search + in-page paper cards), open the
project homepage.
Attention, recurrence & systems
Training, position & memory
Compression, reasoning & multimodal
Evaluation & reports
Please contact me if I miss your names in the list, I will add you back ASAP!