Planning with the Views |
 |
2026-05 |
Github |
 Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration? |
 |
2026-02 |
Github |
EscherVerse: An Open World Benchmark and Dataset for Teleo-Spatial Intelligence with Physical-Dynamic and Intent-Driven Understanding |
 |
2026-1 |
Github |
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics |
 |
2025-12 |
Github |
Towards Cross-View Point Correspondence in Vision-Language Models |
 |
2025-12 |
Github |
 ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints |
 |
2025-11 |
- |
Scaling Spatial Intelligence with Multimodal Foundation Models |
 |
2025-11 |
Github |
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition |
 |
2025-11 |
Github |
Visual Spatial Tuning |
 |
2025-11 |
Github |
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models |
 |
2025-10 |
Github |
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence |
 |
2025-10 |
Github |
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes |
 |
2025-10 |
- |
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions |
 |
2025-10 |
Github |
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs |
 |
2025-09 |
Github |
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes |
 |
2025-09 |
Github |
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture |
 |
2025-09 |
Github |
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning |
 |
2025-08 |
Github |
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations |
 |
2025-09 |
Github |
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models |
 |
2025-08 |
Github |
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis |
 |
2025-08 |
- |
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting |
 |
2025-07 |
Github |
Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models |
 |
2025-07 |
- |
SpatialViz-Bench: An MLLM Benchmark for Spatial Visualization |
 |
2025-07 |
Github |
Spatial Mental Modeling from Limited Views |
 |
2025-06 |
Github |
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks |
 |
2025-06 |
- |
 IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering |
 |
2025-06 |
Github |
 From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes |
 |
2025-06 |
Github |
 PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly |
 |
2025-06 |
Github |
Can Vision Language Models Infer Human Gaze Direction? A Controlled Study |
 |
2025-06 |
Github |
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence |
 |
2025-06 |
Github |
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations |
 |
2025-06 |
Github |
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models |
 |
2025-06 |
Github |
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models |
 |
2025-06 |
- |
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence |
 |
2025-05 |
Github |
 RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics |
 |
2025-05 |
Github |
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models |
 |
2025-05 |
Github |
SpatialScore: Towards Unified Evaluation for Multimodal Spatial Understanding |
 |
2025-05 |
Github |
MIRAGE:A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence |
 |
2025-05 |
Github |
Can Multimodal Large Language Models Understand Spatial Relations |
 |
2025-05 |
Github |
Visuospatial Cognitive Assistant |
 |
2025-05 |
Github |
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning? |
 |
2025-05 |
Github |
Vision language models have difficulty recognizing virtual objects |
 |
2025-05 |
- |
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models |
 |
2025-05 |
Github |
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames |
 |
2025-05 |
- |
 SITE: towards Spatial Intelligence Thorough Evaluation |
 |
2025-05 |
Github |
CameraBench: Towards Understanding Camera Motions in Any Video |
 |
2025-04 |
Github |
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs |
 |
2025-04 |
Github |
From Flatland to Space:Teaching Vision-Language Models to Perceive and Reason in 3D |
 |
2025-03 |
Github |
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLM |
 |
2025-03 |
- |
Open3DVQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space |
 |
2025-03 |
Github |
 STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding? |
 |
2025-03 |
Github |
 CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models |
 |
2025-03 |
Github |
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models |
 |
2025-03 |
Github |
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? |
 |
2025-03 |
Github |
 Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models |
 |
2025-02 |
Github |
FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks |
 |
2025-02 |
- |
iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs |
 |
2025-02 |
Github |
Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics |
 |
2025-02 |
- |
 SAT: Spatial Aptitude Training for Multimodal Language Models |
 |
2024-12 |
Github |
 SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models |
 |
2024-12 |
Github |
 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark |
 |
2024-12 |
Github |
  Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces |
 |
2024-12 |
Github |
 RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics |
 |
2024-11 |
Github |
 An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models |
 |
2024-11 |
- |
 IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos |
 |
2024-11 |
Github |
Is ‘Right’ Right? Enhancing Object Orientation Understanding in Multimodal Language Models through Egocentric Instruction Tuning |
 |
2024-10 |
Github |
 ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models |
 |
2024-10 |
Github |
 DOES SPATIAL COGNITION EMERGE IN FRONTIER MODELS? |
 |
2024-10 |
- |
 Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities |
 |
2024-10 |
Github |
R2D3: Imparting Spatial Reasoning by Reconstructing 3D Scenes from 2D Images |
 |
2024-10 |
Github |
 Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models |
 |
2024-09 |
Github |
 Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning? |
 |
2024-09 |
- |
VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs |
 |
2024-07 |
Github |
 EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models |
 |
2024-06 |
Github |
 TopViewRS: Vision-Language Models as Top-View Spatial Reasoners |
 |
2024-06 |
Github |
 Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models |
 |
2024-06 |
Github |
 GSR-Bench: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs |
 |
2024-06 |
- |
 Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning |
 |
2024-05 |
Github |
 Visually Descriptive Language Model for Vector Graphics Reasoning |
 |
2024-04 |
- |
  SQA3D: Situated Question Answering in 3D Scenes |
 |
2022-10 |
Github |
  Things not Written in Text: Exploring Spatial Commonsense from Visual Signals |
 |
2022-03 |
Github |
 SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings |
 |
2020-03 |
Github |