Back to papers
March 19, 2026cs.CVcs.AIIntermediate

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

AI-Generated Summary

This paper introduces MultihopSpatial, a new benchmark to test how well AI vision-language models can understand complex spatial relationships in images, such as 'the object to the left of the object above the red box.' The benchmark includes a new evaluation metric (Acc@50IoU) that checks both whether the model answers correctly and whether it can precisely pinpoint the location in the image, which is crucial for robots performing real-world tasks. Testing 37 state-of-the-art models reveals that spatial reasoning remains challenging, but training models on their new dataset improves both the models' spatial understanding and their ability to perform physical manipulation tasks.

Difficulty
Intermediate
Categories

cs.CV, cs.AI

AI Tags
vision-language modelsspatial reasoningbenchmarkvisual groundingembodied AIVLA agentscompositional reasoningreinforcement learning