Back to papers
March 19, 2026cs.CVcs.LGIntermediate

Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders

AI-Generated Summary

This paper investigates whether state space models (SSMs) can replace the standard Vision Transformers used in vision-language models (VLMs) that combine images with text. Through systematic testing, the researchers found that SSM-based vision encoders perform comparably to or better than Vision Transformers, especially when trained on detection or segmentation tasks, while using fewer parameters. The study challenges the assumption that larger models or higher image classification accuracy automatically lead to better VLM performance and proposes stabilization techniques to improve robustness.

HF Upvotes

3

Difficulty
Intermediate
Categories

cs.CV, cs.LG

AI Tags
vision-language modelsvision encodersstate space modelstransformersVQAimage understanding