When Vision-Language Models Look but Don't See: Anatomical Bias in Endoscopic Spatial Reasoning
Abstract
Reliable spatial understanding is essential for complete mucosal inspection during upper gastrointestinal endoscopy, particularly under the Systematic Screening Protocol for the Stomach (SSS). Vision-Language Models (VLMs) show promise for assisting navigation but often rely on textual priors rather than visual information, a limitation known as anatomical bias. To evaluate spatial reasoning, we introduce EndoSSS-RP, a benchmark based on the GastroHUN dataset (380 patients, 2,796 images, 3,678 binary questions). Each question, targeting left/right or above/below relationships between gastric surfaces (e.g., anterior vs. posterior wall), is tested on original, flipped, and rotated images to assess model robustness in determining relative position across different orientations. Prompts are structured across three levels: L1 (anatomical terms only), L2 (anatomical terms with visual markers), and L3 (visual markers only). We evaluate four VLMs (GPT-4o, Gemini-2.5-Flash, JanusPro-7B, and LLaMA-3.2) and find that accuracy declines under transformations when using anatomical prompts (L1, GPT-4o: 68% original, 61% flipped, 46% rotated) but perform consistently better with visual–marker only prompts (L3: 87%). These results reveal anatomical bias and underscore the need for models to rely more heavily on visual evidence in endoscopic AI systems.
BibTeX
@inproceedings{bravo2026when,
title={When Vision-Language Models Look but Don't See: Anatomical Bias in Endoscopic Spatial Reasoning},
author={Bravo, Diego and Wolf, Daniel and Hurtado-Tobar, Juan and Gómez, Mart{\'i}n and Romero, Eduardo},
booktitle={Proceedings of IEEE International Symposium on Biomedical Imaging (ISBI)}
year={2026},
pages={1--5}
}