文章摘要
Shi Wentao (施文涛)*,Tong Jianjun**,Ma Jie**,Zhang Rui***.[J].高技术通讯(英文),2026,32(3):308~321
MVQAM: exploring state space models for medical visual question answer Mamba
  
DOI:10. 3772 / j. issn. 1006-6748. 2026. 03. 010
中文关键词: 
英文关键词: medical visual question answer, multimodal fusion, state space model, SLAKE
基金项目:
Author NameAffiliation
Shi Wentao (施文涛)* (*School of Big Data, Fuzhou University of International Studies and Trade, Fuzhou 350202, P. R. Chin) (**Key Laboratory of Data Science and Intelligent Computing, Fuzhou University of International Studies and Trade, Fuzhou 350202, P. R. China) (***Engineering Research Center of Big Data Applications in Modern Service Industries,Fujian Province University, Fuzhou 350202, P. R. China) 
Tong Jianjun**  
Ma Jie**  
Zhang Rui***  
Hits: 62
Download times: 74
中文摘要:
      
英文摘要:
      Medical visual question answering (MedVQA) is a complex multimodal task. In this domain , multimodal fusion models that leverage the attention mechanism and transformer architecture have been extensively utilized. However , attention-based fusion models face limitations due to constraints in training datasets and global receptive fields , which hinder their ability to incorporate additional information effectively. Furthermore , the self-attention mechanism inherent in Transformer models presents challenges for real-time applications because of its high computational cost resulting from quadratic complexity. Recent studies indicate that state space models (SSM) , exemplified by Mam- ba , can effectively model long-range interactions while preserving linear computational complexity. To address the limitations mentioned above , this paper proposes a novel MedVQA framework based on state-space modeling , termed MVQAM. As the core component of MVQAM , the Mamba vision correlation connector (MVCC) module adopts a dual-branch hybrid structure that synergizes convo- lutional neural networks (CNN) and Mamba. Specifically , CNN branch extracts local fine-grained features from medical images through stacked convolutional layers , while the Mamba branch captures long-range contextual dependencies in the feature map. The two branches ’ outputs are aggregated via linear projection to generate comprehensive visual feature representations. This integration not only enhances the efficiency of feature extraction but also reduces the overall computational complex- ity of the framework by leveraging Mamba ’ s linear complexity advantage. Experimental results dem- onstrate that MVQAM surpasses existing multimodal fusion models on both SLAKE-EN and visual question answering in radiology ( VQA-RAD ) datasets in terms of accuracy , convergence speed , and parameter efficiency. These findings emphasize the significant potential of SSM in advancing MedVQA tasks.
View Full Text   View/Add Comment  Download reader
Close