| Shi Wentao (施文涛)*,Tong Jianjun**,Ma Jie**,Zhang Rui***.[J].高技术通讯(英文),2026,32(3):308~321 |
|
| MVQAM: exploring state space models for medical visual question answer Mamba |
| |
| DOI:10. 3772 / j. issn. 1006-6748. 2026. 03. 010 |
| 中文关键词: |
| 英文关键词: medical visual question answer, multimodal fusion, state space model, SLAKE |
| 基金项目: |
| Author Name | Affiliation | | Shi Wentao (施文涛)* | (*School of Big Data, Fuzhou University of International Studies and Trade, Fuzhou 350202, P. R. Chin)
(**Key Laboratory of Data Science and Intelligent Computing, Fuzhou University of International Studies and Trade, Fuzhou 350202, P. R. China)
(***Engineering Research Center of Big Data Applications in Modern Service Industries,Fujian Province University, Fuzhou 350202, P. R. China) | | Tong Jianjun** | | | Ma Jie** | | | Zhang Rui*** | |
|
| Hits: 62 |
| Download times: 74 |
| 中文摘要: |
| |
| 英文摘要: |
| Medical visual question answering (MedVQA) is a complex multimodal task. In this domain ,
multimodal fusion models that leverage the attention mechanism and transformer architecture have
been extensively utilized. However , attention-based fusion models face limitations due to constraints
in training datasets and global receptive fields , which hinder their ability to incorporate additional
information effectively. Furthermore , the self-attention mechanism inherent in Transformer models
presents challenges for real-time applications because of its high computational cost resulting from
quadratic complexity. Recent studies indicate that state space models (SSM) , exemplified by Mam-
ba , can effectively model long-range interactions while preserving linear computational complexity.
To address the limitations mentioned above , this paper proposes a novel MedVQA framework based
on state-space modeling , termed MVQAM. As the core component of MVQAM , the Mamba vision
correlation connector (MVCC) module adopts a dual-branch hybrid structure that synergizes convo-
lutional neural networks (CNN) and Mamba. Specifically , CNN branch extracts local fine-grained
features from medical images through stacked convolutional layers , while the Mamba branch captures
long-range contextual dependencies in the feature map. The two branches ’ outputs are aggregated
via linear projection to generate comprehensive visual feature representations. This integration not
only enhances the efficiency of feature extraction but also reduces the overall computational complex-
ity of the framework by leveraging Mamba ’ s linear complexity advantage. Experimental results dem-
onstrate that MVQAM surpasses existing multimodal fusion models on both SLAKE-EN and visual
question answering in radiology ( VQA-RAD ) datasets in terms of accuracy , convergence speed ,
and parameter efficiency. These findings emphasize the significant potential of SSM in advancing
MedVQA tasks. |
|
View Full Text
View/Add Comment Download reader |
| Close |
|
|
|