基于多源知识跨模态对齐与融合的医学报告生成方法
Medical Report Generation Method Based on Multi-source Knowledge Cross-modal Alignment and Fusion
-
摘要: 针对医学影像和报告在数据内容和数据分布2个层面存在的数据偏差问题, 提出一种基于多源知识跨模态对齐与融合的医学报告生成方法(medical report generation method based on multi-source knowledge cross-modal alignment and fusion, MRGM-MKCAF)。首先, 将放射学相关的医学概念作为先验知识, 并将案例影像和报告作为经验知识, 构建多模态知识库; 然后, 提出一种基于多头注意力机制的跨模态对齐方法, 利用多个注意力头分别关注知识库中多种模态特征的子空间, 并进行跨模态细粒度特征对齐, 获得更加全面准确的编码特征; 最后, 提出一种基于门控机制的解码模块, 对多源编码特征进行动态的特征选取, 去除冗余和无关信息, 进而用于报告生成。在2组数据集上的实验结果表明, 提出的方法显著提高了生成报告的准确性、流畅性和完整性。Abstract: To address data bias in medical imaging and reporting at both the data content and distribution, a medical report generation method based on multi-source knowledge cross-modal alignment and fusion (MRGM-MKCAF) is proposed. First, radiology-related medical concepts were used as prior knowledge, and case images and reports were used as empirical knowledge to construct a multi-modal knowledge base. Subsequently, a cross-modal alignment method employing multi-head attention mechanism was proposed, which used multiple attention heads to focus on the subspaces of various modal features in the knowledge base and performed fine-grained cross-modal feature alignment to obtain more comprehensive and accurate encoded features. Finally, a decoding module based on gating mechanism was proposed to dynamically select features from multi-source encoded features, removing redundancy and irrelevant information, which was then used for report generation. Experimental results on two datasets show that the proposed method significantly enhances the accuracy, fluency, and completeness of generated reports.
下载: