CT-Bench: A Comprehensive Benchmark for Multimodal AI in Computed Tomography Analysis

Qingqing Zhu1, Qiao Jin1, Tejas S. Mathai1, Yin Fang1, Zhizheng Wang1, Yifan Yang1, Maame Sarfo-Gyamfi1,2, Benjamin Hou1, Ran Gu1, Praveen T. S. Balamuralikrishna1, Kenneth C. Wang3, Ronald M. Summers1, Zhiyong Lu1
1: National Institutes of Health, Bethesda, MD, USA, 2: Howard University College of Medicine, Washington, D.C., USA, 3: Radiology, U.S. Department of Veterans Affairs, USA
Publication date: 2026/09/21
https://doi.org/10.59275/j.melba.2026-b141
PDF · Dataset

Abstract

Publicly available computed tomography (CT) resources with lesion-level image, spatial, and textual annotations remain limited, restricting reproducible development and evaluation of multimodal medical AI. We present CT-Bench, a CT lesion image–text resource and visual question answering benchmark constructed by extending DeepLesion with curated lesion-level text, standardized lesion-size representations, structured attributes, and QA tasks. CT-Bench contains 20,335 lesion instances from 7,795 CT studies and 3,793 patients, including CT key-slice images, bounding-box annotations, PACS-derived lesion descriptions, lesion size measurements, and associated metadata. The resource also includes a multiple-choice QA benchmark with 2,850 question-answer pairs across seven lesion analysis tasks: image-to-description matching, multi-slice CT-to-description matching, description-to-image retrieval, description-to-bounding-box localization, lesion size estimation, image-based attribute recognition, and multi-slice CT attribute recognition. Hard negative answer choices are included to evaluate fine-grained image-text alignment and lesion-level reasoning. The resource is available at https://kaggle.com/datasets/cd1661d6d6aeab08b8eb99b58885b4489d76fc5ac07d5aac76cae577e6426e2f with documentation, metadata files, and reproducibility code. Validation includes annotation review by trained annotators and medical experts, hard-negative verification, human-reader comparison, and baseline experiments with general and medical vision-language models.

Keywords

computed tomography · open data · medical image computing · multimodal learning · lesion analysis · visual question answering

Bibtex @article{melba:2026:036:zhu, title = "CT-Bench: A Comprehensive Benchmark for Multimodal AI in Computed Tomography Analysis", author = "Zhu, Qingqing and Jin, Qiao and Mathai, Tejas S. and Fang, Yin and Wang, Zhizheng and Yang, Yifan and Sarfo-Gyamfi, Maame and Hou, Benjamin and Gu, Ran and Balamuralikrishna, Praveen T. S. and Wang, Kenneth C. and Summers, Ronald M. and Lu, Zhiyong", journal = "Machine Learning for Biomedical Imaging", volume = "2026", issue = "Special Issue on MICCAI Open Data 2026", year = "2026", pages = "739--757", issn = "2766-905X", doi = "https://doi.org/10.59275/j.melba.2026-b141", url = "https://melba-journal.org/2026:036" }
RISTY - JOUR AU - Zhu, Qingqing AU - Jin, Qiao AU - Mathai, Tejas S. AU - Fang, Yin AU - Wang, Zhizheng AU - Yang, Yifan AU - Sarfo-Gyamfi, Maame AU - Hou, Benjamin AU - Gu, Ran AU - Balamuralikrishna, Praveen T. S. AU - Wang, Kenneth C. AU - Summers, Ronald M. AU - Lu, Zhiyong PY - 2026 TI - CT-Bench: A Comprehensive Benchmark for Multimodal AI in Computed Tomography Analysis T2 - Machine Learning for Biomedical Imaging VL - 2026 IS - Special Issue on MICCAI Open Data 2026 SP - 739 EP - 757 SN - 2766-905X DO - https://doi.org/10.59275/j.melba.2026-b141 UR - https://melba-journal.org/2026:036 ER -

2026:036 cover