From Classification to Cross-Modal Understanding: Leveraging Vision- Language Models for Fine-Grained Renal Pathology

Zhenhao Guo1, Rachit Saluja2, Tianyu Shi3, Tianyuan Yao4, Quan Liu4, Junchao Zhu4, Haibo Wang5, Daniel Reisenbüchler6, Haichun Yang7, Yuankai Huo4, Benjamin Liechty8, David J. Pisapia8, Kenji Ikemura8, Steven Salvatoree8, Surya Seshane8, Mert R. Sabuncu2,8, Yihe Yang8,9, Ruining Deng4,8
1: New York University, New York, NY 10012, USA, 2: Cornell Tech, New York, NY 10044, USA, 3: Sichuan University, Chengdu, 610207, CN, 4: Vanderbilt University, Nashville, TN, 37235, USA, 5: Carnegie Mellon University, Pittsburgh, PA 15213, USA, 6: University of Regensburg, Regensburg, Bavaria 93053, DE, 7: Vanderbilt University Medical Center, Nashville, TN, 37232, USA, 8: Weill Cornell Medicine, New York, NY 10065, USA, 9: Northwell Health, New Hyde Park, NY 11040, USA
Publication date: 2026/10/09
https://doi.org/10.59275/j.melba.2026-3893
PDF · Code

Abstract

Fine-grained glomerular subtyping is essential for kidney biopsy interpretation, yet it presents a harder setting than the broad organ- or disease-level recognition tasks commonly used in vision-language model (VLM) evaluation. In this task, all categories arise within the same renal microanatomic unit, and clinically meaningful subtypes often differ by subtle, overlapping, and continuum-like morphological cues. Clinically valuable labels are also scarce, making fully supervised image-only paradigms difficult to translate to practice. Although VLMs are promising for data-efficient pathology learning, it remains unclear how their cross-modal representations should be adapted and evaluated for such within-structure discrimination under realistic few-shot constraints. To our knowledge, this study establishes the first fine-grained renal pathology benchmark for few-shot VLM adaptation and systematically examines how supervision level, model design, and adaptation strategy shape both classification performance and multimodal representation structure. Beyond accuracy, area under the receiver operating characteristic curve (AUC), F1 score, image-text alignment, positive-negative similarity separation, embedding geometry, and confidence calibration quantified by expected calibration error (ECE), we ask when performance gains reflect more reliable cross-modal understanding. Across renal pathology cohorts, additional supervision consistently improves discrimination, but reliable adaptation is not explained by alignment alone; stronger performance is associated with larger class separation, more coherent embedding organization, and better awareness of confidence-related failure modes under domain variation. Within this retrospective two-cohort evaluation, these findings move assessment beyond coarse classification and accuracy alone and provide comparative evidence to inform model selection and future translational validation. Our implementation has now been publicly released for reproducibility on GitHub at https://github.com/ddrrnn123/glo-vlms

Keywords

Vision-Language Model · Fine-Grained Classification · Fine-tuning · Few-shot Learning · Digital Renal Pathology

Bibtex @article{melba:2026:052:guo, title = "From Classification to Cross-Modal Understanding: Leveraging Vision- Language Models for Fine-Grained Renal Pathology", author = "Guo, Zhenhao and Saluja, Rachit and Shi, Tianyu and Yao, Tianyuan and Liu, Quan and Zhu, Junchao and Wang, Haibo and Reisenbüchler, Daniel and Yang, Haichun and Huo, Yuankai and Liechty, Benjamin and Pisapia, David J. and Ikemura, Kenji and Salvatoree, Steven and Seshane, Surya and Sabuncu, Mert R. and Yang, Yihe and Deng, Ruining", journal = "Machine Learning for Biomedical Imaging", volume = "2026", issue = "October 2026 issue", year = "2026", pages = "903--932", issn = "2766-905X", doi = "https://doi.org/10.59275/j.melba.2026-3893", url = "https://melba-journal.org/2026:052" }
RISTY - JOUR AU - Guo, Zhenhao AU - Saluja, Rachit AU - Shi, Tianyu AU - Yao, Tianyuan AU - Liu, Quan AU - Zhu, Junchao AU - Wang, Haibo AU - Reisenbüchler, Daniel AU - Yang, Haichun AU - Huo, Yuankai AU - Liechty, Benjamin AU - Pisapia, David J. AU - Ikemura, Kenji AU - Salvatoree, Steven AU - Seshane, Surya AU - Sabuncu, Mert R. AU - Yang, Yihe AU - Deng, Ruining PY - 2026 TI - From Classification to Cross-Modal Understanding: Leveraging Vision- Language Models for Fine-Grained Renal Pathology T2 - Machine Learning for Biomedical Imaging VL - 2026 IS - October 2026 issue SP - 903 EP - 932 SN - 2766-905X DO - https://doi.org/10.59275/j.melba.2026-3893 UR - https://melba-journal.org/2026:052 ER -

2026:052 cover