CF202646853
Multimodal Medical Vision-Language Foundation Model for Healthcare Reasoning
D-36
Doctorate Full Doctorate
Disciplines
Other (Maths)
Laboratory
UMR 5141 Laboratoire de Traitement et Communication de l'Information
Host institution
Institut Polytechnique de Paris Télécom Paris

Description

This PhD project aims to construct a large-scale, longitudinal, multimodal dataset enriched with strong grounding signals and to develop a compact-to-scalable medical vision-language model (VLM) whose internal structure aligns closely with physician workflows.
The research will be organized around two tightly coupled thrusts. The first focuses on dataset construction, involving the collection and harmonization of de-identified Vietnamese hospital data across X-ray, CT, PET, MRI, and clinical reports, complemented by carefully curated public datasets. The second focuses on methodology, starting from clinically competitive, moderate-size backbone models in the spirit of LLaVA-Med, and decomposing the system into interactive expert modules for retrieval, localization, segmentation, quantification, masking, gating, verification, and generation.
The expected outcome is a clinically grounded research framework capable of supporting report generation, medical visual question answering (VQA), localization, interpretation, and decision support. Crucially, this framework provides a realistic pathway from compact, domain-specific modeling toward larger multimodal healthcare reasoning systems, ensuring both practical applicability and clinical relevance throughout the course of the PhD.

Skills required

- Master's degree (or equivalent) in computer science (machine learning, artificial intelligence), or related fields - Strong background in computer science, applied mathematics, and statistics, with an emphasis on machine learning (esp. deep learning) - Proficient programming skills, preferably in Python - Practical experience with machine learning/deep learning frameworks (e.g., PyTorch) - Familiarity with working/analysing healthcare data - Advanced proficiency in English: The candidate should be fluent in spoken and written English

Bibliography

[1] Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tris- tan Naumann, Hoifung Poon, and Jianfeng Gao. “Llava-med: Training a large language- and-vision assistant for biomedicine in one day”. In: Advances in Neural Information Processing Systems 36 (2023), pp. 28541–28564.
[2] Lin Yang, Yunhe Wang, Qi Zhao, Chaitanya D. Kulkarni, Tao Tu, Shekoofeh Azizi, Vera Rabinovich, Yossi Matias, Greg Corrado, Alan Karthikesalingam, et al. “Advancing Mul- timodal Medical Capabilities of Gemini”. In: arXiv preprint arXiv:2405.03162 (2024).
[3] Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, et al. “Lingshu: A gener- alist foundation model for unified multimodal medical understanding and reasoning”. In: arXiv preprint arXiv:2506.07044 (2025).
[4] Jiarui Ye and Hao Tang. “Multimodal Large Language Models for Medicine: A Compre- hensive Survey”. In: arXiv preprint arXiv:2504.21051 (2025).
[5] Armin Berger, Sarthak Khanna, David Berghaus, and Rafet Sifa. “Reasoning LLMs in the Medical Domain: A Literature Survey”. In: arXiv preprint arXiv:2508.19097 (2025).
[6] Huu Tien Nguyen, Dac Thai Nguyen, Trung Thanh Nguyen, Thao Nguyen Truong, Hieu Pham, Johan Barthelemy, Tran Minh Quan, Quoc Viet Hung Nguyen, Thanh Tam Nguyen, and Mai Hong Son. “Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Genera- tion”. In: The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. 2025.
[7] Alistair E. W. Johnson, Tom J. Pollard, Seth J. Berkowitz, Nathan R. Greenbaum, Matthew P. Lungren, Chih-Ying Deng, Roger G. Mark, and Steven Horng. “MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports”. In: Scientific Data 6 (2019), p. 317.
[8] Ha Q. Nguyen, Hieu T. Nguyen, Hieu Pham, Khanh Lam, Linh T. Le, Minh Dao, Vu A. D. Nguyen, Dung T. Nguyen, Cao K. Nguyen, Quang D. Ho, Dinh D. Do, Chinh Q. Dinh, Son T. Nguyen, Thanh T. Nguyen, Duc T. Nguyen, et al. “VinDr-CXR: An open dataset of chest X-rays with radiologist’s annotations”. In: Scientific Data 9 (2022), p. 429.
[9] Ibrahim Ethem Hamamci, Sezgin Er, Chenyu Wang, Furkan Almas, Ayse Gulnihan Sim- sek, Sevval Nil Esirgun, Irem Dogan, Omer Faruk Durugol, Benjamin Hou, Suprosanna Shit, et al. “Generalist foundation models from a multimodal dataset for 3D computed tomography”. In: Nature Biomedical Engineering (2026), pp. 1–19.
[10] Ruiyang Zhao, Burhaneddin Yaman, Yuxin Zhang, Russell Stewart, Austin Dixon, Florian Knoll, Zhengnan Huang, Yvonne W Lui, Michael S Hansen, and Matthew P Lungren. “fastMRI+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data”. In: Scientific Data 9.1 (2022), p. 152.
[11] Sergios Gatidis and Thomas Kuestner. “A whole-body FDG-PET/CT dataset with man- ually annotated tumor lesions”. In: Scientific Data 9 (2022), p. 601.
[12] Khizar Anjuma, Muhammad Arbab Arshad, Kadhim Hayawi, Efstathios Polyzos, Asadul- lah Tariq, Mohamed Adel Serhani, Laiba Batool, Brady Lund, Nishith Reddy Mannuru, Ravi Varma Kumar Bevara, Taslim Mahbub, Muhammad Zeeshan Akram, and Sakib Shahriar. “Domain Specific Benchmarks for Evaluating Multimodal Large Language Mod- els”. In: arXiv preprint arXiv:2506.12958 (2025).
[13] Jiazhen Pan, Che Liu, Junde Wu, Fenglin Liu, Jiayuan Zhu, Hongwei Bran Li, Chen Chen, Cheng Ouyang, and Daniel Rueckert. “MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning”. In: arXiv preprint arXiv:2502.19634 (2025).
[14] Yue Wang et al. “A Survey of Multimodal Hallucination Evaluation and Detection”. In: arXiv preprint arXiv:2507.19024 (2025).
[15] Yun-Wei Chu, Kai Zhang, Christopher Malon, and Martin Renqiang Min. “Reducing Hallucinations of Medical Multimodal Large Language Models with Visual Retrieval- Augmented Generation”. In: arXiv preprint arXiv:2502.15040 (2025).
[16] Sebastian Wind et al. “Multi-step Retrieval and Reasoning Improves Radiology Report Structuring and Accuracy”. In: arXiv preprint arXiv:2508.00743 (2025).
[17] Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. “Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks”. In: Advances in Neural Informa- tion Processing Systems. 2015.
[18] Sidi Lu, Yaoming Zhu, Weinan Zhang, Jun Wang, and Yong Yu. “Neural Text Generation: Past, Present and Beyond”. In: arXiv preprint arXiv:1803.07133 (2018).
[19] Pengyu Wang, Shuchang Ye, Usman Naseem, and Jinman Kim. “MRG-R1: Reinforce- ment Learning for Clinically Aligned Medical Report Generation”. In: arXiv preprint arXiv:2512.16145 (2025).
[20] Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng, Quanli Shen, Xiaobo Zhang, Junjun He, and Shujun Wang. “AOR: Anatomical Ontology- Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation”. In: arXiv preprint arXiv:2505.02830 (2025).
[21] Jiayu Lei, Xiaoman Zhang, Chaoyi Wu, Lisong Dai, Ya Zhang, Yanyong Zhang, Yanfeng Wang, Weidi Xie, and Yuehua Li. “AutoRG-Brain: Grounded Report Generation for Brain MRI”. In: arXiv preprint arXiv:2407.16684 (2024).
[22] Ziyue Wang, Junde Wu, Chang Han Low, and Yueming Jin. “MedAgent-Pro: Towards Multi-modal Evidence-based Medical Diagnosis via Reasoning Agentic Workflow”. In: arXiv preprint arXiv:2503.18968 (2025).
[23] Yannian Gu, Wenhui Lei, Hanyu Chen, Shaoting Zhang, and Xiaofan Zhang. “Interactive Segmentation and Report Generation for CT Images”. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer. 2025, pp. 273– 283.
[24] Sanjay Subramanian, Lucy Lu Wang, Ben Bogin, Sachin Mehta, Madeleine Van Zuylen, Sravanthi Parasa, Sameer Singh, Matt Gardner, and Hannaneh Hajishirzi. “Medicat: A dataset of medical images, captions, and textual references”. In: Findings of the Associa- tion for Computational Linguistics: EMNLP 2020. 2020, pp. 2112–2120.
[25] Alejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen, Jeffrey J. Nirschl, Jeffrey Gu, Ivan Lopez, Josiah Aklilu, Austin Wolfgang Katzer, Collin Chiu, Anita Rau, Xiao- han Wang, Yuhui Zhang, Alfred Seunghoon Song, Robert Tibshirani, and Serena Yeung- Levy. “BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision- Language Models Derived from Scientific Literature”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025.

Keywords

Foundation Model , Multimodal , Vision-Language, Reasoning, Healthcare

Funded offer

Countries

Mexico (Conacyt)

China (CSC)

Dates

Application deadline 30/09/26

Duration36 months

Start date01/10/26

Creation date17/03/26

Languages

Level of french requiredNone

Level of English requiredNone

Miscellaneous

Annual tuition fee400 € / year

Contacts

You must connect to be able to display the contacts.

click here to connect or register (it's free!)