CF202646159
Multimodal Representation Learning for Breast Cancer Decision Support from Imaging and Clinical Data
D-6
Doctorate Full Doctorate
Disciplines
Other (Maths)
Laboratory
UMR 5141 Laboratoire de Traitement et Communication de l'Information
Host institution
Télécom Paris, Institut Polytechnique de Paris Télécom Paris

Description

Clinical decision-making in oncology increasingly relies on the integration of heterogeneous
data sources, particularly medical imaging and structured clinical data [1]. Breast cancer, for
instance, is routinely assessed using imaging modalities such as mammography, ultrasound,
or MRI, alongside clinical variables including patient demographics, comorbidities, and
treatment history [2]. While deep learning has significantly advanced image-based analysis,
structured clinical data---typically stored in tabular form---remains underutilised in multimodal
medical AI, despite its critical role in real-world decision-making.
Recent research has predominantly focused on vision-language models that combine image
encoders with Large Language Models (LLMs). However, these architectures are often
ill-suited to clinical tabular data, which contains continuous variables, missing values,
and ordinal or categorical structure not naturally handled by LLMs [3]. Moreover,
state-of-the-art methods are frequently benchmarked on large, curated datasets and fail to
generalise to real-world settings characterised by small sample sizes, heterogeneous
patient cohorts, incomplete follow-ups, and irregular annotations [4-5].
In oncology applications, these limitations are especially problematic. Many prediction
tasks---such as BI-RADS assessment, tumour grading, or risk stratification---involve ordinal
labels and require nuanced reasoning across both imaging and clinical variables. However,
current multimodal models often overlook the semantics of clinical tabular data, reducing
its contribution to naïve concatenation or late fusion, and ignoring how ordinal and structured
features could guide representation learning.
This motivates the need for new multimodal self-supervised learning (SSL) approaches
that can robustly combine imaging and structured clinical data, while accounting for missing
values, semantic structures, and population biases. Such representations must support
clinical transferability, adapt to small or incomplete datasets, and provide clinically
grounded outputs across diverse patient groups and institutions.

Skills required

Master's degree (or equivalent) in computer science, applied mathematics, biomedical engineering, or a related field Strong interest in medical imaging and healthcare applications Practical experience with deep learning frameworks (e.g. PyTorch, TensorFlow) Solid programming skills (preferably in Python) Familiarity with machine learning, computer vision, or multimodal data processing is a plus Good written and spoken communication skills in English

Bibliography

[1] Benjamin D Simon et al.
“The future of multimodal artificial intelligence models for
integrating imaging and clinical metadata: a narrative review”
. In: Diagnostic and
Interventional Radiology 31.4 (2025), p. 303.
[2] Clayton R Taylor et al.
“Artificial intelligence applications in breast imaging: current status
and future directions”
. In: Diagnostics 13.12 (2023), p. 2041.
[3] Xi Fang et al.
“Large Language Models (LLMs) on Tabular Data: Prediction, Generation,
and Understanding–A Survey”
. In: arXiv preprint arXiv:2402.17944 (2024).
[4] Paul Hager, Martin J Menten, and Daniel Rueckert.
“Best of both worlds: Multimodal
contrastive learning with tabular and imaging data”
. In: Proceedings of the IEEE/CVF
Conference on Computer Vision and Pattern Recognition. 2023, pp. 23924–23935.
[5] Marta Hasny et al.
“TGV: Tabular Data-Guided Learning of Visual Cardiac
Representations”
. In: arXiv preprint arXiv:2503.14998 (2025)

Keywords

Representation Learning, Medical Imaging, Multimodal, Self-supervised Learning, Foundation models

Grant holder offer / non-funded

Open to all countries

Dates

Application deadline 31/08/26

Duration36 months

Start date01/10/26

Creation date14/02/26

Languages

Level of french requiredNone

Level of English requiredC1 (advanced)

Miscellaneous

Annual tuition fee400 € / year

Contacts

You must connect to be able to display the contacts.

click here to connect or register (it's free!)