CF202645816
From text-lake to knowledge graph: frugal algorithms for mapping knowledge (traceability, evidence, diachrony)
D-6
Doctorate Full Doctorate
Disciplines
Other (Computer Sciences)
Laboratory
LORRAINE LABORATORY OF RESEARCH ON INFORMATICS AND ITS APPLICATIONS (LORIA)
Host institution
Université de Lorraine
Doctoral school
Informatics - Automatics - Electronics - Electrotechnics - Mathematics in Lorraine (IAEM-Lorraine) - ED 77

Description

Context:
Scientific corpora as well as public-debate corpora now produce a paradoxical effect: information is abundant, but gaining an overall understanding is increasingly difficult. Important statements circulate under multiple formulations, evolve over time, change speaker, and rely on evidence that can be contradictory. Search engines retrieve documents without making this dynamic visible; automatic summaries remain local; and traditional knowledge graphs often assume data that is already structured.

What does it mean, concretely?
The goal is to build a navigable 'map' of a domain (e.g., a scientific field or a media corpus) or a debate: identify key ideas/claims in a large collection of texts, group formulations that essentially express the same claim, link each claim to its sources (and, when possible, to the cited evidence), and track how these claims evolve over time (emergence, rephrasing, controversy, consensus).

Scientific objective;
Develop and evaluate new algorithms to transform massive, unstructured text-lakes into coherent, auditable, and extensible knowledge graphs. The thesis will focus in particular on disambiguation/canonicalization strategies (detecting when two formulations refer to the same claim) under cost constraints (frugal approaches), and on how a flexible ontology can (i) enable better representations while (ii) reducing dimensionality and (iii) guiding graph construction (ideas, actors, sources, evidence, temporal relations).


Skills required

Candidate profile • Strong command of Python (or equivalent programming tools). • Interest in linguistics, NLP, information extraction, semantic representations and/or graphs (Neo4j, RDF, etc. appreciated). • Master's degree (or equivalent) in computer science / data science / AI (or related fields). • Enjoys data processing, rigorous evaluation (benchmarks, ablations), and scientific writing.

Bibliography

• Lamirel, J.-C. (2012). A new approach for automatizing the analysis of research topics dynamics: application to optoelectronics research. Scientometrics, vol. 93(1), pages 151-166.
• Lamirel, J.-C. et al. (2014). Federating clustering and labeling capabilities based on feature maximization. Neurocomputing, 147, 136-146.
• Lamirel, J.-C. et al. (2020). An overview of the history of Science of Science in China based on bibliographic and citation data: a new method based on clustering with feature maximization and contrast graphs. Scientometrics.
• Lamirel, J.-C. et al. (2023). The CFMf Topic-Modeling Method Based on Neural Clustering with Feature Maximization: Comparison with LDA. Proceedings of ISSI 2023.
• Hogan, A. et al. (2021) — Knowledge Graphs (survey).
• Reimers & Gurevych (2019) — Sentence-BERT.
• Strubell et al. (2019) — coût énergétique du deep learning en NLP (perspective frugale).
• [Jean Zay] : https://www.cnrs.fr/fr/presse/jean-zay-le-supercalculateur-le-plus-puissant-de-france-pour-la-recherche
• [ROMEO] : https://romeo.univ-reims.fr/welcome

Keywords

NLP, information extraction, semantic disambiguation, knowledge graphs, information mapping, frugal AI

Funded offer

Funding type
Contrat Doctoral
Countries

Mexico (Conacyt)

China (CSC)

Dates

Application deadline 31/08/26

Duration36 months

Start date01/10/26

Creation date16/01/26

Languages

Level of french requiredNone

Level of English requiredA2 (elementary)

Miscellaneous

Annual tuition fee400 € / year

Website

Contacts

You must connect to be able to display the contacts.

click here to connect or register (it's free!)