From text-lake to knowledge graph: frugal algorithms for mapping knowledge (traceability, evidence, diachrony)
D-6
Doctorate Full Doctorate
- Disciplines
- Other (Computer Sciences)
- Laboratory
- LORRAINE LABORATORY OF RESEARCH ON INFORMATICS AND ITS APPLICATIONS (LORIA)
- Host institution
- Université de Lorraine
Description
Context:Scientific corpora as well as public-debate corpora now produce a paradoxical effect: information is abundant, but gaining an overall understanding is increasingly difficult. Important statements circulate under multiple formulations, evolve over time, change speaker, and rely on evidence that can be contradictory. Search engines retrieve documents without making this dynamic visible; automatic summaries remain local; and traditional knowledge graphs often assume data that is already structured.
What does it mean, concretely?
The goal is to build a navigable 'map' of a domain (e.g., a scientific field or a media corpus) or a debate: identify key ideas/claims in a large collection of texts, group formulations that essentially express the same claim, link each claim to its sources (and, when possible, to the cited evidence), and track how these claims evolve over time (emergence, rephrasing, controversy, consensus).
Scientific objective;
Develop and evaluate new algorithms to transform massive, unstructured text-lakes into coherent, auditable, and extensible knowledge graphs. The thesis will focus in particular on disambiguation/canonicalization strategies (detecting when two formulations refer to the same claim) under cost constraints (frugal approaches), and on how a flexible ontology can (i) enable better representations while (ii) reducing dimensionality and (iii) guiding graph construction (ideas, actors, sources, evidence, temporal relations).
Skills required
Candidate profile Strong command of Python (or equivalent programming tools). Interest in linguistics, NLP, information extraction, semantic representations and/or graphs (Neo4j, RDF, etc. appreciated). Master's degree (or equivalent) in computer science / data science / AI (or related fields). Enjoys data processing, rigorous evaluation (benchmarks, ablations), and scientific writing.Bibliography
Lamirel, J.-C. (2012). A new approach for automatizing the analysis of research topics dynamics: application to optoelectronics research. Scientometrics, vol. 93(1), pages 151-166. Lamirel, J.-C. et al. (2014). Federating clustering and labeling capabilities based on feature maximization. Neurocomputing, 147, 136-146.
Lamirel, J.-C. et al. (2020). An overview of the history of Science of Science in China based on bibliographic and citation data: a new method based on clustering with feature maximization and contrast graphs. Scientometrics.
Lamirel, J.-C. et al. (2023). The CFMf Topic-Modeling Method Based on Neural Clustering with Feature Maximization: Comparison with LDA. Proceedings of ISSI 2023.
Hogan, A. et al. (2021) Knowledge Graphs (survey).
Reimers & Gurevych (2019) Sentence-BERT.
Strubell et al. (2019) coût énergétique du deep learning en NLP (perspective frugale).
[Jean Zay] : https://www.cnrs.fr/fr/presse/jean-zay-le-supercalculateur-le-plus-puissant-de-france-pour-la-recherche
[ROMEO] : https://romeo.univ-reims.fr/welcome
Keywords
NLP, information extraction, semantic disambiguation, knowledge graphs, information mapping, frugal AIFunded offer
- Funding type
- Contrat Doctoral
- Countries
-
Mexico (Conacyt)
China (CSC)
Dates
Application deadline 31/08/26
Duration36 months
Start date01/10/26
Creation date16/01/26
Languages
Level of french requiredNone
Level of English requiredA2 (elementary)
Miscellaneous
Annual tuition fee400 € / year
Contacts
You must connect to be able to display the contacts.
