CF202647354
Post-training neural architecture optimization for small language models
D-97
Doctorate Full Doctorate
Disciplines
Software Engineering
Laboratory
Laboratoire Intelligence Intégrée Multi-capteurs Département Systèmes et Circuits Intégrés Numériques (LIST)
Host institution
Université Paris-Saclay
Doctoral school
INTERFACES : APPROCHES INTERDISCIPLINAIRES / FONDEMENTS, APPLICATIONS ET INNOVATION - ED 573

Description

Generative AI, and particularly language models (LLM), have sparked a new revolution in AI with applications across all domains. However, LLMs are highly resource-intensive and, hence, difficult to implement on autonomous embedded systems. LLMs can be optimized by modifying their architecture to replace heavy Transformer layers with lighter alternatives. Given the difficulty of training LLM "from scratch," this thesis aims to develop post-training neural architecture optimization methods applicable to small LLM (SLM). Additionally, the thesis seeks to propose performance metrics of different layers of an SLM and their alternatives, to guide the replacement, and thus propose a comprehensive methodology for optimizing SLMs while considering hardware constraints. The work will be valorized through publications in major AI conferences and journals, and the developed codes and methods could be integrated into the tools developed at CEA.

Funded offer

Funding type
CEA

Dates

Application deadline 30/11/26

Duration36 months

Start date01/10/26

Creation date03/04/26

Languages

Level of french requiredNone

Level of English requiredNone

Opportunity to make his thesis in English

Miscellaneous

Annual tuition fee391 € / year

Website

Contacts

You must connect to be able to display the contacts.

click here to connect or register (it's free!)