Nlp Post-Doc - Engineer For Information Mining In Historical Data French 3Rd Republic H/F - INRIA
- CDD
- INRIA
Les missions du poste
A propos d'Inria
Inria, l'institut national de recherche dans les sciences et technologies du numérique, est en appui de l'État pour les stratégies nationales de recherche et d'innovation du numérique en tant qu'Agence de programmes. Inria mène plus de 300 projets de recherche et d'innovation avec ses 3500 scientifiques, ingénieurs et personnels d'appui, en partenariat avec les universités et l'écosystème numérique (entreprises, entrepreneurs, acteurs publics). Ensemble, nous explorons des domaines clés comme l'intelligence artificielle, la cybersécurité, l'informatique quantique, le Cloud, la transformation numérique de la santé, les jumeaux numériques ou encore les technologies numériques pour la défense. Nous construisons des solutions concrètes telles que des logiciels, des startups technologiques, des partenariats avec les entreprises du tissu national et des formations de pointe. Notre objectif : l'impact scientifique, technologique et industriel au service de la souveraineté numérique de la France.
NLP Post-doc / Engineer for Information Mining in Historical Data (French 3rd Republic)
Le descriptif de l'offre ci-dessous est en Anglais
Type de contrat : CDD
Niveau de diplôme exigé : Bac +5 ou équivalent
Autre diplôme apprécié : PhD
Fonction : Ingénieur scientifique contractuel
Niveau d'expérience souhaité : Jeune diplômé
Contexte et atouts du poste
This position is part of the ANR project DECIDON (Débats parlementaires et Espace médiatique (1870-1940) : comprendre la CIrculation du discours politique grâce à des méthodes à forte intensité de DONnées), which is a collaboration between multiple French team, led at Epita Paris by Marie Puren, and including the Inria ALMAnaCH (Inria Paris center) project team as a work package lead for natural language processing techniques. The objective of the project is to enable a better understanding of how the parliamentary debates are being set up in the Third Republic and how thematics can appear and disappear across the newspaper and the parliamentary debate proceedings. The ALMAnaCH team role specifically focus on:
- Development of an annotation interface to support the creation of thematic datasets across the corpus of parliamentary debates and, potentially, the press.
- Development of an easily deployable approach for robust information retrieval across the corpus, targeting topics whose vocabulary may have evolved over time and diverged from contemporary French.
- Evaluation of information retrieval methods for topic-based search on a historically situated corpus (including RAG).
The employee will be supervised by Thibault Clérice, permanent researcher at Inria, and will work in close collaboration with other members of the team and project, namely Florian Cafiero and Marie Puren (EPITA)
They will work at Inria, within the ALMAnaCH project-team. Within the team, they will find researchers connected to the topic outside of the project itself, including Cecilia Graiff, a PhD student in the team, supervised by Benoît Sagot and Chloé Clavel, working on multilingual and cross-cultural automatic analysis of argumentation structures in political debates.
Participation to national meetings and national/international conference are to be expected.
Mission confiée
The post-doc / NLP engineer will design and train models to support historians in constructing focused sub-corpora from large, noisy, OCR'd historical text collections. The work involves:
- Setting up an annotation workflow (true/false positive labeling of keyword occurrences in context) in collaboration with historians;
- Training and evaluating small-scale classifiers - ranging lightweight models to larger pretrained models - capable of distinguishing relevant from irrelevant occurrences of ambiguous or polysemous terms in diachrony (e.g., distinguishing an "anti-parliamentary" use of réforme de l'État from a routine administrative reform);
- Integrating active learning so that model performance improves iteratively as historians annotate, minimizing labeling effort while maximizing corpus quality;
- Prioritizing model deployability: given the scale of the corpus (at least dozens of millions of tokens across noisy OCR output) and the need for the tool to run efficiently and reproducibly within a web interface used directly by non-specialist historians, the research should focus on enabling this on lightway models: fast at inference, and easy to retrain/update rather than relying on large LLM inference at scale;
- Benchmarking against and complementing RAG-based exploration (T5.5), providing a transparent, low-cost alternative for corpus-scoping that historians can audit and reproduce before moving to more exploratory or generative tasks.
The research component focuses on efficient, low-resource sequence classification for historical/noisy text - including handling class imbalance, domain-shift across a ~70-year corpus, and figurative/contextual language. Publishable outputs (methods, benchmarks, and the resulting tool) are expected as part of the project's open, FAIR-data pipeline.
Bibliography:
El Assadi, A., Muennighoff, N., & Lee, J. (2026). The embedder's dilemma: LLMs are better, but at what cost? arXiv.
Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., Villemonte de la Clergerie, É., Seddah, D., & Sagot, B. (2020). CamemBERT: A tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
van Strien, D., Beelen, K., Coll Ardanuy, M., Hosseini, K., McGillivray, B., & Colavizza, G. (2020). Assessing the impact of OCR quality on downstream NLP tasks. In A. Rocha, L. Steels, & J. van den Herik (Eds.), Proceedings of the 12th International Conference on Agents and Artificial Intelligence: ICAART 2020 (Vol. 1, pp. 484-496). SCITEPRESS.
Luo, X., Shinnick, Z., Griesshaber, N., Wang, Y., Yu, J., Shi, F., Torr, P., & Lu, Y. (2026). Pretraining language models on historical text. arXiv.
Bian, D., Puren, M., & Cafiero, F. (2026). How to efficiently explore noisy historical data? Leveraging corpus pre-targeting to enhance graph-based RAG. In D. Alves, Y. Bizzoni, S. Degaetano-Ortlieb, A. Kazantseva, J. Pagel, & S. Szpakowicz (Eds.), Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature 2026 (pp. 241-250). Association for Computational Linguistics.
Clerice, T. (2024). Detecting sexual content at the sentence level in first millennium Latin texts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4772-4783). ELRA and ICCL.
Principales activités
The main activities of the applicant will include:
- carrying out research on the topic outlined above, both in the development of new ideas, positioning with respect to related work and validation of the methodology via experiments and analysis
- Producing a solution that will be deployable for post-lexical information retrieval
- Interact with the project's historians as well as the researchers from the work package on natural language processing
- the presentation of work both internally to colleagues and externally in the form of conference/journal/workshop papers
- interacting and exchanging with colleagues on NLP topics
Compétences
Technical skills and level required :
- Python
- NLP Model Training and evaluation, beyond just LLM
Languages :
- French (Reading understanding to interact with the data)
- English
Relational skills :
- Good organizational skills.
- Good interpersonal skills.
Additional skills considered an asset:
- Knowledge of the French 3rd Republic or similar political systems.
- Interefest in information retrieval
- A strong interest in open science.
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage
Compétences requises
- Python
- Anglais
- Processing
- Français