Post-Doctoral Research Visit F - M Postdoctoral Researcher In Medieval French Philology In The Romam2 Project H/F - INRIA
- CDD
- INRIA
Les missions du poste
A propos d'Inria
Inria, l'institut national de recherche dans les sciences et technologies du numérique, est en appui de l'État pour les stratégies nationales de recherche et d'innovation du numérique en tant qu'Agence de programmes. Inria mène plus de 300 projets de recherche et d'innovation avec ses 3500 scientifiques, ingénieurs et personnels d'appui, en partenariat avec les universités et l'écosystème numérique (entreprises, entrepreneurs, acteurs publics). Ensemble, nous explorons des domaines clés comme l'intelligence artificielle, la cybersécurité, l'informatique quantique, le Cloud, la transformation numérique de la santé, les jumeaux numériques ou encore les technologies numériques pour la défense. Nous construisons des solutions concrètes telles que des logiciels, des startups technologiques, des partenariats avec les entreprises du tissu national et des formations de pointe. Notre objectif : l'impact scientifique, technologique et industriel au service de la souveraineté numérique de la France.
Post-Doctoral Research Visit F/M Postdoctoral Researcher in Medieval French Philology in the ROMAM2 Project
Le descriptif de l'offre ci-dessous est en Anglais
Type de contrat : CDD
Niveau de diplôme exigé : Thèse ou équivalent
Fonction : Post-Doctorant
Niveau d'expérience souhaité : Jeune diplômé
Contexte et atouts du poste
This position is part of the ANR project ROMAM², led by Thibault Clérice at Inria Paris within the ALMAnaCH project-team (Automatic Language Modelling and Analysis & Computational Humanities, led by Benoît Sagot). ROMAM² treats pre-editorial normalisation (PEN) of graphemic automatic text recognition (ATR) output as a dedicated NLP task. PEN is a traceable process in which every editorial inference (abbreviation expansion, post-correction of recognition errors, spelling normalisation) stays anchored to the manuscript it comes from. The project works on two medieval languages: Latin, which is heavily abbreviated, and Old French, whose spelling varies in linguistically meaningful ways.
A central claim of the project is that editorial normalisation is not neutral. Printed critical editions silently expand abbreviations and regularise spelling. In doing so, they erase variation that is evidence for the history of the language. Dees' quantitative geography of Old French and its successors rest largely on such editions, and Morin has shown how this can distort dialectal conclusions. However, no study has yet measured this distortion on a controlled parallel corpus. This postdoctoral position is designed to produce that study and the gold data it requires.
The postdoctoral researcher will be supervised by Thibault Clérice. They will work closely with:
- the project's PhD candidate in NLP, co-supervised by Thibault Clérice, Benoît Sagot and Rachel Bawden. The PhD candidate will use the gold data and evaluation framework produced by the postdoc to train and evaluate normalisation models.
- David Smith (Northeastern University), a specialist in aligning noisy historical data.
- the ANR JCJC project Phil-IA, coordinated by Ariane Pinche at CIHAM (UMR 5648, ENS de Lyon). A regular collaboration is expected on Old French graphemic transcription, digital editing and TEI encoding.
The position is based at Inria Paris, within ALMAnaCH. The team brings together researchers in NLP, language modelling and computational humanities, and offers a rare environment for philologists and linguists who want to work directly with NLP researcher. The postdoc will also benefit from existing community resources developed by the team: the CATMuS dataset (the largest ATR dataset for medieval manuscripts), the CoMMA corpus (3.3 billion tokens of Latin and Old French from over 32,000 manuscripts) and the upcoming work on Biblissima-Textes.
The position is for 18 months, starting March 2027.
Bibliography
- Clérice, T., Bawden, R., Glaise, A., Pinche, A., & Smith, D. (2026). Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC 2026. https://arxiv.org/abs/2602.13905
- Clérice, T., Pinche, A., Vlachou-Efstathiou, M., Chagué, A., Camps, J.-B., et al. (2024). CATMuS Medieval: A multilingual large-scale cross-century dataset in Latin script for handwritten text recognition and beyond. In Proceedings of ICDAR 2024 (LNCS 14806, pp. 174-194). Springer. https://doi.org/10.1007/978-3-031-70543-4\_11
- Clérice, T., Gabay, S., Vlachou-Efstathiou, M., Pinche, A., & Sagot, B. (2026). CoMMA, a Large-scale Corpus of Multilingual Medieval Archives. In Proceedings of the Fifteenth Language Resources and Evaluation Conference. ELRA. https://inria.hal.science/hal-05299220
- Dees, A. (1985). Dialectes et scriptae à l'époque de l'ancien français. Revue de Linguistique Romane, 49(193-194), 87-117.
- Morin, Y. C. (2006). Histoire du corpus d'Amsterdam : le traitement des données dialectales. In Le Nouveau Corpus d'Amsterdam. Actes de l'atelier de Lauterbad.
- Scheer, T., & Brun-Trigaud, G. (2022). L'atlas Dees électronique. Concordial, Grenoble. https://hal.science/hal-03912660
- Kuparinen, O., & Scherrer, Y. (2024). Corpus-based dialectometry with topic models. Journal of Linguistic Geography, 12(1), 1-12.
- Pinche, A. (2022). Guide de transcription pour les manuscrits du Xe au XVe siècle. https://hal.archives-ouvertes.fr/hal-03697382
- Duval, F. (2012). Transcrire le français médiéval : de l'« Instruction » de Paul Meyer à la description linguistique contemporaine. Bibliothèque de l'École des chartes, 170(2), 321-342.
Mission confiée
The postdoctoral researcher, trained in philology or historical linguistics with skills in digital humanities, will build the philological foundations and evaluation resources of ROMAM² and carry out a controlled study of how editorial practices affect the dialectometry of Old French. The work involves:
- Building a multi-layer gold corpus from the Nouveau Corpus d'Amsterdam (NCA, formerly Dees' corpus). Using Tobias Scheer's Atlas Dees Électronique mapping between NCA editions and their source manuscripts, the postdoc will sample each text (~500 words per document, ~100,000 words in total) and transcribe them in eScriptorium, align the edited text with a graphemic transcription of the manuscript, and annotate tokens as abbreviated or not in XML-TEI.
- Abbreviation-aware dialectometry. The postdoc will quantify abbreviation practices across the corpus and compare the resulting feature maps and dialectal distances with those of Dees and his successors (e.g. Scherrer's Dialektkarten). They will assess how far ambiguous abbreviation resolution, of the kind Morin identified in Floovant, changes the conclusions, using several dialectometric methods.
- Auditing and augmenting the training data. In collaboration with David Smith, the postdoc will audit the automatically aligned corpus from the prototype PEN work. The goal is to separate valid alignments (identity, ATR post-correction, abbreviation expansion) from invalid ones (literary variants, spelling variants), and to explore inter- and intra-manuscript alignment between witnesses. This work also yields a corpus for studying abbreviation practices across textual traditions.
- Contributing to the evaluation framework, jointly with the PhD candidate. This includes lossless conversion between ALTO, plain text and TEI that preserves uncertainty markup (, , ), stage-specific metrics, and a fine-grained error taxonomy (overnormalisation, variant insertion, hallucination, morphosyntactic errors, ambiguity collapse, etc.). This builds on the expertise of Ariane Pinche and the Phil-IA project.
- Supervising annotation work carried out by hourly-paid annotators, with the PI, to ensure the linguistic and editorial quality of the gold data (Old French and Latin if the applicant knows Latin).
The research component lies at the intersection of historical linguistics, material philology and NLP evaluation. The postdoc is expected to publish results in both communities: a paper on the impact of normalisation on dialectal attribution in Old French, open datasets (the TEI NCA sample and the new gold PEN dataset), and a proposed panel at the International Medieval Congress (Leeds).
Principales activités
The main activities of the applicant will include:
- carrying out research on the topic outlined in the job descriptions, including developing new ideas, positioning the work with respect to related work in historical linguistics and NLP, and validating the methodology through corpus construction, experiments and analysis
- producing and releasing open, reusable datasets in XML-TEI using the ParamHTRs interface (TEI NCA sample with abbreviation annotation, gold PEN data for Old French and Latin)
- working closely with the project's PhD candidate (starting September 2027), so that the gold data and evaluation framework directly support model development
- collaborating with the ANR Phil-IA project (CIHAM, Lyon) and with the project's external experts (David Smith, Ariane Pinche)
- supervising and checking annotation work carried out by annotators
- presenting work both internally and externally in conference, journal and workshop papers, in NLP and humanities venues (e.g. Revue de linguistique romane, IMC Leeds, CHR, LREC)
- Organizing and taking part in the project's workshops and exchanging with colleagues on NLP and philological topics
Compétences
Required:
- PhD in philology, historical linguistics, medieval studies or a related field
- Strong knowledge of Old French
- Training in palaeography, with experience reading medieval manuscripts
- Working knowledge of XML-TEI
- Ability to do statistics and to process or parse structured data with at least one programming language (Python or R)
Highly appreciated:
- Experience in dialectometry, or a demonstrated interest in the dialects and scriptae of medieval French
- Knowledge of medieval Latin
Soft skills:
- Ability to work in an interdisciplinary team with NLP researchers
- Good written and oral communication in English; French is an asset
- Good organization skills
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage