European Lexicographic Infrastructure for Artificial Intelligence
The European Lexicographic Infrastructure for Artificial Intelligence project (ELEXAI) has been launched in order to upgrade the existing ELEXIS lexicographic infrastructure with the latest developments in AI, in particular those related to the emergence of LLMs. The new infrastructure will offer upgraded technology, language data, tools, and services that are crucial for improving transformer-based models with multilingual knowledge management originating from high-quality lexicographic resources.
In order to achieve this, we will evaluate the performance of LLMs in lexicography and assess their potential to optimize the lexicographic process while safeguarding the high quality and trustworthiness of lexicographic data. One of the main goals is to upgrade the machine-readable knowledge representations (knowledge graphs). Another important goal is to use agentic LLMs for enriching and expanding lexicographic resources (e.g. sense relations, definitions, multiword expressions). This will also provide lexicographic data to the work packages responsible for improving LLMs with lexicographic data, realising a virtuous cycle between lexicography and AI.
To verify the effect of incorporating linguistic knowledge in LLMs, the creation of reliable benchmarks and other means of evaluation of machine-generated output is foreseen as a result. We will systematically evaluate LLM capabilities of understanding, reasoning, and generating natural language, and we will in particular look into figurative language where LLMs are still struggling in medium- and low-resourced languages. Another project goal is to preserve linguistic and cultural diversity in LLMs for the European languages.
As such, the infrastructure design will contribute to the yet unresolved tasks of Natural Language Understanding. The establishment of a new virtual lexicographic infrastructure will be carried out by a broad and diverse consortium, including partners from all relevant fields: lexicography, Computational Linguistics, and AI. For long-term sustainability, the infrastructure will rely on several prominent infrastructural initiatives: CLARIN and DARIAH, two ESFRI Landmark infrastructures, and ALT-EDIC, as the new pan-European initiative dedicated to the development of European open massively multilingual language models.
In the meantime, you can read more about our partners:
Jozef Stefan Institute | CLARIN ERIC | ALT-EDIC | DARIAH ERIC | University of Ljubljana | Lexical Computing | Instituut voor de Nederlandse Taal | University of Galway | Eesti Keele Instituut | Berlin-Brandenburgische Akademie der Wissenschaften | ELTE Hungarian Research Centre for Linguistics | Borys Grinchenko Kyiv Metropolitan University | Det Danske Sprog- og Litteraturselskab | Leibniz-Institut für Deutsche Sprache | Euskal Herriko Unibertsitatea | Università Cattolica del Sacro Cuore | Institute for Bulgarian Language | Macedonian Academy of Sciences and Arts | Universidade Nova de Lisboa | Univerzitet u Beogradu – Rudarsko Geoloski Fakultet | Belgrade Center for Digital Humanities | Mykolo Romerio Universitetas | Accademia Europea di Bolzano | Athena Research Center | Politechnika Wroclawska



