{"id":35018,"date":"2026-08-10T12:19:38","date_gmt":"2026-08-10T10:19:38","guid":{"rendered":"https:\/\/opi.org.pl\/eurodense-jeden-model-dziewiec-jezykow\/"},"modified":"2026-08-12T13:35:23","modified_gmt":"2026-08-12T11:35:23","slug":"eurodense-jeden-model-dziewiec-jezykow","status":"publish","type":"post","link":"https:\/\/opi.org.pl\/en\/eurodense-jeden-model-dziewiec-jezykow\/","title":{"rendered":"EuroDense \u2013 one model, nine languages"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong><strong>Artificial intelligence in nine languages from OPI PIB<\/strong><\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As generative AI systems that use domain-specific data become more widespread, the importance of information retrieval technologies continues to grow. The quality of responses generated by language models relies on their ability to retrieve relevant documents and sources of information. This has led OPI PIB\u2019s AI Lab to develop advanced dense retrieval models capable of searching for content by meaning, rather than depending exclusively on keyword matching.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>\u2018The new EuroDense model, developed at our institute, supports nine European languages, including Polish, English, German, French, Spanish, Italian, Portuguese, Dutch, and Russian. <\/em><em>With 435 million parameters and a context window of up to 8,192 tokens, the model can also effectively analyse longer documents,\u2019 <\/em>says Dr Marek Koz\u0142owski, Head of the AI Lab at the National Information Processing Institute. <em>\u2018Compared with other multilingual models containing fewer than 1 billion parameters, EuroDense achieved the best results in seven of the nine languages and recorded the highest average NDCG@10 score both across languages and tasks,\u2019 <\/em>adds Dr Koz\u0142owski.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Although the market offers a wide range of sophisticated models for English, significantly fewer provide comparable performance in other European languages. EuroDense was created to specifically address this shortfall. It is also the first language model developed by the AI Lab to include languages other than Polish.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong><strong>Trained on 200 million texts <\/strong><\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">EuroDense was developed in a three-stage training process with two knowledge distillation phases followed by fine-tuning.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>\u2018The training data consisted of a multilingual corpus of approximately 200 million texts. <\/em><em>Ensuring high-quality semantic representations across all supported languages was a priority, allowing the model to efficiently handle information retrieval tasks and support applications based on the RAG architecture,\u2019 <\/em>says Dr S\u0142awomir Dadas, Deputy Head of the AI Lab.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong><strong>EuroDense for everyone<\/strong><\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">EuroDense is now available as an open-source model on the Hugging Face platform. This approach enables academic institutions, public administration, and enterprises to leverage the model for building multilingual search systems, smart knowledge bases, and modern AI-powered applications. By releasing EuroDense, OPI takes another step toward developing European artificial intelligence technologies and strengthening its expertise in advanced natural language processing. EuroDense builds upon the experience accumulated during the development of the PolDense family, adapting it to a multilingual environment to meet the needs of users across Europe.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong><strong>The LLMs4EU project<\/strong><\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The PolDense model was developed as part of the Large Language Models for the European Union (LLMs4EU) project, which is being implemented by the Alliance for Language Technologies European Digital Infrastructure Consortium (ALT-EDIC). The project aims to develop and provide artificial intelligence models, tools and services for five key sectors: science, public services, tourism, telecommunications, and energy. It is also intended to encourage public institutions and small and medium-sized enterprises to adopt European AI technologies. The LLMs4EU project is co-financed by the European Union under the Digital Europe Programme and by the Polish Ministry of Digital Affairs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The EuroDense model is available free of charge on the <a href=\"https:\/\/huggingface.co\/OPI-PIB\/EuroDense-435M\" class=\"external\" rel=\"nofollow\" target=\"blank\">the Hugging Face platform<\/a>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The AI Lab at the National Information Processing Institute (OPI PIB) has developed and released EuroDense, a model designed for information retrieval in nine major European languages. It is the first multilingual solution of this kind developed by the OPI PIB research team. The model was designed for modern search engines, Retrieval-Augmented Generation (RAG) systems, chatbots, and AI assistants that require efficient retrieval of information from large document sets.<\/p>\n","protected":false},"author":30,"featured_media":35020,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_links_to":"","_links_to_target":""},"categories":[411],"tags":[835,1099,846,1106,1107,492,838,876,840],"class_list":["post-35018","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news-en","tag-ai-en","tag-ailab","tag-artificialintelligence-2","tag-eurodense","tag-gpt","tag-innovation","tag-it-en","tag-languagemodels-2","tag-opipib-2-en"],"_links":{"self":[{"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/posts\/35018","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/users\/30"}],"replies":[{"embeddable":true,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/comments?post=35018"}],"version-history":[{"count":1,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/posts\/35018\/revisions"}],"predecessor-version":[{"id":35022,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/posts\/35018\/revisions\/35022"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/media\/35020"}],"wp:attachment":[{"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/media?parent=35018"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/categories?post=35018"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/opi.org.pl\/en\/wp-json\/wp\/v2\/tags?post=35018"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}