OPI PIB’s PolDense 1B model tops the Polish Information Retrieval Benchmark

Achievements
The AI Lab at the National Information Processing Institute (OPI PIB) has developed PolDense, a new family of language models designed specifically for information retrieval. The model is the latest addition to an OPI PIB’s portfolio of advanced AI solutions tailored to the Polish language and the needs of domestic users and institutions.

PolDense models have been designed for systems that process large volumes of unstructured data, particularly search engines, chatbots, AI assistants and applications that rely on the increasingly popular Retrieval Augmented Generation (RAG) architecture. The quality of the PolDense models is best demonstrated by their results in the Polish Information Retrieval Benchmark (PIRB), which is one of the leading benchmarks for evaluating the efficiency of Polish-language information retrieval models. According to the public PIRB ranking, PolDense 1B is the top-performing model, with the highest average score among all evaluated models.

The AI Lab develops key components for cutting-edge AI systems

A growing number of solutions powered by large language models rely on the RAG approach, which enables them to generate answers that are delivered based on both the knowledge within the model and up-to-date documents that are relevant to a particular domain or industry and retrieved in real time. The effectiveness of such systems depends largely on the quality of their information retrieval mechanisms. The best results in this field are currently achieved by so-called dense retrievers. These models use deep neural networks to transform queries and documents into compact vector representations, allowing content to be retrieved based on its meaning rather than solely on keyword matching.

This is precisely the type of information retrieval technology currently being developed by the AI Lab at OPI PIB. The team is advancing the PolDense model family for the Polish language, while also working on EuroDense, a model designed to support major European languages.

‘The launch of the PolDense models is another step towards building Poland’s expertise in artificial intelligence. We are creating open technologies that serve researchers, public authorities and businesses to develop cutting-edge, efficient and secure AI tools. The quality of our solutions is evidenced by one of OPI PIB’s new models ranking first in the Polish Information Retrieval Benchmark, outperforming multilingual models such as Llama and BGE,’ said Jarosław Protasiewicz, Head of the National Information Processing Institute.

OPI PIB’s model ranks first in the PIRB benchmark

The quality of the PolDense models is best demonstrated by their results in the Polish Information Retrieval Benchmark (PIRB), which is one of the leading benchmarks for evaluating the performance of Polish-language information retrieval models. According to the PIRB ranking, PolDense 1B is the top-performing model, with the highest average score among all evaluated models. It outperforms significantly larger multilingual solutions, including models with nearly ten billion parameters.  This result confirms that the solutions developed by the AI Lab at OPI PIB are among the world’s leading information retrieval technologies for the Polish language, offering an optimal balance between effectiveness and implementation costs.

‘PolDense 1B’s first place in the PIRB ranking demonstrates that specialised models designed for the Polish language can not only set quality standards for domestic models, but also compete effectively with international and multilingual models many times their size. This success reflects years of dedicated research at the AI Lab focused on cutting-edge neural representation models and specialised information retrieval models for the Polish language,’ said Sławomir Dadas, Deputy Head of the AI Lab at OPI PIB.

Six OPI PIB’s models for different applications

The PolDense family consists of six models based on the ModernBERT architecture and ettin-encoders technology. They range from 17 million to 1 billion parameters, allowing users to select a model suitable both for tasks requiring the highest possible quality and for environments with limited computing resources. The models can analyse texts of up to 8,192 tokens, enabling them to retain the context of longer documents more effectively and identify the most relevant information with greater accuracy.

The largest model in the family, PolDense 1B, scored 64.11 points in the PIRB benchmark, establishing a new performance level for information retrieval in the Polish language. It outperformed significantly larger multilingual models, including Llama-Embed-Nemotron-8B and BGE-Multilingual-Gemma2-9B, which scored 63.73 and 63.26 points respectively.

The models with 400 million and 150 million parameters offer highly competitive results compared with solutions containing several billion parameters. The smallest variants, with 68 million, 32 million and 17 million parameters, have been designed for deployment on CPUs, mobile devices, and edge computing environments.

OPI PIB’s open solutions for science, public administration and business

The National Information Processing Institute has released the PolDense models as open-source solutions. They can be used to develop custom search mechanisms, RAG systems, and document database management tools. The PolDense models not only improve the quality of information retrieval but also reduce implementation costs. By using smaller and more computationally efficient models, organisations can reduce resource requirements without compromising performance.

‘PolDense demonstrates that it is possible to develop specialised Polish-language models capable of competing with the world’s largest solutions and outperforming them in many specific tasks.  Our goal was to create a family of models suitable for a wide range of applications, from large corporate systems to lightweight solutions operating locally. All our models are available free of charge on Hugging Face to support the development of Poland’s AI ecosystem,’ said Marek Kozłowski, Head of the AI Lab at the National Information Processing Institute.

‘We are proud of the success of the PolDense models, but we remain focused on further innovation and future improvement. In the coming months, the AI Lab plans to release EuroDense, a model designed to support information retrieval across nine European languages, added Marek Kozłowski.

PolDense Models developed as part of the LLMs4EU project

The PolDense models were developed as part of the Large Language Models for the European Union (LLMs4EU) project, which is being implemented by the Alliance for Language Technologies European Digital Infrastructure Consortium (ALT-EDIC). The project aims to develop and provide artificial intelligence models, tools and services for five key sectors: science, public services, tourism, telecommunications, and energy. It is also intended to encourage public institutions and small and medium-sized enterprises to adopt European AI technologies. The LLMs4EU project is co-financed by the European Union under the Digital Europe Programme and by the Polish Ministry of Digital Affairs.

The PolDense models are available free of charge on the Hugging Face platform.