Pleias 1.0: the First Ever Family of Language Models Trained on Fully Open Data

Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee, Carlos Rosas Hinostroza, M. Delsart, Irène Girard, Anastasia Stasenko, Ivan P. Yamshchikov · Procedia Computer Science · 2025

Linguistic diversity and strong generalization in foundation language models are typically achieved by training on trillions of data tokens with very large model parameter counts. However, most such training datasets include substantial amounts of copyright-protected or private data that is not explicitly published under the licence that is permissive for LLM training, raising legal and ethical concerns. We introduce Pleias 1.0, a family of comparatively small foundation language models (with at most 3 billion parameters) trained exclusively on public domain or permissively licensed data. We release the model weights and training code so that our results are fully auditable and reproducible. Furthermore, we fine-tune our models for the Retrieval-Augmented Generation (RAG) task and demonstrate that these models – despite their smaller size – can outperform competitors that have orders of magnitude more parameters on RAG evaluations. All models, data, and code are released under open licenses, offering a new standard for transparency and compliance.

Read the paper · More papers on PaperTik