Skip to content

DFM Mimir v1 puts post training data source at the heart of model testing

Share
DFM Mimir v1 puts post training data source at the heart of model testing

Listen to this article

Read by Anchor

In the race among large language models, the usual questions feel familiar: how many parameters, what size the test set is, and who is ahead of whom. But a new paper on the “DFM Mimir v1” model starts with a narrower, more sensitive question: what data may be used after training? This shift is important because the research team describes the prevailing reliance on massive datasets, which may not all be permissible, as an obstacle for researchers who want to work with open-source and ethically vetted data.

The paper, submitted to arXiv on August 13, presents a language model with a Hierarchical Reasoning Model architecture and says it was trained from scratch. It has a billion parameters, a figure the study cites clearly, but its size is not the whole story. The reported results rely on the use of permitted post-training data, not simply on increasing scale or expanding access to data whose source is unspecified.

The page classifies the work as a 20-page technical report on a single model, and the first version is dated 13 August. It attributes it to Peter Schneider Camp, Jakob Nielsen, Gianluca Parmena, Kenneth Involfson and Lukas Galkhe Boisch. These data do not add evidence of performance, but they define what is available to read: a specific research paper, with a published version and identified authors, and not a product announcement that explains commercial terms of use or a public release date.

The comparison here is conditional on what the paper tested.The study says that “Mimir v1” was trained on a mix of 161 datasets and that it outperformed the original “HRM-Text 1B”. It also says it competed with larger models, including “Qwen 3.5 4B” and “Gemma 4 E2B”, across 20 benchmarks covering English, mathematics, programming and Danish. These are results presented by the researchers in a technical report, not an independent judgment beyond the tests they selected.

The strongest signal in the paper concerns the Danish language. The authors note that the model achieves their best announced result in this language, alongside competitive performance in English. Therefore the paper does not suggest that a small model outperforms on every task, nor does it specify a date for a commercial release. What it presents is a specific research case: a model with a billion parameters, post-training data that the team describes as permissible, and results on a set of announced benchmarks.

Openness here is part of the experiment’s subject.The model is available on the Hugging Face platform, according to the paper. This ties the description of “open” to the model’s availability as well, rather than to a general discourse on transparency. In contrast, the attached page does not provide a detailed breakdown of each of the 161 datasets, nor does it demonstrate on its own that any subsequent use of the model meets legal or ethical requirements in every market. Distinction is necessary so that the phrase “permitted data” does not become a broader claim than the study states.

For readers in the region, the paper does not present an Arab project or a regional partnership, and it is not accurate to attribute it to that. However, it poses a question that can be transferred to any program building a local model: is capacity measured solely by size, or also by the data pipeline that can be defended? Mimir v1 does not answer this question for everyone, but it makes the answer testable in a specific model rather than keeping it as a slogan.

The bottom line is that the paper’s value does not lie in announcing the end of the race for larger models. Its value is that it documents an experiment that links performance to constraints on post-training data, and presents its results and limitations together. Anyone who wants to evaluate it accurately should read it as a technical report on a specific model and its metrics, not as a comprehensive proof that small size or permitted data settle every comparison.

Don't miss the next story

Subscribe for updates