Trial of LLMs: Dutch framework evaluates 30 government models and reveals a clash between quality and environmental cost
Listen to this article
Read by Anchor
Government uses of large language models in public work environments are increasing, but current evaluation frameworks rarely combine reflection of public administration values and linguistic requirements for non-English contexts. In this context, a research team including Laurens Samson, Eva Gornischka, Gossa Lo, Yuki M. Asano, and Senay Gibril, presented a paper accepted at the AIES 2026 conference, to develop a systematic evaluation framework called Grip on LLMs. This framework aims to evaluate AI models intended for government use in Dutch, bridging the gap between institutional values and technical measurement tests.
The framework was developed through direct collaboration with experts from a major Dutch municipal organization. The researchers based their methodology on a comprehensive process that included forming an advisory board, conducting user research, and implementing a survey of users of a conversational robot designed for civil servants. These processes resulted in the identification of six key dimensions for evaluation: realistic accuracy, honesty, social bias, energy consumption, financial cost, and training data transparency. The team translated these dimensions into a standard test package that evaluated over 30 models, ranging from multilingual models to those specifically designed for Dutch.
The evaluation results showed that no single model could excel in all six dimensions at the same time, making trade-offs inevitable when choosing the appropriate model. The study revealed that higher quality performance is consistently associated with greater environmental impact and higher financial cost, while social bias remains largely independent of quality and cost levels.The results showed that realistic accuracy and honesty are subject to entirely separate properties, as a model's ability to answer correctly does not necessarily mean it can recognize what it does not know.
In terms of realistic accuracy, the framework measures the model's ability to provide correct answers, while the honesty dimension focuses on the model's recognition of its knowledge limits and its admission of what it does not know without providing misleading information. The detailed results show that improving the level of accuracy and transparency in training data imposes additional financial and environmental burdens, as it requires higher energy consumption and high operating costs, putting government institutions in the face of delicate trade-offs between environmental sustainability, financial efficiency, and model performance.
To turn these theoretical results and field tests into practical tools that non-experts can benefit from, the researchers launched a public overview designed in a simple style suitable for all stakeholders involved in choosing large language models in the government sector, from technical engineers to policymakers. This tool provides a comprehensive vision that helps public institutions make balanced decisions that combine ethical standards, environmental commitments, budget constraints, and linguistic requirements specific to the Dutch language.