At a glance
- What it is: Generative AI experimentation in government: Learning from emerging guidelines
- Who: Piret Tõnurist, of the OECD Public Governance Directorate, and Moritz von Knebel, an external consultant. The work was co-funded by the European Union.
- Where: OECD Working Papers on Public Governance No. 93, 2026. doi.org/10.1787/42815683-en
- Type: a systematic review of official government guidelines, complemented by academic literature and case studies. It is neither an experiment nor an impact measurement.
First reading: what it does and what it finds
The document’s starting point is a fact that anyone who works in government will recognize: civil servants are already using generative AI, often without formal approval or clear oversight. Between that rapid, decentralized adoption and the governance mechanisms, which move considerably more slowly, a gap has opened. The work steps into that gap and chooses a specific angle: it does not look at large-scale deployment but at experimentation, understood as the stage where a government tests, learns and decides whether scaling up is warranted.
The method is a review of the official guidelines that governments have already published on the subject. The text reports fourteen countries: Australia, Canada, Finland, Ireland, Italy, Japan, New Zealand, Norway, Sweden, Switzerland, the Netherlands, the United Kingdom, the United States and Singapore. The annex, which lists the documents one by one, also adds the United Arab Emirates and Korea, and gives a sense of the universe: the United Kingdom appears with seven different guidelines and Australia with seven.
The first finding is about form. The guidelines exist and are plentiful, but they are fragmented and very uneven. They range from general statements of values to detailed manuals with checklists. Many governments published high-level principles on ethical or responsible use, and rather little operational guidance on how to design, run and judge a concrete experiment. The consequence the authors point to is not the one you might expect: that disparity does not produce freedom, it produces uncertainty, and uncertainty produces risk aversion or inconsistent practices across agencies and even within the same one. The lack of shared tools also multiplies duplication, because each team reinvents what another has already tested.
The second finding is the one that gives the title, and it is more uncomfortable. Monitoring and evaluation are the critical weakness. Few governments systematically check whether an experiment delivered the benefits it promised, whether it created new risks or whether it justifies scaling up. When there is evaluation, it tends to rely on basic indicators such as usage or user satisfaction, and not on structured measurements of performance, impact, cost and compliance. The authors add a technical point that makes this more serious: since these systems are probabilistic and the same request can produce different results, evaluating them requires auditing outputs; reviewing the design once is not enough. Without agreed criteria or comparable metrics, a government has no way to shut down an experiment that failed or to prioritize what to invest in next.
From there comes the document’s proposal: a framework for evaluating generative AI experiments across five areas, which are the pilot’s performance and the quality of what goes in and comes out, projected impact and public value, the cost and feasibility of integrating it into the institution, usability and acceptability, and risk management and compliance. It closes with a list of ten priority actions, including building public trust in experiments, training civil servants, preventing projects from becoming fragmented, investing in shared tools and data, and planning evaluation from the start rather than at the end.
Second reading: from Latin America
The most telling fact for the region is not in the findings but in the annex: among the countries whose guidelines were reviewed there is none from Latin America. The list is of high-income countries, with the United Kingdom, Australia, Canada and the Nordic countries accounting for most of the documents. This is not a reproach to the authors, who review what exists and is published, but it does define what this work is for a reader in the region: it is a map of what others did, not a diagnosis of our own situation.
That leaves two possible readings, and it is worth not confusing them. The first is about opportunity. If the central finding is that even the countries with the most resources published principles and fell short on evaluation, then a country in the region that is writing its guidelines today can skip that stage and start with the evaluation framework built in. It is cheaper to do it at the beginning than to add it later, and the document provides the five areas already laid out.
The second is a warning, and it is more mine than the paper’s. The underlying diagnosis, civil servants using these tools without formal approval or oversight, does not require national guidelines to happen: it happens anyway, and it probably happens more where there is less oversight capacity. What changes without guidelines is not the use, it is the visibility of the use. It is worth connecting this with something we have already seen in the region: when ChileCompra announces language models to detect irregularities in public procurement, the question the OECD framework would ask first is not whether the model works, but by what criteria it will be measured whether it worked, and what recourse someone flagged by mistake has.
A final note of caution about what kind of evidence this is. It is a review of documents, not a measurement. It says what the guidelines say and what they lack. It does not say whether the countries that evaluate better get better results, because it did not study that. The five-area framework is a reasoned OECD proposal, not an empirically validated instrument.
The fine print
- It is a review of published official guidelines, so it inherits the bias of what is written down and in English. A country may have evaluation practices without having published them as guidelines, and it would not appear here.
- The authors themselves clarify in the notes that their main table only includes countries whose documents explicitly state principles, and that some cases, such as the United States, issued guidance for specific agencies or at the state level that is not represented.
- The five-area evaluation framework is the document’s proposal, not an agreed standard or something that has been tested in the field.
- The work was co-funded by the European Union. The publication itself clarifies that its content is the sole responsibility of the OECD and does not necessarily reflect the position of the European Union. It is worth keeping in mind when reading a document that evaluates government guidelines, several of them from member states.
Paper keywords: generative AI, public sector, experimentation, evaluation, AI governance, public administration
Automated reading. This text was generated by Claude, an Anthropic model, from the original source, without line-by-line human review. It may contain errors or debatable interpretations; to check any point, see the original source.