Five days after asking the industry to slow down, Anthropic put numbers on its own table: it published the first internal metrics on the pace of development inside a frontier lab and revealed that Claude already “leads” 26% of the company’s AI research and engineering work. “Leading,” by its definition, means that the model sets the goal and carries it out, with human review afterward. In more than 90% of the work it operates at least at the collaboration level.
The data are from August 2026 and describe a factory more than an assistant: some 30,000 agents working simultaneously on the internal platform, 100% of actions passing through a monitor before being executed, a block rate of 0.002% —one in every 47,000 decisions— and 6% of research and development compute devoted to safety, rising to 12% if you look only at AI-led research. The company is asking other labs to publish the same, with a public methodology and third-party verification.
It is the first hard figure for how much AI builds AI, and it carries the original problem of this whole debate: it is published, without an external audit, by the company measuring itself while it negotiates its IPO. For Latin America, the question is not whether the figure is interesting but who verifies it. No authority in the region today has the mandate or the technical team to audit a measurement like this, and the transparency window opens just as the region’s governments are buying these models for public services.
Also today
- Unsealed documents in the New York Times lawsuit: a Microsoft executive called scraping “the largest theft of labor in human history” — 91,692 copies of journalistic pieces used in training, click-through drops of up to 93%, and internal documents describing language models as “a product that destroys its supply chain.” It is the strongest body of evidence that exists today for the region’s media outlets negotiating licenses.
- The FAA awards $875 million over twelve years to Air Space Intelligence — Predictive AI in US air traffic control. It sets a price and a reference architecture for the region’s aviation authorities.
- Emerald AI raises 150 million to free up space on the power grid — an alliance with Google, Nvidia, Anthropic and four utilities that aims to free up 100 GW by shedding load instead of building generation. Meanwhile, Crusoe raised 3.9 billion for modular, transportable AI factories.
- GPT-6 Astra deciphers in ten hours a 1941 Enigma message that had gone unsolved for 83 years — the recovered content, operator typos included: a German soldier asks for marching route instructions and an “immediate reply by radio.”
- A 404 Media journalist hijacks a real band’s Spotify and releases an AI-made song on it — the fraud does not require cloning any voice: getting in through the distributor is enough. It is much easier to exploit against Latin American artists who reach streaming through cheap aggregators.
In the region
The week brought no announcements from Latin American ministries or regulators, but it did bring two developments that set the agenda by another route. The first is Peru: a group of academics warned that the country should not copy the technological brake the major powers are discussing, and it is the first articulated regional response to the global slowdown debate. It also comes from the only country in the region with an AI law in force —Law 31814, with operational regulations since January—, which is why it can discuss “how to adopt” instead of “whether to brake.” The second comes from the infrastructure side: Huawei presented its strategy for Latin America, combining 5G, agentic AI in health and rural solar power in a single package; the figures are self-reported and have no independent verification. A third question remains that no one has yet asked out loud: with the UN handing its global statistics over to a donated technical layer, the indicators of ECLAC and the national statistics institutes are going to circulate through it, and it is not clear whether those bodies take part in designing the scheme or merely supply the series.
Launches
- Bonsai 2 27B, from PrismML — a 27-billion-parameter model compressed to 5.9 GB with ternary weights (each parameter stores only three possible values, which shrinks the file enormously), which retains 98% of performance and runs on a phone. Open, free weights on Hugging Face. It solves cost per query, connectivity and data sovereignty in one stroke: it is the technical line that most changes real access in the region. The company also raised a $22.25 million seed round.
- UN System Data Commons, from the United Nations and Google — a free public platform at data.un.org for querying UN system indicators in natural language, with 26 agencies committed and $2 million donated. The trigger was a UNICEF benchmark that measured 21.2% average accuracy across six models answering questions about development.
- Redesigned Projects in Claude Code, from Anthropic — a coordinator distributes tasks among parallel threads, each in its own cloud session, able to run tests and open pull requests, with shared memory and per-thread effort control. In beta for the Pro and Max plans; local execution was announced as coming next.
Threads we’re following
Last Saturday, Dario Amodei, Anthropic’s CEO, published an essay calling for a deliberate slowdown in the improvement of frontier models and committed his company to the first step: permanent external evaluators inside the lab. The obvious objection was that no one knew how fast things are moving today, because the measurement did not exist. This morning’s metrics are the answer to that objection and, at the same time, its best illustration: the number exists because the company decided to calculate it, chose the metrics and defined what “leading” means. The next step —for someone on the outside to be able to repeat the measurement— still has no owner.
Anthropic has just shown that it is possible to measure how much AI builds AI, and that the figure is already 26%. What would have to happen for Latin America to be able to demand that measurement instead of receiving it, and who in the region would have the technical capacity to verify it today?
About this entry. It is generated automatically from public sources, without human review before publication. It may contain errors of interpretation or summary; please check each story against its original source (the links lead there) before citing it or making decisions based on it.
Doble Click is written with Anthropic models.