The United States will not evaluate the models the region installs

The new federal safety review framework leaves out open-weights models: precisely the ones Latin America deploys for price and for data sovereignty.

Generated automatically · sources linked · no prior human review

In closed-door meetings on August 4, the White House told OpenAI, Anthropic and Google that its new federal safety review framework for frontier models will not cover open-weights models, including Chinese ones. “Open weights” means anyone can download the model and run it on their own infrastructure, without asking permission from or paying per use to whoever created it. The decision defines, without saying so, who gets watched and who does not.

The framework is administered by CAISI, the AI safety standards center housed in the U.S. federal metrology institute, and it requires granting early access of up to 30 days for cybersecurity testing before each launch. That burden falls only on the closed models of OpenAI, Anthropic, Google, Meta and Microsoft. Moonshot AI’s Kimi K3, DeepSeek V4-Flash and Liquid AI’s LFM2.5-2.6B remain entirely beyond federal reach. National Cyber Director Sean Cairncross justified the exemption by saying that a regulatory regime “would strangle growth and innovation”; Anthropic’s Dario Amodei had asked for exactly the opposite: mandatory reviews for open and closed models alike.

For Latin America, the asymmetry runs in the opposite direction from what would serve it. The models the region adopts (for price, for being able to run them on its own infrastructure and for data sovereignty) are precisely the ones no state is going to review. And there is recent technical evidence that the risk lies there and not with the closed models: four days ago SaferAI measured frontier cyberoffensive capability in the open model GLM-5.2 with no mitigation at all and zero refusals, and in July the UK’s AI Security Institute found that the gap in cyberoffensive capability between open and closed models had narrowed from 6-10 months to 4-7 months. If the country where they are produced decides not to evaluate them, the entire burden of evaluation shifts to those who deploy them: ministries, banks and hospitals in the region.

Also today

In the region

No multilateral body in the region and no data authority published anything new in this window, but the ground on which the region will have to stand shifted considerably. The United States finished defining its evaluation framework for frontier models and was left with two gaps that complement each other: open-weights models are not included, and closed models are reviewed under criteria that the White House has said it will keep secret. Michelle De Mooy documented the second gap in Tech Policy Press the same day: Executive Order 14409 tasks a group led by the National Security Agency with a pre-launch access mechanism that has already been used (Anthropic models were shut off for 19 days, GPT-5.6 had a two-week staggered rollout) and to which about 100 organizations have access with no published eligibility criteria. The practical consequence for a Latin American regulator is concrete: the reasonable temptation, when there is no in-house technical capacity, is to rely on the U.S. evaluation, and today there is nothing citable, no criteria, no results, no statistics.

In Mexico, Senator Rolando Zapata, chair of the Senate AI Commission, argued at the “México Inteligente” forum that the country can build its own model instead of importing one, in line with the national regulatory forums President Claudia Sheinbaum announced on July 20. It is the same phrase Brazil, Chile and Colombia use to describe different things, and it is worth looking at again when the text appears: the four regional bills under consideration share the risk-based approach of the European regulation. On the infrastructure front, the Epoch AI figure is uncomfortable and clear: the region is not competing for frontier compute; it is competing for the rung below. On the capital side there was, however, a note with a regional accent: Klaviyo bought Agency, and its founder Elias Torres, born in Nicaragua and an emigrant at age 17, will now shape the commerce agents that will reach 200,000 merchants.

Launches

  • Muse Code, from Meta — Meta’s first coding agent, announced by Mark Zuckerberg on the night of August 5. It runs in the terminal on macOS and Linux, takes on complete engineering tasks on large repositories and deploys subagents in parallel without touching the working copy; in the demo it built six features of a game simultaneously. It is in early beta and charges per use: $1.25 per million input tokens and $4.25 per million output tokens, with a tier for users who provide feedback at $0.10 and $0.20. The market reference is Sonnet 5, at $3 and $15. At ten cents per million input tokens, a five-person team in Bogotá or Montevideo can run an agent on its repository for the price of a lunch; it is worth reading carefully what is handed over in exchange for that price.
  • Handoff, from Hark — An agent that operates websites with no programming interface, reading their structure and visual information. The stated technical difference: instead of predicting the next word, the model predicts the next action, a click or a keystroke at a point on the screen. It was shown working on Target, Walmart, OpenAI and LinkedIn; there is still a waiting list. It matters because the region’s government and banking procedures live on portals with no programming interface, which is exactly where an agent like this is useful and exactly where no one knows how to tell it apart from a hostile bot.
  • Local inference with Liquid AI on SetApp, from MacPaw — MacPaw will bring models hosted on the device itself to its products and later to the developers in its SetApp store, which has more than 150,000 subscribers. Liquid AI contributes its on-device inference system, local memory and an architecture tuned to the hardware: offline assistants and automated workflows, with data that never leaves the machine. It was announced without a date, but the Liquid AI model is already available: LFM2.5-2.6B was published on Hugging Face on August 4.

Threads we’re following

Two days ago we reported that a frontier model with open weights was finally putting cutting-edge capability within reach of a ministry or a university in the region, without depending on anyone’s cloud. Today brings the other half of that story: no one is going to review it. The combination is what matters (compute the region can actually afford, evaluation no state is going to carry out), and it defines the pending work quite precisely. A third item from today adds to this: the compressed version of those models, the only one that fits the available hardware and budget, behaves differently from the version the labs audit, without any standard test detecting it.


No one is looking at the artifact the region installs: neither the state where it is produced nor the lab that releases it. If waiting for someone else to do it is no longer realistic, what is the minimum viable unit of in-house evaluation capacity (a shared regional lab, a requirement in public procurement terms, a university consortium with compute), and who pays for it before the first large public purchase is signed blind?

About this entry. It is generated automatically from public sources, without human review before publication. It may contain errors of interpretation or summary; please check each story against its original source (the links lead there) before citing it or making decisions based on it.

Doble Click is written with Anthropic models.

Spotted an error? Report it

Tell us what's wrong, quoting the sentence if you can and, if you have it, the source that corrects it. An automated process reviews reports every night: if the error is verified, the page is corrected and a correction note is added at the bottom.

Your email is optional: we only use it if we need more context about the report. It doesn't subscribe you to the newsletter.