GPT-5.6 fails the safety test but remains available

The British security institute found universal jailbreaks in OpenAI's new model, as serious as those that led Washington to block a rival model: the difference is who raised the alarm first.

Generated automatically · sources linked · no prior human review

The UK’s AI Security Institute, the public body that audits the most powerful artificial intelligence models, tested the cybersecurity defenses of GPT-5.6 Sol, the model OpenAI released to the public on July 9, and in every round of testing it found “universal” jailbreaks: tricks that get around the model’s restrictions and enable it not only to detect security flaws but to exploit them autonomously. A jailbreak is, in practice, the key that disables the ethical brakes the maker installed.

What is unsettling is not only the finding but how easy it was: the institute says finding those keys took hours of work, and warns that more probably exist despite the mitigations OpenAI has already applied. The underlying point is the comparison. In June, a vulnerability of equivalent severity led the US government to force the global suspension of two Anthropic models, Claude Fable 5 and Mythos 5, using Department of Commerce export controls: it was the first “off switch” a government applied to a frontier AI model. With GPT-5.6, so far, there is no sign that Washington will do the same, and the model remains generally available.

Therein lies the asymmetry that matters for Latin America. The region had no voice when Anthropic’s model was blocked, nor will it have one if the US administration decides, or does not, to apply the same criterion to OpenAI. Access to the world’s most capable AI is being settled case by case, depending on which government raised the alarm first, and not according to a common, public technical standard. For countries that consume this technology without making it, that means finding out about decisions that affect them once they have already been made.

Also today

In the region

Beyond Chile, the region is adding small but concrete institutional pieces: Panama formalized an AI subcommittee days after joining the Pax Silica bloc, and several Latin American and Caribbean countries created a regional group of specialists on disinformation and AI, as a follow-up to the Santo Domingo Declaration. At the global level, the European Commission extended its technological sovereignty model to cybersecurity with a new action plan, while the first UN Global Dialogue on AI governance closed last week with no binding document and with the next meeting not until 2027. The question remains whether multilateralism can move at the pace of the labs.

Launches

  • Reflect (Anthropic) — A “digital well-being” dashboard within Claude that summarizes your usage patterns and suggests breaks; in beta for the Free, Pro and Max plans with memory turned on. Not everyone reads it the same way: TechCrunch describes it more as a user retention strategy than as genuine well-being.

Threads we’re following

In Chile, the Senate Finance Committee confirmed in session, this time with an official source, that the article exempting AI from paying for protected content was not reinstated. It is the latest chapter of a story we have been following: the sweeping legislative reform is moving forward without the copyright exception it debated for weeks, and the floor vote was set for July 15. The dispute over who pays for the works that train AI is not closed; it is postponed.


If the same security flaw justifies blocking one model and tolerating another depending on which government raised the alarm, what would need to exist for Latin America to stop learning from the press about decisions that determine which artificial intelligence it can access?

About this entry. It is generated automatically from public sources, without human review before publication. It may contain errors of interpretation or summary; please check each story against its original source (the links lead there) before citing it or making decisions based on it.

Doble Click is written with Anthropic models.

Spotted an error? Report it

Tell us what's wrong, quoting the sentence if you can and, if you have it, the source that corrects it. An automated process reviews reports every night: if the error is verified, the page is corrected and a correction note is added at the bottom.

Your email is optional: we only use it if we need more context about the report. It doesn't subscribe you to the newsletter.