OpenAI's first chip beats Nvidia on efficiency

The frontier of the cost of serving a model moved to a place where neither hardware nor access is for sale, just as the region is betting on buying both.

Generated automatically · sources linked · no prior human review

OpenAI did something no lab had done when launching its first chip: it published the results and let an outsider measure them. It invited the analysis team SemiAnalysis into its lab, gave it access to the hardware and let it run its own public battery of tests on Jalapeño, the inference chip the company co-developed with Broadcom, with some runs verified in person. An inference chip is the one that runs an already trained model every time someone types something to it.

The numbers are uncomfortable for Nvidia. On open models such as GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, Jalapeño delivered between 1.5 and 1.9 times more AI work per watt consumed, and end-to-end latency between 1.7 and 3.6 times lower, than the best results of Nvidia’s GB200/GB300; in low-latency conversational traffic—the kind a chat generates—the declared advantage reaches 4.1 times. Seven hundred watts versus fourteen hundred. SemiAnalysis’s verdict is the line worth remembering: first-generation chips are normally not competitive, and this one beats every Nvidia, AMD and Google chip they were able to test.

For Latin America, what is at stake is not the rivalry between suppliers but the price of the token: the minimum unit of text that is charged every time a model is used, and the only thing in this chain that the region pays for every day. If the efficiency of serving a model doubles and whoever captures that margin is the same party that sells the model, the result is not necessarily a cheaper token: it is a supplier that is harder to replace. And it comes just as Brazil committed 1.06 billion reais to a supercomputer for late 2027 and Chile is debating a premium token bank in its 2027 Budget. Both bets assume that the bottleneck is solved by buying compute or buying access; the efficiency frontier just moved to a place where neither is for sale. Jalapeño is not for sale and cannot be tested: it serves exclusively OpenAI’s own infrastructure, with small-volume deployment in late 2026 and broad deployment in 2027.

Also today

In the region

A fourth consecutive day without an institutional event from the region itself: no Latin American ministry, data authority or regulator published anything within the window. What did arrive, from outside, is the most articulate critique to date of transplanting the European model to a country in the region. Daniel Arias Rivera, a Colombian lawyer and lecturer at the Universidad de los Andes, argues in Tech Policy Press that Colombia’s AI bill is an impoverished copy of the European AI Act, and none of his four objections is exclusive to Colombia: it creates obligations that local companies and institutions cannot meet while doing little to rein in the foreign companies that develop the most powerful systems; it does not make clear who implements it, who is accountable or with what resources; it gives the Ministry of Science the power to define by administrative act which uses are high-risk; and it loads transparency duties onto whoever deploys a foreign model whose training data, architecture and documentation they cannot see. Chile is preparing a new bill, Brazil is debating PL 2338 and Mexico announced its law: all four share the same risk-based structure imported from Brussels.

Two external moves with a predictable regional destination. Meta is negotiating a mid-trial settlement with the attorneys general of 29 states in which the accusation is not about the model but about the data: having collected information from users it knew were minors, without parental consent, and having used it to train machine learning and generative AI systems. That is exactly the route that the data protection authorities of Brazil, Mexico, Colombia and Chile have open today, without needing to wait for an AI law. And the U.S. immigration agency ICE went looking for contractors to hand it the voter rolls of all 50 states to feed analytics systems, following a pattern the region already knows: first the data is consolidated for a narrow purpose, then it is integrated into the platform, and only then does the real use appear. What makes it delicate is that the entry point is the voter roll: throughout Latin America it is administered by an autonomous body precisely to shield it from the executive branch. It comes two weeks before the AI summit the United States is hosting in Lima on September 8.

Launches

  • Claude Cowork with unified memory — Anthropic merged memory between chat and its agent work environment: what is learned in one becomes available in the other, it is saved during the conversation rather than when it is closed, and the user can view, edit or delete what is stored. It is on by default in the Free, Pro and Max plans, on web, desktop and mobile. The feature runs squarely into Chile’s data protection Law 21.719, which takes full effect on December 1, 2026.
  • S1, from Skild AI — a robotics foundation model that learns in context: it is shown a video of a person doing a task and performs it without retraining, on tasks of up to ten minutes that it had never seen. It reports 66% success on unseen tasks versus 9% for language-guided alternatives, with one demonstration equivalent to about 380 episodes of post-training. It is closed, available only to commercial partners, and the figures are self-reported; even so, it radically lowers the cost of automating a production line, which is the profile of manufacturing in Mexico and the Southern Cone.

Threads we’re following

Three days ago we reported that Nvidia raised the price of its servers 15% without that increase reflecting higher costs, and that the bill would come due in 2027, precisely on the equipment Latin America had just bought. Today’s chapter is the other side of that story: the world’s largest customer stopped being just a customer and now makes the chip that serves its own model. For the region, which is weighing its compute capacity between a Brazilian supercomputer and a Chilean token bank, the question changes shape. It is not whom to buy the hardware from, but what is left to negotiate when whoever sells the model also makes the machine that runs it.


If the cost of serving a model falls by half and whoever captures those savings is the same party that charges for the model, at what point—and by what route—would that drop reach whoever pays for the token in the region?

Correction (September 30, 2026). The original version said Chile’s Law 21.719 had been fully in force since December 1; the correct date is December 1, 2026, when it takes full effect, according to Chile’s Judicial Academy. The original version also said Brazil committed 2.5 billion reais to an open-architecture supercomputer; in fact, the supercomputer costs about 1.06 billion reais, according to Brazil’s National Laboratory for Scientific Computing (LNCC), and the open architecture (RISC-V) refers to the chip strategy announced the same day, not to the supercomputer, according to Agência Brasil.

About this entry. It is generated automatically from public sources, without human review before publication. It may contain errors of interpretation or summary; please check each story against its original source (the links lead there) before citing it or making decisions based on it.

Doble Click is written with Anthropic models.

Spotted an error? Report it

Tell us what's wrong, quoting the sentence if you can and, if you have it, the source that corrects it. An automated process reviews reports every night: if the error is verified, the page is corrected and a correction note is added at the bottom.

Your email is optional: we only use it if we need more context about the report. It doesn't subscribe you to the newsletter.