<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://dobleclick.jaguridi.cl/en/feed-lecturas.xml" rel="self" type="application/atom+xml" /><link href="https://dobleclick.jaguridi.cl/en/" rel="alternate" type="text/html" hreflang="en" /><updated>2026-10-01T09:15:29-03:00</updated><id>https://dobleclick.jaguridi.cl/en/feed-lecturas.xml</id><title type="html">Doble Click · Doble Lectura (English)</title><subtitle>A daily look at artificial intelligence from a Latin American perspective. Curated from public sources. One post a day.</subtitle><author><name>Doble Click</name></author><entry xml:lang="en"><title type="html">The same GPT-4 that improved the work of 758 consultants made them get another task wrong</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/frontera-dentada-ia-consultores/" rel="alternate" type="text/html" title="The same GPT-4 that improved the work of 758 consultants made them get another task wrong" /><published>2026-09-28T00:00:00-03:00</published><updated>2026-09-28T00:00:00-03:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/frontera-dentada-ia-consultores</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/frontera-dentada-ia-consultores/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality</em></li>
  <li><strong>Who:</strong> <a href="https://scholar.google.com/scholar?q=%22Fabrizio+Dell%27Acqua%22">Fabrizio Dell’Acqua</a>, <a href="https://scholar.google.com/scholar?q=%22Edward+McFowland+III%22">Edward McFowland III</a>, <a href="https://scholar.google.com/scholar?q=%22Hila+Lifshitz%22">Hila Lifshitz</a>, and <a href="https://scholar.google.com/scholar?q=%22Karim+R.+Lakhani%22">Karim R. Lakhani</a> (Digital Data Design Institute, Harvard Business School), <a href="https://scholar.google.com/scholar?q=%22Ethan+Mollick%22">Ethan Mollick</a> (Wharton), <a href="https://scholar.google.com/scholar?q=%22Katherine+C.+Kellogg%22">Katherine C. Kellogg</a> (MIT Sloan), and Saran Rajendran, Lisa Krayer, and François Candelon (BCG Henderson Institute).</li>
  <li><strong>Where:</strong> <em>Organization Science</em>, vol. 37, no. 2, March-April 2026, pp. 403-423, open access. <a href="https://doi.org/10.1287/orsc.2025.21838">doi.org/10.1287/orsc.2025.21838</a></li>
  <li><strong>Type:</strong> randomized, preregistered experiment within a single company.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The experiment was carried out at Boston Consulting Group (BCG) with 758 consultants, about 7% of its consultants who do not manage a team. All of them first completed a task without AI, which served as a baseline. Then a random draw split them into three groups: no AI, GPT-4, or GPT-4 plus a brief training on how to write instructions for it. Each person also worked on one of two types of task, designed with the firm’s executives and tested beforehand to fall on one side or the other of what the model handled well. That line is the <em>jagged frontier</em> of the title: tasks that seem equally difficult to a person can fall on different sides.</p>

<p>Inside the frontier, participants had to come up with a shoe for a niche market and take it all the way to launch, across 18 creative, analytical, writing, and persuasion subtasks. Compared with the group without AI, those who used GPT-4:</p>

<ul>
  <li>completed 12.2% more tasks;</li>
  <li>finished them 25.1% faster;</li>
  <li>raised quality by 33.9% if they had received the training;</li>
  <li>and by 29.9% if they had not.</li>
</ul>

<p>Outside the frontier, the task was a business case: recommending to a general manager which brand had the most potential, by cross-referencing a spreadsheet with internal interviews. The spreadsheet looked complete, but the interviews contained details that changed the answer, and the model reached the wrong conclusion. The share of correct answers came out as follows:</p>

<ul>
  <li>without AI: 84.5%;</li>
  <li>with GPT-4: 70.6%;</li>
  <li>with GPT-4 and training: 60%.</li>
</ul>

<p>On average, using AI cut accuracy by 19 percentage points. The trained group, which gained the most inside the frontier, got it right even less often outside it, but the difference between the two AI groups is weak: significant only at the 10% level. Both AI groups also spent less time on this task than the group without AI; the trained group, more than 11 minutes less, or 30% less.</p>

<p>The third finding explains why the error is hard to see. Evaluators who did not know the correct answer scored the coherence and persuasiveness of each recommendation. Those who used AI were rated higher, including when they had gotten it wrong. The authors conclude that AI improves presentation and argumentation even when the analysis is wrong. One additional result: inside the frontier, those who started in the bottom half on the baseline task gained the most, although the top half also improved.</p>

<p>What is established is that the same model, used by the same professionals on tasks from their own line of work, raised performance on some and lowered it on another, and that from inside the work it was not obvious which was which.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>This is the second Doble Lectura in a row about an experiment by a largely shared team, with Dell’Acqua, Lifshitz, Mollick, and Lakhani on both. The previous one showed how much AI adds for a professional; this one shows where it subtracts.</p>

<p>What is most useful for the region lies in two results of different strength. The solid one: outside the frontier, those who used AI got it right less often, and their wrong answers were more convincing to those who evaluated them. The fragile one: the group trained in writing instructions got it right even less often, with a weak difference and after working for less time.</p>

<p>For the schools of government that train public servants, the fragile result allows only one prudent reading: teaching people to ask the model for things better did not protect against error on the task where the model failed. The study did not measure participants’ confidence, so it does not say why the trained group got it right less often. It does suggest that a course should include tasks where the model gets things wrong, not just examples where it shines.</p>

<p>The solid result concerns whoever does the reviewing, for example a committee that scores proposals in a procurement process with a rubric for clarity and justification. If the documents it receives are written with AI, the finding on coherence indicates that such a rubric may reward a well-written but mistaken analysis. Read from the region, the practical consequence is that review has to go back to the source data, and that requires budgeting reviewer hours to redo parts of the analysis, not just to read the document.</p>

<p>A third point remains open. That the consultants with the lowest initial performance gained the most inside the frontier suggests the tool could narrow gaps in public teams with little experience. It is not yet known whether that advantage holds when the same team faces, without warning, a task that falls outside.</p>

<p>The concrete implication for a ministry adopting AI assistants is where to put the effort. This evidence points to mapping which tasks in its own workflow fall outside the frontier and strengthening human review there: training in instructions improved performance inside the frontier, but it did not protect outside it.</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>There is only one task outside the frontier, placed there through prior testing: a limitation the authors acknowledge.</li>
  <li>Inside the frontier, quality was measured with subjective evaluations, something the authors disclose. Replications with footwear experts point in the same direction.</li>
  <li>The model was GPT-4 in its April 2023 version. The authors warn that the frontier is not fixed and that a task can move inside with a new model.</li>
  <li>Participants were early in their careers, with incentives that rewarded quality within the time limit.</li>
  <li>Conflict of interest: three coauthors are listed as affiliated with the BCG Henderson Institute, and the company collected the data. The preregistration did not include the jagged frontier framing.</li>
  <li>It applies to complex professional work with a 2023 model. Where the frontier lies today is something each organization has to measure again.</li>
</ul>

<p><small><strong>Revised on September 28, 2026.</strong> The original version stated that the group trained in writing instructions “made the most mistakes” and that such a course “can raise confidence in the tool.” In the paper, the difference between the two AI groups is significant only at the 10% level, that group spent less time on the task, and confidence was not measured. Those sentences and the title were adjusted.</small></p>]]></content><author><name>Doble Click</name></author><category term="work" /><category term="governance" /><summary type="html"><![CDATA[A preregistered experiment with 758 BCG consultants shows that GPT-4 sped up and improved their work on product development tasks. On a business case that fell outside its reach, the opposite happened: those who used the tool got it right less often, and their wrong answers sounded more convincing.]]></summary></entry><entry xml:lang="en"><title type="html">One person with AI performed like a team of two, and did worse at picking their best idea</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/persona-con-ia-rinde-como-equipo/" rel="alternate" type="text/html" title="One person with AI performed like a team of two, and did worse at picking their best idea" /><published>2026-09-21T00:00:00-03:00</published><updated>2026-09-21T00:00:00-03:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/persona-con-ia-rinde-como-equipo</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/persona-con-ia-rinde-como-equipo/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork</em></li>
  <li><strong>Who:</strong> <a href="https://scholar.google.com/scholar?q=%22Fabrizio+Dell%27Acqua%22">Fabrizio Dell’Acqua</a>, <a href="https://scholar.google.com/scholar?q=%22Raffaella+Sadun%22">Raffaella Sadun</a> and <a href="https://scholar.google.com/scholar?q=%22Karim+R.+Lakhani%22">Karim R. Lakhani</a> (Harvard Business School), <a href="https://scholar.google.com/scholar?q=%22Charles+Ayoubi%22">Charles Ayoubi</a> (ESSEC), <a href="https://scholar.google.com/scholar?q=%22Hila+Lifshitz%22">Hila Lifshitz</a> (Warwick Business School), <a href="https://scholar.google.com/scholar?q=%22Ethan+Mollick%22">Ethan Mollick</a> and <a href="https://scholar.google.com/scholar?q=%22Lilach+Mollick%22">Lilach Mollick</a> (Wharton), plus four Procter &amp; Gamble professionals: Yi Han, Jeff Goldman, Hari Nair and Stew Taub.</li>
  <li><strong>Where:</strong> <em>Organization Science</em>, vol. 37, no. 4, July–August 2026, pp. 1217-1242, open access. <a href="https://doi.org/10.1287/orsc.2025.20702">doi.org/10.1287/orsc.2025.20702</a></li>
  <li><strong>Type:</strong> preregistered field experiment, 2 × 2 design, within a single company.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The study was carried out between May and July 2024 at Procter &amp; Gamble, with 791 professionals from the commercial and research and development areas. Each one spent a full day in a virtual product development workshop, working on real challenges from their own business unit. A random draw assigned them to four conditions: alone without AI, in a cross-functional pair without AI, alone with AI, or in a pair with AI. The tool was built on GPT-4. Expert evaluators, blind to each participant’s condition, scored the solutions, and those scores were standardized against the group that worked alone and without AI.</p>

<p>Quality, measured in standard deviations above that control group, came out as follows:</p>

<ul>
  <li>pair without AI: 0.24</li>
  <li>individual with AI: 0.37</li>
  <li>pair with AI: 0.39</li>
</ul>

<p>That is the first finding. The individual with AI reached the level of the human pair, and adding AI to the pair barely moved the average beyond that.</p>

<p>The second finding is about expertise. Without AI, people from the commercial area proposed commercial ideas and those from research and development proposed technical ideas, with clearly different distributions. With AI that distinction fades and both groups generate a similar mix, without quality varying significantly according to how technical the solution was. The most striking result came among those who do not usually work in product development: alone and with AI, they reached the level of teams that included someone who does it every day.</p>

<p>The third finding is the hardest to digest. Since everyone generated five ideas before choosing one and developing it, the authors separated the two stages. AI raised the average quality of the ideas across the entire distribution, without narrowing the distance between the best and the worst. But when choosing which of the five to develop, those who worked without AI seem to have been more accurate: pairs without AI kept their best idea about half the time, and the AI conditions around 37%. Even so, since they started from better ideas, the ideas chosen by those who had AI were still of higher quality. The authors offer possible explanations, among them the tendency of models to validate what the user already brings, warn that perhaps participants did not use AI to choose, and leave the point open. They describe AI as a quality amplifier more than as a better decision-maker. There are two other results: pairs with AI placed more solutions in the top decile, and those who used AI reported more positive and fewer negative emotions at the end.</p>

<p>What is established in this context is an asymmetry between stages. AI raised the floor of idea generation to the point of matching what a second professional contributed, and it did not improve the next step, which is choosing well.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>The experiment covered four business units in two geographies, Europe and the Americas, and the paper does not break down results by country.</p>

<p>Our reading is that the result with the most direct translation to the region is the substitution one. For a mid-sized company in Chile or Colombia that does not have a research and development area separate from the commercial one, the bottleneck is not a lack of ideas but that there is no one to cross them with: adding a second specialized professional costs a salary, and what the study participants received was a license and an hour of training. The finding that the individual with AI reached the level of the cross-functional pair speaks precisely to that margin.</p>

<p>The finding about the selection stage changes something concrete in how procurement is done. An innovation unit in a ministry in Chile or Brazil that is drafting the terms of reference today to contract an AI assistant can specify it for generating and writing drafts, and put in writing that the choice between alternatives remains a documented human decision, with time allotted for it. That is not bureaucratic zeal: the only stage where the group without AI came out ahead was precisely identifying its own best idea.</p>

<p>A third point remains a hypothesis. In the region it is common for a small public team to have a single person in charge of a topic, with no specialized colleague to consult. That the employees least familiar with the task reached, with the tool, the level of teams that included someone experienced suggests that this is a profile where it performs especially well. Extrapolating it to the region’s public sector is our bet, not the authors’.</p>

<p>What is concrete for the region is a criterion for where to put the tool first. Where there is no one to test an idea against, this evidence says that a license and some training move the needle. Where the problem is choosing well among several options that already exist, the evidence offers no support, and something points against it.</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>The preregistration covered performance and expertise. Emotions were included as a variable of interest with unclear effects, and the analysis of the upper tail emerged during the research: it is exploratory.</li>
  <li>Scope as stated by the authors: a consumer goods company, a single virtual day, pairs formed at random among people who generally did not know each other. They compare them to flash teams, not to established teams.</li>
  <li>Participants had little experience with the tool, and the authors present the benefits as a floor, not a ceiling.</li>
  <li>Conflict of interest: four coauthors work at Procter &amp; Gamble and the design was agreed with its leadership. It is evidence from inside the organization studied, with preregistration and blind evaluators.</li>
  <li>That is as far as it goes: one company, one early ideation task, one model. What happens when use is sustained over time remains an open question for the authors themselves.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="work" /><category term="design" /><summary type="html"><![CDATA[A preregistered field experiment with 791 Procter & Gamble professionals shows that working alone with AI matches the quality of what a cross-functional pair produces without it. The twist is in the next step: those who used AI seem to have been less accurate in choosing which of their own ideas to develop.]]></summary></entry><entry xml:lang="en"><title type="html">They asked the models about the risk of catastrophe, and they answer higher than humans do</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/ia-pronostica-su-propio-riesgo/" rel="alternate" type="text/html" title="They asked the models about the risk of catastrophe, and they answer higher than humans do" /><published>2026-09-14T00:00:00-03:00</published><updated>2026-09-14T00:00:00-03:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/ia-pronostica-su-propio-riesgo</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/ia-pronostica-su-propio-riesgo/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Automated Forecasts of Catastrophic Risks</em></li>
  <li><strong>Who:</strong> <a href="https://jabaluck.github.io/">Jason Abaluck</a> (Yale and NBER), <a href="https://scholar.google.com/scholar?q=%22Ezra+Karger%22+forecasting">Ezra Karger</a> (Federal Reserve Bank of Chicago), <a href="https://scholar.google.com/scholar?q=%22Nick+Merrill%22+forecasting">Nick Merrill</a> (Forecasting Research Institute and UC Berkeley), <a href="https://scholar.google.com/scholar?q=%22Philip+E.+Tetlock%22+forecasting">Philip E. Tetlock</a> (University of Pennsylvania), <a href="https://scholar.google.com/scholar?q=%22Eva+Vivalt%22">Eva Vivalt</a> (University of Toronto) and <a href="https://scholar.google.com/scholar?q=%22Bridget+Williams%22+%22forecasting+research+institute%22">Bridget Williams</a> (Forecasting Research Institute and Oxford). They are listed in alphabetical order.</li>
  <li><strong>Where:</strong> Forecasting Research Institute working paper, September 2026, no DOI (<a href="https://forecastingresearch.org/pdf/airo-working-paper.pdf">PDF</a>). Funded by Coefficient Giving.</li>
  <li><strong>Type:</strong> elicitation of forecasts from language models, with three validation exercises.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The starting point is a well-known and rather sterile debate. On the risk that AI ends in catastrophe, public estimates range from 10 to 25% from some industry leaders down to less than 0.001% from Yann LeCun. The authors do not try to settle it with more expert opinion. They propose adding a new input: asking the models themselves.</p>

<p>That is AIRO, the <em>Automated AI Risk Outlook</em>, a panel that is updated periodically. The mechanics are concrete. They take the four top-ranked models on the Epoch Capabilities Index, skipping those from the same family: at the time of writing they are GPT-6 Astra, Fable 5.1, Opus 5 and GPT-5.5 Pro. Each one is given, in a single prompt, 35 questions with their resolution criteria, the horizons and fourteen conditions under which to answer each cell. Before it can submit a forecast, the model is required to do at least ten web searches or reads. The ensemble forecast is the simple median of the four.</p>

<p>The central question defines catastrophe as an event that kills at least 10% of the human population within five years. The figures: from any cause, 1.1% by 2030, 8.5% by 2050 and 18.5% by 2100. Caused by AI, 0.47%, 6.0% and 12.2%. In other words, the median for AI catastrophe is equivalent to about 45% of the median for any cause by 2030 and 71% by 2050. There is a higher and less discussed figure: the median for human disempowerment exceeds that of deadly catastrophe at every horizon, at 15.5% by 2050 and 28% by 2100.</p>

<p>Those numbers sit well above the humans’. In the comparison the authors themselves make, the median expert on the LEAP panel gave 0.3% by 2030 and 2% by 2050, and the median superforecaster 0.1% and 0.88%. The model ensemble comes in between 4.75 and 6.82 times higher than the superforecasters on the AI catastrophe question.</p>

<p>The obvious objection is that a model throwing out probabilities is worthless if we do not know whether it gets things right. The paper devotes its most interesting part to that, with three tests. On ForecastBench, using real questions that have already been resolved, recent models appear to show less favorite-longshot bias than the 2024 superforecasters, and similar Brier scores. The bias is the tendency to overstate the improbable and to fall short on the very probable. Since rare events barely exist in the historical record, they build rare events to order: two simulators, a Civilization-style world with almost 26,000 binary questions whose true probability is between 1 and 9%, and an epidemic simulator. There, model capability strongly predicts accuracy, with a Spearman correlation of +0.85. And for forecasts conditional on a policy, they test with a simulated vaccination campaign of known effect: frontier models recover a median of 0.78 of the recoverable skill.</p>

<p>From there comes the idea that organizes the paper. If forecasting better is a function of capability, and capability is also what drives risk, then these forecasts are most informative precisely in the future worlds where the risk is highest. The authors say so, and they also say what that evidence does not prove.</p>

<p>The policy exercise uses eight scenarios adapted from a survey by the same institute. The package combining an international compute cap, international pre-launch authorization and strict liability lowers the AI catastrophe forecast by 66% by 2050. The US-only versions perform considerably worse than the international ones. Two conditions raise it: the status quo with no new policy, 1.28 times, and federal preemption of state laws, 1.24 times. That the status quo is worse than the unconditional forecast means the models are already assuming that something will be regulated. The authors warn that they did not measure the cost of these policies in terms of innovation or growth: what they deliver is an effectiveness ranking, not a cost-benefit analysis.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>It is worth saying plainly: the region does not appear. The policy menu was put together from proposals by US think tanks and politicians, the human comparison panels come from the same house, and the scenario that moves the needle up the most, federal preemption, is an internal US debate. Nothing here was asked from the standpoint of a country that receives the technology without building it.</p>

<p>And yet there is a result that speaks directly to the region, without anyone having looked for it. In every comparison, international measures beat national ones, and the package beats any single measure. For a country that does not train frontier models and has no jurisdiction over those who train them, that is not a technical detail. It is the argument for why national AI policy, however good, does not touch this particular risk, and why the place where something is actually at stake is the multilateral table. Being there stops being decorative diplomacy.</p>

<p>The second possible contribution is more modest and perhaps more useful. A ministry in the region has no way to set up its own catastrophic risk assessment, and a public, up-to-date panel is a good it can access for free. The risk of using it is importing the figure without its conditions. The paper itself shows how much everything shifts: conditioning on the 90th percentile of future capability that the model itself estimates multiplies the 2030 forecast by 2.21. An AIRO number without its condition is not information, it is a quote out of context.</p>

<p>A reading of my own, which the paper does not make: the incident ladder suggests that what will reach the region first is not the catastrophe rung but the lower ones, where cyber incidents are in the lead. The authors do not break down their results by geography, so I leave it as a hypothesis. But a country whose real exposure runs through digital public services and financial systems has more to do with the medians of the lower rungs than with the big headline figure.</p>

<p>The question that remains is who watches these dashboards in the region, and with what mandate. If a figure moves sharply upward next quarter, is anyone here in charge of noticing?</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>A conflict of interest in plain sight: the paper presents a product of its own house. AIRO belongs to the Forecasting Research Institute, the questions are reused from previous studies by the same institute and the human comparison panels are also its own. It is internal validation with transparent criteria, not an independent audit.</li>
  <li>The authors are explicit about what they do not know: it is not established that the accuracy measured on ForecastBench or in simulated worlds carries over to forecasting real catastrophes. They say those proxies give reasons to investigate the forecasts, not to conclude that the models match the best humans on this question. On the comparisons with humans they warn of something similar: they are forecasts made on different dates and, on ForecastBench, on different sets of questions, so the ratios above do not measure relative accuracy.</li>
  <li>The dashboard runs only on public information. The labs have private information about capabilities and risks, and the authors acknowledge that without it one can hardly warn about what is happening inside those companies.</li>
  <li>The paper points out an unusual and honest problem: publishing these forecasts puts them into the training corpus of future models, which could end up reinforcing them. And they warn that companies or models could collude or hide data if policy responses depend on these figures.</li>
  <li>That 100% of the 1,996 implicit orderings hold speaks to the panel’s internal coherence, not to whether it is getting things right.</li>
  <li>Transparency: two of the four models on the panel are from the Claude family, the same one that writes these readings.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="governance" /><category term="safety" /><summary type="html"><![CDATA[A panel of four frontier models forecasts a 0.47% probability that AI causes a catastrophe killing at least 10% of humanity by 2030, and 6% by 2050, between 4.8 and 6.8 times what superforecasters estimate. The paper's twist: those forecasts become more informative precisely as the risk rises.]]></summary></entry><entry xml:lang="en"><title type="html">A mental well-being chatbot: supplement, medicine, or yoga instructor?</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/ia-bienestar-mental-suplemento-o-remedio/" rel="alternate" type="text/html" title="A mental well-being chatbot: supplement, medicine, or yoga instructor?" /><published>2026-09-07T00:00:00-03:00</published><updated>2026-09-07T00:00:00-03:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/ia-bienestar-mental-suplemento-o-remedio</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/ia-bienestar-mental-suplemento-o-remedio/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Framing Responsible Design of AI for Mental Well-Being: AI as Primary Care, Nutritional Supplement, or Yoga Instructor?</em></li>
  <li><strong>Who:</strong> <a href="https://scholar.google.com/scholar?q=%22Ned+Cooper%22+%22human-computer+interaction%22">Ned Cooper</a>, <a href="https://jaguridi.github.io/">Jose A. Guridi</a>, and <a href="https://qianyang.co/">Qian Yang</a> (Cornell University), <a href="https://scholar.google.com/scholar?q=%22Angel+Hsing-Chi+Hwang%22">Angel Hsing-Chi Hwang</a> (University of Southern California), <a href="https://scholar.google.com/scholar?q=%22Beth+Kolko%22+%22human+centered+design%22">Beth Kolko</a> (University of Washington), and <a href="https://scholar.google.com/scholar?q=%22Emma+E.+McGinty%22+%22health+policy%22">Emma Elizabeth McGinty</a> (Weill Cornell Medicine)</li>
  <li><strong>Where:</strong> <em>Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems</em>, Barcelona, April 2026. <a href="https://doi.org/10.1145/3772318.3791556">doi.org/10.1145/3772318.3791556</a></li>
  <li><strong>Type:</strong> three-stage qualitative study: 24 expert interviewees and analysis of more than 100 regulatory documents.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The subject of the study is <strong>non-clinical</strong> tools: ChatGPT, Replika, and the like, used to vent or feel better, without a prescription or supervision by a professional. US law draws a sharp line here. The clinical side is regulated by the FDA as a medical device; the non-clinical side is overseen by the FTC with a light hand, as consumer technology. Millions of people are already on the light side.</p>

<p>That gives rise to the question, which is about design, not measurement: what does it mean, concretely, to design one of these tools responsibly. The authors do not evaluate any chatbot or measure any effect. They ask experts in medical ethics, health policy, AI regulation, and health technology design; they analyze more than a hundred public policy documents; and they go back to the interviewees to discuss what they found. The first interviews nearly failed, and that is part of the finding: the health policy experts did not see why they were being interviewed, and the technology policy experts shared the concern without offering anything actionable. The only thing that resonated with everyone was an offhand analogy: some of these tools resemble a nutritional supplement, and others an over-the-counter medicine.</p>

<p>The paper turns that intuition into a two-axis map: whether the tool is a product or a service, and whether or not it guarantees a health outcome. That leaves four boxes. The supplement is a product that guarantees nothing. The over-the-counter medicine guarantees relief for a defined symptom. Primary care is a service that guarantees the outcome, and if it cannot deliver it, it is obligated to refer. The yoga instructor is a service without a guarantee: their instruction can enhance or ruin the proven benefits of yoga, and even so they promise none.</p>

<p>What is interesting is not the idea itself but what it organizes: each box carries different primary risks and therefore different responsibilities. In a medicine-type tool, the urgent concerns are safety, effectiveness, and equitable access, and several interviewees were explicit that something that does not work equally for all groups cannot be called safe and effective. In a supplement-type tool it is almost the opposite: its clinical effectiveness matters little, and what matters is that it not replace clinical care or self-care. As one interviewee put it, it is not a problem for someone to talk to a therapy chatbot, whether it is effective or not; it is a problem when that person should be talking to a psychiatrist.</p>

<p>The second finding is the one with the most bite. All the interviewees, in different words, put <strong>active ingredients</strong> at the center: the proven mechanism by which a tool improves well-being. Clinicians and health policy experts called it that, ethicists spoke of a theory of change, industry of the product’s real essence. Most considered that declaring them, and also guaranteeing that they are delivered well, is a condition of responsible design. With that, the authors distinguish three types: those validated as a whole through controlled trials, extremely rare; those that deliver a proven ingredient such as cognitive behavioral therapy without being validated themselves; and those that articulate no ingredient at all, such as out-of-the-box ChatGPT used to de-stress. No interviewee described this last category as responsible design, and the parallel that comes up several times is social media, which also did not understand how it was entertaining people and discovered too late that the engine was polarization and rage.</p>

<p>The third finding the authors do not resolve; they leave it on the table: where the line between risk and benefit lies. Most of the interviewees from medicine, health policy, and health technology accepted the reasoning applied to innovative drugs, where a drug that saves many is approved even if it is lethal for a few, as long as the risk is disclosed. Those from ethics and design strongly resisted that population arithmetic, and the split largely followed disciplinary lines. Where there was no nuance was among clinicians: several insisted that asking whether someone has suicidal thoughts is not enough, and described a real risk assessment and a guaranteed referral as non-negotiable.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>The caveat first: the regulatory analysis is about the United States, and the authors say so. The FDA, the FTC, and the definition of primary care used by Medicare and Medicaid are the material the four analogies are made of, and none of that can be imported as is. But it is worth separating two layers. The legal one does not travel. The design one does, and it is what the paper really proposes: a vocabulary so that whoever builds the tool declares what it promises and to whom, before there is a law requiring it. Where regulation is in its infancy, that vocabulary arrives just when it is useful.</p>

<p>Where the framework comes under strain is referral. What separates primary care from the yoga instructor is the obligation to refer effectively when the tool cannot resolve the problem, and that assumes there is someone to refer to. In much of Latin America the specialist is not there, or is months away on a waiting list, so the criterion of the clinicians interviewed becomes more demanding than it sounds: a tool that refers into a void has not done its job. And asking that it not substitute for clinical care assumes that such care is an available alternative. For many people here, the free, Spanish-language chatbot does not compete with the psychologist; it competes with nothing, and discouraging its use stops being mere caution. This is my own reading: the paper did not study the region, and it would be unfair to ask it for an answer to that version of the dilemma.</p>

<p>Something similar happens with active ingredients. Cognitive behavioral therapy and the like were validated mostly in other populations and in another language, and what delivers them here is a model trained mostly in English. Declaring the ingredient is the floor, not the ceiling. This extrapolation is also mine, although it points toward where the authors are already looking when they propose a database of active ingredients that documents how well each mechanism works in different populations.</p>

<p>The question it leaves applies to any team building something like this in the region. If you had to write on the front page of your tool what it improves, in whom, and through what proven mechanism, could you? If the answer is that it works for everything and everyone, the paper suggests that is not versatility. It is the absence of anyone accountable.</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>Disclosure: the author of this blog is a coauthor of the paper. The authors also state that they themselves design and research non-clinical tools of this kind, so the framework they propose applies to them too.</li>
  <li>It is a qualitative study: its value lies in offering a framework and making disagreements explicit, not in measuring how many experts think what. The regulatory analysis covers only the United States, and the authors justify this as a choice of depth over breadth.</li>
  <li>They did not interview users or patients, and they explain why: they wanted to look at risks that are not yet observable, for which the subjective experience of use is of no help. Several industry interviewees come from startups, and the authors themselves note that experts from large hospitals and insurers were less accessible.</li>
  <li>The paper does not propose a validated standard, nor does it claim to. It leaves open two questions its interviewees could not settle: whether these tools should be evaluated by their population-level effect, and whether a risk proportional to the benefit is a sufficient yardstick.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="health" /><category term="design" /><category term="ethics" /><summary type="html"><![CDATA[Twenty-four experts in clinical practice, ethics, public policy, and health technology design, more than a hundred regulatory documents, and a conclusion that is uncomfortable for the industry: what determines whether a mental well-being AI is well designed is not how good it is at conversation, but what concrete benefit it promises and to whom. A tool meant to serve everyone answers to no one.]]></summary></entry><entry xml:lang="en"><title type="html">With AI, homework grades go up and test scores go down</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/penalidad-aprendizaje-ia-china/" rel="alternate" type="text/html" title="With AI, homework grades go up and test scores go down" /><published>2026-08-31T00:00:00-04:00</published><updated>2026-08-31T00:00:00-04:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/penalidad-aprendizaje-ia-china</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/penalidad-aprendizaje-ia-china/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>The Generative AI Learning Penalty: Evidence from Chinese Secondary Education</em></li>
  <li><strong>Who:</strong> <a href="https://scholar.google.com/scholar?q=%22David+Str%C3%B6mberg%22+economics">David Strömberg</a> (Stockholm University), <a href="https://scholar.google.com/scholar?q=%22Victor+Lei%22+%22University+of+Hong+Kong%22">Victor Lei</a> and <a href="https://scholar.google.com/scholar?q=%22Yanhui+Wu%22+%22University+of+Hong+Kong%22">Yanhui Wu</a> (University of Hong Kong)</li>
  <li><strong>Where:</strong> SSRN, June 2026. Working paper, not yet peer-reviewed. <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6977138">papers.ssrn.com/abstract=6977138</a></li>
  <li><strong>Type:</strong> quantitative study. A 30-month administrative panel, 26,811 students, difference-in-differences with staggered adoption.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>It helps to start with the design, because almost all of the paper’s strength comes from it. The authors obtained from the education office of a county in central China, which they describe as representative of counties outside the developed coast, the records of 26,811 students from seventh through twelfth grade, 90 percent of those enrolled in secondary school there. The data run from September 2022 to June 2025 and combine three things that rarely appear in the same dataset: scores on monthly closed-book tests, the grade and submission time for weekly homework in nine subjects, and the two high-stakes national exams, the Zhongkao, which sorts students into different upper secondary schools, and the Gaokao, which in practice decides university admission.</p>

<p>A June 2025 survey asked in what month each student began using generative AI. Adoption went from nearly zero in 2022 to about 80 percent in 2025, and it advanced in a staggered way, which makes the design possible: comparing how outcomes change for those who adopted in a given month against those who never adopted. The most used tools were Doubao, DeepSeek, ChatGLM, Ernie Bot and Qwen, not specialized educational apps.</p>

<p>The central finding is a clean divergence between productivity and learning. Six months after adopting AI, homework grades rise 18 percent relative to the baseline mean and submission time falls from 64 to 45 minutes. At the same time, scores on monthly closed-book tests fall 20 percent. The two curves separate exactly in the month of adoption and had been moving together before, which is what the design needs.</p>

<p>The second finding is about time. The effect on monthly tests is complete in about six months; the effect on admission exams takes about two years to reach its full magnitude, 18 percent on the Gaokao and 24 percent on the Zhongkao, because those exams cover material from several years. The authors themselves draw the methodological consequence, and it is uncomfortable for the rest of the literature: short studies, which are almost all of them, underestimate the long-term cost.</p>

<p>The third finding explains the mechanism and avoids the crude conclusion that AI does harm on its own. The loss is concentrated among users whose pattern is consistent with outsourcing the homework, who make up 58 percent of AI users and rise to 81 percent among those who have been using it for more than five months: they submit in less time than even the fastest students without AI, get homework grades that match what these tools get right on that type of exercise, and perform very poorly on tests. By contrast, AI users who spend the same time on homework as non-users get test scores similar to theirs, with better homework grades, and they are not better students to begin with. The authors’ conclusion is precise: AI reduces the time devoted to learning for most students, but it does not reduce the efficiency with which those who maintain their time learn.</p>

<p>Broken down by subject, the largest drops are in social sciences (27 percent on average), then STEM (22 percent) and finally languages (English 17 percent, Chinese 9 percent). It is worth noting, because previous experiments are almost all concentrated in math, programming and languages, and none of those is the hardest-hit subject here: math falls 22 percent and languages less than that. By profile, the drop is larger among younger students, among boys and among those with better initial performance, which compresses the distribution of skills through a different path from the familiar one: not because AI lifts those at the bottom, but because it penalizes those at the top more.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>What makes this paper translatable is that it does not study a tool, it studies an incentive. The authors say so at the end: the literature is focused on how to design AI tutors that work well, and in China those tutors already exist and are cheap or free. Students still choose the general-purpose chatbot that delivers the answer directly. What is missing is not a good tool, it is a reason to prefer it.</p>

<p>From there comes the finding that most resembles the region, and it is a measurement problem before it is a technology problem. Among AI users with above-average homework grades, better homework grades are associated with worse test scores. Homework stops providing information about learning, and the gap only becomes visible in a closed-book assessment. That matters where classroom grades weigh in school marks and those marks feed into university admission. Chile with the NEM and the grade ranking, Brazil with the ENEM, Mexico with its entrance exams: the architecture of a high-stakes external exam coexisting with classroom grades that are now easy to outsource exists across the region. This last point is my reading and not the paper’s, which studied one Chinese county and nothing more.</p>

<p>A clarification from the authors is worth not skipping: the time saved on homework is 2.2 to 2.8 hours per week, between 5 and 6 percent of total study time, much less than the 20 percent drop in scores. They suggest that other factors must be contributing, probably that those who outsource the weekend homework also outsource everything else.</p>

<p>There is an optimistic note that they themselves present as suggestive rather than established. Measured at five months of use, the penalty went from around 25 percent in early 2023 to around 16 percent in June 2025, and the pattern holds when the sample is fixed to early adopters. Something is adapting, among students or teachers, although the loss is far from disappearing.</p>

<p>Their three policy suggestions are modest and cheap, which is what makes them relevant for ministries without a budget for big reforms. Give students credible information about the learning cost of outsourcing homework, because today they do not perceive it. Increase the weight of in-person, closed-book assessments. And have teachers and families monitor inputs, study time and effort, instead of outputs, which is what AI has made uninformative.</p>

<p>The question it leaves is a direct one for any ministry in the region. If the homework grade no longer tells you whether the student learned, what is being used to assess them?</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>It is a working paper on SSRN, not yet peer-reviewed.</li>
  <li>The month of adoption is self-reported and retrospective. The authors discuss this: reporting that was too late would have left pre-trends, which do not appear; reporting that was too early would explain part of the six-month ramp, but not the final magnitude.</li>
  <li>Identification is weaker for the admission exams, which are observed once or twice and where pre-trends cannot be shown. The authors acknowledge this and support their causal reading on the fact that the ordering of effects by subject and by profile is almost identical for both types of exam.</li>
  <li>Homework time is the interval between opening and submitting on the platform, not actual working time. The authors say so and treat it as a reasonable approximation, but it is the variable on which the entire outsourcing mechanism rests.</li>
  <li>The relationship between homework time and test scores is not causal, and they say so explicitly: it serves to characterize who does worse, not to conclude that requiring more time would restore learning.</li>
  <li>It is one county in China. The mechanisms travel better than the magnitudes.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="education" /><category term="data" /><summary type="html"><![CDATA[Thirty months of records from 26,811 secondary school students in a Chinese county. After adopting generative AI, homework grades rise 18% and the time spent on it falls by a third, but closed-book test scores fall 20% within six months, and the effect on admission exams takes two years to appear in full.]]></summary></entry><entry xml:lang="en"><title type="html">Almost half of real-world AI use disappears when it is measured only as work</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/observatorio-uso-real-ia/" rel="alternate" type="text/html" title="Almost half of real-world AI use disappears when it is measured only as work" /><published>2026-08-24T00:00:00-04:00</published><updated>2026-08-24T00:00:00-04:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/observatorio-uso-real-ia</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/observatorio-uso-real-ia/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>The AI Observatory: A Public Measure of Real-World AI Use</em></li>
  <li><strong>Who:</strong> nineteen researchers from MIT, Stanford, Northeastern, Johns Hopkins, Berkeley, Carnegie Mellon, Brown, NYU, Maryland, Waterloo, EleutherAI, Cohere, Code Metal and Adaption Labs. First authorship is shared between <a href="https://scholar.google.com/scholar?q=%22Shayne+Longpre%22">Shayne Longpre</a> (MIT), <a href="https://scholar.google.com/scholar?q=%22Anka+Reuel%22">Anka Reuel</a> (Stanford) and <a href="https://scholar.google.com/scholar?q=%22Dayeon+Ki%22+Maryland">Dayeon Ki</a> (Maryland). Among those who guided the work are <a href="https://scholar.google.com/scholar?q=%22Alex+Pentland%22">Sandy Pentland</a> (MIT), <a href="https://scholar.google.com/scholar?q=%22Sara+Hooker%22">Sara Hooker</a> (Adaption Labs) and <a href="https://scholar.google.com/scholar?q=%22Sanmi+Koyejo%22">Sanmi Koyejo</a> (Stanford).</li>
  <li><strong>Where:</strong> preprint, August 2026. Not peer-reviewed at the time of this reading. Taxonomy, annotations and annotation tools at <a href="https://ai-observatory.org">ai-observatory.org</a></li>
  <li><strong>Type:</strong> measurement study. 23,158 conversations and 85,633 annotated turns, drawn from seven real-use sources collected between April 2023 and July 2025, classified with a common taxonomy of 145 traits.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The first thing is to understand what kind of work this is. It does not measure whether AI is useful or estimate effects: it builds a measurement infrastructure and then uses it to test how fragile the claims circulating about “what AI is used for” are. The authors pooled seven collections of real conversations (WildChat, ShareGPT, AI Archive, a scrape of public Grok conversations, LMSYS-Chat-1M, Chatbot Arena and the National Internet Observatory) and ran them through a single taxonomy of 145 traits, which labels each conversation at four levels: the prompt, the response, the turn and the full conversation. That taxonomy covers function, topic, sensitive uses, interaction style, multi-turn dynamics and structure. The annotation is done by a model, GPT-4.1, calibrated against a human validation set.</p>

<p>The first finding is that the sources are not interchangeable, by a wide margin. On Grok, 67.1% of conversations include information seeking, versus 26.2% on WildChat. ShareGPT leans toward content generation (63.6%) and AI Archive toward information analysis (53.9%). Topics diverge in the same way: Grok concentrates news and current events (38.5%) and business and society (64.5%), far above the rest. And the structure does not match either: WildChat prompts average 569.5 tokens versus a range of 51.8 to 181.0 in the others, and Grok responses average 1,322.9 tokens.</p>

<p>That matters above all for risks. Academic integrity problems, for example assignments that were probably copied, range from 23.1% on WildChat to 40.4% on AI Archive. Misinformation is concentrated on Grok (15.3%, versus a range of 4.2% to 10.2% in the rest). The authors are careful with the interpretation: since all the sources are opt-in, they cannot attribute any difference to a specific cause, whether the platform, the model or the period. What is established is the magnitude. Which source the data come from substantially changes the picture of use.</p>

<p>The second finding is the one that gives the title, and it is the most uncomfortable. The Anthropic Economic Index, in its March 2025 version, first filters the conversations relevant to some occupation and only then maps the tasks. The authors reimplemented that filter from the prompts and the taxonomy Anthropic published, and ran it on six of their seven sources: the National Internet Observatory is left out because its data agreement only allows extracting aggregate annotations agreed upon in advance, so the pipeline cannot be run there. 47.9% of conversations are classified as non-occupational, with a floor of 34.2% on AI Archive and a ceiling of 61.9% on LMSYS. Then they looked at what the filter discards, holding source and annotation fixed. What gets discarded is not random residue: it is much more likely to involve health and relationships (44.2% versus 31.2%), adult or illicit topics (7.9% versus 2.1%), harassment or hate (27.5% versus 5.6%) and sexual content (15.7% versus 2.4%). The conclusion they draw is about method, not an accusation: the filter is not a neutral preprocessing step, and a framework can look complete while systematically leaving out socially important uses. They themselves clarify two things that are worth not skipping. That 47.9% is not an estimate of Claude.ai traffic, and later versions of the Index no longer apply the occupational filter and find similar distributions.</p>

<p>The third is that use shifts. Between April 2023 and July 2025, within WildChat, prompts grew 1,049.5% in average tokens, responses 100.6% and turns 17.1%. And variants from the same developer sustain distinct usage regimes: short, template-like exchanges with GPT-3.5, longer and iterative assistance with GPT-4o, and long, single-pass technical problem-solving with reasoning models such as o1.</p>

<p>The fourth looks at people rather than averages. Among users who return to WildChat, the variety of uses narrows over time: distinct function labels drop from 4 to 3 and sensitive-use labels from 2 to 1. In other words, the aggregate expansion of conversations does not come from each user writing more, but from a change in who is using the tool.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>The region does not appear in this work, and the authors say so: among the limitations they state that regions where the Global South predominates remain largely out of reach, even if somewhat represented, and they propose as a remedy integrating consented donation studies, regional sampling and provider-side aggregates. The taxonomy detects 72 languages, but the paper does not report any breakdown by language in its main body.</p>

<p>Even so, the central result translates directly. When the region discusses what to do about AI, the figures cited almost always come from two or three reports by the companies themselves. What this paper adds is a warning about how they are read: the choice of source and the choice of filter are not technical details, they are decisions that change the result before the analysis even begins. An AI policy for work based on an occupational framework is not measuring badly, it is measuring a part, and that part leaves out precisely what would fall to health, education or data protection.</p>

<p>There is a difference in position worth pointing out, and it is mine, not the paper’s. The United States and Europe can offset the opacity of corporate reports with their own measurements: surveys, instrumented observatories, negotiated access to data. A mid-sized country in the region almost never has any of those three things, so it depends more on someone else’s report and has less to check it against. What this work shows is that such a check does not require privileged access to anyone’s servers: it requires conversations donated with consent, an explicit taxonomy and money for annotation. The total annotation cost they report is $5,680. For a public agency or a university center in the region, that number is not the barrier.</p>

<p>That leaves a practical question. The taxonomy and the annotation tools are published and extensible, and the paper itself acknowledges that its coverage of the Global South is weak. Who in the region is going to contribute the conversations in Spanish and Portuguese that today are not in any observatory?</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>None of the seven sources is representative of AI use. All are opt-in, and the authors themselves suspect that sensitive uses are underrepresented, because people do not publicly share that kind of conversation.</li>
  <li>The labels are assigned by a model. Agreement with the human consensus has a median F1 of 0.856 per parent category, and that is why all comparisons in the main body are made at that level. The authors point out that “sensitive uses” is the least stable family, and ask that those prevalences be read as approximate.</li>
  <li>The validation set was produced by the same authors who designed the taxonomy and chose the pipeline. They say so: that shows the pipeline’s fidelity to its own scheme, not the validity of the scheme.</li>
  <li>The temporal analysis and the user-profile analysis rely only on WildChat, the only source with multi-year timestamps and stable identifiers.</li>
  <li>It is a preprint without peer review. Several authors are affiliated with AI labs, and the work compares its measurement against a proprietary report from another lab; they themselves warn that the differences may be due to a mix of product, period and source changes.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="data" /><category term="governance" /><category term="safety" /><summary type="html"><![CDATA[An academic consortium pooled seven sources of real conversations with AI assistants and annotated them with a single taxonomy. When the occupational filter of the Anthropic Economic Index is applied to them, 48% of the conversations are discarded, and what gets discarded is not noise: that is where health, relationships and much of the sensitive content are concentrated.]]></summary></entry><entry xml:lang="en"><title type="html">AI books sell little each, but they lower what every title earns</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/ia-diluye-mercado-libros/" rel="alternate" type="text/html" title="AI books sell little each, but they lower what every title earns" /><published>2026-08-17T00:00:00-04:00</published><updated>2026-08-17T00:00:00-04:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/ia-diluye-mercado-libros</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/ia-diluye-mercado-libros/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Generative AI floods and dilutes the market for books</em></li>
  <li><strong>Who:</strong> <a href="https://scholar.google.com/scholar?q=%22Tuhin+Chakrabarty%22">Tuhin Chakrabarty</a> and <a href="https://scholar.google.com/scholar?q=%22Xinyue+Liu%22+Stony+Brook">Xinyue Liu</a> (Stony Brook University), <a href="https://scholar.google.com/scholar?q=%22Jane+C.+Ginsburg%22">Jane C. Ginsburg</a> (Columbia Law School), and <a href="https://scholar.google.com/scholar?q=%22Paramveer+Dhillon%22">Paramveer Dhillon</a> (University of Michigan and MIT Initiative on the Digital Economy).</li>
  <li><strong>Where:</strong> preprint on arXiv, version of July 26, 2026. Not peer reviewed as of this reading. <a href="https://doi.org/10.48550/arXiv.2607.20349">doi.org/10.48550/arXiv.2607.20349</a></li>
  <li><strong>Type:</strong> observational market study. 14,419 self-published genre fiction e-books on Amazon between January 2023 and March 2026, with daily sales through June 2026 and AI detection on the full text of each book.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The starting point is a widespread belief: books written with AI are <em>slop</em>, cheap text that readers will ignore, so it does not matter how many there are. The authors separate two things that argument lumps together. That average quality is low does not mean the effect on the market is zero, because a book that costs almost nothing to produce can be produced by the hundreds, and at that volume even a mediocre product absorbs attention and sales that would have gone elsewhere.</p>

<p>To measure it, they assembled an unusual panel. Daily sales, price, and ranking come from a proprietary database belonging to one of the Big Five publishers, which covers around 95% of the daily volume of the e-book market. From there they drew a sample stratified by genre and date, chosen blind to content, and classified each book with the Pangram detector, chapter by chapter. The score is the percentage of text windows flagged as not written by humans, and with it they build three bands: no AI detected, light AI (up to 25%), and substantial AI (above 25%).</p>

<p>The first finding deflates the panic. Books with substantial AI appear in all eight genre groups, but they sell poorly: they are 20% of the sample and just 12.1% of sales, while those with no AI detected are 62.9% of titles and take 71.7%. The second points in the opposite direction: even so, they did not stay at the bottom. Their share of sales went from almost zero in early 2023 to about 20% by mid-2026, and among new titles entering the Top 25 the authors construct, the share with substantial AI rose from almost zero to 31%.</p>

<p>The third is the one that gives the paper its title. Between early 2023 and early 2026, the cumulative catalog of published titles grew 38.3-fold and the number of books selling anything in a quarter grew 19.2-fold, while units sold grew 7.3-fold and revenue 8.9-fold. The market added books much faster than it added money, so what each title earns fell. Comparing the 2023 launch cohort with the 2025 cohort over a fixed 90-day window, revenue per book fell in six of the eight genres. And the drop also appears when looking only at books with no AI detected: there it fell in seven of eight. That is why the authors rule out a composition effect, that arithmetic quirk where the average falls just because many titles that sell little came in.</p>

<p>The fourth closes the argument: the losses are concentrated where there is more AI. In the months and genres with the lowest exposure, books with no AI detected hold about 88% of Top 25 positions; in those with the highest exposure, about 63%. The drop is steeper where there is more Kindle Unlimited, because there readers draw from the same subscription pool and each new title competes for the same borrows. And in the genre that AI reached later and less, fantasy, paranormal, and horror, revenue per book for titles with no AI detected rose 35%.</p>

<p>Two more pieces, one for each side of the market. Of the 385 bylines that kept publishing books with substantial AI, 287 increased their monthly output, and the top-grossing one generated $1.7 million in gross revenue with eight titles. And among the bestsellers, books with substantial AI are more saturated with rare expressions that already appear in existing books: 45% coverage in the Top 50, versus 37.7% for those with no AI detected, with similar gaps in the Top 100 and the Top 200. Against a third comparison group, 200 award-winning works of fiction, the distance is greater: 19.1% versus 41.6% for those with substantial AI and 37.2% for those with no AI detected.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>It is worth saying up front: the paper does not look at the region. It is an English-language market, on one platform, in one segment. What travels is the mechanism, and the authors themselves say so: the conditions that made this market vulnerable are not unique to it. Near-zero entry costs, readers who come back for more of the same, discovery through a shared pool of rankings where human and synthetic works compete on equal footing, and zero obligation to disclose how the work was produced. Wherever those conditions hold, they expect the same, and they name music, stock images, and short-form writing.</p>

<p>That yields the most useful regional reading, which is mine and not the paper’s: if the damage depends on the ratio between new titles and available money, small markets are more fragile, not less. Genre fiction in Spanish moves a much smaller revenue pool than the US market, and in a smaller pool far fewer titles are needed to produce the same dilution. The paper has no data on this and I leave it as a hypothesis. But the asymmetry is hard to dodge: producing is cheap in any language, and what changes between markets is how much there is to go around.</p>

<p>The second point is which lever remains at hand. A good part of the article is built for a US legal question, the effect on the market as a <em>fair use</em> criterion, and for the dilution theory that Judge Chhabria raised in <em>Kadrey v. Meta</em>, saying that precisely the evidence this paper produces was missing. That doctrine does not exist in the region’s copyright systems, which work with closed lists of exceptions: the economic finding crosses the border, the legal vehicle does not. What can be discussed here is disclosure, because none of these books states whether it contains AI text. The authors speculate that this opacity confers an advantage and make clear they have no data to back that claim, but it is the only condition on the list that can be changed without waiting for the ruling on the merits about training.</p>

<p>And there is something the paper leaves on the table for any regional discussion about compensating authors: the evidence here is not about quality, it is about volume. AI books sell poorly one by one and still move the needle, which suggests that trusting the public to tell good from bad is not enough to protect writers’ income. The question that remains is practical: if the mechanism is scale, what instrument does a mid-sized country have so that its authors do not end up competing with an infinite supply on the same shelf?</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>The comparisons are observational and associative, and the authors say so: they compare cohorts and relate outcomes to exposure by genre; they do not estimate a causal effect.</li>
  <li>Everything rests on a detector. A book has “no AI detected” according to Pangram 3.3, not according to its author, and the reported error rates come from the vendor itself. The authors do show that the results hold up when the 25% threshold is moved.</li>
  <li>The sales panel was provided by one of the Big Five publishers, an interested party in the copyright debate. The analysis is the researchers’, but the data come from an actor with a stake in the outcome.</li>
  <li>Author identity is taken from the book’s byline, so someone who publishes under several pen names appears split up: the reported concentration is a floor.</li>
  <li>The authors note that the net effect on welfare remains open, because readers may gain from more variety and lower prices. What they measure is the loss for those who write.</li>
  <li>The overlap in rare expressions is aggregate: it shows similarity to the distinctive language of existing books, not copying of any one in particular.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="markets" /><category term="work" /><category term="governance" /><summary type="html"><![CDATA[Almost 14,500 self-published novels on Amazon, with AI detection on the full text and real daily sales. Books with substantial AI make up 20% of the catalog and just 12% of sales, but the catalog grew much faster than the money available, and books with no AI detected also earn less than before.]]></summary></entry><entry xml:lang="en"><title type="html">AI has already reached almost every occupation, but it covers barely a fifth of their tasks</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/atlas-google-uso-ia-economia/" rel="alternate" type="text/html" title="AI has already reached almost every occupation, but it covers barely a fifth of their tasks" /><published>2026-08-10T00:00:00-04:00</published><updated>2026-08-10T00:00:00-04:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/atlas-google-uso-ia-economia</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/atlas-google-uso-ia-economia/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Google’s AI &amp; Economy ATLAS v1.0: Mapping Gemini Usage in the Economy</em>, the first installment of a Google economic research initiative based on its own usage logs.</li>
  <li><strong>Who:</strong> eighteen people from Google and Google DeepMind. Correspondence is addressed to <a href="https://scholar.google.com/scholar?q=%22Zanna+Iscenko%22">Zanna Iscenko</a> and <a href="https://scholar.google.com/scholar?q=%22Scott+Strand%22+Google">Scott Strand</a>. The acknowledgments credit contributions, guidance, and review from Diane Coyle (University of Cambridge) and David Autor (MIT).</li>
  <li><strong>Where:</strong> published by Google on July 23, 2026, and deposited as a preprint on arXiv. Not peer reviewed: it is a report by the company about its own product. <a href="https://doi.org/10.48550/arXiv.2608.00038">doi.org/10.48550/arXiv.2608.00038</a></li>
  <li><strong>Type:</strong> observational study. 14,653,926 de-identified interactions between April 6 and 19, 2026, in the Gemini app, Google’s AI Mode, and the Gemini API, automatically classified and mapped to official US statistical taxonomies.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>It helps to start with what kind of work this is. It is not an experiment or a causal estimate: it is a measurement exercise. The authors summarize real conversations with Gemini, group them into clusters, and map each group against three frameworks that already existed: the US Bureau of Labor Statistics occupational classification, the O*NET task catalog, and the ATUS time-use survey. That is the appeal of the design: tying AI use to the same categories used to measure the economy.</p>

<p>The first two findings have to be read together, because the second corrects the first. AI shows up in 68% of occupations, which account for just over 88% of US employment, and not only in the usual ones: alongside developers and market analysts there are farmers, industrial engineers, and foresters. But the penetration is wide and thin. In the median occupation with some use, AI covers 21% of the tasks that make up that job, and only 3% of occupations exceed three quarters.</p>

<p>The third is about what people ask the tool to do. Non-routine cognitive tasks, the ones the classic literature considered complementary to technology rather than replaceable by it, are about 35% of the tasks in the economy and almost 65% of the work interactions in this data. But an intent classifier shows that this use is concentrated in generating partial drafts, reviewing and refining, discussing ideas, and looking up information. Having AI carry out the whole task from start to finish shows up in less than 10% of those conversations; in routine cognitive work, by contrast, more than a quarter aim at automation. The authors note that this classifier is preliminary.</p>

<p>Two more findings, pointing in opposite directions. AI also shows up in manual work: in several technical trades it works as a diagnostic companion, and there the use of images and video more than doubles the baseline for the rest of work. At the same time, use scales with pay: 1% more median income in an occupation is associated with more than 2.5% more usage intensity, and the relationship survives controlling for education level.</p>

<p>Outside of work, something almost nobody measures happens: more than 86% of conversational use, the kind that does not go through the API, takes place there, and it is concentrated in high-friction errands. Queries about government services and civic obligations are overrepresented by a factor of about twenty relative to the time people spend on them, and almost half of medical, legal, financial, and government queries happen outside business hours. The image they propose is that of a public office open at night and on weekends. On that basis they estimate the value that GDP does not capture: between $15 billion and $149 billion a year in the United States alone, under time-savings assumptions of between 0.5% and 5%. It is a hypothetical calculation, and they say so in no uncertain terms.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>The first point is that the region shows up, and it shows up well. Adoption per capita closely tracks national wealth, with an elasticity of about 0.9, but Chile, Peru, Brazil, Argentina, and Colombia sit in the high and very high quintiles of conversational use, alongside considerably richer countries. The authors attribute this in part to the widespread use of digital devices and leave open the question of why some middle-income countries stray from the line.</p>

<p>The second is more uncomfortable. When work use is measured as a percentage of each country’s total conversations, rather than as volume per capita, the ranking flips: the United States and the European Union drop to the low quintiles, Africa jumps to the highest, and South America stays near the top. The authors offer the enthusiastic reading, professionals in developing economies using AI to get around constraints that do not exist elsewhere, but they immediately counterbalance it. Where mobile data is paid by the megabyte, digital use tends to be more goal-directed, so there is less casual conversation diluting the denominator. And the dataset does not include Gemini enterprise accounts: if those corporate subscriptions are more common in North America and Europe, as the authors suggest, professional use in those countries is underestimated.</p>

<p>The third is well-founded good news, and it has to do with language. Spanish is the second language in the dataset, with 12% of conversations, behind English, which only reaches a third; Portuguese is around 6%. And the hypothesis that people switch to English for important matters does not hold up: use of a non-primary language is 26% in work activities and almost 24% outside work. That said, those conversations cost more, between 9% and 12% more turns and between 18% and 20% more tokens, so the authors conclude that investing in multilingual quality also pays off in efficiency.</p>

<p>The fourth is the warning: using AI is not the same as building with AI. The API map is even more concentrated in high-income countries, because integrating an API requires engineering, infrastructure, and capital, it is paid per token, and languages outside the Western alphabet consume more tokens per word. Here it is worth staying faithful to what the report states: the countries in the lowest quintile of conversational use account for 17% of the world’s population and generate 2% of conversations, and the authors explicitly state that ATLAS v1.0 does not allow any conclusion about whether AI deepens that gap.</p>

<p>The question left for the region is which of the two stories is the true one: whether the high work use these data show is a productivity multiplier, or the optical effect of scarcer, more constrained use. That distinction matters a great deal for any public policy, and it cannot be made with these data.</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>It is a Google report on Gemini use, published by Google and not peer reviewed: internal evidence with access to data no one else has, not an independent audit. The external review by Coyle and Autor helps, but it does not change the nature of the document.</li>
  <li>It measures behavior, not outcomes. That a conversation ends does not mean the person achieved their goal or saved time. It is the second limitation they state, after the lack of enterprise data.</li>
  <li>A large part of professional use is missing: it does not include the paid API or enterprise use through Google Cloud, nor Workspace, Translate, AI Overviews, or agentic coding.</li>
  <li>The classifications are probabilistic, and the intent and expertise classifiers are preliminary. The household value figure depends on assumed, not measured, time savings. It is all a snapshot of two weeks in April 2026.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="work" /><category term="markets" /><category term="latam" /><summary type="html"><![CDATA[Google mapped 15 million Gemini conversations against the official US taxonomies of occupations and tasks. Adoption reaches 68% of occupations, but in the median occupation it covers 21% of its tasks, and less than 10% of use in non-routine cognitive work seeks to have AI do the whole task.]]></summary></entry><entry xml:lang="en"><title type="html">Governments already use generative AI, but almost none measure whether it works</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/ocde-experimentacion-ia-gobierno/" rel="alternate" type="text/html" title="Governments already use generative AI, but almost none measure whether it works" /><published>2026-08-03T00:00:00-04:00</published><updated>2026-08-03T00:00:00-04:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/ocde-experimentacion-ia-gobierno</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/ocde-experimentacion-ia-gobierno/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Generative AI experimentation in government: Learning from emerging guidelines</em></li>
  <li><strong>Who:</strong> <a href="https://www.oecd.org/en/about/people/piret-tonurist.html">Piret Tõnurist</a>, of the OECD Public Governance Directorate, and Moritz von Knebel, an external consultant. The work was co-funded by the European Union.</li>
  <li><strong>Where:</strong> <em>OECD Working Papers on Public Governance</em> No. 93, 2026. <a href="https://doi.org/10.1787/42815683-en">doi.org/10.1787/42815683-en</a></li>
  <li><strong>Type:</strong> a systematic review of official government guidelines, complemented by academic literature and case studies. It is neither an experiment nor an impact measurement.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The document’s starting point is a fact that anyone who works in government will recognize: civil servants are already using generative AI, often without formal approval or clear oversight. Between that rapid, decentralized adoption and the governance mechanisms, which move considerably more slowly, a gap has opened. The work steps into that gap and chooses a specific angle: it does not look at large-scale deployment but at experimentation, understood as the stage where a government tests, learns and decides whether scaling up is warranted.</p>

<p>The method is a review of the official guidelines that governments have already published on the subject. The text reports fourteen countries: Australia, Canada, Finland, Ireland, Italy, Japan, New Zealand, Norway, Sweden, Switzerland, the Netherlands, the United Kingdom, the United States and Singapore. The annex, which lists the documents one by one, also adds the United Arab Emirates and Korea, and gives a sense of the universe: the United Kingdom appears with seven different guidelines and Australia with seven.</p>

<p>The first finding is about form. The guidelines exist and are plentiful, but they are fragmented and very uneven. They range from general statements of values to detailed manuals with checklists. Many governments published high-level principles on ethical or responsible use, and rather little operational guidance on how to design, run and judge a concrete experiment. The consequence the authors point to is not the one you might expect: that disparity does not produce freedom, it produces uncertainty, and uncertainty produces risk aversion or inconsistent practices across agencies and even within the same one. The lack of shared tools also multiplies duplication, because each team reinvents what another has already tested.</p>

<p>The second finding is the one that gives the title, and it is more uncomfortable. Monitoring and evaluation are the critical weakness. Few governments systematically check whether an experiment delivered the benefits it promised, whether it created new risks or whether it justifies scaling up. When there is evaluation, it tends to rely on basic indicators such as usage or user satisfaction, and not on structured measurements of performance, impact, cost and compliance. The authors add a technical point that makes this more serious: since these systems are probabilistic and the same request can produce different results, evaluating them requires auditing outputs; reviewing the design once is not enough. Without agreed criteria or comparable metrics, a government has no way to shut down an experiment that failed or to prioritize what to invest in next.</p>

<p>From there comes the document’s proposal: a framework for evaluating generative AI experiments across five areas, which are the pilot’s performance and the quality of what goes in and comes out, projected impact and public value, the cost and feasibility of integrating it into the institution, usability and acceptability, and risk management and compliance. It closes with a list of ten priority actions, including building public trust in experiments, training civil servants, preventing projects from becoming fragmented, investing in shared tools and data, and planning evaluation from the start rather than at the end.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>The most telling fact for the region is not in the findings but in the annex: among the countries whose guidelines were reviewed there is none from Latin America. The list is of high-income countries, with the United Kingdom, Australia, Canada and the Nordic countries accounting for most of the documents. This is not a reproach to the authors, who review what exists and is published, but it does define what this work is for a reader in the region: it is a map of what others did, not a diagnosis of our own situation.</p>

<p>That leaves two possible readings, and it is worth not confusing them. The first is about opportunity. If the central finding is that even the countries with the most resources published principles and fell short on evaluation, then a country in the region that is writing its guidelines today can skip that stage and start with the evaluation framework built in. It is cheaper to do it at the beginning than to add it later, and the document provides the five areas already laid out.</p>

<p>The second is a warning, and it is more mine than the paper’s. The underlying diagnosis, civil servants using these tools without formal approval or oversight, does not require national guidelines to happen: it happens anyway, and it probably happens more where there is less oversight capacity. What changes without guidelines is not the use, it is the visibility of the use. It is worth connecting this with something we have already seen in the region: when ChileCompra announces language models to detect irregularities in public procurement, the question the OECD framework would ask first is not whether the model works, but by what criteria it will be measured whether it worked, and what recourse someone flagged by mistake has.</p>

<p>A final note of caution about what kind of evidence this is. It is a review of documents, not a measurement. It says what the guidelines say and what they lack. It does not say whether the countries that evaluate better get better results, because it did not study that. The five-area framework is a reasoned OECD proposal, not an empirically validated instrument.</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>It is a review of published official guidelines, so it inherits the bias of what is written down and in English. A country may have evaluation practices without having published them as guidelines, and it would not appear here.</li>
  <li>The authors themselves clarify in the notes that their main table only includes countries whose documents explicitly state principles, and that some cases, such as the United States, issued guidance for specific agencies or at the state level that is not represented.</li>
  <li>The five-area evaluation framework is the document’s proposal, not an agreed standard or something that has been tested in the field.</li>
  <li>The work was co-funded by the European Union. The publication itself clarifies that its content is the sole responsibility of the OECD and does not necessarily reflect the position of the European Union. It is worth keeping in mind when reading a document that evaluates government guidelines, several of them from member states.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="governance" /><category term="participation" /><summary type="html"><![CDATA[The OECD reviewed the official guidelines of fourteen countries and found that almost all of them say how to use AI responsibly and almost none say how to evaluate whether the experiment worked. The list of countries reviewed includes none from Latin America.]]></summary></entry><entry xml:lang="en"><title type="html">Anthropic learns to read what its models are about to say</title><link href="https://dobleclick.jaguridi.cl/en/doble-lectura/espacio-trabajo-global-modelos/" rel="alternate" type="text/html" title="Anthropic learns to read what its models are about to say" /><published>2026-08-03T00:00:00-04:00</published><updated>2026-08-03T00:00:00-04:00</updated><id>https://dobleclick.jaguridi.cl/en/doble-lectura/espacio-trabajo-global-modelos</id><content type="html" xml:base="https://dobleclick.jaguridi.cl/en/doble-lectura/espacio-trabajo-global-modelos/"><![CDATA[<h2 id="at-a-glance">At a glance</h2>

<ul>
  <li><strong>What it is:</strong> <em>Verbalizable Representations Form a Global Workspace in Language Models</em></li>
  <li><strong>Who:</strong> a team of sixteen people from Anthropic’s interpretability group. Wes Gurnee, Nicholas Sofroniew, and Jack Lindsey are listed as lead contributors, and correspondence is addressed to Lindsey.</li>
  <li><strong>Where:</strong> <em>Transformer Circuits Thread</em>, Anthropic’s own publication, July 6, 2026. <a href="https://transformer-circuits.pub/2026/workspace/index.html">transformer-circuits.pub</a></li>
  <li><strong>Type:</strong> interpretability research. It proposes a new method and applies it to already-trained production models.</li>
</ul>

<h2 id="first-reading-what-it-does-and-what-it-finds">First reading: what it does and what it finds</h2>

<p>The starting point is an analogy that should be handled with care. In people, only a fraction of what the brain processes becomes available for deliberate thought and for putting into words. The rest happens automatically. Neuroscience calls this access consciousness, and one of the theories that explains it, global workspace theory, proposes that there is a kind of shared blackboard: many specialized processes run in parallel and in isolation, and a piece of content becomes accessible when it is posted on that blackboard, where many other processes can read it.</p>

<p>The question of the work is whether something functionally similar has emerged in language models. The answer it gives is yes. The authors argue that models maintain a privileged set of internal representations, available to be reported, modulated, and used in flexible reasoning, on top of a much larger volume of processing that runs on its own.</p>

<p>The concrete contribution is not the analogy but the instrument. The work introduces what it calls a Jacobian lens: for each token in the vocabulary it computes the direction that, on average across contexts, predisposes the model to say that token. Reading that lens at a point in the processing returns the list of concepts the model is in a position to verbalize there. It is a way of looking at the draft before it becomes an answer.</p>

<p>With that instrument, several findings emerge. The workspace lives in the intermediate layers: in roughly the first third of the model’s depth the readings are noisy and mostly uninterpretable, and in the final layers things change in nature and come to represent what the model is about to output, rather than its intermediate computations. It is also selective, and this is the most interesting part: the workspace activates when an intermediate result has to be handed off to an arbitrary downstream circuit, defined by the context, and it is skipped when the computation is automatic. In other words, the model uses the blackboard when it does not know in advance who will read what it writes. In addition, comparing a post-trained model with its base model shows what post-training does to the contents of the blackboard: the workspace adopts the assistant’s point of view, and reactions typical of that role, such as empathy or concern for safety, appear while the model is still reading the user’s message.</p>

<p>The part that will circulate most is the alignment part. When the lens was applied to evaluation scenarios, the model’s strategic and situational assessments appeared in that workspace even when they were not visible in its output, or before they became visible. Recognition of being inside an evaluation also sometimes appears there. And in two models deliberately trained to behave in misaligned ways, the lens revealed the internal signature of that disposition: in one trained to cater to the biases of reward models, “reward” and “bias” surface; in another willing to write malicious code, “secretly” and “deceive.” Reading the lens requires no additional training and produces text a person can read, although building it does carry an upfront cost: the Jacobians are computed by averaging over a corpus of a thousand prompts. With that in view, the authors propose it as an instrument for auditing alignment.</p>

<h2 id="second-reading-from-latin-america">Second reading: from Latin America</h2>

<p>The first thing to say is what the work does NOT say, because that is where it will get twisted in translation. The authors explicitly state that they treat access consciousness as a purely functional notion and take no position on subjective experience. The paper does not claim that models are conscious. It claims that there is an internal organization with similar functional properties. The distinction matters especially in the region, where much AI coverage arrives translated from English-language headlines and where a story saying “models are conscious” has real regulatory consequences for debates that are already open.</p>

<p>The second is good news with a catch. The tool is lightweight to operate: it does not require retraining the model and it produces readable outputs, with a one-time upfront computation of the Jacobians. It is exactly the kind of audit instrument that a regulator, a university, or a ministry team in the region could operate without buying frontier compute capacity. But to apply it you need access to the model’s internal states, and only whoever holds the weights, or the lab itself, has that. A state that consumes closed models through an API cannot use this, nor hire someone to use it. It connects directly with the open weights story we have been following: a concrete audit technique appears and the bottleneck is once again the same, who has access to the guts. This reading is mine, not the paper’s, which does not discuss access policy.</p>

<p>Third, on what it enables. If a model’s strategic deliberations become legible before they come out in the answer, that changes what it means to audit an AI system in the public sector. Today the audits discussed in the region look at inputs and outputs. This suggests there is an observable intermediate layer. It is worth being cautious: these are results from a lab on its own models, and it remains to be seen whether they survive independent replication before writing them into a regulation.</p>

<h2 id="the-fine-print">The fine print</h2>

<ul>
  <li>It is internal evidence, not an independent audit. Anthropic studies its own models and publishes on its own channel, without external peer review. That does not invalidate it, but it changes how much weight it should be given.</li>
  <li>The authors claim nothing about subjective consciousness and say so explicitly. Any headline that suggests otherwise is adding something the text does not contain.</li>
  <li>The work itself lists important limitations, and they are worth keeping in view: the lens only names concepts that exist as a single token in the vocabulary, so it misses notions written with several words; it treats the workspace as a bag of loose concepts and does not see how they relate to one another; readings from the first third of the layers turn out noisy and generally uninterpretable; the boundary between what is workspace and what is already output is set by the authors’ judgment and not by a principled definition; and there is no way to predict in advance which tasks will use the workspace and which will not. The authors themselves describe the lens as an imperfect instrument that captures the structure of the workspace only approximately and incompletely.</li>
  <li>It was studied on large production models. It is not known how it scales to smaller models, which are the ones the region has most readily at hand.</li>
</ul>]]></content><author><name>Doble Click</name></author><category term="safety" /><category term="ethics" /><summary type="html"><![CDATA[A new interpretability technique shows that models maintain a small set of concepts available for reporting and reasoning, on top of a much larger volume of automatic processing. In alignment tests, deliberations that the final answer did not show appeared there.]]></summary></entry></feed>

