At a glance
- What it is: The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork
- Who: Fabrizio Dell’Acqua, Raffaella Sadun and Karim R. Lakhani (Harvard Business School), Charles Ayoubi (ESSEC), Hila Lifshitz (Warwick Business School), Ethan Mollick and Lilach Mollick (Wharton), plus four Procter & Gamble professionals: Yi Han, Jeff Goldman, Hari Nair and Stew Taub.
- Where: Organization Science, vol. 37, no. 4, July–August 2026, pp. 1217-1242, open access. doi.org/10.1287/orsc.2025.20702
- Type: preregistered field experiment, 2 × 2 design, within a single company.
First reading: what it does and what it finds
The study was carried out between May and July 2024 at Procter & Gamble, with 791 professionals from the commercial and research and development areas. Each one spent a full day in a virtual product development workshop, working on real challenges from their own business unit. A random draw assigned them to four conditions: alone without AI, in a cross-functional pair without AI, alone with AI, or in a pair with AI. The tool was built on GPT-4. Expert evaluators, blind to each participant’s condition, scored the solutions, and those scores were standardized against the group that worked alone and without AI.
Quality, measured in standard deviations above that control group, came out as follows:
- pair without AI: 0.24
- individual with AI: 0.37
- pair with AI: 0.39
That is the first finding. The individual with AI reached the level of the human pair, and adding AI to the pair barely moved the average beyond that.
The second finding is about expertise. Without AI, people from the commercial area proposed commercial ideas and those from research and development proposed technical ideas, with clearly different distributions. With AI that distinction fades and both groups generate a similar mix, without quality varying significantly according to how technical the solution was. The most striking result came among those who do not usually work in product development: alone and with AI, they reached the level of teams that included someone who does it every day.
The third finding is the hardest to digest. Since everyone generated five ideas before choosing one and developing it, the authors separated the two stages. AI raised the average quality of the ideas across the entire distribution, without narrowing the distance between the best and the worst. But when choosing which of the five to develop, those who worked without AI seem to have been more accurate: pairs without AI kept their best idea about half the time, and the AI conditions around 37%. Even so, since they started from better ideas, the ideas chosen by those who had AI were still of higher quality. The authors offer possible explanations, among them the tendency of models to validate what the user already brings, warn that perhaps participants did not use AI to choose, and leave the point open. They describe AI as a quality amplifier more than as a better decision-maker. There are two other results: pairs with AI placed more solutions in the top decile, and those who used AI reported more positive and fewer negative emotions at the end.
What is established in this context is an asymmetry between stages. AI raised the floor of idea generation to the point of matching what a second professional contributed, and it did not improve the next step, which is choosing well.
Second reading: from Latin America
The experiment covered four business units in two geographies, Europe and the Americas, and the paper does not break down results by country.
Our reading is that the result with the most direct translation to the region is the substitution one. For a mid-sized company in Chile or Colombia that does not have a research and development area separate from the commercial one, the bottleneck is not a lack of ideas but that there is no one to cross them with: adding a second specialized professional costs a salary, and what the study participants received was a license and an hour of training. The finding that the individual with AI reached the level of the cross-functional pair speaks precisely to that margin.
The finding about the selection stage changes something concrete in how procurement is done. An innovation unit in a ministry in Chile or Brazil that is drafting the terms of reference today to contract an AI assistant can specify it for generating and writing drafts, and put in writing that the choice between alternatives remains a documented human decision, with time allotted for it. That is not bureaucratic zeal: the only stage where the group without AI came out ahead was precisely identifying its own best idea.
A third point remains a hypothesis. In the region it is common for a small public team to have a single person in charge of a topic, with no specialized colleague to consult. That the employees least familiar with the task reached, with the tool, the level of teams that included someone experienced suggests that this is a profile where it performs especially well. Extrapolating it to the region’s public sector is our bet, not the authors’.
What is concrete for the region is a criterion for where to put the tool first. Where there is no one to test an idea against, this evidence says that a license and some training move the needle. Where the problem is choosing well among several options that already exist, the evidence offers no support, and something points against it.
The fine print
- The preregistration covered performance and expertise. Emotions were included as a variable of interest with unclear effects, and the analysis of the upper tail emerged during the research: it is exploratory.
- Scope as stated by the authors: a consumer goods company, a single virtual day, pairs formed at random among people who generally did not know each other. They compare them to flash teams, not to established teams.
- Participants had little experience with the tool, and the authors present the benefits as a floor, not a ceiling.
- Conflict of interest: four coauthors work at Procter & Gamble and the design was agreed with its leadership. It is evidence from inside the organization studied, with preregistration and blind evaluators.
- That is as far as it goes: one company, one early ideation task, one model. What happens when use is sustained over time remains an open question for the authors themselves.
Paper keywords: technology and innovation management, research design and methods, field experiments, implementation of new technology, organization and management theory, organizational processes, economics and organization, organizational economics
Automated reading. This text was generated by Claude, an Anthropic model, from the original source, without line-by-line human review. It may contain errors or debatable interpretations; to check any point, see the original source.