Doble Lectura #24

The same GPT-4 that improved the work of 758 consultants made them get another task wrong

A preregistered experiment with 758 BCG consultants shows that GPT-4 sped up and improved their work on product development tasks. On a business case that fell outside its reach, the opposite happened: those who used the tool got it right less often, and their wrong answers sounded more convincing.

Generated automatically · sources linked · no prior human review

At a glance

First reading: what it does and what it finds

The experiment was carried out at Boston Consulting Group (BCG) with 758 consultants, about 7% of its consultants who do not manage a team. All of them first completed a task without AI, which served as a baseline. Then a random draw split them into three groups: no AI, GPT-4, or GPT-4 plus a brief training on how to write instructions for it. Each person also worked on one of two types of task, designed with the firm’s executives and tested beforehand to fall on one side or the other of what the model handled well. That line is the jagged frontier of the title: tasks that seem equally difficult to a person can fall on different sides.

Inside the frontier, participants had to come up with a shoe for a niche market and take it all the way to launch, across 18 creative, analytical, writing, and persuasion subtasks. Compared with the group without AI, those who used GPT-4:

  • completed 12.2% more tasks;
  • finished them 25.1% faster;
  • raised quality by 33.9% if they had received the training;
  • and by 29.9% if they had not.

Outside the frontier, the task was a business case: recommending to a general manager which brand had the most potential, by cross-referencing a spreadsheet with internal interviews. The spreadsheet looked complete, but the interviews contained details that changed the answer, and the model reached the wrong conclusion. The share of correct answers came out as follows:

  • without AI: 84.5%;
  • with GPT-4: 70.6%;
  • with GPT-4 and training: 60%.

On average, using AI cut accuracy by 19 percentage points. The trained group, which gained the most inside the frontier, got it right even less often outside it, but the difference between the two AI groups is weak: significant only at the 10% level. Both AI groups also spent less time on this task than the group without AI; the trained group, more than 11 minutes less, or 30% less.

The third finding explains why the error is hard to see. Evaluators who did not know the correct answer scored the coherence and persuasiveness of each recommendation. Those who used AI were rated higher, including when they had gotten it wrong. The authors conclude that AI improves presentation and argumentation even when the analysis is wrong. One additional result: inside the frontier, those who started in the bottom half on the baseline task gained the most, although the top half also improved.

What is established is that the same model, used by the same professionals on tasks from their own line of work, raised performance on some and lowered it on another, and that from inside the work it was not obvious which was which.

Second reading: from Latin America

This is the second Doble Lectura in a row about an experiment by a largely shared team, with Dell’Acqua, Lifshitz, Mollick, and Lakhani on both. The previous one showed how much AI adds for a professional; this one shows where it subtracts.

What is most useful for the region lies in two results of different strength. The solid one: outside the frontier, those who used AI got it right less often, and their wrong answers were more convincing to those who evaluated them. The fragile one: the group trained in writing instructions got it right even less often, with a weak difference and after working for less time.

For the schools of government that train public servants, the fragile result allows only one prudent reading: teaching people to ask the model for things better did not protect against error on the task where the model failed. The study did not measure participants’ confidence, so it does not say why the trained group got it right less often. It does suggest that a course should include tasks where the model gets things wrong, not just examples where it shines.

The solid result concerns whoever does the reviewing, for example a committee that scores proposals in a procurement process with a rubric for clarity and justification. If the documents it receives are written with AI, the finding on coherence indicates that such a rubric may reward a well-written but mistaken analysis. Read from the region, the practical consequence is that review has to go back to the source data, and that requires budgeting reviewer hours to redo parts of the analysis, not just to read the document.

A third point remains open. That the consultants with the lowest initial performance gained the most inside the frontier suggests the tool could narrow gaps in public teams with little experience. It is not yet known whether that advantage holds when the same team faces, without warning, a task that falls outside.

The concrete implication for a ministry adopting AI assistants is where to put the effort. This evidence points to mapping which tasks in its own workflow fall outside the frontier and strengthening human review there: training in instructions improved performance inside the frontier, but it did not protect outside it.

The fine print

  • There is only one task outside the frontier, placed there through prior testing: a limitation the authors acknowledge.
  • Inside the frontier, quality was measured with subjective evaluations, something the authors disclose. Replications with footwear experts point in the same direction.
  • The model was GPT-4 in its April 2023 version. The authors warn that the frontier is not fixed and that a task can move inside with a new model.
  • Participants were early in their careers, with incentives that rewarded quality within the time limit.
  • Conflict of interest: three coauthors are listed as affiliated with the BCG Henderson Institute, and the company collected the data. The preregistration did not include the jagged frontier framing.
  • It applies to complex professional work with a 2023 model. Where the frontier lies today is something each organization has to measure again.

Revised on September 28, 2026. The original version stated that the group trained in writing instructions “made the most mistakes” and that such a course “can raise confidence in the tool.” In the paper, the difference between the two AI groups is significant only at the 10% level, that group spent less time on the task, and confidence was not measured. Those sentences and the title were adjusted.

Paper keywords: technology and innovation management, organizational economics, economics and organization, organization and management theory, research design and methods, field experiments, implementation of new technology, organizational processes

Automated reading. This text was generated by Claude, an Anthropic model, from the original source, without line-by-line human review. It may contain errors or debatable interpretations; to check any point, see the original source.
Spotted an error? Report it

Tell us what's wrong, quoting the sentence if you can and, if you have it, the source that corrects it. An automated process reviews reports every night: if the error is verified, the page is corrected and a correction note is added at the bottom.

Your email is optional: we only use it if we need more context about the report. It doesn't subscribe you to the newsletter.