What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
A 2023 experiment with 758 Boston Consulting Group consultants found that GPT-4 helped participants complete 12.2% more tasks and work 25.1% faster on tasks within the model’s tested capability range. Human-rated quality was also substantially higher: early coverage put the gain at about 40%, while a later Harvard summary reported about 32%. Those figures are not a universal 40% productivity increase for enterprise workers.
Which study the 40% claim refers to
The claim comes from “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” a study led by researchers affiliated with Harvard Business School, Wharton, MIT and other institutions, in collaboration with Boston Consulting Group (BCG). The paper first appeared as a working paper in September 2023 and was published online in Organization Science on March 11, 2026. Read the published paper or the Harvard Business School AI Institute summary.
The experiment involved 758 BCG consultants, about 7% of the firm’s individual-contributor consulting workforce. They were highly educated professionals, not a representative sample of workers across industries. The tasks were designed to resemble consulting work; the study did not track months of client projects or company-wide results.
What the consultants did
Participants worked on realistic tasks for a fictional footwear company, including product ideation, market segmentation, marketing copy and persuasive communication. Examples included proposing products, analyzing customer segments and drafting an internal memo. These were controlled experimental assignments, not measurements of revenue, billable work, customer satisfaction or long-term business performance.
#1 Best Overall
Participants first completed baseline work without AI, then were randomly assigned to one of three conditions: no AI access, GPT-4 access, or GPT-4 access plus an overview of prompt-engineering techniques. The study examined tasks where GPT-4 was within its capability frontier and included a separate task outside that frontier. The working-paper record describes the experimental conditions.
What the headline numbers measure
| Measure | Finding on tasks within GPT-4’s tested capability frontier |
|---|---|
| Tasks completed | 12.2% more tasks |
| Completion speed | 25.1% faster |
| Human-rated quality | About 40% higher in early coverage; about 32% in a later Harvard summary |
The quality estimates differ by the version or summary being cited; they should not be presented as one uncontested figure. The later Harvard retrospective describes an average quality advantage of roughly 32%, while early coverage described it as about 40%. Neither figure means participants produced 40% more economic value. Task volume, elapsed time and graders’ quality ratings are distinct measures, and they cannot be added together into a single productivity score.
The phrase “40% performance boost” became a shorthand for a prominent quality result, not a finding that all employees using GPT-4 become 40% more productive. The early headline coverage is useful for understanding how that framing spread, but the published study is the better source for what was measured.
The result that complicates the headline
On a complex managerial task outside GPT-4’s reliable capability range, participants with AI access were 19% less likely to produce a correct solution than participants without AI. This was not just a matter of the model making mistakes: users could also fail to recognize that its plausible-sounding output was wrong. The study therefore found that AI could help substantially on one task and impair performance on another.
For managers, the practical distinction is between work an employee can independently check and work where an error may be hard to spot. Drafting or generating candidate ideas may be useful with review; consequential analysis or decisions need stronger safeguards when users cannot validate the reasoning or facts. The journal’s abstract and publication record describe the outside-frontier finding.
Who gained most—and why expertise still mattered
Lower-performing participants made the largest gains on suitable tasks; contemporary reporting put the improvement for the lowest performers at about 43%. That suggests AI can help some workers produce stronger first drafts or reach a usable answer more quickly. It does not show that skill differences disappear or that inexperienced workers can replace trained consultants. Participants still needed enough judgment to decide whether the tool’s output made sense, especially on tasks where the model was unreliable.
What “Centaurs” and “Cyborgs” mean at work
The researchers used two labels for patterns of human-AI collaboration. They are descriptions of how participants worked, not formal job types or guaranteed recipes for success.
- Centaurs: People divide work between themselves and the AI, switching roles as the task changes.
- Cyborgs: People weave AI into an ongoing workflow, combining human and model contributions more continuously.
Both patterns point to a more useful deployment question than “Should everyone get a chatbot?”: which task should the model handle, which decisions require a person, and where can a reviewer check the result?
Best Value
- Brand New in box. The product ships with all relevant accessories
Higher average quality can mean less variety
The experiment also found that AI-assisted ideas could be higher quality while being less varied or more homogeneous. A tool that helps many employees produce polished, plausible work may also pull their ideas and language toward similar patterns. That trade-off matters when originality is a goal, such as product strategy or creative development. Teams may want a separate stage for human-generated alternatives or other deliberate ways to widen the range of ideas before selecting a final direction.
What the study can—and cannot—say about enterprise AI
The experiment is evidence about particular people, tasks and a particular model at a particular time. It does not establish that an average worker in any industry will see the same gains.
- One firm and occupation: Participants were BCG consultants, not a cross-section of enterprise employees.
- Simulated assignments: The tasks resembled consulting work but did not measure long-term outcomes on real client engagements.
- An early GPT-4 system: The experiment tested GPT-4 as it existed in or around June 2023. The 2026 publication does not make it a benchmark for today’s models or workplace products.
- Task-level outcomes: The study measured completion, speed and assessed work quality—not revenue, profit, headcount savings, customer outcomes or total cost of ownership.
Peer review gives the study a settled publication in Organization Science; it does not remove these limits or establish how current systems perform in other workplaces. Results from later GPT-4-class or other frontier systems require their own testing under comparable conditions.
How a company can test the claim in its own workflow
Rather than assume the study’s percentages will transfer, run a controlled pilot on specific, reviewable tasks. Compare AI-assisted work with a no-AI baseline, and measure more than whether employees say the tool feels faster.
- Select a narrow task: Start with work such as drafting, summarizing or generating candidate ideas, where a qualified person can review the output.
- Set a baseline: Record typical completion time, completion rate, rework and quality without AI.
- Define review criteria: Decide in advance how to assess accuracy, usefulness and originality, and who is qualified to judge each one.
- Compare like with like: Have comparable workers perform the same task with and without the tool, keeping the evaluation consistent.
- Track costs and risks: Include time spent checking and correcting, licensing and implementation costs, data-governance requirements, and any security or compliance incidents.
- Expand only on evidence: Scale a workflow when its measured gains outweigh review effort, errors and costs; reassess if the model, task or process changes.
A useful pilot tracks completion time, rework, error and correction rates, human-rated quality, idea diversity, customer or client outcomes where measurable, adoption by role, and the full cost per completed task. For confidential or high-stakes work, organizations also need approved data controls and clear human-approval rules; the experiment did not evaluate those deployment conditions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




