Case Study · Productivity

BCG found AI made consultants 40% better - and 23% worse, depending on the task

What happened

BCG, working with researchers from Harvard, MIT, Wharton, and the University of Warwick, ran an experiment with more than 750 BCG consultants. On a creative task - developing new products and go-to-market strategy - about 90% of participants using GPT-4 improved their results, with average performance roughly 40% higher than the control group.

But on a task requiring structured business problem-solving, GPT-4 users performed 23% worse than the control group.

The business problem

AI isn't automatically productive - its value depends entirely on the type of task, the model's actual capability at it, and how the output is used.

Why it worked

  • The right question isn't "should we use AI" - it's "where does AI increase output, and where does it increase risk"
  • Different task types need different guardrails, not a blanket policy
  • Measure before rolling out at scale
  • This is the argument for an AI audit before a rollout, not after

Where this points

AI AuditAI Use Case DiscoveryAI Training

Want to know where your organization stands?

Book a Consultation