In 2023 a group of researchers from Harvard Business School, with colleagues at Wharton, MIT and elsewhere, did something rare: they tested AI on real professionals doing real work at scale. The partner was Boston Consulting Group. 758 consultants took part, about 7 percent of the firm's individual-contributor consultants1,3.
Participants were randomly assigned. Some worked as usual. Some were given GPT-4. Some were given GPT-4 plus a short overview of prompting. Then everyone received the same kind of tasks consultants do every day: generating product ideas, analysing markets, writing proposals, preparing client presentations1.
Inside the frontier
On 18 tasks within AI's capabilities, the results were striking. Consultants using AI completed 12.2 percent more tasks, finished 25.1 percent faster, and produced work rated more than 40 percent higher in quality than the control group1.
The biggest gains went to those who started weaker. Consultants below the average improved by 43 percent against their own baseline; those above average by 17 percent1. AI narrowed the gap between strong and less strong performers.
Had it stopped there, this would be a study for every AI tool launch presentation.
Outside the frontier
But the researchers deliberately included a task outside AI's capabilities. At a glance it looked no harder than the others. It required combining quantitative data with subtle details from interview notes, the kind of work the models of the time often got confidently wrong.
On this task, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it1. They trusted the machine's fluent response and set aside their own judgement.
The authors called this the "jagged technological frontier": AI is excellent at some tasks and surprisingly poor at others of similar difficulty, and the line between the two is invisible to the naked eye1.
Two ways of working with the machine
Watching how successful participants used AI, the researchers saw two patterns. The first, which they called centaurs, divided the work clearly: which parts go to the machine, which to do themselves, and checking the machine's output before using it. The second, which they called cyborgs, wove AI into every step, trying, correcting and asking again1.
What both had in common was that they did not hand their judgement to the machine. They knew, or made the effort to find out, where AI does well and where it needs to be doubted.
Why this belongs to the learning team
Plenty of companies have bought AI licences for everyone and treated that as the transformation done. The BCG experiment shows that issuing licences is only the first step, and can do harm if it stops there: on tasks outside the frontier, people with AI did worse than people without.
The capability to teach is not clever prompting. It is recognising which tasks fall inside AI's strong zone and which do not, and the habit of checking before trusting. That capability differs by profession and process, so no generic course can teach it on your behalf.
The frontier in HR work
The study was run with consultants, but the question applies to every office profession, HR included. Some tasks usually sit comfortably inside AI's strong zone: drafting a job description, summarising hundreds of engagement survey comments, suggesting competency-based interview questions, rewriting a notice so it is easy to read.
Others look no harder but slip outside the frontier easily: calculating overtime pay, personal income tax or social insurance under recently changed Vietnamese rules, interpreting a Labour Code provision for a specific case, weighing the right disciplinary outcome for a complicated incident. Here the machine's answer is usually fluent and confident, and that confidence is what makes people skip the check.
The list is not fixed. Each new model version pushes the frontier somewhere new, in directions nobody predicts. The Harvard team's results were later peer-reviewed and published in Organization Science2, but the most important lesson does not depend on any particular model: do not assume, test it on your own work.




