Patrick Boyle on Finance
Patrick Boyle on Finance

Is AI Actually Useful?

Send us a textA new Harvard Business School study analyzed the impact of giving AI tools, to white collar workers at Boston Consulting Group.In the study, management consultants who were told to use Chat GPT when carrying out a set of consulting tasks were far more productive than their colleagues w

Featured Speakers

Patrick Boyle Host

Topics Discussed

Episode Summary

Executive Summary: The episode examines whether generative AI is actually useful in professional work, using a Harvard/BCG experiment to show that GPT-4 boosts productivity and quality on tasks within its capabilities, but can reduce accuracy on tasks requiring deeper human judgment. It argues AI works best when humans validate, steer, and combine it with their own expertise.

Main Topics: AI hype versus real workplace utility (Priority: 5/5): The episode opens by questioning whether generative AI’s surprising capabilities translate into dependable workplace value, noting that models can excel at some complex tasks while failing at seemingly simple ones. The jagged technological frontier (Priority: 5/5): The central concept is that AI performance is uneven: some tasks are easy for models, others nearby in apparent difficulty are not. This makes it difficult for users to know when to rely on AI. Harvard/BCG field experiment on consultants (Priority: 5/5): A pre-registered study with 758 BCG consultants tested AI’s effect on realistic business tasks, comparing no AI, GPT-4, and GPT-4 plus prompt-training groups. Productivity gains inside the frontier (Priority: 4/5): On tasks well suited to AI, consultants with GPT-4 performed better, produced higher-quality work, and completed more of the assigned work, especially when trained in prompt usage. Failure on outside-the-frontier tasks (Priority: 5/5): On tasks requiring careful integration of multiple information sources and human judgment, AI users were more likely to get the wrong answer by over-relying on the model, even though they worked faster. How successful users work with AI (Priority: 4/5): The study identifies two effective collaboration styles: 'centaurs,' who divide tasks between human and machine, and 'cyborgs,' who weave AI into sub-tasks interactively. Broader labor-market and training implications (Priority: 4/5): The episode connects the findings to freelancer displacement, chatbot errors, and the possibility that junior workers may lose opportunities to build expertise if routine tasks are increasingly automated.

Key Arguments: Generative AI is useful, but usefulness depends heavily on task type; it is not uniformly reliable across comparable-sounding problems. The jagged technological frontier explains why AI can solve difficult-seeming tasks yet fail at basic ones like counting or tic-tac-toe decisions. Training users in prompt engineering improves performance, but only modestly and mainly when the task is within AI’s capabilities. AI increased output quality and completion rates for consultants on inside-the-frontier tasks. Less skilled workers benefited more from AI than top performers, suggesting AI can narrow skill gaps in some settings. On outside-the-frontier tasks, AI access caused more wrong answers because users deferred too much to the model. Successful AI use requires human validation and active judgment, not blind acceptance of outputs. AI may devalue some work, especially in freelancing and entry-level knowledge jobs, while changing rather than eliminating human labor.

Data Points: NVIDIA market impact: Biggest single driver of returns in the S&P 500 this year - Used to illustrate the scale of AI enthusiasm in markets Potential automation of business activities: Up to 70% - McKinsey estimate for generative AI’s automation potential across occupations by 2030 BCG consultant sample size: 758 - Participants in the Harvard/BCG experiment Experiment groups: 3 - No AI, GPT-4 access, GPT-4 access plus prompt-engineering training Inside-the-frontier performance gain with AI: 38% better - AI users versus control group on the task AI was suited for Inside-the-frontier performance gain with AI + training: 42.5% better - Trained AI group versus no-AI group on the inside-the-frontier task Task completion rate, control group: 82% - Average completion of assigned tasks without AI Task completion rate, AI group: 91% - Average completion rate with GPT-4 access Task completion rate, AI + training group: 93% - Average completion rate with GPT-4 plus prompt training Performance boost for top performers: 17% - Highest performers on the control task improved with AI access Performance boost for bottom-half performers: 43% - Lower-skilled consultants saw much larger gains from AI access Outside-the-frontier accuracy, control group: 85% correct - No-AI group on the difficult strategic recommendation task Outside-the-frontier accuracy, AI users: 60% correct - GPT-4 users on the task requiring human judgment beyond the model Outside-the-frontier accuracy, AI + training group: 70% correct - Prompt-trained GPT-4 users on the same task Fabrication rate on federal court cases: 69% - Stanford study finding for ChatGPT Fabrication rate on federal court cases: 88% - Stanford study finding for Meta’s Llama2 Freelance market impact timeframe: Within a few months of launch - Short-term effects paper on generative AI and employment Air Canada chatbot case: 1 lawsuit - Example of legal and reputational risk from erroneous chatbot advice

Pivotal Quotes: "Generative AI has hit a tipping point." — Jensen Huang: Referenced at the start of the episode to frame market enthusiasm around AI chips and adoption "the jagged technological frontier" — Patrick Boyle: Core concept used to describe uneven model performance across seemingly similar tasks "The study concludes that while AI can boost the performance of highly skilled knowledge workers, the best approaches to using AI are not yet fully understood and need to be studied before implementing these models carelessly in the workplace." — Patrick Boyle: Summarizes the main lesson from the Harvard/BCG experiment

Implications: AI is likely to raise productivity on some knowledge-work tasks but also increase error risk where human judgment matters. Firms should train workers, validate outputs, and redesign workflows carefully rather than adopt AI blindly.

🔓 Sign Up for Unlimited Episode Search

About Patrick Boyle on Finance

This podcast is all about quantitative finance and financial history. Subscribe to hear about financial markets, derivatives, and how investors use quantitative tools from statistics and corporate finance theory. Included are interviews with some of the most interesting thinkers in finance. Occasional longer form financial documentaries, open up fascinating elements of financial markets history. Patrick Boyle is a quantitative hedge fund manager, a university professor, and a former investment banker. To contact Patrick visit http://onfinance.org Find Patrick on YouTube at: https://www.youtube.com/c/PatrickBoyleOnFinance

View all episodes from Patrick Boyle on Finance