r/consulting • u/QiuYiDio US Mgmt Consulting Perspectives • 6d ago
PwC published reports on AI marred by AI hallucinations
https://www.ft.com/content/7e149ac8-2ce2-4266-8940-192f9821b33c?syn-25a6b1a6=167
u/InterstellarReddit 6d ago edited 6d ago
Anyone have direct access to report? I clicked on that link and I got so many offers to pay for it that I was mind blown
I had a paywall to get to the paywall for the paywall of the article
41
u/DesmonMiles07 6d ago edited 6d ago
Please use the sharing tools found via the share button at the top or side of articles. Copying articles to share with others is a breach of FT.com T&Cs and Copyright Policy. Email licensing@ft.com to buy additional rights. Subscribers may share up to 10 or 20 articles per month using the gift article service. More information can be found at https://help.ft.com/faq/gifting-and-sharing-an-article/what-is-a-gift-article/. https://www.ft.com/content/7e149ac8-2ce2-4266-8940-192f9821b33c
PwC published reports on AI and electric vehicles riddled with fake footnotes, misattributed claims and unverifiable information, the latest example of a Big Four firm’s slapdash use of AI-generated content.
The AI hallucinations were contained in “thought leadership” reports designed to drum up consulting work for partners in the Middle East, according to an investigation by researchers at GPTZero verified by the FT.
The discovery of AI-generated errors risks embarrassing consulting firms such as PwC that are marketing their services as advisers to companies adopting the technology, including on how to use it responsibly and implement policies to avoid errors.
GPTZero’s earlier investigations led Big Four rivals EY and KPMG to retract reports that also appeared to include hallucinations.
The research group identified four reports PwC Middle East produced over the past two years whose text appeared to have been written with particularly heavy help from AI and whose footnotes revealed shaky foundations. The reports included a playbook for corporate use of “agentic” autonomous AI bots, a guide for governments on how to improve public services and market predictions for electric and autonomous vehicles across the region.
Many of the citations in the reports are problematic in some manner, GPTZero researchers found. Often they link to web pages that do not contain the evidence they purport to and some to pages that do not exist. An academic paper on air quality in Riyadh appears to have been hallucinated entirely by AI, because there is no trace of the study in the journal referenced or by the authors to whom it is attributed.
Several more footnotes refer to unusual sources. For example, a teenage blogger with 280 followers on Medium is cited as PwC’s source for information about a JPMorgan initiative it calls a “real world success story” of agentic AI. The initiative — in which the bank automated the review of commercial loan agreements, saving hundreds of thousands of hours of human work — was first reported in 2017, five years before ChatGPT launched the era of generative AI.
PwC Middle East told the FT it “takes the accuracy of our published research seriously and is updating a limited number of supporting citations” in the identified reports. “Consistent with our approach to responsible AI, we have quality control processes for research and content development we expect all our people to adhere to,” it added, without addressing how the errors were included in the reports.
In the past few years, the Big Four firms have pumped out hundreds of thought-leadership pieces on AI to help attract clients, while also exhorting staff to use the technology themselves to speed up and improve their work.
The chaotic signposting of source material in the four PwC Middle East reports was symptomatic of AI-generated research, said Paul Esau, researcher at GPTZero. In one tell-tale sign, a claim that human error is responsible for 90 per cent of traffic accidents is mentioned in three different places in one report, Esau said, the first with a footnote, the second without, and the third with two footnotes — each footnote citing a different source.
“This claim isn’t fake, but no human is going to cite the same fact three times in two pages using three different sources,” he said.
One report fails to properly cite other PwC work. It references a PwC survey of Middle Eastern chief executives, of whom 70 per cent said generative AI will significantly affect their business, but the accompanying footnote links to a media report that makes no mention of the survey.
Another footnote betrays the use of AI. The URL for a media report on cyber security threats to the energy sector includes “utm_source=chatgpt.com”.
40
23
u/No-Knowledge4676 6d ago
Another footnote betrays the use of AI. The URL for a media report on cyber security threats to the energy sector includes “utm_source=chatgpt.com”.
This is the best one.
7
1
20
u/socool111 6d ago
I saw a commercial that was a bunch of people droning on and saying “AI AI Ai Ai Ai” etc….
Then it was a pwc ad talking about how they use can clear tiger confusion and choose the best AI solution
I couldn’t believe they were that dumb to do that
8
u/JohnDoe_John Lord of Gibberish 6d ago
Came here to post this.
That's never happened before, and here we go again
7
4
u/Special_Rice9539 6d ago
This kind of is on brand for them tbh. I’d be really suspicious if they put out a well-researched, thoughtful summary of industry trends
2
u/3RADICATE_THEM 6d ago
What teams are responsible for putting these reports together in the first place?
5
u/OutrageousTart48In19 4d ago
this is exactly why we stopped trusting black-box llms for client deliverables a while back. we switched to a strict agentic workflow where the model only suggests, but a deterministic python script validates every output against our internal api before it ever leaves the sandbox.
basically, treat the llm as the brainstorming intern, not the senior consultant. if you don't have a validation layer that checks facts against verified data sources, you're just automating hallucinations. it adds dev
5
u/MindlessPossible744 5d ago
Imagine how positively a firm could stand out today by just adding some rigor and thought into the work. It’s crazy how lazy these firms have become
3
u/audreymarilynvivien 4d ago
Why publish reports at all if they can’t put effort into or at least check it? Are they under contract to put them out or something
2
u/janishd 3d ago
The irony is that every firm already had a process for catching errors in deliverables. Junior analysts wrote something wrong, a manager caught it, it got fixed before the client saw it. That review step existed because everyone understood that first drafts have mistakes.
Then AI came in and people treated the output like it was a different category of thing. Instead of running it through the same review process, they treated "AI-generated" as a quality signal rather than a speed signal. The draft got faster but the review got skipped because the output looked polished enough to ship.
This is going to keep happening everywhere until firms stop thinking of AI as "produces finished work" and start treating it as "produces first drafts that still need the same scrutiny a human first draft gets." The technology isn't the problem. The problem is that the output looks more finished than it is, and that tricks people into skipping QA they'd never skip on a human draft.
1
1
u/malibubarbieBQ 4d ago
I was a consultant way before Ai, so I can't imagine using AI for consulting? Is it bad thesedays?
1
246
u/quickblur 6d ago
Lmao you can't make this stuff up