Why human judgment — Not AI quality — Is the real driver of business performance
Generative AI has arrived with a compelling promise: an always-available, low-cost adviser that anyone with basic literacy can use. For entrepreneurs and small business owners — who often lack access to expensive consultants or mentors — this seems like a genuine breakthrough. But does it actually work?
A field experiment conducted with hundreds of small business owners in Kenya suggests the answer is far more complicated than the hype implies. The technology did help some entrepreneurs grow their revenues and profits. But it made things measurably worse for others. The deciding factor was not the quality of the AI’s advice — it was the quality of the human using it.
Why business support is so hard to scale
As explained here, helping entrepreneurs improve their performance at scale has long been one of development economics’ most stubborn challenges. The interventions that actually move the needle — one-on-one consulting, personalized mentorship, face-to-face networking — are labor-intensive, expensive, and nearly impossible to replicate at meaningful scale. In lower-income countries, the gap is even starker: qualified business advisers are scarce, and their fees are far out of reach for most small operators.
This is precisely why generative AI seemed so promising. A tool capable of fielding business questions in plain language, around the clock, at near-zero marginal cost, could theoretically democratize access to the kind of guidance that has historically been reserved for well-resourced companies.
The experiment
To put this idea to the test, researchers recruited 640 small business owners across a range of sectors in Kenya — including food and beverage, agriculture, and car-wash services — and ran a randomized controlled trial between May and November 2023. Half of the participants were given access to a version of OpenAI’s GPT-4, configured to act as a Kenyan business adviser and delivered through WhatsApp, the dominant messaging platform in the country. The other half received a conventional online business training guide.
The setup was deliberately open-ended. Entrepreneurs in the AI group could ask whatever they wanted, use the tool as frequently or infrequently as they chose, and act on the advice however they saw fit. This mirrors how AI tools are actually deployed in the real world — not as tightly scoped assistants with narrow mandates, but as general-purpose advisers capable of responding to almost any question.
Four out of five participants had never used ChatGPT or any similar AI tool before the study began.
A striking split beneath the average
At first glance, the results looked unimpressive: on average, access to the AI adviser had no statistically significant effect on business performance. But averages can be deeply misleading.
When researchers looked beneath the mean, they found a sharp divergence. Among business owners who had been performing well before the experiment — roughly the top half by prior business outcomes — AI access led to revenue and profit gains of around 15%. Among those in the bottom half, however, AI access was associated with nearly a 10% decline in revenues and profits.
The same tool, the same quality of advice, produced meaningfully different outcomes depending on who was using it.
Same questions, different choices
To understand why, the researchers examined how different entrepreneurs actually used the tool. It turned out that high and low performers asked similar numbers of questions, raised similar topics, and received advice of comparable quality. The divergence came afterward — in what each group decided to act on.
The AI, like most general-purpose advisers, offered a mix of recommendations: some generic (lower your prices, spend more on advertising) and some more tailored to the specific business and situation. Low performers tended to act on the generic suggestions. They cut prices. They increased advertising expenditure. These moves eroded profit margins and raised costs without generating enough additional business to compensate.
High performers took a different approach. They used the AI to surface ideas specific to their circumstances and were more skeptical of cookie-cutter advice. A cybercafe owner began renting out gaming accessories. A car-wash operator introduced a new detergent that customers were already asking for and started selling cold sodas to people waiting for their vehicles. Another entrepreneur found alternative power sources to work around electricity outages. Each of these improvements emerged from a combination of AI-generated ideas and the owner’s own contextual knowledge — their understanding of their customers, their costs, and their competitive environment.
The AI did not do this work for them. It offered possibilities. The entrepreneur still had to judge which possibilities were worth pursuing.
AI amplifies judgment — For better or worse
The central finding of this research is not that AI is good or bad for entrepreneurs. It is that AI, in open-ended contexts, amplifies whatever judgment the user brings to the table. For those with strong intuition and business acumen, it extends their reach, helping them discover options they might not have considered. For those whose judgment is weaker or whose grasp of their own business situation is less secure, it can actively mislead — lending false credibility to advice that sounds sensible but does not fit the reality they are operating in.
This is a meaningful departure from what researchers have found in studies of AI applied to narrow, well-defined tasks — drafting emails, reviewing code, generating marketing copy. For those kinds of tasks, AI tends to benefit weaker performers the most, because the task itself is specific enough that the output can be used with minimal filtering. Running a business is not like that. It requires constant judgment calls about context, timing, and fit. AI cannot reliably make those calls on its own, and users who cannot make them either will not benefit from AI assistance in this domain.
Anthropic’s own internal experiment with Claude managing a small vending business illustrates the point: left to its own devices, the AI sold items at a loss and quickly drove the business into the red. The lesson is not that AI is useless for business — it is that AI needs a capable human in the loop to translate its suggestions into sound decisions.
What this means for leaders and policymakers
For anyone responsible for deploying AI tools at scale — whether in a company, a government program, or a development initiative — these findings carry several practical implications.
Don’t rely on averages. When evaluating whether an AI tool is working, average outcomes can obscure serious harm to specific groups. A tool that improves outcomes for high performers by 15% while damaging low performers by 10% might look fine on paper if you only look at the mean. Disaggregated analysis is essential.
Design with heterogeneity in mind. High-performing users with strong judgment can often benefit from open-ended AI access. Weaker performers may need more structured support — tools that incorporate contextual data about the user’s specific situation (their financials, their customer base, their competitive landscape) so that generic advice is filtered out before it reaches them. Building this kind of context-awareness into AI products remains a work in progress, but it is the right direction.
Invest in the human side. AI tools are only as good as the people using them. Organizations that want to deploy AI effectively should also invest in building the judgment and literacy that allow users to engage with AI output critically — to recognize when advice is inapplicable, to ask better questions, and to know when to seek human expertise instead.
Audit for unequal effects. Leaders should regularly examine three dimensions: whether different groups are adopting the tool at different rates; whether different groups are interacting with it differently (asking different questions, providing different amounts of context); and whether the tool is producing different real-world outcomes across groups. Each of these is a potential lever for intervention.
The bigger picture
The ease with which the researchers in this study deployed a capable AI business adviser — via WhatsApp, in a matter of weeks — speaks to how quickly and cheaply these tools can now be rolled out. That accessibility is genuinely exciting. But it also means that poorly designed deployments can cause harm at scale just as easily as well-designed ones can create value.
The risk is not merely that AI fails to help some users. The risk is that it makes performance gaps wider — strengthening those who are already strong while undermining those who are already struggling. In a business context, in a development context, in almost any context where people are trying to improve their outcomes, that is the opposite of what the technology is supposed to do.
AI’s potential to raise performance at scale is real. So is its potential to concentrate that uplift among people who need it least. Recognizing that tension — and designing deliberately to address it — is the central challenge for anyone serious about putting these tools to good use.

