In the debate about artificial intelligence and leadership there is one sentence more uncomfortable than most promises of salvation: An average AI today may give a manager more precise feedback on their own communication than an average manager gives their people.
The sentence does not elevate any technology. Its appeal lies elsewhere: it shifts the question. The common question is whether AI replaces managers. The answer is unspectacular — no. The more productive question is: which parts of today's leadership work genuinely require human judgement, and which are repeatable patterns that simply nobody has observed systematically until now?
What AI actually makes visible
Katharina Lange and José Parra Moyano of IMD Lausanne built a language-model-based tool and trialled it with 167 managers from various industries and countries. The setup was deliberately plain: managers coached one another through real professional situations while the tool listened. Afterwards they asked for an assessment — which style did I use, which patterns repeat, what do you see that I do not?
That a language model can classify conversational structures is hardly surprising. What stood out was the size of the gap. Around fifty-five per cent of the responses fell into a learning zone where the feedback was both unexpected and useful: it challenged assumptions and exposed blind spots. For about ten per cent of the participating managers, by contrast, an irritation zone occurred — usually when the feedback contradicted the self-image.
The two figures do not belong to a single distribution: the first refers to responses, the second to people. The authors themselves report them inconsistently, which is why they are kept apart here.
These figures deserve a sober reading. They come from a single tool trial based on self-report, not from a controlled experiment: there is no control group, no published methods section and no peer-reviewed version. A model is only as good as its data, it can hallucinate, and it captures cultural nuance or unspoken signals poorly. The authors' conclusion is explicit: AI complements human expertise, it does not replace it.
An old problem, newly measured
Leadership research knew the finding long before language models. Since the work of Leanne Atwater and Francis Yammarino, self-other agreement has been an established indicator of leadership effectiveness. The most robust finding across many studies: Those who systematically overrate themselves are judged less effective — and respond more weakly to developmental feedback.
Leadership was long a field in which self-image, external perception and actual effect were hard to tell apart. Not out of incompetence, but because reliable, frequent, unvarnished feedback is expensive and rare. What language models change is not the insight that the gap exists. It is the price at which it becomes observable.
Leadership thus shifts from a quality attributed to a person to an observable behaviour within a system. The decisive question is no longer: how did I mean it? But: what effect did my behaviour produce?
Why the mirror is accepted — and where it is not
Research knows two patterns that point in different directions. In six experiments with around 1,900 participants, Logg, Minson and Moore described algorithm appreciation: people follow advice more readily when they believe it comes from an algorithm than when they attribute it to a person. Why that is, the paper does not establish. The plausible reading is that a model does not appear to judge and burdens no social relationship — listening to a machine requires no saving of face. That is an interpretation, though, not a finding.
For managers, a second result from the same paper is the more interesting one: experienced professionals who make forecasts regularly followed algorithmic advice less than lay people did — at the cost of their own accuracy. The more experience someone has, the worse they listen, precisely where listening would have paid.
Against this stands an opposing finding. In five studies with around 2,500 participants, Dietvorst, Simmons and Massey demonstrated algorithm aversion: once people see an algorithm make a mistake, they turn away faster and more decisively than they would from a human. That the rejection is stronger for tasks considered subjective or value-laden — precisely the terrain of leadership — comes from a third paper, by Castelo, Bos and Lehmann.
The findings do not contradict each other. They describe different conditions: as long as no error is visible, the willingness to follow the model prevails; afterwards it flips.
An uncomfortable ambivalence follows: the same manager who gratefully takes in feedback in the learning zone may reflexively dismiss it in the irritation zone. And the same apparent objectivity that builds trust can become a new form of diffused responsibility. Phrases such as “the data supports it” sound sober — and yet imperceptibly shift accountability from a person to a model.
Development or control — a design question
The same instrument can produce two very different social logics. A protected space for reflection, in which managers voluntarily analyse their communication patterns, can be extraordinarily valuable. A system that continuously measures, compares and rates managers will hardly produce better leadership. It will mostly produce conformity. People then speak more smoothly, act more cautiously, risk less.
From which it follows: AI governance must not be reduced to data protection, tool approvals and technical security. Those aspects are necessary but fall short. Governance must also settle which social logic emerges — whether the system encourages reflection or control, whether it increases accountability or shifts it. It must also settle who may see the analyses, whether they serve development or assessment, and whether participation is voluntary.
In practice this logic is decided earlier than many assume — in apparently technical choices. A mirror whose analyses only the manager sees produces development. The same mirror whose data flows into performance objectives produces conformity.
What leadership remains
The reason lies in a property of AI that gets lost in efficiency debates: It amplifies what is already there. An organisation with unclear decision paths does not become more decisive through AI. An organisation that covers up mistakes does not become better at learning through better analytics. The prior question is therefore not which processes get faster, but whether an organisation is mature enough to handle the transparency AI creates.
That is precisely why human judgement becomes more important, not less. An AI can analyse a sequence of conversation, but it does not know the history of a conflict. It can structure a decision option, but it does not carry the social consequences of a wrong decision.
Peter Drucker wrote decades ago that the first task of leadership is to lead yourself. AI does not make that task redundant — it makes it verifiable.
Whether AI replaces managers is therefore not the decisive question. What matters more is whether managers are willing to become more visible through AI. AI does not take responsibility away from leadership. It takes away the excuses.
A note on AI use: this article was written with AI support — not exclusively, but particularly for source research, source checking and linguistic quality assurance.
Sources
- Lange, Katharina / Parra Moyano, José (25 Feb 2026): How AI can coach us to improve how we communicate. I by IMD, IMD Business School, Lausanne. Tool trial with 167 managers; zone classification by self-report; not peer-reviewed. Source of the 55 per cent and 10 per cent figures. imd.org
- Lange, Katharina / Parra-Moyano, José (14 Feb 2025): Research: How AI Helped Executives Improve Communication. Harvard Business Review. Earlier report on the same study; full text paywalled. hbr.org
- Heron, John (2001): Helping the Client: A Creative Practical Guide. 5th edition, Sage, London. Six categories of intervention; model first published 1975.
- Atwater, Leanne E. / Yammarino, Francis J. (1992): Does Self-Other Agreement on Leadership Perceptions Moderate the Validity of Leadership and Performance Predictions? Personnel Psychology 45(1), 141–164. DOI 10.1111/j.1744-6570.1992.tb00848.x
- Fleenor, John W. / Smither, James W. / Atwater, Leanne E. / Braddy, Phillip W. / Sturm, Rachel E. (2010): Self–other rating agreement in leadership: A review. The Leadership Quarterly 21(6), 1005–1034. DOI 10.1016/j.leaqua.2010.10.006
- Lee, Angela / Carpenter, Nichelle C. (2018): Seeing eye to eye: A meta-analysis of self-other agreement of leadership. The Leadership Quarterly 29(2), 253–275. DOI 10.1016/j.leaqua.2017.06.002
- Logg, Jennifer M. / Minson, Julia A. / Moore, Don A. (2019): Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes 151, 90–103. Six experiments, around 1,900 participants. DOI 10.1016/j.obhdp.2018.12.005
- Dietvorst, Berkeley J. / Simmons, Joseph P. / Massey, Cade (2015): Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General 144(1), 114–126. Five studies, around 2,500 participants. DOI 10.1037/xge0000033
- Castelo, Noah / Bos, Maarten W. / Lehmann, Donald R. (2019): Task-Dependent Algorithm Aversion. Journal of Marketing Research 56(5), 809–825. DOI 10.1177/0022243719851788
Correction, 10 September 2026. Checking every figure against its primary source turned up four inaccuracies; they are corrected in the text above. The two percentages were presented as one distribution, although one counts responses and the other people. They were attributed to the authors’ Harvard Business Review piece rather than the freely available version at I by IMD. A finding on the rejection of algorithmic advice was attributed to the wrong papers. And one source that could not be substantiated has been removed. The article’s observation is unaffected.
This article offers a professional assessment and does not replace legal or management advice.