ArticleAI coaching

Does AI Coaching Work? What the Evidence Actually Shows

An honest review of the evidence on AI coaching: meta-analyses, randomized trials, real limits — and how to use it well.

Search for « does AI coaching work » and you will mostly find two kinds of answers: vendors telling you it works brilliantly (theirs, especially), and skeptics telling you a machine can never coach. Neither camp cites much research. That is a problem, because the research exists — including randomized controlled trials — and it tells a more interesting story than either sales pitch.

We build an AI coach for a living, so read us with that in mind. Precisely because of that, we would rather show you the actual evidence, including the parts that limit what we can claim. Here is what the science says, study by study.

First: does coaching itself work?

Before asking whether AI coaching works, it is worth asking whether coaching works. For decades the honest answer was « probably, but the evidence is thin ». That changed with two meta-analyses.

Theeboom, Beersma and van Vianen (2014) pooled the results of coaching studies conducted in organizational settings and found significant positive effects across five outcome categories — performance and skills, wellbeing, coping, work attitudes, and goal-directed self-regulation — with effect sizes ranging from g = 0.43 to g = 0.74. In plain language: moderate to large effects, not rounding errors.

Jones, Woods and Guillaume (2016) ran a second meta-analysis focused on workplace coaching and found an overall positive effect (δ = 0.36), rising to δ = 1.24 for outcomes measured at the individual level. They also found something that matters enormously for this article: coaching delivered through non-face-to-face channels (phone, video, digital) still worked. The medium mattered less than people assumed.

What do numbers like these mean off the page? An effect size in the 0.4–0.7 range puts coaching in the same league as well-established organizational interventions — meaningfully better than doing nothing, visibly short of a miracle. In our own coaching practice, that matches what we see: coaching rarely transforms someone in a quarter, but it reliably moves the needle on specific behaviors — how a leader runs one-on-ones, prepares hard conversations, delegates, decides. Keep that calibration in mind, because it is the fair benchmark for AI coaching too. The question is not « does the AI produce miracles? » but « does it produce effects in the range coaching itself produces? »

So the baseline is established: structured coaching conversations reliably help people change behavior and reach goals. The question is whether an AI can hold up its end of that conversation.

The centerpiece: a randomized trial where an AI matched human coaches

The single most important study on this question is Terblanche, Molyn, de Haan and Nilsson (2022), published in PLoS ONE. It deserves a proper description, because its design is exactly what most AI coaching marketing lacks.

The researchers ran two longitudinal randomized controlled trials, each lasting ten months, with eight measurement points: a baseline survey, six monthly surveys, and a follow-up three months after the final session. In the first trial, 105 participants were coached by trained human coaches while 105 served as a control group. In the second, 134 participants worked with an AI chatbot coach while 134 served as controls — 327 participants completed all eight timepoints across the two studies.

The AI coach was Vici, a text-based chatbot running on Telegram. Two details matter here. First, Vici was a narrow, rules-based system built on goal theory — structured questioning in the spirit of the GROW model — not a large language model. Second, it did exactly one job: help people define and pursue a goal.

The results: both coached groups attained their goals at significantly higher rates than their control groups. And the effect sizes were almost identical — ηp² = .265 for human coaches, ηp² = .269 for the AI. On this outcome, in this design, a simple chatbot rivalled trained human coaches over ten months.

Two honest caveats before anyone over-claims. The outcome was goal attainment, specifically — the study does not show an AI matching humans on deeper outcomes like wellbeing or leadership transformation. And participants were volunteers pursuing self-chosen goals, a favorable scenario for any structured accountability tool. But within those boundaries, this is rare gold in a hype-saturated market: replicated, randomized, longitudinal evidence that AI coaching produces real results.

If you want to hear the lead researcher discuss this work in his own words:

What can AI coaching actually do — and what can't it?

Graßmann and Schermuly (2021) mapped this question conceptually in Human Resource Development Review, and their framework has held up well. Their conclusion: AI coaches can already support core coaching processes — structuring reflection, supporting problem-solving, tracking goals, prompting self-awareness — and they do some things better than humans: perfect availability, perfect patience, no judgment, no schedule.

Their limits list is just as instructive. AI coaches struggle where coaching shades into psychology: reading what is not being said, working with strong emotions, handling crises, and challenging a client who is comfortably deceiving themselves. A skilled human coach notices the flicker of hesitation before your rehearsed answer. An AI mostly works with what you type.

In our own practice building Lumia, we would add a limit the literature underweights: an AI coach is only as honest as the user. A human coach can cross-check your story against how you behave in the room. An AI, for now, cannot — which is why we anchor sessions in an Octagon self-assessment rather than letting the conversation float free.

Can you build a real relationship with an AI coach?

Decades of psychotherapy and coaching research agree that the working alliance — the bond and agreement on goals and tasks between coach and client — is one of the strongest predictors of outcomes. If humans could not form an alliance with an AI, the whole enterprise would have a ceiling.

A 2024 study in Frontiers in Psychology tested exactly this. Fifty-two graduate students with prior coaching experience completed a 60-minute coaching session on leadership, randomly assigned to either a human coach or what they believed was an AI coach (an avatar — in reality simulated, a standard research method for isolating perception). The working alliance ratings were statistically indistinguishable: 74.50 on average for human coaches, 72.73 for the AI condition (p = 0.48).

One single session, one study, a simulated AI — we would not build a cathedral on it. But it converges with a broader pattern in chatbot research: people form functional working relationships with conversational agents faster and more deeply than intuition predicts. We unpack the mechanisms, and the caveats, in our dedicated piece on the working alliance with an AI coach.

What about ethics and privacy?

An honest evidence review has to include the concerns, and the researchers themselves are the loudest voices here. Terblanche (2024), writing in The Journal of Applied Behavioral Science, frames AI coaching as a genuine opportunity to democratize development — coaching for the many, not just the executive floor — while flagging unresolved issues: data privacy, the absence of regulation, the risk of organizations deploying AI coaches as surveillance-adjacent tools, and the temptation to over-extend chatbots into territory (mental health, crisis) they are not built for.

We take these seriously enough to have written a full piece on what happens to what you tell an AI coach. The short version: before trusting any AI coach, know where your words go, who can read them, and whether your employer is in that list.

Who does AI coaching work for — and who should skip it?

Reading across the evidence, a fairly clear profile emerges.

AI coaching works well when:

  • The goal is concrete: prepare a difficult conversation, build a feedback habit, clarify a decision, hold yourself accountable to a plan.
  • You need frequency. Behavior change is a game of repetitions, and a coach available at 6:45 before the board meeting beats a brilliant coach available in three weeks.
  • Budget or hierarchy would otherwise mean no coaching. For the majority of managers who will never be offered an executive coach, the realistic alternative to an AI coach is nothing.
  • You are willing to be honest in writing. The tool amplifies reflection; it cannot extract truth you refuse to type.

Choose a human, or add one, when:

  • The issue involves trauma, clinical distress, or crisis. Coaching of any kind is the wrong tool; AI coaching doubly so.
  • The work is deep identity-level change — the kind where the coach's felt presence and lived experience carry the intervention.
  • You have a track record of gaming systems. An AI is easier to perform for than a person who watches your face.

A composite picture from our sessions, to make this concrete. A sales director has to tell a loyal, underperforming team member that the territory is being reassigned. A human coach would be ideal — and unavailable until the 14th. Instead: one evening session to untangle what is fear of conflict and what is legitimate concern, one 15-minute rehearsal the morning of the conversation, a 5-minute debrief after. The conversation happens a week earlier than it otherwise would have, and the pattern — avoidance dressed up as patience — is now named and trackable. Nothing in that sequence required emotional depth beyond what a well-designed AI handles; all of it required availability that no human coach offers.

We compare the two options dimension by dimension in AI coach vs human coach.

How do you use an AI coach well?

The evidence points to usage patterns that make the difference between a novelty and a practice:

  1. Bring a real situation, not a topic. « I need to tell Sara her promotion is delayed, and I have been avoiding it for nine days » beats « let's talk about communication ». The Terblanche trials worked because participants pursued specific goals.
  2. Work in short, frequent sessions. Fifteen focused minutes before a hard conversation, a five-minute debrief after. Frequency is the AI coach's structural advantage — use it.
  3. Let it know you. Generic advice is what you get when the coach knows nothing about you. This is why Lumia starts from your Octagon profile — your self-assessed patterns across eight leadership pillars — so a session about delegation connects to your relationship with control, not a textbook's.
  4. Close every session with a commitment. Goal-directed structure is precisely the mechanism the RCTs validated. A session that ends without a next action is a chat, not coaching.

We go deeper on all four in how you should use your AI coach, and on session rhythm in particular.

Our honest position: this is not human vs machine

Here is where we land, and notice that it is not the conclusion a vendor is supposed to write: the strongest configuration for most leaders is hybrid. Use an AI coach for what the evidence says it does well — frequency, structure, goal pursuit, rehearsal, always-on reflection. Use a human coach, when you can, for what the evidence says machines still cannot do — deep challenge, emotional attunement, the alliance that comes from being truly seen by another person.

The two are not rivals; they are different layers of the same practice. A leader who sees a human coach monthly and works with an AI coach daily gets more from both. And a leader who can afford neither the executive-coach fee nor the waiting list now has an option validated by randomized trials rather than testimonials — at a price we examine honestly in the ROI of AI coaching.

What the research still hasn't answered

Honesty also means naming the open questions, because they are real.

The RCT evidence comes from narrow systems. Vici was a rules-based, goal-focused chatbot. Today's AI coaches are built on large language models — far more conversationally capable, far less trial-tested. Capability has outrun evaluation: it is reasonable to expect LLM-based coaches to perform at least as well on structured goal work, but « reasonable to expect » is not « demonstrated », and we refuse to pretend otherwise.

Durability is partially known. Terblanche's trials followed participants for three months after the final session — solid by coaching-research standards. Whether gains hold at two or three years, nobody knows, for AI or human coaching alike.

The samples are not you, necessarily. Trial participants were volunteers pursuing self-chosen goals. Senior executives under acute pressure, skeptical conscripts sent by HR, people in genuine distress — all underrepresented. Effects observed in willing populations do not automatically transfer.

Novelty may inflate early results. Some of the engagement with a new AI tool is curiosity. The studies' longitudinal design mitigates this — ten months is long past novelty — but the field needs more replications before the question is closed.

None of these caveats reverses the picture. They bound it. A bounded, replicated, positive result is worth more than an unbounded promise.

The evidence at a glance

| Study | Design | What it found | |---|---|---| | Theeboom et al. (2014) | Meta-analysis of coaching in organizational contexts | Significant effects on performance, wellbeing, coping, attitudes, self-regulation (g = 0.43–0.74) | | Jones et al. (2016) | Meta-analysis of workplace coaching | Overall δ = 0.36; δ = 1.24 for individual-level outcomes; non-face-to-face formats still effective | | Terblanche et al. (2022) | Two 10-month longitudinal RCTs, 8 timepoints, 327 completers | AI chatbot matched human coaches on goal attainment (ηp² = .269 vs .265); both beat controls | | Graßmann & Schermuly (2021) | Conceptual review | AI handles structure, reflection, goals; limits in emotion, crisis, deep challenge | | Barger / Frontiers in Psychology (2024) | Single-session RCT, 52 participants | Working alliance ratings statistically equal for human and (simulated) AI coach | | Terblanche (2024) | Peer-reviewed commentary | Democratization potential is real; privacy and regulation gaps are unresolved |

Stop reacting. Start seeing. The evidence says an AI coach can genuinely help you do that — if you bring it real situations and honest answers. Lumia already knows your Octagon profile, so the work starts from your actual patterns rather than a blank page. Practice it with Lumia, your AI coach.

Frequently asked questions

Is AI coaching as effective as human coaching?

On goal attainment, the best available evidence says yes: in two 10-month randomized controlled trials (Terblanche et al., 2022), an AI chatbot coach matched trained human coaches, and both outperformed control groups. But that finding covers structured goal pursuit specifically. For deep emotional work, crisis situations, or identity-level change, human coaches retain clear advantages the research has not shown AI can match.

What is the strongest scientific evidence that AI coaching works?

Terblanche, Molyn, de Haan and Nilsson (2022) in PLoS ONE: two longitudinal randomized controlled trials over ten months with eight measurement points. Participants coached by the chatbot Vici attained their goals at significantly higher rates than controls, with an effect size (ηp² = .269) nearly identical to that of human coaches (ηp² = .265). It remains the reference RCT in the field.

What are the main limits of AI coaching?

Research points to consistent boundaries: AI coaches struggle to read unspoken signals, work with strong emotions, handle psychological crises, and challenge clients who are deceiving themselves. They also depend entirely on what the user honestly shares. Privacy and regulation remain open questions. For clinical distress or trauma, coaching of any kind — AI included — is the wrong tool.

Will AI replace human coaches?

The evidence suggests complementarity, not replacement. AI coaches win on availability, frequency, cost and structured goal work; human coaches win on emotional attunement, deep challenge and lived presence. Researchers like Terblanche frame AI as democratizing coaching for people who would otherwise get none. The strongest setup for most leaders combines both layers.

How do I get real results from an AI coach?

Bring specific, current situations rather than abstract topics; work in short, frequent sessions rather than rare long ones; use a coach that knows your profile and patterns so advice is contextual; and end every session with one concrete commitment. Goal-directed structure is precisely the mechanism the randomized trials validated — recreate it deliberately.

References

  • Theeboom, T., Beersma, B., & van Vianen, A. E. M. (2014). Does coaching work? A meta-analysis on the effects of coaching on individual level outcomes in an organizational context. The Journal of Positive Psychology, 9(1), 1–18. DOI: 10.1080/17439760.2013.837499
  • Jones, R. J., Woods, S. A., & Guillaume, Y. R. F. (2016). The effectiveness of workplace coaching: a meta-analysis of learning and performance outcomes from coaching. Journal of Occupational and Organizational Psychology, 89(2), 249–277. DOI: 10.1111/joop.12119
  • Terblanche, N., Molyn, J., de Haan, E., & Nilsson, V. O. (2022). Comparing artificial intelligence and human coaching goal attainment efficacy. PLoS ONE, 17(6), e0270255. DOI: 10.1371/journal.pone.0270255
  • Graßmann, C., & Schermuly, C. C. (2021). Coaching with Artificial Intelligence: Concepts and Capabilities. Human Resource Development Review, 20(1), 106–126. DOI: 10.1177/1534484320982891
  • Terblanche, N. H. D. (2024). Artificial Intelligence (AI) Coaching: Redefining People Development and Organizational Performance. The Journal of Applied Behavioral Science, 60(4), 631–638. DOI: 10.1177/00218863241283919
  • Barger, A. S. (2024). Artificial intelligence vs. human coaches: examining the development of working alliance in a single session. Frontiers in Psychology, 15:1364054. DOI: 10.3389/fpsyg.2024.1364054

Last updated: Jul 30, 2026

© 2026 Become Luminous. All rights reserved.