AI coaching

What to look for in an AI public speaking coach: 7 questions before you trust the score

Most AI coaches hand you confident numbers. The useful ones tell you how each number is made and where the check stops.

· 9 min read · The StageMirror team

A woman seen from behind rehearses a pitch to a laptop in a glass-walled meeting room at dusk, one hand raised mid-gesture, the kind of take an AI public speaking coach would review.
AI-generated illustration.

An AI public speaking coach is worth trusting when it tells you how each number is made and where its checking stops. Before you upload a pitch, ask seven questions. Is every metric defined, with a target? Are camera figures labelled as estimates? Does the coaching quote your words? What happens when the analysis fails? What does it refuse to score? Where does your video go, and for how long? What does the free plan include? Plain answers mean you can practise against it. Vague ones are a warning.

Why confident numbers aren’t enough

A readout like “Pace 148, eye contact 64%, overall 7.8” looks like measurement. Behind it is a chain. Speech becomes text, the text is counted and timed, a camera model guesses where your eyes were, and a language model writes about the result. Each link can fail.

The first link matters most. In a 2020 PNAS study, Koenecke and colleagues tested five commercial speech recognisers on interviews with 115 speakers. The average word error rate was 0.35 for Black speakers and 0.19 for white speakers. The lesson outlasts those systems: a wrong transcript makes every number built on it wrong, silently. So don’t ask how accurate a tool is. Ask how each number is made, and whether it tells you when it couldn’t make one.

QuestionA good answerA red flag
1. Defined metrics?Fillers over words, pace with its time base, a target band.A score out of 100, no formula.
2. Estimates labelled?An estimate label, the frames read, a coverage cut-off.Eye contact to the decimal, no caveat.
3. Coaching quotes you?Notes cite your words or a time; unsupported ones are dropped.Advice that fits anyone’s talk.
4. Failure handled?Says what failed, keeps the recording, no second charge.A full scorecard whatever happened.
5. Refusals listed?No score for confidence, charisma or funding odds.A “confidence score” from your face.
6. Data explained?Named processors, no training, a retention window, deletion.“Your data is safe with us.”
7. Free spelled out?Takes a month, features, playback window, card or not.“Free” that is really a trial.
The seven questions at a glance.

Questions 1 and 2: defined metrics, and estimates labelled as such

1. Is every number defined, with a target band?

A usable metric says what is counted, what it’s divided by and what range is good. Filler rate should be filler words over total words, from a published list. Is “like” counted in a comparison? Is a mid-sentence “so”? Pace should say what it divides by: the whole take, first word to last, or only the seconds you speak. A speaker who pauses well looks slower on the first two. And look for bands, not “higher is better”. Our posts on pitch pace and trading “um” for a pause explain why.

2. Are camera numbers labelled as estimates, with coverage shown?

A webcam guesses eye contact from head angle and iris position, gestures from wrist movement and posture from body landmarks. Glasses, low light and a low camera make the guess worse. So ask: is the figure labelled an estimate everywhere? Does it show coverage, the share of frames actually read? Does a low-coverage figure still count towards your score? Eye contact read from a tenth of your frames describes your lighting, not your eyes. Our posts on eye contact on camera and your hands show how these estimates are made.

A stack of blank white index cards held by a black binder clip on an oak conference table, a silver pen beside it and a steel stopwatch standing behind, softly out of focus.
A stopwatch only helps if you know what it timed. The same goes for every number on a readout. AI-generated illustration.

Question 3: does the coaching have to quote your words?

“Strong opening, but watch your pacing in the middle” sounds specific and fits almost any talk. Ask for every claim to point at evidence: a phrase you said, or a time you can jump to. Then ask whether that evidence is checked or merely requested, because a model told to cite quotes can still invent one. The best sign is a tool that drops failed claims and tells you how many.

Also ask where the check stops. A summary is often written from the numbers rather than from quotes, and a real quote still doesn’t prove the model read the context right. A vendor who names those gaps earns more trust.

Questions 4 and 5: failure, and what it refuses to score

4. What happens when the analysis fails?

It will, sometimes: no face found, clipped audio, a service timing out. A good tool says which part failed, shows what did work, marks missing items as unmeasured rather than failed, keeps your recording, and doesn’t charge you twice. A complete scorecard every time is the red flag: if you never see a gap, something is filling it.

5. What does it refuse to score?

Words, seconds and pauses can be counted, and head direction roughly estimated. Confidence, charisma, authenticity, story quality and whether an investor will write a cheque can’t. Emotion is the clearest case. A 2019 review led by Lisa Feldman Barrett found that the way people express anger, disgust, fear, happiness, sadness and surprise varies substantially across cultures, across situations, and even between people in the same situation. Since 2 February 2025, the EU AI Act has also prohibited AI systems that infer people’s emotions in workplaces and education institutions, outside medical or safety uses. Ask for the list of what a product won’t score.

An empty wooden lectern with a gooseneck microphone under a single spotlight on a small dark stage, seen past two rows of empty folding chairs.
Whether a talk lands is decided in the room. AI-generated illustration.

Question 6: where does your video go, and for how long?

A rehearsal can hold your revenue, your customers’ names and your face. In the privacy policy, look for four things.

  • Named processors. Who else receives your audio, video or transcript? An outside speech-to-text service or language model is normal. Silence about it isn’t.
  • Training. Is your data used to train anyone’s models? A model provider’s terms apply too. OpenAI, for example, says data sent to its API isn’t used for training unless the customer opts in.
  • Retention. A window measured in days, which you can shorten, beats “for as long as your account exists”.
  • Deletion. Can you delete one take, or your whole account, yourself?

Question 7: what does the free tier actually include?

Some “free” plans are trials that renew into a subscription. Others lock the useful parts. Four questions settle it.

Free-plan questions

  • How many analyzed takes a month, and what happens after the limit?
  • Are the camera figures and the coaching included?
  • How long can you watch a recording back?
  • Is a card needed, and does anything renew by itself?

A test you can run on any AI public speaking coach

Feed the tool a take where you already know the answers.

StageMirror’s answers, limits included

We build StageMirror, so this is a vendor answering its own checklist.

To run the rigged-take test on us, record your first take and check these answers against your own readout.

Frequently asked questions

Can an AI public speaking coach replace a human coach?

It does different work. An app counts fillers, pace and pauses the same way every time. A person is better at judging whether your story works for a particular audience. Use the app for repetitions and a person for judgement calls.

Are AI eye contact scores accurate?

They are estimates, worked out from head angle and iris position, and glasses, low light and a low camera make them worse. Check the coverage: a figure read from few frames says more about your setup than about your eyes. Where to look on camera.

Does an AI speech coach work if English isn’t my first language?

It can, but transcription accuracy varies by accent and every count rests on the transcript. Read it after each take. Where it is wrong, the numbers are wrong too, and that is the tool’s limit, not a verdict on how you speak.

Is there a free AI public speaking coach?

Free plans differ widely, so check what each one includes. StageMirror’s free plan gives you 5 analyzed takes a month with every metric and the coaching, with no card and no trial.

Sources

  1. Racial disparities in automated speech recognition Koenecke, Nam, Lake, Nudell, Quartey, Mengesha, Toups, Rickford, Jurafsky and Goel, PNAS 117(14), 2020
  2. Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements Barrett, Adolphs, Marsella, Martinez and Pollak, Psychological Science in the Public Interest 20(1), 2019
  3. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 5 European Union, Official Journal, 2024
  4. Data controls in the OpenAI platform OpenAI API documentation, accessed September 2026

More from the blog

AI public speaking coach: 7 questions to ask · StageMirror