Filler words

How to stop saying “um”: trade every filler for a pause

You can’t will an “um” away mid-sentence, because it is usually doing a job: telling a listener you need a moment. Keep the moment and drop the sound: a silence held longer than feels comfortable.

· 9 min read · The StageMirror team

A man stands at a kitchen island at sunrise, lips closed and one palm turned up at the counter’s edge, beside a phone on a slim tripod, a steaming mug and a closed notebook.
AI-generated illustration.

You decide, before the pitch, that you will not say “um”. Then the first time you have to find a word in the middle of a sentence, it comes out anyway, and now you are thinking about that too. The “um” was doing a job, and banning the word doesn’t cancel the job. What works is handing the job to something else: a silence, held on purpose, for longer than feels comfortable.

Why you say “um”, and why willpower doesn’t stop it

The most useful account of fillers comes from the psycholinguists Herbert Clark and Jean Fox Tree. In a 2002 paper in Cognition they proposed that speakers use “uh” and “um” to announce a delay they can see coming: “uh” for a minor one, “um” for a major one. On their reading, a filler is a word you plan and produce like any other, and it tells the listener something specific: hold on, I’m looking for a word, or I’m deciding what to say next.

In the London–Lund corpus of recorded British English conversation, “um” was followed by a further delay 61% of the time and “uh” 29% of the time, which is what you would expect if “um” announces the longer wait. It is a proposal with corpus data behind it, not a law of speech. A second study, by Heather Bortfeld and colleagues, found that disfluency overall rose with planning difficulty but fillers didn’t track it neatly, and its authors suggest fillers also serve the conversation, signalling a delay to the listener.

Fillers can turn up anywhere in a pitch, but they bunch at the joins. In Clark and Fox Tree’s corpus, “uh” and “um” came about three times as often at the start of a phrase as late in it, where you are still choosing what to say next. The start of an answer is one of those joins, and plenty of answers open with “so” or “good question”. The post on investor questions deals with that moment.

That is why forbidding the word rarely works. The delay is still there, with nowhere to go, so it is likely to come out as a stall, a restart or a different filler.

How to stop saying “um”: replace it with a pause

Finish the sentence, close your lips, and hold the silence until the next thought is ready. Then start it. The closed lips are a coaching trick, not a research finding: they give the pause a start and an end you can feel, which makes the silence easier to hold on purpose.

Then hold it for longer than feels right. StageMirror counts a gap as a deliberate pause only when it runs longer than 1.5 seconds; a gap of exactly 1.5 doesn’t qualify. A slow, silent “one thousand one, one thousand two” gets you past it.

Does the silence cost you anything with the listener? One lab study gives a narrow answer. In experiments by Martin Corley and Robert Hartsuiker, people pressed the button for one of two pictures faster when its name came straight after a delay of about a second, whether the delay was an “um”, silence or an artificial tone. The gains were tens of milliseconds. Silence, it suggests, gives the listener the same beat an “um” does. The study didn’t test longer silences, or whether silence makes you look more competent.

Where the silence goes matters more than how many you manage. A pause after a point gives it room. The same silence before a point sounds like searching. Day 3 of the 14-Day Founder Speaker Intensive, “Silence instead of filler”, asks for three placed pauses: after the problem, after your strongest number, and just before the ask.

The target is 2 to 3 deliberate pauses a minute, so the same habit lowers your filler rate and raises your pause count. It also lowers your words per minute, because silence takes time, and Day 3’s recording still expects 130160 words a minute with the silence included. The post on pitch pace covers that band.

A grey-haired woman pauses mid-presentation in front of a dark wall screen, two seated colleagues blurred in the foreground.
A pause after the point, not before it. AI-generated illustration.

What counts as a filler, and what doesn’t

To count your own fillers, or to trust anyone else’s count, you need the rules. Here are StageMirror’s, as the code applies them.

WordCountedNot counted
“um”, “uh”, “er”, “ah”, “actually”, “basically”, “literally”Every time, even when you mean the word.Never exempt.
“you know”, “I mean”, “kind of”, “sort of”Every time. Each word in the phrase counts, so one “you know” adds two.Never exempt.
“like”Everywhere else: “customers like it”, “I’d like to show you”, “we were like, scrambling”.After is, looks, feels and similar verbs (or things, something, stuff), when a word such as a, the, this, my or it follows: “it looks like a market”.
“so”As the first word of the take, or after a gap of more than half a second: “So, we raised…”Mid-sentence: “we grew so fast”.
“right”Followed by a gap of more than half a second, or as the last word: the tag question “…right?”Mid-sentence: “the right time”.
“and”, “well”, repeated words, restartsNever.Not on the list at all.
StageMirror’s filler rules, taken from the lexicon and the counting code.

The “like” rule is the crudest. It reads words literally: “it looks like a market” is exempt, while “it’s like a market” is counted, because “it’s” is not on the verb list. “Like” as an ordinary verb counts too. If your number looks high, search your transcript for “like” before you blame your delivery.

Hedges are a separate list: “maybe”, “perhaps”, “probably”, “possibly”, “just”, plus “I think”, “I guess”, “a bit”, “kind of”, “sort of”. They are counted but not scored against a target. “Kind of” and “sort of” sit on both lists, so they count as fillers regardless.

Why the target is under 2%, not zero

Ordinary speech is full of disfluency. In Bortfeld’s study of pairs doing a matching task, speakers produced 5.97 disfluencies per 100 words, counting repeats, restarts and fillers together. Fillers alone came to 2.56 per 100 words. Two cautions: that was conversation, not a pitch, and their filler list was only um, uh, er and ah, far narrower than the one above. Treat it as context, not a benchmark: ordinary adults talking freely are nowhere near zero.

Reading a script aloud can get you close to zero. In a talk you are composing as you go, chasing zero can swap fillers for restarts and a stiff delivery, which is no improvement. So StageMirror’s bar is under 2%. A three-minute take at 130160 words a minute runs to roughly 390480 words, which leaves room for no more than about 7 to 9 fillers in the whole take.

The 14-Day Founder Speaker Intensive tightens the bar in two steps. The Day 7 checkpoint’s rubric item reads “Filler rate under 3%”; by Day 14 it is “Filler rate under 2%”. Both are scored on the web, automatically, from your recording’s stored analysis rather than by you, and both are strict: a rate exactly on the line misses.

There is one place where zero is the goal. Day 4, “First fifteen, last fifteen”, asks for no fillers at all in the opening and closing fifteen seconds. Those windows are short and scripted, so zero is realistic. You count them yourself on playback, because the readout gives one rate for the whole take. The day-by-day rehearsal plan covers the rest of the two weeks.

How to measure your own filler rate

You can do this by hand. Record three minutes of your pitch, get a transcript and correct it against the audio, because automatic transcripts can leave fillers out. Count fillers by the rules above, count all the words, and divide. For pauses, scrub through the recording and note each silence between words that runs past 1.5 seconds.

A phone held sideways in a clamp on a small tripod, sharp in the foreground, records a softly blurred speaker in a dim, lamp-lit room.
Count it from the recording, not from how the take felt. AI-generated illustration.

StageMirror does the counting on the phone and on the web; the free plan analyzes 5 takes a month. The readout shows your filler rate next to your deliberate pauses per minute, colour-coded against the 2% target: good up to 2%, a warning up to 4%, bad beyond that. Exactly 2% counts as good there, though not on the Day 14 rubric. Progress plots both numbers across takes.

Don’t take the number on trust. OpenAI’s transcription guide lists preserving filler words among the uses of a transcription prompt, and StageMirror sends no prompt. The authors of CrisperWhisper, a Whisper variant retrained for word-for-word transcripts, cite earlier work showing that the original drops many filler words. So some of your “um”s may never reach the transcript, and your rate can read low.

The drill below is adapted from Day 3 of the Intensive, which runs on the web and in the phone app.

Then read the number for what it is. A filler rate counts words in a transcript. It can tell you that you said “you know” eleven times. It can’t tell you what the room made of it.

Sources

  1. Using uh and um in spontaneous speaking Herbert H. Clark and Jean E. Fox Tree, Cognition 84, 73–111, 2002
  2. Disfluency rates in conversation: effects of age, relationship, topic, role, and gender Heather Bortfeld, Silvia D. Leon, Jonathan E. Bloom, Michael F. Schober and Susan E. Brennan, Language and Speech 44(2), 123–147, 2001
  3. Why um helps auditory word recognition: the temporal delay hypothesis Martin Corley and Robert J. Hartsuiker, PLOS ONE 6(5): e19792, 2011
  4. File transcription OpenAI API documentation
  5. CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions Laurin Wagner, Bernhard Thallinger and Mario Zusag, Interspeech 2024
  6. Model card: Whisper OpenAI, openai/whisper on GitHub

More from the blog

How to stop saying “um”: trade every filler for a pause · StageMirror