Can't catch spoken English?
You've never met its weakened sounds.
You know the words. You can read the sentence. And yet, in real conversation, it slips right past you. In most cases the problem isn't vocabulary — in real speech, little words like of and to collapse into weak, blurry vowels. This guide lets you hear it happen in real YouTube clips: play each example, slow it down, and replay it as many times as you need.
August 15, 2026 ・ VoiceGlyph
Reduction: real examples
Let's go straight to a real example.
Example: kind of
First, press play. This is from a Tokyo travel vlog, where the speaker recommends a pastry shop in Shinjuku as a great way to kind of break up your shopping day. Pay attention to how the two words kind of sound in the middle of the sentence.
On paper, the speaker says “kind of.”
But the sound that actually comes out is:
kind of = “kinda”/ˈkaɪn.də/
What happens here is very simple.
- of gets reduced (weakened) — of is a little function word that carries no stress, so it isn't pronounced clearly. Instead of “of,” it collapses into a weak, blurry “uh.”
- The /d/ of kind links into that “uh” — the /d/ connects straight into what's left of of, making “da.” The two words go by as one shape: “kinda.”
As long as you're listening for “of,” it never shows up in the sentence. Kind of (meaning “somewhat” or “sort of”) is filler-level frequent in everyday speech — English even spells the reduced form: kinda. It's worth knowing by its real sound.
Now that you know the sound shape “kinda,” play the card again. Tap kind or of inside the card to replay from that word.
Example: supposed to
One more, this time with to collapsing. In an Osaka travel vlog, the speaker heads to a nighttime teamLab garden, saying it's supposed to be a different vibe from the one in Tokyo. Listen for the two words supposed to.
On paper, she says “supposed to.”
But the sound that actually comes out is:
supposed to = “supposta”/səˈpoʊs.tə/
Let's break it down step by step.
- to gets reduced (weakened) — to is another little function word. Instead of “to,” it collapses into a weak “tuh.”
- The /d/ at the end of supposed disappears — the /d/ is swallowed by the /t/ right after it, and the s connects straight into “tuh.” What goes by is “supposta,” in one breath.
Reduced to is everywhere: it's why have to sounds like “hafta” and used to like “yoosta.” Supposed to (meaning “expected to” or “meant to”) is an everyday expression for plans and expectations — worth knowing by its real sound, not just its spelling.
Now that you know the sound shape “supposta,” play the card again. Tap supposed or to to replay from that word.
So how do you start hearing it?
You know the phrase kind of — everyone does. And still, “kinda” slips past you in conversation. The reason is simple: you had never met that sound shape. And here is the tricky part — a sound you don't know just goes by as “something I missed.” You can't even notice what you don't know.
So the work splits into two stages.
Stage 1: Find out what you don't know.
Check the real sound shape of the part you couldn't catch, and recognize: “I never knew this sound.” The listen-then-learn flow of this guide is built to give you that first stage on the spot.
Stage 2: Repeat until it's familiar.
Unfortunately, knowing alone doesn't make you hear it. You need to listen to the real audio again and again, turning a sound you know into a sound you're used to. That's the stage where the parts you couldn't hear start coming through.
About VoiceGlyph
We build a tool for doing exactly these two stages — learn the sound shape, then replay until it's familiar — right on top of YouTube. When you hit a moment you can't catch in a video you're watching, you can turn those few seconds into a card, keep it with a translation and a listening tip, and replay it any time. The cards embedded in this article are the actual product: play from any point on the waveform, slow it down, tap a word, and check it until it clicks.
See VoiceGlyph
More sound changes
The audio and stills in each card are quoted from videos published under a Creative Commons (CC BY) license. Every card credits the channel and links to the exact scene in the original video.