
Connected Speech: Rules, Examples, and Why Word-by-Word Fails
Connected speech patterns explained with IPA — weak forms, linking, assimilation, elision, the flap — and practice steps for hearing and producing each one.
Connected speech is the set of patterns native speakers use to link, reduce, and reshape sounds across word boundaries. It's one major reason conversational English can sound so different from dictionary-style word-by-word pronunciation.
Picture this: it's your third week at a new job. A colleague leans over and says, "Hey, wanna grab lunch? There's a place round the corner, kind of a hidden gem." You catch "lunch" and "corner." The rest dissolves into noise. You smile and say yes, hoping for the best. Your English score was B2. You knew every word in that sentence. And yet.
You've put in the work. You know your vocabulary. You've studied IPA. You can pronounce every word correctly in isolation. And yet native speakers still tilt their heads, ask you to repeat yourself, or switch to slower, simpler sentences when they talk to you. That gap between knowing English and sounding English is usually connected speech.
This article breaks down exactly what connected speech is, walks you through five core patterns behind it, and shows you the difference between what a dictionary tells you a sentence sounds like and what it actually sounds like coming out of a native speaker's mouth.
The "Word-by-Word" Problem
Imagine you're reading a sentence aloud: "What are you going to do?" You've heard this sentence a thousand times. You know every word. So you say: /wʌt ɑːr juː ˈɡoʊɪŋ tuː duː/. Clean, careful, one word at a time.
A native speaker says: /wʌɾərjə ˈɡoʊnə duː/, or even just /wʌɾəjə ˈɡɒnə duː/ in casual speech.
Those are not the same sentence to the ear. The word-by-word version sounds like a language learner reading from a script. The connected version sounds like a person having a conversation. The difference isn't vocabulary, grammar, or even accent. It's the mechanics of how sounds behave when words flow into each other at natural speed.
The core issue is that dictionaries teach citation form: how a word sounds when pronounced in isolation, slowly and deliberately. Real conversation never works that way. Native speakers don't pause between words; they compress, link, and reduce sounds constantly, and they do it without thinking. If you've been told "your English is good but sounds a bit unnatural," this is almost certainly why.
What Is Connected Speech?
Connected speech refers to the phonological processes that occur when words are spoken together in a continuous stream rather than in isolation. Native speakers use these processes largely automatically because they reduce articulatory effort while maintaining intelligibility. In plain terms, connected speech is faster, smoother, and more efficient.
The gap between citation form (how a word sounds alone) and connected form (how it sounds in a sentence) can be dramatic. The word "and" in isolation is /ænd/. In the phrase "bread and butter," most native speakers say /ˈbrɛd ən ˈbʌɾɚ/: the vowel reduces, the final /d/ often disappears, and the /t/ in "butter" becomes a flap. Three words, five phonological changes.
Textbooks under-teach this for a practical reason: connected speech is harder to represent on a page, harder to drill in a classroom, and harder to grade on a test. But if your goal is to sound natural and to understand native speakers at real conversational speed, it is the highest-leverage skill left to learn.
Key Connected Speech Rules
Weak Forms
Function words, the grammatical glue of English, have two pronunciations: a strong form used when the word is stressed or cited in isolation, and a weak form used in normal connected speech. The weak form usually involves the schwa /ə/, the most common vowel sound in English.
| Word | Strong Form | Weak Form | Example |
|---|---|---|---|
| the | /ðiː/ | /ðə/ | "the cat" → /ðə kæt/ |
| to | /tuː/ | /tə/ | "going to" → /ˈɡoʊɪŋ tə/ |
| and | /ænd/ | /ən/ | "bread and butter" → /ˈbrɛd ən ˈbʌɾɚ/ |
| of | /ɒv/ | /əv/ or /ə/ | "cup of tea" → /kʌpə tiː/ |
| can | /kæn/ | /kən/ | "I can go" → /aɪ kən ɡoʊ/ |
| for | /fɔːr/ | /fər/ | "wait for me" → /weɪt fər miː/ |
The rule is straightforward: grammar words are weak by default unless you're stressing them for emphasis. If you say "I can go" with emphasis on "can," you use the strong form /kæn/, because you're contrasting it with "can't." In normal speech, it's /kən/ every time.
Learners who always use strong forms sound overly deliberate, as if they're reading a legal document aloud.
Liaison (Consonant-Vowel Linking)
When a word ends in a consonant sound and the next word begins with a vowel sound, the consonant naturally links across the word boundary. There is no pause; the consonant becomes the onset of the following syllable.
- "an apple" → /ən ˈæpəl/ sounds like "a-napple"
- "stop it" → /ˈstɒpɪt/ sounds like "sto-pit"
- "take it easy" → /ˈteɪkɪˈtiːzi/ (the /k/, /t/, and /z/ all link forward)
This is why learners who insert micro-pauses between every word sound choppy even when their individual sounds are correct. The pause is the problem, not the pronunciation. Practice phrases as single phonological units rather than sequences of separate words.
Flap T (American English)
In American English, the /t/ sound between two vowels, or between a vowel and a syllabic /l/ or /r/, is typically realized as a flap /ɾ/: a very brief tap of the tongue against the alveolar ridge. It sounds closer to a soft /d/ than a /t/.
- "water" /ˈwɔːtɚ/ → /ˈwɔːɾɚ/
- "butter" /ˈbʌtɚ/ → /ˈbʌɾɚ/
- "better" /ˈbɛtɚ/ → /ˈbɛɾɚ/
- "what do you" /wʌt duː juː/ → /ˈwʌɾə jə/
In many North American accents, /t/ may be realized as a flap in unstressed-vowel environments like these, and producing a crisp /t/ in every position ("water," "butter") is one of the things that can make speech sound careful rather than conversational to American ears. The pattern varies with stress, speech rate, and speaker. Note that British English uses different patterns here; the glottal stop (see below) is more characteristic than the flap.
Reductions

Certain high-frequency multi-word combinations are compressed so often in natural speech that they have acquired informal spellings that reflect their spoken form. These are not lazy or incorrect. They are standard features of informal and even semi-formal spoken English.
| Written Form | IPA (citation) | Reduced Form | Informal Spelling |
|---|---|---|---|
| want to | /wɒnt tuː/ | /ˈwɒnə/ | wanna |
| going to | /ˈɡoʊɪŋ tuː/ | /ˈɡoʊnə/ | gonna |
| got to | /ɡɒt tuː/ | /ˈɡɒɾə/ | gotta |
| out of | /aʊt ɒv/ | /ˈaʊɾə/ | outta |
| kind of | /kaɪnd ɒv/ | /ˈkaɪndə/ | kinda |
| have to | /hæv tuː/ | /ˈhæftə/ | hafta |
Using these reductions in appropriate contexts, like casual conversation and informal speech, signals that you understand how English actually works. Avoiding them entirely in conversation makes you sound overly formal, even in situations where that formality is unwarranted.
Glottal Stop
The glottal stop /ʔ/ is produced by briefly closing the vocal cords rather than using the tongue. In English, it most commonly replaces /t/ in certain environments, particularly before consonants or in syllable-final position.
- "cat flap" → /kæʔ flæp/ (the /t/ is replaced, not released)
- "button" → /ˈbʌʔən/ (very common in both British and American English)
- "network" → /ˈnɛʔwɜːrk/
- "that person" → /ðæʔ ˈpɜːrsən/
The glottal stop is particularly prominent in British English, especially in London varieties, but it appears in American English too. Recognizing it is essential for listening comprehension. Many learners hear a gap or a click where they expect a /t/ and simply don't process the word correctly.
Assimilation
Sounds don't just link or drop. They also shift to become more like a neighboring sound. This is assimilation, and it's why "did you" rarely sounds like "did" followed by "you" at conversational speed.
- "did you" → /ˈdɪdʒu/ (the /d/ + /j/ merge into the affricate /dʒ/)
- "would you" → /ˈwʊdʒu/
- "this year" → /ðɪʃ jɪər/ (the /s/ shifts toward the palatal /ʃ/ before /j/)
Assimilation is entirely rule-governed and happens in careful speech too, not just casual speech. It's how the articulators move efficiently between adjacent sounds.
Elision
Elision is when a sound is dropped entirely rather than reduced or shifted. It's especially common when consonants cluster together at word boundaries, where pronouncing every sound would require unnatural effort.
- "next day" → /neks deɪ/ (the /t/ drops before the /d/)
- "last time" → /læs taɪm/ (the /t/ drops before the /t/ of "time")
- "give me" → /ˈɡɪv miː/ often realized as /ˈɡɪmiː/ in fast speech
Elision is not sloppy speech. It's a predictable feature of fluent English at normal conversational speed. Learners who pronounce every consonant in a cluster sound stilted, just as learners who drop too many sound unclear.
Standard IPA vs Spoken IPA: See the Difference
The difference between dictionary transcription and real speech becomes concrete when you see full sentences side by side. The table below shows the same sentences in Standard IPA (citation form, how a dictionary would transcribe each word individually) and Spoken IPA (how the sentence actually sounds in natural, connected speech).
| Sentence | Standard IPA | Fluent speech (what you'll hear) |
|---|---|---|
| I am going to get a cup of water. | /aɪ æm ˈɡoʊɪŋ tuː ɡɛt ʌ kʌp ɒv ˈwɔːtɚ/ | /aɪəm ˈɡoʊnə ˈɡɛɾə kʌpə ˈwɔːɾɚ/ |
| What are you going to do? | /wʌt ɑːr juː ˈɡoʊɪŋ tuː duː/ | /ˈwʌɾərjə ˈɡoʊnə duː/ |
| I want to get out of here. | /aɪ wɒnt tuː ɡɛt aʊt ɒv hɪər/ | /aɪ ˈwɒnə ɡɛ ˈdaʊɾə hɪər/ |
| She is kind of tired today. | /ʃiː ɪz kaɪnd ɒv ˈtaɪərd təˈdeɪ/ | /ʃɪz ˈkaɪndə ˈtaɪərd təˈdeɪ/ |
| Can you give me a hand with that? | /kæn juː ɡɪv miː ʌ hænd wɪð ðæt/ | /kənjə ˈɡɪvmiː ə hænd wɪð ðæʔ/ |
| He told me to stop it. | /hiː toʊld miː tuː stɒp ɪt/ | /hiː ˈtoʊldmɪ tə ˈstɒpɪt/ |
One note on the fine print: the right-hand column above shows the fullest phonetic detail, including the flap /ɾ/ and the glottal stop /ʔ/. IPAtranslator's Spoken IPA mode captures the changes that matter most for learners, weak forms, linking, and contractions, and deliberately writes /t/ as /t/ rather than /ɾ/ to keep the transcription easy to read. Treat the flap and glottal stop rows as ear training: that is what you will hear from native speakers, even when the simplified transcription keeps the standard symbol.
The gap between the two columns is the gap between sounding like a careful language learner and sounding like a fluent speaker. Standard IPA is not wrong; it's the correct transcription of each word in isolation. But Spoken IPA reflects what your ears actually receive in real conversation.
IPAtranslator's Spoken IPA mode generates a connected-speech transcription automatically, with weak forms, linking, and contractions applied for you. There is no need to work through each rule manually. See our guide to using IPAtranslator to learn more.
Why This Matters for Fluency
There are two directions this affects you, and both matter equally.
For listening comprehension, if you've been trained to expect dictionary pronunciation, native speech at normal speed will sound like a wall of sound with no clear word boundaries. When someone says /ˈwʌɾərjə ˈɡoʊnə duː/, your brain is searching for /wʌt ɑːr juː/, and it won't find it. Connected speech rules are the decoding key. Once you know them, that wall of sound starts to resolve into recognizable units.
For speaking, the issue is the opposite: your pronunciation may be technically correct but rhythmically wrong. English has a strong stress-timed rhythm: stressed syllables come at roughly regular intervals, and unstressed syllables (including all those weak forms) are compressed to fit. If you give every syllable equal weight and duration, your speech sounds robotic and metronomic regardless of how accurate your individual sounds are. Connected speech is what creates the characteristic rhythm that makes English sound like English.
The confidence dimension matters too. Many intermediate learners develop a frustrating loop: they speak carefully to be understood, careful speech sounds unnatural, native speakers adjust their own speech in response, and the learner never gets exposure to real connected speech. Understanding these rules breaks that loop.
This matters well beyond casual conversation. Connected speech is present in interviews, pitches, and presentations too, just usually with slightly less reduction than relaxed speech. If you speak every word fully separated and stressed, the way a dictionary would render it, you can end up sounding stiff even when every word is technically correct. If you're preparing a specific talk, our guide to speaking in public confidently walks through that scenario in depth, and our deep dive on how to pronounce "to" shows the weak-form pattern on a single high-frequency word.
How to Practice Connected Speech
A pattern that comes up constantly among intermediate learners: they spend months drilling individual IPA sounds, reach a point where every symbol is recognizable, and then sit down with a native-speaker podcast, and still can't parse a sentence at normal speed. The sounds are right. The words are right. But the stream doesn't resolve into language.
The reason is usually the same: they practiced citation form, not connected form. They trained their ears on dictionary pronunciation and then encountered speech that follows different patterns. Citation forms remain useful for learning unfamiliar words and building a sound inventory. The gap isn't ability; it's that connected forms also need deliberate practice.
This is why the practice method below starts with connected speech rules from the beginning, not after you've "mastered" individual sounds. Waiting until your sounds are perfect before addressing connected speech is like learning to read individual letters perfectly before attempting words. The skill you actually need is the integrated one.
The most effective approach is to isolate one rule at a time rather than trying to apply all five simultaneously. Spend a week focused only on weak forms, then move to liaison, then flap T. Layering rules gradually is far more effective than attempting a complete overhaul of your speech patterns at once.
Shadowing is the highest-leverage practice method available. Find a short clip of a native speaker (a podcast, a TV scene, a YouTube video) and repeat what you hear simultaneously, matching rhythm and reduction as closely as possible. Don't transcribe first; listen and mirror.
Use a Spoken IPA tool to make the invisible visible. Paste any sentence into IPAtranslator and compare dictionary-style IPA with the connected-speech cues it highlights: weak forms, linking, and contractions. Treat flap T, glottal stops, assimilation, and elision as listening concepts rather than features the tool fully transcribes. Our deeper comparison of standard IPA vs spoken IPA walks through the difference in more detail.
Record yourself reading the same sentence in citation form and then in connected form. Your ear will catch discrepancies that your mouth doesn't notice in real time. The gap between what you think you're saying and what the recording captures is often significant, and that gap is exactly where your practice energy should go.
Finally, practice with phrases, not words. Drilling "water" in isolation will never teach you the flap T. You need "a bottle of water," "the water's cold," "what do you want to drink?" The connected speech rules only apply across word boundaries, so isolated word practice won't transfer.
Frequently Asked Questions
No. Connected speech consists of recurring, well-documented phonological patterns of fluent English; linguists have studied them for decades. How strongly they apply varies by accent, speaker, emphasis, and speech rate. Calling it "lazy", though, misunderstands how language works. It is efficient, not careless.
Work on both in parallel, but don't wait until your individual sounds are "perfect" before starting connected speech; that moment never comes. A reasonable approach is to focus on connected speech rules once you're at B1–B2 level and can produce most individual sounds accurately.
Yes, significantly. The flap T /ɾ/ is primarily American; British English more commonly uses the glottal stop /ʔ/ in similar environments. Weak forms and liaison operate similarly in both varieties, but the specific reductions and their frequency differ. Be clear about which variety you're targeting.
Record yourself and compare to native speakers saying the same phrases. Alternatively, use a Spoken IPA tool to see which connected-speech patterns are likely in a phrase, then listen for them in your own recording. The tool helps you notice likely patterns. It is a study aid, not a judge of whether your pronunciation is right or wrong.
Many standard IPA tools show citation form only, with each word transcribed in isolation. IPAtranslator's Spoken IPA mode is designed to show connected-speech transcription for any sentence you input, focusing on weak forms, linking, and contractions; finer phonetic detail such as flap T and glottal stops is taught in guides like this one as a listening skill.
Conclusion
Fluency isn't about knowing more words. It's about understanding that English sounds different in motion than it does on a page, and training your ear and mouth to work with that reality rather than against it. Weak forms, liaison, flap T, reductions, and the glottal stop are not exceptions or corner cases. They are how English is actually spoken.
One way to keep going: take a sentence you said out loud today, run it through IPAtranslator's Spoken IPA mode, and see how its weak forms, linking, and contractions play out. Then use that transcription as your practice target.
Author

Categories
More Posts

ChatGPT IPA Prompts for Connected Speech: What Works, What Fails
Copy-paste ChatGPT prompts for connected speech IPA — and the places AI output tends to go wrong. Three prompts plus a verification workflow.


Why Language Models Struggle with Connected-Speech Transcription
LLMs tend to pattern-match IPA rather than apply phonological rules — why their connected-speech output slips on weak forms, assimilation, and boundaries.


How to Speak in Public Confidently: A 7-Step Guide
Rehearse out loud, mark your script's rhythm, steady your pacing and nerves — then learn how to speak in public confidently in 7 steps with a free-to-try tool.

Newsletter
Join the community
Subscribe for English pronunciation tips and product updates