Part III · Chapter 10
Bone, Color, and Music
What the skeleton leaves out
For nine chapters we've worked with the bone. It was worth it: with skeletons and classes you can see kinships that used to be invisible. But every method focuses, and in focusing it leaves things out. Someone who looks only at the skeleton of a word can stop hearing its music: its tone, its rhythm, the feeling it's said with.
This part of the manual brings back what we left out. The discipline that studies it is called psychophonology: the study of the sound of language as something that happens in three places at once, in structure, in time, and in the mind. It's a proposal of mine, and it works in alliance with endolinguistics. Endolinguistics studies the bone; psychophonology studies the flesh and the voice.
One "no," three times
Let's start with an experiment you can do right now. Say the word no in three ways:
- A dry, hard, short no. A door slamming shut.
- A long no that hesitates, that rises and falls. A no that's still making up its mind.
- A small, soft, almost whispered no. A no that's half a plea.
It's the same word: an n and an o. For grammar, all three are identical. And yet whoever hears you knows perfectly well which one you said, and knows something about you that the word itself doesn't say.
Where is the difference? Psychophonology says that every sound of language lives on three planes at once:
- The plane of structure: which sounds they are, in what order. Here the three no's are the same.
- The plane of music: rhythm, length, the melody of the voice, where it rises and where it falls. Here all three are different.
- The plane of the mind: the intention, the emotion, the tension of the person speaking, and what it stirs in the person listening. This is where the difference comes from.
The three planes aren't layers you can peel apart and study separately. They're three ways of looking at a single act. When someone says no, they say all three things at once, and the listener receives all three at once.
Consonants draw, vowels color
In chapter 2 I told you that consonants are like an outline and vowels like a color. Now we can take that image seriously.
A consonant is an event: a cut, an edge, an instant. That's why consonants can be counted and combined, and why codes are built from them. A vowel is a state: the mouth held open in a shape, sustained, sounding. Vowels don't cut the air: they color it.
And vowels can move without jumps. Between ee and ah there are endless in-between positions; you can glide from one to the other without a break. Between p and t, on the other hand, there's nothing: either your lips close or they don't. Consonants are like the keys of a piano; vowels are like the sound of a violin, which can slide up little by little.
That's why psychophonology says the vowel is an operator of quality. It doesn't change a word's skeleton: it gives it a color. Three vowels mark the corners of that color space, and they're the same three at the corners of every vowel chart linguists use (written here as they sound in Spanish or Italian):
- A (as in father): the mouth at its widest. Amplitude, openness, outward.
- I (as in see): the tongue high and forward, the mouth nearly closed. The small, the sharp, the precise.
- U (as in moon): the tongue high and back, the lips rounded. The deep, the dark, the round.
All the other vowels live between those three.
Bouba or kiki?
That vowels have color isn't only an endolinguistic idea. It's one of the most solid findings in the experimental psychology of language.
You almost certainly said the round one is bouba and the spiky one is kiki. You're not alone: people who speak very different languages, and people of very different ages, give the same answer in the vast majority of cases. The round, open vowels of bouba go with roundness; the i and k of kiki, with sharpness.
There's an older experiment. In 1929 the linguist Edward Sapir invented two words, mil and mal, and told people that both were names of tables: one big and one small. Which one is big? Most people picked mal for the big table and mil for the small one. The open a goes with big; the closed i, with small.
Now look at two word-pieces you surely know, from Greek: macro (big, large-scale) and micro (small). They have the same skeleton, m · k · r. Only the vowel changes, and the vowel seems to carry the size: a for big, i for small.
These tendencies aren't laws. They're leanings: you notice them when you look at many words and many people, and there are plenty of exceptions. But they're real, and they say something important: in the deepest layer of language, where the body makes the sounds, sound and sense aren't completely separate.
Rhythm decides what gets heard
Let's go from the color of a single vowel to the music of a whole sentence.
Languages have different rhythms. Spanish tends to give each syllable roughly the same time: ca-ra-me-lo sounds like four even beats. English tends to squeeze the unstressed syllables between the stressed ones: caramel gets compressed, and one vowel nearly disappears. The same bone, put into a different rhythm, sounds different.
That's why, when you speak Spanish with an English rhythm, even if you pronounce every sound correctly, people can tell. You squeeze syllables where Spanish keeps them even; you swallow vowels Spanish wants you to say in full. What we call an accent isn't just putting a vowel wrong here or a consonant wrong there. It's that your inner rhythm, the rhythm of your first language, keeps sounding underneath the new one.
Endolinguistics calls that inner rhythm endorhythm. It isn't the rhythm you can record with a microphone, but the underlying disposition all your rhythms come from: a tempo, a way of moving through time, as personal as the timbre of your voice. You carry it with you from one sentence to the next and from one language to another.
When the music is the word
In some languages the music doesn't accompany the word: it is part of the word. They're called tonal languages. In Mandarin Chinese, the syllable ma said on a high, level tone means "mom"; with a rising tone, "hemp"; with a tone that dips and rises, "horse"; with a sharply falling tone, "to scold." Four different words with the same sounds. What a consonant or a vowel would do in English, the pitch of the voice does in Mandarin.
And here's a beautiful discovery from historical linguistics. Many tones were born when a language lost a consonant. The consonant disappeared, but the difference it used to mark wasn't lost: it moved into the melody of the vowel next to it. It's called tonogenesis, the birth of tone. It's as if the bone, wearing away, had turned into music.
What the voice says without words
There's an even finer level: the small tremors of the voice, the uneven pauses, the throat tightening a little, the almost imperceptible hurry. Psychophonology calls it micromusicality. It sits below what we consciously notice, and yet we read it instantly.
Think of a trembling voice. If the person is cold, the tremor is even, mechanical, rhythmic: the body shivering. If the person is holding back tears, the tremor is different: irregular, with uneven pauses, with tension gathering in the throat. You can hear both. Both are tremors. And they mean opposite things. The difference isn't in any word or any sound: it's in the micromusicality.
What to do with all this when you learn a language
This part may sound theoretical, but it has very practical consequences.
- Pick a reference accent and stick with it at first. Don't mix the Spanish of Madrid with the Spanish of Mexico City in the same week. Your ear needs a stable model.
- Listen before you speak. The rhythm and melody of a language are learned by ear, not by eye. Play audio in the language even if you don't understand all of it: your body is learning the other language's endorhythm.
- Shadow. Play a recorded sentence and repeat it on top of it, almost at the same time, imitating the melody, not just the sounds. It's called shadowing.
- Exaggerate the music. At first, exaggerate the rises and falls of the new language, as if you were acting. Your English rhythm will keep pulling you toward squeezed, swallowed syllables; with Spanish, practice giving every syllable its full, even beat.
- Listen in parallel. Use the bridge sentences from chapter 8: hear the sentence in the new language and then the English bridge, one right after the other. You're not just training words: you're training the switch in rhythm.
In the next chapter we go to the deepest question of this part: how a thought, which makes no sound, becomes a voice.