Pitch tells you which note a voice is singing. Formants tell you who is singing it. They are the resonant frequency regions created mainly by the shape of the vocal tract, and they carry both vowel identity and the apparent size of the person producing the sound.
Shift pitch without doing anything about them and the whole vocal tract seems to shrink or grow. Shift them on purpose and you change the character of a singer while the musical note stays exactly where it was. Formant shifting is that second operation, and it is the most useful vocal control most producers never touch.
Quick answer: pitch shift changes the note, formant shift changes the apparent size of the singer. Keep formants preserved for natural transposition and harmony, shift them by -2 to -5 semitones for darker ad-libs or +2 to +4 for brighter ones, and let them move with the pitch when you want the varispeed sound.
What a formant actually is
When you sing, your vocal folds produce a buzzing tone at the pitch of the note. That buzz is harmonically rich and, on its own, sounds nothing like a voice.
The sound then travels through your throat, mouth and nose, and that cavity resonates at particular frequencies. Those resonant peaks are formants. They boost some harmonics of the buzz and suppress others, and the pattern they produce is what your brain decodes as a vowel and as a particular person.
Two facts follow from this and they explain everything else in the article:
Formants are mostly independent of pitch. You can sing “ah” at any note in your range and the formant pattern stays broadly similar, because your mouth shape has not changed. That is why you recognise a singer across their whole range.
Formants are determined by physical size. A larger vocal tract resonates lower. That is why a child and an adult singing the same note still sound obviously different, and it is why moving formants reads as changing the size of a person rather than changing the note.
Pitch shift versus formant shift
| Operation | The note | Apparent character | Typical use |
|---|---|---|---|
| Pitch shift only | Changes | Changes too, usually unintentionally | Old-school transposition, the chipmunk sound |
| Pitch shift with formant preservation | Changes | Tries to stay put | Natural transposition and harmony |
| Formant shift only | Stays | Changes | Character, ad-lib contrast, size effects |
| Both together | Changes | Changes with it | Tape and varispeed style effects |
The second row is what a modern pitch corrector does by default when it is doing its job well, and the third row is the one worth learning.
Why two semitones already matters
The chipmunk effect is obvious at an octave, which is why people assume formant drift is only a problem at extremes. On an exposed vocal it is audible far earlier.
A harmony voice shifted up a third with no formant preservation already sounds like the lead singer’s head got smaller. Nobody listening consciously identifies it, but the part reads as artificial and gets buried in the mix to hide it.
The same is true of correction. A tuner moving a note two semitones without holding the formants makes the singer subtly change size from word to word. That inconsistency is one of the things people are hearing when they say tuning sounds “processed” but cannot say why.
The reverse is also true. Perfect formant preservation at an extreme shift can sound detached and unnatural, because a real voice an octave higher would have different formants. The best setting is not always maximum preservation. It is whatever suits the layer.
Practical starting ranges
| Goal | Formant shift | Notes |
|---|---|---|
| Natural harmony | 0 to +/-1 st | Preservation on. Vary each layer slightly |
| Subtle double separation | +/-0.3 to 1 st | Add timing and pan differences too |
| Brighter ad-lib | +2 to +4 st | Watch for thinness and sibilance |
| Darker ad-lib | -2 to -5 st | Watch for mud and heavy plosives |
| Creature or robot | +/-6 to 12 st | Artifacts are the point. Combine with talkbox |
NEOTUNE Pro offers formant shifting to +/-12 semitones with six voice models and throat modelling, sitting in the same window as the correction. If you want a dedicated vocal-tract modeller instead, Antares Throat is the long-standing one. Start at zero, tune first, then change the character. See NEOTUNE Pro
Formants and generated harmony
Generated harmony voices start life as copies of one source, which means they share one throat. Identical timing and identical formants across four voices is exactly what gives the trick away.
Give the upper parts slightly brighter formants and the lower parts slightly darker ones, but keep the moves small: half a semitone to two semitones. Then vary timing by 10 to 30 ms, vary the level by a dB or two, and pan them away from the lead.
Formant shifting alone will not fake real backing singers, because it cannot produce the phrasing differences between two human beings. It is one of four tools, not a solution. More on this in vocal harmony plugins.
Formants in rap, drill and hyperpop
Rap and drill ad-libs. Lowering the formants by 2 to 5 semitones creates size and menace without touching the melody, which means the ad-lib still sits correctly in the chord. This is the single most effective ad-lib trick that does not involve reverb, and it is covered with the rest of the chain in UK drill vocals.
Hyperpop. Pitch and formants routinely move independently so a very high musical line does not have to sound like plain varispeed. Shifting pitch up while pulling formants down gives you a high note from a large voice, which does not exist in nature and is precisely why the genre uses it. The full stack is in how to make hyperpop vocals.
In both cases, automate by phrase and check the result in the full mix. A formant effect that is spectacular soloed can vanish entirely behind drums.
Common artifacts and what causes them
Metallic or hollow vowels. The shift is too extreme for the algorithm, or the source is not clean. Reduce the amount or improve the input.
Phasey, smeared highs. Usually a stereo or reverberant source. Formant processing wants a clean mono monophonic signal.
Consonants smeared or lisping. Sibilance and plosives are noise, not pitched resonance, and they do not survive heavy formant processing well. Split them off: duplicate the vocal, high-pass the duplicate around 5 kHz, and blend the unprocessed consonants back over the shifted body.
Unstable, wobbling tone. The pitch detector is losing the fundamental. Fix that first. If the vocal warbles at zero formant shift, the formant control is not your problem. See why does autotune sound warbly.
Muddy low shifts. Lowering formants pushes energy downward. High-pass after the shift, not before.
Common mistakes
Shifting formants before the tuning is stable. Detection first, character second, always.
Using formant shift as a gender switch. It moves size cues. Identity also lives in range, articulation, breath and phrasing, which one control cannot touch.
Shifting the lead and the double by the same amount. Then they are still the same person. Differences are the point.
Reverb before the shifter. Formant processing on a reverberant source produces the phasey artifacts above. Effects after.
Leaving preservation on for a deliberate varispeed effect. Sometimes you want the chipmunk. Turn it off.
Frequently asked questions
What is formant shifting? Moving the resonant frequency regions of a voice, which changes the apparent size and character of the singer while the musical note stays the same.
Does lowering formants lower the pitch? Not with an independent formant control. The note stays where it is while the resonant envelope moves down, so the singer sounds larger without singing lower.
What is the difference between pitch shift and formant shift? Pitch shift changes which note is sung. Formant shift changes who appears to be singing it. Doing both together gives you the tape or varispeed sound.
Is formant shifting a gender changer? It moves size and gender-coded cues, and it is often convincing for a couple of semitones. A voice’s identity also depends on range, articulation, breath and phrasing, so do not expect a complete transformation from one control.
Should formant preservation always be on? Use it for natural transposition and harmony, then compare by ear. For varispeed or exaggerated effects, letting the formants move is the sound you want.
How many semitones of formant shift is too much? Beyond about 5 semitones most algorithms start producing audible artifacts on an exposed vocal. That is a limitation on natural results, not on creative ones.
Why do my harmonies sound like one person? They share one throat. Vary the formants slightly per voice, offset the timing by 10 to 30 ms and pan them apart.
Why does formant shifting sound metallic on my vocal? The source is probably not clean or monophonic. Remove reverb and doubles upstream, high-pass the rumble, and make sure detection is stable before shifting.
NEOTUNE Pro shifts formants to +/-12 semitones with six voice models and throat modelling, in the same window as the correction so you can hear both decisions together. $29.99. See the plugin.
Related: vocal harmony plugins and autotune settings for rap vocals.
The post Formant Shifting Explained appeared first on Producer Sources.