your voice
find the range you actually have, then run it through a rack. the pitch is read off the raw mic with the same algorithm a tuner uses, so it tells you what you sang rather than what you meant.
wear headphones before turning monitoring on. your mic hearing your own speakers is a feedback loop, and it gets loud fast. output starts muted for exactly this reason.
turn the mic on and sing anything.
a note only counts once you've held it steady for a few frames, so a cough doesn't become your record.
your range
sing as low as you can hold, then as high. the bar fills in as you go.
lowest
—
highest
—
span
—
octaves
—
the rack
reverb, delay and a bit of grit — the same three boxes behind most of what you've heard sung into a microphone.
wind the tone down to about 2 kHz and you get the telephone sound — that's not an effect so much as a bandwidth limit, the same one phone lines have had since they were built for speech and nothing else.
harmony
copies of your voice, shifted to sit above or below whatever you happen to be singing. in key, the harmoniser works out the right interval for each note instead of transposing everything by the same amount — so hold one note and you get the chord the key actually wants.
what's coming out — the harmony voices show up as extra peaks above your own note
in key, the interval changes with the note. a third above C in C major is E — four semitones. a third above D in the same key is F, only three, because that's what the key has. watch the numbers above move as you change note: that's the difference between a harmoniser and a transposer.
the looper
sing a phrase, stop, and it repeats forever. everything after that gets trimmed to the same bar and stacked on top — a bassline, then a rhythm, then a melody over your own backing.
the first take sets the bar length; every overdub waits for the top of it before it starts recording, so layers stay lined up instead of drifting. loops are captured after the rack, so whatever reverb you dialled in is baked in.
what a tuner actually does
finding the pitch of a voice is harder than it sounds. a sung note is a stack of harmonics, and the loudest part is often not the fundamental at all — so the obvious approach, picking the strongest repeating pattern, lands an octave out embarrassingly often.
this uses YIN: instead of hunting for the strongest match, it measures how badly the signal disagrees with a delayed copy of itself, then divides each candidate by the average of everything shorter. that division is the whole trick — it penalises the half-and-quarter-speed impostors that make naive detectors jump octaves, which is why the note above stays put while you hold a vowel.