Why a broken radio sounds like speech
A spirit box is a radio that refuses to settle. It sweeps across the band faster than any station can hold, so instead of a broadcast you get fragments: a syllable from one station, a consonant from the next, a burst of noise between them.
Your ear will not accept that. Speech perception is aggressive — it is built to pull words out of a noisy room full of other people talking, and it does not switch off simply because there is nothing there to pull. Given fragments at roughly the rate of speech, it will assemble words. Given a question, it will assemble an answer.
The effect has a name. It is auditory pareidolia, it is extremely well documented, and it is stronger when you already know what you are hoping to hear. That last part is worth sitting with before you use this instrument.
Instrument 07 sweeps a bank of synthesised sources rather than live radio, because rebroadcasting live radio is not ours to do and, more to the point, because the effect has never depended on the source being real.
Everything you hear is made here
There is no audio file anywhere in this instrument. The static, the stations, the whistles and every word are generated by your own browser, in real time, out of oscillators and filters. That is a plain statement of fact and it is the first thing to know about the instrument.
The words are made with a model of a vocal tract. A buzzing source — a wavetable shaped like the pulse a pair of vocal folds actually produces — is fed through four resonant filters, and the frequencies of those filters are driven to the formants of each speech sound in sequence. That is not an approximation of how a voice works; it is how a voice works. A vowel is its first two or three formants, and moving them is what turns one sound into another.
We could have used the speech engine built into your phone. We did not, for one decisive reason: its output cannot be reached. No browser will route it into the audio graph, so it cannot be filtered, swept or put behind the receiver at all — it would arrive over the top of the instrument as a clean assistant voice from a different piece of software. A synthesised tract lives inside the receiver, comes through the same bandwidth limit as everything else, and sounds the same on every device.
It is also articulate rather than intelligible, and that is the point rather than a limitation. A word you have to reach for is a word you believe.