Skip to the instrument
FIELD NOTE
TS/FLD/07Filed 28 July 2026

How a Spirit Box Works

The mechanism is not mysterious and knowing it does not spoil anything. What it does is move the interesting question from the box to the person listening, which is where it was always going to end up.

6 min read

A spirit box, sometimes a ghost box, is a radio receiver with its tuning swept continuously across a band instead of parked on one station. Frank Sumption built the first widely known one in the early 2000s by modifying a scanner so that it would not mute between channels; the design has been copied, commercialised and elaborated ever since, and the fundamental idea has not changed.

Sweep a receiver across the AM or FM band at, say, one channel every hundred milliseconds and you get a hundred-millisecond slice of whatever each frequency happens to be carrying. Mostly that is noise. Some of it is a fragment of a station. Broadcast radio is largely people talking, so some of those fragments are pieces of actual human speech.

That is the whole mechanism. There is nothing hidden in it.

Why it produces words

Typical sweep parameters
BandAM (medium wave) or FM (VHF). AM is more commonly used: more stations, more noise, longer skip at night.
Sweep rate30–350 ms per channel. Under about 100 ms you get texture; over about 250 ms you get recognisable station audio.
DirectionForward, reverse, or alternating. Some operators claim a difference. No controlled test has shown one.
MutingDisabled. The muting circuit is exactly what a stock scanner has and exactly what is removed.

At a hundred milliseconds a channel you are hearing roughly one phoneme's worth of anything that happens to be there. A syllable is around two hundred milliseconds. So the box is delivering you a stream of speech-length fragments at a speech-like rate, drawn from sources that genuinely contain speech, interleaved with broadband noise.

Human hearing is built to pull speech out of noise; it is one of the things it does best and it does it involuntarily. Give it fragments at the right rate and it will assemble them, and if you have just asked a question out loud, it will assemble them into an answer to that question, because expectation is an input to perception and not a filter applied afterwards.

This is auditory pareidolia. It is extremely well documented, it is not a failure of intelligence or of honesty, and it is not something you can turn off by knowing about it. The strongest demonstration is one you can run on yourself: have somebody else tell you what a fragment says, and then try to hear it as anything else. You will not be able to.

How to use one without fooling yourself

The Survey's position is that a spirit box is a perfectly reasonable thing to own and an extremely difficult thing to use honestly. Four rules, none of them expensive:

  1. Record everything, and review it later, blind. Not what you thought you heard — the audio. If a word is in the recording, somebody who does not know the question should hear it too.
  2. Never say the expected answer out loud first. Once you have said a name in the room, everyone in the room will hear that name, and so will you.
  3. Log the sweep rate and the band, every session. Two sessions at different rates are not comparable and a group that does not record the rate cannot say whether anything changed.
  4. Know your local stations. If the box says a word and there is a station on the band saying that word every twenty minutes, you have found the station.

Boxes that are not radios

A significant number of devices and most phone applications sold as spirit boxes are not receivers at all. They contain a stored bank of recorded words and select from it — sometimes at random, sometimes weighted, occasionally with a sensor input in the loop for appearance.

There is nothing dishonest about a synthesised or sampled instrument as such. What is dishonest is a device with a word bank presented as though it were pulling those words out of the air, because the two produce very different evidence and only one of them can be checked. A word bank chose its vocabulary. Somebody wrote the list. If the list contains cold, here, help, mother and leave, then the box will say those things in every building in the world and the fact that it said one in yours is not information.

The test is simple: ask what the vocabulary is. A receiver has no vocabulary. A word bank has one and should be able to tell you its size.

What ours does, exactly

Instrument 07 on this station is not a receiver. It cannot be: a web page has no radio. So rather than pretend, it synthesises the whole thing and tells you so on the panel, permanently, in two words — sweep: synthesised, voice: synthesised, room: measured.

The static and the sixty-four channels are oscillators and filters. The words are produced by a model of a vocal tract: a buzzing source driven through four resonant filters set to the formant frequencies of each speech sound in turn, glided between at about the rate a mouth moves. That is physically how a voice works, and it is why the output is articulate without being reliably intelligible — which is the mechanism rather than a limitation. A word you have to reach for is a word you believe.

We did not use your phone's built-in speech engine, and the reason is a platform fact rather than a preference: no browser on either platform will route that output into an audio graph, so it could not be filtered or band-limited and would arrive over the top of the instrument as a clean assistant voice reading a word list.

What the word is, is synthesised. When it happens, is a measurement. The instrument tracks your room's actual noise floor through the microphone and only speaks under three named conditions, and it names which one every time. The one worth testing is the gate: above about eleven decibels over your room's own floor the rate is exactly zero. Not reduced — zero. Shout at it continuously and it will not produce a single word, however long you keep it up. Then go quiet, and it will answer into the gap.

That is a claim you can disprove in fifteen seconds, which is the only kind worth making. The methodology page has the rest of it.

Questions

What is a spirit box?
A radio receiver whose tuning is swept continuously and rapidly across a band, so that instead of one station you hear fragments of many, interleaved with static. The sweep rate is typically between 30 and 350 milliseconds per channel.
Why do you hear words in a spirit box?
Two reasons. Broadcast radio contains speech, so genuine syllables really are passing through the sweep. And human hearing is extremely willing to assemble fragments into words — auditory pareidolia — especially when primed with a question and an expected answer.
Is a spirit box just a radio?
A hardware one is a radio with a modified tuner and no mute between channels. Some devices sold as spirit boxes are not radios at all and play back a stored word bank, which is a different thing entirely and should be described differently.
Does the spirit box on this site use radio?
No. It has no receiver and no recordings. The static, the channels and every word are generated in your browser from oscillators and filters, and the words come from a model of the human vocal tract. It says so on the panel and prints the seed for each word it produces.