A formant is a resonance of the vocal tract — the thing that makes one vowel sound different from another. Three of them are the entire instrument here.
There is no word list in this app. No phoneme bank. No dictionary. No alphabet. Nothing is ever chosen from a set of options, because there is no set of options.
Instead the phone builds a voice — an artificial throat and mouth — and hands the controls to the sensors.
Human speech is two things stacked. A sound source at the bottom: vocal folds buzzing, or breath hissing through a narrow gap. And a filter above it: the shape of your throat and mouth, which lets some frequencies through and blocks others.
Change the shape and you change the vowel. That is all a vowel is — a mouth shape. Two measurements describe nearly all of them, called the first and second formant, F1 and F2.
F1 tracks how open your jaw is. F2 tracks how far forward your tongue sits.
Every vowel in English is a point on that two-dimensional map. EE and OO share almost the same F1 and sit at opposite ends of F2. AH is wide open and central. That map is the grid on the main screen.
The sensors are not picking sounds off a list. They are holding the mouth. Where they move it, that is the sound it makes. Move through the map in the right way and you get syllables, because syllables are what movement through that map is.
Left alone, this would drone forever. A steady sensor makes a steady tone, and a steady tone is not speech.
So the output is gated by change, not by level. When the sensors sit still the gate closes and the room goes quiet, however strong the readings are. When they move, it opens. Only what changes gets a voice.
That is the same principle as clearing a channel: hold back everything constant so anything that is not constant can be heard.
Three modes, chosen before you open a session.
LISTENER — you watch and tap Mark. The detector is off, so nothing suggests anything to you.
MONITOR — the detector marks moments on its own and the Mark button is hidden. You are shown nothing until the session closes, so its findings cannot cue your ears while you listen.
BOTH — both run at once and neither can see the other. Use this one.
It has no dictionary, no word list and no phoneme templates. It never names anything, because the moment it named a word it would be handing you its own guess and you would hear that guess ever afterwards.
It looks for one thing: the gate opening in short bursts, at a rate a mouth opens at, with the formants moving far enough between bursts that the shape genuinely changed. A steady tone never qualifies however long it lasts. Neither does rapid flickering that keeps the same shape.
Any detector finds things. The number on its own tells you nothing at all, so the app never shows it on its own.
After the session the same detector is run again on your own recording, forty times, with each channel's timing rotated by a different random amount. Every channel keeps its exact distribution, its exact variability and its exact tendency to jump. The only thing destroyed is the coordination between channels — which is precisely the thing the whole instrument is claiming.
You are then told how many events your real session had, and how many the scrambled versions had. If the two numbers are the same, the events were a property of your sensors and not of anything organising them.
In BOTH mode you also find out how often you and the detector independently landed on the same moment, and how often two unrelated lists of moments would agree by chance.
This is the strongest thing the app produces. You cannot see the detector while you listen, and it has no vocabulary to be influenced by. If a person hearing words and a machine that has never heard of words keep pointing at the same instants, that agreement is not something either one of you could have arranged.
The whole session records automatically from the moment you open it. There is no record button and nothing to remember to press.
While it runs, a Mark button sits under the display. Tap it the instant you hear something — or hit the space bar. Do not type anything yet. Typing while you listen pulls you out of listening, and worse, it fixes a word in your head before you have heard it a second time.
When you close the session, every mark comes back with the audio, a button to play from two seconds before it, a box to write what you heard, and a rating.
It rates how much interpretation was needed, not how convinced you feel.
1 · SUGGESTION — heard only because you were expecting it, or only after someone named it aloud.
2 · SHAPE — word-like rhythm and vowel shape, no settled identity. You could name two or three words it might be.
3 · CANDIDATE — one word came unprompted and stayed the same on replay, but you had to be listening for language.
4 · CLEAR — unambiguous on replay, the same word every time, and you would expect another person to reach it without being told.
5 · UNDENIABLE — someone else, told nothing and not warned a word is there, named it independently.
A five cannot be given to yourself. The app will not accept one until the blind box is ticked, and it drops back to a four if you untick it. That is the only rung on this ladder that somebody else has to hand you, and it is the only one worth anything to a person who was not in the room.
Most marks in an honest session are ones and twos. A session of all fours is a reason to distrust the process rather than to celebrate.
Each entry also carries what the data was doing in the three seconds around it — how many times the gate opened, at what rate, which vowels the tract passed through, and the largest sensor excursion. That part is not a judgement and does not depend on your ears.
This is the least testable thing you could build, and that is worth saying plainly rather than burying.
The other tools produce countable things — letters, words, scores — so a run can be compared against an empty room and the comparison settles something. This produces continuous sound. There is nothing to count and no score to compute. Whatever you hear in it, you supplied the hearing.
What it does give you honestly: no vocabulary of anyone's choosing, a complete record of every parameter, and audio you can export and hand to someone else. The trace file holds the exact F1, F2, voicing and pitch for the whole session, so any moment that struck you can be found again and checked against what the sensors were doing.
Record a still, empty room first. If it sounds like voices when nobody is there, that tells you what your ears do with this signal, and it is the most useful thing you will learn all night.