Vision paper
We are building the voice input layer for people.
Proposition
For the past few years almost all of the money and the best people in this industry have gone into one thing: making models answer more accurately, more quickly, more like a person. That work has gone well.
But the difficulty an ordinary person has with AI was never at that end. The difficulty is this: when the thought arrives, you cannot type.
Walking somewhere, driving, out of the shower, awake at three in the morning, standing in a queue, running, on the way back from dropping the kids off. In those moments you have something to say and there is no screen in front of you — or there is one and both your hands are busy. By the time you sit down the sentence has gone, or what is left of it is a blur.
Input is the scarce half. And almost nobody is working on it.
What we build
And in time the video input layer too.
This is ambient AI — we also call it Aether AI: a layer that stays close to you, is always there, and asks nothing of your attention. It does not think on your behalf. It only makes sure that when you have something to say, there is somewhere to say it.
The two get filed under the same heading often enough that it is worth being exact.
Edge AI means moving the computation onto the local hardware. That is not us. Our hardware does exactly one thing locally — it picks up sound. It does not read, does not understand, does not infer anything. It gets your voice cleanly and that is all.
Turning it into text happens on a phone or a computer, using a large model in the cloud. The hardware is not clever; the hardware only has to hear well. That division is deliberate, and it is what lets the hardware be small, frugal with power and cheap — and it lets the model be swapped for a better one whenever a better one shows up.
One-way
You speak, it captures, it turns into text. The AI's reply is text — you read it when you are back at a screen.
We do not do live spoken conversation. That is not a shortfall in what we can build. It is a choice.
At the moment a thought arrives, what you need is not to lose it, not a chat. A conversation asks you to stop, to give it your attention, to wait for a reply and then carry on — while you are walking, or driving, or in the middle of something else. A thing that requires your concentration comes looking for you at precisely the moment you have none to give.
Capture and conversation are two different products. We only do capture.
That choice makes the whole layer lighter: no wake word, no always-listening, nothing talking in your ear, and no reason to put down whatever is in your hands.
Quiet-Voice Engine
This is the first principle behind every piece of hardware we make, and it is almost too simple to call a technology.
Sound pressure falls with distance. Halve the distance and roughly twice as much of it reaches the microphone — 6 dB. Background noise does not get quieter because you moved the microphone closer; it is spread evenly through the whole space. So every halving lifts not a little signal but the entire signal-to-noise curve.
An ordinary headset's microphone sits ten to fifteen centimetres from your mouth, which outdoors leaves it guessing. A phone held in your hand does get to two or three centimetres — but first you have to take it out, unlock it and find the right screen, and a thought does not survive those three steps.
So for any piece of AI hardware the central question is the same one: how do you get the microphone closer to the mouth without asking the person to do anything extra?
On AI sunglasses the answer is to put the main mic at the bottom centre of the left rim, port facing down at your mouth — worn, that is four to seven centimetres. If it is genuinely loud, or you simply do not want to be overheard, take the glasses off and hold them up: the same microphone comes to two or three centimetres. Nothing switches modes. You shortened the distance with your hand.
We call this the Quiet-Voice Engine. What it means is: at that distance you can speak very quietly — too quietly for someone half a metre away to hear — and it still hears you clearly.
Only you
Not intruding on anyone else.A piece of hardware worn on the body that records the whole room will always, to everyone else in it, be a problem. Our test is the level difference between the main mic and the reference mic: when you speak they are 7 to 9 dB apart; someone half a metre away is almost equidistant from both, under 1 dB, and never reaches the threshold. No voiceprint, no model — it is decided by where the holes are.
And protecting you.This half gets mentioned far less. Hold a small device up to your mouth in public and you can speak at something close to a whisper. What you said stays between you and the AI.
Which is why it cannot record a meeting.A few people brainstorming, an interview, a lecture — this layer cannot do it, and a voice recorder should. That is not a shortcoming. It is the other face of the same decision.
Philosophy
It does not think.It does not judge for you, does not offer advice, does not try to become another AI. There is no shortage of clever things on the market. What is missing is a quiet one.
It does one thing.It carries what you said out of your head intact. Get that wrong and nothing else stands up; get it right and everything else can be somebody else's job.
Restraint is not doing fewer things. It is refusing to become something you have to learn.
Universal connector
The text can go into any coding agent, any framework, any box you could have typed into. You pick the destination; our job is getting the words there.
There is a judgement about position underneath that: the large companies will never unify each other's agents. Not one of them has any reason to hand a user smoothly to a competitor. It is work only somebody unaligned can do.
So we deliberately do not build a model, and we do not intend to. We want the wire, not either end of it.
Memory
What gets captured is stored as Markdown. An AI can read it and so can a person, and the whole thing exports whenever you want. Move to any other AI application and you take it with you.
We do not lock you in. A product that holds people by the cost of leaving ends up spending its effort on the height of the wall rather than on what is inside it.
But the record itself turns into something else: an archive of how you think. What occurred to you and when, which questions keep coming back, which one you have been circling for three years without resolving. What that is worth after a year is not remotely what it was worth on the first day.
Form
The input layer is not tied to any one form. The shell can change; the first principle does not. The microphone has to be able to reach the mouth, and it records only the person wearing it.
This is the form we care most about at the moment, and not because it is cool. It is because when the sun comes out you were going to put sunglasses on anyway — the one wearable nobody has to be talked into. A ring, a necklace, a pendant, a companion object: each has to win a hard argument first, which is persuading someone to change what they carry on their body. Sunglasses never have to have that argument.
The main mic sits at the bottom centre of the left rim, port facing down at your mouth; the reference mic sits at the top corner of the right rim, port facing up and away. Playback, volume and calls all belong to the phone, exactly as with any other Bluetooth headset. No camera, no wake word, and not one light anywhere on it that can come on — something that physically cannot record the people around you does not need a light apologising for it.
Then add a prescription. Glasses with a correction in them are worn something like ten times as many hours as sunglasses. And the property that matters most in an input layer is not how clever it is. It is whether it is there when the thought arrives.
There will be a version with a camera later. At that point the layer stops being only voice.
The pendant carries its own microphone. When you want it, you lift the pendant to your mouth — which is as close as this layer can get.
The rest of the time it is a pendant: worn at the chest, or clipped to a bag. Take it off, say a sentence, put it back — three seconds in total. Its real advantage is that it changes nothing about what you wear.
The same principle in a shell that invites more affection. It sits on the desk, in a bag, in a child's hands; when you want to say something you pick it up and bring it close.
Something you were already willing to pick up satisfies the closest-to-the-mouth condition without being asked to.
The simplest one of all. They can be an ordinary pair of Bluetooth earphones, or they can add a piece of local storage and a transcription application. When you want to keep something, take one out and hold it to your mouth; the rest of the time they are in your ears; when they are not in use they go into a small cloth pouch worn as a pendant.
Properly in-ear beats every other form in a noisy environment, and when it is loud outside it is still the best thing to reach for. Not while exercising, though — running or riding, a good pair of glasses suits better.
Forms will keep changing, and they should. The first principle does not.
Why now
AI wearables.They have reached their own first-iPhone moment. Every device sold is one more entry point.
Speech to text.It crossed the usable line. Accuracy on casual speech, on accents, in noise has finally reached the point where you do not go back and fix it afterwards. Transcription went from roughly readable to directly usable, and only then does capture mean anything.
Agentic AI.It became reliable. So the things you captured finally have somewhere worth being delivered.
The moat
Position.What we want to occupy is the wire between a person and every AI, rather than any one of the AIs.
Stateful, not stateless.Anyone can copy a pipe. Nobody can copy a year of one person's memory — replacing it means starting from nothing.
Scarce data.A structured record of how people actually direct AI in real life. That happens away from the screen, and only an entry point somebody has grown used to wearing can reach it.
Neutrality.No model, no side taken, so anyone can connect.
More use → deeper memory → higher cost of leaving and scarcer data → more people using it.
Last
The way people speak to machines — plainly, the way an ordinary person reaches AI at all — is not going to grow on a single object.
It is on the phone and on the computer. It is also in AI glasses, in AI earphones, in a voice ring, in a voice necklace, in a small voice pendant, and in a creature that keeps you company.
The entry point will not belong to the cleverest AI.
It belongs to whatever is nearest when you want to speak.