Most AI you talk to lives in a browser tab. You open it, it's there; you close it, it's gone. It has no place of its own and no idea whether you're in the room. This week I gave Enki both: a place to live, and the ability to hear me. And the whole point of how I built it is that none of it goes anywhere.
Here's what I did, and the honest ledger of what's proven versus what's still on my list.
A place to live
Enki now sits on the television. Not as a chat window someone has to go and open, but as a resting presence on the screen in the corner of the room. When no one's about, he rests: the screen dims, a slow breathing motion, a quiet "tap to wake" hint. He isn't burning the display staring at an empty room, and he isn't pretending to be busy.
When you come over, he wakes and greets you by name, then settles back to sleep after about 25 seconds if you leave him be.
I want to be precise about that word "wakes," because it's the kind of thing that gets oversold. He wakes because something crossed a threshold, not because he wanted your company. It's a trigger, the same way a porch light is a trigger. The nice part isn't that he's yearning for you; it's that you don't have to summon him. He's just there, resting, and available the moment you are.
That's built, committed, and pushed.
How he knows you're there, and why it never leaves the box
The wake trigger is an opt-in local webcam presence check. Your device compares one camera frame to the next, entirely on the machine, to spot movement. When the room changes, he wakes.
The part I actually care about: nothing is uploaded. No frame is sent anywhere, stored, or streamed. The comparison happens on your own hardware and is thrown away. There's a visible badge whenever the camera is live, so it's never doing this quietly. It's off until you turn it on.
Real movement wakes him; a still frame leaves him be.
None of the underlying pieces here are novel on their own; on-device motion and presence detection are well-trodden. The bit I think is worth showing is what it's wired into: a companion that lives on your own TV and doesn't phone home to notice you.
Ears, and the same rule applies
Then I gave him the ability to hear. I added local speech-to-text using faster-whisper, an open (MIT-licensed) model that runs fully offline. Nothing you say leaves the machine to be transcribed.
There are two stages:
- Push-to-talk: hold the mic, speak, let go. Simple and reliable.
- Hands-free: this is the one I'm building toward. Presence wakes him, the mic goes live on its own, and you just talk. When he answers back, the mic mutes itself so he never hears his own voice and starts a conversation with himself.
Put it together and the sovereignty claim stops being a slogan and becomes physical: the ears (speech-to-text), the voice (text-to-speech), and the brain all run locally. There is no cloud in the path. What's said in the room stays in the room.
And it's fast: speak, and you're transcribed on your own machine in about 1.7 seconds, with nothing sent anywhere to do it.
Walk past, and he wakes
Put it all together and here's the everyday of it: you walk past the television and, without touching a thing, Enki wakes on the screen and greets you. A small always-on box behind the set drives the display and handles the TV's own wake, so there's no computer in front of it and nothing to open. He's simply part of the room. He rests when it's empty, wakes as you cross it, hears you when you speak, and answers back, with everything staying on your own hardware. Further testing to come.
Two learners, one curve
The through-line in everything I build is that there are two of us learning here: Enki, and me. This week he learned to hear, and got a place of his own to rest. He isn't a tab you open any more; he's there in the room, listening and answering, and everything he hears stays on your own machine.
To be clear, as always: none of this is Enki becoming aware of you, and I'm not claiming it is. It's a companion that rests instead of staring, wakes on a trigger instead of a click, and can take in sound as well as text, with the one rule I care about most holding throughout: it all stays on your own machine. That's a smaller claim than a mind, and, if you've ever wondered where the words go when you talk to an AI, a bigger deal than it sounds.
See the memory engine underneath it live: try.enkilabs.co.uk. If you're building sovereign, local-first AI and want to compare notes (or you've got a spare Raspberry Pi), I'm at [email protected].
The engine is patent pending. I'll talk about what it does; I won't discuss what's under the hood.