Archived copy of the Nexus post as it stood on 6 October 2026, before it was rewritten. The current version is at kanzie.com/#post-nexus.

Nexus: The Personal AI That Lives in My House

My life runs in several languages. I'm a Swede living in Bavaria, my career happened in English, and my son is half Chinese. Every week brings a German letter from an authority, a Swedish bank asking for something, an English work thread and a family calendar that doesn't care about any of it. None of it is hard on its own. Together it's a constant background hum of things I shouldn't forget.

So I built Nexus: a personal operations system that lives on a small server in my house, reads the context of my life and tries to get ahead of it. This is the long version: how it's put together, where it lets a language model think and where it very deliberately doesn't, what it does on its own and what it asks me, and why part of it is a tiny island full of animals.

Last updated 5 October 2026: the Village's camera now films like a director (parachute drops, bench stories, animals that wave at you), and the island got golden hours, readable nights, weather and new sound. The whole post is also a good deal shorter. The changelog has the earlier rounds.

Three layers

Nexus, the brain, is a durable ledger of everything that needs attention: tasks, commitments, deadlines, correspondence and situation reports. It's the single source of truth for what I've promised and what's still open. It's plain SQLite, and every change to a task is written together with its change-feed entry in one transaction, because an assistant you can't trust to remember is worse than none.

Hermes Agent by Nous Research is the engine. It hosts a small team of agents with their own skills and schedules, delegates research to sub-agents and talks to me through the messengers I already use. It has operational access to my servers and network, so it can carry out a work instruction end to end. For Sit.Rep, the part of Nexus that investigates and fixes problems in the house, I deliberately took most of that power away again (more below). The few standing facts worth keeping live in a short, fixed part of its persona; everything else goes through a recall index with strict rules about who may see what. Neither may overrule the ledger.

The surfaces are how I live with it: Telegram and Signal on the move, a web dashboard at my desk, push notifications for what can't wait, a native iOS app, and the Village. Everything runs at home, each service in its own container on one Unraid server, next to a small local language model. Which jobs use it and which go to the cloud is in who gets to read what.

SIGNALS Mail & correspondenceinbox, sent mail Calendartoday, tomorrow, travel Home & servermonitors, backups Memessages, check-ins ENGINE & BRAIN Hermes Agentagents · skills · schedulesresearch sub-agents Nexus ledgertasks · commitmentsSit.Rep cases · timelines SURFACES Telegram & Signalchat · briefs · nudges DashboardSit.Rep · loops · load iOS appnative companion Push (ntfy)only what needs me now The Villageagents at work, live SIT.REP — HOW A PROBLEM TRAVELS Signal One case Investigate Plan I approve Execute Verify I review condition gone for good → the case clears itself · comes back → reopens as a regression
Signals flow into the engine and the ledger; I meet the results wherever I am.
Context first: Nora and the ledger

Most "AI assistants" wait to be asked. Nexus starts from context. Nora, my email agent, is a scheduled Hermes persona. Every morning she reads my inbox, my recent sent mail and my calendar, and through the day she runs short checks for anything urgent. Whatever she finds, she reconciles against the ledger: a new commitment becomes a task, one I've clearly finished gets closed, and I get a short brief in Telegram, whether the letter was in German, Swedish, English or Chinese.

Those runs never send, reply to or archive anything. Mail only goes out when I ask: she drafts a reply, shows it to me, sends it only after I've said yes to that exact draft, and then checks my Sent folder to confirm it went. The fully proactive loop, where Nexus investigates, plans, waits for my approval, acts and checks its own work, started with the house's infrastructure (that's Sit.Rep) and now covers personal matters too, behind more walls.

The Village

Agent orchestration is invisible by nature, so I gave it a place to live. The Village is a tiny 3D diorama island on a wooden plinth, and it's a mirror of my actual infrastructure. Every service is a building, every scheduled Hermes job is a little felt animal, and every on-demand Sit.Rep agent is a small robot. The rule that keeps it honest: nothing animates unless the data says so. An animal only goes to work when its real job is running.

The Nexus Village at golden hour: a low sun in a peach and blue sky over cobbled paths, a market cart, houses with smoking chimneys and power lines on wooden poles, with Zelda the raccoon saying in a speech bubble that she is off to a weekly new-release scan.
Golden hour in the demo village. Every sign, bubble and resident here runs on simulated data.

Buildings map one-to-one to what runs on the server, grouped by service family, so the whole media stack is one cinema, not five sheds. The Grand Webby Hotel has one room per website, lit while it runs. Tap any building and it tells you which containers live there and how long they've been up.

ON THE SERVERIN THE VILLAGE SCHEDULED JOBRESIDENT Nexus dashboardTown hall Agent runtimeClock tower My websitesGrand Webby Hotel Side web projectLodge in the Alps Media stackCinema Photo libraryGallery Smart homeSmart cottage DNS filterGatehouse Reverse proxyBridge Internet tunnelTunnel Code host & CILibrary Local LLMsObservatory Text-to-speechBandstand Push & messengersPigeon loft Array & parityMine Email, remindersNorafox WatchdogSnugglescorgi Security scansFrankbadger Deep healthKaibeaver Parity baselineChesterfieldtortoise Media jobsZeldaraccoon DNS policyLindamole Sit.Rep gateLesshedgehog Memory reviewsHazelsquirrel My chatsHermesowl Sub-agentsthe chicks Sit.Rep agentsrobotson demand (no job)Pipotter, sulking Cables = the real network Sky = the real sun and moon
The mapping, simplified. Unknown containers get a catch-all workshop until they earn a building.
Meet the residents

Every animal has a personality card: a name, a job in plain words, a few quirks and a way of speaking.

  • Nora, the fox, runs the post office, does yoga before sunrise and hums while she sorts. When a mail job fans out into sub-agents, a line of chicks gets busy with her.
  • Snuggles, the corgi, is the watchdog: short legs, big opinions.
  • Frank, the badger, guards the gatehouse at night and runs the security scans. Gruff, dry, secretly warm.
  • Kai, the beaver, keeps the mill running and does the daily deep health check. Never without a fishing rod.
  • Chesterfield, the tortoise, watches the storage array's parity, and takes the long way everywhere on principle.
  • Zelda, the raccoon, runs the cinema and its media jobs, and steals everyone's snacks.
  • Linda, the mole, keeps the addresses at the bridge pointed the right way (DNS).
  • Less, the hedgehog, is the town crier. She runs the Sit.Rep gate and rings her bell when something needs doing.
  • Hazel, the squirrel, lives in the observatory and keeps watch over what everyone remembers.
  • Hermes, the owl, is the one I actually talk to. While a chat is live, he's up on the pigeon loft, flapping away.
  • And Pip, the otter, whose job went to Hermes. He is not taking it well, keeps his sou'wester on in case someone calls, and now and then sits down somewhere for a good cry.
A three-by-three grid of Village residents, each standing on a yoga mat with a caption: Nora the fox in a postmaster cap (post office, mail), Snuggles the corgi (watchdog), Frank the badger in a police-style cap (gatehouse, security), Kai the beaver in a yellow hard hat (mill, deep health), Zelda the raccoon in a green visor (cinema, media jobs), Linda the mole in round glasses (bridge, DNS policy), Less the hedgehog in a crier's cap with a bell (town crier, Sit.Rep gate), Hazel the squirrel in a purple beanie (observatory, memory) and Pip the otter in a yellow sou'wester (no job, sulking).
Who's who: the original code-built figures on the yoga lawn. Read on for what replaced them.

When a job starts, its animal announces it and walks to work. Tap one and you get a line in its own voice and the facts: which job, for how long, how many tool calls and tokens, when it runs next. If an animal is asleep, tap the Zzz and it goes out to play for a bit.

Close-up of Nora the fox in a navy postmaster cap outside the post office, with a speech bubble saying she is working on an urgent mail check, with its duration, tool calls and tokens.
Tap a resident and it tells you what it's doing: the job, how long, how many tool calls and tokens. Never the content.
From code to felt

The first cast was built entirely in code: charming from a distance, a bit stiff up close. I wanted a set of handmade felt toys from one toy maker, so the residents were remade as 3D felt plush characters. It started with a single image of Nora as a felt fox postwoman. From that, Meshy generated a concept for every other resident with Nora as the style anchor: the same stitched matte felt, button eyes and light, and a pose that would rig cleanly.

A three-by-four sheet of felt plush concept images on an off-white background: Nora the fox postwoman with a satchel full of letters, Snuggles the corgi with a whistle and bandana, Frank the badger in a navy guard uniform with a lantern, Kai the beaver in a hard hat with a tool belt, Chesterfield the tortoise with a miner's helmet and pickaxe, Zelda the raccoon with a green visor and a film-reel bag, Linda the mole with round glasses and a scroll, Less the hedgehog in a tricorn with a bell, Hazel the squirrel in a purple beanie with a book, Pip the otter in a yellow sou'wester, Hermes the owl with flying goggles and a satchel, and a yellow chick in an eggshell hat.
The approved concept sheet. Nora (top left) is the style reference; every other resident was generated to match her.

Each concept went through Meshy's image-to-3D. Its automatic rigging gave up on the chibi proportions (it gave the corgi crossed legs), so every model is finished in a headless Blender script instead and rigged to the Village's own skeleton, so all the existing walks and poses just work. Ears, tails and satchels hang on short spring-driven bone chains, so a tail trails the walk and swings wide in a turn. The mouths are stitched, like on the toys, so "talking" just squashes the muzzle a little. Expressions are shape keys (blink, happy, sad, talk). Each character stays under 1 MB, model and texture together, and the Simple quality tier keeps the original code-built animals and never downloads the models at all. All twelve residents are felt now, and the robots followed with stitched screen faces.

Five felt plush residents lined up on a cobbled path in the Village by day: Less the hedgehog with her tricorn and bell, Snuggles the corgi with her bandana and whistle, Nora the fox in her postmaster cap with a satchel of letters, Hermes the owl with goggles and a satchel, and Frank the badger in uniform holding a lantern. A speech bubble above Nora says she has a new task.
The first five felt residents in a development build of the demo village.
Robots for the work nobody scheduled

The agents Sit.Rep dispatches per case (an investigator, an executor and a verifier, more below) are little robots. A new one rises out of a charging dock, its antenna flickers on, and it waddles off to the building of the case's domain: Kai's mill for infra, Frank's gatehouse for security, Nora's post office for life admin. There it scans, tinkers or stamps its clipboard, depending on its role. A robot whose plan is waiting for me walks to the notice board, pins the plan up and waits under a question mark while Less rings her bell. Tap it and "Review plan" opens that exact plan in the dashboard. That's what the Village is for, day to day: the place I look when a specific agent needs me.

A small white robot with a big round eye stands next to the village notice board, which has paper plans pinned to it, with a question-mark thought bubble above its head.
An investigator robot has pinned its plan to the notice board and is waiting for a decision.

Some Hermes jobs have no animal, like the one that feeds new mail into the recall index. Rather than let them run invisibly, each gets a worker robot while it runs, in a hard hat and hi-vis vest, tapping at a terminal or carrying parcels in the part of the village its job belongs to. When a robot's work is done, it turns to face you, announces its self-destruct sequence, counts down from ten and goes up in a flash, a fireball and a puff of felt, bolts and springs. If nobody is watching, it just quietly goes.

Wires, gems and a very slow tortoise

The network is part of the diorama too. Old-town wires droop from pole to pole into a switchboard on the town hall, which plays the router, and pulses of light travel them, coloured by what they carry. A toolbar button turns the island into an exploded diagram: the turf stays put, the soil drops away, and you see the cables run underground from every building to the router hub.

The exploded network view: the grassy top layer of the island floats above the brown soil layer, and in the gap coloured cables run from every building to a glowing router hub in the middle.
The network view splits the island in two to show the cabling between the buildings, the router, my devices and the internet.

My favourite detail is Chesterfield's long walk. A parity check on a large array takes most of a day, so while one runs, a cart heaped with glowing data blocks creeps along a little railway at the real progress point, with Chesterfield walking beside it carrying a lantern. Ten milestone lamps switch on as the check passes each 10%, a sign reads the time left and the error count, and Less announces every step. At night he dozes next to it and the cart keeps going.

The array mine: a glowing mine entrance in a rock, a curved rail track with lamp posts, a cart full of glowing coloured blocks, and a sign reading Parity check 75 percent, about 8 minutes left, 0 errors.
The mine during a (simulated) parity check: lamps for each 10%, a cart of data blocks at the real progress point.
A real sky, and a village that lives

The sun and the moon sit where they really are over our lake in Bavaria, at the real time, and the moon shows its real phase. When an animal has no work, it has a life: a deterministic schedule, weighted by personality and time of day, sends them to yoga in the morning, lunch at the market, fishing off the jetty and boule at dusk. Some head up the Alps to go snowboarding.

Hazel the squirrel in a red helmet and goggles riding a blue snowboard down the snowy piste straight towards the viewer, arms out for balance, with orange piste markers and pine trees around her.
Hazel, off duty, carving down the piste.

When two or three meet, they chat, and those lines are written by a local language model: Google's small Gemma 4 under Ollama on my own hardware, so they cost no cloud tokens. The model only ever sees the characters' personality cards, what they're doing and the time of day; there's no code path from Hermes, Sit.Rep or the server into that prompt. The lines are cached and rate-limited, and anything that doesn't validate (wrong length, an unknown speaker, something that looks like a link or code) is thrown away for a template line in the same voice.

Pip the otter in a yellow sou'wester sitting beside a house, crying big blue tears.
Pip, unemployed and inconsolable.
The eventful life of the residents

For a while the chat was pleasant and forgettable. Since 2 October, every resident has a soul: a short, hand-edited file with their voice, a few quirks, a private goal and a secret. Nora signs off mid-errand with "must dash!" and keeps every unsigned letter she's ever received in a hatbox under her bed. Frank says "seen worse" about almost everything.

They also remember: up to a hundred and fifty memories each, the least important fading first. They have feelings about each other (fondness, trust, a one-line opinion, the odd running joke) that drift a little with every conversation, and they have moods. Every night around three, each one rewrites their outlook on life, a short paragraph you can read on their About me tab.

Nora the fox in her postmaster cap outside the post office, with her About me tab open in a felt speech bubble: a mood chip saying content, a line saying her mind is on those unsigned letters on the sorting table, a dated Outlook on life paragraph about the great sack of village life, and a Friends list starting with Less and an affinity bar.
Nora's About me tab in the demo: her mood, what's on her mind, last night's outlook on life and her friends.

The model only writes the words. Nexus decides in plain code who talks and about what: a story beat that's due, village news, a rumour, gossip, small talk. Gemma gets the souls, the speakers' memories and feelings and that topic, and returns a few lines plus how the talk nudged their moods. Nexus checks it before anything is saved, and every nudge is capped, so one bad afternoon can't make two old friends into enemies. They can talk about their work, but only the way you'd talk about yours at a dinner party: Nexus turns its signals into fixed story phrases first ("Nora had an enormous mail sack"), and no names, machines or numbers ever reach the prompt.

And there's a plot. I wanted a family-friendly soap opera, and I got one: eight story arcs over weeks, with unsigned love letters on Nora's sorting table, a boule rivalry, a ring that might mean a proposal, a missing lemon cake and Pip's comeback. Some twists wait for real work, like a big mail day for Nora or a rough night at Frank's gate, and how an arc ends depends on how the residents actually feel about each other by then. In the recorded season, Nora chose Kai's morning fish over the poetry, and Hermes took it graciously. The chicks, meanwhile, decided early on that a fast orange animal is probably their mother:

"Fast orange. Peep. Mum?" — the chicks

When the chicks found out that Kai had made Chesterfield a ring box in secret, they told Frank, who turned out to be the best secret-keeper on the island:

"Hm. Ring box. Small things." — Frank
"Peep! Kai made it! Secret for someone." — the chicks
"Seen worse. Keeps to quiet." — Frank
The Village relationships panel over the island: twelve residents in a ring, Nora, Snuggles, Frank, Kai, Chesterfield, Zelda, Linda, Less, Hazel, Pip, Hermes and the chicks, with Chesterfield selected and green lines of different thicknesses to everyone else. A line underneath says Chesterfield likes Frank most and the chicks least.
The friendship web. Tap someone and you see how they feel about everyone else; here, Chesterfield likes Frank most.

On my own dashboard I can also whisper a rumour, and it travels from conversation to conversation, changing a little with every retelling. The season's first was a whale seen in the pond at dawn:

"Someone whispered to me a whale was seen at dawn, dear." — Hazel
"Stamp, sniff, sort, that tale sounds like a large sack." — Nora

The demo here can't reach my server, so it replays a recorded season: sixteen days and 240 conversations, generated offline by the real engine with made-up signals and rumours. It moves on one season day every fifteen minutes, so if you leave it open, moods and friendships change in front of you. About 90% of them passed the checks first time, which is why the checks and the fallback lines matter: a dull line is fine, a broken one on screen isn't.

The Workshop: watching Nexus get built

Nexus is built by Claude agents on a virtual machine in my house, and since 3 October that machine has a building of its own: the Workshop, out on the west rim. Twice a minute it reports how busy it is and what's running, with a key that can post exactly one kind of validated report. You can read the load from across the island: the chimney smoke gets darker, the flywheel blurs, and from 70% the roof rattles and the shutters bang. At the start of October that machine went down twice under too many agents, and now I see it coming.

The work is done by Claude's crew: felt gnomes from the same toy maker, one per agent, in knitted hats coloured by role (rust for code, teal for verifying, mustard for building, plum for testing, green for deploying). The coder hammers at the bench, the tester ticks off the easel and hops for a green run, and the verifier peers through a magnifier, then gives a thumbs up or a slow shake of the head. When the machine is struggling, they all clutch their hats. They talk like developers too (commits, flaky tests, rollbacks), and tap one and it tells you which release it's working on and what it thinks of it. That's the most honest picture yet of how I work with agents.

The Workshop in late-afternoon light: a timber-framed building with a teal slate roof and a brick chimney, a tool wall under a striped awning and a sign saying Open. Four felt gnomes in orange dungarees and knitted hats work in the yard, in orange, plum, mustard and teal. A speech bubble above the plum-hatted gnome reads: Tester, Claude's crew, ticking off the test list, v1.96, Village: the Workshop and Claude's crew, Verifying, and the quote Green, green, green!
Tap a gnome and it tells you what it's building. In the demo, the crew works on real past releases.

New work arrives by air: a sealed blueprint tube floats down on a little parachute, the foreman catches it and a RECEIVED stamp thumps onto it. When a release goes live, Less calls it out on the square and a pearl and gold mist rolls across the island; my open dashboard quietly reloads onto the new version under it. When one is merged, a gnome wheels the release cart over to the town hall. The yard got so busy that the smart cottage in front of it moved to the empty meadow on the east side. Release titles come from Nexus's own notes and are filtered at both ends, so paths, addresses and keys never leave the machine.

A new coat of paint

My reference was the Link's Awakening remake, where every house looks like a painted toy you could pick up. So every building got real materials: plaster, coursed stone, planks and logs, tiles and shingles, with a soft sheen and worn edges. The other half of the work made it lighter: compressed models and textures, simpler animals in the distance, and on a device that can't keep up the Village turns down one thing at a time, the least visible first. The residents also stopped walking through each other: everyone looks a couple of seconds ahead, slows early and steps round whoever is in the way, while anyone sitting or working counts as a fixed obstacle. And on my iPhone, the Village now installs as its own Home Screen app and runs edge to edge.

A camera that directs

The Village has a "stalker mode" (the film-camera button) that roams the island at animal height and goes wherever something is happening. It used to arrive late and wander. It now weighs every event by how important it is, how soon it can be framed and how much of it will be left, re-decides four times a second, and turns, dashes or makes a smooth crane hop to get there. Events are on screen about 2.5 seconds after they start, down from almost five, and a well-composed shot is held for five to eight seconds, with nothing in front of the lens.

Then it learned the big moments, the way a director would. Each session opens with an establishing shot over the square, backlit at golden hour. On a release, the camera stays on Less, the town crier, through "Hear ye!", rises for the mist, then cuts to a reaction shot of her ringing and the residents cheering. When an order parachutes into the Workshop, it is already up on its crane, follows the tube down from high above and lands on the foreman catching it.

A high-angle shot over the Workshop at golden hour: a red and white parachute carrying a blueprint tube floats down towards the timber workshop and its yard on the island's grassy rim, with the low sun in the sky.
The parachute shot: high over the Workshop as a new order floats down to the crew.

On the square's benches the residents tell stories. Friends sit half turned towards each other, and when the camera rests on them, two or three talk an arc forward, a step per visit: a recap, then tension, a reveal. The camera films it with a two-shot and a single on each speaker, line by line. The local model writes the lines while the camera is still on its way; if it's slow, hand-written lines for the same story step play at once.

Night on the town square: Frank the badger with his lantern and Kai the beaver in a hard hat sit side by side on a bench, Hazel the squirrel in a purple beanie on the next bench, lit windows behind them. Frank's speech bubble says: Clocks don't just invent an extra hour. Somebody is up to something.
A bench conversation two-shot on the square.

And now and then, an animal notices the camera. Its head follows the lens, it waves, or jumps and waves, and carries on. That's at most once every 40 seconds across the island, never while someone works, and with reduced motion it's just a nod.

Kai the beaver in a white hard hat and denim dungarees stands on a path between houses in warm evening light, facing the camera with a raised paw, under a speech bubble saying his daily deep health check failed on its last run and that he found a crack in the beam.
Kai notices the camera and waves.
Every hour of the day

The light got its own pass. Golden hour is warm and flattering, dawn and dusk glow rose and peach, the blue hour is lavender, and nights are readable: lamps throw soft pools on the cobbles, lit windows spill onto the street, and a soft light from the lens keeps faces out of the dark. At night the release mist glows moonlit blue, lilac and gold.

The town square at night: Less the hedgehog in her tricorn stands on the cobbles with her bell, Linda the mole beside her, street lamps throwing warm pools of light, lit windows, the notice board, the clock tower and the neon sign of the Grand Webby Hotel behind. Speech bubbles say a cottage thermostat issue is fixed and verified, and that a new DNS check has come in.
Night on the square: lamp pools, window glow and residents you can still read.
The whole island at night seen from above, a soft lilac and gold mist rolling across it with sparkles, lit windows and lamps glowing through it, a countdown bubble and two speech bubbles from residents.
A release going live at night, under the moonlit mist.

The island has ambient life now: a flock of birds crossing every minute or so, butterflies over the grass, moths at the lit doorways, falling leaves, fish rising in the pond and smoke from the chimneys of whoever is home. The weather is rare and gentle: mist over the pond and meadows on about half the mornings, thinning as the sun climbs, and at most one short drizzle an hour, with rings on the pond.

Early morning in the Village: a pale lilac sky, low mist drifting over the pond and the meadow beyond, wooden power poles, a grassy mound and the clock tower, with residents waking up on the square.
Morning mist over the pond. It comes on about half the mornings and thins as the sun climbs.

The sound is still off by default, and still synthesised in the browser with no audio files: birdsong, crickets, the mill and the stream, the crier's bell. Every big moment now has a sound of its own (a whoosh as the parachute opens, a shimmer for the mist, a small cheer at the reveal, a pip-pip when someone waves), and every effect comes in a few slightly different versions, played in a shuffled cycle, so repeats never sound identical. A guard mutes anything that starts to loop.

Finally, the residents say what's happening. When a job fails, a container goes down or a plan needs my OK, the animal or robot it belongs to says so in a speech bubble: a plain fact first, then a line in its own voice. The facts come from templates over a short allowlist of fields, through the same privacy filter as the Workshop's titles. Mail contents and senders, secrets, addresses and paths are never spoken, and a personal case is only ever "a personal case". In stalker mode, a small card shows who you're watching, what they're doing and what they're thinking.

Nora the fox in her postmaster cap working at a table outside the post office at night. Above her, the stalker info card reads: Nora, post office, email, working on Urgent Check for 20 seconds, 8 tool calls, 5k tokens, and her thought, Stamp, sniff, sort. Nearly through the pile. To the left, Snuggles' speech bubble says the higher-ups need a watchdog sweep.
Nora's stalker info card, with Snuggles talking over it. All of it from the demo's made-up world.
Private by construction

A live view of my house's machinery is only fun if it can't leak anything. So the Village works from a whitelist: job names, states, durations and counts. Session titles, prompts, message contents and mail subjects are dropped on the server before anything reaches the browser, and Sit.Rep robots carry a case number and a role, never a title. Nexus never gets the Docker socket, only a read-only proxy, and nothing polls unless someone is looking.

I used Nexus to get properly familiar with Opus 5.5, and it generated all of the graphics: the island, every building, the original animals and the robots, down to Pip's tears, plus the Blender pipeline for the felt cast. It's the same way I used Fable on Me Or Them, my network analyser.

The version on this site is the public demo: the same code, fed entirely by a simulated world with made-up containers, jobs and cases. Left alone, a showcase director makes sure every fun moment comes round regularly: boule, yoga, a soap-opera beat, Pip, the Workshop's busy spells, a parachute order and a release going out under the mist. There's a live demo on this page. Go and say hello to the whole cast, and be nice to Pip.

Deterministic by default, a model where judgement is the job

Behind the cute animals sits the one design rule I keep coming back to. A language model is non-deterministic: ask it the same thing twice and you can get two answers. That's what you want when reading a messy situation, and exactly what you don't want when deciding whether a disk is full or a command may run. So Nexus uses reasoning only where there's real judgement to be made, and plain, deterministic code everywhere else.

Code does all the watching, counting and permitting. The monitors are scripts with no model attached; they report facts, never prose. Code decides whether a sighting is new or noise, clears cases that disappear and reopens ones that come back, decides whether a model wakes up at all, starts the per-case agents and enforces the token cap. Every permission is code: per-role keys into Nexus, a gatekeeper on the server with a fixed allowlist, the risk rating of a plan, the hash that binds my approval to one exact plan, and the step that writes an approved event into my calendar.

Language models come in where the work is interpretive: reading a job's output and working out which leads are real, investigating a case and writing the plan, carrying out an approved plan strictly inside its box, judging whether a check really passed, talking a case through with me. A plan that fails the check goes back, however well it's argued. The hand-off always looks the same: code decides when a model wakes and what it may touch, the model does the thinking, and code checks the result before it counts.

DETERMINISTIC CODEsame input, same answer, every time LANGUAGE MODELSonly where judgement is the job Monitorsscript-only jobs, no model Fingerprintssame problem → same case Auto-cleargone 3 runs · back = reopen Wake gatenothing to do → 0 tokens Dispatcherwho runs, when · backoff Role keys, gatekeeperallowlist · no shell Dossier check, riskrules, not opinions Approval hashexact plan + my answers Exec authorisebyte for byte in the plan Reading job outputdigests → next step MailNora's briefs → the ledger Investigationevidence · vendor docs Dossiers & plansquestions · rollback Executioninside the approved plan Verificationis it really fixed? Case chatanswers · proposals only Mail readerlocal model · quarantined Village chatterlocal model · 0 cloud tokens Code decides when a model wakes and what it may touch · the model thinks · code checks the result
Where the model lives, and where it doesn't. Violet is code that gives the same answer every time, teal is where judgement is the job.
Sit.Rep: the system that watches itself

A home full of services drifts: backups age, certificates expire, a container restarts at 3 a.m., a disk fills up. Sit.Rep (situation report) keeps track of all of it, and the watching is done by plain scripts, each looking at one corner of the house. Each finding arrives with a fingerprint the monitor chooses (detector : resource : condition) and a flat set of facts. The same fingerprint always lands on the same case, so a problem seen a hundred times is one card with a timeline, not a hundred alerts, and only a material change (an error count jumping from 2 to 12) gets my attention back. Facts are compared with a little tolerance: numbers fall into coarse buckets, lists are compared as sets, and volatile values like "checked at" are never compared. Every monitor also reports the full list of fingerprints it sees, so a case missing from three runs in a row clears itself; if it returns, it reopens as a regression. Most problems that fix themselves, I never see.

Detector signalfingerprint + facts, no prose reject → back to a new plan triage investigating decision approved executing verifying review closed I approve plan vN I close cleared fyi acknowledged dismissed failed gone 3 runs nothing to do I retry the same plan drift, error or averifier fail: whyis on the card needs me agents at work resting · back if the facts change or it recurs
The case lifecycle. Orange is my turn, teal is the agents' turn, grey is resting.
Lanes, states and "Needs you"

Every case sits in one of four lanes (Infra, Home, Security, Life admin) and moves through an explicit state machine, from triage through investigating, decision, executing and verifying to review. Every transition is one transaction with a version check, so an agent and I can never act on a stale view of the same case. Acknowledging isn't a mute button: the facts at that moment become the baseline, and a material change brings the case back. The dashboard shows one number: Needs you, the cases waiting for a decision, in review or failed. If it's zero, I can stop looking.

The gate: wake the model only when there's real work

The first Sit.Rep coordinator was an AI agent on a fifteen-minute timer that spent most of its time rediscovering that nothing had happened: 30 to 44 million tokens a day. Now a small script asks Nexus first whether there's anything to do, and if not, the model never wakes up. That took a normal day to about 1.5 million tokens, a 96% cut with nothing lost, and because a silent tick now costs nothing, the gate runs every two minutes, so an approved fix starts almost as soon as I tap approve. The dashboard shows an ETA for it, with a "Start now" button and an Abort that works until the executor picks it up.

Agents with real boundaries

The uncomfortable truth about version one: a single coordinator investigated, executed and reported on itself, with broad access to the server, and the safety rules were text in a prompt. Since 29 September those rules are walls. A script-only dispatcher starts narrow agent profiles in Hermes:

  • The investigator gathers evidence, reads vendor guidance and writes a dossier: what's affected, evidence with sources, a step-by-step plan, how to verify it and how to roll it back.
  • The executor runs an approved plan and nothing else, checks first that the world still looks as it did when I approved it, and journals every step.
  • The verifier starts fresh with only the plan, its criteria and the executor's claims, re-runs every check itself, and runs on a different model family so they don't share blind spots.
  • The chat is what I talk to about a case. It has no tools; it can only propose, as a button only I can press.

None of them has memory, a shell or file access. Each role has its own key into Nexus, held by the tool server so the model never sees it, and its own SSH key that can only start a small gatekeeper script: a short allowlist of read-only commands, no pipes or shell tricks, and redacted output. The executor can go one step further only by asking Nexus whether that exact command, byte for byte, is part of the plan I approved. Refusals turn into a Security case, so if an agent tries to step out of its box, I hear about it. The pipeline also runs under a hard daily token cap, with a separate one for personal cases.

HERMESNEXUSTHE SERVER Dispatcherscript · every 2 min · 0 tokens Investigatorevidence → dossier + plan Executoronly the approved plan Verifierfresh context · re-checks Chatno tools · proposals only Monitorsscripts · facts, not prose Role keyseach role, its own endpoints Dossier checknames · evidence · rollbackquestions, not guesses Plan-bound approvalhash of plan + my answersstale facts → refused Exec authorizeexact command in the plan? Sit.Rep casestimeline · review · Needs you Gatekeeperone key per role, forcedcommand · no shell Read-only allowliststatus · bounded logs Redactionsecrets out of every output Audit log + watcherrefusals → Security case Containers & arraywhat's actually running asks
Where the walls are: role keys into Nexus, one gatekeeper on the server, and approval bound to the exact plan.
Plans I can actually approve

An early case taught me the most important lesson: the agent proposed a tidy fix whose first step quietly depended on a decision only I could make. So plans are now checked by code before I see them. A plan that doesn't name what it touches, has a claim without evidence, lacks a rollback at medium risk or hides a decision in a step goes straight back. Each step carries an action class, and the risk is computed from those classes, not taken from the model's own opinion. When I approve, I approve that exact version: a hash of the plan, my answers and the facts as they were. If any of them changes before the executor starts, the approval is void.

What it does on its own, and what it asks me

The line is simple. Nexus does everything that doesn't change a system on its own: watching, investigating, researching, writing plans, clearing cases that fixed themselves. Anything that does change a system waits for me, and finished work lands in review with what changed and the evidence. In practice my part is a two-minute read and one tap.

The first live plans taught me the opposite lesson too: the agents asked far too much. Since I approve every execution anyway, a wrong guess costs me one edit while a question costs a round trip, so the agents now decide and state: a plan lists its assumptions with how sure it is, and every question comes with a suggested answer. I tap "Looks right" or change a value, and a change re-plans the case. When I tick Remember, my answer is kept as a fact for future cases, and every case counts the questions it asked, so I can see whether this is working. Picking the right investigator is deterministic too, and when I re-route a case, Nexus keeps it as a visible, removable example for similar cases.

Personal cases: the same loop, more walls

My life admin is where the loop pays off most, and where a mistake or a leak would hurt most. So personal cases get their own personal investigator with no access to the server at all. It can read my calendar, work out where I'll be on a given day and estimate a drive time locally, so no address goes to a routing service. The walls: it waits for my explicit consent before looking into anything; to every other agent a personal case simply doesn't exist; a push about one says only that something needs me; and it has its own token budget.

Calendar actions: the model proposes, code acts

The first personal action Nexus can take is putting things in my calendar. A plan's calendar steps (title, times, place, calendar) are part of the hash my approval is bound to. Then no language model is involved: Nexus writes the events itself, reads each one back to check it, and only then hands me the case for review. It can only change events it created, and it holds its own calendar-only permission, never the broader one Nora uses for mail.

Recall: what Nexus already knows

Agents kept rediscovering things, so the recall index is one search over closed cases, excerpts of my mail and standing facts: plain SQLite full-text search, no embeddings. Access is scoped per role on the server: the infra investigator only ever finds infra cases, and mail and personal facts are visible to the personal investigator and me, nothing else. Mail excerpts are redacted, kept twelve months and never include an address, and the access log records who searched, never what for.

The quarantined mail reader

Mail is the easiest way to attack an agent: anyone can put text in front of it by sending me an email. So the personal investigator never reads mail itself. It asks a question, and a quarantined mail reader answers it, with Gmail-read tools and nothing else: no web, no shell, no memory, a hard turn cap. Its Gmail access and its Nexus key live in two separate processes, so neither holds both. Its answer must fit a strict schema, and Nexus strips anything that reads like an instruction before storing it. If a mail tried to instruct the reader, the answer is marked "injection flagged" and I get a push naming only the case. Details in a plan that came from mail are labelled "from mail", so I know what to double-check. Since 1 October the reader runs on a local model, so my mail never leaves my network for it.

Who gets to read what

Nexus knows a lot about me: my mail, my calendars, where I'll be, the standing facts about my family, the letters from the tax office and the bank. An assistant like that is only worth having if I know exactly which model reads which part. So for every feature the question isn't only "can a model do this?" but "which model, where, and what does it see?"

Most of Nexus never shows my data to a model at all. The ledger, the recall index and the case timelines are plain SQLite in my house; the monitors, the gate, the dispatcher and the calendar writer are plain code; the Village works from a whitelist. The parts that do put personal data in front of a model are few, and I can name them: Nora's mail briefs, my chats with Hermes, the personal investigator, the quarantined mail reader and the summary a task gets when I open it.

At home runs one small model, Google's Gemma 4 (e4b) under Ollama: the mail reader, the task summaries and the Village's chatter and souls. In the cloud run the bigger models: the investigators, the executor and the case chat on an OpenAI model, and the verifier on Kimi, from Moonshot, a different family on purpose. Public web lookups go through a separate agent that holds no personal data.

The rule I'm working towards: the most sensitive raw material stays at home by default, and a cloud model gets the least it needs. The personal investigator in the cloud never opens my inbox; it asks, and what travels back from the local reader is short, checked, redacted and labelled "from mail". It isn't all local yet, and I'd rather say so: Nora's brief and my chats with Hermes still run on a cloud model and read real mail, and the personal investigator sees my calendar and at most eight standing facts, picked by plain code and listed on the case.

What makes this a choice rather than a rewrite is the layering: every role is its own Hermes profile with its own model, key and tools, and Nexus talks to roles, not models. In its first day and a half, the mail reader ran on three different models, and its walls didn't move. Part of the routing is now a setting, sorted by sensitivity. Your take, a short judgement Nexus writes on a case, runs locally for personal, money, health, family and legal items, and if the local model fails it doesn't quietly fall back to the cloud: I get a Retry button. Infra, home and security items may use the cloud model, because its judgement is better and the data is about machines, not me, and one switch, "Keep every take local", overrides both. A page called Where your data goes lists every model Nexus uses and what each one receives, so "who reads this?" is a page I can open, not something I have to remember.

Local has a price. Gemma e4b runs on a few processor cores, no graphics card, so it's slow, and it can't yet write a plan that passes Nexus's own checks. For the hard thinking, the cloud models are still clearly better, so that's where it stays, behind the walls above.

Which Nexus roles run on the local model at home and which on cloud models Two columns. At home, on a small local Gemma 4 model running on the server's processor: the quarantined mail reader, task summaries (local first, cloud if the local model is down), the Village chatter and souls, the mail watcher in testing, and Your take for personal items, which never falls back to the cloud. In the cloud, on larger models behind Nexus's walls: the investigator and executor, the case chat with no tools, the personal investigator, which sees answers rather than my inbox, Nora's briefs and my chats with Hermes on an OpenAI model, the verifier on Kimi as a second opinion from another model family, and web lookups that hold no personal data. Underneath: every role is its own profile with its own model, key and tools, so moving a role is configuration, not a rewrite; and in Settings, personal items go to the local model with no cloud fallback, infra may use the cloud, and one switch keeps everything local. AT HOME · SMALL LOCAL MODELGemma 4 e4b · on the server's processor IN THE CLOUD · LARGER MODELSone Hermes profile per role · behind the walls Mail readerquarantined · mail stays home Task summarieslocal first · cloud if down Village chatter & soulsno personal data Mail watcherin testing · unsure mail only Your take: personalnever falls back to the cloud Investigator · executorOpenAI · infra work Case chatOpenAI · no tools Personal investigatoranswers, not my inbox Nora · Hermes chatsOpenAI · reads mail VerifierKimi · a second opinion Web lookupsno personal data at all Every role is its own profile: model · key · tools. Moving one is configuration, not a rewrite. Settings: personal → local, no cloud fallback · infra may use the cloud · one switch keeps it all local
Who reads what, as of October 2026. Teal runs on the small model at home; grey runs on cloud models, behind Nexus's walls.
A local model, measured first

Before trusting a small local model with my mail, I had the candidates measured: Gemma 4 (e4b) against Qwen3 8B on my own machine, on the jobs Nexus actually hands its models, scored by Nexus's own validators, 404 trials in all.

  • Reliable: verification, tool calls, mail extraction in English, German and Swedish, and short summaries. Neither model ever followed an injected instruction (0 of 24), but Gemma flagged every attempt, Qwen a quarter.
  • Advisory only: routing. Gemma was right 98% of the time, but the deterministic router is right by construction.
  • Not ready: writing plans. None of Gemma's twelve passed the plan check first time.

Gemma was also two to two and a half times faster, so it won. One gotcha worth passing on: Ollama's OpenAI-compatible endpoint silently truncates every prompt to a 4,096-token context, with no error. Create a model variant with a larger num_ctx and call that instead.

What the layers buy besides privacy

Cost. The cheapest token is the one that's never sent. The gate cut a normal day by 96%, investigations now take about 70,000 tokens instead of 150,000, and the local model costs no cloud tokens at all.

Resilience. When a provider runs out of credits or rejects its key, the cases that depend on it are paused, not failed, with a note saying why, and they resume on their own when a short check every 30 minutes sees it's back. A banner says which provider is out and my phone gets a push when it happens and when it recovers. Because Nexus knows which role runs on which provider, only the work on that provider stops, and the mail reader at home can't run out of credits at all.

Freedom to swap, and a second opinion. A new model can be tried on one role, measured first, without touching the others. And two models trained differently, like the executor and its verifier, are less likely to share a blind spot or fall for the same trick in an email.

One model, for stability. When the Village souls arrived, they asked Ollama for a slightly different model and context size, which it treats as a different model. It kept swapping 6.6 GB in and out until a virtual machine on the server reset. Now every local caller sends exactly the same model and context, and a test pins it.

How I live with it

Nexus doesn't ask me to open another app. It lives in Telegram: a short morning brief, a nudge when something really can't wait, and a conversation I can pick up anywhere. For what needs me now, there's ntfy, a self-hosted push service: a case that needs a decision, a failed run, finished work waiting for review. Routine notices never push, there's at most one push per case per state every six hours, an overflow folds into one "N more items need you", and a push never carries facts or output, just a title and a link. For a personal case, not even the title.

At my desk, the dashboard is the calm overview: open loops in the order I want to tackle them, a deadline horizon, the calendar, and the Sit.Rep panel with a dossier per case. The case view has one action area at the top, saves as I type and updates live. It also keeps an honest eye on my load and wellbeing with a one-minute check-in; it's reflective, not diagnostic, but surprisingly good at telling me to slow down. And in my pocket there's a native iOS app for the daily loop.

Meanwhile, in the rest of Nexus

Most of the recent work went into the Village, but a few things changed in Nexus itself.

Cheaper investigations. From about 150,000 tokens to 70,000, by taking fewer steps: no looking up their own tools, several read-only checks at once, and shorter replies from Nexus.

One tap to close or acknowledge. Done, Won't do or Not relevant on any open case, and an acknowledged security or drift case reopens on any change.

A Settings page. Most knobs that used to live in code are settings now, in three tabs: Nexus (the Sit.Rep budget, notifications, mail and recall), the Village, and the public demo. Each resident has a card with an editor for their soul, which is checked before it's saved.

Better tools for working with Claude. The agents take their screenshots on a real graphics card on my home server (every new picture in this post came from there), and leave reports for me in a private files area.

Security. On 2 October a set of separate agents audited the whole system: every agent role now needs its own narrow key on every internal route, the chat agent has no tools at all, and anything from my mail is treated as untrusted quoted text wherever it appears. None of it is visible, which is the point.

Learnings

Agents and jobs get out of hand quickly, and quietly. Almost every real problem I've had was dispatching gone wrong, not a model saying something stupid. A single test run of an investigator once burned 148,000 tokens and still came back wrong, which is why every profile has a hard turn cap. A test run of mine left the gate's state file owned by the wrong user, and the coordinator was blind for seven and a half hours. My first hand-over to the new dispatcher lasted five minutes: a schema mismatch sent every case back to triage and the next tick sent a fresh investigator into the same wall. I rolled it back, and now the dispatcher backs off a case after two failed investigations and flags it to me as "Investigation stuck" instead of retrying forever.

Seeing it beats reading about it. None of those announced themselves in a log I'd have read in time. That's the serious case for the Village: a robot climbing out of the dock for the same case every two minutes is something you notice from across the room. It works the other way round too. Wiring up the Village's feed is how I found two monitors that had never run on their schedule at all, because the scheduler needs an Apply click I had never made. And some script jobs finished in under a second and never showed up on the island, so now each one gets a few minutes of visible work. A job you can't see isn't one you trust.

Lean on the harness. Before Nexus I had more or less hand-rolled my own agent loop and struggled to make it perform. This time I let Hermes do what it's built for: a profile per role with its own model, skills, turn cap and nothing it doesn't need; jobs that run a script first and only wake a model when told; a tool server that holds the keys. The flip side is that the harness configuration is part of the product: part of that failed hand-over was a single flag that was never passed.

Independent verification earns its keep. Nexus is built by AI agents. An Opus 5.5 session orchestrates: it plans, hands work to implementing agents (Opus for anything visual, Sonnet for the plumbing) and deploys only when I ask. Every feature then goes to a separate verifier that didn't write it, and those verifiers have caught real security holes, like a calendar client that would have followed a redirect to any host with my credential. An agent marking its own homework is the same mistake as the coordinator that reported on itself.

If you've read this far: the island is waiting. Open the Village demo and see who's busy.

Changelog
  • September 2026: first published.
  • 29–30 September: Sit.Rep's agents with boundaries, decide-and-state plans, re-routing, the personal investigator, the recall index, calendars and the quarantined mail reader.
  • 1–2 October: paused-not-failed outages, the felt cast, worker robots, local Gemma 4, and Village Souls with a recorded season in the demo.
  • 3 October (v1.16.6, v1.16.7): the Workshop and Claude's gnome crew, the release mist, a new coat of paint, lighter downloads, residents that step around each other.
  • 4 October (v1.17.0–v1.17.2): the rest of Nexus caught up, and a new section, who gets to read what.
  • 5 October (v1.18.0, v1.18.1): the Village demo is now Nexus 1.146.3: a camera that directs (parachute shot, bench stories, waving animals) and stalker mode that starts at once, golden hours and readable nights, ambient life and weather, new sound, and residents that say what's happening, and softer nights. The post is about a third shorter.