One of my side projects earns me a bit of Trinkgeld - pocket money, basically. Nothing huge, but real people use it, and recently one of them ran into the same bug that had been annoying me for a while.
Normally the fix would go like this: find a free evening, sit down at my laptop, and vibe code my way through it. Which realistically means a bug like that lives on prod for a few weeks, at least.
This time I fixed it from my phone, sitting in my doctor's waiting room. Twenty minutes is the usual wait here in Berlin, and it turned out to be enough. But the bug is actually the second half of this story.
We bet on the wrong thing
A couple of years ago, Abhi Saha and I got talking about where agentic AI was heading, and what role smartphones would play in it. We both figured the answer was hardware: phones would need to get powerful enough to run serious AI on the device itself.
We were wrong. Not about phones mattering, but about what they'd need.
So before I touched the bug, I wanted to fix the thing that actually kept bugs alive for weeks: needing to find time and sit down at a laptop. That's where Hermes came in. The goal was simple i.e. go from an idea to a preview deployment just by messaging an app like Telegram or WhatsApp.
Saha has written up his own setup, which is worth reading alongside this one.
You don't need heavy machinery
This is the part I wish someone had explained to me earlier.
When people hear "AI agent running at home", they picture a beast of a machine with a big graphics card. That's what you'd need to host a model - the actual AI brain.
But I'm not hosting a model. The models stay where they are, built and run by OpenAI, Anthropic and others. What runs at home is the agent: the part that takes my message, remembers context, works with my code, and calls those models whenever it needs to think.
Think of it like a restaurant. The model is the kitchen. The agent is the waiter - it takes your order, knows what you like, runs back and forth, and brings the result to your table. You don't need a kitchen at home to order food. You just need a good waiter.
The setup, briefly
My waiter is Hermes, an open-source agent running on a Raspberry Pi 5 - a computer about the size of a credit card. I talk to it through a Telegram bot. It works on a copy of my code inside a sandbox, so it can't touch anything else on the Pi.
Everything before the Promote step is a draft. Only my tap puts a change in front of visitors.
The bug
SilkAway has a light theme and a dark theme. In dark mode, moving between pages showed a brief white flash before the dark colours came back.
It didn't break anything. But I'd noticed it, and so had one of my real users, which is when a cosmetic bug stops being cosmetic.
What I sent
I opened Telegram and messaged my agent the way I'd message a colleague:
When I am on silkaway, the transition b/w page is very snappy, especially in dark mode. The screen goes white and then theme is applied. Very bad ux on each page navigation. Go through the code base, identify the issue and prepare fix and push
That was the whole instruction. No file names, no hints about where the problem lived.
What it found
The agent traced the flash to timing. Here's what happened on every page:
- The browser starts loading the page.
- The page arrives without knowing I prefer dark mode.
- The browser paints it, in the default light colours.
- The app's code starts running in the browser.
- It checks my saved preference.
- Only then does it switch to dark.
The gap between steps 3 and 6 was the white flash. The page was painted before anyone told it I wanted dark mode.
Developers call step 4 hydration: the moment a page's code wakes up in your browser and takes over. The theme was being decided after hydration. It needed to be decided before.
The fix
The fix moves that decision earlier. A few lines of script now run at the very top of the page, before anything is painted.
The script checks two things: whether I've picked a theme before, and if not, what my device is set to. It applies dark mode immediately, so the browser never gets the chance to show the light version first.
The agent also set a dark background colour as a fallback, so even the page's first frame is right. Then it checked its own work: the code compiled, a production build passed, and the script sat exactly where it needed to.
Three rounds in preview
Every change the agent pushes gets its own preview: a private copy of the site containing only that change. It isn't live, and no visitor ever sees it.
The flash was gone in the first preview. I then spent three quick rounds testing, about a minute each, before shipping.
The whole thing took around 17 minutes from my message to a merged pull request. Not instant, but it fit neatly into a Berlin waiting room.
Nothing ships until I say so
This is the part I'd insist on.
The agent can write code, create a branch and open a pull request. It cannot put anything on my live site. When a change is merged, Vercel builds it, then waits. Nothing reaches production until I tap Promote.
It showed up in the demo. After merging, the live site still flashed white, because I hadn't promoted the build yet. That wasn't a glitch. It was the safety net doing its job.
I set it up this way because GitHub doesn't enforce branch protection on private repositories on my plan. So the gate lives in Vercel instead.
Before you try this
A few honest caveats.
- It isn't instant. Seventeen minutes for a small fix. Quicker than me learning what hydration is, slower than a colleague who already knows.
- You still have to review. The pull request is the approval step. Merge without reading it, and the safety net is only as good as your attention.
- Previews don't get my secrets. Database and AI keys live in production only, so features that need them won't fully work in preview. Visual fixes like this one preview perfectly.
- The sandbox protects the machine, not the keys. The agent runs isolated on a Raspberry Pi, but anything I hand it, it can use. So it only gets narrow credentials that expire.
What's next
This is the first post in a short series:
- How I set up the agent: a Raspberry Pi, a sandbox, a Telegram bot, and the evening I locked myself out of my own machine
- What I learned about security: mostly that a green light on a dashboard doesn't mean the thing is working
One last detail. The 67-second video posted on LinkedIn was cut by the same agent that fixed the bug.
