← back to writing··4 min read

Four boxes and a Pi: the homelab as a DevOps classroom

A Raspberry Pi monitoring hub, an Ubuntu box behind a Cloudflare tunnel, a Windows tower serving LLMs off a consumer GPU, and a NAS taking nightly backups — stitched together with a mesh VPN and systemd. What operating a five-machine fleet end-to-end teaches that tutorials can't.

My “infrastructure” is four machines and a NAS: the laptop I work on, an Ubuntu mini-PC running my dev database stack behind a Cloudflare tunnel, a Windows tower serving LLMs off a consumer GPU, and a Raspberry Pi that watches all of them. A mesh VPN stitches the fleet together; the NAS takes nightly backups. None of it is a product. It's the platform under my projects — and the cheapest DevOps education I've found.

Tutorials teach you commands. A fleet you can't afford to lose teaches you operations. These are the lessons that stuck.

The monitoring host should be your most boring machine

The Pi runs Uptime Kuma for service checks and Beszel for host metrics — and nothing else. That's a deliberate policy, learned the obvious way. The Ubuntu box hosts experiments and the database; the Windows box runs GPU workloads that occasionally need a hard reboot. If monitoring lived on either of them, every interesting failure would take down the thing that was supposed to report the failure.

The process that died when I closed my laptop

The incident that reorganised how I run anything long-lived: I kicked off a five-hour embeddings backfill on the Ubuntu box over SSH — ssh uwh "node backfill.js &" — saw it start, closed the laptop, went to bed. In the morning: no process, no error, a half-written table.

Nothing crashed. The process was parented to my SSH session; when the lid closed, the connection dropped and SIGHUP cascaded down and killed the child. The ampersand backgrounds a process — it does not detach it.

bash
# wrong: parented to the SSH session, dies when it drops
ssh uwh "node backfill.js &"

# survives disconnects: new session, no stdio, out of the job table
ssh uwh "cd ~/jobs && setsid nohup node backfill.js \
  > backfill.log 2>&1 < /dev/null & disown"

# trust nothing — verify the parent is init
ssh uwh 'ps -o pid,ppid,comm -p $(pgrep -f backfill.js)'
# PPID must be 1. Anything else and it's still coupled to you.
detach or it didn't happen

There's a graduation rule hiding in there: nohup survives logout, not reboot. Anything that must survive both stops being a shell command and becomes a systemd user unit — systemctl --user enable --now backfill.service — with restart policy, journald logs, and a supervisor that isn't me.

If a process matters, it deserves a supervisor. If it doesn't have one, you've just volunteered.

Windows told me the key was installed

Adding the Windows box to the fleet produced my favourite failure of the bunch. ssh-copy-id reported success. Password prompt on the next login anyway. Ran it again — success again, password again.

Windows OpenSSH reads keys for administrator accounts from C:\ProgramData\ssh\administrators_authorized_keys, not from the user's ~/.ssh/authorized_keys. ssh-copy-id happily writes the file Windows will never read and exits zero. The tool verified its write; nothing verified the outcome. I keep a small script now that writes the key to the right path and then proves it with a non-interactive round-trip login.

Verify the effect, not the exit code.

Same LAN, slower path

Once every machine has a mesh-VPN address, the tailnet name becomes muscle memory — and at home, that habit quietly routes traffic through WireGuard encrypt/decrypt to reach a box two metres away on the same switch. For a health check, irrelevant. For streaming tokens from the GPU box, a real tax.

So the rule is: check which network you're on first, then pick the path. LAN IP at home, tailnet name everywhere else. The overlay is a fallback, not a default. Knowing why — where the crypto happens, when NAT traversal falls back to a relay — is the part that transfers to any networking problem at work.

What the fleet actually teaches

The NAS takes nightly restic backups of everything that matters, and occasionally I restore one on purpose. A backup you've never restored is a hypothesis, not a backup.

Add it up and the homelab compresses the whole ops feedback loop into one room: you're the developer, the platform team, and the pager. systemd, DNS, WireGuard, tunnels, exit codes versus effects — each one stops being an abstraction the first time it pages you. When the cloud equivalent breaks at work, it's the same failure wearing a uniform.

You don't need Kubernetes to learn operations. You need one small fleet you can't afford to lose, and a pager that points at you.