A workflow is running, and I need to close my laptop. Maybe I’m at a meeting, the end of the day, somewhere without a connection. If the agent is running on the laptop, closing the lid ends the run. For a while, I planned around it: before I left, I made sure agents were wrapping up or at a stopping point.
Then I stopped. The agents now run on a server at home. Each client project gets its own virtual machine. I reach all of it over a private network from my MacBook. This post covers why the setup looks the way it does, what it’s built from, and what I learned building it. It is not a Proxmox tutorial. There are plenty of those.
The Workflow This Has to Support
When I’m in heavy agentic development on a project, I can have as many as five worktrees going at once, each working a separate Linear issue. Those sessions can run for four or five hours. That’s the ceiling for how much I’m running concurrently, and it’s the normal shape of a busy day, not an exception.
In a given week I might also have two or three distinct client projects I could switch between. Usually I’m focused on one, but I need other environments ready: clients on a support agreement, where I go in for a day, work a couple of issues, and go back to my main project. Each of those clients has its own infrastructure, its own credentials, its own toolchain.
Where the Laptop Breaks Down
Two problems, and both come from the workflow rather than the hardware.
The first is tethering. Long agent runs, several worktrees in flight, and a week split between home and the office. If the agent lives on the laptop, the run depends on the laptop staying open and online.
The second is that one machine serves every client. In one VM I can be logged into a cloud CLI with access to one client’s infrastructure, and in another VM logged in with access to a different client’s. On one laptop, those logins share a home directory. I wasn’t after heavy-duty sandboxing. I wanted a thin layer that makes it unlikely an agent working on client A wanders into client B, and a VM boundary gives me that for free.
Before touching the hardware, I wrote a product spec for the setup: numbered requirements, a decisions log, and acceptance criteria. The acceptance criteria say what done looks like: from the office, I can attach to a client VM and find the agents I left running still running. I can rebuild any VM from a template in under an hour. And a client’s secrets are never visible from another client’s VM.
What I Landed On
A PC at home runs Proxmox. It has a 12-core Ryzen, 128 GB of RAM, and three disks: one for virtual machine disks, one for backups, and one for the host operating system and images.
On top of that sits one thin Ubuntu template built from the official cloud image with six packages installed. Every client project gets a full clone of it: 20 virtual CPUs, a 96 GB memory ceiling with a 32 GB floor, and a 500 GB thin-provisioned disk that can grow online and never shrinks.
Every VM and the host join my Tailscale network with tags. My devices can reach everything. The VMs can reach nothing: not my laptop, not the host, not each other. Nothing is exposed to the public internet, by design, permanently.
A working day starts by opening Ghostty on the Mac, running herdr, and picking a machine from the sidebar. Agents, panes, and worktrees live in the VM. I can sleep the Mac, quit the terminal, or drive home, then reattach to exactly what I left. If a VM reboots, herdr restores the layout and resumes the Claude Code and Codex sessions it had recorded.
Each VM has its own GitHub key, its own agent logins, its own client credentials, and its own untracked local shell config. One dotfiles repository configures every machine, Mac included. A new VM is a scripted, resumable provision that pauses for the steps a human has to do. A finished project is a scripted teardown that revokes keys, removes the network node, deletes backups, and retires the VM ID for good.
Why It’s Shaped This Way
The decisions that map most directly to the two problems above, and why each beat the alternatives:
- One large VM per client, not containers. The VMs are the isolation, and agents use worktrees to stay out of each other’s way. One VM is active almost all the time, so it gets nearly the whole machine. Memory ballooning is what makes two running VMs survivable.
- The VM is the security boundary, so the inside is permissive. Passwordless sudo, because unattended agents stall on a sudo prompt. Security updates apply automatically but never reboot, so a patch can’t kill a running agent.
- Thin template plus a per-VM checklist, no configuration management. A template with everything pre-baked goes stale within a month. The checklist is short, and it’s the same for every VM.
- No GUI, no RDP, no mosh, nothing public. I work from fixed places. I don’t roam.
- The MacBook runs no client project code. Project code and agents live in the VMs.
- The human is the concurrency limit. I have a “RAM limit” of my own: I can only do so much parallel work at once. If two VMs are running, I am doing one task on each, not running full five-worktree agent workflows on both. One active VM is normal, two is fine for light use, and three running multi-agent workflows is unsupported.
Cloning and sizing a new VM is two commands. Everything after that (console password, network join, keys, bootstrap, logins) is the per-VM checklist:
qm clone 9000 100 --name client-a --full 1 --storage tank
qm set 100 --cores 20 --memory 98304 --balloon 32768
The Tools, and the Job Each One Does
Proxmox owns the hardware. My initial instinct was that Proxmox would be more complicated than installing Ubuntu and running containers inside it. Thinking about it more, it’s easier: once Proxmox is set up I can spin up as many VMs as I want, experiment, and tear them down, and that’s a quicker turnaround than reinstalling Linux on the physical machine over and over. I roll on and off projects many times a year, so that turnaround matters.
ZFS, one pool per disk, no redundancy. The cheapest, slowest disk holds the host OS because the OS is the cheapest thing to rebuild. The backup pool gets a whole disk so it can never fill the host pool, and it carries a hard quota so a backup that would overflow it fails instead of filling the disk.
A cloud image and cloud-init for the template. Thin on purpose. A golden image with all my tools pre-baked would be stale in a month. The template gets the guest agent, SSH, a firewall, Tailscale (installed but not joined), git, and curl. Everything else arrives per VM. Templates are versioned and rebuilt, never patched in place.
Tailscale for all access. Two tags, one for the host and one for the VMs, and a policy that denies everything not explicitly granted. There is exactly one grant, and a tests block that makes the admin console refuse any future edit that would break the isolation:
"grants": [
{ "src": ["me@github"], "dst": ["tag:pve", "tag:devvm"], "ip": ["*"] }
],
"tests": [
{ "src": "me@github", "accept": ["tag:pve:22", "tag:pve:8006", "tag:devvm:22", "tag:devvm:3000"] },
{ "src": "tag:pve", "deny": ["me@github:22", "tag:devvm:22"] },
{ "src": "tag:devvm", "deny": ["me@github:22", "tag:pve:22", "tag:pve:8006", "tag:devvm:22"] }
]
The Proxmox firewall for what Tailscale cannot see. Tailscale’s rules only govern Tailscale addresses. The VMs also sit on my home LAN for internet access, and so do the host and my laptop. So every VM’s network interface carries an outbound ruleset that allows DHCP and DNS to the router and drops everything else on the LAN. It lives in the template and clones inherit it, so a VM can never reach the host or my Mac by a local address either:
[OPTIONS]
enable: 1
policy_in: ACCEPT
policy_out: ACCEPT
[RULES]
OUT ACCEPT -dest 10.0.0.1 -p udp -dport 67:68 -log nolog # DHCP to the router
OUT ACCEPT -dest 10.0.0.1 -p udp -dport 53 -log nolog # DNS to the router
OUT ACCEPT -dest 10.0.0.1 -p tcp -dport 53 -log nolog # DNS to the router
OUT DROP -dest 10.0.0.0/24 -log nolog # nothing else on the home LAN
OpenSSH with keys only. One dedicated key on the Mac for logging into VMs, separate from any repository key. Tailscale SSH is deliberately off, because two login paths is two things to reason about. Each VM generates its own GitHub key and uploads it for both authentication and signing, so retiring a project revokes exactly one machine.
herdr for daily driving. herdr is a terminal multiplexer, and the multiplexing itself isn’t novel. tmux would do that part fine. What herdr adds is the agentic lens: a sidebar that rolls up agent status across workspaces, each VM as a saved machine in that sidebar, and hooks that record Claude Code and Codex session IDs so that after a VM reboot the first attach brings the conversations back, not just the shells. Sessions survive disconnects. That lens is why I use it instead of tmux.
Claude Code and Codex in every VM. Each logged in separately per machine. Claude Code leads hand work to Codex workers as real CLI processes running with a filesystem sandbox and no interactive approvals. Blocked operations fail instead of waiting for an interactive approval.
Git worktrees inside a VM. Concurrent agents on the same project each get a worktree. Five at once is the design load. Worktrees are checkout isolation, not a security boundary. The VM is the security boundary.
Dotfiles with GNU Stow and an idempotent bootstrap. The same repository configures macOS and Linux from the same files. Homebrew on both. The two agent CLIs come from their standalone installers because both are macOS-only Homebrew casks. Machine-specific settings stay in untracked local files: one for shell tokens, one for conveniences like project aliases, plus the SSH config and Codex’s local config.
1Password and its CLI for anything secret. The provisioning script accepts only op:// references. The console password for a new VM is set from a terminal outside any agent session, so it never lands in a transcript.
A knowledge base that is itself part of the system. A product spec with stable requirement IDs, runbooks that cite the requirements they satisfy, inventory pages with a last-verified date, an append-only log, and CI checks. The rebuild runbook has to be readable when the host is dead, so the authoritative copy lives in git, not on the host.
What I Learned Building It
Here are a few lessons I learned.
Write the contract first, and give every requirement an ID.
I wrote a product spec the way I’d write one for a client: goals, boundaries, numbered requirements, acceptance criteria, a decisions log, a change log. It was reviewed before anything was built. The payoff is that agents cite requirement IDs across sessions and days, and runbooks cite them too. When the live system and the spec disagree, that’s recorded as drift and then either the system is fixed or the spec is amended. Neither side outranks the other by default. The spec still encoded a wrong idea once: it said “three VMs” when I meant three at a time. I corrected it, and the requirement was rewritten.
Every new Linux machine found something in the dotfiles.
The same repository configures the Mac and every Linux machine, and each new machine turned something up. A notification hook tried to spawn a desktop session bus on a headless host. A tool integration kept re-adding a hardcoded /Users/rob path into tracked config through a symlink. The bootstrap script didn’t install Claude Code or Codex on Linux at all. A personal Homebrew tap failed on Linux and broke the first automated build. The rule for all of this: if bootstrap fails, fix the dotfiles repo, not the VM. No VM-only workarounds.
Check what your checks prove.
tailscale ping runs below the access rules, so a successful ping only proves the tunnel exists. A DNS failure is not proof of isolation. A tee after bootstrap was hiding the bootstrap’s exit code. A safety hook that echoed BLOCKED and exited 1 was advisory: force pushes and .env reads had been blocked in name only. The isolation probes now require the specific exit code that means “timed out,” and the hook now actually blocks.
The third VM was scripted, and the runbook stayed the source of truth.
I built the first VM by hand, the second from the runbook alone as a test of the runbook, and the third with a script that cites and executes the runbook step by step. The script pauses for the steps a human has to do: choose the console password, add the machine in herdr (it requires an interactive terminal), do the logins, confirm the attach.
Teardown is half the lifecycle.
When I pointed out that I roll on and off projects many times a year, the plan changed. A three-VM build became a recurring provision-and-decommission workflow. Decommissioning revokes the VM’s GitHub keys, removes its network node, removes the saved machine, deletes its backups (or keeps one final archive for 90 days if the project may return), and retires the VM ID permanently. The VM’s page in the knowledge base is marked retired, not deleted.
What It Feels Like Now
On my first office day with the new setup, I closed the laptop to go to a meeting, and later to head out, without planning ahead to make sure the agents were wrapping up or at a stopping point. The work is where it was, doing what it was doing.
The other benefit is organizational. Every project has its own Linux environment, and nothing from any other project is in it. On a laptop that served every client, I had accumulated tools I couldn’t remove because some older project might need them someday, and runtime versions I was switching between, with global packages that belonged to one version and not the other. In a VM I install the version the project needs and leave it. That benefit has nothing to do with agents, and I’d value it regardless.
Closing the Lid
The workflow keeps running whether the laptop is open, closed, or on its way home with me. That was the point, and it works.