A small cloud VPS, operated through Git and systemd

Published: September 13, 2026Read time: ~8 minutes
Google CloudOpenTofuPodmansystemdRestic

My homelab has a cloud side as well. A Google Cloud VM runs Headscale, VictoriaMetrics, Grafana, and the backends for some of my personal projects. It gives those services a home outside the local Kubernetes cluster, while keeping the deployment small enough to understand as one Linux machine.

I use the same idea of keeping desired configuration in Git, with a different mechanism here. The gitops-vps repository holds a Compose file and service configuration. systemd prepares the host and runs the deployment workflow, and rootless Podman runs the containers under a dedicated application account.

A machine I can replace, data I can keep

The infrastructure code lives in my gcp-vps-infra-greenfield OpenTofu project. It describes a custom VPC, subnet, firewall rules, a static external address, a service account, and a Debian VM with a separate state disk. The bootstrap script prepares the application user, storage mount, runtime tools, and systemd units.

The disk boundary matters more to me than the VM size. On the running host, the boot disk is 20 GB and a separate 30 GB disk is mounted at /srv/state. Headscale's database, Grafana's data, VictoriaMetrics storage, Caddy's state, and the QR logger's data all live under that mount. The Git checkout at /opt/gitops-vps can be replaced during a deployment without treating those directories as disposable files.

That separation does not make the state disk a backup. It gives the application a persistent place to write while the operating system and checkout have their own lifecycles. The backup routine takes the next step and copies that state off the VPS.

I had to account for what happens when the mount is missing. A directory named /srv/state can still exist on the boot disk, and letting containers write there would put data in the wrong place. My homelab-state-ready.service combines RequiresMountsFor=/srv/state with an explicit mountpoint -q check. The application unit requires that service, so having the expected path is not enough: it must actually be mounted before the stack starts.

Rootless container permissions are another detail in the bootstrap. I fix the application account at UID 2000 and its subordinate UID range at 200000. Grafana writes as UID 472 inside its container, so the script calculates the corresponding host owner for its state directory:

GRAFANA_CONTAINER_UID=472
GRAFANA_HOST_UID=$((SUBID_START + GRAFANA_CONTAINER_UID - 1))

With that range, the directory owner is host UID 200471. Assigning every state directory to vps-app would miss this case. Keeping the mapping stable also gives files copied with numeric ownership the same meaning on a replacement VM. The storage layout, container user, and backup options have to agree for this to work.

Following a revision into running containers

At boot, homelab-app.service waits for the state mount, repository, runtime secrets, and the application user's runtime directory. It then runs podman-compose up with recreation and orphan removal enabled. The Compose file describes six services: Caddy, Headscale, the Spotify backend, the QR logger, VictoriaMetrics, and Grafana. All six were running when I checked the host on September 13, 2026.

For an update, homelab-deploy.service runs the deployment script and then restarts the application unit. The script takes a lock, validates the checkout, fetches the configured branch, records the previous and target commits, and moves the disposable checkout to the selected revision. It renders runtime secrets and performs a Compose dry run before systemd restarts the stack.

The validation happens before the checkout is reset. It checks the expected repository path and origin, rejects a symbolic-link checkout, and refuses local working-tree changes. The nonblocking flock prevents two deployment runs from moving the checkout at once. These checks matter because the later reset and clean deliberately replace local files in that directory.

The service connects preparation to activation with these two lines:

ExecStart=/usr/local/libexec/homelab/deploy
ExecStartPost=+/usr/bin/systemctl restart homelab-app.service

The deployment script runs as vps-app; the + prefix lets the follow-up command restart the system-level application unit with elevated privileges. If preparation exits unsuccessfully, systemd does not proceed to that restart. There is still no automatic rollback if the newly started application fails, which is why the recorded commit IDs and a separate check of the containers matter.

The live deployment script follows production. The greenfield template currently names greenfield-test, so that template describes a rebuild path with a branch setting that differs from the running host. Keeping that distinction visible is more useful than assuming the local template is an exact copy of production.

This GitOps workflow runs through a service: the deployment unit fetches and applies Git state when invoked. There is no periodic Git sync timer on the current VPS. Argo CD continuously reconciles my Kubernetes platform; this machine uses an explicit deployment step. A successful Compose dry run also checks less than an application health check, so I still inspect service status after deployment.

Services at the edge, credentials at runtime

Caddy is the public entry point for Headscale and the personal backends. Headscale provides the coordination service for my Tailscale clients. The containers share a private Compose network, so Caddy can route requests to service names without publishing every container port on the host.

The runtime account retrieves the deployment key and application secrets through Google Secret Manager. The deployment key is temporary, and application credentials are rendered under /run/homelab and mounted into containers as read-only files. The repository holds references to those files, rather than the credential values.

The secret-rendering script downloads into temporary files first, rejects empty payloads, and checks that the Spotify and YouTube credentials are nonempty JSON objects. Only after validation does it move them to the filenames used by Compose. That avoids replacing a usable credential file with a partial download. The Git deploy key has an even shorter lifetime: it is removed after the fetch, before the script recreates the application containers.

Grafana is bound to 127.0.0.1:3000 on the VPS, which keeps its host port available through an SSH tunnel instead of a public bind. For metrics ingestion, Caddy requires a trusted client certificate and forwards only POST /api/v1/write to VictoriaMetrics. Other requests on that endpoint receive a 404.

Bringing the cluster's measurements here

The Argo CD project configures vmagent in the local cluster to send samples to this endpoint. VictoriaMetrics stores those samples, and Grafana queries the storage service over the Compose network. The VPS query API returned 39 up series during the September 13 check, confirming stored scrape-status data; that count is not a claim that all 39 targets were healthy.

The Compose configuration sets VictoriaMetrics retention to 48 hours and its allowed memory percentage to 30. That keeps this a modest operational view of the lab. It also makes the tradeoff explicit: I have a short metrics history, and this single VPS is still one failure domain for the services hosted here.

Backups start from my workstation

The backup scheduling lives on my workstation. Its homelab-backup.timer runs daily at 20:00 local time with up to 30 minutes of randomized delay and persistent catch-up. A separate maintenance timer runs weekly on Sunday morning. Both timers were present and scheduled when I checked them; they are not VPS timers.

The backup script first checks SSH access, the state mount, local space, and access to the Restic repository. It then stops homelab-app.service on the VPS, flushes the state filesystem, and uses rsync to update a local mirror while preserving ownership and filesystem metadata. After the copy, it starts the application stack again and creates a Restic snapshot from the mirror. An exit trap attempts to restart the remote stack if the copy fails.

The copy uses rsync -aHAXx --numeric-ids: hard links, ACLs, extended attributes, and numeric ownership are preserved, and the transfer stays within the state filesystem. Numeric IDs matter for the rootless Grafana ownership described earlier. The script also checks a marker and the resolved local mirror path before using --delete-delay, so the directory being updated must be the prepared backup destination.

Failure handling starts before the remote stop command. The script marks the stack as potentially needing a restart, then clears that flag only after a successful start and active-state check. Its cleanup trap can therefore attempt recovery when an intermediate SSH or rsync operation fails. It also verifies access to the Restic repository before interrupting the remote services; discovering a bad backup password after stopping the stack would waste that interruption.

Stopping the stack gives the file copy a quiet source, at the cost of a service interruption for the duration of that copy. Restarting before the Restic snapshot keeps the snapshot work outside that interruption. The mirror is updated in place; the Restic repository is what provides versioned backups beyond the current mirror.

This arrangement makes the workstation, its available storage, and its ability to reach the VPS part of the backup system. Persistent scheduling helps after missed runs, but a timer and a snapshot still do not prove a complete restore. The boundary I have built is concrete: Git tracks deployment configuration, the state disk holds live data, and the workstation pulls a separate backup copy.