
Overview
Proteus is a GitOps tool for a single bare metal Linux machine. You commit a manifest to git and it will reconcile the machine to match, the way ArgoCD does within Kubernetes. Nothing but the name exists yet, and this is just some of my ideas, what I would like it to become and the pain that spawned the project.
Coming from Kubernetes and working with the reconciliation pattern that tools like ArgoCD give you, going back to the roots of bare metal Linux is bleak. I recently tried to set up a small media server on an old NUC, and the current options and their compromises disheartened me.
In my mind I wanted a fully declarative machine. Pick an operating system, pick your services, config, any binaries you want installed, network setup, quadlets, and boom, you’re up and running. From mooching around I didn’t see anything that ticked these boxes, and the ones that came close had compromises and complex setup.
I started looking at bootc, as it’s probably the closest thing to my internal idea, Dockerfiles for full operating systems, but every change needs an image built first. Talos has the same problem and only runs Kubernetes. Ansible and friends start to drift once the box has been online for a while. I didn’t want to patch lots of tools together to reach my desired state.
I eventually landed on NixOS, which is great. It’s an awesome distro, you can configure services, users, your disks, pretty much anything you need on Linux, and with comin it pulls from git too. My problem is the Nix language and the learning curve. Everything was close, but I did not feel at home.
Nothing gives you ArgoCD-style reconciliation for a whole host without prebuilt images or a new language to learn.
| Tool | Where it’s great | Falls short for me |
|---|---|---|
| NixOS (with comin for git pull) | Fully declarative, generations give exact rollback | The Nix language and learning curve |
| openSUSE MicroOS | btrfs snapshot per update via transactional-update, auto rollback | No git sync, no manifest format |
| Talos | API-driven, minimal, machine config split into install-time and runtime | Only runs Kubernetes |
| Kairos | Immutable OS with Kubernetes-flavoured config, several distro bases | Image-based, built around k3s |
| bootc | OCI images for the whole OS, A/B updates | Needs an image build for every change |
| Fedora CoreOS (Ignition) | Declarative first boot, A/B updates | Ignition runs once, so later changes need a new image or manual work |
| Ansible pull, Salt, Puppet | Converge a live system from a repo | Drift, no transactional rollback |
In Comes Proteus
Like any over-ambitious engineer, I immediately wanted to tackle this with my own solution, bringing common K8s patterns into the bare metal world. What follows are design ideas and general thoughts on what I want out of Proteus.
To start, I only ever really need a box to run containers, which quadlets handle perfectly, so they would be the core of all services. Custom services that aren’t containerised get an additional way to be configured.
I want to cover all bases of software, and some software won’t have full configuration in its container image, so my idea is to bridge something similar to ConfigMaps in Kubernetes. You specify what’s in a file and a place to mount it, and it’s mounted as a read-only copy. You can also set binary data as a ConfigMap for anything custom.
The basic idea is a configuration file with all possible options and enough customisability to cover 99% of needs. I’m assuming btrfs on the host, since snapshots do a lot of the rollback work later on.
Bootstrap
If you are like me, you hate faffing around configuring a new machine physically, getting it plugged in, hooked up to a monitor, moving your keyboard around, just to skip to the end and enable sshd. Ten minutes of cables is ten minutes I don’t spend on the homelab.
My idea is to package a pure bootstrap image. There’s no way around BIOS level and booting, so a bit of intervention is required. After that, the bootstrap image should bring the device up, allow ssh connections with a set of default credentials, and allow remote setup.
Since the whole machine reduces to one configuration file, the initial setup is pointing it at a git repo and somewhere to look for a file. Easy peasy.
Manifest Draft
kind: Machine
name: homelab
template:
base: { distro: fedora, release: 44 }
packages: [kernel-core, systemd, systemd-networkd, podman]
---
kind: Quadlet
name: jellyfin
wave: 2
container:
Image: docker.io/jellyfin/jellyfin:10.10
PublishPort: 8096:8096
Volume: /srv/media:/media:ro
I would like the manifest to be simple fields, setup OS, your packages, and then any quadlets you want to run. Everything is deny/disable by default to keep things lean. I haven’t worked out how the root gets assembled from that yet, and it comes up again under rollback.
With ArgoCD you commonly use the App of Apps pattern, where one app handles the creation and tracking of all others. Similarly, the Machine would be the app, and the Quadlets and services would be the apps it tracks. The relation being one machine holds many apps.
Many machines raise the question of a MachineSet, the equivalent of Argo’s ApplicationSets. Im not sure on this, but I will draft the relationships properly as design goes on.
Reconciliation
Once the device points at a valid git repo with the necessary file structure and valid syntax, it polls the repo for updates. Every time HEAD moves, it diffs the changes and updates itself to match, much like ArgoCD.
Steps would look something like:
- Fetch the new commit
- Render the Machine object manifest
- Validate syntax, values, each unit and container digest, and fail on error
- Diff the state
- Order the changes by
wave. ArgoCD calls these Sync Waves, and we would sync secrets and network changes first, then apps - Snapshot the root subvolume, or build a new root if the changes are breaking
- Apply the changes wave by wave, health checking along the way
- On success, report synced and everyone’s happy. On failure, roll back and mark the failure
Rollback and Recovery
Generic config is the easiest part to roll back. It can happen at file level: restart, done. With Proteus I wanted to be sound of mind, though, because nobody wants to go out and physically fix the device after a dodgy update. It’s crucial to me that Proteus checks changes before it trusts them.
Two ideas for recovery; for a normal change, Proteus snapshots the root subvolume, and if the health checks fail it restores the snapshot.
The second is for what I’d call a “breaking change”. Here I’d like something like branching in git. Proteus builds a new root from the manifest on the box itself, with no registry or anything in the middle. This is for things that would usually require a full re-install: distro changes, version changes, kernel updates. Everything in that realm.
I’m not sure how achievable that is. It would then run checks against the new root, including whether it boots. If the checks pass, it reboots into the new root and uses systemd-boot counting to determine if the boot worked. If it didn’t, we roll back a commit and return to normal.
What happens if the rollback breaks? No idea.
With Metal Comes Rust
I’m still new to Rust, and I think a project I can work towards for a long time will help me learn. It may be ambitious, die, take off, take months, take years, but I’d like to see it reach an MVP.
I have no doubt lots of my ideas will get crushed once I start and feel the limitations, and some of what I’ve imagined might not be possible. The end goal is a GitOps tool for bare metal, and if it gets that far I’ll be happy.
The MVP so far:
- Parse and render a configuration schema, with a CLI tool to read and write quadlets
- Apply changes from git
- Run very generic health checks on updates
- Ship a bootstrap image
- Run one container with no manual steps
I have more ideas around secret management and diskless setups, but they sit outside the MVP.
Lots of research is required, and I may figure out why this hasn’t been done before. Hopefully another post with a demonstration follows, if this one doesn’t get forgotten.
Thanks for reading!