muster in practice: Active Directory lite, from the directory up
How a small Linux fleet gets one directory, Kerberos logins, trustworthy security logs, a rendered WireGuard mesh, typed host policy and SSH host certificates, and what a real rollout taught.
The first article on muster described what it is and why it was built the way it was: one directory of people, groups and machines, Kerberos for authentication, and central rules for who may log in and use sudo, assembled from stock parts (OpenLDAP, MIT Kerberos, SSSD and WireGuard) plus a small amount of Rust that renders the directory into local configuration and ships security events. It also set out what muster refuses to do, and the decisions and found failures behind the design.
This post is the other half. It walks through the running system one picture at a time, follows a login, an alert and a policy change from end to end, and closes with what a real rollout taught that the design could not.
🔺 The shape of it
Everything runs over a private WireGuard mesh with IPv6 addresses. Two hubs carry the identity services. Each hub is an ordinary host that runs one small Firecracker microVM, and inside that VM run slapd and the KDC. The two directories replicate in OpenLDAP mirror mode, so either hub can serve a login on its own; writes go through one designated write hub, which alone runs kadmind. Lighter hosts, the light hubs, carry read-only directory copies for the day both hubs are unreachable.
Every other host runs SSSD for logins and one agent, musterd. The
agent renders the directory into local state and sends security events
to a collector on each hub. An administrator works through one CLI from
a workstation and never edits a host by hand.
Figure 1. The fleet. Hosts use either hub for LDAP and Kerberos and send security events to both collectors; the light hubs hold read-only copies.
The mapping to Active Directory is close enough to borrow its
vocabulary. A hub is a domain controller. A directory copy is a
read-only DC. A computer account is a host/ principal and its keytab.
Group Policy becomes typed host policy that moves through rings. The
security event log becomes a hash-chained record stream with two
independent copies.
Figure 2. Active Directory terms on the left, the muster part that does the same job on the right.
🔺 One account, every login path
A person is one entry in the directory and one Kerberos principal. SSH
keys live on the person’s entry. Groups decide where they may log in:
membership of access-all, or of access-web-01 for a single host.
Sudo rules are directory entries too, scoped to people or groups and to
hosts.
On the host, three login methods meet in one place. A key login asks
SSSD for the person’s keys. A password login goes through pam_sss to
the KDC. A Kerberos login presents a ticket for the host’s own
principal, which sshd checks against the local keytab. All three end in
the same PAM account check, which applies the access filter: in the
right group, and neither suspended nor tombstoned. Revoking a person
therefore means one directory change, and it covers every method.
Figure 3. An SSH login on web-01. Each method ends in the same pam_sss account check, and SSSD’s cache carries logins through an outage. SSSD caches what it reads. That is what lets someone who logged in yesterday still log in while both hubs are down, and it is also why a change can take up to five minutes to reach a host. The trade is deliberate: a fleet that cannot reach its directory keeps working, and a revocation lands within one cache interval.
People are never deleted. A person who leaves is tombstoned: keys removed, groups left, principal disabled, entry kept. The entry keeps its uid number, so no later account can inherit files the old one owned on some host.
🔺 Identity services in a box
The directory and the KDC are the crown jewels, so they run as stock
daemons inside a microVM with a read-only root image, no shell and a
single state disk. A small PID 1, muster-init, starts slapd and the
KDC in order, and only after the guest clock has been set from the host
through ptp_kvm. Mirror-mode replication orders conflicting writes by
timestamp, so a hub whose clock runs slow can have its write silently
discarded by the other. The clock comes first.
The VM sits on a host-only bridge. An nftables table forwards exactly one thing into it: the hub’s mesh address, on the LDAP and Kerberos ports. The public interface forwards nothing. Nothing in muster listens on a public address.
Figure 4. Inside a hub host. Only the mesh address is forwarded into the microVM; the public interface reaches nothing.
Admin rights need two keys at once. An admin has a separate
alice/admin principal with its own directory entry, and must also be
a member of cn=admins. The directory ACL checks both. So does
kadmind, whose ACL file is rewritten from the directory every ten
seconds. Removing the group membership revokes admin rights within
seconds, even from someone holding a valid admin ticket.
Figure 5. Admin rights need the /admin principal and membership of cn=admins. Removing the membership revokes within seconds, even with a valid ticket.
🔺 Logs you can check
Each host’s agent reads the journal for sshd, sudo, su, logins and SSSD. On a hub it also reads the KDC and kadmind lines from the VM console and the directory’s own change log. Every event is parsed, stripped of anything secret, and appended to a per-host hash chain: each record carries the blake3 hash of the one before it. Records wait in a local spool and go to both collectors over Kerberos-authenticated, encrypted connections.
A collector stores each host’s chain as it arrived, checks the links,
and runs the alert rules. Because the chain is linked, muster log verify names the first record that was changed or removed. Because
there are two collectors, an attacker on one hub cannot quietly rewrite
history.
Figure 6. Events are parsed, chained and spooled on the host, then sent to both collectors, which check the chain and raise alerts.
Alert rules are only useful if they are quiet. The first live week
proved it. One collector held over three thousand unacknowledged
alerts. Most came from internet scanners trying user names like admin
against public sshd, and from the administrator’s own root logins with
the break-glass key. The rules now leave unknown-name scanning to
fail2ban and recognise the break-glass key by its fingerprint. A
suspended person, a tombstoned person or an unknown name at the KDC
still raises an alert every time. The noise went away and the real
signals stayed.
🔺 A mesh rendered from the directory
The WireGuard mesh itself is directory data. Each mesh host is a node with a public key, an endpoint and the prefixes it routes. The agent plans the host’s peers from the directory with pure code, then applies them.
The rule that matters most is to fail toward no change. If no hub answers, the plan is refused and nothing changes. If a key is invalid or two nodes claim overlapping prefixes, the plan is refused. A missing entry never removes a peer; removal needs an explicit tombstone.
Figure 7. The WireGuard renderer plans from the directory and changes nothing unless every check passes. Rollout went in stages. First diff mode on every mesh host: plan every minute, report differences, change nothing. Only after 48 hours of empty diffs on every host does apply mode start, one host at a time, leaves first and hubs last. During that window a planned test stopped both hub VMs for two minutes, every host reported “no hub reachable”, and the clock started again. Annoying, and exactly the behaviour you want.
🔺 Policy as types, not scripts
Host hardening comes from the ComplianceAsCode project, lifted into typed rules. A rule is a resource type plus parameters: a sysctl value, an sshd setting, a file mode, a package. There is no script primitive. The checks, the fix, the rollback and the facts a rule needs all derive from its type.
Every host collects facts and evaluates its rules in audit mode. The results go to the collectors like any other event, and a high-tier rule that starts failing raises an alert. Because the collectors keep each host’s facts, a what-if run answers “what would this change fail?” without touching a single host.
Figure 8. Policy is lifted from upstream content into a bundle that every host audits. Enforcement moves through rings, and every change can roll back. Enforcement is off by default and moves through rings: a canary host, then leaves, light hubs, and hubs. Each change takes a snapshot, applies, waits through a confirm window and checks health gates, and rolls back if anything fails. A guard refuses any rule that would touch root’s keys, sshd’s break-glass settings or local accounts, whatever the rule claims.
🔺 Certificates instead of trust on first use
Clients should not have to trust a host key the first time they see it.
muster runs an SSH host CA with an offline master: an ECDSA key in a
SoftHSM2 token on a USB stick, plugged in only to create the master and
to sign a trust list. Each hub’s VM holds an intermediate key and a
small signer. A host’s agent asks a signer for a certificate for its
own names, valid for at most 90 days and renewed at 30. The agent also
renders the verified trust list into @cert-authority lines in the
host’s known_hosts, and revocations into a KRL.
The master never signs a host. It signs only the list of intermediates hosts should trust. A tampered list, or one signed by another key, changes nothing on a host.
This is the only certificate authority muster runs, and it is a narrow one. It is not X.509, which the design refuses; it signs host names, never people; and it is a different key from the fleet’s break-glass CA for root certificates, which muster never touches.
Figure 9. The offline master signs only the trust list. Hub signers issue host certificates; the break-glass CA stays separate.
🔺 The floor under everything
Every system like this needs a way in when it is broken. Here that is
break-glass: root’s own key in authorized_keys, an SSH CA for root
certificates where a host trusts it, and local accounts. muster never
writes any of it. The client role hashes those files before and after
every run and fails on any difference. A static check refuses any role
task that would write there. Policy enforcement refuses rules that
touch it.
So when every hub is down, every directory login falls back to the cache, and if the cache cannot help, break-glass still works, because it never depended on muster in the first place.
Figure 10. Break-glass access sits below muster and does not depend on it. With every hub down it still works.
🔺 Building it from scratch
Setting muster up is a sequence of stages, each closed by an acceptance script that tests the real fleet, not a lab. Build and freeze a release with checksums, and prove it reproducible with a forced rebuild. Bootstrap the write hub, then hand the directory to the second hub. Take the first admin tickets and change the temporary passwords. Start the collectors and run the hub acceptance: either KDC alone, replication both ways, admin rights that need both keys, no public listener.
Then enroll hosts one at a time, canary first. The login acceptance covers key, password and Kerberos logins, sudo rules arriving and leaving, access refused and revoked, break-glass from inside and outside the mesh, and logins with both hubs stopped. Security logs, policy in audit, the WireGuard renderer, directory copies, enforcement and the SSH CA follow, each with its own gate.
Figure 11. The setup sequence. Each gate is an acceptance script that must pass before the next stage.
🔺 What the rollout taught
The design held. The surprises were in the seams between muster and the operating system, and each one now has a fix and a test.
- OpenSSH 10 penalises sources. A host that sees many failed logins from one address drops every connection from it for a while. When tailnet traffic reaches the mesh masqueraded to a single address, one client’s mistakes lock everyone out, root’s break-glass over the mesh included. That address now sits on an exemption list, and the tailnet policy limits which devices can reach the mesh at all.
- sshd keeps the first value it reads. A host whose own hardening drop-in sorted after muster’s would have had password login turned on. The client role now sets password login per host and checks the effective result rather than trusting the file it wrote.
- Hostnames are not mesh names. sshd accepts a Kerberos ticket only
for
host/<system hostname>unless told otherwise. On real hosts the hostname was a provider name and the keytab held the mesh name, so every Kerberos login failed. Turning off the strict acceptor check fixed it. The dev lab never saw it because its hostnames matched. - Never delete, even in tests. A test that deleted its throwaway users freed their uid numbers, the next run reused them, and the hosts’ SSSD still mapped them to the old names. Tombstoning test users made the problem disappear, as the design always said it would.
- Terminals talk back. On Ubuntu 26.04 systemd wraps a
susession in terminal context sequences. A test that read the last line of output read the closing sequence instead of the answer. It now looks for a marker. - The new sudo ignores the directory. Ubuntu 26.04 selects sudo-rs,
which reads only
/etc/sudoers. Directory sudo rules need classic sudo, so the client role selects it.
Most of these were found by the acceptance scripts on their first real run, which is the best argument for writing them before the rollout rather than after.
🔺 What it leaves out
muster is deliberately narrow. It issues no X.509 certificates, keeps no general logs, joins no Windows domain, and does not configure BGP. It does not defend against a compromised admin workstation holding an admin ticket, or an attacker who controls both hubs. Inside those limits it gives a small Linux fleet what a Windows shop takes for granted: one place to decide who gets in, a record of who did that nobody can quietly edit, and hosts that stay the way they were meant to be.