I told you so.

muster, from the shell to the browser

A week after the first rollout: DNS from the directory, single sign-on with passkeys, a page for first passwords, device enrollment, and the seams that only real clients find.

The first three pieces described muster as it went live and how it is built: one directory, Kerberos logins, logs nobody can quietly edit, a WireGuard mesh rendered from the directory, typed host policy and SSH host certificates. That covered the shell. This one covers the week after, when the same identity had to reach a browser, a phone and a laptop on someone else’s network, and when most of the surprises came from the edges rather than the core.

Six phases went live in two days, a week after the first rollout. The order mattered more than the speed: each one had to stand on the one before it, and each had a gate that tested the real fleet, not a lab.

🔺 Machines that do not give up

The hubs run their directory and KDC inside small microVMs. Until this week, a VM whose boot failed would power itself off. That is a reasonable default in a lab and a poor one in production: a hub that stops trying is a hub somebody has to notice, log in to and restart.

The rule now is simpler. A system never powers off on failure. It keeps trying until it succeeds. The VM’s init restarts each failed step with a backoff, forever, and says so in its log. A slapd that dies is restarted without a reboot.

Writing that rule down found a real bug the same day. The VM’s clock sync failed on every boot in the test. chrony needs four readings of the host’s clock per poll, and the poll setting allowed at most four, so a single slow reading meant no sync, and no sync meant no Kerberos. One setting changed, eight boots out of eight failing became none out of eight. The fix was found because the VM no longer hid the failure by switching itself off.

🔺 DNS from the directory

The mesh has its own DNS zone. It used to be generated from the fleet inventory file and copied to the two name servers. The directory already knew every host and its address, so the zone moved there too: musterd on each name server renders it from the directory copy that already runs on that host.

It went live in three steps, and the middle one did the work.

  • Diff. The renderer computes the zone it would write, compares it with the zone being served, and reports the difference without touching anything. It ran this way for about ten hours.
  • Fix the inputs, not the output. The first diffs were refused. Two hosts appeared both in the directory and in a hand-kept list of hosts without a directory entry. Then the renderer could not read the live zone at all, because the old generator wrote relative names (ca CNAME hub-a) and the parser expected fully qualified ones. Then it wanted to delete seven names: six aliases that existed only in the inventory file, and one workstation with no directory entry. Each was fixed where it came from: the aliases became directory attributes on their hosts, and the workstation, which has no directory entry, went back on the hand-kept list.
  • Apply. When the diff showed only what was meant to change, nine new SRV records for the KDCs and LDAP, the renderer was allowed to write. Both name servers now serve the same zone, record for record.

Two writers for one file is the classic way to lose a week, so the old generator was switched off for those hosts in the same change.

🔺 One sign-in for the web

Web services in the fleet now sign people in through Keycloak, using the same directory and the same Kerberos passwords. Keycloak does not keep a password. It asks the KDC, as kinit does, so a password changed anywhere works everywhere at once.

Figure 1. Web sign-in. Keycloak asks the KDC to check the password, reads people and groups from the directory, and requires a passkey; the web service only introspects the token.

Figure 1. Web sign-in. Keycloak asks the KDC to check the password, reads people and groups from the directory, and requires a passkey; the web service only introspects the token. The first attempt to install it stopped before it changed anything. The hub already ran a Keycloak for an unrelated service: same unit name, same port, same install directory, same system user. Running the role would have replaced it. muster’s instance is now separate in every name it uses, and the role writes a claim file on first run and refuses to touch any unit, port, user or path it did not create. Its database got its own PostgreSQL cluster, because the host’s default cluster was a read-only replica.

The second attempt failed on a hyphen. Debian’s PostgreSQL unit takes the cluster name from the unit instance, and systemd turns a dash in an instance name into a slash. A cluster named muster-sso could be created but never started. It is muster_sso now, and the role refuses a dash.

The third attempt found four ordinary bugs in the role, the kind a dry run cannot see: a folded YAML line that split a shell pipe, a root umask that made unpacked files unreadable to the service, a link task that handed the data directory to root, and a task result that overwrote a variable of the same name. Each now has an offline check that runs with the other role tests. Through all three attempts, the unrelated Keycloak kept its process and kept answering.

🔺 Passkeys, and where they stop

Every web sign-in now ends with a passkey: a key pair that lives in a phone, a laptop or a security key and never leaves it. The first sign-in asks for one. After that the password alone opens nothing on the web. A passkey is bound to the sign-in site’s name, so a look-alike page cannot use it.

It is worth being exact about what that does not change. A passkey works only in the browser. Shell access still uses a key in the directory, a Kerberos ticket from a password, a short-lived SSH certificate issued against that ticket, or the password itself. The closest shell equivalent already works without any change: a FIDO2 SSH key on the same hardware token, created with verify-required, so every login needs the token, its PIN and a touch. Making the certificate path require the passkey is a design question for the next phase, not something to bolt on.

🔺 The first password

Every new account starts with a temporary password that expires on first use. In a terminal that is invisible: kinit says the password has expired and asks for a new one. Keycloak does not. It reads the directory read-only and asks the KDC to check a password; when the KDC answers “expired”, Keycloak reports a wrong password and stops.

The check that found this was written down before the cut-over: an auditor whose first password has expired must still reach the audit dashboard through single sign-on. It failed, the cut-over stopped, and nothing was switched.

The fix is a small page on the enrollment site: name, the temporary password, a new one twice. It does no Kerberos of its own. It runs MIT’s own kinit as a child process, first with nothing on its input to ask the KDC whether the password has expired, then, only if it has, with the passwords on its input so kinit performs the change. A password that is not waiting to be changed is refused with the same words as a wrong one, so the page cannot be used to change anyone’s settled password or to learn which names exist. It throttles below the KDC’s lockout, so a typo costs a pause, never a locked account.

Figure 2. The first-password page probes first and changes a password only when the KDC says it has expired; every other case gets one answer.

Figure 2. The first-password page probes first and changes a password only when the KDC says it has expired; every other case gets one answer. Running it against a real KDC before release found three bugs the in-process tests could not: an account without preauthentication answered the probe in a way the page read as “realm unreachable”; the KDC’s refusal text already ended in a full stop and the page added another; and the test’s own failure counter read the KDC’s count through a path that always came back empty, so two of its checks had passed by accident.

🔺 Enrolling a device

Laptops and phones join the mesh as road warriors: WireGuard tunnels to the hubs, gated by a directory group. Until now an administrator added each one by hand. The enrollment site lets a person sign in with their passkey, give a device name and its public key, and get the tunnel configuration back. A device stays suspended until an administrator approves it, unless the owner’s group is on an auto-approve list, which is empty.

The first real enrollment failed with “the request was not understood”. The cause was a header. The site sent Referrer-Policy: no-referrer, which is a sensible default, and under it a browser that follows the Fetch standard sends Origin: null on a form submission. The service compared the Origin with its own address and refused everything else. The first-password page has the same check, and it had been tested only with curl, which sends whatever it is told. The fix accepts Origin: null only when the browser’s Sec-Fetch-Site header, which a page cannot set, says the form came from the same site. Until that ships, the site sends a same-origin referrer policy, so browsers send the real Origin.

🔺 Road warriors, mesh first

A laptop reached the mesh through its WireGuard tunnels for everything except the directory. The hubs forwarded their directory VM only from the mesh interface, so a laptop’s Kerberos and LDAP traffic went over the tailnet instead. That worked until an administrator tried to change a password from a laptop: kinit succeeded, kadmin hung. The tailnet’s policy allowed the KDC and LDAP ports from that device and silently dropped the admin ports.

The quick fix added the two admin ports to the tailnet’s policy, for that one device and the write hub only. The lasting fix was not more tailnet ports. The hubs now forward their directory from the road-warrior interface too, but only for devices in the same set that already gates everything else from that interface, so a revoked device loses the directory through its tunnels as it loses the rest of the mesh. The tailnet path answers to the tailnet’s own policy, which allows only a handful of ports. The laptop now uses its tunnels for all of the mesh, the first hub preferred and the second as standby. When neither tunnel has a fresh handshake, its failover script withdraws both routes and traffic falls through to the tailnet. The fallback was tested by pulling both routes and watching the KDC and kadmin still answer through the tailnet.

Figure 3. A road warrior reaches the mesh, directory included, through either tunnel, gated at the hub; the tailnet carries it only when neither tunnel is up, under the tailnet’s own policy.

Figure 3. A road warrior reaches the mesh, directory included, through either tunnel, gated at the hub; the tailnet carries it only when neither tunnel is up, under the tailnet’s own policy.

🔺 The doctor that only watches

The mesh runs BGP over WireGuard between hosts that all render their tunnels from the directory. A new part of musterd, the mesh doctor, reads each host’s WireGuard handshakes and BGP sessions every minute and reports any session that is down, drifting or flapping. It can also repair: re-resolve an endpoint, restart a session within a budget, and eventually render the BGP sessions themselves from the same directory data as the tunnels.

It is live in observe mode on every router except the OpenBSD one, where it is not installed yet. It reports and changes nothing. Repair waits, host by host, until its rendered view has matched that host’s hand-kept configuration for a full day, because a tool that fixes things is only safe once it has proved it understands them.

🔺 What the week taught

  • A client is not a browser. Two pages passed every test written with curl. The first person to use one of them hit the bug, and it was in both. The browser adds its own headers, and policies change them.
  • Install next to things, not over them. The Keycloak collision was found by reading the host before acting, and the claim file now makes the role refuse what it did not create.
  • Names carry rules you did not write. A dash in a cluster name, a relative name in a zone file, a header the browser rewrites under a policy. Each came from a convention two layers down.
  • Widening an inventory widens every play. Marking one host “managed” so a single role could reach it also put it in the target list of every playbook aimed at all hosts. None of them ran, but a production mail relay was one command away from a hardening profile it was never meant to get. Hosts managed for one purpose now sit in a group the others exclude.
  • Run the real thing before release. The dev KDC test cost seventeen seconds and found three bugs, one of them in the test.
  • A flaky test is a hypothesis, not a verdict. The full suite failed twice this week. The first time, two suites failed on timing alone and passed on a rerun. The second time, one flake cascaded through two later suites, and beside it sat a real regression in a dev harness. Reading the logs told them apart; rerunning blindly would not have.

🔺 What is still open

The shell does not yet ask for the passkey. Device enrollment is live, but no device has completed it yet, and the browser extension that will carry it on managed machines waits for its signing key, which is offline. The mesh doctor watches and does not yet mend. One workstation still runs no agent at all. Each of these is written down with the decision it needs, which is most of what makes a week like this one possible.