Runbook 26 — Automated SSH host-CA (online, Tang-sealed)¶
Goal¶
Make SSH host-certificate signing reliable and hands-off without weakening the part that
matters. Today both CAs live on the YubiKey and every host cert is signed by hand (runbook 20,
scripts/rotate-host-certs.sh) — so a rebuilt host or a routine rotation silently lapses. This
runbook splits the two CAs by what they grant:
- User-CA — grants ACCESS (forge an
admincert → root on the fleet). Signed rarely and deliberately. Stays on the YubiKey, manual. Unchanged. It is the offline recovery root: no online compromise can reach it, so any online compromise is revocable + re-establishable. - Host-CA — grants IDENTITY (impersonate a host's sshd → MITM; never a login). Signed
frequently. Becomes an online, automated CA, its key Clevis+Tang-sealed to
tang1and decrypted per-signature onctrl.
Property: signing needs ctrl (the signer) AND tang1 (the tang) live + on-LAN. A
stolen / off-LAN ctrl fails closed (clevis can't reach tang1 → no key → no signing). The
accepted cost is that a live, on-LAN ctrl compromise can mint host certs — but such a
compromise already yields fleet-root via the control cert, so the only marginal loss is
host-identity forgery, which is bounded (28-day certs + KRL) and recoverable (revoke +
fall back to the still-trusted YubiKey host-CA). Tang-sealing (not TPM/vault) is deliberate: a TPM
travels with a stolen box; tang1 stays behind.
This is the simpler of two designs considered. The rejected alternative put the signer on
tang1itself (sealed to a 2nd tangd onctrl) so even a livectrl-hack couldn't sign — at ~3× the moving parts. We chose to separate identity from access via the offline user-CA and accept the live-ctrlresidual. If host-identity forgery ever needs containment separate from fleet-root, see that variant in git history / the design discussion.
Architecture¶
host-CA PRIVATE key ──Clevis+Tang seal──► /etc/lab/host-ca.jwe (on ctrl, 0640 root:almalinux)
│ decrypt needs tang1 (10.20.10.11) live+on-LAN
host-ca-autosign.timer (weekly, on ctrl) ──────────┤
self-update (ff-only to a SIGNED origin/main) → load-control-key (tang1-gated) → gather
pubkeys → clevis decrypt → ssh-keygen -s (/dev/shm, shred) → deploy (site.yml --tags sshca)
→ ntfy every sign
host-CA PUBLIC key ──► 3rd line of ssh_host_ca.pub ──► common @cert-authority (fleet trusts it,
alongside the YubiKey host-ca@A/@B)
The renewer is scripts/host-ca-autosign.sh, wired by the control role
(host-ca-autosign.{service,timer}, vars hostca_* in roles/control/defaults/main.yml):
28-day certs, renewed at <25 d left (floor ≈21 d — a healthy weekly run renews every cycle), or immediately when a host key changes
(rebuild) or no cert exists. It reuses the existing gather (playbooks/gather-host-pubkeys.yml),
deploy (site.yml --tags sshca), principals/serial scheme (rotate-host-certs.sh), and the
tang-gated scripts/load-control-key.sh.
The host-cert deploy is gated to this signer.
host_certs/<host>-cert.pubis gitignored and local to each checkout, and/opt/myos(this autosign) is its sole authority. Thecommonrole's "Deploy this host's signed host certificate" task therefore runs only when the caller passes-e common_host_cert_deploy=true(defaultfalse) — whichhost-ca-autosign.shand the break-glassrotate-host-certs.shdo, and nothing else does. This stops a converge from any other checkout (the agent's~/src/myos, a git worktree) from redeploying its own stalehost_certsand reverting a fresh cert — the failure mode behind the 2026-07-26caddyincident, where an agent-checkout converge kept reverting caddy's host cert to a 3-week-old copy until it nearly expired. The othersshcatasks (CA pubkeys, KRL,@cert-authorityknown_hosts — all git-tracked and identical across checkouts) still run from any checkout; only the per-checkout host-cert copy is gated.
Self-refresh (step 0 of the script). So the /opt/myos signer checkout tracks main
without a manual git pull, the script first git fetches origin and fast-forwards only
if origin/main's tip is a commit signed by an allowed signer (scripts/allowed_signers —
operator keys, not the agent), then re-execs once under the refreshed code. This keeps the
/opt/myos privilege boundary intact: the checkout stays almalinux-owned and agent-unwritable,
and a main push (which the T2 agent can do) still can't steer the root-adjacent CA signer,
because unsigned/agent-signed tips are refused (WARN + ntfy) and the signer stays pinned to
its current checkout. It fails safe, never blocking renewal: a failed fetch (offline/forgejo
down) just proceeds on the current checkout. Effective only once signed commits are enforced on
main (real signers populated in the allowed_signers overlay + operator commits actually
signed + forgejo branch protection requiring signed commits — runbook 28); until then
verify-commit fails and the signer safely never auto-updates. The /opt/myos read-only Forgejo
credential that enables the fetch is set up in runbook 15 §b (ctrl-opt-myos-ro).
Who signs the tip, in practice (2026-07-26). A merged PR's tip is a merge commit created by Forgejo, not by any human — so it was unsigned, and the signer froze after every merge. Two trusted signers now exist, and the split is the point:
Signer Format Signs Trusted via Forgejo (control-plane) OpenPGP the merge commits it creates a gpg keyring on the verifying side Operator (YubiKey) SSH hand-made commits — optional, currently unused scripts/allowed_signersForgejo 10 has no SSH signing (
SIGNING_FORMATarrived later), hence OpenPGP — harmless, becausegit verify-commitpicks its backend from the signature itself, so both formats verify side by side (the verifying side can hold both trust stores).In practice only the Forgejo row is used, and that is by design (2026-07-27). Branch protection blocks direct pushes, so every tip of
mainis a merge commit created by Forgejo — never a hand-made commit. The operator therefore never has to sign anything: their control point is the approval click, not a signature.scripts/allowed_signersstays placeholders (trusting nobody on the SSH side) unless someone deliberately wants hand-made commits to verify — it is an available slot, not an outstanding task.⚠ That rests on one invariant: the tip must always be a Forgejo-created commit.
rebase,rebase-explicitandfast-forward-onlymerges do not create one — the tip becomes the PR branch's own commits, which are agent-signed and deliberately untrusted, so the guard fails and the signer silently freezes. The repo is therefore restricted to the merge-commit style only (all others disabled 2026-07-27). Squash is also disabled pending a test — it is Forgejo-created and probably signed, but that was never verified; test it on a throwaway PR before re-enabling, and confirmgit log --show-signature -1 origin/mainstill reports a good signature.The Forgejo signature is only worth trusting because merges are gated.
mainhas branch protection requiring an operator approval and forbidding direct pushes, so the signature attests "an approved pull request was merged". Without that gate it would degrade to "somebody merged" — and since the agent can call the merge API, that would silently hand the agent the power to steer the root-adjacent CA signer, i.e. exactly the boundary this guard exists to hold. If you ever remove the approval requirement, remove Forgejo from the trust store in the same change. The signing key itself is unattended and regenerable (it signs merges, it is not a CA); replacing it costs one re-import on the verifying side.
One-time bootstrap (operator — credential-bearing)¶
Run on ctrl, with tang1 up. The raw key never enters the Ansible vault (ctrl auto-opens
the vault via its TPM, which would let an off-LAN ctrl bypass the tang gate). Escrow it to
KeePass + offline only — the deliberate 2-way exception to the 3-way policy.
- Generate + seal (transient, in
/dev/shm):umask 077; cd /opt/myos ssh-keygen -t ecdsa -b 384 -N '' -C host-ca@online -f /dev/shm/hostca # match the CA family eval "$(scripts/load-control-key.sh)" # (only if you'll run ansible after) clevis encrypt tang '{"url":"http://10.20.10.11"}' -y < /dev/shm/hostca > /etc/lab/host-ca.jwe chmod 640 /etc/lab/host-ca.jwe; chgrp almalinux /etc/lab/host-ca.jwe clevis decrypt < /etc/lab/host-ca.jwe | diff - /dev/shm/hostca && echo "JWE round-trips OK" - Escrow
/dev/shm/hostca(the raw private key) + its fingerprint (ssh-keygen -lf /dev/shm/hostca.pub) into KeePasslab.kdbx("SSH host-CA online key") and an offline copy. Then shred:shred -u /dev/shm/hostca. - Trust it additively — append the public key as the 3rd line of
ansible/roles/common/files/ssh_host_ca.pub(keephost-ca@A/@B— permanent break-glass):cat /dev/shm/hostca.pub # ecdsa-... host-ca@online → paste as a new line, then: git add ansible/roles/common/files/ssh_host_ca.pub && git commit -m 'sshca: trust online host-CA (additive)' && git push cd ansible && ansible-playbook -i inventory.ini playbooks/site.yml --tags sshca # fleet trusts both now - ntfy token — create a write-only ntfy token for the signer and vault it as
vault_hostca_ntfy_token(same pattern as the grafana token; CLAUDE.md ntfy notes):ssh ntfy 'podman exec ntfy ntfy token add --expires=never hostca' # → tk_… # add `vault_hostca_ntfy_token: <tk_…>` to the gitignored vault (ansible-vault encrypt_string) - Deploy the control role (installs clevis, the token file, the timer):
ansible-playbook -i inventory.ini playbooks/site.yml --tags … --limit ctrl(or a full run). - (Optional) ctrl forgejo push key so the signer can push renewed certs (else it commits locally + warns; deploy still works from the local checkout).
Roll out + verify¶
# canary first (one low-risk host)
cd /opt/myos && HOSTCA_RENEW_DAYS=99 scripts/host-ca-autosign.sh # force-sign everything? no — limit:
# for a single canary, temporarily point the script at one host, or just let the timer run.
ssh-keygen -L -f /etc/ssh/ssh_host_ed25519_key-cert.pub shows Signing CA = host-ca@online;
a not-yet-renewed host still shows host-ca@A/@B. From a client whose known_hosts has only the
three @cert-authority lines, ssh to both verifies with no TOFU prompt.
- Fail-closed: stop/block tang1, run scripts/host-ca-autosign.sh → clevis decrypt fails →
it signs nothing, logs FAIL-CLOSED, ntfys, exits non-zero. Restore tang1 → next run heals.
- No vault bypass: grep -ri vault_ssh_host_ca ansible/ returns nothing — the tang gate is the
only path to the key.
- Renewal + idempotency: systemctl list-timers on ctrl shows host-ca-autosign.timer; a fresh
matching cert is skipped (no git churn); a <25 d or key-changed cert is re-signed —
with 28-day validity and a weekly timer that means every run renews (floor = 21 d left),
keeping a clean week between the renewal floor and the alert below (rb42 P3 threshold
separation; was <21 d, which parked the healthy floor at ~14-15 d, flush against the alert).
- ntfy: every run that signs pushes a host-CA signed: … notice to lab-alerts (priority 5 +
"HOST KEY CHANGED" if a host key rotated — verify the rebuild).
- Monitoring: the obs alert fleet-ssh-cert-expiring fires at 14 d
(obs_alert_cert_days_left) — a true "signer/tang1 broken" signal: the healthy floor is 21 d,
so 14 d means a full week of renewals was missed. (This doc previously said 7 d — drift, fixed
2026-07-14.)
Central tang access logging (chokepoint audit)¶
tang1 ships its syslog to the obs Alloy syslog-gateway (10.20.20.50:514 → Loki) — see
provision/tang1-setup.sh §5b. Every NBDE decrypt is a socket-activated
tangd@<…-PEERIP:port>.service instance, so the peer IP + operation are captured for every
key-unlock in the lab (pve vmpool unlock, ctrl control-key load, host-CA sign) — independent of the
per-sign ntfy, and catching any rogue clevis decrypt that bypasses the signer. Query in Grafana
Explore: {job="syslog"} |= "tangd". (We do not alert on raw /rec — legitimate unlocks are
frequent; the host-CA ntfy is the proactive signal.) Caveat to confirm at deploy: how verbosely
the installed tangd logs the operation/key-id; the peer IP comes from the unit instance name
regardless.
Disaster recovery¶
- Lose
tang1: renewals stop → obs alert at 14 d left (≈1 week after the loss). Break-glass = re-sign with the YubiKey host-CA (scripts/rotate-host-certs.sh, still trusted). Rebuildtang1(runbook 24); the online host-CA public key is unchanged → no fleet re-trust needed. - Lose
ctrl: rebuild (runbook 15), then regenerate/etc/lab/host-ca.jwefrom the KeePass raw key (re-run bootstrap step 1'sclevis encrypt). Deployed certs stay valid for weeks; the YubiKey host-CA is the interim signer. - Lose both / the raw key: KeePass + offline escrow restores the key; worst case, retire the
online host-CA (drop its line from
ssh_host_ca.pub) and revert to YubiKey-only signing. - Revoke a mis-issued cert: add it to the KRL (
roles/common/files/revoked_keys.krl,--tags krl) and/or drop the online host-CA line to distrust all its certs at once.
Security model (what each single-box compromise yields)¶
| Compromise | Can mint host certs? | Notes |
|---|---|---|
| ctrl, live on-LAN | Yes | …but already has fleet-root via the control cert. Marginal loss = host-identity forgery; bounded (28 d + KRL), revocable, and logged (ntfy + tang audit). |
| ctrl, stolen/off-LAN | No | clevis decrypt can't reach tang1 → fail-closed. |
| tang1 | No (no JWE/signer there) | tang is an on-LAN oracle; it gates off-LAN/stolen, not on-LAN access. |
| the JWE alone, off-LAN | No | inert without tang1. |
| a fleet host | No | the CA key is never on fleet hosts; that host can assert only its own identity. |
| the YubiKey user-CA | n/a | offline; the recovery root — this is what stays untouchable. |