`POST /renew` is authenticated by mTLS with the very leaf it renews, so once
that leaf lapsed the host could never renew it and the tunnel stayed down until
an operator re-paired by hand. Production hit exactly this: the Mac slept
through its 8h renewal window, the 24h leaf expired, and the agent then logged
`client certificate has expired; renew before dialling` 6380 times over 8 days
without recovering.
Three layers independently refused an expired leaf, so all three had to move:
- agent `buildTlsOptions` gains an opt-in `expiredGraceMs`. Absent/0 keeps the
historical fail-closed behaviour, and the TUNNEL dial never passes it — only
the renew transport does. Past the window it throws the new
`CertExpiredBeyondGraceError`.
- agent rotator routes an already-expired leaf to a separate recovery endpoint
and treats "beyond grace" as TERMINAL: report once via `onExhausted`, stop
retrying, and name the fix (re-pair) instead of spamming warnings forever.
- control-plane `assertPresentedCertTrusted` grants a bounded grace on
`notAfter` only. Chain validation, SPIFFE identity, `notBefore`, and the
registry `active`/account checks are all unchanged, so revocation still bites.
- new `deploy/nginx/recover-mtls.conf` (:8472). The strict enroll vhost cannot
host this: under `ssl_verify_client optional` nginx answers a bare 400 as soon
as a presented cert fails verification, so an expired leaf never reaches the
location — and the directive is server-level, not per-location.
Grace defaults to 30 days on both ends and is configurable (`RECOVER_URL`,
`expiredRenewGraceMs`). The honest trade is recorded in the code: a stale stolen
leaf stays reusable for the window, which widens an existing exposure (an
unexpired stolen leaf already renews indefinitely) rather than opening a new one.
Verified: agent 296/296, control-plane 286/286, tsc clean on both.