Building Blue Canoe · 10 of 19
Going Live Was the Short Part
The preparation, staged controls and production lessons behind Blue Canoe's mail cutover.

This account was drafted on 19 August 2026. Operational states, versions and test results describe that period unless a later update is explicitly dated.
The mail cutover was one of the shortest parts of the mail project.
That was not because mail is simple. It was because most of the dangerous questions had already been dragged into daylight before production depended on the answers.
The network edge had been replaced. Authoritative DNS had moved onto the hidden-master design. DNSSEC was validating where registrar access allowed it. The three mail nodes had been built, storage prepared, users and aliases imported, replication checked and all 30 mailboxes migrated.
By the time the public records changed, the cutover was the final stage of a much longer dependency chain.
Give each node a clear job
The production topology uses mx1 as primary, mx2 as secondary and mx3 as the tertiary edge. The first two carry the main mail service and replicated mail storage. The third provides a further mail route and the public support endpoints around the platform.
The worker could reach each node over the backend network. mx3's recipient synchronisation produced a known view of the hosted domains and accepted recipients rather than turning the tertiary server into an indiscriminate relay.
Clear roles made validation easier. A service absent from mx3 could be an intentional design choice rather than a fault. A queue on any node, however, was something to inspect and explain.
Security features need an order too
DANE was one of the original reasons for rebuilding DNS, but we did not publish every strong transport control at the first possible moment. TLSA binds the identity of a live endpoint into DNSSEC; MTA-STS asks supporting senders to enforce valid TLS against a published MX policy. Both can turn a configuration mistake into failed delivery.
The order was deliberate: stabilise DNS, establish DNSSEC, bring conventional mail into production, settle the MX identities and certificates, observe TLS reporting, then introduce DANE and MTA-STS as separately testable stages.
By 19 August, both were live for blue-canoe.net. Independent testing reported valid DANE for mx1, mx2 and mx3. MTA-STS was in enforce mode with a cautious one-day cache lifetime. Before enforcement, 11 parsed TLS reports recorded 63 successful sessions and no failures. Google's first three complete enforce-mode reports then recorded another 12 successful sessions and no failures.
A green test can still be testing the wrong thing
TLS provided a useful reminder. After the platform was live, an external test appeared to show protocol support that did not match the intended posture. We did not dismiss it and we did not immediately redesign the mail system.
The effective Postfix settings were inspected, the protocol configuration was corrected where required, services were reloaded and the endpoints were tested again. mx1 and mx2 then passed the intended TLS checks and test messages delivered successfully.
The important point was not that a setting once needed correction. It was that an external observation disagreed with our expectation and therefore became work until the discrepancy was explained.
Production supplied the better tests
A brief power interruption provided one of them. All three nodes recovered and their queues were empty. Later, normal operation supplied quieter but more useful tests: SRS forward and reverse handling, working bounce paths, user-driven spam and ham feedback, certificate renewal dry-runs, external DANE validation and daily production reports that could be checked at a glance.
The report also caught a transient PTR lookup timeout on mx3. One missing lookup produced two visible failures because FCrDNS could not be evaluated without the PTR. Unbound had not restarted, the journal was clean and a later 75-check validation passed. We recorded the event and did not change a healthy resolver merely because a report had once been red.
That same report email then revealed a separate problem in itself: Gmail reported a DKIM body-hash mismatch. Oversized 8bit HTML lines were being changed after signing. Encoding the MIME bodies as base64 before signing fixed the fault; the candidate message then passed DKIM, SPF and DMARC, and the released report moved to v0.1.7.
At the point of this August account, the proposed live operations page and broader central logging remained future work. The reporting foundation was already producing text, JSON, HTML and retained evidence.
The execution window was supposed to be boring
The most visible production act was changing DNS so that mail flowed to the new platform. The meaningful work was everything that made that act reversible, observable and unsurprising.
There were corrections after launch, and there will be more. SRS, feedback handling, reporting, DANE and MTA-STS each had their own failure paths and therefore earned their own evidence rather than being smuggled into one triumphant definition of 'finished'.
A network is never finished, and neither is a mail platform. The useful state is not finished. It is understood well enough that the next change can be made deliberately.
Going live was the short part because planning had already done most of the work.
September update
On 14 September, the public mail-trends page was installed and its initial publication checked. That is a specific step beyond the August reporting position, not a claim that every proposed operations or central-logging feature is complete. The installation record still left the first normal scheduled refresh awaiting observation. Operating Blue Canoe returns to that work in One Green Report Is Only One Day.