829b4e3a

By: Michael Lynch <git@mtlynch.io>

Widen the shutdown budget so the final sync can finish

Litestream's shutdown-sync-timeout defaulted to 30s and fly.toml's
kill_timeout was also 30s. Those two numbers being equal meant Fly's
SIGKILL was scheduled to land at exactly the moment Litestream was still
allowed to be uploading. Any final sync that took the full budget lost
the race, and Fly won.

That race matters more here than it would in most deployments, for two
reasons that compound:

  - sync-interval is 10m, so at any given moment up to ten minutes of
    writes exist only in the local WAL and have never been sent to
    object storage.

  - There is no Fly Volume by design, so the machine's disk disappears
    with the machine. The final sync is not an optimization; it is the
    only thing standing between a routine deploy and ten minutes of lost
    comments, uploads, and sessions.

Raise shutdown-sync-timeout to 60s and kill_timeout to 90s. The budget
is now 15s of request draining plus 60s of syncing, leaving 15s of slack
for process teardown inside the 90s Fly allows.

This costs nothing in the common case. Shutdown normally completes in a
second or two; the larger ceiling only gets used when object storage is
slow, which is precisely the situation where giving up early is most
expensive. The tradeoff is that a pathological deploy takes longer to
roll, which is acceptable for a single-machine app that already accepts
brief downtime during deploys.

Both values carry comments pointing at each other, since changing either
one in isolation reintroduces the race.

Co-Authored-By: Claude <noreply@anthropic.com>

Suite timing

Time to Start Worker time Duration Time to finish Idle
Config 32s 3s 3s 36s 32s
Eval 1m45s 1m36s 1m36s 3m22s 1m09s
Build 3m13s 7m36s 5m19s 8m33s 2m52s
Suite 32s 9m17s 8m01s 8m33s 4m34s

Timeline

0s2m2m20s2m40s3m3m20s6m20s6m40s7m7m20s7m40s8m8m20s