Untitled Paste

03 Oct 2026 18:50
2 views
Expires 02 Nov 2026 18:50
plaintext 2,473 chars · 24 lines
Then my explanation doesn't cover what he's seeing — so I tested the fixed code rather than argue. The repo ships its own probe harness (test/server_probe.mjs) that owns the server process and reports how it dies, and its probe 1 is exactly my probe's pattern: connect, send a HELLO the server must reject, read the rejection, then RST the socket.

I built 0.14 here and ran it: 60 vanishing peers, server survived, still accepting new players afterwards. So the code as it stands does not die from a proto-check-style connection.

Which leaves one likely explanation: the binary actually serving his players isn't the cda94ef build. The classic is rebuilding without restarting — the running process keeps the old image until it's bounced — and the commit just before this one (22392b0) added the systemd unit / rc.d script for the game server, so it's worth checking which path that unit actually execs.

Two ways to settle it on his box:

- Whatever supervises it records the exit status. SIGPIPE (exit 141) = the old binary is still live and the fix simply isn't running. Anything else — segfault 139, a core file — is a different bug and I'll chase that instead.
- He can reproduce it without touching production: cmake --build build --target probe-test spawns its own server on 127.0.0.1:14443 and reports the exit signal. If that passes on his machine while production still dies, it's the deployment, not the code.

I'm running the full probe suite (all four probes) against our local 0.14 build now and will report what it finds.


-------------------


Full suite finished on the local 0.14 build: 10 passed, 0 failed — including the two that matter most here:

- probe 2, the deterministic version of the crash (a peer that stops reading until the server's write blocks, then resets mid-write) — server survived, and so did the other player's connection
- probe 3, a player leaving a finished match — the dangling-slot guard the fix added — survived, server still serves new players after

So every pattern that used to kill it is handled by that commit. That leaves the deployment: if production still dies when a connection drops, the process serving the players isn't running this code. The exit signal settles it — 141 (SIGPIPE) means the old binary is still live; anything else and I'll dig into that instead. Worth also checking which path his service unit execs, since that unit was only added a commit earlier.