plaintext
2,485 chars
· 14 lines
Here's the whole thing, from the upstream author's own write-up in the fix commit (cda94ef, v0.14, pushed today at 20:14):
The crash is a SIGPIPE bug in the server. The server had no SIGPIPE handler, and its connection sockets set only timeouts — nothing to stop a write to a departed peer from killing the process. It's thread-per-connection, so when one client disappears, the server's next write to it raises SIGPIPE and the default action terminates the whole process: no error, no core file, every live match gone at once, and daemon -r quietly respawns it — so the only visible trace is "it restarted again". The ideal trigger, in the author's words: a peer that stops reading until the server's write parks in write(), then resets.
That's why it looked like "crashes when proto-check joins." It was never proto-check. Any connection that drops does it — a browser tab, a bot restart, my version probes. My probe just happened to be the connection that dropped most visibly, so it got the blame. Same story for this afternoon's silent outages.
Two more connection bugs were found while chasing it, and they're the likely source of the weird game behaviour (matches ending instantly, "opponent hit a mine", spurious results):
- Two threads wrote to the same TLS connection with no lock — a chat relay or board broadcast from one thread goes out on the other player's socket — which corrupts the record stream and shows up as tlsv1 alert internal error.
- Only the mid-match disconnect path cleared the match's connection slots. A match that ended normally left a slot pointing at a connection its owner was about to free, for the remaining player's thread to write through — a use-after-free.
(1/2)
[2026/10/03 20:43] Cherry: Where the fixes are: all of it is in that one commit — SIGPIPE ignored, a per-connection write lock, and reference-counted connections. The terminal client had the identical SIGPIPE gap too (on our side that's the one that could leave your terminal in raw mode / the pane looking wedged), and our client is already rebuilt at 0.14.
What still needs doing: iRobbery's server is still running the old binary. He needs git pull → rebuild → restart the service — restart, not just rebuild; the running process keeps the old code until it's bounced. The repo even ships a deterministic repro (test/server_probe.mjs, reports "killed by SIGPIPE" before the fix), so he can confirm it on his own box before and after. (2/2)