How Erlang and Network Failures Caused a Crypto Exchange Outage
Summary
This incident statement explains separate technical failures that caused trading-platform downtime. A bot’s repeated requests triggered an Erlang vulnerability that brought down the exchange’s web nodes twice; the matching engine remained unaffected. The exchange reports that it restarted the web nodes in under ten minutes each time and applied a code workaround while reporting the underlying issue to the Erlang team.
A separate failure involved a Linux kernel panic on the main web node while it handled a network-card interruption. Because that node also managed load balancing, rebooting it caused another outage, despite other web nodes and the matching engine remaining online. The statement says the exchange planned to replace that load-balancing setup to remove a single point of failure. It offers an operational account of how front-end availability can fail while matching infrastructure remains available, but gives no independent incident analysis, trading impact measurements, or evidence that the planned architecture change was completed.
Key ideas
- Repeated bot requests triggered an Erlang vulnerability that took down web nodes while leaving the matching engine unaffected.
- A separate Linux kernel panic on the main load-balancing web node caused additional downtime.
- The statement distinguishes front-end availability failures from matching-engine availability.
- The proposed remedy was a load-balancing design without a single point of failure.
- The account does not quantify customer trading impact or confirm completion of the planned change.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.