Case files
Ecommerce case studies, written as postmortems.
Symptom, root cause, fix, outcome. That’s how an engineer writes up an incident, so that’s how we write these ecommerce case studies. Some client names are missing because they asked. None of the technical detail is.
Revenue 14% below forecast for six weeks. Checkout worked every time the owner tested it. The gateway dashboard was clean. Support was quiet, which is exactly why nobody caught it.
The host cut the outbound request timeout from 30s to 8s and told nobody. The gateway’s webhook takes 9–12s to answer, so every callback timed out and the orders sat in pending payment instead of moving through.
We raised the timeout, put failed webhooks into a retry queue, and added a reconciliation job that re-queries the gateway for anything stuck longer than 15 minutes. Monitoring now flags pending-ratio drift within 30 minutes, so the next quiet change on the server doesn’t cost six weeks.
$180k recovered in the 30 days after the fix. Some of it processed straight out of the pending queue, the rest recaptured.
3:17am UTC alert on the busiest sales night of the year. Database connections exhausted. The load balancer was still sending shoppers to a node that had stopped answering.
A plugin writing session data held database connections open under load. The health check never touched the database path, so as far as the balancer knew that node was fine. It couldn’t take a single checkout.
We widened the connection pool and moved those writes to object cache. The health checks were rewritten to exercise the full checkout path, database included. Then we drained the sick node and recycled it.
Checkout back up in 41 minutes, an estimated $340k of peak-night revenue protected, and a load test that now runs before every major sale.
More ecommerce case studies go up in Field Notes as we finish the work, alongside diagnostics and plugin-conflict field guides. If your symptom is on this page, the write-up probably is too.