architecture·ThotExperiment / Headero·6 min read
Splitting a Node monolith without stalling user APIs
Chat spikes and nightly cron jobs were slowing down login and scroll. Here’s the order we pulled things out, and why a gateway came before any service.
01 · The symptom
One Node process served chat, discovery, sessions, profile views, admin tooling and three cron jobs. They shared one scaling policy. When chat got busy, or the expiry job ran, every user-facing endpoint paid for it.
02 · Gateway first
Before extracting anything, the Node app became a gateway. Every later move was then a routing change — the Flutter client never knew.
// gateway routes (illustrative) route('/chat/*', proxy(CHAT_SVC) // go route('/discover/*', proxy(DISCOVERY_SVC) // go route('/sessions/*', proxy(SESSIONS_SVC) // spring boot route('/views/*', proxy(VIEWS_SVC) // go
03 · Order of extraction
- Chat — loudest traffic, cleanest boundary. Go.
- Discovery & sessions — the hot read paths. Mongo aggregations with Redis ranking; Mongo truth with Redis cache-first.
- Views — became the view → notify → profile-open loop.
- Admin — its own HPA so it can’t starve users.
- Crons — to Lambda, reaching session APIs only through an internal NLB.
04 · What it bought
10 msp50 at the service
50 msp50 end-to-end scroll
9 → 20weekly min in app
Because every extraction was a routing change at the gateway, each one could be rolled back the same way. Chat spikes and nightly jobs no longer stall the APIs users touch, and the admin console can no longer starve them.