Three processes and a restart policy
Small always-on services do not fail in interesting ways. They get killed during a patch window, or they leak until the allocator gives up. What helps is unglamorous:
- One supervisor per process, with a restart policy and a backoff.
- Logs that rotate on size, not on hope.
- A health endpoint that says something true about dependencies.
None of that needs a platform team. It needs thirty minutes and someone who will run the checklist.