I built a 24/7 self-healing agent system
Five machines running agents around the clock. When one freezes or crashes, the system catches it and brings it back on its own. I built the orchestration and the recovery around the models. Next I want to build the models themselves.
Nothing important depends on me being awake.
Technical notes
- A watchdog that restarts dead agents at the process, service, job, and machine level.
- Durable job state, so interrupted work picks back up after a restart.
- Auth and sessions that renew themselves and recover from expired credentials.
- Memory and state shared across all five machines.
- Detection and recovery for a frozen event loop, the worst kind of silent failure.
- Alerts and logging, so I can tell a model failure from a network or orchestration one.