Keeping your automations healthy
Automations finish quietly. A flow that starts failing, an agent whose task lookup errors a third of the time, a job whose runs have slowly gone from clean to “finished, but with errors” — none of that pages anyone. The evidence is sitting right there in your Runs log, but reading it every day is nobody’s job.
The Jobs hygiene reviewer makes it somebody’s job. It’s an agent that reads how your automations have been finishing, on a schedule, notices the patterns worth caring about, and drops a proposal in your Inbox — “this looks like it’s breaking, here’s the fix” — for you to approve or wave off. It never changes anything itself. Silence from it means your automations are healthy; a proposal means something deserves a look.
What it watches
Section titled “What it watches”Every automation run — a flow or an agent — finishes in one of a few states: it completed cleanly, it completed with errors (it got to the end, but something went wrong along the way), or it failed outright. There’s also a quieter signal: an agent run can finish green and still have hit an error it recovered from on a retry — fine once, worth noticing when it happens run after run, because a bug that self-heals today won’t forever.
On each of its scheduled passes (daily by default), the reviewer reads the recent run outcomes across your automations and looks for patterns, not one-off blips:
- A specific flow or agent that has failed repeatedly — a rate, not a single flaky run.
- A tool an agent keeps erroring on, even when it recovers each time.
- A job whose share of completed-with-errors runs is climbing versus its clean ones.
A single failure isn’t a finding. It’s the trend that earns a mention.
How it tells you — and what it won’t do
Section titled “How it tells you — and what it won’t do”When the reviewer spots something genuine, it doesn’t fix it. It proposes. A card lands in your Inbox naming what it saw — the affected flow or agent, the rate, the window — and a suggested next step. You Approve, Decline, or ignore it. This is the same propose-and-confirm model every agent uses when it wants to change something: nothing happens until a person says so.
Approve a finding and the reviewer’s one and only write happens — it files a maintenance task describing the pattern and the suggested fix, so the work is captured somewhere you’ll see it. That’s the whole extent of what it can do. It reads run history and it writes a task once you’ve agreed. It never edits a flow, changes a setting, touches your data, or “fixes” anything on its own.
Most blips fix themselves before you see them
Section titled “Most blips fix themselves before you see them”Not every failure is a problem. A lot of them are a model provider having a bad second — a server error, a moment of being over its rate limit, a connection that dropped. Nothing was wrong with your flow; it just asked at an unlucky moment. Routario handles those in two places, so they never reach you.
A stumble on one AI step is retried on the spot. If the model returns a server error, says you’re going too fast, or the connection never lands, the step waits a couple of seconds and asks again, then waits a little longer and asks once more before giving up on that model. If the model is genuinely unavailable — or simply takes too long to answer — the step moves on to the next model in its fallback list rather than sitting there retrying something that isn’t coming back.
A run that failed before it did anything is simply run again. If a flow died partway and every step it had finished was a read — looking a customer up, reading a ticket, classifying some text, working something out — then re-running it changes nothing except giving it another chance. Routario notices those within a couple of minutes and quietly starts them over, up to two attempts.
One re-run at a time
Section titled “One re-run at a time”Two attempts is the ceiling, and each one has to earn its place. Routario starts another only when the previous attempt failed the same clean way:
- Never two at once. While a re-run is going, nothing else starts for that run, however often Routario checks in the meantime.
- A re-run that reaches the end closes the matter. Once an attempt finishes, that run is done — it is never tried “once more” afterwards. The work happened once, and that is the point.
- A re-run that got further than the original stops the chain. If an attempt managed to send or write something before it failed, nothing more starts on its own: another try would repeat what that one already did.
- An attempt that breaks still counts. A re-run that hits an internal error instead of finishing is marked failed and reported like any other failure, rather than sitting there looking like it is still working.
Each attempt also starts clean, rather than inheriting the circumstances of the run before it. If somebody had confirmed a send for the original run, that confirmation is not carried over — the new attempt asks again before doing anything irreversible.
One more thing worth knowing if you switch this on: Routario waits about a minute and a half after a failure before trying again, which is long enough for the run to finish unwinding and for a provider’s bad moment to pass. It also only picks up failures from the last six hours, so it never resurrects last week’s.
What is never re-run automatically
Section titled “What is never re-run automatically”This is the part worth trusting, so it’s worth being precise about. A run is only restarted when nothing had been written yet. Re-running starts the flow from the beginning, so anything already done would be done twice — and Routario will not do that to you.
A run is left alone, for you to look at, if any completed step had already:
- Sent something out — an email, a WhatsApp or Slack message, a webhook, a write into a connected system.
- Changed something in Routario — created or updated a record, a task, a contact, a note.
- Handed work to an agent or another flow. Even though starting one is itself harmless, whatever it went on to do is not — the agent may well have replied to a customer before the flow tripped over.
It also leaves alone anything that would fail identically the second time — a missing required field, an answer that couldn’t be read. Those aren’t bad luck; they’re something to fix, and retrying them would only hide them.
A run interrupted by a restart picks up where it stopped. Routario restarts from time to time — an update, a move to another machine. A flow that was in the middle of a step at that moment doesn’t start over and doesn’t silently die either. Every step it had already finished is kept, and the step it was on is looked at with the same care as above: if that step was a read, it is simply done again and the run carries on; if the step had already finished and only the bookkeeping was lost, the run continues with the next one. If the interrupted step was one that sends or writes, the run is stopped and marked with a plain reason naming the step — because whether the email went out or the record was written in that last instant is exactly what can’t be known from outside — so you can check and re-run it yourself. A run is resumed this way at most twice; a step that keeps taking the process down with it is a bug to fix, not something to retry forever.
You also hear directly when a run goes wrong
Section titled “You also hear directly when a run goes wrong”The reviewer looks for patterns over time. Separately, and much more simply, a single bad run tells you itself.
When one of your agents finishes a run badly — it failed outright, or it finished but something it tried didn’t work — the agent’s owner gets a notification in the bell. No pattern needed, no review cycle: the run knows it went wrong, so it says so.
The same is true of a flow. A flow that fails tells the person who created it, wherever it broke — an agent step, a lookup, a send, anything. That matters because a flow is usually the thing a customer is actually waiting on: if a reply never got drafted, the flow failing is the event, not whichever step happened to break.
Three things shape what you actually see, all deliberate:
- One notification per agent, and per flow, per day. Something having a bad afternoon is one ping, not forty. The first bad run of the day tells you; the rest are waiting on that agent’s or flow’s page when you get there.
- It points, it doesn’t ask. The notification names what broke and why, and that’s it — there are no buttons on it. Something went wrong is information, not a decision, and your Inbox is for decisions.
- A problem that fixed itself stays quiet. If a first attempt failed and a later one worked, that’s a working automation, not a broken one, and you won’t hear about it.
- You’re not told while a retry is still coming. If a failed flow is one of the ones Routario is about to restart on its own, it says nothing yet. You hear about it only once the retries are used up and it’s still failing — so a ping means it really stuck, not that something wobbled.
You’ll see something like “Learning judge finished with errors” for an agent, or “Draft on ticket failed” for a flow, with the reason underneath. If it had been retried first, the notification says so — “Auto-retried 2× and still failing” — which is your signal that this one isn’t bad luck.
Notifications go to the person who created the agent or flow. If that isn’t set, they go to the owner of the Job it belongs to.
A worked example
Section titled “A worked example”The reviewer’s morning pass reads the last week of runs and notices your Growth agent’s List Tasks step errored in 6 of its 20 runs — each time it recovered and finished green, so no run looked broken on its own. It proposes:
Jobs hygiene reviewer — needs a decision Growth agent’s List Tasks tool errored in 6 of 20 runs this week (all recovered). File a maintenance task to fix the status filter? Approve · Decline
You approve. A task — “Growth agent: List Tasks errors ~30% of runs — check the status filter” — appears in your tasks with the run details, and you (or whoever owns that agent) picks it up. Nothing about the Growth agent changed; you just found out about a slow leak before it became an outage.
How you get it
Section titled “How you get it”The Jobs hygiene reviewer is a maintenance job set up by whoever runs your Routario instance, rather than something you install from the jobs catalogue yourself. Once it’s in place it runs on its own schedule and its proposals arrive in the Inbox like any other. If you’d like it watching your workspace, ask your Routario administrator or contact.
Where to go next
Section titled “Where to go next”- Tell at a glance whether a Job is healthy — the same run outcomes as you see them yourself, on the Jobs board.
- How agents ask before they write — the approval model the reviewer’s proposals use.
- See what happened — the Logs page — the same run history the reviewer reads, in the Runs lens.
- See an automation’s run history — the same outcomes for one agent, flow or Job, where you can filter and export them yourself.
- Jobs — what a job is and how one bundles agents and flows toward an outcome, including what a Broken reference on its Elements list means.