Almost every published claim about AI agent reliability has no denominator. “Agents are brittle.” “Most pilots never reach production.” “Hallucination is the blocker.” Nobody says out of how many, on what, measured when.
So here is ours. On 27 July 2026 we read the scheduler on our own production agent host — 19 specialist agents on one Fedora 42 VPS, 4 vCPU and 15 GiB, running natively under systemd — and wrote down what every scheduled job was doing. This post is that reading, including the part where our own arithmetic does not tie out.
What is the AI agent failure rate in production?
Six of forty-seven scheduled jobs were in an error state.
| Status | Jobs | Share of defined jobs |
|---|---|---|
ok — last run succeeded |
34 | 72% |
error — last run failed |
6 | 13% |
| Never run | 7 | 15% |
| Total defined | 47 | 100% |
Which number you call “the failure rate” depends on the denominator you pick, and the choice is not neutral:
- 13% of defined jobs were failing.
- 15% of the 40 jobs that had ever executed were failing.
- 30% of defined jobs were not doing their job — failing or never run.
That third number is the one an operator should care about and the one you will never see in a vendor deck. A job that has never run is not a success.
The cadence behind those 47 jobs: 12 daily, 9 weekly, 8 continuous interval loops, plus twice-daily engagement jobs and a one-shot. Twenty-seven of the 47 are dispatched by a single orchestrator agent into other agents rather than being scheduled independently, which matters for blast radius — the architecture write-up covers that layer, and the schedule and uptime record covers how the jobs are laid out and how a failure surfaces.
What actually broke
Every failing run left a real error string. Grouped by class, with the counts recorded against jobs:
| Failure class | Jobs | Verbatim error |
|---|---|---|
| Delivery channel misconfig | 8 | platform 'discord' not configured/enabled · no delivery target resolved for deliver=none |
| Model entitlement drift | 3 | HTTP 404: This model is unavailable for free. The paid version is available now - use this slug instead: meta-llama/llama-3.3-70b-instruct · Error code: 400 - AccessDenied.Unpurchased: Access to model denied |
| Upstream 5xx | 2 | RuntimeError: HTTP 500: Internal server error |
| Provider unconfigured | 1 | No LLM provider configured. Run 'hermes model' to select a provider |
Read that table again with one question in mind: which of these is a model quality problem?
None of them. The agent reasoned fine. In the largest class it did the work correctly and then had nowhere to put it.
Where our own arithmetic disagrees, and why we are publishing it anyway
The class counts total 14 job-level occurrences. Only 6 jobs were in an error state at the moment we looked. Those two numbers do not reconcile, and we are not going to quietly pick whichever one reads better.
The difference is what each was counted against. The 34/6/7 tally is a snapshot: the status of each job’s most recent run at survey time. The class counts were taken across recorded error output, which includes runs that failed earlier and jobs that were reconfigured afterwards. A job that hit a revoked model slug in June and now runs green still contributed its error string to the taxonomy.
Both are real readings of the same machine, at different scopes. If we had normalised them into one tidy figure before publishing, the tidy figure would have been the invented part. When you read anyone’s agent reliability numbers — including ours — the question to ask is what was the denominator, and was the snapshot the same as the history?
Delivery configuration is the failure mode nobody plans for
The biggest class was not glamorous. The agent researched, drafted, or assembled exactly what it was asked for, and the result went nowhere because a destination was not enabled or no delivery target resolved.
This is uniquely dangerous because of how it looks in monitoring. The model call succeeded. Tokens were spent. The run took a plausible amount of time. Any dashboard built on model-level telemetry reports a healthy system, while the business outcome — a post published, a digest delivered, a lead routed — never happened.
The fix is structural, not clever: treat “produced output but delivered nowhere” as a hard failure, and make the delivery target part of the job contract rather than a runtime lookup that can quietly resolve to nothing. Our publishing agent carries the same rule as an explicit refusal — it will not silently skip a destination that failed — which is the pattern documented in the content pipeline case study.
Model entitlement drift is a supply-chain risk, not a bug
Three jobs died because a model slug they depended on stopped being available on the terms they were using it on. Nothing in our code changed. A provider withdrew a free tier and told us so in an HTTP 404 body.
If your agent names one model and that name is a third party’s commercial decision, your agent has an unmanaged dependency. Teams that would never pin a single unversioned package in production routinely hardcode a single model slug and never think about it again.
Two mitigations, in order of value:
- A fallback cascade. Primary, secondary, and a local or self-hosted last resort. The job degrades instead of stopping.
- Fail loudly on substitution. If the agent falls back, that should be visible in the run record. Silent degradation to a weaker model is how quality drifts without anyone choosing it.
This is the same argument as the context-engineering post, now backed by dated observation rather than someone else’s incident report.
The seven jobs that never ran
Seven scheduled jobs had never executed. No error, no output, no alert — they simply were not doing anything.
This class does not appear in failure-rate discussions because it does not generate a failure. It generates silence, and silence is what every run-outcome monitor interprets as fine. We found them by reading the scheduler’s job definitions against last-run timestamps, which is a cheap check nobody thinks to run until the first time it embarrasses them.
If you take one operational habit from this post, take that one: alert on absence, not only on errors. A job whose last successful run is older than its own cadence is a failure regardless of what its status field says.
What to do about each class
| Class | Standing fix | Where it belongs |
|---|---|---|
| Delivery misconfig | Delivery target validated as part of the job contract; “delivered nowhere” treated as failure | Job definition |
| Entitlement drift | Fallback cascade with visible substitution | Provider config |
| Upstream 5xx | Bounded retry with backoff, then surface; never an infinite loop | Runtime |
| Provider unconfigured | Startup check that refuses to schedule work with no provider | Deploy step |
| Never run | Sweep comparing last-run age against declared cadence | Monitoring |
None of the five is a prompt change. That is the point: the reliability work on an agent system is dependency management, contract design and monitoring — the same work as any other production service. The production readiness diagnostic scores a system against these classes if you want to check your own.
What this number does not prove
Being explicit, because the number is quotable and it would be easy to let it imply more than it can:
- It is one host, read once. One machine, one survey, 2026-07-27. It is not an industry failure rate and we have no basis to claim it is representative.
- It counts job outcomes, not answer quality. A job that ran successfully and produced a mediocre draft counts as
okhere. Nothing in this reading measures whether the agents were substantively right. - It says nothing about cost. This host carries no per-task cost ledger, so we publish no cost or ROI figure from it — for this post or any other.
- It has no before/after baseline. We cannot tell you the system got more reliable, only what it looked like on the day.
What it does establish is narrower and, we think, more useful than the usual claim: on a real fleet doing real scheduled work, the things that broke were dependency drift, configuration and upstream outages. The model being wrong was not on the list.
If you are scoping an agent and want the failure classes designed out before the first run rather than discovered in month two, that is what the free agent blueprint is for.