Blog · Original data · 2026-09-27

AI Agent Failure Rate in Production: 6 of 47 Jobs, and Not One Was the Model

We read the scheduler on our own agent host: 47 jobs, 34 healthy, 6 failing, 7 never run. Every failure was dependency drift, config or an upstream 5xx.

Get my free blueprint

What is the AI agent failure rate in production?

On the one production agent host we can publish figures for, 6 of 47 scheduled jobs were in an error state when we read the scheduler on 27 July 2026 — 13% of defined jobs, or 15% of the 40 that had ever run. A further 7 had never run at all. No failure was caused by the model producing a wrong answer.

This is one host, read once. It is a denominator most vendors do not publish, not an industry rate — and it says nothing about how often the agents were substantively wrong when they did run.

By Amit Kumar7 min readPublished

Almost every published claim about AI agent reliability has no denominator. “Agents are brittle.” “Most pilots never reach production.” “Hallucination is the blocker.” Nobody says out of how many, on what, measured when.

So here is ours. On 27 July 2026 we read the scheduler on our own production agent host — 19 specialist agents on one Fedora 42 VPS, 4 vCPU and 15 GiB, running natively under systemd — and wrote down what every scheduled job was doing. This post is that reading, including the part where our own arithmetic does not tie out.

What is the AI agent failure rate in production?

Six of forty-seven scheduled jobs were in an error state.

Status Jobs Share of defined jobs
ok — last run succeeded 34 72%
error — last run failed 6 13%
Never run 7 15%
Total defined 47 100%

Which number you call “the failure rate” depends on the denominator you pick, and the choice is not neutral:

  • 13% of defined jobs were failing.
  • 15% of the 40 jobs that had ever executed were failing.
  • 30% of defined jobs were not doing their job — failing or never run.

That third number is the one an operator should care about and the one you will never see in a vendor deck. A job that has never run is not a success.

The cadence behind those 47 jobs: 12 daily, 9 weekly, 8 continuous interval loops, plus twice-daily engagement jobs and a one-shot. Twenty-seven of the 47 are dispatched by a single orchestrator agent into other agents rather than being scheduled independently, which matters for blast radius — the architecture write-up covers that layer, and the schedule and uptime record covers how the jobs are laid out and how a failure surfaces.

What actually broke

Every failing run left a real error string. Grouped by class, with the counts recorded against jobs:

Failure class Jobs Verbatim error
Delivery channel misconfig 8 platform 'discord' not configured/enabled · no delivery target resolved for deliver=none
Model entitlement drift 3 HTTP 404: This model is unavailable for free. The paid version is available now - use this slug instead: meta-llama/llama-3.3-70b-instruct · Error code: 400 - AccessDenied.Unpurchased: Access to model denied
Upstream 5xx 2 RuntimeError: HTTP 500: Internal server error
Provider unconfigured 1 No LLM provider configured. Run 'hermes model' to select a provider

Read that table again with one question in mind: which of these is a model quality problem?

None of them. The agent reasoned fine. In the largest class it did the work correctly and then had nowhere to put it.

Where our own arithmetic disagrees, and why we are publishing it anyway

The class counts total 14 job-level occurrences. Only 6 jobs were in an error state at the moment we looked. Those two numbers do not reconcile, and we are not going to quietly pick whichever one reads better.

The difference is what each was counted against. The 34/6/7 tally is a snapshot: the status of each job’s most recent run at survey time. The class counts were taken across recorded error output, which includes runs that failed earlier and jobs that were reconfigured afterwards. A job that hit a revoked model slug in June and now runs green still contributed its error string to the taxonomy.

Both are real readings of the same machine, at different scopes. If we had normalised them into one tidy figure before publishing, the tidy figure would have been the invented part. When you read anyone’s agent reliability numbers — including ours — the question to ask is what was the denominator, and was the snapshot the same as the history?

Delivery configuration is the failure mode nobody plans for

The biggest class was not glamorous. The agent researched, drafted, or assembled exactly what it was asked for, and the result went nowhere because a destination was not enabled or no delivery target resolved.

This is uniquely dangerous because of how it looks in monitoring. The model call succeeded. Tokens were spent. The run took a plausible amount of time. Any dashboard built on model-level telemetry reports a healthy system, while the business outcome — a post published, a digest delivered, a lead routed — never happened.

The fix is structural, not clever: treat “produced output but delivered nowhere” as a hard failure, and make the delivery target part of the job contract rather than a runtime lookup that can quietly resolve to nothing. Our publishing agent carries the same rule as an explicit refusal — it will not silently skip a destination that failed — which is the pattern documented in the content pipeline case study.

Model entitlement drift is a supply-chain risk, not a bug

Three jobs died because a model slug they depended on stopped being available on the terms they were using it on. Nothing in our code changed. A provider withdrew a free tier and told us so in an HTTP 404 body.

If your agent names one model and that name is a third party’s commercial decision, your agent has an unmanaged dependency. Teams that would never pin a single unversioned package in production routinely hardcode a single model slug and never think about it again.

Two mitigations, in order of value:

  1. A fallback cascade. Primary, secondary, and a local or self-hosted last resort. The job degrades instead of stopping.
  2. Fail loudly on substitution. If the agent falls back, that should be visible in the run record. Silent degradation to a weaker model is how quality drifts without anyone choosing it.

This is the same argument as the context-engineering post, now backed by dated observation rather than someone else’s incident report.

The seven jobs that never ran

Seven scheduled jobs had never executed. No error, no output, no alert — they simply were not doing anything.

This class does not appear in failure-rate discussions because it does not generate a failure. It generates silence, and silence is what every run-outcome monitor interprets as fine. We found them by reading the scheduler’s job definitions against last-run timestamps, which is a cheap check nobody thinks to run until the first time it embarrasses them.

If you take one operational habit from this post, take that one: alert on absence, not only on errors. A job whose last successful run is older than its own cadence is a failure regardless of what its status field says.

What to do about each class

Class Standing fix Where it belongs
Delivery misconfig Delivery target validated as part of the job contract; “delivered nowhere” treated as failure Job definition
Entitlement drift Fallback cascade with visible substitution Provider config
Upstream 5xx Bounded retry with backoff, then surface; never an infinite loop Runtime
Provider unconfigured Startup check that refuses to schedule work with no provider Deploy step
Never run Sweep comparing last-run age against declared cadence Monitoring

None of the five is a prompt change. That is the point: the reliability work on an agent system is dependency management, contract design and monitoring — the same work as any other production service. The production readiness diagnostic scores a system against these classes if you want to check your own.

What this number does not prove

Being explicit, because the number is quotable and it would be easy to let it imply more than it can:

  • It is one host, read once. One machine, one survey, 2026-07-27. It is not an industry failure rate and we have no basis to claim it is representative.
  • It counts job outcomes, not answer quality. A job that ran successfully and produced a mediocre draft counts as ok here. Nothing in this reading measures whether the agents were substantively right.
  • It says nothing about cost. This host carries no per-task cost ledger, so we publish no cost or ROI figure from it — for this post or any other.
  • It has no before/after baseline. We cannot tell you the system got more reliable, only what it looked like on the day.

What it does establish is narrower and, we think, more useful than the usual claim: on a real fleet doing real scheduled work, the things that broke were dependency drift, configuration and upstream outages. The model being wrong was not on the list.

If you are scoping an agent and want the failure classes designed out before the first run rather than discovered in month two, that is what the free agent blueprint is for.

Topics:ai agent reliabilityproduction ai agentsoriginal data

Frequently asked questions

What percentage of AI agent jobs fail in production?

On the host we surveyed, 6 of 47 scheduled jobs were in an error state — 13% of defined jobs, or 15% of the 40 that had ever executed. That is one machine read on one day, so treat it as a published denominator rather than an industry benchmark. The more useful figure is the breakdown: every one of those failures was a dependency, configuration or upstream problem, not a model problem.

Do AI agents fail because the model gets things wrong?

Not in the errors we recorded. Across the failing jobs on our own host, the error strings were revoked model entitlements, delivery channels that were not enabled, upstream HTTP 500s, and one host with no provider configured at all. None of them were the model reasoning badly. Model quality is the part teams worry about and the part that had not yet cost us a scheduled run.

What happens to a scheduled agent job when a model slug is revoked?

It stops, and the error names the replacement. We recorded it verbatim as `HTTP 404: This model is unavailable for free. The paid version is available now - use this slug instead` and as `AccessDenied.Unpurchased: Access to model denied`. Three jobs died that way without a line of our code changing. The mitigation is a fallback cascade — primary, secondary, and a self-hosted last resort — with the substitution recorded in the run log so quality cannot drift silently.

How do you detect an AI agent job that silently does nothing?

By alerting on absence rather than on errors. A job that has never run emits no error, so any monitor that watches run outcomes will report nothing wrong forever. We found seven of them by reading the scheduler definitions against last-run timestamps, which is a check that belongs in a recurring sweep rather than in a person's memory.

Is a 13% job failure rate bad for a production agent system?

It depends entirely on what the failing jobs do and whether anyone finds out. None of ours were in the path that ships work to a customer, and all of them surfaced in a daily health sweep. The number worth managing is not the failure rate itself but the time between a job failing and a human knowing — and whether the failure classes are ones you have a standing fix for.

Amit Kumar

Founder of I Am Agent Man. Builds and runs production AI agents on Hermes, OpenClaw, and MCP — self-hosted, model-agnostic, with persistent memory and hard cost ceilings.

Get your agent blueprint — free.

Tell us what eats your week. We’ll come back with a concrete plan for the first agent to build and what it’ll automate — before you spend anything.

Goes straight to hello@iamagentman.com — we read every message ourselves. Prefer to answer three questions instead?Build your blueprint.

Talk to the studio

One message · reply within one business day