Why we publish the schedule instead of an uptime percentage
Every agent vendor will tell you their system is reliable. Almost none will tell you how many scheduled jobs they run, how many were working the last time someone looked, or how they find out when one stops. An uptime percentage is easy to quote and answers none of that — it usually describes whether a web process was accepting connections, which is not the same thing as the work getting done.
So this write-up is the schedule itself. On we read our own production host over SSH, read-only, and wrote down every scheduled job, its cadence, its owner, and the status of its last run. Nothing was tidied first.
AI agent scheduled jobs reliability, as a tally
Forty-seven jobs were defined. Thirty-four had succeeded on their last run, six were in an error state, and seven had never executed at all. That last group is the one worth sitting with: a job that has never run produces no error, so it looks identical to a healthy job in any monitor built on run outcomes.
| Cadence | Jobs | What lives here |
|---|---|---|
| Daily | 12 | Work on things that accumulate overnight: the morning feed, a trends pull, the agency digest, the health sweep, lead collection |
| Weekly | 9 | Audits and sweeps that would be noise if run daily: the technical SEO audit, the retention sweep that asks whether each output track still earns its place |
| Interval loops | 8 | Continuous watchers that move an artifact between pipeline stages and should not wait for a clock |
| Twice daily + once | Remainder | Engagement jobs that need two passes a day, plus a single one-shot |
Twenty-seven of the 47 are not scheduled independently at all. They are dispatched by one orchestrator agent into the agent that owns the work — lead collection, the daily QA sweep, the publisher check, the weekly SEO audit, the retention sweep, the health sweep, and two pipeline watchdogs. That concentration is deliberate, and it is a trade: one place to change the cadence of most of the fleet, one place whose failure is felt widely. Thearchitecture write-up covers how the agents underneath it are separated.
How a failure surfaces when nobody is looking
Unattended work needs someone to notice, and the design decision that matters is that the noticer is not allowed to fix anything.
- Two interval watchdogs. One watches the trust pipeline; one watches for verified work waiting to be published. Both exist to catch an artifact that stalled between stages — the failure mode where nothing errored and nothing moved.
- A daily health sweep. An agent reads what every other agent is doing and reports fleet state. Its mandate forbids it from fixing what it finds, so a monitoring agent can never quietly resolve the thing you needed to see.
- A weekly retention sweep. A second non-acting agent audits whether each output track still earns its keep, and recommends only.
Those three roles are why the 34/6/7 tally existed to be read at all. A fleet with no reporting layer does not have a better record — it has an unknown one.
What zero restarts does and does not prove
Fourteen long-lived services — four agent gateways, a dashboard, and nine hosted applications — reported zero restarts, with the oldest continuously active since 13 July 2026 and the host itself up nine weeks and three days. There is no container runtime and no orchestration platform on this machine; the services are native systemd units behind a Caddy reverse proxy, with the dashboard reachable only over the private tailnet.
What that proves is bounded: no process crashed, no supervisor thrashed, nothing leaked badly enough to be killed. What it does not prove is that the work succeeded. On the same day, six scheduled jobs were failing. Process health and work health are different measurements, and quoting the first while the second is unmentioned is the most common way agent reliability gets oversold. The failure classes behind those six — and why none of them was the model being wrong — are broken down inthe failure rate write-up.
What we would carry into your build
- Alert on absence, not only on errors. Compare the age of each job's last success against its own declared cadence. Seven never-run jobs is what happens without it.
- Separate the watcher from the fixer. A reporting agent that can also remediate will eventually hide the thing it remediated.
- Treat "produced output, delivered nowhere" as a failure. It is invisible in model-level telemetry and worthless to the business.
- Publish the denominator internally. A reliability claim with no job count behind it cannot be checked by the person who inherits the system.
What this case study does not prove
One host, read once, on 27 July 2026. It is not an availability guarantee and not an industry benchmark. Nothing here measures whether the agents' output was substantively correct — a job that ran and produced a mediocre result counts as ok in this tally.
There are no cost figures, because this host carries no per-task cost ledger, and no client outcomes, because there is no client-side measurement to draw on and client work is under NDA. Those are different claims and we are not going to blur them. If you want the timeline side of the question rather than the operating side,the production timeline breakdown covers how long it takes to get a system to the point where it has a schedule worth reading.