Case studies

AI agent case studies with architecture details

Five production agent systems, documented the way an engineer would document them: the pipeline, the gates, the schemas, the schedule, the jobs that failed and why. All first-party, all read off the host.

Get my free blueprint

AI agent case studies with architecture details

These five case studies document AI agent systems this studio built and runs on its own production host, each written up with its architecture, its governance rules and its observed failures. Every figure was read directly off that host on 27 July 2026. No client systems appear here, because client work is under NDA.

No cost figures and no ROI percentages appear in any of them: the surveyed host carries no cost ledger and there is no client-side outcome measurement to draw on. These are capability studies, not return-on-investment claims.

We are a new studio, so we are not going to show you a wall of client logos we do not have. What we can show you is the agent infrastructure we built for ourselves and run in production every day — because the systems we run for clients are built the same way, from the same patterns, with the same guardrails.

Every number on these pages was read directly off our own production host on. Nothing is estimated and nothing is rounded up. Where we cannot prove something, we say so instead of implying it.

What the architecture detail covers

Each write-up names the agents in the chain, the boundary between them, and the rule enforced at each boundary. The content pipeline write-up shows a publishing agent that refuses to ship anything the verification agent has not signed off. The research ledger write-up shows fabrication blocked at the data layer rather than in a prompt, with findings and claims kept in separate append-only files. Thelead engine write-up shows the fixed row schema and the deduplication key that let 5,231 collected rows be narrowed to 145 worth contacting.

The newest one goes at the operating layer rather than the architecture: the47-job schedule and the uptime record behind it lists every cadence, names the two watchdogs and the non-acting health agent that surface a stalled job, and publishes the status tally exactly as it read — 34 succeeding, 6 failing, 7 that had never run.

What none of them answer is how long your own build would take, because that depends on your systems rather than ours — the production timeline breakdown covers that separately.

What you will not find here

No cost claims. Our host carries no per-task cost ledger, so any dollar figure would be a guess dressed as data. When we have measured cost data, we will publish the methodology with it.

No client outcomes. Client engagements are covered by NDA. We do not publish client names, client metrics, or client architecture without written permission, and we are not going to anonymise a client into a "leading B2B SaaS company" to imply proof we cannot show you.

No ROI percentages. These are capability studies — how a system is built, governed, and kept running. They demonstrate engineering judgement, not return on investment. Those are different claims and we are not going to blur them.

Our own production host

A 19-Agent Operating System With Enforced Separation of Duties

The fleet that runs our own agency: a governed pipeline, not a swarm. 972 sessions in 68 days, zero service restarts.

specialist agents
19specialist agents
scheduled jobs
47scheduled jobs
sessions in 68 days
972sessions in 68 days
service restarts
0service restarts
Read the write-up

Our own production host

A Content Pipeline Where Nothing Publishes Without Passing QA

Four agents, one hard gate. 457 sessions on the drafting agent alone, and a publisher that refuses to post unverified work.

agents in the chain
4agents in the chain
drafting sessions
457drafting sessions
publish cadences per week
2publish cadences per week
documented runbooks
6documented runbooks
Read the write-up

Our own production host

A Research Agent That Cannot Fabricate

721 findings and 310 separately-ledgered claims, each with a source URL, a confidence score, and a verification outcome.

findings logged
721findings logged
claims ledgered
310claims ledgered
daily collection runs
2daily collection runs
entries carry a source
100%entries carry a source
Read the write-up

Our own production host

A Lead Engine That Built a 5,231-Row Pipeline Across Four Markets

5,231 rows across four market segments, narrowed to 145 contact-verified. The filtering is the product.

rows in master pipeline
5,231rows in master pipeline
market segments
4market segments
contact-verified
145contact-verified
collection cadence
dailycollection cadence
Read the write-up

Our own production host

A 47-Job Schedule and the Uptime Record Behind It

The schedule is the system. 47 jobs across three cadences, 27 of them dispatched by one orchestrator, and nine weeks without a restart.

scheduled jobs
47scheduled jobs
ok, error, never run
34/6/7ok, error, never run
uptime at survey
9w 3duptime at survey
restarts in 14 services
0restarts in 14 services
Read the write-up

Questions

Why are all your case studies about your own systems?

Because they are the only production systems whose numbers we can publish without a permission problem. Client engagements are under NDA, and we will not anonymise a client into a vague descriptor to imply proof we cannot actually show. The systems we run for clients use the same architecture and guardrails documented here.

Why is there no ROI or cost data in these case studies?

Our production host does not carry a per-task cost ledger, and we have no client-side outcome measurement to draw on. Publishing a dollar figure or a percentage saved would mean estimating and presenting it as data. These are capability studies: how the systems are architected, governed, and kept running.

How were these numbers verified?

They were read directly off our own production host on 27 July 2026 — systemd service state and restart counts, scheduled job definitions with their last-run status, session counts on disk, and line counts in the research and lead ledgers. The full survey is documented internally and every figure on these pages traces back to it.

Can you build something like this for us?

Yes — that is the business. The engagement starts with a free agent blueprint that scopes which workflow to automate first, which framework fits, and what the guardrails need to be. You see the plan before any invoice.

Want a system like this?

Tell us the workflow that eats your week. We'll scope the first agent and the guardrails it needs — free, before any invoice.

Goes straight to hello@iamagentman.com — we read every message ourselves. Prefer to answer three questions instead?Build your blueprint.

Talk to the studio

One message · reply within one business day