Free tool

AI agent vendor scorecard

Twelve questions that separate a team who has run agents in production from a team who has built demos. Four of them are deal-breakers. Score your shortlist — including us.

Get my free blueprint

Most AI agent RFPs still use a checklist designed for traditional software: does your product have feature X, are you SOC 2, what does it cost. Those questions generate checkbox answers, and every vendor passes.

The questions that actually differentiate are the ones a capable team can answer concretely and an evasive team cannot. "Do you evaluate your agents?" gets a yes from everyone. "What is your eval methodology under production traffic, and what did it catch before release?" gets a specific answer from a small number of vendors and a brochure answer from the rest.

That difference is the whole tool. Score each criterion on what they actually said, not what you hope they meant.

Evidence

Can they describe their eval methodology under real production traffic?deal-breaker

A full-marks answer: They name specific evaluations, what each one checks, the pass threshold, and a failure the evals actually caught before release.

Can they point to an agent of theirs running unattended today?deal-breaker

A full-marks answer: A named system, how long it has run, how often it runs, and who is on the hook when it breaks. Not a demo video.

Will they tell you how their agents have failed?deal-breaker

A full-marks answer: Specific failure classes with frequencies, and what they changed afterwards. A vendor with no failures has no production.

Architecture

Is the component that verifies output independent of the one that produces it?

A full-marks answer: A separate verification step that checks against a spec it reads itself, rather than the builder confirming its own work.

What happens when their model provider has an outage or revokes a model?

A full-marks answer: A named fallback cascade across providers, and evidence it has been exercised. Model access changes without warning.

How do they detect work that completed but was delivered nowhere?

A full-marks answer: Delivery treated as part of success, with alerting when a run finishes but the output never lands.

Control

What hard spend ceiling is enforced in code?deal-breaker

A full-marks answer: A per-task and per-day limit that stops execution, plus a circuit breaker on repeated failure. "We monitor it" is a 0.

Where does a human approve before the agent acts?

A full-marks answer: Named approval gates on specific high-risk actions, configurable by action type — not a blanket on/off switch.

What can you see after the agent has acted?

A full-marks answer: Per-action logs you can query yourself, including model used and output, retained long enough to investigate.

Commercial

Where does your data live, and is it ever used to train a model?deal-breaker

A full-marks answer: Self-hosted or your-infrastructure by default, contractual confirmation that your data is never training data.

Can they show the arithmetic behind their price?

A full-marks answer: A scoped fixed price with the drivers named. A range with no methodology is a range they invented.

What do you own, and can your team take it over?

A full-marks answer: You own the code and config, with handoff documentation and no runtime dependency on the vendor to keep it running.

Why four criteria are deal-breakers

A weighted average lets a vendor bury a fatal gap under a good average. Four of these cannot be compensated for, so scoring zero on any of them should stop the evaluation regardless of the total.

No eval methodology means nobody knows whether the agent works — including them. No production reference means you are the first, and the first agent a team ships is where they learn what they did not know. No failure stories is the strongest signal of all: every team running agents in production has watched them break, so a vendor with nothing to report either is not running anything or is not being straight with you. And no enforced spend ceiling means one loop can produce an invoice nobody approved.

Data residency is on the list for the same reason. If your data can become someone else's training set, no other strength on the scorecard matters to your security team.

How to run the scoring honestly

Score what the vendor said, not what you inferred. If you find yourself filling in the gaps on their behalf, that is a 1 at best. The most common way buyers get this wrong is rewarding confidence: a fluent answer with no specifics is still a zero.

Run it live on a call rather than sending it as a questionnaire. Written responses get polished by whoever writes the proposals; spoken answers reveal whether the person in front of you has actually operated the thing.

And score more than one vendor. A single scorecard tells you little in isolation — the value is in the spread between candidates on the same twelve questions.

Score us with it

We built a buyer-side tool knowing it would be used on us, so it would be strange not to invite that. Our answers to the four deal-breakers are on thecase studies: a documented failure taxonomy with frequencies, a fleet of agents running unattended with restart counts, enforced per-task and per-day spend ceilings, and self-hosted deployment by default.

If a vendor cannot show you the equivalent, that is the answer to your question.

Questions

What should I ask an AI agent development company before hiring them?

The questions that produce specific answers rather than brochure answers: what is your eval methodology under production traffic and what did it catch, which of your agents runs unattended today and for how long, how have your agents failed and what changed afterwards, and what hard spend ceiling is enforced in code. Those four separate teams who have operated agents from teams who have demoed them.

Why is "how have your agents failed?" a good interview question?

Because every team running agents in production has watched them break, so an inability to answer means either there is no production system or the vendor is not being straight with you. A strong answer names failure classes with rough frequencies and describes what changed as a result — model access being revoked mid-flight, work completing but never being delivered, a runaway loop caught by a spend ceiling.

What are the red flags when evaluating an agentic AI vendor?

No eval methodology, no reference to an agent running unattended, no failure history, and no spend ceiling enforced in code rather than monitored by a human. Any one of those is disqualifying on its own. Softer flags: verification performed by whatever produced the output, no named fallback when a model provider fails, and a price range with no methodology behind it.

Should I send this scorecard to vendors as a questionnaire?

Run it live on a call instead. Written responses get polished by whoever writes proposals, while spoken answers reveal whether the person in front of you has actually operated the system. Score what they said rather than what you inferred — a fluent answer with no specifics is still a zero.

Is this scorecard biased toward your own strengths?

It is built from the failure classes we have hit in our own production fleet, so yes, it reflects what we think matters. We publish our answers to all four deal-breakers in our case studies so you can score us on the same twelve questions. If a vendor cannot show you the equivalent evidence, that is useful information regardless of who wrote the scorecard.

Want us on your shortlist?

Score us with the same twelve questions. Then tell us the workflow and we'll scope the first agent — free, before any invoice.

Goes straight to hello@iamagentman.com — we read every message ourselves. Prefer to answer three questions instead?Build your blueprint.

Get my agent blueprint — free

No retainer to start · reply within 24 hours