Most AI agent RFPs still use a checklist designed for traditional software: does your product have feature X, are you SOC 2, what does it cost. Those questions generate checkbox answers, and every vendor passes.
The questions that actually differentiate are the ones a capable team can answer concretely and an evasive team cannot. "Do you evaluate your agents?" gets a yes from everyone. "What is your eval methodology under production traffic, and what did it catch before release?" gets a specific answer from a small number of vendors and a brochure answer from the rest.
That difference is the whole tool. Score each criterion on what they actually said, not what you hope they meant.
Why four criteria are deal-breakers
A weighted average lets a vendor bury a fatal gap under a good average. Four of these cannot be compensated for, so scoring zero on any of them should stop the evaluation regardless of the total.
No eval methodology means nobody knows whether the agent works — including them. No production reference means you are the first, and the first agent a team ships is where they learn what they did not know. No failure stories is the strongest signal of all: every team running agents in production has watched them break, so a vendor with nothing to report either is not running anything or is not being straight with you. And no enforced spend ceiling means one loop can produce an invoice nobody approved.
Data residency is on the list for the same reason. If your data can become someone else's training set, no other strength on the scorecard matters to your security team.
How to run the scoring honestly
Score what the vendor said, not what you inferred. If you find yourself filling in the gaps on their behalf, that is a 1 at best. The most common way buyers get this wrong is rewarding confidence: a fluent answer with no specifics is still a zero.
Run it live on a call rather than sending it as a questionnaire. Written responses get polished by whoever writes the proposals; spoken answers reveal whether the person in front of you has actually operated the thing.
And score more than one vendor. A single scorecard tells you little in isolation — the value is in the spread between candidates on the same twelve questions.
Score us with it
We built a buyer-side tool knowing it would be used on us, so it would be strange not to invite that. Our answers to the four deal-breakers are on thecase studies: a documented failure taxonomy with frequencies, a fleet of agents running unattended with restart counts, enforced per-task and per-day spend ceilings, and self-hosted deployment by default.
If a vendor cannot show you the equivalent, that is the answer to your question.