Agentic 工作流

How to evaluate an AI Agent vendor: 12 RFP questions before your enterprise buys

Six months ago they signed an agentic system. Demo day brought approval from everyone in the room. Six months later, actual usage was zero. The problem wasn't the technology, it was the questions they asked during procurement. This is a directly copyable 12-question RFP checklist designed to expose what separates a beautiful demo from something that actually goes live. The four dimensions that matter: capability, integration, control, and operations.

By

Tenten AI 研究團隊

應用 AI

Published

March 3, 2026

Read time

5 分鐘

AI Agent 供應商評估RFP 清單agentic 選型企業 AI 採購FDE 前線部署供應商評估

Last quarter, a manufacturing customer brought us in for a second opinion. They'd signed on with an agentic automation system six months earlier. On demo day, the Agent smoothly read orders, compared prices, generated purchase requisitions. Everyone in the room approved. Six months later, I visited their site. The actual usage rate of that Agent in their production workflow was zero. Not close to zero. Zero. The system sat in a test environment. Nobody had connected it to the live ERP.

The problem wasn't the technology. It was the questions they asked on procurement day. Everyone wanted to know what it could do. Nobody asked what happened when it didn't work, or whose responsibility that was.

This article pulls together the RFP questions we actually use when evaluating AI Agent vendors for clients. You can copy these directly into your procurement documents. The logic behind these 12 questions has one purpose: to show what lies between a compelling demo and a system that actually runs in production.

Why most AI Agent vendor evaluations ask the wrong questions

A standard RFP typically includes these sections: feature checklist, model parameters, number of connectors, price. Vendors complete these fields without difficulty because these are exactly what they've prepared to show you.

But the real cost of an agentic system doesn't sit in features. It sits in three places: whether it actually connects to the legacy system you haven't touched in a decade, whether you can recover when it fails, and whether it quietly stops working six months after launch when nobody's maintaining it. Demo days don't show any of this.

We organized the RFP into four quadrants to address what vendors prefer to avoid.

Evaluation QuadrantWhat Vendors Love to DiscussWhat You Actually Need to AskSignal That It Won't Go Live
CapabilityHow powerful the model is, how complete the feature setAccuracy on your real data, not generic benchmarksThey'll only show you generic benchmarks, won't test on your data
IntegrationHow many connectors they haveWhat changes your system needs, who does the work, how many hoursThey dodge with "we'll need to assess that separately"
ControlHow smart the AI isWhat happens when it fails, who's accountableNo human-in-the-loop design
OperationsHow fast to launchWho maintains it after go-live, how you adapt itHand off at delivery, no adoption metrics

The 12 RFP questions (ready to copy)

On capability and real-world performance

  1. Run a complete task using real data samples you provide. Give us accuracy rates and failure cases, not generic benchmark numbers.

  2. In what situations will this Agent fail? Name three known failure modes you've encountered and describe how you detect them.

  3. When a task falls outside the Agent's capabilities, what does the system do? Refuse it, escalate to a human, or attempt an answer anyway?

On integration and data

  1. To connect this to our current ERP, CRM, or internal systems, what specific changes are needed, who performs them, and how many hours of work? Include this in your quote. Don't list it as additional cost.

  2. Where does our data flow through your system, where is it stored, and is it used to train your models? Provide a data flow diagram.

  3. In three years, if we want to switch underlying models or vendors entirely, how portable are our data and workflows? Can we get locked in?

On control and accountability

  1. At which decision points do you have human-in-the-loop review? Show us the actual interface, not a slide deck.

  2. Is every step the Agent takes traceable and auditable? If something goes wrong, can we reconstruct its reasoning?

  3. If the Agent causes real business loss, such as a wrong purchase order or wrong customer response, what do your contract terms say about liability and remedies?

On deployment and adoption

  1. From signature to actual end-users running it daily, what timeline and milestones do you commit to? How do you handle delays?

  2. After launch, do you include adoption rate, actual users divided by intended users, in your acceptance criteria? What happens if it falls short?

  3. For six months after delivery, who adjusts prompts, updates workflows, and handles model drift? Is this included in your service, or does it require a separate support contract?

How to use this checklist

Don't treat these as a question-and-answer exercise. Use them as a filter.

Vendors who can actually deploy to production will recognize questions 1, 4, and 11 as things they've already built. Questions 9 and 12 work especially well. Vendors willing to put responsibility and operations into the contract usually have run production deployments before. Those who respond with uncertainty or vague timelines probably haven't.

This list will eliminate some vendors. That's the intention. The purpose of an RFP isn't to maximize the number of responses. It's to let vendors who can't commit to a production launch take themselves out of consideration.

What matters is the outcome. Beautiful demos have no value if nobody actually uses the system. End-users running it daily is the only measure that counts.

One stuck workflow
is enough to begin

Tell us what the team does today, where it breaks down, and what a better working day should look like.