To evaluate an AI employee vendor, run an RFP that asks 40 specific questions across eight areas: scope and fit, integrations, security and identity, governance and audit, deployment and data residency, pricing, support and rollout, and references. The checklist below gives each question plus a one-line note on what a good answer looks like, so procurement and security reviewers can score responses consistently.
In case you didn't know, An AI employee is software that owns a defined job end to end inside your systems, such as reviewing purchase requests or resolving invoice exceptions. If you need the background first, read our complete guide to AI employees.
Why AI employee RFPs need different questions
Standard SaaS questionnaires assume software that waits for a person to click something. An AI employee acts on its own, at machine speed, across many cases. That changes what a security reviewer needs to know: who the agent acts as, what it is allowed to change, and whether you can reconstruct a single decision months later.
Gartner predicts over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The same release warns about "agent washing," where existing chatbots and RPA tools are simply rebranded as agents. The questions below are built to surface both problems before you sign a contract.
A practical tip: for every question, ask for an artifact as well as a written answer. A redacted decision log, a sample escalation record or an actual pricing scenario is much harder to overstate than a paragraph.
Who should review the answers
Split the checklist by dedicated owners so each section gets an expert reader:
- The process owner scores scope and fit, since they know which exceptions matter.
- Enterprise architecture or the IT team scores integrations.
- The CISO's team takes security and identity plus deployment and data residency.
- Internal audit or compliance takes governance and audit.
- Procurement owns pricing, rollout terms and reference calls.

Agree the scoring rubric before responses arrive, so no one is grading against a vendor's framing.
Scope and fit
- Which specific job will the AI employee own, and where does its authority end? A good answer names the job, the start and end of the process boundaries, and the typical decisions it hands to a person.
- What outcome metrics will you report, and how are they calculated? Look for completion rate, cycle time and correction or escalation rate, with definitions, rather than "tasks automated."
- How does the agent handle a case type it has never seen? The right answer is that it escalates with context and plausible options instead of guessing or silently failing.
- Who writes and maintains the agent's operating rules? Prefer a process owner editing plain-language rules over a dependency on vendor engineers for every change.
- Which comparable jobs is your platform running in production today? Expect named job types with volumes, not a list of possible use cases.
Integrations
- How does the agent connect to our ERP, CRM, ITSM and other core systems? A good answer lists native API connectors and states which of your systems are already supported.
- Can the agent operate systems that have no API? Strong vendors support browser-based action on legacy portals and explain how those sessions are logged.
- Do you support the Model Context Protocol or other standard connectors for internal tools? Look for a concrete answer on MCP or an equivalent, with examples.
- Where does the agent write its outputs and reasoning? The best answer is inside your existing system of record, not only in a separate vendor console.
- What happens when a connected system changes its UI or API? Expect monitoring that detects breakage, a defined fix process and escalation of affected cases in the meantime.
Security and identity
- Does each AI employee have its own identity, separate from human and shared service accounts? The answer should be yes, with a unique non-human identifier per agent.
- How are the agent's permissions scoped, and who approves them? Look for least-privilege access per role, reviewed by your system owners before go-live.
- Can you revoke one agent's access without affecting others? A good answer describes per-agent credentials that can be disabled in minutes.
- How do you defend against prompt injection in documents, emails and web pages the agent reads? Expect specific guardrails, input handling and testing, with reference to the OWASP Top 10 for LLM applications or similar.
- Is our data used to train models shared with other customers? The answer should be no by default, stated in the contract.
Identity is where most first-generation agent deployments fall short. See AI agent identity management and least-privilege agent identity for the underlying model.
Governance and audit
- For a single case, can you show us everything the agent saw, decided, attempted and changed? Ask for a redacted sample; a good one includes inputs, sources, action, result and approval state.
- How are approval gates and confidence thresholds configured? Look for per-role settings that your process owner can tighten or loosen.
- How are changes to the agent's rules tested and approved? Strong answers describe versioning, regression tests against known cases and review before release.
- Can we see the full change history of the agent's operating procedure? Expect who changed what, why, and which version was live on a given date.
- How does your platform support our regulatory obligations, such as SOX, GLBA or 21 CFR Part 11? A good answer maps specific platform controls to your obligations rather than claiming blanket compliance.
The NIST AI Risk Management Framework is a useful reference when scoring this section.
Deployment and data residency
- Which deployment models do you support: multi-tenant SaaS, your own cloud account, on-premises? The more options, the easier it is to match the deployment to each job's data sensitivity.
- Where is our data stored and processed, including logs and model inference? Expect named US regions and a clear answer for each data type.
- Which third-party model providers process our data, and under what terms? Look for a sub-processor list and contractual limits on retention and training.
- Can we bring our own encryption keys and keep our network controls? In a bring-your-own-cloud model the answer should be yes.
- Which independent attestations do you hold, and can we see the reports? Ask for a current SOC 2 Type II report and recent penetration test summary under NDA.
For how the three models compare, see on-prem, BYOC or SaaS deployment.
Pricing
- What is the billable unit? A good answer defines it precisely, whether seats, tasks, compute or a base fee plus usage.
- What will this job cost at our current volume and at twice that volume? Expect a written scenario for both, not a rate card alone.
- Are integration, implementation and change requests included? Look for a clear split between one-time and recurring costs.
- How do you charge for escalated or failed cases? Strong vendors explain whether you pay for work that ends up with a human.
- What usage reporting do we get, and can we set budget limits? Expect per-job usage dashboards and alerts before you exceed a threshold.
Pricing models vary widely across the category. The deployment and pricing guide explains the per-seat, per-task, usage-based and hybrid models.
Support and rollout
- What does the first 90 days look like, week by week? A good answer has named milestones from system access to live work.
- Will the pilot run on live data, and what are the success criteria? Look for a defined scope, agreed metrics and a clear go or no-go decision point.
- Who on your side is accountable after go-live? Expect a named team, response-time commitments and an escalation path.
- How do you handle an incident where the agent takes a wrong action? Ask for the runbook: detection, stopping the agent, reversal and root cause.
- How do we expand to new jobs after the first one? Strong answers show reuse of integrations and operating knowledge, not a fresh project each time.
References
- Can we speak to a customer running a similar job in production? Insist on production, not a pilot, and ideally in your industry.
- Can you share a case study with before-and-after metrics? Look for specific numbers with a time period and volume attached.
- What was the escalation rate at launch and what is it now? A credible vendor shares both and explains what changed.
- Have any customers stopped using the platform, and why? An honest answer is a good sign; a claim of zero churn deserves follow-up.
- Can we review a sample security questionnaire you completed for a regulated customer? Expect a redacted version that shows the depth of past reviews.
How to score the responses
Score each question from zero to two: zero for no answer or a marketing answer, one for a written answer without evidence, two for a written answer backed by an artifact. Weight security and identity, and governance and audit, more heavily if the job touches money movement, customer data or regulated records.
Run the same scenario through every shortlisted vendor. Give each one the same job description, the same sample volume and the same five tricky exception cases from your own backlog, then compare how each would handle them.

How Zamp answers these questions
Each AI employee runs on an Agent Operating Procedure that the process owner writes, with every change tested on a branch against evals and reviewed before merge. Each agent has its own least-privilege identity, and the decision audit trail logs input, sources, action, result and approval state for every case. Zamp supports SaaS, BYOC and on-premises deployment, and usage is metered in Agent Compute Units. The full security model is in the trust and security overview, and the biopharma procurement case study shows escalation and cost figures from a live deployment.