
AI Chatbot Vendor Checklist: 12 Questions to Ask Before You Buy
October 5, 2026
A chatbot demo usually shows the easiest version of the buying decision: a clear question, a polished answer, and a smooth booking or support flow. It rarely shows what happens when a customer requests another person’s account details, a policy document becomes outdated, or the integration fails halfway through a transaction.
This AI chatbot vendor checklist gives you 12 questions to ask before you buy, along with the evidence that makes an answer credible. It covers ownership, security, model choice, retrieval-augmented generation, integrations, human handoff, analytics, compliance, maintenance, pricing, and your ability to leave.
Whether you are evaluating Chatbot360 for your business or comparing other providers, use the same scenarios and acceptance criteria across your shortlist. A compelling presentation should earn a vendor a place in your evaluation—not replace the evaluation.
Table of contents:
- Questions 1–3: Data, security, and compliance
- Ask for an architecture map, not just a feature list
- Questions 4–6: Models, knowledge, and answer reliability
- Questions 7–9: Workflows, human handoff, and analytics
- Questions 10–12: Maintenance, pricing, and exit options
- Turn the checklist into a buying decision
- Frequently Asked Questions
Questions 1–3: Data, security, and compliance
1. Who owns our data, and what rights do you retain?
Do not accept “your data is yours” as a complete answer. Ask the vendor to distinguish uploaded documents, customer messages, generated responses, conversation metadata, embeddings, and configuration. Ownership matters, but so do the contractual rights to use, retain, and disclose that information.
Specifically ask whether your information can be used to train or improve the vendor’s models or any upstream provider’s models. Clarify whether restrictions apply by default, require an opt-out, or depend on your subscription.
- Request contractual language covering data use and confidentiality.
- Identify retention periods for live systems, logs, and backups.
- Confirm who can request exports and deletion, including after cancellation.
- Ask which derived data is retained and whether it remains linked to your organization or customers.
A useful test is to request the deletion of one test conversation and ask the vendor to explain where copies may remain and for how long.
2. How do you prevent unauthorized access and unsafe actions?
Encryption is important, but it does not answer whether one customer can retrieve another customer’s order details. Evaluate tenant isolation, administrator permissions, authentication, audit logs, secrets management, and access controls on retrieved content.
For transactional bots, authorization must be enforced by the underlying application or integration—not by instructions telling the model to behave. A chatbot should not be able to issue a refund merely because a user persuaded it that approval was unnecessary.
Request a security overview, relevant independent assessment evidence, vulnerability-management practices, and incident-notification terms. Ask the team to demonstrate prompt-injection defenses and permission enforcement using documents and customer records with different access levels. A certificate can support due diligence; it cannot substitute for testing your actual workflow.
3. Can you support our specific compliance obligations?
A blanket claim of being “compliant” is too broad to evaluate. Start with your use case: what data enters the chatbot, whose information it contains, where those people are located, and whether regulated activities are involved.
Then request the relevant contractual and operational evidence. Depending on your situation, this may include a data processing agreement, subprocessor list, international transfer arrangements, regional hosting options, or a business associate agreement where applicable.
Ask whether data-location commitments cover inference, logs, backups, and support access—not only the primary database. Have your legal and security teams assess the actual documents. A vendor’s controls can support your compliance program, but they do not automatically make your deployment compliant.
Ask for an architecture map, not just a feature list
The prisms in the accompanying image offer a useful way to think about chatbot architecture: separate components can look like one polished system from the outside. The chat interface, retrieval service, model provider, integration layer, and analytics platform may each process different information under different conditions.
Ask shortlisted vendors to draw the path of a single customer message. Mark where identity is checked, documents are retrieved, information leaves the platform, actions are authorized, and logs are stored. Include third-party services and the people responsible for each boundary.
This diagram makes vague answers easier to challenge. If a provider promises regional data storage but sends prompts elsewhere for inference, the distinction should be visible before contract negotiations—not discovered during implementation.
Questions 4–6: Models, knowledge, and answer reliability
4. Which models can we use, and who controls model changes?
Ask which models are available for your plan, whether different tasks can use different models, and whether you can bring your own provider account. Model choice affects cost, latency, data handling, and the behavior of the chatbot.
Flexibility also creates maintenance obligations. Find out whether model versions can be pinned where supported, how deprecations are handled, and what happens during a provider outage. An automatic fallback may keep the service running while changing answer quality or data-processing arrangements.
If the Chatbot360 platform is on your shortlist, request the same model and change-control details you request from every other vendor. Do not infer architectural capabilities from the appearance of the chat interface.
5. How does your RAG system find and maintain trustworthy information?
Retrieval-augmented generation, or RAG, gives the model relevant source material when preparing an answer. It can improve grounding, but its usefulness depends on what gets indexed, how it is retrieved, and whether the user is allowed to see it.
Bring a small set of realistic documents to the evaluation: a current policy, an obsolete version, a table, and a restricted internal page. Ask the vendor to show how its system handles each one.
- Can answers cite the exact source passage rather than a generic homepage?
- How are conflicting documents and effective dates handled?
- How quickly do changes and deletions reach the retrieval index?
- Are document permissions enforced during retrieval?
- What happens when a connector stops syncing or a document cannot be parsed?
A manual upload process may suit a stable FAQ. It is a poor fit for policies that change frequently unless someone owns the update routine.
6. How do you test unsupported answers and unsafe behavior?
Ask for the vendor’s evaluation process, then run your own tests. Include ambiguous requests, missing information, hostile instructions, outdated policies, and questions outside the chatbot’s scope. The goal is not to trick the product for entertainment; it is to discover how it fails before customers do.
Look for appropriate clarification, refusal, or escalation when the available evidence is insufficient. A polished answer with a citation is still wrong if the cited passage does not support the claim.
Ask the vendor to demonstrate a question the chatbot should not answer, an action it should not take, and a situation it should hand over to a person.
Confirm whether you can keep a regression test set and rerun it after prompt, model, connector, or knowledge-base changes. Avoid relying on an unexplained confidence score as your only safety mechanism.
Questions 7–9: Workflows, human handoff, and analytics
7. What do your integrations actually allow us to do?
An integration logo may mean anything from a basic webhook to a maintained, bidirectional connection. Specify the task you need: reading order status, updating a contact, checking appointment availability, or creating a support case with custom fields.
Ask which operations are supported, which plan includes them, and who maintains the connection when the external system changes. Test rate limits, expired credentials, incomplete records, and downstream outages.
For actions that change business records, check validation, approval steps, and duplicate prevention. A retried request must not create two bookings or issue the same refund twice.
When assessing Chatbot360 for connected customer workflows, bring a task-level integration brief. “Works with our CRM” is less useful than “creates a lead with consent status, assigns an owner, and reports failed submissions.”
8. What happens when a customer needs a human?
Human handoff is a workflow, not a button. Ask what triggers escalation, where the conversation goes, and what the customer sees while waiting. Distinguish live transfer from ticket creation and callback requests.
The receiving agent should get enough context to continue without making the customer repeat everything. That may include the transcript, verified identity, issue category, attempted actions, and relevant source references. Any generated summary should remain distinguishable from the original conversation.
Test handoff outside business hours and when no agents are available. Confirm that the customer can explicitly ask for a person and that the bot does not keep answering after ownership has passed to an agent.
9. Can we measure outcomes rather than just activity?
Conversation volume and response speed describe activity. They do not prove that customers received correct answers or completed their tasks.
Ask vendors to define “resolved,” “contained,” and “deflected.” A session that ends without escalation might represent success, abandonment, or a frustrated customer switching channels.
Request access to conversation-level records and the ability to review outcomes by intent, channel, and knowledge source. Useful measures include verified task completion, escalation reasons, repeat contact, customer feedback, and the cost of a successfully completed task.
For lead generation, connect chatbot events to qualified leads or booked appointments where appropriate. Also check analytics retention, export access, and redaction controls so measurement does not become an unmanaged store of personal information.
Questions 10–12: Maintenance, pricing, and exit options
The operational view matters as much as the launch experience. Once the chatbot is live, your team needs visibility into failures, clear ownership of updates, and costs that can be reconciled against actual usage.

Ask to see the administrator’s daily workflow, not only the customer-facing widget. A system that looks simple to visitors can still require substantial work behind the scenes.
10. Who maintains the chatbot after launch?
Get a written responsibility split for content updates, prompt changes, connector repairs, incident investigation, and quality reviews. “Managed service” means little unless the scope and response commitments are explicit.
Confirm whether your team has a staging environment, configuration history, approval controls, and rollback options. Ask how the vendor communicates releases that may change behavior and whether regression testing is included.
Support terms deserve separate scrutiny. Identify coverage hours, severity definitions, initial response commitments, and escalation contacts. A promise to acknowledge a critical issue is not a promise to restore service within the same period.
Assign an internal owner even if the provider handles most maintenance. Someone in your business must decide whether the chatbot’s policies and answers are still correct.
11. What will the complete operating cost include?
Ask vendors to price the same workload assumptions, including normal activity, seasonal peaks, and unusually long conversations. Understand the billing unit: messages, conversations, tokens, resolutions, seats, actions, or a combination.
For a fair assessment of Chatbot360’s fit with your operating budget, request a written quote with inclusions, limits, and overage rules. Apply the same requirement to every shortlisted provider.
- Separate implementation and migration fees from recurring charges.
- Check model usage, retrieval storage, connectors, and channel fees.
- Identify charges for analytics, security features, support, and additional environments.
- Include internal administration time and the human work left after automation.
- Ask about usage alerts, spending caps, and service behavior when a limit is reached.
Clarify whether failed actions, retries, spam, and test traffic are billable. These details can matter more than a low advertised starting price.
12. How can we leave without losing essential assets?
Discuss exit options before signing, while you still have negotiating leverage. Ask which documents, transcripts, metadata, prompts, configurations, and evaluation sets can be exported, in what formats, and at what cost.
Not every asset will be portable. Proprietary workflows may require rebuilding, and a retrieval index may need to be recreated in the replacement system. Knowing these limits is more useful than a vague promise of “no lock-in.”
Review cancellation notice periods, renewal terms, post-termination access, migration assistance, and deletion schedules. Request a sample export during the pilot and check whether it is usable outside the vendor’s interface. An export button is not enough if the resulting files omit the fields your business needs.
Turn the checklist into a buying decision
Send the 12 questions before the sales call and ask vendors to attach evidence. Separate what is available now from roadmap commitments, custom development, and features restricted to another plan. Record contractual commitments separately from informal sales answers.
Start with hard gates. If a vendor cannot meet an essential access-control requirement or process your data under acceptable terms, attractive pricing should not compensate for that gap.
Then run a bounded pilot using representative tasks and sanitized or appropriately authorized data. Agree on what success looks like before testing.
| Pilot scenario | Evidence to capture | Acceptance condition |
|---|---|---|
| A policy answer changes after a source update | Source version, sync status, answer, and citation | The answer reflects the approved update within the agreed refresh window |
| A customer requests restricted information | Identity context, retrieval behavior, and audit record | Unauthorized content is neither retrieved for that user nor disclosed |
| An external action times out | Integration logs, retry behavior, and customer message | No duplicate action occurs, and failure is communicated accurately |
| A conversation requires an agent | Routing destination, transferred context, and customer experience | The case reaches the correct destination with a usable continuation path |
| Your team exports the pilot data | Export files, field definitions, and associated charges | Required records are readable and usable outside the platform |
Keep commercial and technical findings together. A missing capability might be acceptable if it is unnecessary for your use case; an undocumented workaround for a critical control usually is not. Evaluate Chatbot360 against your pilot acceptance criteria rather than treating any product page as proof of implementation fit.
Frequently Asked Questions
Should we buy a chatbot before cleaning up our knowledge base?
You can evaluate platforms while improving your content, but avoid a broad rollout if essential policies conflict or lack owners. Start with a narrow set of approved sources. A pilot should reveal content gaps, yet no retrieval system can reliably resolve business rules your organization has not settled.
How long should an AI chatbot pilot run?
Base the duration on coverage rather than a fixed calendar target. The pilot should exercise your main intents, a source update, a failed integration, an escalation, and an administrative change. If seasonal workflows matter, simulate them instead of assuming ordinary traffic represents your busiest period.
Does a smaller business need enterprise security features?
Requirements should follow the sensitivity of the data and actions, not company size alone. A small business handling sensitive records may need stronger controls than a large organization publishing a public FAQ bot. Prioritize least-privilege access, appropriate authentication, retention controls, and incident handling before paying for features you do not need.
Is choosing the most powerful model enough to improve chatbot quality?
No. A capable model can still answer from obsolete content, retrieve the wrong document, or act through a poorly designed integration. Evaluate the complete system on your tasks. A more expensive model may help with complex reasoning, but it will not fix missing permissions or an unclear refund policy.
What if a vendor cannot share its full security reports?
Some providers restrict detailed reports to qualified buyers under a nondisclosure agreement. Ask what alternative evidence is available, such as an executive summary, a scoped assurance report, a completed security questionnaire, or a review with your security team. Persistent refusal to provide meaningful evidence is different from a controlled disclosure process.
Can we launch without allowing the chatbot to take actions?
Yes. A read-only deployment can be a sensible first stage, especially while you validate knowledge quality and escalation. Keep its limitations clear to customers. Add transactional capabilities only after authorization, validation, duplicate prevention, and failure handling have been tested.
If you want help turning this checklist into a practical shortlist and pilot brief, discuss your chatbot requirements with the MarketingV8 team. Bring your intended workflows, data constraints, and current support tools so the conversation starts with what your business actually needs.