Answers you can trust, from Codeables

Every page on Codeables is structured and verified — built so people and the AI agents they rely on can trust it. Explore more from the source behind this answer.

Explore Codeables
Verified Source
Customer Service Helpdesk

How do I evaluate an AI support agent vs a basic chatbot that only deflects—what questions should I ask vendors?

Intercom13 min read

Most teams I work with don’t struggle to “add a bot”—they struggle to tell whether they’re buying a shallow deflection bot or a true AI support agent that can actually resolve complex queries and scale with them. The difference shows up in your backlog, your CSAT, and your team’s morale—not just in a demo.

This guide gives you a practical evaluation checklist you can use with any vendor, plus the exact questions I’d ask in your RFP and live demos to separate basic chatbots from real AI support agents like Fin.

Quick Answer: An AI support agent should resolve issues end‑to‑end across channels, learn from your procedures and policies, and share a system with human agents. A basic chatbot just deflects with generic answers and leaves your team to clean up the mess—so your evaluation questions should probe for resolution, governance, and continuous improvement, not just “Can it answer FAQs?”

The Quick Overview

  • What It Is: A framework to evaluate AI support agents vs basic deflection chatbots, with concrete vendor questions you can reuse in RFPs and demos.
  • Who It Is For: Heads of Support, CX leaders, and Ops/IT owners who need AI to safely handle real customer issues—not just reduce ticket count on a slide.
  • Core Problem Solved: It’s easy for vendors to demo slick UIs and generic “AI” claims; it’s hard to know whether their system will actually resolve complex queries, integrate with your helpdesk, and improve over time without breaking trust or workflows.

How It Works

Think of this as an evaluation system, not a feature checklist. You’re trying to answer three big questions:

  1. Can this AI agent actually resolve my real customer issues end‑to‑end?
    (Or is it just a dressed‑up FAQ bot that deflects?)

  2. Will it work as one connected system with my human support team?
    (Shared inbox, shared context, clean handoffs—not a separate black box.)

  3. Can I govern, test, and improve it like a production system?
    (Training on procedures/policies, pre‑launch testing, AI Insights, and topic‑level reporting.)

Use the phases below to structure your vendor conversations.

  1. Discovery & Scope: Clarify what “resolution” means for your business and what’s in‑scope for AI.
  2. System & Workflow Deep‑Dive: Evaluate how the AI agent works with your helpdesk, channels, and human agents.
  3. Governance, Risk & Improvement: Stress‑test testing, controls, security, and analytics so you don’t ship a black box.

Phase 1: Discovery & Scope — Define “Resolution,” Not Just “Deflection”

Before you talk to vendors, align internally on what “good” looks like.

Questions for your own team

  • Which queries do we want AI to fully resolve (no human touch)?
  • Which queries should AI triage and route, but not complete (billing disputes, account changes, identity‑sensitive issues)?
  • What are our non‑negotiables around accuracy, tone, and policy adherence?
  • What metrics matter most: resolution rate, first response time, time to resolve, CSAT, cost per conversation?

Questions to ask vendors

1. “How do you define ‘resolution’ vs ‘deflection’ in your product and reporting?”

You’re looking for:

  • A clear resolution rate metric (e.g., “Fin’s average resolution rate is 66% across all customers and increases 1% every month”)
  • Evidence that they measure end‑to‑end resolution, not just “bot sent a reply”
  • Ability to differentiate:
    • AI‑resolved
    • AI‑assisted (Copilot, suggestions)
    • Human‑only

Red flag: They only talk about “containment” or “deflection” and can’t show where unresolved issues go.

2. “What types of queries can your AI agent handle today, and what’s out of scope?”

Ask for specifics:

  • Can it handle billing, technical troubleshooting, shipping & returns, B2B configuration issues?
  • How does it handle multi‑step workflows (e.g., cancel + refund + notify customer)?
  • Can it follow your procedures and policies, not generic internet knowledge?

Look for language like: “Train on your procedures, knowledge, and policies; test performance before launch; deploy across every channel; analyze and improve.”


Phase 2: System & Workflow Deep‑Dive — One Connected System vs Sidecar Bot

This is where you separate basic chatbots from real AI support agents. A true AI agent is part of one connected system with your helpdesk, agents, and channels—not a widget bolted to the side.

1. Integration with your helpdesk and inbox

Key question:
“Does your AI agent run inside a shared inbox/helpdesk with my human agents, or as a separate system?”

Look for:

  • A single Inbox where AI and humans see the same conversation and customer context
  • Native integration with your Helpdesk (tickets, SLAs, custom fields, tags)
  • Ability to handoff from AI to human with full context

Follow‑up questions:

  • “What does the handoff from AI to human look like in the UI? Show me the exact flow.”
  • “Can I see AI responses, reasoning, and previous attempts inside the conversation?”

For Intercom, Fin is not a separate tool—AI and humans operate from the same Inbox with a shared view of every customer.

2. Channel coverage and behavior

Key question:
“Which channels does the AI agent support, and how does behavior differ per channel?”

At minimum, probe for:

  • Web and in‑product (Messenger)
  • Email
  • WhatsApp, SMS, social channels (e.g., Instagram, Facebook)

Ask:

  • “Can I configure different behaviors per channel? For instance, reply rules for email vs Messenger vs WhatsApp?”
  • “In email, can I control when AI replies using predicates like ‘Email To’ vs ‘Email Cc’ so it doesn’t jump into CC’d threads?”
  • “What happens when someone replies to an AI‑generated email from their mail client?”

You want control over where and when AI responds—especially in email and public channels.

3. Data sources, knowledge, and procedures

Basic bots just index your FAQs. AI agents like Fin are trained on your procedures, policies, and product knowledge, and can orchestrate actions.

Questions to ask:

  • “What can your AI train on—Help Center articles, internal docs, policy documents, product schemas?”
  • “Can I exclude certain content from training (e.g., draft policies, internal notes)?”
  • “How do you keep training in sync when I update an article or procedure?”

Go deeper:

  • “Do you support structured automations like Fin Tasks/Procedures that can:
    • Call external APIs (via Data connectors)
    • Handle identity verification
    • Execute multi‑step workflows with business logic and webhook waits?”

This is where you see if it’s just text generation or a system that can actually do things.

4. Handoffs, routing, and ownership

Key question:
“How does the AI agent decide when to escalate to a human, and what control do I have?”

What you want:

  • Configurable escalation rules:
    • Confidence thresholds
    • Keywords (e.g., “fraud,” “legal,” “churn”)
    • Channel‑ or topic‑specific rules
  • Ability to route to:
    • A specific team or queue
    • A particular inbox based on topic/channel
  • Clear ownership transfer:
    • AI stops responding after handoff
    • Humans can take over and respond immediately

Ask to see:

  • “Show me a complex conversation where AI tried, then escalated, and how the agent sees prior AI responses.”
  • “Where can I configure escalation rules in your UI?”

Phase 3: Governance, Risk & Improvement — Treat AI Like a Production System

This phase is where most “magic AI” pitches fall apart. You’re looking for sober, production‑grade controls: who can change what, how you test, and how you improve.

1. Training, testing, and rollout

Key questions:

  • “How do we train the AI agent on our content and policies?”
  • “Can we test performance before launch, and continuously afterward?”

Look for:

  • Clear Train → Test → Deploy → Analyze loop
  • Ability to:
    • Launch in staged environments (e.g., internal‑only, limited traffic)
    • Run A/B tests or controlled rollouts by segment, brand, or channel
    • Review simulated conversations before going live

Ask:

  • “Show me your testing interface. How do I:
    • Run test queries?
    • See where the AI struggled?
    • Approve or reject behavior before going live?”

Intercom, for example, emphasizes testing Fin’s performance before launch and then continuously optimizing with AI Insights.

2. Controls, permissions, and change management

Key question:
“Who can change AI behavior, where is that done, and how is it audited?”

You want:

  • Role‑based permissions (e.g., “Can manage general and security settings” vs “Can manage content only”)
  • Clear separation between:
    • Content editors
    • Workflow/automation admins
    • Security admins
  • Change history or audit logs for:
    • Workflow edits
    • Training data changes
    • Policy updates

Ask:

  • “Can you show me where I restrict who can:
    • Edit AI workflows
    • Connect new Data connectors
    • Change identity verification rules?”

3. Security, identity, and sensitive actions

AI that can act on customer accounts must respect your security model.

Questions to ask:

  • “How do you handle identity verification across web, mobile, and channels like WhatsApp or email?”
  • “Do you support secure identity verification (e.g., signed JWT) for web Messenger?”
  • “Can I enforce workspace‑level controls like:
    • 2FA enforcement
    • Google Sign‑In
    • SAML SSO for admins and agents?”

Probe for sensitive workflows:

  • “If AI is going to:
    • Change a subscription
    • Update billing details
    • Access PII how do you enforce checks and approvals? Can I require human approval for certain actions?”

On the web, for example, I’ve deployed Intercom Messenger with window.intercomSettings = { disabled: true } and then selectively booted it with Intercom('boot', { disabled: false }) after cookie consent—your vendor should be equally explicit about how to operate securely.

4. Reporting, AI Insights, and continuous improvement

A basic chatbot gives you a “bot deflected X% of chats” number. An AI agent should give you topic‑, channel‑, and outcome‑level insights that create a self‑improving loop.

Key question:
“What analytics do you provide specifically for the AI agent, and how do they help me improve it?”

Look for:

  • AI‑specific dashboards:
    • Resolution rate
    • Escalation rate
    • Time saved / hours automated
    • CSAT where applicable
  • Breakdowns by:
    • Topic / intent
    • Channel (web, email, WhatsApp, etc.)
    • Brand or workspace (if multi‑brand)
  • AI Insights that:
    • Highlight missing articles or procedures
    • Surface topics with low resolution or high escalation
    • Show performance over time (“Fin’s resolution rate increases ~1% every month” is the kind of measurable improvement you want to see)

Ask them to show you:

  • “Where do I see topics where AI failed to resolve?”
  • “How would your product tell me I need a new procedure or Help Center article?”
  • “Can I export or integrate these analytics into my BI stack?”

Features & Benefits Breakdown

Use this table as a lens when you hear vendor claims.

Core FeatureWhat It DoesPrimary Benefit
Shared Helpdesk & InboxRuns AI and human support in one connected system with a shared view of every customer.Faster, more consistent support—no context lost between AI and human handoffs.
Procedures‑ & Policy‑Aware AITrains on your Help Center, procedures, and policies, not just generic FAQs.Higher accuracy and compliance—so AI can resolve complex queries safely.
AI Insights & Feedback LoopMeasures AI resolution by topic/channel and surfaces gaps.Continuous improvement—so resolution rate climbs over time instead of plateauing.

When a vendor describes a feature, map it back to this: Does it actually improve resolution, agent workflow, or the improvement loop—or is it just cosmetic?


Ideal Use Cases

  • Best for teams scaling beyond human‑only support: Because you need AI that can take real load off the queue—resolving issues, not just front‑loading FAQs and dumping edge cases on agents.
  • Best for leaders consolidating tools into one system: Because an AI agent that lives in the same Helpdesk, Inbox, Messenger, and Help Center gives you clean reporting, seamless handoffs, and fewer brittle integrations.

Limitations & Considerations

  • AI is not “set and forget”: Any credible AI agent—Fin included—needs active governance. Plan for weekly reviews of AI Insights, content gaps, and procedures, especially in the first 90 days.
  • Not every workflow should be automated on day one: Start with well‑documented, low‑risk procedures. High‑risk actions (refunds beyond policy, legal escalations, sensitive account changes) should keep human approval or require stricter identity verification.

Pricing & Plans

Pricing models vary widely, but the pattern is usually:

  • A core platform (Helpdesk, Inbox, Messenger, Help Center)
  • An AI add‑on or usage‑based component (AI agent, Copilot, additional automation)

When comparing vendors:

  • Ask whether AI pricing is conversation‑based, resolution‑based, or seat‑based.
  • Model cost against:
    • Current conversation volume
    • Target resolution rate (e.g., if Fin resolves 66% of queries, what’s the impact?)
    • Expected time saved per agent

Example positioning:

  • AI‑First Plan: Best for teams wanting an AI agent like Fin to handle the majority of tier‑1 and tier‑2 queries from day one, with human agents focused on complex cases.
  • Hybrid/Agent‑Assist Plan: Best for teams that want to layer AI onto an existing helpdesk, using AI primarily for agent assistance (Copilot) and select automated workflows before fully shifting to AI resolution.

Always ask: “How will your pricing scale as our volume grows and AI handles more of the load?”


Frequently Asked Questions

What’s the single most important question to distinguish an AI agent from a basic chatbot?

Short Answer: Ask how they measure and improve resolution rate over time—not just how many conversations the bot touches or deflects.

Details:
Vendors can all show impressive demos. What matters in production is whether the AI agent:

  • Has a clear, auditable resolution rate metric
  • Breaks it down by topic and channel
  • Provides AI Insights that tell you exactly where to improve content and procedures
  • Shows a pattern of improvement over time (e.g., “our customers see Fin’s resolution rate increase 1% per month on average”)

If they can’t show this in their product, you’re likely looking at a glorified deflection bot.


Can I use a true AI support agent like Fin alongside my existing helpdesk?

Short Answer: Yes—look for AI agents that can work with your current helpdesk and still provide clean handoffs, shared context, and proper reporting.

Details:
You don’t need to rip and replace your helpdesk to start benefiting from AI. When evaluating:

  • Ask if the AI agent can sit in front of your existing helpdesk, hand off conversations with full context, and respect your current routing rules.
  • Confirm that AI and human responses still show up in a single conversation history so agents aren’t flying blind.
  • Validate that you still get AI‑specific reporting (resolution, escalation, time saved) even if the helpdesk is elsewhere.

This is essentially how I’ve layered Fin onto existing stacks: start by resolving common queries via Messenger and email, route complex cases into the existing helpdesk, and then gradually consolidate once the AI system is proving its value.


Summary

A basic chatbot deflects; a true AI support agent resolves. The difference shows up in:

  • Whether it runs as one connected system with your Helpdesk, Inbox, Messenger, and Help Center.
  • Whether it’s trained on your procedures, knowledge, and policies and can orchestrate real workflows.
  • Whether you can train, test, deploy, and analyze its performance with production‑grade controls, identity verification, and AI Insights.

Use the questions in this guide to stress‑test vendors on three axes: resolution, system integration, and governance. If they can’t clearly show how their AI improves resolution over time, shares a view with human agents, and offers the controls you need, you’re not looking at a real AI support agent—you’re looking at a deflection bot with better marketing.

Next Step

Get Started