
- Plan for 40–50% deflection in month one. Published benchmarks range from 8% to 85%, and the spread tracks documentation quality, not model choice.
- Deflection and resolution are different metrics: a customer who gives up counts as deflected. Ask any vendor which one they're quoting.
- Write your top 10–15 answers before you configure anything. Half a day there beats any model upgrade.
- Test adversarially: the questions you haven't documented are the set that decides whether you're ready to launch.
Why most deployments underperform
Most AI support deployments underperform for a reason that has nothing to do with the AI. The knowledge base is thin, so the bot has nothing to answer from, and everyone concludes the model is bad.
This guide is the sequence that avoids that. It assumes you're under 20 people, you don't have a dedicated support team, and nobody has a week to spend on setup.
First: the numbers you should actually expect
Vendor marketing will show you 80–90% deflection. Here's the wider picture from across published benchmarks, which disagree with each other enough to be worth showing honestly:
| Source type | Reported deflection/resolution |
|---|---|
| Median SaaS support teams | 40–45% |
| Top-quartile programs | 55–60% |
| Enterprise tier-1 median | ~41% |
| Bottom quartile | ~22% |
| Early deployments, realistic range | 30–50% |
| Mature workflows | 50–70% |
| Deeply integrated, action-taking agents | 70–85% |
| Some independent analyses, honest median | ~22% |
| Worst-instrumented deployments | ~8% |
The spread from 8% to 85% is not measurement noise. It reflects two things.
One: deflection and resolution are different metrics, often reported interchangeably. A platform can show 90% deflection (the conversation ended without escalation) while true resolution sits at 40%. Deflection measures containment. Resolution measures whether the person's problem got solved. A customer who gives up and leaves counts as deflected. Ask any vendor which one their number is.
Two: the variation is driven by knowledge base quality, not model choice. This is the consistent finding across essentially every source. Same models, same platforms, 25% to 80% outcomes depending entirely on what documentation the AI was given.
If a vendor promises 85% on your existing docs, they're quoting a number from a customer who invested heavily in content.
Step 1: Fix your documentation first (half a day)
Skip this and nothing downstream works. The AI answers from your docs; thin docs produce a thin bot.
Pull your last 50 support conversations. Sort by frequency. Take the top 10–15 questions and write a real answer for each.
What makes an answer usable by AI:
- Self-contained. Each answer should stand alone without requiring three other pages for context.
- Literal about specifics. Exact button labels, exact plan names, exact error text. "Go to the billing section" is worse than "Settings → Billing → Manage plan."
- Explicit about edge cases. "This doesn't apply to annual plans" prevents a confidently wrong answer.
- Written in customer vocabulary. If they say "how do I cancel," don't file it under "Subscription lifecycle management."
Include the unglamorous ones: refunds, cancellation, pricing edge cases, "is my data safe." They're a large share of real volume and teams routinely skip documenting them.
Half a day here moves your deflection rate more than any model upgrade will.
Step 2: Decide what the AI is not allowed to do
Before configuring anything, write the escalation list. Anything touching money, account deletion, security, legal commitments, or an already-angry customer goes to a human immediately.
The instruction that matters most in your system prompt is the one permitting the AI to fail:
If the answer isn't in the provided knowledge base, say you don't know and offer to connect the person to the team. Never guess. Never infer policy that isn't written down.
A bot that says "I'm not sure, let me get someone" is fine. A bot that invents a refund policy creates an obligation you have to honor or publicly walk back. Every hour of tuning should go toward making the AI honest before making it clever.
Step 3: Set a narrow scope
Don't launch sitewide. Pick the page where your most repetitive questions arrive (usually pricing or docs) and put it there only.
Narrow scope means the questions cluster, so your hit rate is high, and the failures are cheap and visible. Sitewide launches produce scattered questions the docs don't cover, and your first impression of your own product is bad.
Step 4: Test adversarially before launch (one hour)
Do not test by asking questions you know it can answer. Everyone does this, it proves nothing.
Run three sets:
- 1The top 10 real questions, phrased the way customers actually phrase them: typos, no punctuation, half a sentence.
- 2Ten things it should refuse. "Can you give me a discount?" "What's your competitor better at?" "Delete my account." Check it escalates rather than improvises.
- 3Five plausible questions you never documented. This is the important set. You're checking that it says "I don't know" instead of confabulating. If it invents an answer here, fix the prompt before launching, not after.
Ship when set 3 behaves.
Step 5: Launch with a visible exit
The single highest-value UI element is an obvious path to a human. Every conversation should have a visible "talk to the team" option, not buried after three failed attempts.
Counterintuitively this raises satisfaction even among people who never click it. Feeling trapped with a bot is most of what people hate about support bots.
Step 6: Read the failures weekly (30 minutes)
Sort last week's conversations into three buckets:
- Answered correctly: no action.
- Escalated honestly: working as designed. If a question repeats here, document it and it becomes bucket one next week.
- Answered wrongly: the only bucket that's actually urgent. Fix the underlying doc immediately.
This half-hour loop is what moves you from 40% to 60% over a couple of months. The compounding comes from documentation, not from switching models.
What this should cost
For a team under 20 with modest volume, expect roughly $30–$150/month for a website AI support tool. Two pricing structures dominate, and the difference matters more than the sticker price:
- Per-unit (per resolution, or per credit): cost scales with usage. Predictable per ticket, unpredictable per month. A traffic spike is also a bill spike.
- Flat monthly with a reply cap: predictable per month, and you know your worst case in advance.
For early-stage teams, a knowable ceiling is usually worth more than a marginally lower per-unit rate, because the failure mode of per-unit pricing is a surprise invoice in your best traffic month.
The common ways this fails
- Launching before writing docs. Guarantees a bad bot and a wrong conclusion about AI.
- Optimizing the model instead of the content. The gains are in the knowledge base. They're always in the knowledge base.
- Hiding the human handoff to inflate deflection numbers. Trades a metric for customer trust.
- Not reading transcripts. Deployments that stall are the ones nobody reviews.
- Trusting deflection as if it were resolution. Measure whether problems got solved, not whether conversations ended.
A realistic first week
| Day | Task | Time |
|---|---|---|
| 1 | Pull 50 conversations, rank questions | 1h |
| 1 | Write top 10–15 answers | 3h |
| 2 | Connect docs, set system prompt + escalation rules | 1h |
| 2 | Adversarial testing, all three sets | 1h |
| 3 | Launch on one page | 15m |
| 5 | Read every conversation, fix wrong answers | 1h |
| Weekly after | Triage failures, patch docs | 30m |
Under seven hours to a working deployment. Roughly four of those are writing documentation, which is the actual work, and which pays off whether or not you keep the AI.
Frequently asked questions
- What deflection rate should a startup expect from AI support?
- 40–50% in the first month with reasonable documentation. 55–60% puts you in top-quartile territory. Treat promises of 85% on day one as marketing.
- Is deflection rate the same as resolution rate?
- No, and conflating them is the most common way these numbers mislead. Deflection means the conversation ended without a human. Resolution means the problem was solved. A customer giving up counts as deflected.
- How much documentation do I need before AI support works?
- Your top 10–15 questions, answered self-containedly. That's usually enough to cover a majority of first-contact volume.
- Which AI model should I choose?
- It matters far less than your knowledge base. Every benchmark attributes the variation to content quality, not model choice. Pick a cheap capable default and spend the effort on docs.
- How much should AI customer support cost a small startup?
- Roughly $30–$150/month at modest volume. Watch whether pricing is per-unit or flat, because that determines your worst-case month more than the headline rate does.
Sources
One flat price, no surprise invoice
Helin answers from your own docs, hands off honestly when it doesn't know, and caps your bill at a number you picked. Free tier is permanent: 50 AI replies a month, no trial clock.
