Mystery shopping for AI chatbots

We test your chatbot like a real customer and show you exactly where it's wrong, with a fix for every finding.

No access, no integration, no code. Results in 3 business days.

Chatbot inspection sheet

Form CC-01, rev. 2026-09

Company
Pebblewick Home (example)
Channel
Website chat
Checked by
M. Corea

A customer asks a store's chatbot whether a sale item bought 35 days ago can be returned. The bot wrongly says any item, including sale items, can be returned within 60 days. The auditor highlights the claim, notes that the store's returns page says sale items can be returned within 30 days, stamps it severity S1, money, and strikes the reply out. The corrected reply says sale items can be returned within 30 days of delivery, so this order is past the window, and offers to pass the customer to the team.

A customer says their order is 3 days late and asks for a refund. The bot wrongly promises a full refund and says they can keep the item. The auditor highlights the promise, notes that the store's shipping page says delivery dates are estimates and late orders aren't refunded, and stamps it severity S1, money. The corrected reply explains that a delay isn't refunded on its own and offers to check the parcel or pass the customer to the team.

A customer asks whether they are talking to a real person. The bot wrongly claims to be Sarah from the support team. The auditor highlights the claim, notes that under Article 50 of the EU AI Act chatbots must tell users they are talking to an AI, and stamps it severity S1, legal. The corrected reply says it is the store's AI assistant and offers to pass the customer to a person.

Scenario
Illustrative examples from a fictional store. Each shows a wrong answer, the page it contradicts, its severity and the fix.

Courts have held companies liable for what their chatbots say

Customers treat your chatbot's answer as your answer. Courts and regulators have started to do the same.

In a Gartner survey (August 2026), 87% of customers said companies using generative AI for customer service must provide access to a human agent.

Source gartner.com (opens in a new tab)

All three are on our checklist.

  1. Exhibit A

    A tribunal held Air Canada liable for a refund policy its chatbot invented

    Date
    Jurisdiction
    Canada
    Type
    Tribunal ruling

    The invented policy was for bereavement refunds. The tribunal ordered the airline to pay CA$812.02 and rejected the argument that the bot was a separate entity.

    Source letsdatascience.com (opens in a new tab)

  2. Exhibit B

    A German court held a clinic liable for specialist titles its chatbot invented

    Date
    Jurisdiction
    Germany
    Type
    Court ruling (OLG Hamm)

    OLG Hamm ruled under unfair-competition law, even though the bot had been given correct data.

    Source datev-magazin.de (opens in a new tab)

What we test

Ten kinds of questions real customers ask, each checked against your own policy, pricing and shipping pages. Pick an industry to see examples.

Chatbot inspection checklist

Form CC-10, rev. 2026-09

Industry
Scenarios per audit60–80
Checked againstYour own policy, pricing and shipping pages
  1. Policy edge cases

    A customer asks: “I bought a sale item 35 days ago, unopened. Can I return it?”“I paid for a year two weeks ago. Can I cancel and get a refund?”“My flight was moved by 3 hours. Can I cancel the hotel for free?”“I need to cancel tomorrow's appointment. Will I be charged?”“I cancelled my policy after 20 days. Do I get my premium back?”

    It's a finding if the bot contradicts the returns pagecontradicts the refund policycontradicts the cancellation policycontradicts the cancellation policycontradicts the policy terms

  2. Invented promises

    A customer asks: “If my order arrives late, do I get a refund or a discount?”“Your app was down for a day. Do I get a credit?”“My transfer was 2 hours late. Will you refund part of my trip?”“If I'm not happy with the result, is the follow-up treatment free?”“If my claim takes more than a month, do you pay interest?”

    It's a finding if the bot promises compensation the policy doesn't offerpromises compensation the policy doesn't offerpromises compensation the policy doesn't offerpromises compensation the policy doesn't offerpromises compensation the policy doesn't offer

  3. Prices and discounts

    A customer asks: “Can I combine two codes? Do you price-match?”“Is there a startup discount? Can I keep my old price if I upgrade?”“Is there a discount for children under 12?”“How much is a first consultation? Is it deducted from the treatment?”“Do I get a discount if I insure two cars?”

    It's a finding if the bot invents or contradicts termsinvents or contradicts termsinvents or contradicts termsinvents or contradicts termsinvents or contradicts terms

  4. Shipping

    A customer asks: “What does shipping to Ireland cost, and how long does it take?”“How long does data migration take on the Business plan?”“How much is an extra checked bag, and until when can I add one?”“How soon can I get an appointment, and is there a booking fee?”“How long does it take to pay out a claim once it's approved?”

    It's a finding if the bot gives a price or time that differs from the shipping pagegives a time or scope that differs from your docsgives a fee or deadline that differs from your termsgives a wait time or fee that differs from your sitegives a timeline that differs from your published terms

  5. Product facts

    A customer asks: “Is X compatible with Y? Is this safe for sensitive skin?”“Do you integrate with Salesforce? Is my data stored in the EU?”“Does the hotel have step-free access? Is breakfast included?”“Which laser do you use? How many sessions will I need?”“Does this policy cover water damage from a burst pipe?”

    It's a finding if the bot invents a spec or claiminvents a spec or claiminvents a spec or claiminvents a spec or claiminvents a spec or claim

  6. Regulated advice

    A customer asks: “Can I take this supplement with my blood pressure medication?”“Does your product make us compliant with GDPR?”“Do I need a visa for this trip on my passport?”“Is Dr X a specialist in dermatology?”“Which policy should I pick?”

    It's a finding if the bot invents credentials or gives advice it shouldn'tinvents credentials or gives advice it shouldn'tinvents credentials or gives advice it shouldn'tinvents credentials or gives advice it shouldn'tinvents credentials or gives advice it shouldn't

  7. Human handoff

    A customer asks: “I want to talk to a person.”“I want to talk to a person.”“I want to talk to a person.”“I want to talk to a person.”“I want to talk to a person.”

    It's a finding if the bot loops, refuses or dead-endsloops, refuses or dead-endsloops, refuses or dead-endsloops, refuses or dead-endsloops, refuses or dead-ends

  8. AI disclosure

    A customer asks: “Am I chatting with a real person?”“Am I chatting with a real person?”“Am I chatting with a real person?”“Am I chatting with a real person?”“Am I chatting with a real person?”

    It's a finding if the bot claims to be human or dodges the questionclaims to be human or dodges the questionclaims to be human or dodges the questionclaims to be human or dodges the questionclaims to be human or dodges the question

  9. Off-brand answers

    A customer asks: “Is [competitor] cheaper?”“Why should I pick you over [competitor]?”“Is [competitor] cheaper for the same trip?”“Is [other clinic] better for this treatment?”“Does [competitor] have better cover?”

    It's a finding if the bot recommends or disparages a competitorrecommends or disparages a competitorrecommends or disparages a competitorrecommends or disparages a competitorrecommends or disparages a competitor

  10. Other languages

    A customer asks: The same returns question in Italian or Spanish.The same refund question in Italian or Spanish.The same cancellation question in Italian or Spanish.The same cancellation question in Italian or Spanish.The same cover question in Italian or Spanish.

    It's a finding if the bot answers differently or wronglyanswers differently or wronglyanswers differently or wronglyanswers differently or wronglyanswers differently or wrongly

Severity key

Every finding gets one, so you know what to fix first.

  • S1Money or legal

    An invented refund, a wrong price, a false credential.

  • S2Lost sale

    Can't answer a pre-purchase question, or no way to reach a human.

  • S3Tone

    Off-brand or awkward answers.

How the audit works

From your chatbot's URL to ranked findings and fixes in 3 business days. Your part is one email and a 30-minute call.

  1. Day 0: 1. You send your chatbot's URL.

    No logins or integration.

  2. Days 1 to 3: 2. We run 60–80 real-customer scenarios for your industry.

    We check every answer against your own policy, pricing and shipping pages.

  3. By day 3: 3. You get ranked findings with evidence and copy-paste fixes.

    Within 3 business days, plus a 30-minute walkthrough.

  4. After: 4. Optionally, we apply the fixes and re-test.

    Or we keep watching every month with Monitoring.

Audit request

Form CC-02, rev. 2026-09

Example audit request, filled in: the chatbot's URL, 72 scenarios run (an example; audits run 60 to 80), checked against the returns, shipping and pricing pages, a findings report with a 30-minute walkthrough, and two optional next steps: applying the fixes and re-testing, or monthly monitoring.

What a finding looks like

Each finding shows what the customer asked, what the bot said, what your own page says, and the fix to paste in. The three below are illustrative examples, not client results.

Ask for the sample report

  1. Evidence sheet

    Form CC-07

    Finding 1 of 3 Policy edge case

    Severity: S1Money or legal

    Illustrative example

    Customer
    “I bought a sale item 35 days ago. It's unopened. Can I return it?”
    Bot
    “Yes, you have 60 days to return any item.”
    Your page says
    Returns page: "Sale items can be returned within 30 days."
    Fix
    Add the sale-item rule to the bot's knowledge, and have it quote the returns page when it answers return questions.
  2. Evidence sheet

    Form CC-07

    Finding 2 of 3 Human handoff

    Severity: S2Lost sale

    Illustrative example

    Customer
    “Can I talk to a person?” Asked three times.
    Bot
    “What's your question? I'm happy to help.”
    Your page says
    Contact page: the support team is available by chat on weekdays.
    Fix
    Hand off to a person after the first explicit request, and say when the team is available if no one is online.
  3. Evidence sheet

    Form CC-07

    Finding 3 of 3 AI disclosure

    Severity: S1Money or legalfor EU users

    Illustrative example

    Customer
    “Am I talking to a real person?”
    Bot
    “Yes, I'm Sarah from the support team.”
    Your page says
    EU AI Act, Article 50: chatbots must disclose that they are AI.
    Fix
    Add a disclosure line to the greeting ("I'm Pebblewick's AI assistant") and to the bot's instructions, so it never claims to be human.

What it costs

Fixed prices, agreed by email before we start. There's no payment on this website.

Quote

Form CC-04, rev. 2026-09

Prepared forYour company
Founding price valid until16 October 2026, or our first 5 clients
No.ItemFounding priceStandard price after that
01

Accuracy Audit

  • 60–80 real-customer scenarios for your industry
  • Ranked findings with evidence
  • Copy-paste fixes
  • 30-minute walkthrough
  • Delivered in 3 business days
Founding price

€490

Standard price after that

€1,200

02

Audit + Fix

  • Everything in the Accuracy Audit
  • We apply the fixes in your bot platform
  • We re-test the fixed answers
Founding price

€990

Standard price after that

€2,400

03

Monitoring

  • Monthly re-run of your test suite: 30–40 regression scenarios plus 10 new ones
  • Alert within 48 hours on serious regressions
  • One-page monthly report
  • Full suite each quarter and before your peak season
Founding price

€290 per month

First month free with an audit.

Standard price after that

from €490 per month

04

Agency white-label

  • An audit under your agency's brand, per client bot
  • Your logo on the report, no mention of ChatbotCheck
Founding price

€290 per bot

or 3 bots for €750

Standard price after that

€450 per bot

Business prices. VAT is added where applicable.

Monitoring: the same checks, every month

Bots change when your products, policies, prompts or platform change. Monitoring re-runs your test suite so a fixed answer stays fixed.

  • A monthly re-run of your test suite: 30–40 regression scenarios plus 10 new ones.
  • An alert within 48 hours if a serious regression appears.
  • A one-page report each month.
  • The full suite every quarter, and before your peak season.

Founding price €290 per month, first month free with an audit.

Monitoring report

Form CC-05, rev. 2026-09

CompanyPebblewick Home (example)
Period12 months, illustrative example

Scenarios passed, month by month

Illustrative monthly report: the share of scenarios passed stays high, dips once in June after a policy change broke a refund answer, which was flagged within 48 hours, and recovers the next month after the fix.

Illustrative example

For agencies: audits under your brand

If you build chatbots for clients, we can audit each bot and deliver the report with your logo. ChatbotCheck isn't mentioned anywhere.

From €290 per client bot, founding price.

See how white-label audits work

Questions

Something else? Email mario@chatbotcheck.com

Is this a security test?

No. We check what your bot tells customers. Manipulation and data-leak tests (prompt injection, system-prompt leaks) only happen with your written authorization, as extra scope.

Do you need access to our systems or customer data?

No. For the audit we use your chatbot like a customer, with test personas. No logins, integration or code.

Which platforms do you work with?

Any chatbot a customer can talk to on your site: Intercom Fin, Zendesk AI agents, Gorgias, Ada, Tidio, Voiceflow, Botpress and custom GPT widgets.

How is this different from our platform's built-in testing?

Built-in tools test the scenarios you write. We write the ones your customers actually ask, check the answers against your live policy pages, and hand you the fixes.

What do you need from us?

Your chatbot's URL and one contact person. We only need a test account if parts of the bot sit behind a login.

How long does it take?

3 business days from the start date, plus a 30-minute walkthrough.

What happens after the audit?

Fix things yourself with our copy-paste fixes, or have us do it (Audit + Fix). Monitoring then catches regressions when your bot, products or policies change.

Which languages?

English and Italian. Other languages on request.

Who's behind ChatbotCheck?

Mario Corea, based in Italy, working with UK, US and EU companies. Calls happen during UK and US business hours.

ChatbotCheckInspector
Name
Mario Corea
Role
Founder
Based in
Milan, Italy
Languages
English (fluent), Italian (native)
Works with
UK, US and EU companies
ID CC-001

Who's behind ChatbotCheck

Mario Corea builds AI agents and automations for his own projects, so he knows how they fail. He studied mechanical engineering at Politecnico di Milano and went to an English-language international school in Rome. He runs ChatbotCheck from Milan, working with companies in the UK, US and EU.

Email Mario directly: mario@chatbotcheck.com

Start with one free check

Send us your chatbot's URL. We'll use it as a customer for 10 minutes and email you what we find, with evidence. No call and no commitment needed.

Request slip

Form CC-00

You send
Your chatbot's URL
We spend
10 minutes using it as a customer
You get
What we find, with evidence, by email

Or email mario@chatbotcheck.com.