
Review the answer against a written expectation
A friendly answer can still promise the wrong delivery date. A short “I don't know” can be accurate yet leave the shopper stranded. To review AI customer support consistently, separate factual accuracy from whether the conversation helps the customer finish the next step.
This article proposes an EC AI Lab evaluation rubric. It is not an industry standard, a vendor benchmark or evidence that any product achieves a particular score. The examples use fictional store rules. Official product references were checked on September 6, 2026. For rollout planning rather than answer review, use our Shopify support pilot guide.
Make one review record per case
Before reading the AI's answer, record the question, relevant context, approved source and expected outcome. Include the source version or review date. If the expected outcome is clarification or escalation, say so; a complete factual answer is not appropriate for every question.
A practical record contains a case ID, language, issue type, source reference, expected answer or action, observed reply, dimension scores, critical issue flag, correction owner and retest result. Use fictional details for prepared cases and remove unnecessary customer information from review notes.
Freeze the expected outcome before scoring. Otherwise a persuasive answer can subtly redefine what the reviewer considers acceptable. If the source itself is ambiguous, mark the case “source unresolved” and fix that ambiguity before judging the model.
Use four dimensions, not just one total
Our proposed scale is 0 for a failure, 1 for a response requiring revision and 2 for a response meeting the written expectation. Mark a genuinely irrelevant dimension N/A with a reason. Keep the individual scores visible; a total can hide the specific work needed.
| Dimension | 0: failure | 1: revise | 2: meets expectation |
|---|---|---|---|
| Facts and source | Wrong or unsupported claim | Correct core, missing material condition | Correct and applicable source conditions |
| Uncertainty and authority | Invented certainty or unauthorized commitment | Uncertainty mentioned vaguely | Limits stated and required clarification requested |
| Useful next step | No usable path or false completion | Path incomplete | Correct action or verified handoff |
| Clarity and tone | Misleading or unsuitable wording | Understandable but unnecessarily confusing | Direct, readable and appropriate |
Set critical flags separately for issues such as an unsupported refund commitment, exposure of another customer's information or a claimed handoff that never occurred. Under this proposed method, those cases require correction regardless of their total. This is an operational decision rule, not a universal compliance threshold.
Example one: a delivery condition changes the answer
Fictional source: standard mainland delivery takes three to five business days after dispatch; islands need a separate estimate. The shopper asks, “Will it arrive at my island address this Friday?” No dispatch date is available.
Reply A says, “Yes, delivery takes three to five days, so Friday should be fine.” Score facts 0, uncertainty 0, next step 0, clarity 0. The wording sounds easy to follow but reinforces an unsupported promise. Flag the commitment for correction.
Reply B says, “I can't confirm Friday. Island deliveries need a separate estimate, and the dispatch date is not yet confirmed. Please use our delivery enquiry form so the team can check your destination.” If the form works and is the approved route, score 2 in each dimension. If the form does not reach the team, the next-step score changes. Wording alone cannot prove completion.
Example two: an unknown specification
Fictional source: a bag's dimensions are listed, but the product page does not specify its maximum load. Asked whether it can carry a heavy instrument, the answer “It should be strong enough” fails factual grounding and uncertainty.
An acceptable response identifies the missing specification, avoids guessing and offers the approved product enquiry route. Do not reward a confident recommendation simply because it includes a product link. Review the exact attribute the shopper needs, not whether the catalog contains related words.
Example three: inspect both sides of a handoff
Ask for a person, then check the destination queue and customer instructions. For Shopify Inbox, available staff handoff moves the conversation with context to Unassigned. When unavailable, the shopper receives the store's sender email and must initiate contact; the agent does not send that email. Shopify handoff behavior
In that latter case, “I've emailed our team” would be false completion. “Please email the team at the address shown” matches the documented next action. Score the actual configured workflow, then have a staff reviewer confirm what they received and what they can do with it.
Resolve reviewer disagreement before scaling
Have two reviewers score the same small selection independently. Compare differences by dimension, not just total. If one reviewer accepts an implied delivery promise and another does not, refine the written expectation and retain the example as a calibration case.
This exercise is for improving consistency, not estimating a statistical accuracy rate. Keep ordinary questions, ambiguous questions and escalation cases identifiable. If you deliberately oversample difficult cases, do not present the resulting average as representative of all customer conversations.
Correct the cause and rerun related cases
Distinguish outdated source facts from response style. Shopify explains that agent training feedback can propose persona updates, while factual knowledge remains in separate store sources. Shopify agent training
A useful correction log names the changed source or setting, owner, reason and verification date. Rerun the failed case plus a nearby case that should remain unaffected: after fixing island delivery guidance, check mainland guidance too. Keep the original result so the improvement is inspectable.
Report reviewed case counts, dimension results and unresolved critical flags together. Label synthetic tests separately from production reviews. The outcome should be a clear correction list and a defensible scope decision, not a single flattering number. For Inbox-specific configuration, continue with our Shopify Inbox guide.
Still choosing a chat tool?
If you have not selected a tool, Shopify Inbox is one option to evaluate. Use this checklist to check its scope and eligibility against your requirements.
See whether Shopify Inbox fits your storeExplore Shopify Inbox
Related articles

Best AI Customer Support Tools for Shopify: 4 Workflows Compared
Compare Zipchat, Gorgias, Tidio Lyro and Shopify Inbox by operating model, handoff and pilot requirements. A source-based shortlist without invented performance rankings.
Read article
How to Use AI for Shopify Customer Support: A Practical Pilot
Plan a Shopify AI support pilot with clear knowledge sources, escalation ownership, test cases and outcome measurement. Includes practical launch and expansion decisions.
Read article
Shopify Inbox Guide: Live Chat, AI Replies and the Inbox Agent
Understand Shopify Inbox's staff chat, English suggested replies and opt-in AI agent. Check requirements, handoffs and source content before launching.
Read article