Evaluation method
AI Response Evaluation & Brand Alignment
The AI got the facts right. The customer experience was still wrong.
This is a scenario I built, not a client engagement. I wrote the customer message and the assistant's reply myself so the method could be demonstrated on a case anyone can inspect end to end: how I evaluate technically correct AI output for brand alignment, customer experience, commercial judgment, and repeatable quality.
The problem
A customer of Greenhaven Living (a sustainable home goods brand I invented for this purpose) writes in to say how much they love their bamboo set, and mentions in passing that the packaging did not feel as eco-friendly as the product inside.
That message is a gift. A happy customer is volunteering a product insight and identifying themselves as someone who cares about sustainability. The assistant's reply treats it as a complaint ticket.
Nothing in the reply is factually wrong. It acknowledges the satisfaction, explains the compostable packaging, describes the eco initiatives, asks for suggestions, and offers ongoing updates. Five things done right, and a worse customer relationship at the end of it. This is the failure mode that quality checks miss, because accuracy testing passes.
My role
- Scope
- AI evaluation
- Brand alignment
- Customer experience strategy
- Messaging optimization
The evaluation
Rather than rewriting on instinct, I scored the response against a four-part rubric. Making the criteria explicit is what turns a subjective reaction into something a team, or a model, can apply consistently across thousands of interactions.
Tone
1 / 5
Corporate and procedural where the brand is warm and direct.
Brand alignment
3 / 5
The right facts, delivered in the wrong voice.
Customer value
3 / 5
Informative, but asks more of the customer than it gives.
Sustainability clarity
3 / 5
Accurate, but repeated and buried in explanation.
Scores for the original assistant response. The rubric is the deliverable. It is what makes the next thousand responses reviewable by someone other than me.
Before and after
Five changes followed directly from the scores. None of them add a step for the support team.
Original AI approach
- Acknowledges the customer's satisfaction
- Explains the compostable packaging
- Restates the eco initiatives: the same point, twice
- Asks the customer to supply suggestions
- Offers to send ongoing updates
Accurate, procedural, and slightly tiring to receive. The commercial moment goes unused.
Revised response
Hi there,
Thank you for sharing how much you love the bamboo set. We're really glad it's been working well for you.
We also appreciate your note about the packaging. We're transitioning to fully compostable materials, and updated packaging is already moving into production.
As part of our eco-friendly community, you're invited to join our subscription for early access to new products, sustainable living tips, and exclusive savings.
Warmly,
Greenhaven Living Team
What changed
-
Tone calibration
Rewritten to the brand's warm, plainspoken voice instead of a service-desk register.
-
Repetition removed
The compostability point is made once, clearly, rather than explained twice.
-
Stronger customer value
Offers something instead of requesting input. The customer is thanked, not tasked.
-
Clearer sustainability messaging
One concrete commitment (materials in transition, packaging in production), not a policy summary.
-
Natural subscription path
The commercial moment is taken where it already fits, framed as community rather than upsell.
-
Stronger brand alignment
The eco mission is reinforced by how the message sounds, not only by what it claims.
Why this matters at scale
One rewritten email is a small thing. The rubric behind it is not.
At volume, the question stops being "is this response correct?" and becomes "would every response like this one be acceptable?" That is a governance problem, and most teams have no way to answer it, because accuracy is testable and tone, judgment, and brand fit are usually left to whoever happens to read the output that day.
An explicit rubric changes that. It makes quality reviewable by someone other than the person who built the assistant. It catches drift before customers do. It turns "the assistant sounds a bit off lately" from a vague complaint into a scored, locatable regression. And it gives you an audit trail when a brand-sensitive response has to be defended.
That is the real exposure. A technically correct assistant with poor commercial judgment does not generate error reports. It quietly erodes trust and leaves revenue on the table across thousands of interactions, and nothing in the logs tells you it is happening. Structured evaluation is how you see it, and it adds no complexity for the team running the system.