This year we shipped three AI features into production products. None of them were chatbots, and all of them started with a spreadsheet of test cases.
Build the evaluation harness first
Before we design a single screen, we write down what ‘good’ looks like and collect examples. That harness runs on every change, so quality never silently drifts.
Show your work
People trust AI output when they can check it. Citations, confidence and an easy way to correct the result do more for adoption than any amount of polish.
Scope ruthlessly
The best AI features we built did one job very well. The worst ideas we prototyped tried to do everything.
Tagged
AI, Evaluation

Written by
Kofi Mensah
AI Engineer
Kofi builds retrieval systems and evaluation tooling, and makes sure every AI feature we ship can explain itself.
View profile



