Example scenario, not a client project. An app for financial advisors is getting AI that drafts a note after a client meeting. In the first cycle, advisors rewrite most of the notes. After a change — the AI asks about two key agreements before writing the note — in the third cycle most notes are accepted with minor corrections.
Is this for you?
- You have a prototype and want to know whether people will understand it before you pay to build it.
- You're introducing an AI feature and don't know how people will react to its mistakes.
- After a diagnosis you have a list of fixes, and want to check whether they help.
- You want to validate an idea in two weeks, not two quarters.
What I do
- Usability testing — 1:1 sessions in which users do real tasks.
- Validation sprint (about 2 weeks) — idea → prototype → tests → decision.
- Live-product test — observation and data from real use after a change.
The AI layer
- I test on real data, because AI behaves differently than on the examples in a presentation.
- I deliberately test the bad scenarios: a wrong, uncertain or incomplete AI result.
- I measure trust in practice: how many results people accept unchanged, how many they correct, how many they reject — and whether that changes over time.
- I check whether people know when they can rely on AI, and when they can't.
What you get
- A test plan — tasks, participants, metrics.
- Recordings and key highlights from the sessions.
- Results by metric — before and after the fixes.
- An improved prototype, or a list of changes, after each cycle.
- A recommendation — build, change or drop.
The metrics I look at
- Time to first value — how long until the first real benefit.
- Task completion — how many participants finish the key task.
- Drop-off points — where people stop.
- Return visits — whether people come back to the solution (in a live-product test).
- Acceptance of AI results — unchanged / corrected / rejected.
How it runs
- Goal and metrics — what we want to know, and how we'll recognize it.
- Recruitment — usually 5–8 people per cycle [to be confirmed].
- Sessions — remote or on site.
- Fixes — between cycles, usually within a few days.
- Next cycle — until the results are good enough, or it's clear the idea needs to change.
Roughly: one cycle about 1 week, a validation sprint about 2 weeks [to be confirmed].
How we'll know it worked
- The metrics improve from cycle to cycle.
- The build decision rests on results, not opinions.
- The biggest problems were found before the build, not after launch.
What this doesn't cover
I don't run large-scale A/B tests or performance tests. I can plan them together with your analytics team.