Example scenario, not a client project. The HR department of a manufacturing company spent three months testing an agent that answers employees' questions about leave and benefits. Before deciding on a rollout across five plants, criteria were set: answer accuracy, the share of cases handed to a person, employee satisfaction. Two criteria were met, one wasn't — the agent struggles with local plant regulations. Recommendation: roll out at two plants, and at the rest once the data is filled in.
Is this for you?
- A pilot "works" in one team, and the board asks whether to roll it out company-wide.
- No one has agreed how you'll know the solution is ready to scale.
- The design is going to a build team or a vendor, and you're worried something will get lost on the way.
- The people whose work will change don't know it yet, or are worried about it.
- After launch, no one will check whether the AI is learning and whether people use it.
What I do
- Scaling decision criteria — measurable thresholds that tell you: scale, fix or stop.
- A handoff package for the build team — a spec of what survived testing: flows, rules, states, AI patterns.
- A post-launch feedback loop — how to collect corrections and signals from people, and how to turn them into changes.
- Material for the people whose work changes — what changes, what stays, how to work with AI in the new process.
The AI layer
This service covers four gaps that research identifies as the reasons rollouts stall:
- Learning loop — MIT NANDA identifies a system's lack of learning as the main reason pilots don't make it to production. I design where the system gets its feedback and who reviews it.
- Preparing people — 48% of employees want training, and 45% want AI integrated smoothly into their work (McKinsey Superagency). The material is built on the designed process, not on general AI training.
- Scaling criteria — instead of "it seems to work": quality, adoption, time, cost of error.
- Increasing autonomy — a plan for when AI can take on more (from Workflow & roles).
What you get
- A scaling decision card — criteria, current results, recommendation.
- A design handoff package — a spec of flows, rules, states and AI patterns, ready for the build team.
- A feedback loop design — what we collect, who reviews it, how often, what it changes.
- A plan for the first 90 days after launch — metrics, reviews, decision points.
- Material for teams (as an option) — a guide to the new process, examples, frequently asked questions.
How it runs
- Current state — what the pilot showed, what data there is, who is involved.
- Criteria — agreed with decision-makers before we look at the results.
- Assessment — results against the criteria.
- Handoff package and loop — prepared with the build team.
- Material for people — tested with a few people from the team.
- Post-launch review — optionally after 30 and 90 days.
Roughly 2–5 weeks [to be confirmed].
How we'll know it worked
- The scaling decision is made on criteria agreed up front.
- The build team doesn't come back with questions about the basic rules.
- After launch, use of the solution grows, and people's corrections make it into the system.
- People in the new process know what the AI does and what they do.
What this doesn't cover
I don't run the technical rollout, organization-wide change management, or large-scale training. I deliver the design, the criteria and the materials those efforts can build on.