← All Insights/Design sprint2025 · Updated 2026-08-27

Making AI Agents Real: From Use Case to Working Sprint Prototype

The prototype includes the agents and the interface a person uses to supervise them. Four weeks on sanitised data produce a decision and a path to production.

Making AI Agents Real: From Use Case to Working Sprint Prototype

The 2025 version of this article described a facilitation sprint. Interviews and alignment workshops, then a prototype. That was honest when building was the expensive step.

It is the wrong sequence now. An LLM chooses the steps after you state the outcome. Two weeks of slides produce a product the operator will never use. The product is the supervision layer: the plan before irreversible work, and a receipt after. Override lives on the same screen as the answer.

Diagram comparing GenAI, Copilots, and Agentic AI capabilities

Week one used to be a value-versus-feasibility matrix, as if every job were a feature to automate. The split that matters now is different. Some jobs can be stated as an outcome an agent can pursue. Some still need a person on every step. Mixing those two in one backlog is how you get a chatbot nobody supervises.

So the four weeks changed. Week one decides what the agent is for, and who will supervise. It also names the actions that must not be autonomous, including anything that writes to a client record. Week two designs the agents and that interface together. Architecture without the supervision screen is a drawing of a system nobody will operate. Week three builds both on sanitised sandbox data. Week four proves it. Audra Eval and an ISO 27001:2022 governance brief. The demo is someone taking over, then the path to Build and a 2-5 year Run.

The nine deliverable titles did not change. What you walk away with did. Not a chatbot. Working agents, and the screen a person uses when the model is wrong. That is already how the Sprint is sold. This article no longer pretends the middle weeks are workshops.

The older article also promised UX from the beginning, as a facilitation skill. That hid the real change. UX is not a workshop input. It is week three's build: confidence and override, with citations on the same screen, tested with the operator on sanitised data. If that screen is missing, the sprint produced a demo.

We still start on sandbox data with no live customer records. That constraint is how a four-week sprint stays four weeks, rather than four weeks plus a data-access review.

If a use case has been a slide for a year, write to hello@windmill.digital. The door is the AI Strategy & Readiness.

Continue

Working on a related product or operating question?

Explore the four ways Windmill can help shape direction, prove value, build for production or evaluate a live system.