31 Jul 2026

7 min read

Running a Product Discovery with AI.

We’ve been actively rebuilding our Product Discovery workshops around AI… this is how.

Like every agency working with these tools, we are still working out what AI actually changes about how we build digital products. Even though we do not have a tidy answer for this, what we have is a new way of working reshaped around AI… and a rule of thumb to judge it: by the results it provides. No matter how smart it may sound on paper, if it doesn’t work, there’s no point.

Anthropic recently published a piece on harness design, the scaffolding you build around an AI model so it can run long, complex tasks without falling apart. We read it expecting tactics for agents. What we found, instead, was a fairly close description of how we had ended up working. Without ever using the term internally, we had indeed built a harness of our own.

This is not a finished system, and it is not a one-off project. It is where we have got to, end to end, from first kickoff to build plan. The old way was not broken. But a tool like this asks for more than being bolted onto how we already worked, so we rethought the work around it instead. So that is what we did, continuously learning where the shape is right and where it is not.

Harness Design

How Anthropic pushed Claude further in frontend design and long-running autonomous software engineering.

Learn more

What we built

Discovery throws off a large amount of context, and most of it scatters by default: a deck here, a Figma file there, a transcript nobody reads through again, and the tiniest details across our brains. On a long project with a rotating team, that is exactly the context an AI could be working from. Instead, until recently, it was leaking away between working sessions.

So we built our process into a plugin: a set of skills, one per stage of the work, that holds the line on how we move from "what should we build" to "here is how to build it." Brainstorm, define, plan the build, delegate the execution. Each stage knows what the one before it produced, and what the one after it needs. We built the plugin ourselves, and we keep iterating on it.

We pointed all of it at a single repository: not a codebase, but a Discovery repo. Research, analytics, personas, journeys, prioritised features, architecture calls, the risk log, all of it, with every decision written down alongside its reasoning, in the moment. It is closer to the project's brain than a filing cabinet. We make decisions against it, argue with it, and come back to it when nobody can remember why we made a certain decision three weeks ago. Because it is all structured text in one place, we can ask it questions, and so can the AI. It is only ever as good as what we wrote down, but writing it down is a matter of discipline, and no longer a game of memory.

We feel its advantages the most during onboarding: someone joining halfway through a project no longer has to book time with whoever holds all the facts in their head and a Notion doc. Instead, newcomers can point Claude at the repo and ask. This does not replace a proper handover, but does allow for it to start from a better place.

The combination of that plugin that encodes the process and a repo that holds the state, is our harness… here is how it plays out.

No plan, no build

The first skill in the chain has one job: nothing gets built on a plan nobody has validated. It forces a quick intake, restating the problem, naming the user and the goal, listing what we do not know, and confirming before anything moves forward. The intake scales: a small change gets two minutes, a rebuild gets a long session. This is what the paper calls the planner, turning a vague prompt into a spec, and just as much, the planner that is told not to over-specify, because every premature decision becomes a foundation later work is stuck building on, even when it was wrong.

Kickoff preparation goes in first, deliberately light, entailing a read of the client's current site, a scan of a handful of competitors, a first cut at personas. We do not do the deep research yet, because we do not know the real problem yet. Then comes the kickoff, and the discipline we hold hardest: we do not present our conclusions. We listen, check the diagnosis the client already signed up to, answer the blocking questions, and keep our hypotheses in our pocket. We have not earned the right to grade our own thinking until the client has shared theirs.

Define hard, cut harder, hold loops open

As we sit round by round with the client, most of the product gets defined through a technical audit, an analytics deep dive, persona work, the domain model, user journeys, prioritised features, overall architecture. On one recent rebuild we also ran a customer survey, going through dozens of long, unguarded replies from the client's most engaged users, and fed every line into the repo.

From there, we cut. Saying no to a feature is worth more than saying yes: separate what the product actually needs from what is just designer habit dressed up as a decision, and a must-have list can lose a third of its items. Ten architecture decisions might come down to the three that genuinely drive scope; the rest are build-time choices that do not belong in a discovery brief.

We also leave some things open on purpose. A few strategic loops, positioning, brand direction, the bets only the founders can own, we choose not to close remotely. Arrive at the in-person discovery session with everything decided, and you have wasted the one thing that room is good for.

A clickable prototype runs in parallel off the same evolving brief, so the thinking always has something to visually refer to. Before anything reaches the client, we put our own work through deliberately hostile review, (at one point a critique written entirely as a sceptical founder,) with one goal: to find where our artefacts talk past the client rather than to them.

The art of improving UX

Great tools aren’t made of AI alone, and PostHog is one of our favourites.

Learn more from Tiago

A workshop you can click through

We arrive at the client’s office with two things: a product that is mostly defined, and a prototype they can click through. Instead of slides about the product, they experience the product in their hands. That invariably changes the angle of discussions from arguing about abstractions, to walking through real flows, opening the loops we had held back on, and working through them live by changing the prototype on the spot when a decision moves past it.

Afterwards, the outcomes fold back into the repo, the wireframes tighten against the decisions made in the meeting room, and the work routes into the build plan, turning "what to build" into "how to build it.”

The build plan is written so the people and AI agents alike can follow it through implementation without re-arguing Discovery.

When the build runs, each task passes through two separate checks: “does it meet the spec”, and “is it any good”, before it counts as done. Whatever made the work does not sign it off. That is not a quirk of ours. It is the separation Anthropic builds into their coding agents, and what Martin Fowler's catalogue of generative-AI patterns calls Evals: you do not let a model judge its own output, because its self-assessment is "woefully over-confident, preferring to make up a plausible answer rather than admit ignorance."

A very happy ⅖ of Significa after the first new Product Discovery workshop.

None of this is new

Strip away the specifics and a handful of ideas carry the same principles this new way of working delivers: 1) state should live in files, not memory; 2) whoever makes the work should not be the one who grades it; 3) you define "done" before you build, not after; and finally 4) assumptions go stale, so you keep testing them rather than trusting the ones you started with.

None of those are new to us. Our handbook has named them for years: integrity, doing the work right rather than merely fast; being critical-thinking partners rather than yes-people, challenging the work instead of rubber-stamping it; transparency, working openly enough that a decision and its reasoning sit together. What changed is not the values. It is that we stopped relying on remembering to live them, and built a harness that makes them the path of least resistance.

There is one value this proverbial harness not only protects, but frees: playfulness. The odd idea is usually the first thing a deadline kills, because exploring it gets expensive. When ten variations of a layout or a gesture cost minutes instead of days, more portions of the odd idea are likely to survive the deadline.

How we collaborate

Transparency, trust, collaboration, quality: these are non-negotiable for us.

Learn more

What the harness doesn’t hold

For all the scaffolding, the call that matters most on a project like this is one no skill in the chain would make for us. A brand's voice is often the whole game: the specific, slightly odd way a client actually talks to their customers, versus the version that reads fine in a deck. AI is fluent, and fluency pulls toward the clean, competent, generic middle, sanding the weird off everything until what comes back reads fine and means nothing.

Keeping the weird in is a taste decision, made and defended by people, over and over. The harness drafts the options and catches the obvious mistakes, but it cannot tell you which of ten competent variations is the one with a pulse. That judgement is the actual work, and it is still ours.

This way of working is more demanding than the old one. The drudgery falls away, and what is left is judgement, all day, with a machine that is happy to be confidently wrong and the catching left to us. A workshop day leaves us more drained than it used to, because the effort did not disappear, it moved, from production to discernment. Figma's Yuhki Yamashita makes the same point: when anyone can build, speed stops being the edge, and craft becomes "choosing, not accepting.”

There is an upside buried in that cost. A stretch of almost pure judgement, with your own past reasoning sitting in the repo to argue with, is how the eye gets sharper. You go back, see where you were wrong, and adjust.

The principles of building products people love.

Lessons learned at MS, Google and Uber, amidst other insights.

Tune in on Yuhki Yamashita

A harness is never complete

As the tools improve, we strip out scaffolding we no longer need and add structure where the work has moved. We get plenty wrong. The point was never that ours is right. It is that when ours is wrong, the repo tends to show us where, and we change it. That is truest of the plugin itself: a Claude Code plugin the whole team keeps rebuilding, a little different every time one of us runs it on real work. It gets better because the people using it do, which is the part the tool does not touch. Keeping a long, many-handed effort coherent is the same problem whether the agent is a model or a room full of people. The people are still where the best of it comes from.

The question we would leave you with, then, is not whether to use AI in product work. That is settled. It is this: which of your principles depend on people remembering to follow them, and what would it take to build them into the work instead? The tool we’ve built is new. That part is not.

Tiago Duarte

CPO

Author page

Tiago has been there, seen it, done it. If he hasn’t, he’s probably read about it. You’ll be struck by his calm demeanour, but luckily for us, that’s probably because whatever you approach him with, he’s already got a solution for it. Tiago is the CPO at Significa.

We build and launch functional digital products.

Get a quote

Related articles