Turning a v0 brief into a traceable decision without inventing an interview, a customer, or willingness to pay
The transcript below is synthetic, written for this chapter. No real support agent was interviewed.
Product Engineer: Think about the last ticket that gave you the most trouble to classify. What made it feel different? Agent (INT-001): It came in marked "urgent," but the text was about billing, not the product. I had to open another tab to find the customer's plan before deciding where to send it.
A line like that can turn into a product in minutes if nobody separates what was said from what was concluded.
| Observation (what was said) | Interpretation (what it seems to mean) | Hypothesis (what still needs a test) |
|---|---|---|
| The agent reopened another tab to check the customer's plan before classifying a ticket marked "urgent" | The category the customer declared ("urgent") does not match the real type of the problem (billing vs. product) | If the system suggested the real type from the text, the agent would spend less time switching tabs — but we don't know whether this is rare, common, or whether it meaningfully affects assignment time |
Chapter 01 of this series produced a v0 brief for DeskPilot, a teaching example of assisted triage for B2B support: the user is the support agent, the potential buyer is the support manager, and the affected person is the end customer. The brief had a falsifiable hypothesis, an owner still to be named, and an illustrative target of a 20% reduction in median time to correct assignment. It also had a deliberate hole: no interview, no real episode, no evidence beyond structured assumption. My argument here is that product research turns that assumption into a traceable decision; AI can help synthesize interviews, but it does not manufacture need, does not conduct an interview for you, and does not replace a willingness-to-pay test. This chapter asks the reader for two decisions: which opportunity to investigate first, and what result would actually justify building anything.
As in the previous chapter, no real customer shows up here. All six interviews in examples/interviews.synthetic.json were generated for this exercise and labeled "synthetic": true. That is not a footnote detail: a checker in examples/evidence-check.mjs rejects any record that tries to pass as real evidence, exactly as it would reject dressed-up data from an actual pilot.
The core discipline of discovery fits in a three-column table, repeated for every relevant interview excerpt: what the person said, what it seems to mean, and what still needs to be tested. Most discovery mistakes come from jumping straight from the first column to a roadmap, skipping the second and third.
Compare this with the v0 brief from chapter 01, which said: "hypothesis: category and summary suggestions, reviewed by an agent, may reduce median time." That sentence was an outcome hypothesis, written before any real (or synthetic) episode had been examined. This chapter's job is to confront that hypothesis with concrete episode reports — even synthetic ones — and see whether it survives, changes shape, or turns out to matter less than an opportunity nobody had named yet.
Two other synthetic interviews show why the "interpretation" column cannot become the "hypothesis" column without going through a test. INT-003, a senior agent, says reclassification is rare ("1 in 30 tickets") and does not see it as a central problem; she cares more about the time spent explaining bad automated suggestions to trainees. INT-004, a mid-level agent, says he reclassifies "4 or 5 out of every 20 tickets" and has already seen that turn into a formal customer escalation. A hasty interpretation would average the two estimates and conclude "reclassification happens in about 15% of tickets." That would destroy the most useful piece of information in the interviews: frequency may depend on seniority, ticket type, or how each agent defines "reclassify." The correct hypothesis is not an average number; it is a segmented, separately testable question (examples/opportunity-tree.md, Opportunity C).
A well-written brief can describe a real problem and still lack the intensity, evidence, or reversibility needed to justify a pilot. Teresa Torres's opportunity solution tree, accessed 2026-09-28, organizes that test in four layers: a business outcome at the root, opportunities (observed needs and pains) that lead to it, candidate solutions for each opportunity, and tests that verify the riskiest assumption before building anything. Torres recommends three to four story-based interviews before the first tree — a number that guides qualitative depth, not statistical representativeness. Finding a pattern across six synthetic interviews proves nothing about the market; it only proves that the method can structure contradictory signal without erasing it.
The initial segment needs to stay bounded without inventing a market size. For DeskPilot, the scope remains the one set in chapter 01: a single B2B support team with manual triage, a single authorized intake channel, tickets without especially sensitive data. What changes now is the level of detail: the volume to measure is "eligible tickets received per week in the target channel," not "the entire B2B support market"; the current alternative to compare against is the manual flow described by the agents themselves, not an assumption about "how every support org works." None of these synthetic interviews license a sentence like "B2B companies lose X% of productivity to manual triage" — that would require a real sample and a measurement method this chapter does not yet have.
Nielsen Norman Group describes the thinking-aloud method, accessed 2026-09-28, as robust and cheap, but explicitly not statistical: Jakob Nielsen writes that the method "doesn't lend itself to detailed statistics, unless you run a huge, expensive study." The same caution applies here. Interviews about past episodes reveal pattern and cause; they do not estimate prevalence. An internal report that cited "78% of agents prefer automated suggestions" from six conversations would be borrowing statistical precision the sample cannot support.
The full script lives in examples/interview-guide.md. The structure avoids the most common and least useful discovery question: "would you use an AI tool for this?" That question measures hypothetical enthusiasm, not behavior. Instead, each interview walks through one specific, recent episode in six steps:
1. Trigger: what made you notice that ticket was different?
2. Flow: walk me through what you did, step by step.
3. Workaround: is there anything you do on your own that the system doesn't offer?
4. Frequency: roughly how often does this happen?
5. Consequence: what happened next, for the customer and for you?
6. Purchase decision (buyer only): what mattered in the last tool you approved or turned down?
Recruitment, consent, minimization, and anonymization need to exist even when the interviews are synthetic, because the protocol is what changes once the team actually talks to real people. The guide defines: an invitation through the support manager with an explicit option to decline, consent that records purpose and retention period, and anonymization by code (INT-00N) rather than name. No identifiable end-customer data should be requested during the interview; if an agent mentions a specific case, the record strips the customer name and ticket number before it is stored. A team that skips this protocol "because the pilot is small" is deferring a privacy problem, not avoiding one.
The six synthetic interviews (four agents, two managers) feed three opportunities recorded in examples/opportunity-tree.md, all tied to the same outcome from chapter 01: reduce median time to correct assignment without raising reassignment rates, tenant incidents, or trainee onboarding load.
Opportunity A — billing vs. product ambiguity. This comes straight from the opening excerpt: INT-001 keeps a personal spreadsheet of keywords because the existing automated triage gets this kind of ticket wrong. Three competing solutions: (1) manual — a checklist to distinguish billing from product; (2) deterministic rule — a required "request type" field on the intake form; (3) AI-assisted — a category suggestion with the relevant ticket excerpt highlighted, reviewed by the agent. The riskiest-assumption test builds nothing: a senior agent reviews 30 historical ambiguous tickets and checks whether the hypothetical suggestion would have matched the final human decision, before any model is trained or called.
Opportunity B — invisible contractual priority. INT-002 describes asking colleagues which ticket to prioritize between accounts of different sizes, using a list kept off to the side by sales; she has already missed an SLA by following arrival order alone. Here the deterministic solution (syncing the CRM's SLA field straight into the triage screen) is tested before the AI solution (a summary that includes priority), because the riskiest assumption is not about a model — it is about whether the CRM's SLA data is accurate enough to display without review.
Opportunity C — disagreement about reclassification frequency. This grows out of the contradiction recorded between INT-003 and INT-004. Instead of resolving the disagreement with an average, the tree proposes measuring directly: a required reclassification-reason field, no AI involved, segmented by agent seniority. Only after seeing whether INT-003's pattern (rare, but costly in training) or INT-004's pattern (frequent, with escalation) dominates the real data does it make sense to evaluate an AI solution with a confidence signal.
The prioritization table in opportunity-tree.md scores each opportunity by reported intensity, number of sources, and reversibility of the tested solution. Opportunity A scores highest; C scores lowest — not because it matters less, but because the evidence is contradictory by design and needs more testing before any priority makes sense. This score is an ordering heuristic, not a definitive ranking. It decides what to investigate first; it never replaces the test defined for each riskiest assumption. A higher score built on one or two people's accounts is still a better-organized guess, not proof.
The pilot plan defines two tests that happen before any production code. In Opportunity A's concierge test, a senior agent manually reviews a sample of ambiguous tickets for one week and writes down, by hand, which category they would have suggested — without building the automation. This measures whether the idea would work in hindsight, without the cost or risk of exposing a model to live traffic.
The willingness-to-pay test charges nothing. It is a structured conversation with the support manager, presenting the concierge test's result and asking, in this order: what budget exists today for this kind of delay; what would need to be true to approve a paid pilot; who else needs to sign off. A specific answer about budget and approver counts as a willingness-to-pay signal. A line like INT-006's — "sounds like that would solve a lot of my problems" — with no mention of budget, approver, or condition, counts as praise, not demand. The plan records that distinction explicitly, because it is the easiest mistake to make right after a good conversation.
The final gate weighs three dimensions, drawing on the risk split used by outcome-oriented teams as described by Marty Cagan/SVPG, accessed 2026-09-28: viability (does the manager have real budget and authority, or does this depend on an unmapped external approval), usability (did the senior agent agree with the retrospective suggestion for most of the sample), and data access (can the team authorize anonymized historical tickets and a CRM with reliable SLA data). Advance requires all three favorable. Change segment applies when viability or data access fails on this team but the same problem shows up with comparable force on another queue available for discovery. Abandon applies when usability fails and no simpler solution resolves the ambiguity — in that case, chapter 01's decision log and this chapter's method remain valid as a case study, even if the specific product changes direction. None of these three paths is a process failure; they are the process working.
The buyer is not the user, and confusing the two is the most expensive mistake in this chapter. The support manager (INT-005, INT-006) decides budget and approval; the agent (INT-001 through INT-004) decides, day to day, whether to accept or reject a suggestion. A team that only interviews managers "because they're the ones who buy" will never discover that the biggest source of delay is an extra tab the agent opens to check the customer's plan — and a team that only interviews agents "because they're the ones who use it" will never discover that the manager would not approve budget without a measured reduction against a baseline. Both interviews are necessary, and they answer different questions.
AI-assisted synthesis comes in after the six interviews, not instead of them. An automated summary generated from the interview JSON might write something like "agents report occasional reclassification." That sentence is technically true and editorially dangerous: it erases the difference between "1 in 30" and "4-5 in 20," and INT-003's link between reclassification and trainee onboarding, which is the most actionable piece of the interview. The contract adopted here is simple: any AI-generated synthesis must point back to the source interview's id, and a disagreement between two sources (like INT-003/INT-004, marked with a mutual contradicts field in interviews.synthetic.json) must remain visible in the final document, not resolved by averaging or automatic consensus. The checker in evidence-check.mjs rejects the opportunity tree if no contradiction pair survives — a mechanical way to stop editorial convenience from erasing the research's most interesting signal.
Symptom: the question already describes the solution — "would you use an AI suggestion to categorize tickets faster?" — before the agent describes the episode itself. Cause: the interviewer wants validation, not discovery; the question asks for a polite opinion about an already-formed idea. Response: ask about the episode (trigger, flow, workaround, frequency, consequence) before any mention of AI; if the interviewee brings up AI on their own, that is data, but it should never be prompted. This chapter's guide explicitly bans questions like "would that be useful?" for the same reason: almost anyone polite will say yes.
Symptom: the brief lists "customer: support manager" and treats the manager's answers as evidence that agents will adopt the suggestion. Cause: it is easier to book a meeting with whoever approves budget than with whoever does the work every day; commercial language around "the customer" hides the fact that at least two roles have distinct interests. Response: keep the user, buyer, and affected-person columns separate (inherited from chapter 01), and check, for every claim in the v1 brief, which role it actually comes from. A sentence like "the customer would love this feature" should answer: which customer — the one who uses it, the one who pays, or the one affected by the outcome?
Symptom: an automatically generated summary reconciles divergent reports into a single consensus sentence, and nobody notices the information that got lost. Cause: language models are optimized to produce coherent text; coherence and fidelity to the source don't always coincide, especially when two sources disagree. Response: require every synthesis to point back to the source excerpt and id, manually review any summary that combines more than one interview, and treat preserved disagreement as valid data — often the most valuable finding of the round, because it reveals that the problem is not uniform across agents.
Symptom: a verbal expression of enthusiasm ("that would solve a lot of my problems") becomes, in the next meeting's notes, "customer confirmed interest in buying." Cause: ending a conversation on a positive signal feels good, and the pressure to show progress pushes toward the most optimistic possible reading. Response: record the literal answer, require specific mention of budget, approver, and success condition before counting it as a willingness-to-pay signal, and never book a verbal promise as revenue, a contract, or pipeline. This chapter's checker mechanically rejects any line in the pilot plan that claims a payment as already completed outside an explicit restrictions section — a second layer of protection against the same language habit.
This chapter's success criterion is not "DeskPilot validated demand" — that would manufacture a result no synthetic interview can produce. The criterion is: the v1 brief distinguishes observation from hypothesis in every opportunity; every opportunity compares at least three alternatives; every riskiest-assumption test has an explicit owner, deadline, criterion, and decision; and at least one real disagreement between sources remains visible rather than averaged out. node examples/evidence-check.mjs all verifies these four conditions structurally and fails if any is missing (see verification.md). That is process discipline, not a product measurement: the checker has no idea whether "Opportunity A" is real, only whether the document describing it follows the minimum contract to be taken seriously.
Once the team finally interviews real people, the bar changes: chapter 01's outcome (median time to correct assignment, with pending tickets visible in the denominator) is still what needs to be measured after any pilot, not the number of interviews conducted or the elegance of the opportunity tree. A well-built tree that points at the wrong segment is still a failed discovery effort; a modest tree that avoids an expensive pilot with no real demand is a successful discovery effort, even without producing a single line of code.
Interviewing has real cost even when nothing is charged: agent and manager time pulled away from work, research fatigue if the team keeps coming back every week for another conversation, and expectation risk — an interviewee may assume the company has already decided to build something just because someone asked. This chapter's guide caps the round at four to six agents and two managers, avoids hypothetical AI questions, and sets short deadlines (5 business days for the concierge test, one 30-minute meeting for willingness to pay) so discovery does not turn into an endless process.
Data safety starts at the interview, not at the pilot: no identifiable end-customer data should be collected at this stage; if an agent mentions a specific ticket, the customer name and ticket number are stripped before the record is stored. The synthetic interview JSON contains no real data at all — but the protocol it documents (interview-guide.md) is the same one that would apply to real data, including recorded consent, a retention period for the notes, and the interviewee's right to pause or end the conversation.
Reversibility here means two things. First, each test (concierge, willingness to pay) is designed to leave no operational trace if it fails: the concierge test never touches a real ticket, and the willingness-to-pay conversation generates no charge or contract. Second, the gate itself is reversible by design: "change segment" and "abandon" are valid outcomes, recorded in the decision log, not failures to hide. A team that only knows how to record "advance" is, in practice, forbidden from discovering that the right opportunity is somewhere else.
verification.md?Basic — rewrite a leading question. Take "would you use an AI suggestion for this?" and rewrite it as a sequence of questions about a past episode (trigger, flow, workaround, frequency, consequence). Criterion: the new sequence does not mention AI until the interviewee brings it up first. Worked solution: replace it with "think about the last ticket that gave you the most trouble to classify; what made it feel different? what did you do, step by step? is there anything you do on your own that the system doesn't offer? how often does this happen? what happened next?" — the same structure used in interview-guide.md.
Intermediate — separate buyer from user in a real sentence. Take the sentence "the customer would love this feature" (or an equivalent you've heard in a meeting) and rewrite it identifying which specific role supports it: user, buyer, or affected person. Criterion: the rewritten sentence names the role and the type of evidence still missing for the other two roles. Worked solution: "the support manager (buyer) said they would approve budget if there were a measured reduction; we still don't know whether agents (user) would accept the suggestion day to day, or whether end customers (affected) would notice a difference in the response."
Advanced — design a willingness-to-pay test without charging. For one of DeskPilot's three opportunities (or a real opportunity in your own context), write the three questions for the willingness-to-pay conversation and the criterion that separates praise from a buying signal. Criterion: the criterion requires specific mention of budget, approver, and success condition; no real charge occurs; the "advance" versus "keep validating" decision is written before the conversation, not after hearing the answer. Worked solution: see examples/pilot-plan.md, "Teste de disposição a pagar" section — INT-006's line is the documented counterexample of praise without commitment.
An opportunity solution tree is a structure linking a business outcome to observed opportunities, candidate solutions, and assumption tests, avoiding the jump straight from a problem to a favorite solution. A concierge test is a manual simulation of an automated solution, performed by a person, before building any system. Willingness to pay is a specific signal of budget, approver, and success condition — not verbal enthusiasm. A gate is the explicit decision point between advancing, changing segment, or abandoning an opportunity, with a criterion defined before seeing the result. Synthetic evidence is data generated for teaching purposes, labeled as such, that should never be reused as proof of real demand.
This chapter ends with one prioritized opportunity (billing vs. product ambiguity), a better-understood triage flow, and a set of gates not yet exercised against real data. No real interview happened; no customer confirmed payment; chapter 01's illustrative 20% target still has no real baseline behind it. In the planned chapter 03 of this series, the task turns technical: translating the prioritized opportunity into a verifiable domain model — a state machine, a permission matrix, and an event contract — so AI can suggest category and summary inside rules it cannot invent on its own.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares
Turn a chatbot request into a falsifiable hypothesis, a decision agreement, and a first product experiment with DeskPilot, a synthetic support triage example.
A rule, an AI suggestion, and a human agent's decision, for the same synthetic ticket: only one of them can ever change the ticket's state. From that constraint, DeskPilot gains a typed…
Um roteiro de entrevistas, uma árvore de oportunidades e um teste de disposição a pagar para decidir se o DeskPilot merece ser construído, usando apenas evidência sintética rotulada.
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.