A testable brief for moving beyond tickets without pretending to have authority or results
The ticket said, “Add a chatbot to speed up support.” The team shipped a well-built widget. It loaded quickly, passed its tests, and captured conversations. On Monday, the support agent still read every request, worked out which queue it belonged in, and looked for an available colleague. The support manager asked why assignment still took so long. The developer replied that the chatbot met the ticket. Both were right.
This is an illustrative scene, not a Lemon customer story or a report from a running product. It exposes a common gap: correct software can leave the motivating problem untouched. The product question comes before the stack: what decision would get a request to the right owner sooner without increasing errors or quietly shifting work to support?
My argument is that Product Engineering begins when an engineer shares responsibility for defining a problem, choosing an intervention, and observing its consequences with product, design, support, and business partners. Technical depth still matters. So do boundaries: the new title alone grants neither access to users nor authority to change their workflow. Both must be agreed with the people who bear the effects.
We will create the first artifact for DeskPilot, a teaching example of assisted triage for B2B support. At this point it has no customers, interviews, revenue, production application, or proven benefit. This chapter delivers a v0 brief that colleagues can challenge and an offline checker that rejects missing decision fields. The next chapter will ask whether the proposed problem and audience survive contact with actual users.
“Add a chatbot” specifies a deliverable. It says little about the person experiencing the delay, the behavior that should change, or the evidence that would count as progress. Start with two competing interventions:
| Candidate | What changes | How it might fail |
|---|---|---|
| Without AI: required form fields, fixed categories, and agent-reviewed routing rules | Less ambiguous input, understandable rules | Poor fields burden customers; rules become stale |
| With AI: category and summary suggestions reviewed by an agent | Long requests may become easier to understand | A wrong suggestion may bias the agent toward a wrong assignment |
The outcome of interest in both cases is time to correct assignment: from arrival to the first assignment confirmed as correct by a review rule agreed before the test. Our opening hypothesis is falsifiable: “In an authorized pilot with eligible tickets from a B2B support team, agent-reviewed suggestions reduce median time to correct assignment compared with the non-AI flow without increasing reassignments or tenant incidents.” The v0 brief proposes a 20% relative reduction in the median over a two-week period as an illustrative experiment target. It is neither an observed result nor an approved pilot criterion. Support and the pilot owner must review volume, seasonality, baseline quality, and acceptable error cost before adopting any operational threshold.
A shipped feature can be demonstrated by a deployment, tests, and technical acceptance. Adoption means observing whether agents actually use it, when they decline it, and why. Outcome requires a comparison against a baseline and a look at adverse effects. “The model returned a valid category” proves a data contract. “The agent accepted the suggestion” records a behavior. Neither proves a better experience for the customer. The distinction affects design: suggestion, review, assignment, and correction should be separate events, each subject to permissions and a retention decision.
Product Engineer is not an automatic promotion from software engineer. These labels describe responsibilities, and teams draw their lines differently. A software engineer working deeply on reliability, data, or infrastructure can have substantial product influence. Fullstack describes reach across technical layers. Product Engineering usually adds closer involvement in shaping the user experience, deciding what to build, and observing whether it helps. The PostHog Product Engineer Handbook, accessed September 27, 2026, describes PostHog's own practice. It does not establish an industry-wide standard or show that one title is more valuable.
| Role | Typical contribution to support triage | Decision to agree, not assume |
|---|---|---|
| Software engineer | State integrity, authorization, APIs, operations | Which technical trade-offs fit the delivery window |
| Fullstack engineer | End-to-end flow across UI, backend, and data | Where simplification remains safe |
| Product Engineer | Connects the problem, prototype, measurement, and iteration | Which technical experiment will teach the team something |
| Product Designer | Researches and designs the interaction, errors, and cognitive load | How an agent notices, reviews, and rejects a suggestion |
| Product Manager | Handles priority, value, viability, and strategic alignment | Which outcome merits resources and when to stop |
| Growth engineer | Tests acquisition, activation, or expansion when relevant | Whether a funnel is the bottleneck and how to measure it responsibly |
One person may cover several columns in a small team. That does not erase the expertise of their partners or grant their authority by implication. Marty Cagan's product versus feature teams framework emphasizes collaboration across product, design, and engineering around outcomes. He also states that his terminology is not standardized. Use the framework to ask who can make which decision and who will carry its consequences, not to rank colleagues by how “product-minded” they sound.
AI-native has two distinct layers here. First, AI can assist the engineer: draft a small function, propose test cases, or summarize documentation. The engineer still defines the requirement, inspects the diff, tests boundaries, and owns the running system. Second, AI can be part of the user's experience: DeskPilot might suggest a category and short summary. Errors in that layer reach an agent and possibly the end customer. Human review and backend validation become part of the product contract. The 2025 DORA report on AI-assisted software development, accessed September 27, 2026, offers research context about AI and development organizations. It cannot establish that DeskPilot will help, or that an individual developer will advance in their career.
DeskPilot is a proposed tool for B2B support teams, not a validated market. The direct user is an agent who reads the request, looks at authorized context, and decides whether to use a suggestion. The potential buyer is a support manager who must justify cost, quality, and fit with the team's workflow. The affected person is the end customer waiting for a correct handoff and an answer. A single delay lands differently on each of them.
| Person | Possible goal to investigate | Risk from a rushed solution | Missing evidence |
|---|---|---|---|
| Support agent | Find the right queue without losing control | Reviewing opaque suggestions takes longer; automation erodes judgment | Workflow observation, time per step, authorized interviews |
| Support manager | Reduce waiting and rework at a predictable cost | Operations and maintenance cost more than they save; metrics become targets | Volumes, total cost, quality policy, pilot criteria |
| End customer | Get an accurate response with privacy intact | A fast but wrong assignment adds a handoff; sensitive data is exposed | Consented feedback, reassignments, reviewed samples |
“Possible” does real work in that table. We have spoken to none of these people. This is a map of questions, not a completed discovery exercise. Speed can oppose accuracy: a fast wrong assignment creates another transfer. Agent autonomy can pull against the manager's desire for consistent routing. Additional context can improve a suggestion while increasing processing cost and data exposure. A brief should make these conflicts visible before a compelling demo hides them.
The initial scope is also reversible: one B2B support team, one authorized intake channel, and requests that can be handled without especially sensitive data. We do not know the market size. Even the current process needs observation. For the teaching exercise, we assume a manual queue as a synthetic continuity fixture and explicitly invite correction after investigation. If an existing rules-based flow already works well, AI might have no role here.
A one-page brief should fit a conversation with the people running support. Short does not mean vague. The worked document lives in examples/product-brief.md, with an interest map in examples/stakeholder-map.md and a decision record in examples/decision-log.md. Those files use synthetic information. None of their lines represents an actual customer interview.
| Field | Starting position, pending discovery |
|---|---|
| Audience | User: B2B support agent. Potential buyer: support manager. Affected: end customer. |
| Situation | A request arrives through an authorized channel; an agent must identify category and owner. |
| Problem | Time to correct assignment may be high; frequency, causes, and cost are unmeasured. |
| Current alternative | Manual queue assumed for the exercise; verify the real process before designing a product. |
| Non-AI alternative | Structured fields, fixed categories, routing rules, and human review. |
| Hypothesis | Agent-reviewed category and summary suggestions may lower median time without increasing reassignments or customer harm. |
| Missing evidence | Channel volume, reliable timestamps, definition of “correct,” baseline, authorized observations, reassignment rate, cost per case, data policy. |
| Final owner | Assign a pilot owner with support and product before using real data; the prototype author cannot appoint themselves. |
| Risks | Personal data, incorrect suggestions, acceptance bias, review effort, cross-tenant exposure, recurring cost. |
| Out of scope | Automatic customer replies, SLA changes, priority decisions, billing, production deployment, or product-market-fit claims. |
| Reversible decision | Test an offline prototype with fixtures; if a pilot is approved, enable suggestions only for a consenting group with an immediate off switch. |
Missing fields create decision failures. If a brief says “user: support” but omits buyer and affected customer, a team may optimize seconds in an interface while transferring risk to the customer. If it says “use AI” without a non-AI option, there is no honest comparison. If it omits an owner, no one can stop a faulty pilot. The local examples/check-brief.mjs accepts the complete fixture and rejects an intentionally incomplete one. It checks document structure and the separation of hypotheses from evidence. It cannot prove the hypothesis or authorize data processing. verification.md records the commands and results actually run.
Here is the core of the brief in copyable form. Every value remains a tutorial hypothesis:
Problem: time to correct assignment may be high; baseline unknown.
User: support agent. Potential buyer: manager. Affected: end customer.
Without AI: structured fields + routing rules + human review.
With AI: suggested category and summary; agent confirms or rejects.
Outcome: median time to correct assignment per eligible ticket.
Guardrails: reassignment rate, severe errors, and review burden.
Missing evidence: volume, timestamps, baseline, interviews, consent.
Owner: to be agreed with support and product before a real pilot.
Rollback: disable suggestions; retain manual flow and minimal redacted log.
Even that compact block needs operational definitions before it becomes a test. People running support must define and sample-check what “correct assignment” means. Tickets still unassigned at a fixed observation cutoff remain in the eligible denominator and appear in the pending share. A descriptive median among assigned tickets uses only observed times and reports that smaller denominator; alone, it can look better while hard cases remain pending. For a decision, the team should also report the share correctly assigned by a deadline set before the pilot, or use time-to-event analysis with explicit censoring. We need to know whether an arrival timestamp marks the actual arrival or a delayed import, whether complex tickets cluster in particular shifts, and whether reassignment is an error or a normal part of service. A good brief names useful ignorance instead of smoothing it over.
The first increment need not be a chatbot. To learn whether the brief captures a decision, synthetic request fixtures, a described manual flow, optional suggestion examples, and a deterministic checker are enough. This avoids integration work before anyone knows the problem exists and keeps the non-AI alternative eligible to win.
In the offline toy, examples/check-brief.mjs reads the document and fails if it lacks stakeholder roles, a falsifiable hypothesis, an owner, a non-AI alternative, missing evidence, or rollback. An intentionally invalid fixture gives the checker a meaningful negative case. It does not measure service time. That gap is explicit. Later chapters may use TypeScript, a pinned Node.js LTS release, PostgreSQL, and a simple web UI as a teaching stack. Picking those tools now would not validate the opportunity.
An authorized pilot would need a stronger contract. Deterministic backend rules would control tenant boundaries, authorization, SLA, states, and assignment. AI would only suggest a category and summary from permitted text. An agent would see the suggestion, correct or reject it, and confirm an action. The backend would validate the resulting command regardless of what the UI or model says. There would be no automatic external reply, real billing, or deletion of external data. Personal data would be minimized, logs redacted, and retention and access approved before ingestion. A prompt cannot enforce tenant isolation.
Every effect needs an owner, permission, evidence, and a reversal path. A category change requires the agent's permission, an auditable event without sensitive request text, and a way to restore the previous value where feasible. A real pilot requires organizational consent, an operational owner, and an off switch. Even a “suggestions only” interface has an effect: it can bias an agent toward a wrong decision. The agent must be able to spot uncertainty and keep using the manual route.
We still have no evidence that a model-based architecture is needed. Compare improved fields and a simple taxonomy against the observed current process first. If most delay comes from unavailable staff, better classification will not clear the queue. If unclear routing policy is the cause, a rule update may beat a model at lower cost and risk. The team must keep that possibility open until it sees the work being done.
A ticket team often receives a specified solution; a product team is more likely to receive a problem and negotiate one. Real organizations mix both modes. You can do useful discovery in a ticket team without claiming permission to change policy, contact customers, or spend budget.
In a ticket team, propose a bounded conversation with the requester and a support lead: “Can I use a defined window to map the workflow and compare a non-AI option before estimating the chatbot? I need supervised access to agents, no sensitive data, and feedback on Friday. I will bring a brief and open questions; priority remains your decision.” Record investigation time, authorized people, permitted data, spending limit, decisions you can make alone, and who approves a pilot. If user access is denied, use documented workflow and mark the missing evidence. Imagination does not become an interview.
In a product team, PM, designer, engineer, and support lead can shape the hypothesis and guardrails together. Engineering can own feasibility and reversible implementation. Design can study how an agent detects and fixes mistakes. PM can maintain priority and strategic fit. Support can describe operational load and escalation. A named pilot owner makes a go/no-go call after hearing those views. Collaboration should not spread responsibility so thinly that no one can stop a harmful experiment. In either setting, agree on a feedback cadence: review the brief before building, observe each round, and record a decision to continue, change, or stop.
A useful autonomy agreement answers seven questions: what problem may we investigate; which people and data may we access; what may we decide; what requires approval; which effects are forbidden; who resolves conflict; when do we revisit the agreement? It protects the engineer and the team. “Act like an owner” without these answers often increases accountability without increasing the ability to act.
Symptom: someone is called a Product Engineer but sees work only after prioritization and cannot talk with support. Cause: the label changed while access, decision rights, and responsibility did not. Response: ask for a small discovery window, an operational counterpart, and an explicit decision boundary. Record any refusal and the available alternatives. If the organization does not authorize contact, work within the real boundary and state what remains unverified. The user otherwise gets an interface built on assumptions no one can correct.
Symptom: the team compares models and token prices before seeing why requests wait. Cause: a technology demo has become the problem statement. Response: map the flow, establish a baseline, try a form or rules, and state the condition under which AI would outperform them. The delay might come from staffing or queue policy; a classifier cannot solve either by itself. For the end customer, irrelevant automation may make service less legible.
Symptom: “100% of tickets classified” becomes a success claim while reassignments rise. Cause: output is easy to count; correctness and harm take observation. Response: report technical validity, adoption, and outcome separately; define denominator and guardrails before the pilot; review an error sample. A prototype can work technically and still fail to improve service. That is a useful finding and may prevent an unnecessary recurring cost.
Symptom: engineering changes categories or workflow without consulting agents, a support manager, design, privacy, or the people handling complaints. Cause: “ownership” has been mistaken for the right to take over another person's decision. Response: map affected people, seek review before creating effects, name the final owner, and define escalation. A bad suggestion reaches agent and customer; excessive retention raises organizational risk; a new fee affects the buyer. Consultation will not remove disagreement, but it will reveal the trade-off.
Before measuring benefit, define the unit (eligible ticket), denominator (every eligible request arriving in the channel and window), observation cutoff (a fixed interval after each arrival), method (arrival, first assignment, sample-checked confirmation or correction), segments (type and complexity), and limits (seasonality, agent differences, missing records). Count tickets without a correct assignment at the cutoff and report their share of all eligible tickets. The median of completed times includes only tickets with an observed event and states how many entered that calculation. Record agent review effort and ignored suggestions. Without both views, a lower median may only indicate easier requests.
Where feasible, compare the observed current flow, a non-AI alternative, and AI suggestions with human review. A small sequential comparison can teach a team something, but it cannot control every difference. Do not call correlation causation. Decide in advance what would stop the trial: significant assignment errors, data exposure, loss of agent trust, unacceptable total cost, or no improvement after a suitable window. The brief's 20% over two weeks is a provisional teaching target, not an approved decision threshold; operational limits require a baseline and local agreement.
For the offline brief, the success criterion is narrower: required files exist, the valid fixture passes, the incomplete fixture fails, and the document does not present a hypothesis as observed evidence. That verifies decision discipline, not commercial value. The next meaningful step is to observe a few authorized interactions and revise the brief, including if the original problem proves wrong.
A suggestion costs more than a model call. It creates integration, processing, human review, monitoring, incident response, training, and opportunity costs. The non-AI alternative also costs maintenance, but may be cheaper and more predictable. Estimate from local data; do not publish return on investment derived from a fixture. A continue decision should account for the customer's service, not only a price per request.
Support tickets can contain names, contracts, passwords, or sensitive business details. Even a proof of concept should use synthetic data until there is an authorized basis for handling real requests. A real pilot needs data minimization, backend tenant and role enforcement, redacted logs, retention, consent, and a deletion procedure owned by the appropriate team. This chapter claims no certification or legal compliance. Legal and privacy review belongs to qualified owners. If an external provider is considered, document data flow and terms before sending anything.
Reversibility is a feature: keep the manual flow, enable suggestions for a limited group, disable them without breaking request states, retain only the minimum evidence needed to investigate errors, and avoid a migration that forces every agent through a model. The cost of going back belongs in the same discussion that approves moving forward. A feature that can only be disabled through a risky deployment weakens operational control when it is most needed.
verification.md?Basic — rewrite a ticket. Turn “add a chatbot” into one observable problem statement and one falsifiable hypothesis. Include a user, a measure, and a guardrail. Pass condition: a support colleague can describe an observation that would refute the proposal. Worked response: “Assignment may take too long because agents must interpret free text; compare median time to correct assignment and reassignments with the current flow.” The baseline is still missing. That is an honest open question.
Intermediate — negotiate an investigation. Draft a request for 30 minutes of authorized observation, a brief-review meeting, and a list of decisions reserved for the owner. Pass condition: the recipient knows which people, data, and time are involved and can approve discovery without approving a pilot. Worked response: ask for supervised access to a queue without sensitive data, report observations by Friday, and leave SLA changes, spending, and deployment with existing owners. If access is refused, record the gap and review documentation.
Advanced — design a comparison. Specify events, denominator, three conditions, and a stop rule for DeskPilot. Pass condition: unassigned requests remain in the denominator, error severity is captured, complexity differences are reported, and no fictional percentage appears. Worked response: record arrival, assignment, confirmation, and reassignment; compare observed baseline, improved forms/rules, and reviewed suggestions in agreed windows; segment by request type; stop on data exposure or a severe error. Suggestion acceptance is diagnostic, not the outcome.
An outcome is an observed change for users or the business, including harmful side effects. An output is what the team shipped. A baseline is a measured description of the existing flow, with a method and window. A guardrail prevents improvement in one measure from hiding harm elsewhere. A tenant is one organization's data boundary in a shared application. A rollback is a procedure to restore a safe state, not a hope that nothing goes wrong.
The first useful result here is a decision someone can reject with evidence. DeskPilot has moved from zero to a candidate problem, three distinct stakeholders, a possible comparison, and written boundaries. Customers, interviews, and a baseline remain missing. In the planned chapter 02 of this series, we will investigate who experiences the delay, how they work now, and whether this scope warrants a product. Until then, the brief remains v0.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares
An interview guide, an opportunity solution tree, and a willingness-to-pay test to decide whether DeskPilot deserves to be built, using only labeled synthetic evidence.
A rule, an AI suggestion, and a human agent's decision, for the same synthetic ticket: only one of them can ever change the ticket's state. From that constraint, DeskPilot gains a typed…
Um roteiro de entrevistas, uma árvore de oportunidades e um teste de disposição a pagar para decidir se o DeskPilot merece ser construído, usando apenas evidência sintética rotulada.
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.