A state machine, a permission matrix and an event contract for DeskPilot, before any AI suggestion becomes a command

A DeskPilot ticket got resolved on a Friday, close to end of business. The following Monday, the customer reopens it: the problem came back. The system recalculates the SLA deadline from the moment of reopening, as if the ticket were brand new. The ticket is normal priority, with a 16-business-hour budget (960 minutes). It had already burned through roughly 13 business hours between creation and the first resolution. Recalculating "from zero" hands back the full 16 hours again — a deadline more generous than the ticket should get, quietly hiding the fact that most of the original budget was already spent.
That looks harmless, even kind. It isn't. An SLA exists so a manager, an agent and a customer all know how much time remains before an escalation kicks in. If the calculation treats every reopening as a fresh ticket, the SLA number stops meaning what the customer contract says it means. Worse: if nobody turned that rule into testable code, the behavior — right or wrong — depends on whoever happened to write that specific backend branch, or worse, on whichever model got asked to "summarize the rule" in a side conversation.
That is this chapter's subject: turning language, exceptions and responsibility into contracts that an automated test rejects when violated — before deciding what an AI is even allowed to suggest inside that contract. The thesis is direct: knowing the product means converting a rule into something verifiable before delegating any decision to a model. The decision this chapter asks of the reader is twofold: what is a deterministic rule, what is a suggestion, and who has the authority to change each one.
Before any architecture discussion, this chapter's core fits in one function. It takes a command, an authenticated actor and a ticket's current state, and returns whether the transition is allowed — no database read, no AI call, no dependency on the system clock:
export function applyCommand(cmd: TicketCommand, now: { iso: string }): TransitionResult {
const { actor, ticket } = cmd;
if (actor.tenantId !== ticket.tenantId) {
return { ok: false, error: "cross_tenant_denied" };
}
if (actor.status === "suspended") {
return { ok: false, error: "actor_suspended" };
}
if (actor.role === "ai") {
return { ok: false, error: "ai_role_cannot_command" };
}
if (cmd.expectedVersion !== ticket.version) {
return { ok: false, error: "version_conflict" };
}
const rule = TICKET_TRANSITIONS[ticket.state]?.[cmd.type];
if (!rule) {
return { ok: false, error: `invalid_transition:${ticket.state}->${cmd.type}` };
}
if (!rule.allowedRoles.includes(actor.role)) {
return { ok: false, error: `actor_not_permitted:${actor.role}` };
}
// ...applies the transition and returns the new ticket
}
Notice the order of the checks: tenant isolation first, actor status second, and only then — before the transition table is even consulted — the check that blocks any actor with an "ai" role. That order is not cosmetic. An actor from another tenant, or an AI role, should not even learn which transitions exist for that ticket; the function returns before that information ever leaks. The full code lives in examples/domain.ts, with 17 automated tests across examples/domain.test.ts and examples/events.test.ts, all run in this session via node --test examples/ (see verification.md).
The point isn't TypeScript syntax. It's that "who can do what, when" has an answer that fits in a decision table, testable, independent of any language model. If that answer only exists scattered across controller conditionals, a code comment, or — worse — the hope that a prompt "talks" the model into never suggesting the wrong action, it isn't a rule. It's a wish.
The common intuition is: "the AI only suggests, the human decides, so we're fine." That sentence is only true if the system receiving the human's decision also verifies who that human is and what they're authorized to do. A suggestion accepted by an agent, without the backend confirming that agent's role, is, in practice, the AI deciding — just with a coat of human-looking varnish in the middle.
Chapter 02 of this series landed on Opportunity A: ambiguity between billing and product issues during ticket triage, with six synthetic interviews suggesting agents lose time switching tabs to figure out a problem's real type before classifying it. The natural temptation, when building the first version, is to jump straight to "the AI suggests the category and the agent clicks accept." That's a legitimate product choice. What isn't legitimate is letting that acceptance touch ticket state without a testable rule somewhere that says: accepting a suggestion changes the suggestion record; changing the ticket requires a separate command, issued by the agent, checked against their role. The gap between those two sentences is the gap between a product with a contract and a product running on luck.
DeskPilot v1 models seven concepts. None of them are new compared to what chapter 02 described; what changes is that each one now becomes a type with explicit fields, leaving no room for free text to decide anything:
tenant_id.role (agent, manager, system, ai) and a status (active/suspended).ownerId field in this version.The confusion that shows up most often in AI-assisted products is treating ticket and suggestion as the same clock. They aren't. The ticket state (new, triaged, assigned, resolved, reopened) describes the real work being done by humans and the backend. The suggestion state (pending, accepted, edited, rejected, expired) only describes the lifecycle of an AI proposal — whether it was accepted, edited, rejected, or expired unanswered. Accepting a suggestion is evidence that a human agreed with it. It is not, by itself, a ticket transition:
export function applySuggestionCommand(
actor: User,
suggestion: Suggestion,
action: "ACCEPT" | "EDIT" | "REJECT" | "EXPIRE"
): { ok: boolean; suggestion?: Suggestion; error?: string } {
if (actor.tenantId !== suggestion.tenantId) {
return { ok: false, error: "cross_tenant_denied" };
}
if (actor.role === "ai") {
return { ok: false, error: "ai_role_cannot_command" };
}
const rule = SUGGESTION_TRANSITIONS[suggestion.state]?.[action];
if (!rule) {
return { ok: false, error: `invalid_suggestion_transition:${suggestion.state}->${action}` };
}
if (!rule.allowedRoles.includes(actor.role)) {
return { ok: false, error: `actor_not_permitted:${actor.role}` };
}
return { ok: true, suggestion: { ...suggestion, state: rule.to, reviewedBy: actor.id } };
}
Notice what this function returns: an updated Suggestion, never a Ticket. Turning "suggestion accepted" into "ticket triaged" requires a second, independent call to applyCommand, with the agent's own authenticated actor. It's verbose on purpose. The suggestion acceptance does not touch ticket state test exists specifically to stop someone, in a future rushed refactor, from merging the two functions "for convenience."
The table below summarizes examples/permissions.md. It covers ticket transitions; suggestions follow the same principle in a smaller table in the same file.
| From → To | Command | agent | manager | system | ai |
|---|---|---|---|---|---|
| new → triaged | TRIAGE | ✅ | ✅ | ✅ | ❌ |
| triaged → assigned | ASSIGN | ✅ | ✅ | ✅ | ❌ |
| assigned → resolved | RESOLVE | ✅ | ✅ | ✅ | ❌ |
| assigned → assigned (owner change) | REASSIGN | ❌ | ✅ | ✅ | ❌ |
| resolved → reopened | REOPEN | ✅ | ✅ | ✅ | ❌ |
| reopened → assigned | ASSIGN | ✅ | ✅ | ✅ | ❌ |
Two columns deserve attention. First, REASSIGN (changing the owner of an already-assigned ticket) requires manager or system — an agent can't pull themselves off a ticket alone, because that affects team workload, a trade-off chapter 02 assigned to the manager as buyer and owner of that call. Second, and more important: the ai column is empty across the whole table. That absence isn't an if (role !== "ai") throw scattered through the code — it's baked into the shape of the decision table itself. There's no configuration flag to accidentally "turn off" the AI's permission, because the permission never existed. That's the difference between a security rule someone can forget to apply and a rule that's impossible not to apply, because it's part of the data's shape.
The system role deserves its own note, since it's easy to read it as a second way for automation to sneak in through the back door. It isn't: system names backend jobs that already went through their own authorization path before this function ever runs — a nightly suggestion-expiry job, a migration script backfilling reopenCount, a webhook handler reacting to a verified external event. Those processes still carry an actor with tenantId and status, still get checked against the same table, and still can't do anything the table doesn't list for them; system never gets ACCEPT, EDIT or REJECT on a suggestion, for instance, because a job has no business overriding a human review. What system is not, and must never become, is a label someone attaches to a model's output to make the permission check pass.
One more design choice worth naming: applyCommand takes now: { iso: string } as an explicit argument instead of reading the system clock itself. That's what keeps the function pure and the tests deterministic — every test above pins its own timestamp and gets the exact same result on any machine, at any hour, forever. A backend calling this function in production supplies the real clock at the call site; the function never asks for it. The same discipline shows up in the SLA math below: every date that matters is an input, never something the function reaches out and fetches.
Back to the ticket reopened on Monday. Each tenant's Policy declares an explicit Calendar — not an assumption about timezone or holidays:
export interface Calendar {
businessStartMinuteUtc: number;
businessMinutesPerDay: Record<number, number>;
holidaysIso: string[];
}
The wrong version — the one that opened this chapter — lives in domain.ts, exported on purpose so the test can compare the two and prove the difference, never meant to be called from production code:
export function naiveReopenDueAt(reopenedAtIso: string, priority: Priority, policy: Policy): string {
return addBusinessMinutes(reopenedAtIso, policy.slaMinutesByPriority[priority], policy.calendar);
}
The correct version computes how much of the SLA budget was already consumed, in business minutes, between creation and the first resolution, and grants only what's left, with a configurable floor (reopenFloorMinutes) for the case where the budget is nearly exhausted:
export function computeReopenDueAt(params: {
createdAtIso: string;
resolvedAtIso: string;
reopenedAtIso: string;
priority: Priority;
policy: Policy;
}): string {
const totalBudget = params.policy.slaMinutesByPriority[params.priority];
const consumed = businessMinutesElapsed(params.createdAtIso, params.resolvedAtIso, params.policy.calendar);
const remaining = Math.max(totalBudget - consumed, params.policy.reopenFloorMinutes);
return addBusinessMinutes(params.reopenedAtIso, remaining, params.policy.calendar);
}
The naive reopen grants more time than the calendar-aware reopen after heavy SLA consumption test fixes a synthetic scenario — a ticket created on a Monday, resolved the following Monday, reopened at that same instant — where the normal budget (960 minutes) has been mostly consumed. It confirms naiveReopenDueAt returns a later (more generous) date than computeReopenDueAt, exactly the pattern from the opening bug: recalculating from zero hides how much time was already spent. A second test, computeDueAt skips the weekend for a Friday-afternoon urgent ticket, fixes an urgent ticket created on a Friday at 19:30 UTC, thirty minutes before close of business: since the two-hour budget doesn't fit in what's left of the day, the deadline has to land the following Monday — never inside the weekend, even though a naive sum of "two calendar hours" would comfortably still land Friday night.
Worth stating the model's honest limit: businessMinutesElapsed sums whole-day budgets between creation and resolution, an approximation good enough to prove the contrast between naive and correct calculation, not a minute-level measurement fit for production — documented in examples/rules.md, under Explicit limits.
This chapter's prompt cites Martin Fowler on bounded context (2014-01-15, accessed 2026-09-28): a bounded context is a logical model separation, and the article itself gives "multiple contexts within the same application" as a valid example. DeskPilot v1 takes that reading literally. Inside the same process and the same database, three vocabularies live apart:
Ticket, Suggestion, Assignment. Vocabulary: state, priority, SLA, owner.Tenant, User, role, status. Vocabulary: authentication, authorization, isolation.The confusion between "billing" as a ticket topic and "billing" as a system is, in practice, where Opportunity A comes from. Modeling bounded contexts here doesn't mean three databases or three deployments — it means the Ticket type doesn't import billing concepts, and a Suggestion categorized "billing" is a triage label, not a query into the billing system. Getting that boundary wrong is like mixing two languages in the same sentence and expecting the reader to know which word belongs to which grammar.
Every AuditEvent carries event_id, tenant_id, ticket_id, actor_id, occurred_at, schema_version and correlation_id. The schema in examples/events.schema.json closes the door with "additionalProperties": false — adding a free-text field requires a reviewed schema change, never a silent leak:
{
"required": [
"event_id", "tenant_id", "ticket_id", "actor_id",
"occurred_at", "schema_version", "correlation_id", "type"
],
"additionalProperties": false
}
The tests in events.test.ts confirm this both ways: a well-formed event passes, and an event with a ticket_text or ai_summary field gets rejected with unexpected_property. This is the defense against this chapter's fourth required failure, covered next: product telemetry is not the place to reconstruct what a customer wrote.
Symptom: a command arrives at the backend accompanied by text like "the model confirmed this ticket should be resolved"; someone decides that sentence, coming from an AI response, is enough to execute RESOLVE.
Cause: confusing a language model's output — which is text, subject to injection and misinterpretation — with authorization, which needs to come from an authenticated actor with a verifiable role.
Fix: applyCommand never reads prompt or ticket content to decide permission. The role check uses actor.role, a structured field coming from the authenticated session, never from text. The ai role can never carry a state-changing command, even claiming manager-shaped text test simulates exactly that attack: an actor with role: "ai" attempts RESOLVE, and the function rejects it with ai_role_cannot_command before even looking at the transition table. This maps directly onto LLM01:2025 Prompt Injection in the OWASP Top 10 for LLM Applications (2025 edition, accessed 2026-09-28): the taxonomy names the attack class, but the real defense is this test, not a checklist.
Symptom: the ticket reopened on Monday that opened this chapter. A deadline recalculated "from zero" looks safer (more time!), but breaks the SLA's original promise and erases how much time was already spent.
Cause: summing calendar hours, or resetting the SLA budget on every event, instead of consulting a declared business calendar and subtracting what's already been consumed.
Fix: computeReopenDueAt and businessMinutesElapsed, tested against naiveReopenDueAt in the same synthetic scenario. No real legislation or contractual obligation was inferred here — the calendar is per-tenant configuration data (Policy.calendar), to be confirmed with a real customer before any pilot, exactly as this pack's constraints require.
Symptom: the "accept suggestion" button in the UI, by itself, already moves the ticket to assigned or resolved, with no second confirming command from the agent.
Cause: treating "one click" convenience as equivalent to "an authorized decision," merging the suggestion's lifecycle with the ticket's lifecycle.
Fix: applySuggestionCommand returns only an updated Suggestion; never a Ticket. The suggestion acceptance does not touch ticket state test checks this explicitly, confirming the "ticket" key doesn't even exist in the return value. This separation is also the practical answer to LLM06:2025 Excessive Agency: the AI never has, structurally, the agency to close its own suggestion loop into a real action.
Symptom: a telemetry event carries "summary: customer requested a refund on card ending 1234" to "make dashboards easier later."
Cause: treating the audit log as a convenient place to store business context, without noticing that same log feeds analytics, exports and, eventually, third parties.
Fix: the closed schema in events.schema.json, tested by an event carrying ticket text is rejected and an event carrying an AI-generated summary is rejected. The defense isn't a redaction policy written afterward — it's the structural absence of any free-text field on the AuditEvent type.
The tests prove that, given an actor and a ticket with the fields described, the transition function obeys tenant isolation, role, status, optimistic concurrency and the state table — and that the event contract rejects content outside its closed field list. They don't prove that the DeskPilot UI stops a malicious agent from forging a session, that the real AI provider will never produce a biased suggestion, or that the configured business calendar is correct for a real customer. The PostgreSQL row security policies documentation (accessed 2026-09-28) is relevant here for a specific reason: "table owners normally bypass row security as well" — in other words, if this isolation were implemented only as a database RLS policy running under the table owner's role, it wouldn't protect anything. This chapter's isolation test runs entirely in application code; a complementary RLS policy, if adopted later, would need to be tested with an application role that lacks BYPASSRLS and ownership — never treated as the only protection.
Cost: formalizing seven types, two state machines and one event schema before any screen exists is work that doesn't show up in a quick demo. The payoff comes later: every new rule (a fifth priority tier, a new role) is a line in the decision table and a test, not an archaeology dig through scattered conditionals.
Security: this chapter's central risk isn't "the AI will invent a wrong answer" — it's "someone will treat the AI's answer as a command." The defense is structural (the ai role sits outside every permission table), not behavioral (politely asking the model not to do that).
Reversibility: reopening a ticket doesn't erase the prior resolution event; it adds a new one, with reopenCount incremented. The reopen preserves ticket identity and increments reopenCount without deleting history test verifies this at the ticket projection level; in a real backend, the original resolution AuditEvent would remain its own immutable row in the log — reversibility comes from never overwriting history, only ever adding to it.
Basic. Add a "critical" priority to the Priority type and to Policy.slaMinutesByPriority, with a smaller budget than "urgent". Acceptance criterion: a new test creates a "critical" ticket and confirms computeDueAt returns an earlier deadline than an "urgent" ticket created at the same instant. Worked solution: just extend the Record<Priority, number> on the test POLICY object and the Priority type in domain.ts; no other function changes, because all the calculation logic is already parametric in priority.
Intermediate. Add a triaged → new transition called RETRIAGE, allowed only for manager, used when an initial triage turns out wrong before any assignment happens. Acceptance criterion: a happy test confirms the transition for manager; an adversarial test confirms agent gets actor_not_permitted:agent. Worked solution: add an entry to TICKET_TRANSITIONS.triaged.RETRIAGE with allowedRoles: ["manager"]; no change is needed in applyCommand, because the function already reads the table generically — that's the payoff of centralizing rules in a data structure instead of a growing switch.
Advanced. Implement businessMinutesElapsed with minute-level granularity (instead of whole-day) for the resolution day, and write a test showing a case where the current (whole-day) version and the new version disagree by more than 60 minutes of remaining SLA budget. Acceptance criterion: the test fails against the current implementation and passes against the new one, documenting the difference numerically, not just qualitatively. Out of this chapter's scope, logged as an honest open item in examples/rules.md.
version) to detect whether two commands tried to change the same record at the same time, rejecting whichever arrived with a stale version.Chapter 04 (../product-designer-serie-04/) takes this domain as input and builds the first complete vertical slice: import a ticket, review, assign and audit, with two tenant fixtures, a development login explicitly isolated from production, and a simple web UI over TypeScript, Node.js and PostgreSQL — a didactic stack, with versions and a lockfile to be pinned during that implementation session, not an endorsement from this repository. No real DeskPilot benefit has been demonstrated so far; what this chapter delivers is the guarantee that, when AI enters the picture in the next chapter, it will find a contract that already knows how to say no.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares

A rule, an AI suggestion, and a human agent's decision, for the same synthetic ticket: only one of them can ever change the ticket's state. From that constraint, DeskPilot gains a typed…

An interview guide, an opportunity solution tree, and a willingness-to-pay test to decide whether DeskPilot deserves to be built, using only labeled synthetic evidence.

Turn a chatbot request into a falsifiable hypothesis, a decision agreement, and a first product experiment with DeskPilot, a synthetic support triage example.
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.