A vertical slice of DeskPilot with persistence, tenant isolation and idempotency, before any AI suggestion exists
Friday afternoon, an engineer finishes wiring up DeskPilot's triage screens: queue, ticket detail, assignment form, history. He records a four-minute walkthrough — import a ticket, open it, assign it to an agent, watch the event show up in the history — and sends it to the team. Everyone likes it. Monday morning, he opens the laptop to keep going and the queue is empty. There is no ticket at all, not even the one "assigned" on Friday. The entire state lived in React variables in the browser; restarting the dev process wiped everything, and nobody had noticed because the recorded video hid that detail just as well as any production system would.
That is not a cosmetic bug. It is the difference between an MVP and a demonstration. An MVP ships a real slice of value that survives the real world — a process restart, a second user, a double-click, an attempt to reach something that is not yours. A demonstration ships the appearance of value along the exact path the presenter chose to walk. This chapter's thesis is direct: a vertical slice with feedback, persistence and verifiable rules produces better evidence than a collection of screens or prompts. The decision this chapter asks of the reader: what is the smallest complete flow that lets an agent finish a triage with confidence — and what, specifically, proves it finished?
This chapter picks up the domain built in chapter 03 (applyCommand, the permission matrix, the event contract) and carries it into the first end-to-end run: DeskPilot v0.1, local, with AI entirely switched off. Nothing here depends on a model to work — that is the baseline chapter 05 will measure an AI suggestion against, to find out whether it actually helps.
commandId, one row in the databaseBefore any screen, this chapter's executable core is an HTTP route that applies a command idempotently. The POST /tickets/:id/commands handler in examples/deskpilot/src/api/server.ts does exactly one thing before any domain logic runs: it rejects the request if it does not carry a client-generated command identifier.
const commandId = body.commandId ? String(body.commandId) : "";
if (!commandId) {
// A missing client-generated id is a caller bug, not something the
// server can make idempotent on the caller's behalf.
return send(res, 400, { error: "missing_command_id" });
}
const result = await repo.applyIdempotentCommand({
tenantId: actor.tenantId,
commandId,
ticketId: parts[1],
type: body.type as never,
actor,
expectedVersion: Number(body.expectedVersion),
assigneeId: body.assigneeId ? String(body.assigneeId) : undefined,
nowIso: new Date().toISOString(),
});
On the other side, PgRepository.applyIdempotentCommand (in src/db/pg-repository.ts) checks whether that commandId was already applied for that tenant before touching any row in tickets. If it was, it returns the recorded outcome without calling applyCommand again. If not, it applies the transition, writes the audit event, and only then inserts a row into applied_commands with a composite primary key (tenant_id, command_id) — it is that uniqueness constraint, not an in-memory check, that stops two concurrent requests from winning the same race:
try {
await client.query(
"insert into applied_commands (tenant_id, command_id, ticket_id, result_version, applied_at) values ($1,$2,$3,$4,$5)",
[input.tenantId, input.commandId, input.ticketId, next.version, input.nowIso]
);
} catch (err) {
const pgErr = err as { code?: string };
if (pgErr.code === "23505") {
return { ok: true, replayed: true, ticket: current };
}
throw err;
}
This is the opposite of "add one more screen." It is a line of defense against a real, common behavior: an agent double-clicks "Assign" because the interface did not give fast enough feedback, or the browser resends a request after a network failure. The test retrying the same commandId (double-click) does not duplicate the assignment or its audit event, in test/api/api.test.ts, reproduces exactly that scenario and passes both against the in-memory repository and against real PostgreSQL (test/flow/rls.test.ts).
The temptation of an MVP is to optimize for the impression of completeness: four polished screens, each looking finished on its own. The problem is that "looking finished" and "completing a task" are different properties, and only the second one is testable. A vertical slice forces the right question at every layer: does the data survive a restart? Does the backend — not the UI — decide whether this actor can see this ticket? Does a duplicate click produce a duplicate effect?
Those three questions map directly to the four required failures this chapter develops below. The intuition behind the thesis is that each question only has an honest answer once a complete, executable flow exists — not when there is a UI prototype with mocked in-memory data, and not when there is a document describing "how the system should behave."
Chapter 03 left a pure domain (applyCommand, applySuggestionCommand, calendar-aware SLA math) with no I/O dependency at all. This chapter copies that file without changing a single rule — the migration note at the top of examples/deskpilot/src/domain/domain.ts is explicit about that — and builds three layers around it:
src/db/pg-repository.ts, src/db/memory-repository.ts): two implementations of the same DeskPilotRepository interface, one real (PostgreSQL) and one offline (in-memory), so the HTTP contract tests run without Docker.src/api/server.ts): a framework-free HTTP server — the only production dependency in this package is pg, version 8.23.0 pinned in package.json/package-lock.json. Routes: POST /dev/login, GET /tickets (queue), GET /tickets/:id (detail), GET /tickets/:id/audit (history), POST /tickets/:id/commands (idempotent command).src/web/index.html, src/web/app.js): four screens — queue, detail, assignment form, history — in semantic HTML and plain JavaScript, no build step.Two fixture tenants (tenant-a, tenant-b) and four synthetic users are created by seed/seed.ts, running on the database owner connection (never the application connection). No name, email or field that would resemble real customer data appears in any fixture — more on that in Failure #3.
The dev login is the most sensitive part of this layer, so it gets an explicit guardrail instead of a comment: the /dev/login route only responds when the DESKPILOT_ALLOW_DEV_LOGIN environment variable is exactly the string "true", checked on every single request (not a constant loaded once at boot, so there is no code path that could "forget" to reread the configuration). Outside that case, it returns 503.
The complete task this chapter targets — "import ticket → review → assign → audit" — needs exactly four screens, not one more. Each one, under src/web/, explicitly covers the three states prototypes usually skip: empty, loading and error.
The queue screen (GET /tickets) starts in a loading state (renderStatus(queueStatus, "loading", "Carregando fila…")), announces the change in an aria-live="polite" region for screen reader users, and explicitly distinguishes "empty queue" (its own message, not an error) from "failed to load" (an error message, with a retry suggestion). That distinction matters: an agent who sees "queue is empty" knows there is no pending work; an agent who sees "failed to load" knows to retry or flag someone — conflating the two is a common way to hide a bug behind a generic "nothing here" message.
The detail screen loads the ticket and the audit history in parallel (Promise.all) and moves focus to the section heading once it finishes, satisfying WCAG 2.2's Success Criterion 2.1.1 Keyboard (every function operable by keyboard, including navigating between screens) without ever requiring a mouse.
The assignment form disables the submit button on click (the user-visible defense against a double click) and only re-enables it once the request returns — but since any client-only defense can fail (the tab can hang, the click can land before a re-render, the user can have two tabs open), the commandId generated at submit time is the defense that does not depend on the button's state at all. Success Criterion 2.1.2 No Keyboard Trap guarantees focus never gets stuck inside that form — testable by simply tabbing into the field and back out without a mouse.
The history screen lists audit events in chronological order, with no free-text ticket field anywhere — the same closed contract (assertClosedAuditEvent) chapter 03 defined for AuditEvent still holds here, and the UI simply has nowhere to render a field the backend never sends.
The DeskPilotRepository interface (in src/domain/repository.ts) is the contract that separates domain from storage. It declares applyIdempotentCommand as a method that must produce the same recorded result when called twice with the same commandId — that guarantee is not optional or Postgres-specific; the in-memory implementation has to satisfy it too, and the same test file (test/api/api.test.ts) runs against it.
The tenant-isolation choice has two deliberately redundant layers:
PgRepository method opens a transaction, runs SELECT set_config('app.tenant_id', $1, true) (the parameterized equivalent of SET LOCAL app.tenant_id = $1, which does not accept bind parameters), and only then runs the actual query, scoped WHERE tenant_id = $1.migrations/0001_init.sql enables row-level security on every tenant-scoped table and creates a tenant_isolation policy comparing tenant_id against that same current_setting('app.tenant_id', true).alter table tickets enable row level security;
-- (users, suggestions, audit_events and applied_commands get the same alter)
drop policy if exists tenant_isolation on tickets;
create policy tenant_isolation on tickets
using (tenant_id = current_setting('app.tenant_id', true))
with check (tenant_id = current_setting('app.tenant_id', true));
Why two layers for the same guarantee? Because PostgreSQL's own documentation is explicit about a detail that is easy to forget: "the table's owner is typically not subject to row security policies" (PostgreSQL, Row Security Policies, accessed 2026-09-28). If a migration team tests RLS while connected as the owner — the most common role during development — the policy will never fail, even when it is broken. That is why migrations/0001_init.sql creates a separate role, deskpilot_app, with no BYPASSRLS and no ownership of any table, and only that role is what test/flow/rls.test.ts runs against. One of the five tests there exists only to prove the opposite case: with no app.tenant_id set at all, a session authenticated as deskpilot_app sees zero rows from any tenant — the same documentation explains why ("If no policy exists for the table, a default-deny policy is used"), and a missing session value produces the same practical effect.
The idempotency trade-off is narrower than it looks at first glance. Stripe's documentation on idempotent requests (accessed 2026-09-28) describes a richer contract than this example implements: Stripe's idempotency key saves the body and status code of the first request — including 500 errors — and can be pruned automatically after at least 24 hours. This chapter's applied_commands table does not expire, does not return the original body byte-for-byte, and is not scoped per API key. It is enough to prove the property this chapter asks for — a retry does not duplicate an assignment or an event — but it is not a general-purpose idempotency implementation, and a reader building this for a real external API should read that API's specific contract rather than copy this table.
Failure 1 — Demo without persistence. Symptom: the flow works flawlessly live, but any restart (process, tab, laptop) erases the progress. Cause: the entire state lives in UI variables or an in-memory array in the Node process, never reaching a durable store. Response: real PostgreSQL, with a versioned migration (migrations/0001_init.sql) and a reproducible seed (seed/seed.ts); verification.md documents a specific test — triage and assignment over HTTP, then docker restart series-ai-pe04-pg, then a direct re-read through the repository — confirming the ticket and both audit events survive the container restart.
Failure 2 — Authorization only in the UI. Symptom: swapping the id in the URL, or calling the API directly and skipping the screen, lets one tenant read or change another tenant's ticket. Cause: the "is this your tenant?" check only lives in the React component deciding what to render, never in the handler that actually runs the query. Response: two independent layers, described above — a tenant_id = $1 filter on every PgRepository query, plus an RLS policy at the database level tested against a role that does not own the table. The test tenant B cannot read tenant A's ticket by id returns 404, not 403: the resource's existence is not revealed to the wrong tenant either.
Failure 3 — Sensitive seeds. Symptom: the seed script used "just for local testing" contains real names, emails or phone numbers copied from a customer spreadsheet, or a hardcoded password committed to the repository that later shows up in an environment that should never have had access to it. Cause: short-term convenience — copying a real value is faster than inventing a plausible synthetic one, and a fixed password avoids configuring an environment variable. Response: seed/seed.ts contains only synthetic identifiers (tenant-a, agent-b1) with no name, email or phone field at all; the deskpilot_app role's password is read from DESKPILOT_APP_PASSWORD, never written into the .sql file (the :'deskpilot_app_password' placeholder in migrations/0001_init.sql is substituted at runtime by migrate.ts); .env.example documents the variables with no real secret, and .env sits in this package's .gitignore.
Failure 4 — Too many components, no completed task. Symptom: the design board has twelve screens, each reviewed and approved on its own, but no end-to-end path has ever been tested all the way through. Cause: per-component review rewards local visual polish, not task completion; it is possible to "finish" every screen without ever having completed the triage task once. Response: this chapter's slice has exactly four screens (queue, detail, assignment, history), each mapped to one step of the same task, and the acceptance criterion documented in editorial-spec.md is "the complete task passes," not "the screens exist" — the full happy path: triage, assign, then read history test is the criterion, not a component checklist.
examples/deskpilot/ runs 25 automated tests in total: 14 in test/domain/ (13 inherited from chapter 03 unchanged, plus one new one about retry at the domain level), 6 in test/api/ (the HTTP contract against the in-memory repository, no Docker) and 5 in test/flow/ (the same contract against real PostgreSQL, with the deskpilot_app role). verification.md records the exact command and output for each block, from this session, on 2026-09-28.
What this proves: that the ticket and suggestion state machine stays correct after gaining an HTTP layer and a database; that tenant isolation is redundant across two independent layers, and the database layer fails closed when misconfigured, not open; that retries and double-clicks never duplicate an effect, both in memory and in real Postgres; that the history survives a container restart.
What this does not prove: that a real support agent can use the interface unassisted — examples/usability-protocol.md defines the think-aloud script and the timing instrument, but no session with a real participant happened for this package, and that is marked as pending, not as a result. It does not prove full WCAG 2.2 conformance — the UI applies three specific success criteria (2.1.1 Keyboard, 2.1.2 No Keyboard Trap, 4.1.3 Status Messages, the last one via the aria-live region in src/web/index.html), which is different from a formal conformance claim, which by the W3C's own definition requires a date, a level and a list of pages. And it does not prove real demand for DeskPilot exists — testing the application and validating demand are different questions, and only the first has evidence in this chapter.
This slice's operating cost, today, is essentially zero outside engineering time: an ephemeral Docker container (--rm, no named volume, dynamic host port), no managed service, no paid API call — because there is no AI call anywhere in this MVP's path. That is deliberate: the deterministic baseline needs to be cheap to rerun repeatedly, because it will be recreated for every test session and compared against the AI-assisted version in chapter 05.
On the security side, this package's explicit limits are: minimized PII (no fixture contains anything resembling a real person's data), a dev login isolated behind a flag read on every request, a database secret kept out of version control, and two tenant-isolation layers tested adversarially. What is not here: real pilot consent, a retention decision beyond a retentionDays field on the Policy inherited from chapter 03 (not yet enforced by an expiry job), and any claim of regulatory compliance.
Reversibility is high by construction: the container has no named volume, so stopping it erases all state without leaving anything on disk; the migration can be reapplied from scratch against a fresh container at any time; and the single production dependency (pg) is pinned in a lockfile, so reinstalling on a clean machine reproduces the exact same dependency graph.
npm install from the lockfile, no node_modules copied in.series-ai-pe04-pg container started with --rm, no named volume, dynamic port.npm run migrate and npm run seed run against the owner connection (DATABASE_URL), never against DATABASE_APP_URL.npm test (domain + API, offline) and npm run test:rls (real Postgres) both green.commandId does not duplicate an assignment or an event.docker ps -a --filter name=series-ai-pe04-pg empty).Basic. Run the full setup from examples/deskpilot/README.md on a clean machine: npm install, start the container, migrate, seed, run npm test and npm run test:rls. Acceptance criterion: all three test commands finish with fail 0, and you can explain, without looking at the code, why npm test does not need the container. Commented solution: if npm run test:rls fails with a connection error, the most common cause is .env pointing at a stale port — docker port series-ai-pe04-pg 5432/tcp changes on every docker run (and, as this chapter documented in verification.md, it can even change on a docker restart); re-read the port before rerunning the tests.
Intermediate. Add a test in test/api/api.test.ts confirming the queue screen correctly shows the empty state when a new tenant, with no seeded tickets, logs in. Acceptance criterion: the test fails before any code change (confirming there is no empty fixture tenant to test this against today) and passes once you seed — just for the test, without touching seed/seed.ts — a third tenant with no tickets. Commented solution: the simplest approach is to instantiate a MemoryRepository directly in the test and never insert a ticket for that tenant; listQueue already returns an empty array for a tenant with no rows, so the test is checking the API's contract ({tickets: []}, not an error), not an implementation gap.
Advanced. Chapter 05 is going to introduce a role that can only write Suggestion, never read or write Ticket. Modify migrations/0001_init.sql to create that role (deskpilot_suggester, for instance) with GRANT INSERT on suggestions only, and write a test in test/flow/ confirming that an attempt by that role to read tickets fails for lack of SQL privilege (not because of RLS) — the call should return a Postgres permission error, not an empty list. Acceptance criterion: the test explicitly distinguishes "no privilege on the table at all" (missing GRANT) from "privilege granted but no row visible under RLS" (the case already covered by the existing tests) — those are two different defenses, and conflating them in this preparation for chapter 05 would repeat, in miniature, this chapter's Failure #2.
commandId) more than once produces the same effect as applying it once.The next chapter picks up exactly where this one stops: the suggestions table and the getSuggestion method already exist in the DeskPilotRepository contract, but nothing in this package writes a suggestion yet. Chapter 05 introduces a typed category-and-summary suggestion in shadow mode, with an offline adapter and an optional real provider, a labeled dataset with a separate holdout, and the explicit decision that final priority, assignment and external communication remain human commands validated by the same applyCommand this chapter already tested. No domain rule from this chapter changes; what changes is that, for the first time in the series, there will be a model output to compare against this deterministic baseline — and the baseline is only an honest baseline because this chapter measured what it actually does, not what it looked like doing in a four-minute video.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares
Uma demo de triagem impecável, gravada numa sexta-feira, evapora na segunda: nada tinha sido persistido. A partir desse incidente, uma fatia vertical completa do DeskPilot — importar, revisar,…
A rule, an AI suggestion, and a human agent's decision, for the same synthetic ticket: only one of them can ever change the ticket's state. From that constraint, DeskPilot gains a typed…
DeskPilot's tests were green. A synthetic provider outage, an actual local restore, and a release review exposed what those tests did not establish: whether a support agent can safely use the pilot.
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.