DeskPilot v0.3 exercises provider failure and backup recovery, then leaves a real release gate blocked
On the sixth consecutive AI provider failure, DeskPilot blocks another suggestion request. The ticket queue still works. Yet the support agent sees no degradation notice in the web UI. That gap alone prevents an invitation to real users. These are results from an injected-failure test, not a production incident: five attempts fall back and the sixth receives 429 tenant_suspended before the adapter runs.
In brief: local drills exercised fallback, atomic review with audit and outbox, tenant isolation, and recovery of a ticket and its audit trail after restore. Cost reservation does not enforce the provider's final bill, the UI does not show the suggestion, and the pilot has not run. The current decision is no invitation to real users until those gates, privacy operations, and consent review are complete.
The lesson is uncomfortable: keeping a ticket endpoint alive is only one part of keeping a support workflow usable. If a manager cannot see why suggestions stopped, if one tenant can run up uncontrolled costs, or if a backup restores data but not the application's login role, a green test suite gives an incomplete launch answer.
Product ownership includes operating failures, limiting spend, and protecting trust before expanding access or automation. The question is precise: what must be true before inviting actual users, and which observation stops the pilot? A technical gate can establish necessary conditions. It cannot establish customer value or product-market fit.
DeskPilot is a runnable teaching product for assisted B2B support triage. Support agents use it; tenant managers would buy it; end customers are affected by its output. Rules decide tenant boundaries, authorization, SLA, state transitions, and assignment. AI proposes a category and summary; a human reviews; the backend validates commands. There is no automatic customer message, live billing, production tenant, or completed pilot in this package.
Start with two ratios:
const ticketFlowSli = ticketFlow.good / ticketFlow.total;
const suggestionQualitySli = suggestionQuality.good / suggestionQuality.total;
The numerator and denominator need written definitions. ticket_flow asks whether eligible ticket operations completed; suggestion_quality asks whether eligible suggestion attempts produced an acceptable result. If the provider fails while tickets continue to move, the signals diverge. Averaging them into a single “50% healthy” score would conceal the operator's next action. A ticket-flow breach calls for investigation of the app and database path. A suggestion breach calls for checking the provider, tenant budget, and kill switch while preserving manual triage.
The proposed teaching targets are 98% for ticket flow and 70% for suggestion quality. These are proposed SLOs, not a customer SLA or an observed production result. For any window, error budget is (1 - SLO) × eligible events. The local dashboard keeps counters in process memory, so restart resets history and multiple processes would disagree. Before production use, the team would need durable collection, a specified time window, low-volume handling, and an on-call owner. The Google SRE Workbook's SLO guidance informs the ratio and budget method, not the choice of DeskPilot targets.
Chapter 05 left a typed, shadow-mode suggestion in DeskPilot v0.2; chapter 06 found that synthetic acceptance was 81.8% while its stricter synthetic clean-close metric was 63.6%. Neither number describes customers. Chapter 06 also named a concrete gap: its analytics plan included suggestion.reviewed, but the review route did not emit it.
The v0.3 review route calls one repository operation for the agent's accept, edit, reject, or expire action. PostgreSQL locks the suggestion row, applies the human decision, inserts a closed-schema audit event, and enqueues suggestion.reviewed in one transaction. The audit event answers who did what. The analytics payload answers a product measurement question; its later delivery still needs bounded retry and downstream deduplication.
const result = await repo.applySuggestionReview({
tenantId: actor.tenantId,
ticketId,
actor,
action,
nowIso: new Date().toISOString(),
});
// PgRepository locks and commits review, audit, and outbox together.
The PostgreSQL flow tests now inject a failed write and verify rollback; they also replay the same action by the same actor and verify no second event. SELECT ... FOR UPDATE serializes competing reviews of one suggestion. This proves the local transaction boundary in the tested path. It does not make delivery to an external analytics sink exactly once: a sink can accept an event before the worker fails to mark it delivered.
Other additions live in src/ops/: bounded delivery in outbox.ts, per-tenant request and cost controls in budget.ts, and SLI/error-budget calculation in sli.ts. Additive migrations 0003_reliability.sql and 0004_budget_reservations.sql add RLS-protected operations tables and an in-flight cost reservation. The new routes expose an operations view and an authenticated budget reset. Domain rules still own ticket state; AI still has no authority to apply a ticket command.
An analytics sink can remain unavailable longer than one request. The deliberately broken naiveRedeliverForever retries in one loop without delay, attempt ceiling, or dead-letter state. Worse, its table row never records the attempts. The operator sees a pending row with attempts=0 while delivery calls continue.
The real drain processes a bounded batch once per call. A failed row gets exponential backoff capped by policy; after the attempt ceiling it moves to dead_letter. The default ceiling is five attempts; one test chooses three to make the boundary clear. Against a fake sink that rejects ten deliveries, that test dead-letters on attempt three. The naive comparison finally succeeds on call eleven. The counts are test outcomes, not provider reliability measurements.
const attemptsAfter = event.attempts + 1;
const terminal = attemptsAfter >= policy.maxAttempts;
await repo.markOutboxFailed(event.id, {
nowIso,
nextAttemptAtIso: terminal ? null : backoffFrom(nowIso, attemptsAfter, policy),
terminal,
error: String(error),
});
The owner of a dead letter must inspect its cause and explicitly authorize reprocessing. A row left in dead_letter is visible failure, not successful delivery. The current worker is designed for one drain process per tenant. It has no atomic claim such as FOR UPDATE SKIP LOCKED; running competing workers would require a new concurrency design. A downstream sink also needs its own deduplication contract because delivery can succeed and the following “mark delivered” write can fail.
A bare suggestion_quality_sli: 0.40 asks the on-call engineer to remember a target and a runbook. buildDashboard instead returns each signal's status (ok, warning, breach, or no_data), budget consumption, and an action for non-ok states. The runbook ties ticket-flow breaches to the app and database path; suggestion breaches to provider status and per-tenant controls. The endpoint is scoped to a tenant.
An action string alone does not page anyone or establish incident response. The release checklist therefore names the operator, alert channel, and evidence required for each gate. An empty denominator yields no_data, not a flattering 100%.
The local drill took a custom-format pg_dump from PostgreSQL 17.6, stopped the disposable source container, and restored into a fresh container. The restore reported nine grant errors because deskpilot_app did not exist on the destination cluster. Table data appeared when queried as the superuser; the application role could not yet log in. This was an observed local finding, not an imagined failure.
A single-database dump does not carry cluster-level roles. Running the idempotent migration against the restored database recreated deskpilot_app and reapplied grants. The follow-up read used that application role under RLS and recovered a ticket in triaged, version 1, plus its ticket.triage audit event. The runbook now requires restore, migrate, then read as the app role before declaring recovery. That test says something real about this fixture. It does not measure recovery time, durability under load, or a production backup policy. See PostgreSQL row security documentation for why an owner or superuser read alone is weak isolation evidence.
A synthetic ticket, a polished demo, and a passing suite cannot prove that agents want this workflow or that buyers will pay. The release matrix separates operational readiness from commercial validation. The pilot protocol begins with “pilot not executed.” Recruiting, opt-in, data purpose, stop criteria, and feedback channels are plans pending authorized participants and infrastructure. Nobody has reported customer outcomes, revenue, satisfaction, or product-market fit from this chapter.
Product expense, AI expense, and support expense need separate owners and budgets. This example implements an AI request-count limit and a daily cost ceiling per tenant, checked before an adapter call, plus a per-tenant kill switch after five consecutive fallbacks. The proposed $5/day and 500 requests/day are fixtures. They are not pricing, contractual limits, or measured tenant consumption. The cost ceiling and kill switch answer different questions: spending too much does not automatically mean the provider is broken. Only a manager or system actor for the same tenant may reset the switch; the reset records actor and time.
Admission is now atomic per tenant and day. mutateTenantBudget locks the row, checks spent_usd + reserved_usd + $0.05, and reserves before the adapter call. With a synthetic $0.20 ceiling, a PostgreSQL test starts 20 concurrent requests: four enter and 16 are refused. This proves the admission invariant in that test. It does not prove a hard billed-spend ceiling. The provider may charge more than the $0.05 reservation; the high_cost fixture reports $0.42. An interrupted job can also leave a reservation unreconciled. A contracted cap would need a provider-enforced maximum, reconciliation for interrupted work, and tests of those cases.
The threat model names six vectors: untrusted input, cross-tenant leakage, prompt injection, unauthorized data mutation, excessive permissions, and unbounded consumption. Schema checks, human review, RLS tests, optimistic versioning, role checks, and request controls address parts of those risks. OWASP's 2025 LLM Top 10 supplies categories, not a security certification. Automatic log redaction, self-service export/deletion, retention operations, production authentication, and legal or contractual review remain release work. Nothing here asserts regulatory compliance.
The migration is additive. The kill switch is reversible through a named actor. A bad disclosure or a commercial promise based on synthetic data is harder to reverse; that is why privacy and evidence gates sit beside uptime in the decision matrix.
A release review should ask what someone does at 09:14, not merely whether a number exists at 09:14. The runbook therefore separates observation, decision, and recovery. For a suggestion-quality breach, the on-call person first checks whether ticket commands are still succeeding. If they are, the incident is a degraded assistant, not a full support outage. They examine provider responses and the tenant's request and spend counters, suspend suggestions for the affected tenant when needed, and tell support staff to use manual classification. A manager or authorized system actor alone can reset the switch after checking the cause. The reset actor and timestamp remain visible.
If ticket flow itself breaches, that same response would waste time. The operator checks the app path and PostgreSQL health, confirms whether commands can still persist with their audit trail, and uses the documented rollback or restore procedure. These are different failure domains; their alerts should lead to different people and different first commands. The local example returns action strings from /ops/slo, but no external pager, escalation rotation, or durable event store exists. The presence of an action field is a design test; staffing and alert delivery are still go/no-go evidence to collect.
One drill must also test reversal, not merely failure detection. The local backup drill intentionally destroyed its disposable source container after dumping it, restored into a fresh target, reproduced the missing-role error, ran migrations, and read a ticket plus audit row through the app role. That sequence caught a problem a “backup file exists” check would miss. For an actual pilot, the team would choose recovery-point and recovery-time targets with the buyer, schedule repeated restores, check the backup's storage and encryption permissions, and keep at least one operator who can execute the procedure without the original developer present. None of those operational commitments is demonstrated by this single local drill.
The outbox has an equally important recovery question. A dead-letter row is not permission for an automated replay loop. An operator first decides whether delivery failed before or after the sink accepted the event, whether the sink deduplicates by event ID, and whether replay would inflate analytics. Only then does an authorized person requeue it. If an event has already reached the sink but its local delivered marker failed, retry can duplicate it. The tutorial's fixed attempt ceiling controls load; it does not give exactly-once delivery. The operations owner must specify duplicate handling before a live metrics pipeline consumes these rows.
The pilot protocol asks a product owner to recruit support agents and a tenant manager through an authorized channel, explain what data the prototype will process, obtain explicit opt-in, and keep a way to leave the pilot. Support agents would perform ordinary triage with manual fallback available; the manager would review task outcomes and operating burden. The end customer whose ticket content is processed also matters. A legal or privacy owner must review controller/processor roles, retention, access, deletion/export requests, and any contractual restriction before real content enters the system. This is a review list, not legal advice or a claim of compliance.
The first session should be intentionally small. Record which tickets were eligible, which had a suggestion, which suggestions were reviewed, and which later reopened or received an actual customer score. Keep the denominator and observation window with each number. Record time spent by the agent and support burden separately from model cost; a fast model call can still create expensive human cleanup. Qualitative notes from the agent explain why a suggestion helped or hurt. The buyer's feedback addresses whether the workflow solves a paid problem. Neither can be inferred from 76 synthetic offline tests.
An observation form should preserve the difficult cases. If an agent edits a category, ask what evidence prompted the edit. If a customer reopens a ticket, record whether the original suggestion contributed or whether the cause was unrelated. If no satisfaction score arrives, mark it unknown rather than silently treating it as positive. A support lead should review sampling and eligibility rules before calculating a rate, because excluding hard tickets after the fact can make any pilot look good. Comparing assisted and unassisted triage also requires accounting for which agents, tenants, and ticket types enter each path. A before-and-after slide without that context is a prompt for investigation, not a causal result.
The agent needs a visible way to disagree with the model without losing control of the ticket. The manager needs to know that the disagreement was captured, that the ticket still followed deterministic authorization and SLA rules, and that an outage did not convert an optional suggestion into a required step. These are product-design requirements, not merely backend properties. They explain why the current missing suggestion UI is a genuine blocker even though the endpoint tests pass. Shipping the pilot without that screen would test a different workflow from the one the team claims to be evaluating.
The plan also needs a stop rule before it needs a success claim. Cross-tenant exposure, an unaudited state change, unexpected spend, or a failed manual fallback should stop the affected pilot activity immediately. A provider outage alone may permit continued manual work if that path and its communication are proven. The product owner records the pause decision, the operator records technical evidence, and the privacy owner joins if data may have crossed a boundary. Resumption requires the named owner to verify the fix and the rollback path, not a timer silently clearing an alert.
After any injected or real failure, a blameless postmortem should distinguish trigger, detection, customer-visible effect, containment, recovery, and a prevention experiment. The restore failure gives a concrete example: the trigger was a pristine cluster without deskpilot_app; detection was pg_restore grant errors; impact in this local drill was an unusable app role even though a superuser could read restored tables; containment was not pointing an app at that target; recovery was rerunning idempotent migrations and verifying through the app role. The action item belongs in the runbook and in a repeatable release gate. Blame adds no missing role to a backup.
At the end of an authorized pilot, product, engineering, and operations should choose to scale, pause, or simplify based on outcomes and burden. Scale requires evidence that agents complete the task better without disproportionate rework, that the buyer values that improvement, and that incident and support costs fit the business. Pause fits inconclusive samples or unresolved safety gates. Simplify fits a finding that manual classification plus good queue UX solves the problem better than suggestions. The current decision is earlier than all three: the real pilot has not started.
| Gate | Owner | Required evidence | Present state |
|---|---|---|---|
| Ticket flow under injected provider failure | Engineering | Manual ticket command succeeds; fallback observable | Tested locally; UI notice missing |
| Outbox and tenant isolation | Engineering | Bounded retry, dead letter, atomic review, app-role RLS | Local transaction and replay tests pass; multi-worker delivery remains out of scope |
| Cost cap and kill switch | Engineering + product | Concurrent admission, authorized reset, billed-spend control | 20-call admission test passes; provider maximum and interrupted-job reconciliation remain open |
| Backup recovery | Operations | Fresh restore; app-role ticket and audit read | Local drill passed after migration fix |
| Human review UI | Design + engineering | Agent sees suggestion, provenance, and fallback state | Blocked: UI renders no suggestion fields |
| Privacy and consent | Product + privacy owner | Data purpose, retention decision, opt-in, export/deletion process | Plan exists; not approved or exercised |
| Support and incident response | Operations | Named on-call, actionable alerts, stop/rollback drill | Local runbook and drills; real staffing unverified |
| Commercial signal | Product | Authorized pilot use and buyer feedback | Pilot not executed; PMF unknown |
The source examples/release-checklist.md has the full release matrix and evidence owner per line. The table here is a decision aid, not a substitute for that checklist. Current answer: no-go for real users. The missing review UI blocks safe use; provider-side billed-cost control and reconciliation are also open. A future opt-in pilot should stop on cross-tenant exposure, failed audit or restore evidence, uncontrolled spending, or inability to perform the task manually. These thresholds must be approved by the accountable team before recruitment.
The updated local run reports 76 passing offline tests and 16 passing PostgreSQL flow tests against postgres:17.6, including rollback/replay of the review transaction and concurrent budget admission. A separate restore drill recovered a ticket and its audit trail through deskpilot_app after the migration repair. These are useful, bounded facts about a synthetic implementation. They do not validate UI usability, actual provider recovery, traffic-level SLOs, a billed-cost cap, real-user benefit, or customer demand. Commands, environment, observed outputs, and teardown are in verification.md.
To reproduce the offline part, use Node 26 as recorded for this run and the chapter's lockfile. PostgreSQL integration needs a disposable 17.6 instance and separate owner/application credentials; follow examples/runbook.md for the exact setup. Keep all fixture data synthetic. The sample's role and migration steps matter: testing as the database owner can bypass RLS and give a false pass.
node --test test/ops/outbox.test.ts in examples/deskpilot/. Predict what happens if the policy passed to a failed drain has maxAttempts: 1. Acceptance: the first failure becomes dead_letter, with no scheduled retry. The comparison test explicitly passes its own three-attempt policy, so changing only DEFAULT_RETRY_POLICY will not alter that test's expectation.pg_restore plus migrate.ts procedure. Never place a plaintext application password in an article or backup.SLI is a measured service indicator, usually good eligible events divided by total eligible events. SLO is a proposed target for that indicator. Error budget is the allowance of bad eligible events implied by the target. Outbox stores pending side effects for later delivery; transactional guarantees require the business write and enqueue to share a transaction. Dead letter is a terminal delivery state requiring review. Kill switch disables a capability until an authorized reset. RLS is PostgreSQL's row policy enforcement; tests must use the role that the app actually uses.
Chapter 08 inherits DeskPilot v0.3, the local checks, a blocked human-review UI, and an unexecuted pilot protocol. It must carry those gaps forward as evidence, not turn them into a career story about a successful launch.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares
Todo teste do DeskPilot está verde desde o capítulo 03. Nenhum deles diz o que acontece quando o provedor de IA cai no meio de uma triagem, quando um tenant gasta sem teto, ou quando alguém tenta…
A rule, an AI suggestion, and a human agent's decision, for the same synthetic ticket: only one of them can ever change the ticket's state. From that constraint, DeskPilot gains a typed…
Uma demo de triagem impecável, gravada numa sexta-feira, evapora na segunda: nada tinha sido persistido. A partir desse incidente, uma fatia vertical completa do DeskPilot — importar, revisar,…
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.