A practical method for setting boundaries, measuring quality, and starting document extraction with control
An operations specialist opens a review queue. The extractor has filled in an invoice ID, amount, and currency. Everything looks plausible, so she approves it. Only later does she notice that the invoice belongs to another organization. The language extraction was good; the tenant association was wrong. In a deployed product, this would be an integrity and privacy incident. This is a synthetic scene, not a customer story or observed failure.
It reveals a decision made in the wrong order: selecting an extractor before deciding who may see and change its output. A better prompt cannot replace a missing authorization boundary. This chapter's thesis is practical: AI-native architecture begins with testable constraints and consequences. Model autonomy is a decision inside the system. We first identify properties that must hold even when extraction is wrong. We then test whether AI adds enough value on documents that simple rules cannot handle.
The working artifact is RelayOps, a teaching product for multi-tenant document ingestion, extraction, and review. The chapter includes an architecture brief, quality scenarios, three ADRs, and an offline CLI. The starting point is zero: no deployed service, pilot, customer, validated demand, or benchmark. These files describe a proposed system and synthetic fixtures. They are not operational evidence.
We need one contract before drawing boxes. Tenant identity comes from an authenticated context controlled by the server. A document may claim an owner, but that text does not grant access. If the claim disagrees with the authenticated context, ingestion fails. A rule-based parser proposes fields; validation rejects invalid formats; every accepted result enters needs_review. Only an authorized person could approve it in a future product UI.
expected CLI context: acme
input file (one field per line):
Tenant: acme
Type: invoice
Invoice ID: INV-001
Amount: 12.00
Currency: USD
result: fields with source line and rule; status needs_review
expected CLI context: other
input file: examples/fixtures/invoice-valid.txt (declares Tenant: acme)
result: contract error: tenant mismatch; no approval
The first case uses the real fixture format and demonstrates local traceability. The second probes mismatch rejection; the CLI receives its expected tenant as an argument and authenticates nobody. Neither proves database isolation: there is no database, API, authentication, queue, or UI yet. We will maintain that distinction through the series. Evidence from a local function does not become evidence about a product merely because a diagram shows the rest of the product.
Proof must grow with the system. First, unit tests check that extract rejects a document claiming another tenant and that a valid proposal records the source line and rule for each field. Next, an API must derive tenant_id from an authorized session and prevent request text, document content, or a model proposal from overriding it. Once persistence exists, an integration test must query both organizations before and after each attempt, including retries and concurrent requests. Finally, an operator exercise should show that review appears in the right tenant and displays only the permitted original. These steps address different failures. Passing five parser tests today supports only the first claim. Recording the remaining tests prevents a correct JSON object from being mistaken for end-to-end integrity.
RelayOps would serve B2B operations staff who import documents, inspect extracted fields, make corrections, and approve. An operations manager would be a potential buyer seeking predictable throughput and bounded total cost. People and organizations named in those documents would bear the consequences of leaks, wrong associations, or excessive retention. “Potential” matters: we have not interviewed buyers or validated demand.
A software architect owns system structure and trade-offs. A staff engineer may lead cross-team decisions and their implementation. In this series, “AI-native architect” names the judgment needed when probabilistic components take part in a product: where a model may propose data, which contracts surround that proposal, how to measure task and system behavior, when a human acts, and who can stop effects. It is not a standardized job title or a promised career progression. Using AI to help design a diagram or ADR is a separate activity. That may speed up exploration, but accountability for authorization, evidence, and operations remains with people.
The useful question is not “Which agent should we build?” It is which capability benefits from interpreting ambiguity, and which boundaries must remain deterministic? Interpreting heterogeneous documents might eventually justify AI. Tenant identity, permission, schema, state transitions, spending limits, and external dispatch do not need creativity. A running non-AI baseline also gives us a comparison. If rules plus human review do the job, a model must show a net benefit before it is introduced.
Extraction produces a proposal about a document. A valid schema says that fields and types satisfy a contract. Human review says someone accepted or corrected content. None of those events alone proves that a document belongs to the right tenant, its amount matches the original, no job was lost, or its processing cost was acceptable.
Separate three layers. Input: document content and user identity have different sources. Identity comes from the server; document text is untrusted data. Proposal: rules, and perhaps AI later, produce fields with provenance. Effect: backend code validates transitions and records a reviewable item; an authorized human approves. This division creates test points. A bad proposal remains visible for correction. A cross-tenant proposal is rejected before it changes state.
The AWS Well-Architected Framework, checked on September 27, 2026, offers questions for examining security, reliability, cost, and operational trade-offs. It does not certify this design. Google SRE's SLO chapter distinguishes a measured indicator from a chosen objective. That distinction governs the numbers below: they are teaching targets, not measured results. Anthropic's Building effective agents, published December 19, 2024 and checked again on September 27, 2026, distinguishes predefined workflows from agents that direct their own process. We use it for that conceptual distinction; any API or tooling detail would need a separate current check.
“Secure and fast” is too vague to guide architecture. A quality scenario specifies a stimulus, operating environment, response, and measure. This table summarizes quality-scenarios.json, which contains the full scenarios. The values need negotiation with users and measurement under representative conditions. We have neither promised nor met an SLO.
| Attribute | Stimulus and environment | Expected response | Proposed measure and missing evidence |
|---|---|---|---|
| Isolation | Document claims tenant B in a request authenticated as A | Reject before persistence; record a reason without sensitive content | Zero cross-tenant reads or writes per adversarial attempt; real API, DB, and identity are missing |
| Time to review | 100 invoice imports/hour for one hour | Make needs_review item visible | Receipt-to-visible p95 ≤120 s across all accepted imports; unmeasured target |
| Lost jobs | Worker fails after acknowledgment during 1,000 synthetic accepted imports | Retry or expose dead-letter state | Zero jobs unaccounted for after reconciliation; no durable queue yet |
| Recovery | Worker restarts on an in-flight job with the same idempotency key | Resume with one review item | Reconcile within 15 min, zero duplicate approvals; unmeasured target |
| Extraction quality | At least 100 labeled, human-adjudicated synthetic documents | Traceable fields or explicit fallback | Per-field precision ≥0.95 on supported set and 100% unsupported fallback; unmeasured targets |
| Spending ceiling | Tenant reaches teaching cap during a synthetic month | Do not start another paid call; retain review path | USD 100 per tenant/month, adding charged processing and estimated review across all accepted documents; provisional target, neither approved nor measured |
Every measure needs a denominator, unit, time window, and method. “p95 over five hand-picked files” would have little decision value. “Zero errors” in fixtures would not prove a rate on unseen documents. Before promising an external service level, a consented pilot must observe document types, review times, corrections, spending, and incidents with minimized logs.
The context diagram identifies actors and affected parties. It represents a proposal, not a deployment.
In the container view, the synchronous request authenticates, authorizes, validates the request shape, and returns a receipt. Extraction is asynchronous: it may take longer and need a retry. Human intervention is a separate state transition, never an automatic consequence of model completion.
Boundary 1 is operator to API. Credentials and tenant context come from the server; document text has no authority. Boundary 2 is API to queue. A future job must carry a server-assigned tenant, source checksum, and idempotency key. Boundary 3 is worker to store. Its output is a proposal; backend code checks tenant, schema, version, and state before writing. Boundary 4 is review to approval. The person's permission, item version, and transition must be checked in the transaction. This chapter's CLI covers only a slice of boundary 3 in memory. Production isolation requires tests with two actual accounts, scoped database access, and redacted observability.
examples/baseline.mjs uses Node.js 22+ and built-in libraries. The example README records the checked environment and commands. There is no installation, lockfile, API key, or network because there is no external dependency. The parser accepts narrow Key: value text. It is not OCR, a PDF reader, or a general-purpose document parser. That limitation is an experimental advantage: we can observe what simple rules catch before testing whether a model adds value.
node --test examples/check-contracts.test.mjs
node examples/baseline.mjs examples/fixtures/invoice-valid.txt acme
node examples/baseline.mjs examples/fixtures/invoice-invalid.txt acme
node examples/baseline.mjs examples/fixtures/purchase-order.txt acme
These are reproducible commands, not a fabricated transcript. The chapter's verification record should contain the actual session results. The second argument stands in for an authenticated tenant supplied by a trusted caller. A real API must derive it from authentication and check authorization. Typing it into a CLI does not recreate that boundary.
A valid invoice yields needs_review, fields invoice_id, amount, and currency, plus source line and rule for each field. A malformed line yields contract error and a nonzero process status. A purchase_order may have a valid tenant and type but no supported extraction rules; it yields needs_review, fields: {}, and an “unsupported document type” issue. An explicit fallback is better than plausible unsupported fields.
{
"tenant": "acme",
"documentType": "purchase_order",
"status": "needs_review",
"fields": {},
"issues": ["unsupported document type: purchase_order"],
"provenance": {}
}
The JSON illustrates the contract for the corresponding fixture; inspect the actual output when running it. A three-letter currency code validates syntax, not whether it is valid or appropriate. Two decimal places validate amount syntax, not line-item arithmetic or an agreed price. A source line explains where text came from, not whether the document itself is genuine. A real review screen should show the original beside the proposal, allow correction, and preserve a versioned record.
We still need real authentication, safe upload, a durable queue, idempotency, a transactional store, retention policy, pilot consent, access control, log redaction, and an incident owner. We also need a labeled set that reflects permitted real documents. A passing local test suite verifies the local contract; it does not measure time to review, job loss, or model quality. Those gaps determine the next work. A diagram does not close them.
ADR-001 chooses rules for narrow invoices with human review. Alternatives are adding AI on day one or asking staff to type every field manually. Rules require maintenance by document type and fail on broader variation. Manual entry consumes operator time. AI could expand coverage, but introduces variable costs, evaluation work, data handling, and incorrect outputs. Revisit when: the unsupported share, correction time, and per-field error rate have been measured on a consented set. If rules create too much work and AI improves the task while preserving budget and integrity, test it behind the same contract. Rollback means switching AI proposals off and keeping the baseline and review queue. We have no measurement that already warrants the change.
ADR-002 favors clear API and worker module boundaries with an initially simple store rather than independent services. Services may scale components and teams separately; they also add deployment, network, distributed contracts, and operating work. A monolith may suffer contention between import and extraction. Revisit when: a representative load test measures receipt p95, queue depth, resource use, and correlated failures. Separate a service only when scale or failure isolation addresses a measured bottleneck better than a separate worker process or capacity adjustment. Reversibility calls for stable module contracts and versioned messages. No load graph exists yet.
ADR-003 assigns tenant, authorization, schema, state transitions, budget, and effects to backend code; approval belongs to a human. A fixed classify → extract → validate → review workflow is sufficient initially. An agent choosing tools dynamically makes sense only if variable tasks show measurable gains and the runner enforces step, time, permission, and cost limits. Automatic approval might reduce human labor, but it increases risk at the point of consequence. Revisit when: proposals have been compared with blinded human review and correction rates by category. Never infer system integrity from a model score. Rollback means stopping model calls and leaving items pending. Anthropic's workflow/agent distinction helps name the options; it does not prescribe autonomy for RelayOps.
Each ADR records rejected options, costs, reversibility, and a condition for reopening the choice. A decision record is useful when evidence can change a future decision. Otherwise it is decorative history.
1. A diagram with no decision criterion. Symptom: API, queue, and database boxes have arrows, but nobody knows why a queue is needed or when to split a service. Cause: treating a representation as a decision. Response: attach a quality scenario to each major component. Here, a queue would protect an acknowledged import from disappearing; it still needs an ack/retry test. If a validated load shows synchronous processing is enough, the queue can wait. Operators care about a trustworthy receipt and a visible item, not the number of boxes.
2. Selecting AI before writing requirements. Symptom: a team buys model capacity, then asks how to handle an unrecognized document. Cause: mistaking linguistic flexibility for a requirement. Response: run the baseline, label unsupported types, measure review cost and per-field errors. Introduce AI as a hypothesis for reducing measured work. It might not be worthwhile: if most traffic fits rules and review is quick, integration cost may exceed the benefit.
3. Equating model quality with system integrity. Symptom: demos show the right amount and ID while tenant or item version is wrong. Cause: assessing an answer in isolation and omitting boundaries. Response: adversarial tests for cross-tenant access, illegal transitions, duplicates, and failure after ack; human review for content; schema and provenance metrics separate from accuracy. An excellent model can feed a store that accepts a wrong key. A secure API can still show an incorrect value that needs correction. Measure both layers independently.
4. Architecture without a budget or operations owner. Symptom: the plan permits unbounded calls, with nobody owning caps, incidents, or retention. Cause: treating cost and operations as an appendix. Response: set per-tenant and per-window limits before paid calls; account for review time, storage, and investigation as well. Name who changes a limit, receives an alert, stops processing, and replays work after a failure. Without owners, a “scalable” design can scale only the bill and the error queue.
The brief maps each hypothesis to a discriminating test, owner, and unresolved risk. It also orders the work. First, submit a document whose tenant disagrees with the caller and prove local rejection. Then build authentication and a store, and repeat with two actual identities. In parallel, build a labeled synthetic set for rules and review before obtaining consented pilot documents. Measure receipt-to-visible-item latency over an explicit window. Inject failure between acknowledgment and persistence; reconcile jobs. Simulate a spending cap without real billing. Finally, test a load that could distinguish internal modularity from services.
Evaluation needs two ledgers. Task quality: eligible documents, correct fields against adjudicated labels, human correction time, and share sent to review. System integrity: cross-tenant attempts rejected, duplicate effects, jobs without terminal state, illegal transitions, and cost per tenant. A scripted stub can test a runner and contracts; it cannot measure the accuracy of a real model. Synthetic fixtures help reproduce bugs but cannot stand in for market document diversity.
A failed measure should reopen a specific decision. High unsupported share may reopen ADR-001. Bad p95 under worker contention may reopen ADR-002. Frequent correction plus high cost may keep ADR-003 restrictive. A small or biased sample yields no conclusion: enlarge or change the test instead of filling the gap with an opinion.
For cost, separate provider fees, compute, storage, and operator minutes. Model prices change; any future simulation needs a date, currency, charging unit, cache assumptions, and sensitivity to volume. Without a pilot, “cost per document” is a scenario, not data. The spending ceiling belongs in the backend: when it is reached, an item must remain available for manual review rather than triggering attempts to evade the cap.
For security, this chapter uses synthetic fixtures only. A future pilot needs explicit consent, less collected PII, redacted logs, retention and deletion decisions, scoped queries and cache keys, and purpose-bound access records. We make no compliance claim. Instructions embedded in a document are untrusted content, never authorization. Backend code must not accept a tenant selected by document text or a model-generated tool call.
For reversibility, retain the baseline, version schemas and proposals, preserve origin, and separate side effects. Turning off AI should leave a needs_review queue, not lost work. Approval and external dispatch are different steps; this chapter dispatches nothing. A future rollback requires knowing which version processed each document, who approved it, and which effects happened. The local CLI does not implement that ledger.
Write a scenario for a caller authenticated to acme who submits a document declaring Tenant: other. Specify stimulus, environment, response, and measure. Pass condition: another engineer can turn the text into a test checking rejection and no writes in either tenant.
Worked answer: “Given an API authenticated as acme and a store with two organizations, when an import declares other, the API returns a contract error before enqueueing; write counts for acme and other remain zero for that request.” The current CLI can only test local rejection. node examples/baseline.mjs examples/fixtures/invoice-valid.txt other uses the fixture declaring Tenant: acme with other as the expected context, so it fails. To prove zero writes, implement and instrument a store in a later chapter. Do not describe the absence of a database as a passing database test.
Create a synthetic fixture with Tenant: acme and Type: purchase_order, or another unknown type. Run the CLI. Confirm status: needs_review, fields: {}, and an issue naming the type. Pass condition: no invoice fields are invented and nothing is approved automatically.
Worked answer: examples/fixtures/purchase-order.txt is an existing input. Run the command in the README. The type !== 'invoice' branch returns an explicit issue. That proves parser fallback, not that an operator can resolve the document without an interface and training. To estimate value, label a batch of orders and measure manual review time.
Define a load envelope and an experiment that could show ADR-002 is wrong. Pass condition: state volume, duration, import p95, backlog, CPU/memory, correlated failures, a decision threshold, and a cheaper alternative to try before splitting services.
Worked answer: a hypothetical test might say: “Over one hour at 20 synthetic imports/s with two workers, if receipt p95 exceeds 500 ms due to extraction CPU contention despite separate pools, evaluate an independent worker service.” These are teaching assumptions, not observed RelayOps throughput. First test whether a separate process in the same deployment and a concurrency cap solve the problem. Services become compelling when repeated measurements, operating cost, and ownership favor separation. Never cite the hypothetical number as a benchmark.
Tenant: organization whose context determines data access. Provenance: traceable origin of a field, such as a line and rule. ADR: decision record with alternatives and revisit condition. SLI: measured service indicator; SLO: target value for an indicator. Idempotency: repeating an operation without duplicating its effect. Trust boundary: point where data crosses different levels of trust or permission.
This chapter leaves a design that tests can contradict. The next chapter develops identity and isolation: tenant and document enter a real transaction, and the synthetic cross-tenant attempt becomes more than a text example. The series index organizes that planned path; it does not establish publication of a future chapter. For now, the most useful decision remains runnable offline: extraction can help people without receiving approval, access, or spending authority.
An honest architecture review must also record what would change this proposal. If consented real documents fall outside the parser, the non-AI baseline remains a useful control, but its coverage must be measured by type and field. If the proposed queue fails to improve review time under representative load, simplify the design. If the financial cap blocks necessary review, operations must renegotiate the budget and manual path before enabling more paid calls. If persistence isolation fails, no extraction result justifies a pilot. The next ADR should follow that evidence, with an owner and review date, instead of keeping components because they look coherent on a diagram.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares
A microservices diagram appears before any operational measurement. The final AI-native Architect chapter uses virtual-time simulation to test per-tenant serving order, records the trade-off, and…
A RelayOps deploy cuts HTTP errors and, the same week, increases the number of wrong extractions accepted as correct. Chapter 5 builds a versioned dataset, deterministic runners, graders, a release…
A document asks, in plain text, for the agent itself to export the tenant's archive and self-promote its own role. The obvious defense — instruct the model to refuse — runs at the wrong layer.…
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.