More services is not a maturity criterion
Someone sketches, at a RelayOps architecture review, a diagram with three new boxes: ingestion-service, extraction-service, review-service, each with its own deploy, its own repository, its own queue. The argument is "this is how serious systems scale" - the diagram is clean, the arrows are straight, and nobody on the committee has a number that contradicts it. The right response is neither to accept the diagram nor to reject it on principle. It is to ask: which measured bottleneck does this diagram solve? Nobody on the committee can answer, because nobody has measured anything yet.
examples/scale/run.mjs simulates three deterministic scenarios with a virtual clock and invented stage costs. Under its instant burst, four simulated workers, and FIFO order, modeled queue wait exceeds modeled processing time. In the dominant-tenant fixture, two small tenants wait more than twice as long because their jobs arrive in the input after the dominant tenant's burst. That is a property of FIFO under this fixture, not a measured bottleneck in deployed RelayOps. A paired scheduler test is warranted; infrastructure scaling needs operational evidence still missing.
This final chapter revisits chapter 7's open question: is "one service, one worker pool, one queue" still appropriate? The simulation tests serving order, not actual capacity. Fair scheduling reduces the small tenants' virtual wait in the same fixture and increases the dominant tenant's wait. The ADR records a local, reversible choice; its service-separation trigger depends on future pilot measurements.
Chapter 7 closed out security and resilience: deterministic authorization guards that never read the document as a source of authority, a circuit breaker that never discards a job, a schema-compatibility check for rollback, and a local restore drill that measures recovery instead of assuming it. This chapter does not reopen any of those files. It assumes RelayOps already knows how to defend against a hostile document and a provider failure, and asks a different question: does the SHAPE of the system - one process, one shared queue - stay right as volume grows, as one tenant dominates the shared queue, or as document types vary enough that "one pipeline fits all" stops being true? The only thing this chapter reuses from earlier ones is the shape of contracts already fixed - the durable queue's queued/leased vocabulary from chapter 4, the reserve-before-spend discipline from chapter 6, the per-stratum gate shape from chapter 5 - never a literally imported file.
Before any full scenario, the distinction fits in a comparison of two functions with the same signature:
// Before: blind to tenants, order = arrival order in the shared queue
function leaseNextNaive(queue, now) {
const job = queue.shift();
if (!job) return null;
job.status = 'leased';
job.leasedAt = now;
return job;
}
// After: every tenant with a non-empty queue gets visited, in rotation,
// before any tenant gets a second turn - Deficit Round Robin
function leaseNextFair(scheduler, now) {
return scheduler.leaseNext(now); // examples/src/tenant-scheduler.mjs
}
The first function is what any shared queue does by default: shift() on an array, blind to whose job it is. The second visits tenants in rotation and only grants a turn to whoever has a non-empty queue and enough credit - a tenant that enqueued 200 jobs cannot, by that alone, take the turn of a tenant that enqueued 5. Neither function changes how long a document takes to process once it is picked up; both change only WHEN each tenant gets served. It is exactly that distinction - capacity versus order - that the rest of this chapter measures, compares, and decides on.
Primary sources for this section: Google SRE — Handling Overload and Anthropic — Demystifying evals for AI agents. The analogy to RelayOps scheduling is this chapter's interpretation.
The Google SRE book describes overload protection based on a purely local utilization signal: as utilization approaches a configured threshold, the system starts rejecting requests by criticality, and an adaptive-throttling mechanism where each client tracks its own recent acceptance rate and self-regulates - because, in the book's own words, "it's almost equally expensive to reject a request... as it is to accept and run it" for some services, so moving the decision to the client side keeps the server from spending resources just to say no. This chapter's fair scheduler uses exactly that shape: a local deficit counter per tenant, no central coordinator, no timer, no background process - the same "local signal before a shared resource" discipline chapter 7's circuit breaker already applied to the provider call, now applied to the order in which tenants get served.
The second intuition comes from a different place: Anthropic's guide to evaluating agents notes that teams delay building evals thinking they need hundreds of tasks, when "20-50 simple tasks drawn from real failures is a great start" - because, early on, the effect of a change tends to be large, and a large effect does not need a large sample to show up; more mature teams measuring subtler effects need bigger samples. This is exactly the logic examples/scale/scenarios.json applies by declaring, in the file itself, that each small tenant's 11 jobs in the dominant-tenant scenario are enough to show an order-of-magnitude change (13x faster) but not enough for a reliable p99 of that tenant alone - the small sample is not hidden, it is sized to the effect it needs to prove.
RelayOps gains five new pieces in this chapter, independent of any chapter 1-7 file:
examples/scale/scenarios.json, examples/scale/run.mjs): three synthetic scenarios with declared seeds and sample windows. Legacy warmupDocs excludes initial input IDs from the sample; those jobs still execute, with no worker prewarm phase. Five invented stage costs produce model p50/p95/p99 values; host hardware does not determine them.examples/scale/real-trial.mjs): four worker threads execute synthetic CPU work. One Apple M3 Pro / Node v26.8.1 run showed less wait for both small tenants; the dominant tenant improved in that run despite regressing in the virtual model. Results vary between runs and do not establish production capacity.examples/src/tenant-scheduler.mjs): createNaiveFifoScheduler (the "before") and createFairScheduler (the "after," Deficit Round Robin) - this chapter's one minimal, discriminating evolution, measured on the same scenario, same seed, before and after.examples/scale/fitness-functions.mjs): five checks - tenancy, contracts, eval gate, budget, recovery - each with its own comment stating what passing it does NOT prove about production.examples/adrs/006-evolution.md): simulated scheduling behavior, alternatives with explicit operational costs, a local decision, a planned expand/contract migration, and a review trigger requiring real metrics.examples/portfolio/): context/container/sequence diagrams, the ADR log for the whole series, a claim-by-claim evidence index (demonstrated, simulated, pending pilot), and a pilot plan explicitly marked pending.Portfolio navigation: architecture review, evidence index, pilot plan, and ADR-006. Prior chapters' traces, costs, evals, and runbooks need checking in their own chapters; this index neither reproduces nor cross-validates them.
None of these pieces calls a network, database, real provider, or Docker container. The 22 tests cover local contracts; suite wall time says nothing about RelayOps capacity.
This chapter separates capacity (workers and throughput) from serving order (who gets the next slot). With finite workers and FIFO, an earlier burst can delay later tenants. Adding workers changes absolute wait but not the selection rule; with enough workers for every job, queue wait would vanish. The relative effect depends on workload and needs measurement. The fair scheduler conserves jobs under leaseNext(), pinned by does not lose or duplicate a single job across a full drain. weight also has its own test (Weight is proportional, not cosmetic): an early bug made different weights alternate 1:1; the current implementation grants slots roughly in proportion to weight.
Symptom: a diagram with new services shows up before any measured bottleneck - the justification is "this is how serious systems scale," never "we measured X and Y would fix it."
Cause: conflating an architectural shape (how many services, how many agents) with a measure of maturity. Neither is the same thing as solved capacity or reduced risk - and the shape is easy to draw before any evidence exists, because it requires measuring nothing.
Response: examples/adrs/006-evolution.md compares alternatives and selects fair scheduling within the current process because it is reversible and improves serving order in the simulation. Service separation remains deferred because operational measurements are missing.
Symptom: a number generated from invented stage costs gets cited as real-provider latency or the trial gets called a "cloud benchmark."
Cause: the stub exists to measure the calling code's CONTRACT and BEHAVIOR - the order jobs get served in, whether any get lost, whether fairness works - not to measure how long a real model takes to respond. The two numbers look comparable because both are called "latency," but they measure completely different things.
Response: DOC_TYPE_PROFILES feeds INVENTED costs to a virtual clock. "Invoice" costs less than "contract" only in this fixture. examples/portfolio/pilot-plan.md requires per-stage percentiles under real providers and traffic before identifying a RelayOps bottleneck. Queue dominance here follows from the model's instant burst.
Symptom: a simulation number is cited without its seed, sample window, denominator, or virtual-millisecond unit.
Cause: "13x faster" without a unit or context suggests empirical improvement that this model cannot establish.
Response: examples/scale/scenarios.json fixes seed, excluded initial input IDs (legacy key warmupDocs), and window. verification.md records paired results for the same jobs: acme-small (n=11) 976.003→72.433 virtual ms (13.47x); initech-small (n=11) 1027.072→77.703 (13.22x); globex-dominant (n=198) 469.164→569.398 (+21.36%). sampleSizeLimitation rules out a stable p99 claim for either small tenant. Host hardware neither explains these values nor validates real capacity.
Symptom: an architecture portfolio shows only the approved final result - the scheduler that worked, the ADR that got accepted - without recording what was tried and rejected, the bug found along the way, or the cost the accepted decision actually imposes on someone.
Cause: a portfolio that only shows success optimizes for looking successful, not for being auditable - and a reader who only sees the final version cannot tell whether the decision was made well or just presented well.
Response: examples/portfolio/evidence-index.md separates test-demonstrated contracts, simulated results, a local synthetic measurement, and missing pilot evidence. ADR-006 records the dominant tenant's +21.36% virtual-wait regression; the recorded local trial did not repeat it. allowedRegressionTenantIds names the fixture gate's exception. The earlier weight bug remains documented in scheduler code.
The 22 tests in examples/scale/fairness.test.mjs check that FIFO delays a small tenant until lease grant 51 after 50 dominant jobs; fair scheduling grants it on call 2; neither scheduler loses or duplicates jobs; 2:1 weights produce roughly proportional service; and the fitness functions catch injected violations. New tests check paired jobs in the virtual model and conservation in the local synthetic CPU trial.
What these tests do NOT prove: the virtual simulation uses one process; real-trial.mjs uses worker threads for synthetic CPU work, without multi-node isolation or a real pipeline. No test measures representative contention between tenants or restores in-flight leases after restart. checkRecoveryFitness checks only counts after JSON serialization. The ADR's operational trigger remains pending.
This chapter's decision is cheap to adopt and cheap to reverse: createFairScheduler is a new data structure inside the same process, with no new deploy, no new network boundary, no production configuration key to migrate - reverting to createNaiveFifoScheduler is swapping a function call, not an infrastructure rollback. That cheap reversibility is exactly why the ADR accepts the change now and defers service separation: separation, if it ever happens, is expensive to undo (a new network boundary, a versioned contract between two deploys, in-flight jobs crossing a boundary), and that is why the ADR already writes the expand/contract plan and the rollback BEFORE separation becomes necessary, not after - the same "document before you need it under incident pressure" discipline chapter 7's rollback runbook already applied to a code revert.
On security, the scheduler decides serving order, never authorization. checkTenancyFitness checks IDs in synthetic state; authorization stays in chapter 7's authorizeToolCall. In the paired simulation, the dominant tenant pays another 100.234 virtual ms per job on average. This is a model trade-off accepted for testing fairness; real operational cost still needs a pilot.
node --test examples/scale/fairness.test.mjs passes locally before any change to the scheduler, the trial, or the fitness functions.Basic — identify the bottleneck from the per-stage report. Verifiable criterion: run node scale/run.mjs --scenario higher-volume and confirm queueMs (p50, p95, or p99) is at least 10 times larger than any of the other four stages (parseMs, extractMs, validateMs, reviewMs) in the same sample window. Worked solution: the test 'under the higher-volume scenario, queue wait dominates...' in fairness.test.mjs already pins this comparison with a 10x threshold, over the same scenario configuration - running the command should not produce a different result than what the test already proves, since both use the same seed and the same simulation logic.
Intermediate — apply fairness and prove the small tenant makes progress. Verifiable criterion: using createFairScheduler with a dominant tenant (50 jobs enqueued first) and a small tenant (1 job enqueued after), the small tenant's job must be granted by, at most, the 2nd call to leaseNext() - never after the 51st, which is what the naive queue would produce on the identical scenario. Worked solution: the two corresponding tests in fairness.test.mjs ('a dominant tenant burst of 50 jobs delays a small tenant's 1 job until lease call #51' for the naive queue and 'the same 50-dominant-then-1-small burst gives the small tenant its lease on call #2...' for the fair one) run exactly this scenario side by side, proving both halves of the comparison against the identical job set.
Advanced — write a separation ADR with a trigger, a migration/rollback, and the additional evidence required before production. Verifiable criterion: the ADR must name at least three measurable (not subjective) conditions that, together, would justify reopening the decision to separate the service, an expand/contract plan that runs both paths (single process and new service) in parallel before either is turned off, and a rollback that preserves reconciliation of in-flight jobs. Worked solution: examples/adrs/006-evolution.md already contains this full structure - the three conditions in "Review trigger" (a defined and exceeded operational floor for queueMs, CPU/memory contention that is measured rather than assumed, more than one consented tenant with a contractual isolation requirement), the six-step expand/contract migration, and a rollback that reuses chapter 7's reconciliation discipline (findJobsNeedingReconciliation, cited as a conceptual reference, never re-imported) instead of inventing a new one.
This is the last chapter of the series. There is no scoped chapter 9. The next step is a consolidated review of the portfolio — six recorded ADRs, eight example sets, and the gaps declared by each chapter — followed by the still-pending pilot in examples/portfolio/pilot-plan.md. The examples demonstrate local contracts for human approval, isolation, idempotency, evaluation, budgets, and authorization. This chapter adds a comparison of tenant serving order under synthetic load. None of these artifacts demonstrates RelayOps performance with real tenants, providers, or traffic; the pilot must produce that evidence.
Produtos gratuitos e pagos para transformar ideias em uma base que você consegue executar.
13 produtos disponíveisContinue explorando tópicos similares
Start AI-native architecture with testable constraints. Build an offline baseline, measurable quality scenarios, and reversible decisions for RelayOps.
A RelayOps deploy cuts HTTP errors and, the same week, increases the number of wrong extractions accepted as correct. Chapter 5 builds a versioned dataset, deterministic runners, graders, a release…
A document asks, in plain text, for the agent itself to export the tenant's archive and self-promote its own role. The obvious defense — instruct the model to refuse — runs at the wrong layer.…
Checklist de 47 pontos para encontrar bugs, riscos de segurança e problemas de performance antes do lançamento.
Templates testados em produção, usados por desenvolvedores. Economize semanas de setup no seu próximo projeto.