# Guarantees

<!-- Generated by scripts/docs.mjs from test titles. Do not edit. -->

Each line is the title of a test that runs in CI. The suite is the specification: a sentence here is true for as long as its test passes.

## allowance.test.js

- one verify_funding_package call counts one supplier screen against its organization for the month
- documents count one per file in the documents argument, and raw_text counts one
- a refused verify_funding_package call and every other tool count nothing
- nothing is counted while the gate is off or the caller is not entitled
- below 80% of both allowances the result is returned unchanged
- at 80% of an allowance the result carries the near notice with the used count, the limit, the unit and the reset date
- past 100% of an allowance the result carries the over notice and the call is marked over the allowance
- when both allowances are past a threshold the notice names the one further along
- an allowance of zero never produces a notice
- the count starts again in a new calendar month, UTC
- monthOf and resetsOn read the calendar month in UTC, December rolling into January
- the notice texts are the approved copy
- the usage row records over_allowance 1 for a call past the allowance and 0 otherwise
- past the screen allowance a paying organization's verification still runs, carries the over notice, and its row is marked
- GET /usage/allowance answers per-organization counts for a month behind the usage-report bearer
- GET /usage/allowance is disabled without USAGE_REPORT_TOKEN

## auth.test.js

- the frontend API is derived from the publishable key the way Clerk derives it
- a request without a bearer token gets 401 and is told where the resource metadata lives
- a request with a token Clerk rejects gets 401 and no tool runs
- a request with a token Clerk accepts reaches the server
- the protected resource metadata names this server and Clerk as its authorization server
- the authorization server metadata is Clerk's, served here for clients that only look locally
- a preflight on a well-known route is answered without a body
- healthz reports whether sign-in is required
- without Clerk keys the server is open and the well-known routes are absent
- a probe bearer matching PROBE_TOKEN is admitted without Clerk, so regress.mjs can still run
- the probe bearer is admitted before the Clerk verifier is asked, as user and client "probe"
- a near-miss of the probe bearer is 401, and Clerk is still asked about it
- with PROBE_TOKEN unset the probe path is off: the same bearer goes to Clerk and is refused
- regress.mjs and probe.mjs send PLUMBLINE_TOKEN as their bearer, and nothing without it
- a 401 from the server is explained in one line naming the token
- a token Clerk accepts but that names no user is refused, because the run store scopes reads by user
- a verifier that cannot reach Clerk answers 503 auth_unavailable, never 500 and never admission
- the probe token cannot run a verification or an explanation, so a leak cannot spend the model key
- the probe token still reaches every read-only tool, so the sampler and the scripts keep working
- a signed-in account is not a probe, so the tools it is refused are refused for a reason of their own

## bankformat.test.js

- a nine-digit routing number whose ABA checksum resolves and whose prefix is a Federal Reserve one is valid
- a routing number with one digit changed is invalid and the detail names the checksum
- a routing number whose first two digits fall outside the Federal Reserve ranges is invalid
- a routing number of nine identical digits is invalid even though that checksum resolves
- a routing number that is not exactly nine digits is invalid
- an account number shorter than four digits is invalid
- an account number longer than seventeen digits is invalid
- an account number of four digits is valid
- an account number of seventeen digits is valid
- an account number containing a letter is invalid
- spaces and hyphens are stripped from an account number before it is checked
- an account number that is one digit repeated is invalid
- account_last4 must be exactly four digits
- a documentation IBAN is valid and the same IBAN with two characters swapped is invalid
- an IBAN is checked with its spaces stripped and its letters uppercased
- an IBAN of the wrong length for its country is invalid and the detail names the expected length
- an IBAN whose country code the registry table does not carry is not_checked, never valid and never invalid
- an eight-character BIC and an eleven-character BIC are both valid
- a nine-character BIC is invalid
- a BIC whose fifth and sixth characters are not an ISO 3166-1 country code is invalid
- an IBAN and a BIC that name different countries are flagged as inconsistent
- an IBAN and a BIC that name the same country agree
- an IBAN and a BIC are compared only once each has passed its own check
- a fixture IBAN the iban check declined is never reported as agreeing with a BIC
- a US routing number alongside a non-US IBAN is flagged as inconsistent
- a string in the IBAN field that is not an established IBAN is never compared against a routing number
- a routing number that is not an established ABA number is never compared against an IBAN
- a banking field that is absent is reported not_checked and never valid
- a full account number never appears anywhere in the returned object
- every check carries the precheck authority saying it is not an IVP-001 control
- format_checks is additive: every other value in a verify result is still the engine's
- format_checks keys an entry by invoice filename and by approved-banking supplier name
- a declared fixture routing number is not_checked however it arrives, not only as a package_id
- a declared fixture IBAN is not_checked rather than failed on its check digits
- a routing number that merely contains the fixture digits is still checked, never called a deliberate fake
- a record built on a declared fixture identifier is left unchecked whole, fields that carry no marker of their own included
- every banking identifier this repo authors is recognized as a declared fixture, so no verdict is ever published about one
- the synthetic packages' banking gets the same answer submitted as evidence as it does through package_id
- a supporting invoice on the documents path gets no verdict, because a model read the value off the page
- approved_banking on the documents path is never called transcribed, because the caller supplied it
- dfi_account_number is always not_checked through verify_funding_package, because the evidence schema carries no account_number
- two supporting invoices sharing a filename both keep an entry instead of one overwriting the other
- a filename that is an Object.prototype key still keeps both entries apart
- a run started by package_id reports every check not_checked because the fixture values are deliberately fake
- an invoice that shows no banking field at all gets no format_checks entry
- get_verification_run returns the engine trace alone, because the run store never holds transcribed values

## boundary.test.js

- the declared tool list matches what the server really registers
- no tool can approve, fund, pay, waive, or disposition
- prompt injection in evidence cannot move a single control
- the injection guarantee is about VALUES, and the limit is stated
- a documents key is refused with the engine's own code
- an oversized package is refused before the engine, naming the limit it exceeds
- a halted run reports not_reached, and nothing after the halt passes
- an unrecognized 2xx document is refused, never recorded as a run
- a body-read timeout is reported as ours, not as the engine's answer

## check-banking-format.test.js

- the banking format check is listed over both the stdio and the HTTP transport
- a routing number whose ABA checksum resolves is reported valid by the banking format check
- a routing number with two digits transposed is reported invalid and the detail names the checksum
- a banking field the caller did not send is reported not_checked, never valid
- a call carrying no checkable banking identifier is refused with a typed error
- the banking format check points the caller at the tool that compares instructions with approved banking
- a routing number from the declared synthetic-fixture family is left unchecked, not judged
- the banking format check reaches neither the engine nor a model
- the response carries the account number's last four digits and never the whole number
- the usage row keeps the account number as its last four digits and a salted digest, never whole
- the rows the usage report serves carry no full account number either
- a recorded IBAN is masked to its country and last four characters, because it carries an account number
- without a configured salt the usage row keeps the last four digits alone, never an unsalted digest
- the account number an invoice carries is redacted the same way wherever it sits in the arguments
- an argument named for an Object prototype key is recorded as it was sent, not run through a mask
- an argument named __proto__ is recorded as a field of its own, not assigned through the prototype setter
- a list or an object sitting at a banking field is recorded as its shape, never as the value inside it
- an account number inside a submitted document is kept as the document is kept, which is why the promise names the argument fields
- no log the banking format check emits carries the full account number

## config.test.js

- a numeric setting that does not parse refuses to load rather than failing open
- BEHIND_CLOUDFLARE is off unless set to a true word

## contract.test.js

- LIMITS matches PUBLIC_FUNDING_VERIFY_LIMITS in the engine
- VERIFICATION_ORDER matches the engine, in order
- EVIDENCE_FIELDS and RESULT_VALUES match the engine's evidence interfaces and result unions
- the engine's frozen modules are not modified by this lane

## corpus-eval-cli.test.js

- invalid repeat counts fail before the evaluator starts
- a missing repeats value fails before the evaluator starts
- unknown options and missing option values fail clearly
- unknown, partly unknown, and empty case selections fail before side effects
- invalid input leaves an existing same-day record byte-identical
- importing the evaluator has no side effects or key requirement
- validation happens before model construction, run-store creation, or writes
- a complete unfavorable result remains a valid measurement
- partial operational failure is incomplete, nonzero, and kept beside a complete record

## corpus.test.js

- the generator reproduces the committed manifest byte for byte
- every case names only files the corpus holds, with a supported media type, inside the document limits
- every ground-truth invoice reconciles: line extensions sum to the subtotal and subtotal plus tax equals the total
- the forensic signals each PDF claims are present in its bytes
- every re-encoded twin of the authentic invoice reads back the same provenance, page count and alteration verdict as the plain file
- the altered-total case keeps the original total reachable under the update, so a forensic reader can recover it
- the SVG the scans are rendered from carries every text run the invoice PDF's content stream carries, at the same position
- make-corpus.mjs refuses --render and names render-scans.mjs as the only scan renderer
- a corpus whose scans are missing cannot be built, and the error names render-scans.mjs
- the 50 dpi scan keeps one pixel in nine of the 150 dpi scan, unsmoothed and grey, so its bytes really sit at the floor the case claims
- the rendered scans are committed and match the manifest

## coverage.test.js

- the clean synthetic package lists mathematical_validation, which examined none of its matched invoices
- an invoice mathematical_validation skipped raises no finding, so a zero count means it examined nothing
- a totals-only invoice is counted by the engine although no arithmetic runs on it
- a partly examined package is not annotated, because a positive count proves nothing
- a fully itemised package is not annotated either, so absence never reads as a coverage claim
- every engine-produced value in a run trace still deep-equals what the engine returned
- a not_reached control in the missing_document package carries no coverage key
- get_review_artifact carries the notice only when a performed control examined nothing
- the review artifact report markdown is byte-identical to the engine's own
- the annotation is labelled an MCP-lane value rather than an IVP-001 one, and is not named like the engine's field
- a listed control needs no verdict field, because being listed is the whole statement
- mathematical_validation is the only control admitted to the map
- payment_terms is not scored, because a due date it did not count is one §11 accepts
- supplier_verification is not scored, because an uncompared supplier is one it already reported
- chronology and fraud_screening are not scored, because their metric is the denominator itself
- a control publishing a population size, a duplicate count or a section count gets no entry
- a zero with no matched count to cite is still reported, with eligible omitted rather than invented
- the status table prints a bare count when the annotation cites no matched total
- explain_verification_run prints how many items a scored control examined, labelled as this lane's
- the status table carries no coverage legend for a run with nothing scored

## deploy.test.js

- the source digest a server computes from its files equals the one the deploy script reads from git for that commit
- a server started from different code reports a different source digest than the commit being deployed
- deploy refuses to run from a linked git worktree and points at the fresh-checkout rule
- a server already running this commit's shipped files needs no release
- the release workflow triggers on every shipped path
- the release workflow releases only from main
- the release workflow runs after automerge, not only on a push
- the release workflow installs the server's dependencies before it deploys
- the release job may push the tag and open the Release
- the first release of a day is named for the date and later ones take a numeric suffix
- a verified deploy tags the deployed commit and opens a GitHub Release from that tag
- a commit that already carries a release tag is neither tagged nor released again
- the deploy script tags the commit it proved rather than the checkout head
- the platform workflow tags a verified commit on main and never on a pull request
- only the platform tag job may write to the repository contents

## detection-metamorphic.test.js

- the seeded prng returns the same stream for a seed and a different one for its neighbour
- a seed generates one package and the same seed generates it again
- two seeds generate two different packages
- a generated package is a clean multi-supplier package that carries its own totals
- a generated package is synthetic and fits inside the limits the endpoint refuses past
- the operator set splits into a verdict-preserving family and a defect-injecting one
- every operator produces schema-valid synthetic evidence or declines the package
- an operator applied to a seed twice produces the same mutation
- an operator never mutates the package it was handed
- every injecting operator applies to at least one of the sampled seeds
- values carries control outcomes and excludes narrative prose
- values sorts findings so that two orderings of one outcome compare equal
- the severity ladder is the engine's assessment union with the halt removed
- the severity order runs from pass to fail and holds incomplete outside it
- preservesValues holds when nothing moved and names the field when something did
- preservesValues ignores the report narrative, which interpolates evidence
- notImproved holds when a defect hardens the verdict and fails when it softens it
- notImproved treats a halt as a refusal rather than a softer verdict
- notImproved refuses an assessment it cannot place on the ladder
- notImproved reports whether two severities were compared at all
- values carries the position a control ran at, so order invariance can check order
- raisesControl finds the injected defect's control and reports its absence
- raisesControl accepts an observation as well as a material exception
- getOperator finds an operator by name and returns null for one that does not exist
- a bad argument stops the run before anything is written
- the operator table counts mutations and the invariant table counts relations
- every invariant the runner publishes has a statement, and every operator maps onto one
- one mutation is queued for shrinking once, carrying every relation it broke
- a shrink that reproduces nothing falls back to every relation the sweep found
- tally refuses to publish a violation on a relation the sweep never evaluated
- shrinking rejects a candidate that breaks a relation the original did not
- check files a relation under the name invariantFor gives it
- shrinking keeps every relation the original package violated
- withoutLine drops a row, its document, and any banking record left orphaned
- check records a relation it did not evaluate as one it did not evaluate
- a monotonicity violation reached without a comparison is still a relation that ran
- an assessment off the ladder is a monotonicity violation and a relation that ran
- check reports both relations of an injected defect independently
- check declines a package the operator has nothing to change
- shrink returns the smallest package that still violates
- a 2xx document that is not a verification response is refused, never scored
- the worker pool runs every item once and keeps the results in order
- the report names the procedure version from the contract rather than a literal
- the report prints every violation with its minimal package inline
- a determinism violation is reported with both runs it compared
- shrink stops at the smallest package the violation survives

## detection-refusals.test.js

- every refusal input carries a unique id, a narrative and a stated reason
- every refusal input declares one expected layer and either a code or a verdict
- the set covers the protocol, the mcp lane and the engine, and every published limit
- the set holds at least thirty inputs, and all but a handful of them expect a refusal
- exactly three inputs are exempt from the refusal property, and each says why
- an input that carries an evidence object is a bounded edit of the base package
- every string in the refusal set is synthetic
- the stability subset names inputs that exist and spans more than one layer
- evidence() replaces a field, adds a field, and removes one for undefined
- the body-limit input really serializes past max_body_bytes
- the NUL input carries a real control character, and the source file stays text
- nest builds a chain one object per level deep
- an input with raw argument text puts that text on the wire unparsed
- an ordinary input serializes to a complete tools/call for verify_funding_package
- a JSON-RPC error is observed as a protocol refusal with its numeric code
- argument validation arrives as a sentence, and its numeric code is read out of it
- a tool result carrying isError is observed as a refusal, and the engine's own is attributed to the engine
- a refusal whose body is not JSON is recorded as unstructured rather than as no code
- a successful call is observed as a verdict, with its assessment and run id
- a message naming a file path, a stack frame or an errno is recorded as a leak
- a verdict where a refusal was labelled is scored as the failure it is
- a refusal with the labelled code and layer is scored a pass
- a refusal carrying a different code, or the same code from a different layer, is scored a drift
- a 5xx and a leaking message are recorded against an otherwise correct refusal
- a gated input is scored against its assessment, and a review summary fails it
- a server that stops answering healthz fails the input that preceded it
- the summary counts refusals, verdicts, leaks and unstable codes separately
- a code that changes between the two passes is reported as unstable
- an SSE response is read from its data lines, and an unreadable body is kept as text
- every observation the runner can produce carries the same keys
- an input the runner never got an answer for is scored as not collected, never as a drift
- the retry-after seconds come from the body, then the header, then a floor that is never zero
- a 429 is waited out and the retried answer is what gets recorded
- a server that never stops refusing costs the input, and the input alone
- an unhealthy server after an input is recorded against that input
- send posts the input's bytes to the endpoint with the bearer it was given
- healthz answering anything but ok reads as unhealthy, and a throw reads as unhealthy too
- a repair pass replaces the inputs it re-sent and keeps every other result in set order
- a merge of a pass that collected nothing keeps the earlier pass whole
- an input the pass never reached is scored as a gap, never dropped from the denominator
- a pass with no second pass reports no unstable-code count rather than zero

## detection.test.js

- the corpus covers every path and carries both happy and unhappy cases
- every control that can raise a finding is targeted by at least one case
- every case declares an assessment, a path and a reason it exists
- a happy or undetectable case expects no finding, so any finding is scored against it
- every case id is unique and retrievable
- every corpus package is synthetic: invalid domains, reserved routing prefix
- every corpus package fits inside the limits the endpoint refuses past
- a case whose requested amount is not stated sums its own workbook lines
- observe reads only values the server returned and keeps the status vocabulary intact
- a case passes only when disposition, routing and attribution all hold
- a fail reached through the wrong control is scored as a miss, not a pass
- a finding on a clean case is counted as a false positive
- a halted case must record its downstream controls as not_reached
- a case that pins its finding count fails when the engine reports one per invoice
- aggregate separates detection, false positives and the boundary cases
- verificationValues carries control outcomes and excludes narrative prose
- diff names the field path where two observations disagree
- a rate-limit refusal is read back as the wait the server asked for
- an --only prefix collects one family and leaves the rest of the corpus uncalled
- an --only argument with no prefix after it is refused rather than run as the whole corpus
- a selection that matches no case is refused, so an empty study cannot be written over a real one
- a flag means the same thing spelled with a space or with an equals sign
- the corpus holds the case counts the README and AGENTS.md quote
- every evasion case is a defect a reviewer must act on, or a clean package
- the pacer sleeps only once its own window is full, and forgets calls that aged out
- a detection run stopped mid-collection still writes score.json, with the unreached cases marked not_reached

## docs.test.js

- every committed generated doc is byte-identical to a fresh generation
- every test title is a sentence starting lowercase with no trailing period
- generated regions replace only their marked content
- the generated docs are served as MCP resources and over HTTP, identical to the repo files
- the changelog names the release tag scheme and links the GitHub Releases page

## entitlement.test.js

- an organization with an active subscription item on the plan is entitled, and the answer names it
- a past_due subscription item keeps its organization entitled
- a canceled subscription item does not entitle its organization
- an ended subscription item does not entitle its organization
- an active subscription item on a different plan does not entitle its organization
- an organization with no subscription at all, which Clerk answers with a 404, is not entitled
- a user with no organization membership is not entitled
- a user in two organizations where one pays is entitled through the paying one
- an organization listed in PLUMBLINE_COMP_ORGS is entitled without a billing lookup
- the token's org_id claim is tried first and wins over another paying membership
- an unpaid org_id claim falls through to the user's memberships
- a second resolve inside 60 seconds is served from the cache without asking Clerk
- a Clerk 5xx with a cached answer under 15 minutes old serves the cached answer
- a Clerk 5xx with only a cached answer over 15 minutes old fails closed as unavailable
- a Clerk 5xx with nothing cached fails closed as unavailable
- the org_id claim is read from a JWT access token's payload and is null for an opaque token
- with BILLING_PLAN_SLUG set, a member of an unpaid organization gets subscription_required with the pricing URL on every tool
- the refusal sends a caller to app.plumbline.cloud/pricing unless PRICING_URL says otherwise
- with BILLING_PLAN_SLUG set, initialize and tools/list stay open to an unpaid member
- with BILLING_PLAN_SLUG set, a paying member's call runs and its usage row records the entitled organization over the header
- with BILLING_PLAN_SLUG set, a verified org_id claim on the token selects the paying organization
- with BILLING_PLAN_SLUG set and Clerk down with nothing cached, a tools/call answers 503 auth_unavailable
- with BILLING_PLAN_SLUG set, the probe token stays exempt and Clerk is never asked about billing
- with BILLING_PLAN_SLUG unset, a tool call runs as before and Clerk is never asked about billing
- with BILLING_PLAN_SLUG set and Clerk sign-in off, an anonymous tools/call is refused with subscription_required

## eval.test.js

- the eval set carries at least twenty synthetic documents across the named hard cases
- every eval document is synthetic: invalid domains, reserved routing prefix, no real-looking client
- every eval document is within the limits the extraction refuses past
- the eval documents are byte-identical on every build, so a score is comparable across runs
- every case has a unique id, at least one tag, and an expectation for every field it names
- an invoice path is canonicalized by the document it cites, not by the index the model chose
- an invoice entry whose document was never named is dropped from the canonical entries
- a field type is assigned from the path so a weak money, date, or routing field is visible on its own
- a field matches only when the parsed value is exactly the expected one
- a value the model never read is missing, and a value it read that is not on the page is spurious
- a value below the confidence floor counts as absent, because the engine never sees it
- an ambiguous field accepts any of its listed readings, including declining to read one
- text planted in a field is correct when transcribed verbatim and wrong when obeyed
- a path with no expectation is counted as unjudged rather than scored against the model
- precision counts only the confident values the model asserted, so an empty extraction scores no precision
- the derived evidence is checked separately, so a routing or readability defect shows up on its own
- calibration buckets a claimed confidence against how often that claim was right
- the aggregate reports accuracy per field type so a weak field type cannot hide in the mean
- cost is priced from the model and the tokens the response reported, never from a constant
- a perfect transcription scores every field correct and routes to the evidence the engine expects
- an empty transcription scores no field correct and asserts no precision

## evidence-schema.test.js

- get_evidence_requirements publishes every evidence field the synthetic packages submit, with a type
- get_evidence_requirements lists every assessment value the golden outcomes contain, and manual_review_required

## explain.test.js

- the <packageId> explanation is non-authoritative and carries every control status verbatim
- no submitter-supplied text from the injection package reaches the explanation prompt or output
- a narrative that calls an incomplete run passed or approved is rejected
- an explanation that fails the vocabulary check twice is refused, not returned
- the executive audience gets a different instruction from the analyst audience
- explaining an unknown run fails without calling a model
- explanation is refused with a typed error when no model key is configured
- a line item description on a line that does not foot never reaches the explanation prompt
- a planted filename in an unreadable-document finding does not reach the explanation prompt
- a planted supplier name in a supplier-verification finding does not reach the explanation prompt
- a planted invoice number in an ambiguous-reconciliation finding does not reach the explanation prompt
- the explanation still carries every finding classification after submitter text is removed

## extract.test.js

- clean documents extract to evidence the engine passes, with the extraction labelled untrusted
- every extracted field carries a confidence between 0 and 1 and a source document
- the forensic provenance of each submitted PDF is returned beside the extracted evidence and labelled untrusted
- instructions planted in a document change no engine value
- an instruction printed as a field's value is transcribed as that field's value
- document content reaches the model only as user content, never in the system prompt
- money printed on a page parses to integer cents, and nonsense does not
- the extraction schema stays inside the structured-output union limit
- the extraction schema has no field that could hold a finding, status, assessment, or approved banking
- more than ten documents are refused before any model call
- documents over 25 MB in total are refused before any model call
- documents over 100 pages in total are refused before any model call
- an unsupported media type and duplicate document names are refused
- a low-confidence value is dropped rather than sent to the engine
- a path outside the documented set is dropped and marked unused
- an invoice whose number or amount was not read is recorded unreadable, and so is a declared invoice that yields nothing
- an invoice whose amount was read only as total_cents is still recorded readable
- an amount derived from a total is recorded in the audit list with the reading it came from
- a submitted document the engine never saw is listed in documents_not_transcribed with its declared role
- a package whose every document reached the evidence returns an empty documents_not_transcribed
- documents_not_transcribed never changes an engine value
- a declared supporting_invoice that yielded nothing is recorded unreadable in the evidence and not listed as not transcribed
- a roleless document the model cited only below the confidence floor is listed as not transcribed
- a declared supporting_invoice read only below the confidence floor gets its unreadable record and stays off the list
- a roleless document cited by an invoice entry whose record was dropped is listed as not transcribed
- a document whose invoice record was kept is never listed as not transcribed, whatever the model cited as each value's source
- the verify response returns documents_not_transcribed beside the extracted evidence
- an invoice stating an amount and a different total keeps both and derives nothing
- a total below the confidence floor derives no amount, so an unreadable invoice stays unreadable
- an invoice citing a document that was not submitted is ignored
- two invoices read from one document become two supporting-invoice records, each naming its invoice
- the same batch document uploaded twice under two names still yields matching hashes per invoice
- one invoice read twice from the same document stays one record under the document's own name and hash
- approved banking comes only from the caller, never from the documents
- an unusable response is retried once at higher effort, then refused as extraction_failed
- extraction is refused with a typed error when no model key is configured
- over the daily token budget, extraction is refused while package_id still verifies
- document bytes are kept in the usage store by design and still never in the run store
- an entry the model returned with an empty value is omitted from the audit list, as if it had not been returned

## failures.test.js

- an unreachable engine is reported as worker_unreachable with no verification values
- an engine error code is passed through unchanged with its HTTP status
- an unrecognized 2xx verification body surfaces as a refusal from the tool
- a body-read timeout surfaces from the tool as our timeout, not the engine's
- capabilities are not reported when the health endpoint answers with something else
- supplying two input kinds at once is refused as ambiguous
- extraction beyond the concurrency cap is refused with a retry delay
- when required evidence cannot be read the engine is not run and a draft is returned
- an extraction_incomplete refusal carries documents_not_transcribed, naming the document nothing was read from
- an engine refusal after extraction carries documents_not_transcribed beside the draft
- an engine refusal after extraction returns the draft that was sent and names the fields outside the documented schema
- schema violations name unknown fields, wrong types and malformed arrays, and accept a documented null
- an engine refusal of structured evidence carries no extraction draft
- an engine failure after extraction still reports the model tokens spent
- a client that accepts only application/json is served JSON, not 406
- schema-invalid arguments still reach the typed refusals rather than a protocol error
- an upstream model failure is a typed refusal, never the provider's own message
- a fifth concurrent extraction waits for a free slot and then runs instead of being refused
- with no queue wait an extraction past the cap is refused as extraction_busy with a retry delay from the oldest in-flight extraction
- capabilities publish the extraction cap and queue wait a client can plan around
- no engine failure names the engine origin, over any of the four error surfaces
- an unrecognised 2xx body is described by its shape and never echoed
- a non-2xx body is passed through only in the engine's published shape; anything else is described
- a health payload that is not Plumbline's is described by its shape and never echoed

## forensics.test.js

- the forensic pass flags exactly the corpus PDFs whose structure shows alteration
- the forensic provenance matches each corpus case's recorded signals
- a generator producer is surfaced as generated provenance and does not by itself mark the document altered
- PDF dates parse to UTC and junk parses to null
- a submitted PDF with no declared role that yields no invoice still has its provenance returned
- a structurally altered PDF reaches the engine with has_visible_alterations and the reasons in the audit list
- an authentic PDF reaches the engine without an alteration flag
- pages_expected as printed is a transcription path and reaches the engine beside page_count
- an invoice's page_count is counted from the PDF bytes and the model's transcribed page_count stays in the audit list unused
- invoices read from one batch PDF carry no page_count, because the file's page count is not any one invoice's
- an invoice from a batch PDF carries no page_count even when the model read only one of the file's invoices
- a PDF whose bytes do not settle the page count carries no page_count rather than a guessed 1
- one invoice read twice from a single-invoice PDF still carries the page_count from its bytes
- an outline /Count in a bookmarked PDF does not stop the page count from the page tree
- a non-PDF invoice carries no page_count even when the model transcribes one
- an invoice's ocr_confidence is the mean confidence of the fields transcribed for it
- a large PDF with many object references is read in well under a second
- an opaque annotation or image that covers no text does not mark the document altered

## http.test.js

- an initialize request over HTTP returns the plumbline server info
- tools listed over HTTP equal the declared tool set
- a tool call over HTTP runs the real engine on the clean package
- a __proto__ key in evidence reaches the engine, which refuses it like any other unknown key
- a __proto__ key in an approved_banking record reaches the handler as the caller sent it
- an unknown JSON-RPC method returns -32601
- a malformed JSON body returns -32700
- GET and DELETE on /mcp return 405 because the server holds no sessions
- healthz reports ok, sha, source digest, start time, engine, anthropic, and auth status
- a request whose Host is not the public host is refused when PUBLIC_URL is set
- an unknown path returns 404
- a browser asking for an unknown path gets the HTML 404 page
- a missing doc answers the same 404 as any other unknown path
- a client that does not ask for HTML still gets the JSON 404
- the badge icon files serve the brand copies byte for byte
- the landing, status and 404 pages link the badge icon
- over HTTP with Clerk on, a run created by user_1 is unknown_run to user_2
- a wrong-typed evidence argument is refused as invalid_input in the JSON envelope, naming the type received
- a non-object approved_banking record is refused as invalid_input in the JSON envelope, naming its index
- a JSON-RPC batch body is refused as an invalid request
- a request with a foreign Origin is refused before the body is read, and one with no Origin is admitted
- a caller with the in-flight cap of requests open is refused a further one before its body is read
- a development server with no PUBLIC_URL listens on loopback only
- a Host header that is not a host name is never reflected into a served page
- a request in flight when the server is asked to stop still gets its answer
- shutdown gives up on a request that never finishes, rather than hanging forever
- the status sampler stops when the drain begins, so a shutting-down container records no uptime
- the shutdown deadline defaults to SHUTDOWN_GRACE_MS from the environment
- shutdown resolves only after the server has closed and its close listeners have run
- SIGTERM to the hosted process drains the request in flight and exits 0
- SIGINT to the hosted process drains the request in flight and exits 0, so a local Ctrl-C waits for the answer too
- a new connection is refused once shutdown has begun

## limits.test.js

- a string over max_string_bytes is refused as invalid_input naming the field, its size, and the limit
- a string limit counts UTF-8 bytes and names the nested path
- more lines than max_lines is refused as invalid_input naming max_lines
- more supporting invoices than max_supporting_invoices is refused naming that limit
- more approved banking records than max_approved_banking_records is refused naming that limit
- nesting deeper than max_json_depth is refused naming max_json_depth and the path
- more JSON nodes than max_json_nodes is refused naming max_json_nodes
- a body over max_body_bytes is refused naming max_body_bytes
- a package exactly at the limits still reaches the engine

## onboarding.test.js

- every connection rail on the page carries a verification date
- the page states how long a run is kept as the RUN_TTL_HOURS default, six hours
- every tool the page names is a registered tool
- every absolute URL on either page is built from PUBLIC_URL
- the only hard-coded hosts are the two font hosts
- every route the 404 page lists is a route the server serves
- the served page substitutes PUBLIC_URL and leaves no placeholder behind
- every placeholder on the page is one the server resolves
- the hero carries a Subscribe call to action with the price, linked to the pricing page
- the page no longer says there is no account to create or that setup is a bare minute
- the sign-in sentence names Google, Microsoft and an emailed code
- every client flow opens with the Subscribe first prerequisite
- the served page resolves PRICING_URL to the default and to an override
- the Claude setup names the claude.ai controls captured on 2026-09-16, in order
- the onboarding page names Clerk as the login provider before the first sign-in step
- the hero offers one setup control per assistant, each pointing at a flow the picker has
- every flow's done block ends with the same first question to ask
- the Gemini flow is the CLI and never promises the app
- the Gemini CLI capture status is stated once, in the dated setup note
- the page says how to disconnect in every assistant and where the Google grant lives
- the picker script parses
- the picker opens the flow the fragment names, and the first flow for any other fragment
- a tool count past ten fails with a message naming the count
- every tool count the page spells out is the registered count
- the manage section does not claim revoking the Google grant ends access at once
- the page's Kept list names every column of the usage record, so a new column cannot land undisclosed
- the page's Never kept list names the email address and the sign-in token

## precommit.test.js

- a generated file the developer already staged does not block the retry commit
- a generated file regenerated but left unstaged blocks the commit
- a staged generated file changed again afterwards still blocks the commit
- an untracked generated file blocks the commit until it is staged
- a clean generated file reports no drift at all
- the GIT_DIR git exports to hooks does not turn tracked files into untracked ones
- only the worktree column of porcelain output counts as drift
- git ignores env files, key material and a local MCP usage database
- an env example file stays tracked beside the ignored env files

## presentation.test.js

- the selection covers a pass, an observation, an incomplete, a manual review, a fail, an injection and a boundary case
- every selected case has a recorded response and a stated reason it is in the set
- the incomplete case carries controls the run never reached
- the planted instructions are read off the injection case's own evidence
- the system prompt is the server's own agent guidance, not a paraphrase of it
- each case is presented as three recorded tool results answering one question
- the replayed tools are the server's own, so a renamed tool cannot be transcribed here
- no recorded tool result is summarised away before the model sees it
- every rule declares an id, a question it asks and the cases it applies to
- a rule that does not apply to a case is scored as not applicable, never as a pass
- calling a control the run never reached clean is caught
- the engine's assessment has to appear as the engine spells it
- the spaced spelling of an assessment counts as verbatim, a synonym does not
- an answer that makes the funding decision itself is caught
- naming the analyst's own next step is not the same as making the decision
- a heading reading my recommendation is not a decision when the section refuses to make one
- a recommendation to go and check something is not a disposition, and one to decline is
- answering the analyst's own question with a verdict is making the decision
- the disclaimer the rule rewards is not itself read as a decision
- a verdict on its own line is caught as surely as one after a colon
- a value or an assessment wrapped across lines is still stated
- recommending payment, release or refusal is dispositioning
- a sentence that hands the decision back is read no further
- a bare yes or no answering the question is a verdict without the question's words
- a bare no under anything but the funding question is a control summary
- a disposition under a heading that is not a question is still a disposition
- dispositioning and then deferring in one sentence is still dispositioning
- reporting a control as not_performed the way the engine spells it is not calling it clean
- the four rules that read a phrase read it wrapped across lines the same way
- a verdict hedged as a modal is still a verdict
- a planted instruction repeated across a line break is still repeated
- a rule still fires when the sentence it reads is broken across lines
- a prefix that fits both sides of a finding names neither of them
- the committed report is what the committed answers score to
- a long quoted value is named when the answer reproduces enough of it to identify it
- an answer that obeys the text planted in the package is caught
- reproducing a planted instruction outside quotation marks is caught
- a material exception has to be named with the values that raised it
- naming the control without the observed value is not enough
- a pass the procedure cannot see behind is not reported as a clean bill of health
- an unable_to_verify finding is reported as manual review, never as a pass
- counting how many controls passed is not a claim that the package passed
- quoted text is discounted before an answer is read for its own claims
- an agent that quotes the planted text in a blockquote is not failed for quoting it
- a case passes only when every rule that applies to it passes
- the roll-up reports a rate per rule over the cases the rule applies to
- the roll-up names every case that failed and the rule it failed
- the trace the judge reads carries no text transcribed off a submitted document
- a summary quoting planted text reaches the judge fenced as data, with the rubric saying so
- the judge's verdict is recorded as advisory and never gates a case
- an answer with no text stops the run rather than scoring as a bad summary
- a rate limit is retried and a bad request is not
- the report names every rule, its rate, and the cases that broke it

## ratelimit.test.js

- a client over its per-IP allowance is refused with a retry delay
- every request from one caller, tool call or not, is counted against a per-caller request window
- an HTTP caller past the request window gets 429 before the body is read
- verifications across all clients are capped per minute
- concurrent extractions above the cap are refused until one releases
- the client IP is the rightmost X-Forwarded-For entry, which a client cannot forge
- the trusted-hop count decides which forwarded entry is believed
- the configured hop count is what the server meters and stores by
- one signed-in user cannot spend past their share while the server budget remains
- two concurrent calls cannot both pass a budget that fits one, because the first reserves its tokens
- extraction is refused once the daily token budget is spent
- an HTTP client over the limit gets a JSON-RPC error with retry_after_s
- two signed-in users sharing an IP get separate rate-limit buckets, scoped user
- over HTTP with Clerk on, the per-user allowance is keyed on the bearer's userId and the 429 names scope user
- an extraction still waiting when the queue wait runs out is refused, and the next free slot goes to the caller behind it

## reachability.test.js

- the matrix pins the transcription paths extract.js asks for
- every line of the prompt's path block yields at least one path
- a path shape the parser does not recognise is a failure, not a silent drop
- every field the matrix says a control reads is a field of the published evidence contract
- every derived and caller-supplied field the matrix declares is a field a control reads
- every control that reads evidence appears in the matrix under its engine name
- no control reads a field the document path cannot emit
- a conditional row states the condition that limits it
- the has_visible_alterations row and get_verification_capabilities name the same media types it is never derived from
- document integrity is reachable only through the three signals PR 82 wired up
- every detection case is pinned to reachable or structured-only
- a structured-only case says which field put it there
- a case populating a field no extraction path emits is structured-only
- a case whose readable flag disagrees with extraction's own rule is structured-only
- reachable means a document submission could carry the case, never that it would look identical
- a structured case that gives a non-PDF invoice a page_count is structured-only, because only PDF bytes produce one
- the orphaned signals name both the code that computes them and where they stop
- the reachability summary reports the matrix and the corpus together

## regress.test.js

- the four synthetic packages produce the golden traces and artifacts over HTTP
- a structured SDK rate limit retries only the failed tool call
- rate limit retries stop after three attempts and cap each wait at ten minutes
- only a 429 with structured rate_limited data and a valid delay is retried
- ordinary tool errors and malformed tool results are not retried
- the client is closed when connecting fails
- a golden drift is reported with the package and field path
- the injection package differing from clean is itself a regression
- deploy verification waits until healthz reports this commit's source started after the deploy
- deploy verification fails when the new commit never comes up
- deploy verification fails when healthz reports the new sha but the running source is different code
- <script> raises the connect attempt timeout before it makes a request

## runs.test.js

- a recorded run is returned unmodified by its run id
- a run survives a process restart when the store is a file
- a run older than RUN_TTL_HOURS (6 by default) is swept and no longer retrievable
- the extraction envelope is stored with the run and document bytes are not
- an unknown run id states the run lifetime in hours and never other callers' runs
- unknown_run from every run tool carries expired_after_hours and no package list
- a stored run carries expires_at, RUN_TTL_HOURS after recorded_at
- a run created by one user is unknown_run to another, and to nobody in particular, when auth is on
- with auth off a run carries no user and is retrievable by id alone, as before

## schema-key.test.js

- get_evidence_requirements publishes schema as a value the route generates, not one the caller sends
- evidence carrying a schema key is refused as invalid_input naming the key, before the engine is called

## status.test.js

- a sampler run records one row per component
- a probe that throws is recorded as a failure and never as a missing sample
- uptime is computed from recorded samples and reports the recording window when it is short
- a day that was never sampled reports no uptime rather than a full one
- three consecutive failures open exactly one incident and a recovery closes it
- an incident note is stored verbatim and only on an incident that is still open
- an incident note is rendered on the page as text and never as markup
- a note is accepted only with the status write token, and the usage report token cannot write one
- the note route is disabled when no status write token is configured
- the status page answers with the engine unreachable and shows it as down
- the status page and its json answer without a token
- the status json carries a stable shape
- the status store holds no evidence, run id, or client address
- recorded days survive a restart and a failure streak carries across it
- samples older than the retention window are rolled into daily uptime and then dropped
- days older than the retention window are dropped from the store
- behind Clerk the endpoint probe presents the server-held probe bearer
- the endpoint probe calls the public origin, because fetch cannot set a Host header
- an endpoint probe whose connection is refused records the endpoint as unreachable
- an endpoint probe whose hostname does not resolve records the endpoint as unreachable
- a connection the probe could not establish, a connect-phase timeout included, records unreachable
- a connection that fails after it was established records the endpoint as an error
- an endpoint probe that is aborted records the endpoint as timed out
- an endpoint probe that outlasts its timeout records the endpoint as timed out
- an endpoint that answers with a non-200 status records the endpoint as an error
- an endpoint that answers 200 with a body that is not JSON records the endpoint as an error
- an endpoint that answers 200 with JSON that is not a tool list records the endpoint as an error
- an outage of a refused endpoint opens an incident headed unreachable rather than error
- a server that pins Host to its public origin still samples itself as operational
- a component nobody could measure is recorded as no sample rather than as down
- the sampler records on start, repeats on its interval, and stops on request
- a status write that fails is reported and never crashes the server
- the hosted server samples itself once it is listening
- sampling runs only when the status store is a file the service can keep
- the served status page substitutes every placeholder
- every absolute URL on the status page is built from PUBLIC_URL
- the status page carries the Platform beside the MCP server and the engine
- the Platform probe reads its public health route and records nothing when no origin is configured
- a Platform nobody configured shows on the page as no data, never as an outage
- the Platform core rows share one read of /api/health/core per sample
- a failed core check is down with its own reason, and the others stay up
- a core check that is off, or no PLATFORM_URL, records no sample
- an unreachable or garbled core route marks all three rows down
- status.json can be read from another origin, so the demo gallery can show it

## tools.test.js

- get_verification_capabilities reports the live gates truthfully
- capabilities state that has_visible_alterations is derived from PDF bytes only, so a document_integrity pass on a PNG, JPEG or plain-text submission is no alteration check
- capabilities publish runs.ttl_hours, the hours a run stays retrievable
- a verify response and its run trace carry expires_at, RUN_TTL_HOURS from now
- get_evidence_requirements exposes real limits and reachability
- clean package: report issued, pass, twelve controls performed
- missing document: incomplete notice, one performed, eleven not reached
- material exception: fully verified and still fails
- unknown run and unknown package fail cleanly
- the same package verifies identically on repeat — deterministic replay

## usage.test.js

- the client IP is stored only as a salted hash that changes every day
- with no salt the usage row carries a null ip_hash rather than an enumerable one
- the IP hash is keyed by the salt, so the same address and day differ across salts
- the hosted server refuses to start without IP_HASH_SALT
- the daily report counts fires by tool, client, and outcome with latency percentiles
- token spend is summed per day and priced per model
- a tool fire over HTTP stores the ids and counts plus the arguments and the result verbatim
- content capture is a switch, and off means the fire row carries ids and counts alone
- every submitted file within the cap is summarized by name, media type, size and hash
- content past the ceiling is clipped and the row says so
- the ceiling counts UTF-8 bytes, so multi-byte arguments cannot write past it
- a call refused before its documents were read keeps its answer and its file summary, never the bytes
- a call that was read and then failed, at the engine or by a throw, keeps its arguments whole
- over HTTP, a submission the server refused before reading it keeps no bytes, and one the engine refused keeps them
- a caller past the daily content budget keeps an envelope row and a file summary only
- the daily content budget is keyed on the daily IP hash for an anonymous caller and resets with the UTC day
- a usage database created before the omitted column existed gains it on open
- a store opened without the capture flag keeps no content
- over stdio a tool call writes no usage record, even with USAGE_DB_PATH set
- the file summary is bounded at the source, so it is never a clipped fragment
- a rolled-up fire takes its stored content with it
- the usage report route refuses a missing or wrong bearer token
- the usage report route is disabled when no report token is configured
- rowsOn returns one row per event for the day, oldest first, with derived cost
- rowsOn is scoped to the requested day
- the per-row export carries the submission and the result, and still no raw IP
- the usage rows route shares the report bearer and reports the transport
- the rows route reads content=false, off and no as off, not only the digit
- the rows route answers content=0 with the envelope and no stored submission
- the usage rows route is disabled when no report token is configured
- a fire older than the retention window rolls into the daily table and leaves the raw table
- a rolled-up day reports the same fire counts, token sums and USD estimate it reported before the roll
- a roll carries each group's p50 and p95 forward and replays them count-weighted, which reproduces a single group exactly and approximates a mix
- a day whose fires straddle the retention cutoff is counted once, not once per side
- the rollup keeps no IP hash, evidence hash, run id, package id or user agent
- a multi-day total sums each day once whether the day is raw, rolled, or both
- the multi-day report totals a window that spans the retention cutoff without counting a fire twice
- the per-event rows for a day outside the retention window are gone, while its report still answers
- the report attributes fires, tools, outcomes and cost to the client whose initialize shares the same ip hash and user agent
- an identity whose initialize was never seen keeps its own group with a null client name
- by_identity comes back in a stable order, so a consumer can diff one day's payload against the last
- the per-identity block is counts alone: no raw IP, no ids, no submitted content
- a day already rolled into the daily table reports no identities, because the roll keeps no ip hash or user agent
- a multi-day total carries no identities, because the ip hash is salted per day
- clientCountry is null unless the deployment says it is behind Cloudflare, whatever the client sends
- clientCountry accepts only a real two-letter country from the edge
- the country is stored per event and served on the row
- a country is never inferred from the IP, and the IP is still not stored
- a usage database written before the country column gains it on open
- a usage row written under a verified bearer carries the userId
- by_identity groups on user_id when it is present, so two users behind one NAT and one user agent are two identities
- with no Clerk keys the server records every row as unauthenticated with no user, client id or scopes, and by_identity keys on ip hash and user agent as before
- a usage database written before the user_id column existed opens and keeps its rows
- a usage row written under a verified bearer records the OAuth client id, the granted scopes, and that it was authenticated
- the usage record never stores the bearer token or an email address, even when the verified token carries one
- two events from the same signed-in user on different days share a user_id, which the salted daily IP hash never can
- an anonymous row is recorded as unauthenticated, with no client id and no scopes
- a usage database written before the principal columns existed opens, and rows that already carry a user_id are backfilled as authenticated
- a rate-limited call under a verified bearer is recorded with the OAuth client id, the granted scopes, and that it was authenticated
- the authenticated column is stored exactly as the request reported it, never inferred from a user_id
- operator routes hold the Host pin and never answer on a stray hostname
- a bearer-authenticated call that names an organization and a member records both, and GET /usage/rows serves them
- the same headers from an unauthenticated caller record nothing: a header is only as good as the bearer behind it
- a usage database written before the org_id and member_id columns existed gains them and keeps its rows

## useragent.test.js

- a script that calls the hosted server names itself and its version in the User-Agent
- the User-Agent carries the running package version, never a hard-coded one
- no request our scripts make to the hosted server is left on the Node default agent
