Testing
Binding for implementation. This page defines the test layers, where each kind of test lives, the MIME
conformance corpus, property tests and fuzzing, the integration harness against a local workerd with its
fakes, fault injection and time control, the cross-tenant attack suite, the search and triage quality
gates, the live end-to-end suite against staging, coverage, how every edge-case row maps to a test, and
the CI checks. Rust workspace defines the xtask commands and the base CI
pipeline; this page adds what they must contain.
| Requirements | PRD release criteria 1–6, NFR-QUAL-1, NFR-QUAL-2, NFR-QUAL-3, NFR-SEC-1, NFR-SEC-2, NFR-OPS-1 |
| Edge cases | Every row marked S or S+I in the edge-case register |
| Code | crates/core/src/** (#[cfg(test)] modules), crates/conformance/, crates/worker/tests/it/, crates/worker/tests/live/, fuzz/, xtask/ |
1. Principles
- A change ships with a test that fails without it (AGENTS.md definition of done).
- Test where the rule lives. Pure rules are tested natively in
core. Orchestration is tested natively againstplatform::fakes. Platform behaviour and wiring are tested against workerd. Only what needs real mail providers runs live. - Deterministic by default. Clock, randomness, DNS, models and vector search are fakes behind platform traits, so a failing test fails the same way every time.
- No real data. Fixtures use RFC 2606 names (
example.com,example.net,example.org,*.example,agents.example) and.invalidfor the simulator; no real people, addresses or messages (CONTRIBUTING.md). - Names are contracts. Test names in the edge-case register exist verbatim in the code, and
cargo xtask tracefails when one is missing (section 11).
2. Test layers
┌─────────────┐ live:: staging, real Gmail/Outlook/SES nightly, release
┌─┴─────────────┴─┐ eval real Workers AI + Vectorize nightly, release
┌─┴─────────────────┴─┐ it:: local workerd (wrangler dev), fakes every PR
┌─┴─────────────────────┴─┐ worker logic on platform::fakes (native) every PR
┌─┴─────────────────────────┴─┐ conf:: MIME corpus (native) every PR
┌─┴─────────────────────────────┴─┐ core:: unit + property tests, fuzz smoke every PR
└─────────────────────────────────┘
| Layer | Prefix | Location | Runs with | Covers |
|---|---|---|---|---|
| Unit | core:: | #[cfg(test)] modules in crates/core | cargo test --workspace | Every pure rule: parsing, caps, sanitising, classification, verdicts, threading, tokens, addresses, query parser, fusion, citation verifier, triage rules, policy, DNS parsing, domain state machine, SSRF classification, fencing, crypto envelope, receipt builder, SLO rules, JWK thumbprints (core::jwk::), JWT signing (core::jwt::), HTTP message signature bases (core::httpsig::), notification rendering (core::notify::) |
| Property | core:: | same modules, proptest | cargo test --workspace | Section 4 |
| Worker logic | worker::, platform:: | #[cfg(test)] modules in crates/worker and crates/platform | cargo test --workspace | Handlers, Durable Object logic modules (mailbox::Mailbox<P> and the others) over platform::fakes with rusqlite standing in for Durable Object SQLite (Rust workspace) |
| Conformance | conf:: | crates/conformance | cargo test --workspace | The MIME corpus (section 5) |
| CLI | cli:: | #[cfg(test)] modules and tests/ in crates/cli | cargo test --workspace | Configuration, output, setup (including pmail setup ses) and deploy against recorded Cloudflare and AWS API fakes (CLI and setup) |
| Fuzz | target name | fuzz/ | cargo xtask fuzz | Section 8 |
| Integration | it:: | crates/worker/tests/it/ | cargo xtask itest | Section 6 |
| Attack suite | it::security:: | crates/worker/tests/it/security/ | cargo xtask itest | Section 7 |
| Evaluation | – | crates/conformance/golden/, xtask | cargo xtask eval-search, eval-agentic, eval-triage | Section 9 |
| Live | live:: | crates/worker/tests/live/ | cargo xtask live | Section 10 |
3. Unit and worker-logic tests
coretakes time, randomness and lookups as arguments (Design conventions), so its tests pass fixed values:now_ms = 1_791_540_000_000, fixed 32-byte keys, fixed random bytes.corebuilds forwasm32-unknown-unknowntoo (CIwasmjob); tests run natively.- Worker-logic tests build a
Platformbundle fromplatform::fakes: an in-memory clock that tests advance, a seeded RNG,rusqlitefor D1 and Durable Object SQLite (with FTS5), maps for R2, captured queue sends, scripted AI and Vectorize, a DNS zone map, a recording HTTP client and a recording mail sender. They exercise transactions, outbox writes, state machines and policy without workerd. - Error-code tests compare
ErrorCode::http_status()andretryable()with every row of Errors (core::errors::catalogue_matches_reference), and everyErrorCodevariant with the catalogue in both directions. - Snapshot-style assertions compare JSON structurally (
serde_json::Value), never as strings.
4. Property tests
proptest (pin at build time), native only. Each property runs 1,024 cases in CI and 65,536 nightly
(PROPTEST_CASES).
| Test | Property |
|---|---|
core::query::f1_* | For any input string: parsing never panics; it returns a typed tree or invalid_query with a position; the FTS5 expression built from any tree quotes every term, contains no bare FTS5 operator, column filter or NEAR, and never contains the raw input (F1); parse(print(tree)) == tree |
core::address::a1_case_and_dots (property part) | Normalisation is idempotent, case-insensitive on the local part, keeps dots, converts the domain to an A-label; random Unicode local parts are refused with address_unsupported (A1, A3) |
core::thread_token::a2_round_trip | Mint then verify returns Valid for random identities and sequence numbers; every single-bit flip returns Invalid or Absent; a token for one identity never verifies for another (Threading) |
core::refs::f5_* (property part) | Plate normalisation: AB12CDE, ab12 cde and AB12 CDE normalise equally; normalisation is idempotent (F5) |
core::ssrf::refuses_private_ranges | Every address inside each blocked range, including IPv4-mapped, NAT64 and 6to4 embeddings, is refused; addresses outside them pass (Security) |
core::injection::e1_* (fence property) | After core::injection::fence escaping, no content, including content containing the nonce or runs of < and >, can terminate a MAIL_CONTENT fence (Search) |
core::citations::f11_* (property part) | A sentence survives verification only if every cited ID is in the evidence set and every quoted phrase occurs in the cited source after normalisation (F11) |
core::keys::format_round_trip | Generated keys match the key regex; parsing rejects every other shape |
core::mime::b2_caps (property part) | Random nesting and part counts never exceed depth 32 or 500 parts in the parsed tree and never panic (B2) |
5. MIME conformance corpus (conf::)
5.1 Layout
crates/conformance/
corpus/
mime/ b2_*, b4_*, b5_*, b6_*, b7_*, b8_*, b9_*, b13_* structure, charsets, TNEF, nesting
auth/ d1_*, d9_* DKIM, ARC, DMARC, forged Authentication-Results
dsn/ d4_*, g6_* DSNs, MDNs, auto-replies
threading/ c1_*, c2_*, c7_*, c8_* headers and subjects
sanitize/ b7_*, b11_*, e1_* remote content, hidden text, injection text
attachments/ b10_*, b12_* risky types, archive bombs, extraction inputs
dns/ zone files (TOML) for the fake resolvers: DKIM keys, DMARC and SPF records
keys/ DKIM signing keys generated for tests only (file names end in .test-only.pem)
golden/ the evaluation set (section 9)
src/ loader, expectation checker, generators
THIRD_PARTY.md origin and licence of every imported fixture
Each case is a pair: <case>.eml and <case>.toml.
# crates/conformance/corpus/mime/b5_shift_jis_subject.toml
id = "b5_shift_jis_subject"
edge = ["B5"]
source = "authored" # authored | generated:<generator> | derived:<origin> | imported:<project>
licence = "FSL-1.1-ALv2"
[envelope]
from = "sender@example.net"
to = "bookings.acme@agents.example"
[expect]
subject = "ご予約の確認"
flags = []
text_contains = ["予約番号 BK-2291"]
attachments = 0
kind = "normal"
verdict = "none" # with the zone fixtures in dns/
5.2 Sources and licensing
| Source | Rule |
|---|---|
| Authored | Written for this repository, under the repository’s licence (FSL-1.1-ALv2). The default |
| Generated | Produced at test time by a seeded generator in crates/conformance/src/gen/ (large messages, deep nesting, part floods, the 25 MiB message for spike S4). Not committed when larger than 1 MiB |
| Derived | Built from published standards examples (RFC example messages), with every address rewritten to RFC 2606 names; the origin is named in source |
| Imported | Test fixtures from open-source projects under Apache-2.0 or MIT (for example the mail-parser test suite), each listed in THIRD_PARTY.md with its upstream path and licence |
Never: real mail dumps, archives of public mailing lists, or anything containing a real person’s address. DKIM-signed fixtures are signed with the test-only keys and verified against the zone fixtures, so signatures are reproducible.
5.3 Runners
- Native (
cargo test -p pylota-mail-conformance): parse each case withcore, compare with[expect]. Test names areconf::<dir>::<id>, generated from the files, so the register’s wildcards (conf::mime::b2_*) match every case with that prefix. - Through workerd (
it::conformance::corpus_via_workerd): inject every case through the local email endpoint and compare the Message object returned by the API with the native expectations. This catches differences between native and wasm builds and gives the verdict-parity check of spike S4.
6. Integration tests against workerd (it::)
6.1 What cargo xtask itest does
As defined in Rust workspace, plus the details below:
- Build the Worker with
worker-build --releaseand theitest-hookscargo feature. - Render
deploy/wrangler.itest.toml: the production bindings, local resources,PM_ENV = "local",PM_PLATFORM_DOMAIN = "agents.example",PM_API_HOST = "localhost",PM_CONSOLE_HOST = "console.localhost"(the test client sends thatHostheader on console paths),PM_SIGNUP = "open",PM_WEB_BOT_AUTH = "on"(local only: the S13 gate applies to real deployments), the SES variables (PM_SES_REGION = "eu-west-2",PM_SES_INBOUND_*) and the Google, GitHub and Stripe client settings naming resources on the fake server, random test secrets written to.dev.varsin a temporary directory,PM_ITEST_FAKES_URL = "http://127.0.0.1:8798", a randomPM_ITEST_TOKEN, queue consumers withmax_batch_timeout = 1, and noAIorVECTORSbinding (both are served by fakes). Tests inject provider events through the productionQ_DELIVERYproducer binding, which exists for dead-letter redrive. - Apply D1 migrations:
npx --yes wrangler@4.139.0 d1 migrations apply pylota-mail --local --persist-to target/itest/state --config deploy/wrangler.itest.toml(a fresh directory per run). The rendered file setsmigrations_dir = "../migrations/d1", because Wrangler resolves it against the file’s own directory,deploy/. Until v1.0 there is one file,0001_init.sql. - Start
npx --yes wrangler@4.139.0 dev --local --port 8799 --persist-to target/itest/state --test-scheduled --config deploy/wrangler.itest.toml, write its PID totarget/itest/wrangler.pid, and capture stdout and stderr totarget/itest/worker.log. Wait forGET /health. - Seed: insert the platform domain row and one platform key directly with
wrangler d1 execute --local(the harness knowsPM_KEY_PEPPERbecause it generated it). Every other fixture (tenants, identities, domains, keys) is created through the public API. - Run
cargo test -p pylota-mail-worker --features itest-hooks --test it -- --test-threads=1withPM_ITEST_URL=http://127.0.0.1:8799. Theittest target declaresrequired-features = ["itest-hooks"], socargo test --workspacenever builds it. - Stop wrangler; delete
target/itest/stateunless--keepwas passed.
6.2 Test hooks
Compiled only with itest-hooks, honoured only when PM_ENV = "local", and refused unless the request
carries x-pm-test-token: <PM_ITEST_TOKEN>. cargo xtask build-worker refuses the feature and fails if
the release bundle contains /__test/.
| Hook | Does |
|---|---|
| Invocation sync | At the start of every fetch, email, queue, scheduled, alarm and Durable Object request, read GET {fakes}/state (clock offset and fault-plan version) into isolate state |
POST /__test/inbound | Runs the email() handler code with a synthetic message (mail_from, rcpt_to, raw_base64) and returns { "outcome": "accepted" | "rejected" | "tempfail", "smtp": "550 5.1.1 …" }. Used for cases the local endpoint cannot carry (no Message-ID, B3) and to observe reject and temporary-failure outcomes (A6, J1) |
POST /__test/alarm | { "class": "mailbox" | "domain" | "job" | "quota" | "notifier", "object_id" }: runs the object’s alarm handler now, executing every purpose due at the fake clock |
GET /__test/routes | The router table (method, pattern, permissions, scope, idempotency) for the attack suite |
POST /__test/rpc | Sends a raw RpcEnvelope to an object, for owner-mismatch tests |
POST /__test/delivery-event | Publishes a provider event payload to pm-delivery-events through Q_DELIVERY |
POST /__test/mailbox-schema | Sets an object’s meta.schema_version back by one, for J9 |
6.3 Fakes
All external services are served by one fake server inside the test process
(crates/worker/tests/it/fakes/, a blocking HTTP server on 127.0.0.1:8798; the HTTP server crate is
pinned at build time). With itest-hooks, platform::itest provides implementations of the platform
traits that call it:
| Trait | Fake behaviour |
|---|---|
Dns | Zone maps per resolver (First, Second), mutable by tests; per-resolver errors and disagreement (H1, H7); seeded from crates/conformance/dns/ |
Ai::run (embeddings) | Feature hashing of normalised tokens into 1,024 dimensions, L2-normalised: deterministic and similarity-preserving enough for hybrid tests |
Ai::run (rerank) | Cosine similarity of the same vectors |
Ai::run (triage) | Schema-valid output looked up by raw_sha256 from the labelled set, a default rule-based output otherwise; scriptable invalid JSON and timeouts (FR-TRI-4) |
Ai::run (planner) | Scripted tool-call sequences per question from crates/conformance/golden/questions.toml, including a hostile script that tries to widen scope (F10) |
Ai::to_markdown | Text from a sidecar fixture, or a scripted failure or timeout (B12) |
VectorIndex | In-memory namespaces with metadata filters (equality and sent_at ranges), mutation IDs, a configurable processing lag and processedUpToDatetime; scriptable failures (F14) and a “keep one vector” mode for probe tests (F6) |
MailSender (live tenants) | Records every StructuredEmail; scripted outcomes: accepted with a messageId, a coded error (E_RATE_LIMIT_EXCEEDED, E_DAILY_LIMIT_EXCEEDED, E_HEADER_NOT_ALLOWED, E_RECIPIENT_SUPPRESSED, …), an exception, or a timeout |
HttpClient | Routes requests by host to fake handlers: Cloudflare API (zones, including zone creation with scriptable error 1105 and zone-hold refusals; routing rules with the 200-rule limit; sending subdomains including preview_enabled; event subscriptions), SES (SendEmail, which checks the SigV4 signature against test credentials; email identities with scriptable DKIM and MAIL FROM status; the account’s sending status; receipt rules with the 200-rule and 500-recipient caps), S3 (GetObject and DeleteObject on the inbound bucket, with SigV4 checks and scriptable NoSuchKey), SQS (ReceiveMessage and DeleteMessage on the backstop queue), SNS certificates, Google and GitHub OAuth (token, user and email endpoints with scriptable claims and unverified addresses), Stripe (Checkout Sessions and the objects billing reads), RDAP, the scanner, and webhook receivers. Any other host gets HttpError::Connect: integration tests never reach the internet |
| SNS push | The fake server signs SES notifications with a test key (SignatureVersion 2, or 1 and tampered variants on request) and POSTs them to the Worker’s /hooks/ses/inbound and /hooks/ses. A test can skip the push and leave the notification only in the SQS fake, for the backstop cron (N1–N3) |
TCP sockets (the platform wrapper over connect()) | Routes by host name to a scripted SMTP server in the fake process on ports 465 and 587. Scripts can omit STARTTLS, answer 535, refuse some RCPT TO with 4xx or 5xx, close the connection after the final ., or exceed each timeout. Accepted messages can be handed to the platform domain’s inbound path, unchanged or with a rewritten From or a foreign DKIM d=, for alignment probes and DSNs. TLS is simulated; certificate checking is proved by spike S12 and live, not here (N14–N20) |
- Test tenants use the real simulator and loopback code (
*@simulator.invalid, L2, L3); only live tenants use the fake mail sender. The real Cloudflare transport is exercised by spike S1 and the live suite: the localsend_emailsimulation cannot serialise binary attachments (Cloudflare Email Service local-development docs, read 2026-10-09). - Webhook receiver fakes record each request (headers, body, signature check result) and can return any status, delay past the 15 s timeout, redirect, or stream an oversized body.
- The SSRF guard runs unchanged: the DNS fake answers webhook and SMTP relay hosts with a fixed public address that is never contacted, because the fake HTTP client and the socket fake route by host name.
- The groups that use these fakes:
it::ses::*(SNS push, SQS, S3 and SES fakes),it::smtp::*(the SMTP server fake),it::forwarding::*(the mail sender fake plus inbound injection at the platform address),it::domains::*(DNS, Cloudflare API and SES fakes),it::oauth::*(OAuth fakes and one cookie jar per simulated browser),it::checkout::*andit::signup::*(Stripe fake),it::totp::*,it::landing::*,it::onboarding::*andit::abuse::*(fake clock),it::hosts::*(the twoHostvalues),it::identity_keys::*,it::assertions::*,it::http_signatures::*andit::well_known::*(fake clock for overlap windows and expiry; the Rust SDK’sverify_assertionruns natively in the test process against the JWKS that workerd serves), andit::notify::*(fake clock for holds, windows, the 09:00 run, time zones and cooldowns;/__test/alarmwith classnotifier; notification emails are observed the same way as console sign-in mail, which the system identity also sends (Console › Requesting a link or code);/__test/delivery-eventfor a hard bounce on a notification; the DNS fake to make the platform domainfailing). - Another value of a deployment variable or secret. A test that needs one (
PM_CONSOLE=offforit::console::disabled,PM_WEB_BOT_AUTH=offforit::http_signatures::disabled_and_policy,PM_BILLING=offforit::notify::billing_off_quotas,PM_NOTIFICATIONS=off, or the secretPM_MASTER_KEY_NEXTfor the master-key rotation tests) callsrestart_runtime_with(&[(name, value)]), which restarts wrangler likerestart_runtime()(section 6.6) with the value overridden in the renderedwrangler.itest.tomlor.dev.vars, and restores the original on exit.
6.4 Injecting inbound mail
The default path is the endpoint wrangler dev provides for email handlers (Cloudflare docs, read
2026-10-09): POST http://127.0.0.1:8799/cdn-cgi/local/email?from=<envelope from>&to=<envelope to> with
the raw RFC 5322 message as the body; the message must have a Message-ID header. One call per envelope
recipient, as Email Routing invokes the handler once per recipient (A9). The
documentation does not say how a setReject is reported to the caller, so tests read the outcome from
/__test/inbound or from the inbound_rejected log line; S1 records the endpoint’s actual response.
6.5 Time control
- Platform clock.
platform::itest::ClockreturnsDate.now() + offset, with the offset set byPOST {fakes}/clock { "advance_ms" }or{ "set_ms" }and synchronised at each invocation (6.2). - Alarms. Objects arm alarms at absolute times computed from the fake clock, which may be far in the
real future; tests run them with
/__test/alarm. Purposes not yet due at the fake clock do not run. - Cron.
GET /cdn-cgi/local/scheduled?cron=<expression>&time=<ms>triggersscheduled()with that cron andscheduledTime(Cloudflare docs, read 2026-10-09);timeis the fake clock. - Queue delays.
platform::itestproducers andIncoming::retryrecord the requested delay at the fake server (GET {fakes}/queue-log) and send withmin(requested, 1)second. Retry-schedule tests (J4, G3, G8) assert the recorded delays. - Waiting. Asynchronous effects are awaited by polling the API with backoff (50 ms doubling to 1 s) for at most 20 s; a test never sleeps a fixed time.
6.6 Fault injection
POST {fakes}/faults arms a fault plan; decorators in platform::itest consult it before each call:
{ "target": "r2.put", "match": { "key_prefix": "t/" }, "mode": "error", "count": 3 }
| Target | Modes | Used by |
|---|---|---|
r2.put, r2.get, r2.delete, r2.list | error, timeout | J1, erasure retries |
d1.query (with match.sql_prefix) | error | J7 (directory lookup), job retries |
do.call | error, timeout | Partial tenant search (F15) |
transport.send | the mail sender’s scripted outcomes | G2, G3, G10 |
ai.run, ai.to_markdown | error, timeout, invalid_output | F12, B12, FR-TRI-4 |
vectorize.upsert, vectorize.query, vectorize.delete | error, lag | F14, F6 |
doh.query | error, answer (per resolver) | H7 |
http.send (by host) | error, status, delay, redirect | Webhooks, SSRF tests, SES, S3, SQS, OAuth and Stripe failures |
tcp.connect (by host) | error, timeout | SMTP relay connection failures (502 upstream_error at create; RetryLater on a send) |
Runtime restarts. For J2, the harness helper restart_runtime() kills the
wrangler process from target/itest/wrangler.pid mid-test and starts it again with the same
--persist-to directory, then asserts the queue retry produced exactly one stored message.
6.7 Isolation and logs
- Each test creates its own tenant (
slug = "t" + 10 hex of the test name's hash), so tests do not see each other’s data. Tests that change global state (clock, fault plans, platform domain records) reset it in a guard on exit;--test-threads=1keeps them serial. - The I5 log-scrubbing test (I5) runs last: it reads
target/itest/worker.logand fails if any canary string appears. Every fixture plants canaries: a unique token in each body, subject, display name, filename and attachment text, and every address used by the suite, plus every key and webhook secret the suite created.
7. Cross-tenant attack suite
Proves NFR-SEC-1 (zero cross-tenant access) and FR-KEY-3. Rules are in Security › Authorisation.
Fixture. Two tenants, A (victim) and B (attacker), each with two identities, a domain, a webhook, a
key of each level holding every permission valid at that level, threads with messages and attachments,
a held thread, an erasure request, an export, an identity signing key on each identity (one rotated, so a
retiring key exists too) and policy.web_bot_auth.allowed = true. Tenant A’s mail contains a unique
canary term. A third tenant C is a test tenant.
“Every permission valid at that level” follows Security §4.6:
a tenant key holds every permission except tenants:manage and platform:ops, so it holds
identities:sign; an identity key holds the same set without the tenant-only permissions
(members:read, members:manage, suppressions:manage, audit:read, usage:read), plus
usage:read implicitly for its own workspace.
Attacker key classes (each with full permissions for its level):
| Class | Key |
|---|---|
foreign_tenant | Tenant key of B (with identities:sign) |
foreign_identity | Identity key of B’s first identity (with identities:sign for that identity) |
sibling_identity | Identity key of A’s second identity, attacking A’s first identity |
mode_mismatch | Test-mode key of C, attacking live tenant A (L4) |
revoked, expired | A’s own tenant key, revoked or expired |
Matrix. it::security::cross_tenant_matrix reads GET /__test/routes and, for every route with a
path parameter and every attacker class, calls the route with A’s resource IDs (and, for POST/PATCH,
a valid body). The enumeration includes the identity-key routes (…/keys, …/keys/rotate,
…/keys/{kid}/revoke with A’s kid), POST …/assertions and POST …/http-signatures on A’s
identities; the scope check answers before any signing rule, so they give the same
404 identity_not_found as a missing identity. For each call it also makes a control call with the same key and a random non-existent
ID of the same type. It asserts:
- The status is
404with the route’s*_not_foundcode, or403 scope_deniedfor a route above the key’s level on its own tenant, or401for revoked and expired keys. - The attack response equals the control response byte for byte, except
request_idand theRequest-IdandRateLimit-*headers (indistinguishability). - No side effect: D1 row counts for A, A’s mailbox state (via a platform key), A’s outbox and A’s
webhook receiver are unchanged, and no
audit_logrow names A. - The same matrix runs over MCP: every tool, with the same attacker keys, returns an error and no data.
Additional suites.
| Test | Attack |
|---|---|
it::security::route_table_complete | Every route in /__test/routes appears in the matrix, has a scope rule and a non-empty permission list (except Scope::Public, GET /v1/me and GET /v1/tenants/{tenant_id}, which carries foreign_permissions instead; GET /v1/usage needs usage:read, which tenant and identity keys hold implicitly); a route added without them fails this test. The Scope::Public set is exactly the list of Security §4.7, the two /.well-known/ key routes included |
it::security::body_scope_ignored | For every POST, PATCH and list route: tenant_id, identity_id and identity_ids naming A in bodies and query strings, sent with B’s keys |
it::security::search_canary_isolation | B searches for A’s canary in every mode, including agentic with a question that asks for “all tenants”; zero hits and no evidence from A |
it::security::vector_foreign_id_dropped | The Vectorize fake returns one of A’s vector IDs to B’s semantic query; the mailbox read-back drops it and rpc_owner_mismatch_total does not move (the ID is simply not found in B’s mailbox) |
it::security::rpc_owner_mismatch | /__test/rpc sends an envelope with B’s IDs to A’s mailbox; internal_error, rpc_owner_mismatch logged, metric incremented, alert fired |
it::inbound::a2_forged_token_ignored | Mail to B’s address with a token minted for A’s thread files into B’s mailbox only |
it::security::webhook_filter_scope | B creating a webhook with identity_ids of A gets 404 identity_not_found |
it::security::mcp_tools_follow_key | Tools listed and callable only with their permission; mail_sign_assertion and mail_sign_http_request are never listed to a platform key |
it::identity_keys::paused_withdraws_jwks, it::assertions::erasure_tombstones_kid | Without a key: a paused identity’s JWKS answers the same 404 identity_not_found as an unknown ID; an erased identity’s kid is never published again (O1, O7) |
it::notify::one_click_unsubscribe | An unsubscribe token for a person of B, altered to name A’s workspace or another kind, changes nothing and gets the same page as an expired token (O18) |
The suite is part of cargo xtask itest and therefore a required check on every pull request. Timing is
not asserted in CI (too noisy); both code paths do the same D1 read by construction.
8. Fuzzing
The fuzz project lives in fuzz/ (cargo-fuzz, libFuzzer, nightly toolchain only) with the targets listed
in Rust workspace. The five required by the edge-case work:
| Target | Invariants checked beyond “no panic” |
|---|---|
mime_parse | Depth ≤ 32 and parts ≤ 500 in the output; every output string is valid UTF-8; time per input under 1 s |
query_parse | The FTS5 expression from any successful parse quotes every term; errors carry a position inside the input |
address_parse | Normalisation is idempotent; a validated username matches ^[a-z0-9][a-z0-9._-]{0,23}$ |
sanitize | Output contains no <script, no on*= attribute, no remote src; sanitize(sanitize(x)) == sanitize(x); derived text contains none of the hidden-text code points |
dsn_parse | The classification is one of the defined kinds; a DSN’s recipients are syntactically valid addresses or absent |
- Seeds come from
crates/conformance/corpus/. - CI runs the five for 60 seconds each on every pull request (
fuzz-smoke); nightly runs every target for 10 minutes. - A crash is minimised (
cargo fuzz tmin), committed as a regression input undercrates/core/tests/fuzz_regressions/<target>/, and replayed bycore::fuzz_regressions::replay_allin the normal test run. - Before a release every target must have run clean for 24 cumulative hours on the release commit (Security).
9. Search and triage evaluation
The golden set, the labelled queries, the agentic questions, the triage labels and the metric definitions are owned by Search › Quality evaluation and Triage › Evaluation set. This section defines how the harness runs them.
9.1 Files
crates/conformance/golden/
generator.toml seed and template mix; the mailbox (about 5,000 messages in four identities of
tenant `acme`) is generated deterministically at run time and never committed
hard_cases/ hand-written .eml files added to the generated mailbox
queries.toml labelled queries with graded relevance per message key
questions.toml agentic questions with gold facts and gold supporting message keys, including
unanswerable and steering questions
triage.toml labelled triage set
Message keys are stable names (brightwell_invoice_88213) mapped to message IDs at load time. All content
is synthetic, on reserved domains, under the repository’s licence (FSL-1.1-ALv2). Accepted scores (the baseline) are recorded in
docs/src/project/quality.md, as the search and triage designs specify.
9.2 Running
cargo xtask eval-search, eval-agentic and eval-triage:
- Start
wrangler devwithout--local, with theAIbinding (always remote) and the Vectorize binding set toremote = trueagainst a dedicated indexpm-mail-chunks-evalin the CI Cloudflare account; D1, R2, Durable Objects and queues stay local. If the pinned Wrangler cannot bind Vectorize remotely, the Worker uses the Vectorize REST fallback withPM_CF_API_TOKEN(Rust workspace). - Generate the golden mailbox and inject it through the local email endpoint (section 6.4); wait until
semantic_coverage = 1.0for every identity. - Run every query, question or labelled message through the public API (triage through
POST …/messages/{id}/triage); writetarget/eval/<suite>.jsonwith per-item results and totals. - Compare with the baseline in
quality.mdand fail on a gate.
9.3 Metrics and gates
| Suite | Gate | Source |
|---|---|---|
search (eval::search) | Hybrid recall@10 ≥ 0.90 and no drop of more than 0.01 against the baseline (NFR-QUAL-1); keyword zero-result rate 0 on exact-reference queries | Search §13.2 |
agentic (eval::agentic) | Citation precision after verification ≥ 0.98 (NFR-QUAL-2); steering failures 0; no answered status on unanswerable questions (FR-SRCH-9) | Search §13.3 |
triage (eval::triage) | Category accuracy ≥ 0.85 (NFR-QUAL-3) and no drop of more than 0.01 against the baseline | Triage §13 |
Pull requests run the same pipelines with the scripted fake model, so prompts, fencing, budgets, the
verifier and schema validation are checked without network access. In addition,
it::search::golden_keyword_recall loads the golden mailbox with the fake AI inside cargo xtask itest
and asserts that keyword recall@10 does not drop against the baseline: keyword search is deterministic,
so this is a required pull-request check. Updating the baseline is a reviewed change with the score
deltas in the pull request description.
10. Live end-to-end suite (live::)
PRD release criterion 3. cargo xtask live runs
cargo test -p pylota-mail-worker --features live --test live -- --test-threads=1 against the staging
deployment (its own zone, platform domain, D1, R2, Vectorize and queues, Architecture).
| Test | Does |
|---|---|
live::inbound::gmail_to_identity, live::inbound::outlook_to_identity | Send from the Gmail and Outlook test mailboxes (their APIs) to a staging identity; message.received within 120 s; verdict: pass, DKIM and DMARC pass |
live::outbound::to_gmail, live::outbound::to_outlook | Send through the API; read the message in the test mailbox; Authentication-Results there shows aligned DKIM pass; replying from the mailbox threads into the same thread |
live::thread::c7 | A reply to our message matches by token, then by learned Message-ID (C7) |
live::inbound::b1_oversize_rejected | A 26 MiB message sent through SES from the test AWS account is rejected before the Worker (B1) |
live::delivery::bounce_unknown_address | Send to an unknown address on the staging platform domain: our own email() rejects with 550 5.1.1, Email Sending reports a bounce, a suppression is created |
live::delivery::ses_simulator | On an SES-transport domain, bounce@, complaint@ and success@simulator.amazonses.com produce bounce, complaint (with suppression and abuse counting) and delivery (SES mailbox simulator, AWS docs read 2026-10-09) |
live::delivery::complaint_event_path | Publish a cf.email.sending.message.complained payload for a real sent message to staging’s pm-delivery-events through the Queues HTTP API; the recipient becomes complained and suppressed |
live::domains::change_and_reply_via_retiring | Move an identity from its platform address to a zone subdomain, then to a zone apex, then roll back by promoting the retiring address; reply to an old thread through the retiring address at each step (C3, build plan M20) |
live::domains::failure_fallback_recovery | Delete the DKIM record of a staging tenant zone through the Cloudflare DNS API, verify twice, assert failing and a sent_via_fallback send with thread continuity; restore the record and assert domain.recovered |
live::transport::j5_ses_failover | Switch a staging domain to SES per the runbook and send (J5) |
live::domains::dns_records_external_host | Connect a staging domain hosted at a DNS provider other than Cloudflare with dns_records, publish its records through that provider’s API, wait for healthy, receive from the Gmail test mailbox through SES, and send with aligned DKIM and SPF (build plan M20) |
live::erasure::counterparty_live | Counterparty erasure of the Gmail test address with one held thread: the receipt lists the hold, probes are zero, and no object is left under the erased keys (checked through the Cloudflare R2 API) |
live::mcp::client_round_trip | An MCP client built on rmcp (the conformance dev-dependency) connects to /mcp with a staging key, lists tools, searches, and sends with an idempotency_key; a repeat call returns the original result |
live::ops::metrics_reach_analytics_engine, live::ops::restore_drill | Observability |
live::ops::fresh_deploy_rehearsal | NFR-OPS-1: a person who did not build the service deploys a fresh Cloudflare account from self-hosting.md alone; the hands-on time is recorded and must be at most 15 minutes (Build plan › M20, step 12) |
Secrets. Live tests read credentials only from the GitHub Environment staging, which requires a
reviewer and is limited to main and release tags; forks never receive them. The environment holds: a
Cloudflare API token scoped to the staging account, a staging platform key with a 90-day expiry, OAuth
credentials limited to the two dedicated test mailboxes (Google Workspace and Microsoft 365, holding only
synthetic mail), AWS credentials for the staging SES resources, and an API token for the external DNS
provider that hosts the dns_records test domain. The secret names are listed in
Build plan › Human prerequisites (STAGING_*). The harness never prints secrets,
redacts them from failure output, deletes test messages from the mailboxes after each run, and the
credentials are rotated every quarter.
Schedule. Nightly, and on every release candidate before the production rollout (section 12).
11. Edge-case mapping and coverage
11.1 Naming
| Prefix | Meaning | Example |
|---|---|---|
core::<module>::<row>_<name> | Native unit or property test in crates/core | core::address::a4_reserved_and_confusable |
conf::<dir>::<row>_<name> | Corpus case in crates/conformance | conf::mime::b5_shift_jis_subject |
it::<area>::<row>_<name> | Integration test against workerd | it::inbound::a6_reject_codes |
live::<area>::<row>_<name> | Live test against staging | live::transport::j5_ses_failover |
cli::<module>::<name> | Native test in crates/cli | cli::setup::ses_region_check |
platform::<module>::<name> | Native test in crates/platform | platform::config::startup_rules |
sdk::<module>::<name> | Native test in crates/sdk | sdk::coverage::every_operation |
xtask::<name> | A check run by cargo xtask over the workspace, the docs or the built bundle | xtask::size_budget |
eval::<suite> | An evaluation run (section 9): its Covers: line names the quality requirement | eval::search |
- The row ID (
a6,j7) starts the last segment, socargo test a6_finds every test for a row. A test that covers several rows, or a requirement rather than one row, may omit it (for exampleit::ses::retired_rule_synccovers N7 and N29); the register names it explicitly. - A
*in the register (conf::mime::b2_*,it::send::g1_*) means at least one test with that prefix. - Rows owned by
I(integrator) have no service test. Rows owned byS+Ihave the service-side test named in the register. - Every test function carries a doc comment line
Covers: <IDs>, for example/// Covers: FR-OUT-1, G1.
11.2 cargo xtask trace
Parses docs/src/project/edge-cases.md and docs/src/project/prd.md, collects test names and Covers:
lines from the source tree (including generated conf:: names), and fails when:
- a test named in an
SorS+Irow does not exist (PRD release criterion 2); - a
P0requirement has no test with it inCovers:(PRD release criterion 1); - a
Covers:line names an unknown row or requirement.
It prints the traceability matrix as Markdown into the CI summary.
11.3 How rows are exercised
| Rows | Harness support |
|---|---|
| A6, B3, J1 | /__test/inbound outcome and SMTP reply |
| A9, A10, B14 | Local email endpoint called once per envelope recipient; the same raw message twice |
| C4, E4 | Concurrent requests from one test; fake clock |
| D5, D10, E5 | Fake clock across hourly and 30-minute windows; many injected messages |
| G2, G3, G4, G10 | Mail sender fake outcomes; queue-delay log; simulator timeout@ for test tenants |
| G6, G8 | /__test/delivery-event; recorded retry delays |
| H1, H4, H6, H7 | DNS fake per resolver; RDAP fake; Cloudflare API fake errors; /__test/alarm for checks |
| I1–I7, F6 | Stateful Vectorize fake; local R2; JobRunner alarms; probe failure mode |
| J2 | restart_runtime() |
| J4 | Webhook receiver fake failing; recorded delays against the 72-hour schedule |
| J7 | d1.query fault on the directory lookup |
| J8 | Forced dead-letter delivery (a consumer fault beyond max_retries); fake clock for the 15-minute alert |
| J9 | /__test/mailbox-schema |
| L1–L4 | Test tenants with the real simulator and loopback paths |
| B1, C7, J5 | Live (B1 and J5 also need real providers). C7 has an it:: part too, and J5’s API part is it::domains::transport_patch |
| N1–N7, N10, N11, N26–N29 | SNS push and SQS fakes; S3 fake with NoSuchKey; SES fake identity, account and receipt-rule state; a generated 39 MB message for N5; seeded domain rows for the identity count |
| N8, N9, N17, N21–N25 | DNS fake per resolver (MX hosts, doubled names, parent NS); Cloudflare API fake zone errors and zone deletion; /__test/alarm for checks |
| N12, N13 | Forwarding simulated by injecting the outbound copy at the identity’s platform address |
| N14–N16, N18–N20 | SMTP server fake scripts; probe and DSN messages handed to the inbound path |
| N30 | cli:: with a recorded AWS API fake |
| W20–W23 | OAuth fakes; a separate cookie jar per simulated browser |
| W24–W26 | Stripe fake and signed webhook payloads; the return page’s refresh loop |
| W27, W28, W30 | Fake clock (TOTP steps, key rotation plus 8 days, the 7-day ramp) |
| W29, W31–W34 | Plain requests; W33 sends two concurrent creates |
| O1–O13 | Fake clock for verify_until and signature expiry; the Rust SDK verifier run against the JWKS served by workerd; tenant policy per test; restart_runtime_with for PM_WEB_BOT_AUTH=off (O9); restart_runtime_with setting the secret PM_MASTER_KEY_NEXT for O8 |
| O14–O26 | Fake clock for the 2-minute hold, the 10-minute windows, the hourly and 09:00 runs, time-zone changes and the 24-hour cooldowns; /__test/alarm with class notifier; notification emails observed like console sign-in mail; /__test/delivery-event for a hard bounce on one (O17); the DNS fake for a failing platform domain (O25); restart_runtime_with for PM_BILLING=off (O23) |
11.4 Coverage
cargo llvm-cov(cargo-llvm-cov, pin at build time) runs oncargo test --workspacein CI.pylota-mail-coremust keep line coverage at or above 85%; the job fails below it. Other crates are reported, not gated.- Coverage never replaces the traceability check: a covered line without a named test for its rule is not “tested”.
12. CI workflows and required checks
The base jobs are those in Rust workspace › CI pipeline. This design adds the security and traceability jobs:
| Workflow | Jobs | Trigger |
|---|---|---|
ci.yml | fmt, clippy, test (with coverage), layering, wasm, itest (includes the attack suite and the deterministic keyword recall), fuzz-smoke, deny, audit (cargo audit), openapi, docs, trace (cargo xtask trace) | Every pull request and push to main |
codeql.yml | CodeQL for Rust | Every pull request, weekly |
nightly.yml | All fuzz targets for 10 minutes each; property tests at 65,536 cases; eval-search, eval-agentic, eval-triage; live:: suite; cargo audit on main | Nightly |
release.yml | The full ci.yml gate; the three evaluations; CLI binaries; cargo xtask release; SBOM (cargo cyclonedx --format json for the Worker and the CLI); signed SHA256SUMS; build provenance (actions/attest@v4); deploy to staging; the live:: suite; then the GitHub Release, cargo publish, and the production rollout (10% → 50% → 100%, Architecture) | Tag v* |
Required checks to merge into main: fmt, clippy, test, layering, wasm, itest,
fuzz-smoke, deny, audit, openapi, docs, trace, codeql.
Required to publish a release: all of the above on the tagged commit; the three evaluation gates
(section 9.3); the live:: suite green on staging with that commit deployed; no open crash from the
nightly fuzz run; the security pre-release checklist.
GitHub Actions are pinned to full commit SHAs and each job declares least-privilege permissions
(Security › Supply chain).
13. Tests of the test infrastructure
| Test | Proves |
|---|---|
xtask::trace_detects_missing_test | A fixture register naming a non-existent test fails cargo xtask trace |
xtask::itest_refuses_release_hooks | cargo xtask build-worker fails when itest-hooks is enabled or the bundle contains /__test/ |
it::harness::hooks_need_token | Hooks without x-pm-test-token return 404 |
it::harness::no_internet_egress | A request to an unregistered host fails with HttpError::Connect |
it::harness::fake_clock_alarm | An alarm armed 90 days ahead runs through /__test/alarm after advancing the fake clock, and not before |
conf::loader::every_case_has_expectations | Every .eml has a .toml with id, edge, source, licence, and every imported case is in THIRD_PARTY.md |
conf::loader::no_real_domains | Every address in the corpus is under an RFC 2606 or .invalid name |