Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Observability and SLOs

Binding for implementation. This page defines what the deployment records about itself (logs, metrics, alert state), how service-level objectives are measured, which alerts fire and how they reach a person, the dead-letter queue consumers, /health and pmail doctor, and the runbooks. Nothing here may record mail content or clear-text addresses (Security › Logging rules).

RequirementsFR-OPS-3, FR-OPS-4, FR-PRV-6, FR-DLV-3, FR-DLV-5, FR-DOM-9, FR-IDN-6…8, FR-CON-14, FR-CON-15, FR-BILL-13, NFR-REL-1…4, NFR-PERF-1…6, NFR-PRV-1, NFR-OPS-2
Edge casesD4, D5, G3, G8, I5, J4, J5, J6, J8, N1, N4, N10, O24, O25
Codecrates/worker/src/log.rs, crates/worker/src/metrics.rs, crates/worker/src/ops/ (alert evaluator, DLQ consumer), crates/core/src/slo.rs (alert rule evaluation, pure)

1. Signals

SignalWhereRetentionHolds
Structured logsWorkers Logs (one JSON object per console.log line)7 days (Cloudflare)IDs, codes, counts, durations, pseudonyms (section 2)
MetricsWorkers Analytics Engine dataset pylota_mail_metrics, binding METRICS3 months (Cloudflare)Counters and observations with low-cardinality labels (section 3)
Alert stateD1 audit_log rows alert.fired / alert.resolved, plus an alert_fired metricLife of the deploymentAlert key, severity, values
EventsWebhooks (domain.failing, quota.warning, webhook.disabled, erasure.failed, identity.paused, …)Webhook eventsTenant-facing conditions
TracesWorkers traces7 days (Cloudflare)Staging only (below)
Exact countersD1 usage_daily, TenantQuotaPrivacyUsage and caps

Analytics Engine from Rust. workers-rs 0.8.7 exposes the binding: Env::analytics_engine(name) returns AnalyticsEngineDataset, written with write_data_point(&AnalyticsEngineDataPoint) or AnalyticsEngineDataPointBuilder::new().indexes(..).add_blob(..).add_double(..).write_to(&dataset) (docs.rs, worker 0.8.7, read 2026-10-09). Limits: 20 blobs, 20 doubles and one index per data point, 16 KB of blobs, an index of at most 96 bytes, 250 data points per invocation (Analytics Engine limits, read 2026-10-09). wrangler dev does not write local data to Analytics Engine, so with PM_ENV = "local" the metrics sink writes each data point as a log line instead (event = "metric"), which the integration tests read.

Required Worker configuration. The generated wrangler.toml must contain:

[observability]
enabled = true
head_sampling_rate = 1

[observability.logs]
invocation_logs = false        # invocation logs record request URLs and the email recipient (FR-PRV-6)

[observability.traces]
enabled = false                # production; staging sets true with head_sampling_rate = 0.1

[[analytics_engine_datasets]]
binding = "METRICS"
dataset = "pylota_mail_metrics"

Cloudflare’s invocation log for a fetch is the method and full URL, and for the email handler it is the recipient address (Workers Logs docs, read 2026-10-09). URLs such as /v1/identities/lookup?address=… contain addresses, so invocation logs are off. Automatic traces record URLs and handler attributes the same way, so traces run only on staging, whose traffic is synthetic. Workers Issues needs Wrangler 4.134.0 or later. The pinned 4.139.0 supports it, but v1 does not depend on it.

METRICS is listed in Configuration › Bindings and in the template in Rust workspace.

2. Structured logs

2.1 Schema

worker::log writes one JSON object per line. Only these fields exist; values are typed (section 12 of Security):

FieldTypePresentMeaning
tsRFC 3339, msalwaysPlatform clock
levelerror | warn | info | debugalwaysFiltered by PM_LOG_LEVEL
eventsnake_case namealwaysSection 2.2
envproduction | staging | localalwaysPM_ENV
version, commitstringalwaysBuild metadata
handlerfetch | email | queue | scheduled | alarm | rpcalwaysEntry point
request_idreq_…alwaysGenerated per invocation (Design conventions); returned in Request-Id; carried in RPC envelopes and queue bodies
tenant_id, identity_idopaque IDwhen knowntenant_id is the tenant pseudonym: slugs and names are never logged
key_idkey_…authenticated requestsNever the key string
route, method, statuspattern, verb, intfetchThe matched route pattern, never the raw path or query
codeerror or reason codeon failureErrorCode or a reason from Errors
queue, attemptname, intqueue handlers
message_id, thread_id, job_id, domain_id, event_id, delivery_id, export_id, erasure_id, user_idopaque IDswhen relevantuser_id is the person’s usr_ ID (console requests, notifications), never their address
address_phph_ + 16 hexwhen an address must be correlatedPseudonym of the address the event concerns (counterparty or envelope recipient): HMAC-SHA256(PM_HASH_KEY, address) truncated
transport, provider_code, smtp_codeshort codesoutbound and deliveryE_RATE_LIMIT_EXCEEDED, 550, 5.1.1; never smtpResponse text
mcp_method, toolJSON-RPC method, MCP tool nameMCP requestsNever the arguments (MCP)
query_hash16 hexsearch requests and MCP search toolshex(HMAC-SHA256(PM_HASH_KEY, q))[..16] (MCP)
count, bytes, duration_msintwhen relevant
detail^[a-z0-9_.:-]{1,64}$optionalMachine string only
{"ts":"2026-10-09T10:12:03.412Z","level":"info","event":"inbound_accepted","env":"production",
 "version":"1.0.0","commit":"abc1234","handler":"email","request_id":"req_01J9Z5…",
 "tenant_id":"ten_01J9…","identity_id":"idn_01J9…","message_id":"msg_01J9…","bytes":48213,"duration_ms":41}

2.2 Event names

EventLevelEmitted when
http_requestinfo (5xx: error)Every fetch response
inbound_accepted, inbound_staged, inbound_rejected, inbound_tempfailinfo / warnemail() outcome; inbound_rejected carries the SMTP code
inbound_processedinfopm-inbound consumer committed or deduplicated a message
outbound_transportinfo / warnA transport call and its classified outcome
delivery_event, delivery_orphanedinfo / warnA provider event applied or parked (G8)
webhook_attemptinfo / warnOne delivery attempt
index_job, triage_resultinfo / warnIndexing and triage outcomes
search_request, agentic_requestinfoMode, scope, duration, status; the query only as query_hash
job_step, job_failedinfo / errorJobRunner transitions
domain_check, domain_transitioninfo / warnDomainMonitor
dlq_itemerrorDead-letter consumer recorded an item
alert_fired, alert_resolvederror / infoAlert evaluator transitions (the audit actions are alert.fired and alert.resolved)
rpc_owner_mismatcherrorDurable Object owner check failed (Design conventions)
config_invaliderrorRequired variable or secret missing or malformed
secrets_reseal_progressinfoMaster-key rotation sweep (Security)
identity_key_changedinfoAn identity key was created (explicitly or on first signing), rotated or revoked: identity_id, key_id or user_id of the actor, detail = create, rotate or revoke; never key material (Agent signing keys)
signature_mintedinfoAn agent assertion or HTTP signature was minted: identity_id, key_id, detail = assertion or http_signature. Never the token, the signature, the audience, ext, the URL or the headers (Security › Logging rules)
notification_sent, notification_failed, notification_deferredinfo / warn / infoThe Notifier submitted a notification email, had it refused or saw it fail, or held an item back: tenant_id, user_id, message_id once known, code on failure, detail = the kind (usage, new_mail, needs_person, account) or the deferral reason. Never an address or the email’s text (Notifications)
notifier_handoff_failedwarnThe pm-webhooks consumer could not hand a new-mail event to the tenant’s Notifier; the delivery work is unaffected: tenant_id, event_id, code (Webhooks)
notification_unsubscribeinfoPOST /console/notifications/unsubscribe: tenant_id and user_id when the token verified, detail = the kind, code = expired or invalid otherwise; never the token
usage_alertinfoTenantQuota asked the Notifier for a usage alert: tenant_id, detail = {feature}:{threshold}
panicerrorPanic hook; source location only
metricdebugPM_ENV = "local" only: a metrics data point

3. Metrics

3.1 Data point layout

Every metric is written to METRICS with the same layout, so one SQL shape reads them all:

ColumnHolds
index1tenant_id, or platform for deployment-level metrics. This is the sampling key
blob1Metric name
blob2PM_ENV
blob3 … blob8Labels, in the order listed for the metric in section 3.2 (unused positions empty)
double1Value: the count for counters; the observed value for observations (milliseconds, bytes)
  • Counters are aggregated in memory per invocation by (index1, name, labels) and written once at the end of the invocation with double1 = count. Observations (durations, sizes) are one data point each. A per-invocation cap of 240 points protects the platform limit; the overflow is counted in metrics_dropped_total.
  • Labels are low-cardinality codes. Message, thread and identity IDs are never labels; domain_id and tenant_id are, because their counts are bounded by configuration.
  • Writes never fail a request: an error from the binding is logged once per isolate and ignored.

Reading (Analytics SQL API dataset events.analyticsEngine."pylota_mail_metrics", account scope; sample weights are applied to COUNT, SUM and AVG automatically, per the SQL API datasets page read 2026-10-09):

SELECT blob4 AS domain_id, SUM(double1) AS value
FROM events.analyticsEngine."pylota_mail_metrics"
WHERE accountTag = '<ACCOUNT_TAG>'
  AND timestamp >= NOW() - INTERVAL '1' HOUR
  AND blob1 = 'bounces_total' AND blob2 = 'production'
GROUP BY domain_id

3.2 Catalogue

MetricKindLabels (blob3, blob4, …)Emitted by
http_requests_totalcounterroute, method, status_class, codefetch
http_msobservationroutefetch
rate_limited_totalcounterbucketfetch
inbound_received_totalcounterresult: accepted, staged, rejected_unknown, rejected_retired, rejected_suspended, tempfail_suspended, tempfail_storageemail()
inbound_r2_retries_totalcounter–email()
inbound_processed_totalcounteroutcome: stored, deduplicated, quarantined, hidden, throttled, dsn_applied, parse_degradedpm-inbound
inbound_ingest_msobservation–pm-inbound: received_at → commit
inbound_lost_totalcounter–Global retention staging step: a staging object still unrouted after its re-queue
inbound_raw_missing_totalcounter–pm-inbound: a pointer whose raw object is missing and whose message is not in the mailbox (Inbound)
inbound_orphan_raw_total, inbound_staged_unroutable_totalcounter–email() and pm-inbound (Inbound)
inbound_dropped_totalcounterreason (unknown_recipient, tenant_suspended, identity_gone), source (routing, ses)email(), pm-inbound (Inbound); on SES domains unknown recipients are dropped without a bounce (Domains on any DNS host §4.6)
ses_sns_rejected_totalcounterendpoint (delivery for /hooks/ses, inbound for /hooks/ses/inbound), reason (version, signature, cert_host, topic, timestamp)Both SNS endpoints and the SQS backstop: a message refused with 403 invalid_signature (N1)
ses_auth_disagreement_totalcountercheck (dkim, dmarc)pm-inbound: SES’s verdict differs from our own check on the same message
ses_object_lost_totalcounter–pm-inbound: an S3 object was missing while its ses_ingest row was still queued (N4)
scanner_error_total, auth_dns_cache_miss_totalcounter–pm-inbound (Inbound)
backscatter_totalcounter–pm-inbound (D4)
inbound_throttled_totalcounter–mailbox (D5)
thread_token_invalid_total, thread_token_previous_key_totalcounter–mailbox (Threading)
quarantine_totalcounterreasonmailbox
send_api_requests_totalcounteroperation, result (accepted, deduplicated, or the error code)fetch
send_api_msobservationoperationfetch, accepted sends only
outbound_queue_to_transport_msobservationtransportpm-outbound, first transport attempt
transport_outcomes_totalcountertransport (cloudflare, ses, smtp, simulator), outcome (accepted, rejected, retry, uncertain), provider_code (for smtp, the SMTP reply code)pm-outbound
recipients_submitted_totalcountertransport, domain_idpm-outbound
bounces_totalcounterdomain_id, bounce_typepm-delivery-events
complaints_totalcounterdomain_idpm-delivery-events
delivery_events_totalcountertransport, typepm-delivery-events
delivery_orphaned_totalcounter–pm-delivery-events (G8)
delivery_unknown_type_total, delivery_unroutable_totalcounter–pm-delivery-events (Outbound)
reconcile_ambiguous_totalcounter–mailbox (Outbound)
uncertain_total, reconciled_totalcountertransportmailbox
fallback_sends_totalcounterdomain_idmailbox
suppressed_recipients_totalcounterreasonmailbox
provider_quota_errors_totalcountertransport, provider_codepm-outbound (G3)
backup_objects_totalcounterresult (copied, skipped, error)JobRunner backup job (Privacy)
webhook_attempts_totalcounterresult (succeeded or the attempt’s error code)pm-webhooks (Webhooks › Metrics)
webhook_delivery_latency_msobservationevent_class (inbound for message.received and message.quarantined, other), first_attempt (succeeded, failed)pm-webhooks: the event’s occurred_at to the first successful attempt, written once per delivery
webhook_dead_totalcounterevent_classpm-webhooks: a delivery’s 13th attempt failed (not written for endpoint_disabled)
webhook_disabled_totalcounterreasonpm-webhooks
outbox_undispatched_age_msobservationowner (mailbox, domain, job)outbox dispatch
webhook_ssrf_blocked_totalcounter–pm-webhooks, webhook create and update
identity_keys_totalcounterop (create, rotate, revoke)fetch and console: the identity-key handlers; a key created lazily by a first signing request counts as create (Agent signing keys)
signatures_totalcounterkind (assertion, http_signature), result (ok, or the error code, for example rate_limited, identity_paused, policy_denied, web_bot_auth_disabled)fetch: POST …/assertions and POST …/http-signatures
well_known_requests_totalcounterendpoint (identity_jwks, directory), status_classfetch: GET /.well-known/jwks/{identity_id}.json and GET /.well-known/http-message-signatures-directory
notifications_sent_totalcounterkind (usage, new_mail, needs_person, account)Notifier: an email accepted by the outbound pipeline (202) (Notifications)
notifications_failed_totalcounterkind, reason (the error code of a refused or unfinished submit, for example unavailable or timeout; or the message’s reason when an accepted notification ends failed or rejected, for example domain_failing_no_fallback)Notifier, for submits; the system identity’s mailbox, for accepted notifications that end failed or rejected. A bounce or complaint is not counted here: it pauses the person’s preferences (O17)
notifications_deferred_totalcounterreason (cap_person, cap_workspace, paused, platform_domain)Notifier: an item held back by a daily cap (into the next digest, O24), skipped while the person’s preferences are paused, or kept for the hourly retry while the platform domain is failing (O25)
notification_unsubscribes_totalcounterkind (empty when the token cannot be read), result (ok, expired, invalid)fetch: POST /console/notifications/unsubscribe (O18)
usage_alerts_totalcounterfeature (inboxes, sends, triage, custom_domains, storage_gb, seats), threshold (80, 100)TenantQuota: a NotifierRequest::UsageThreshold sent, once per threshold per period, or after the 24-hour cooldown for counts (Notifications §4)
search_requests_totalcountermode, scope, result (ok, degraded, partial, code)fetch
search_msobservationmode, scope, fanout (1, 2-10, 11-100)fetch
agentic_requests_totalcounterstatusfetch
agentic_ms, agentic_first_evidence_msobservation–fetch
ai_calls_totalcounterpurpose (embed, rerank, triage, planner, markdown), resultworker
index_jobs_totalcounterkind, resultpm-index
triage_totalcounterstatuspm-index
job_steps_totalcounterkind, step, resultJobRunner
erasure_msobservationscopeJobRunner: created_at → completion
retention_purged_totalcounterstore (raw, messages, vectors, events)JobRunner
domain_checks_totalcounterresolver, outcomeDomainMonitor
domain_transitions_totalcounterfrom, toDomainMonitor
mailbox_size_bytesobservation–mailbox, hourly at most
dlq_itemsobservationqueueAlert evaluator, every minute: open items per queue (J8)
rpc_owner_mismatch_totalcounterclassDurable Objects
alert_firedcounteralert, severityAlert evaluator
panics_total, config_invalid_total, metrics_dropped_totalcounter–any

4. Service-level objectives

Windows are rolling 30 days unless stated. “Good” and “total” are counted from the metrics above.

IDObjectiveSLI: good / totalNotes
NFR-REL-10 acknowledged inbound messages lostinbound_lost_total + inbound_raw_missing_total + ses_object_lost_total = 0Any non-zero value pages
NFR-REL-2≥ 99.9% of valid inbound accepted(accepted + staged) / (accepted + staged + tempfail_storage) from inbound_received_totalRejections of unknown, retired and suspended addresses are correct behaviour and excluded
NFR-REL-3Inbound accepted → webhook delivered: p95 ≤ 30 s, p99 ≤ 120 sshare of webhook_delivery_latency_ms{event_class=inbound, first_attempt=succeeded} ≤ 30,000 (target 95%) and ≤ 120,000 (target 99%)Measured to delivery, as the PRD states. Deliveries whose first attempt failed at the endpoint are excluded from the SLI so an integrator’s outage does not burn the service budget; they stay visible on the dashboard
NFR-REL-4≥ 99.99% of webhooks delivered within 24 hcount of webhook_delivery_latency_ms ≤ 86,400,000 / (count of webhook_delivery_latency_ms + webhook_dead_total)webhook_dead_total counts only deliveries whose 13th attempt failed; a delivery ended because its endpoint was disabled (attempt error endpoint_disabled) is not counted, so disabled endpoints are excluded
NFR-PERF-1Send API p95 ≤ 500 msshare of send_api_ms ≤ 500 (target 95%)
NFR-PERF-2Queued → transport p95 ≤ 60 sshare of outbound_queue_to_transport_ms ≤ 60,000 (target 95%)First attempt only; quota back-off is measured by provider_quota_errors_total
NFR-PERF-3Keyword search, one identity, p95 ≤ 200 mssearch_ms{mode=keyword, scope=identity} ≤ 200 (95%)
NFR-PERF-4Hybrid search, one identity, p95 ≤ 800 mssearch_ms{mode=hybrid, scope=identity} ≤ 800 (95%)
NFR-PERF-5Tenant search over ≤ 10 identities, p95 ≤ 1 ssearch_ms{scope=tenant, fanout ∈ {1, 2-10}} ≤ 1,000 (95%)
NFR-PERF-6Agentic p95 ≤ 8 s; first evidence ≤ 1.5 sagentic_ms ≤ 8,000 and agentic_first_evidence_ms ≤ 1,500 (95%)
NFR-PRV-1Erasure ≤ 24 h, receipt always producederasure_ms ≤ 86,400,000 (100%)erasure_overdue alert at 20 h
NFR-OPS-2RPO ≤ 1 min (indexes), ≤ 15 min (blobs); RTO ≤ 4 hRestore drills on staging (Restore from PITR)See the R2 note in that runbook

Measured outside production: NFR-QUAL-1…3 by the evaluation harness, NFR-SEC-1 by the attack suite, NFR-SEC-2 by cargo xtask build-worker, NFR-OPS-1 by deploy rehearsals (Testing). NFR-COST-1 is checked by reviewing Cloudflare usage after a week of idling on staging.

5. Alerts

5.1 How alerts reach a person

ClassEvaluated byDelivered by
A: metric alertsCloudflare Custom Alerts (beta), which run a SQL API query on a schedule with threshold, anomaly or SLO detection and deliver to email, webhooks or PagerDuty (Cloudflare Notifications docs, read 2026-10-09). Workers Analytics Engine datasets are queryable through the SQL APIThe deployer’s chosen destination
B: state alertsThe Worker’s alert evaluator (section 5.4), every minute, from exact state in D1 and the objectsAn alert.fired audit row, an error log line and one alert_fired data point; one Class A Custom Alert (state alerts) forwards every alert_fired point
C: event alertsThe service, as part of normal behaviourWebhook events to the integrator’s endpoints

Custom Alerts are created in the Cloudflare dashboard from the queries in deploy/observability/alerts/; no creation API was found in the documentation on 2026-10-09, so pmail doctor cannot check that they exist. Where Custom Alerts are not available on the account, the operator runs pmail doctor --json on a schedule: it lists firing state alerts and exits non-zero when any is firing.

5.2 Burn-rate rules

Availability SLOs use multi-window burn rates. A Custom Alert with SLO detection fires when (1 − bad/total) × 100 is below its target over both the short and the long window, so each rule’s target is 100 × (1 − burn × (1 − SLO)).

SLOBurnLong / short windowCustom Alert targetSeverity
NFR-REL-2 (99.9%)14.4×1 h / 5 min98.56page
NFR-REL-2 (99.9%)6×6 h / 30 min99.40page
NFR-REL-2 (99.9%)3×24 h / 2 h99.70ticket
NFR-REL-4 (99.99%)14.4×1 h / 5 min99.856page
NFR-REL-4 (99.99%)6×6 h / 30 min99.94page
NFR-REL-4 (99.99%)3×24 h / 2 h99.97ticket

Latency SLOs (95% and 99% targets) have budgets too large for those burn rates, so they use two rules each: page when the good share is below 90% (95% targets) or 97% (99% targets) over 1 h and 5 min, and ticket when it is below the target over 6 h and 30 min. Each rule needs at least 100 events in the window (the Custom Alert “minimum event count”).

5.3 Alert list

AlertClassConditionSeverityRunbook
dlq:{queue}BThe oldest open dlq_items row of a queue is older than 15 minutes (J8)pageDLQ growth
bounce_rate:{domain_id}Abounces_total / recipients_submitted_total > 2% over 1 h for a domain with ≥ 50 recipientspageBounce spike
complaint_rate:{domain_id}Acomplaints_total / recipients_submitted_total > 0.1% over 24 h for a domain with ≥ 200 recipientspageComplaint spike
inbound_reject_spikeAAnomaly detection on inbound_received_total{result=rejected_unknown}: spike, 15-minute evaluation window, 24 h baseline, minimum 50 eventsticketDomain failing (routing checks)
inbound_tempfailAinbound_received_total{result=tempfail_storage} > 0 over 5 minutespageDLQ growth (storage path)
inbound_lostBinbound_lost_total or inbound_raw_missing_total > 0pageRestore from PITR (re-ingest step)
ses_object_lostBses_object_lost_total > 0: an S3 object was deleted before every recipient was ingested (N4)pageSES account and receiving
ses_sending_pausedBThe 15-minute SES platform check finds account sending paused. Every SES domain uses the fallback address meanwhile (N10)pageSES account and receiving
ses_rule_missingBThe 15-minute SES platform check finds the receipt rule set PM_SES_RULE_SET inactive or without the rule pm-deliver (only when SES receiving is configured)pageSES account and receiving
ses_identities_90pctBSES identities in the region reach 9,000, 90% of the 10,000 per Region (quotas, read 2026-10-09). Counted as domains rows with ses_region set and not removed, plus the platform identity. There is no quota.warning event for thisticketSES account and receiving
webhook_failing:{webhook_id}Bconsecutive_failures ≥ 10 on an enabled endpointticketIntegrator API down
webhook_disabled:{webhook_id}B + CEndpoint disabled with failing (webhook.disabled event)ticketIntegrator API down
provider_quotaAprovider_quota_errors_total > 0 over 15 minutes: the first quota error (G3)pageQuota exhausted
provider_quota_80BOnly when PM_DAILY_SEND_QUOTA is set: today’s (UTC) sends in usage_daily, summed over live tenants, reach 80% of it (G3)ticketQuota exhausted
quota_warningCquota.warning at 80% and 100% of a tenant or identity cap– (tenant-facing)Quota exhausted
domain_failing:{domain_id}B + CDomain enters failing or suspended (domain.failing, domain.suspended)ticketDomain failing
inbound_throttledAinbound_throttled_total > 100 over 1 h: one or more senders exceed inbound.per_sender_per_hour and their excess is stored throttled (D5)ticketAbusive identity (the affected mailbox’s rate_windows rows name the sender; add a receive-block if it is abuse)
stripe_webhook_errorsBOnly with PM_BILLING=stripe: at least one billing_events row with outcome starting error: received in the last hour (the detail lists each type and code)ticketBilling design › Stripe webhook (fix the endpoint’s event list, or the customer mismatch)
mailbox_size:{identity_id}BMailbox SQLite size > 70% of 10 GB (7,516,192,768 bytes), reported by the mailbox’s size check (at most hourly, after a write; Data model › Mailbox notes)ticketAbusive identity (archive or split)
abuse_pause:{identity_id}B + CAn identity paused with abuse_threshold (identity.paused)ticketAbusive identity
erasure_failed:{erasure_id}B + CErasure request failed (erasure.failed)pageErasure failure
erasure_overdue:{erasure_id}BErasure still running 20 h after creationpageErasure failure
rpc_owner_mismatchBrpc_owner_mismatch_total ≥ 1pageCompromised key (treat as a security incident)
uncertain_spikeAtransport_outcomes_total{outcome=uncertain} > 5 over 15 minutespageEmail Sending outage
delivery_orphanedAdelivery_orphaned_total > 10 over 1 hticketEmail Sending outage
panicsApanics_total > 0 over 5 minutesticketParser bug
config_invalidAconfig_invalid_total > 0pagepmail doctor
webhook_secret_unavailableAwebhook_attempts_total{result=secret_unavailable} > 0 over 15 minutes (a sealed secret no longer opens: wrong or rotated PM_MASTER_KEY, Webhooks)pageSecurity › Rotation procedures (PM_MASTER_KEY)
webhook_ssrf_blockedAwebhook_ssrf_blocked_total > 20 over 1 hticketIntegrator API down (an endpoint’s DNS now points at a blocked range)
notification_send_failuresAnotifications_failed_total > 0 in each of 3 consecutive hours, or > 20 in one hourticketDomain failing for the platform domain first (system mail has no fallback, Notifications §7), then Email Sending outage
SLO burn rulesASection 5.2page / ticketThe runbook of the failing path

5.4 The state alert evaluator

The * * * * * cron runs ops::alerts::evaluate:

  1. Read the conditions:
    • dlq_items: per queue, COUNT(*) and MIN(first_seen_at) of open rows; also written as the dlq_items metric;
    • webhook_endpoints with enabled = 1 AND consecutive_failures >= 10, and those disabled with failing in the last minute;
    • domains in failing or suspended;
    • erasure_requests with status = 'failed', or status = 'running' and created_at older than 20 h;
    • when PM_DAILY_SEND_QUOTA is set: SELECT SUM(u.value) FROM usage_daily u JOIN tenants t ON t.id = u.tenant_id WHERE u.day = ?today AND u.metric = 'sends' AND t.mode = 'live' against 80% of the quota. The roll-up runs every 15 minutes, so this alert can lag by up to 15 minutes; SES sends are counted too, which only makes it fire earlier;
    • when SES is configured: SELECT COUNT(*) FROM domains WHERE ses_region IS NOT NULL AND state <> 'removed', plus one for the platform identity, against 9,000 (ses_identities_90pct);
    • conditions reported by objects and crons since the last run (mailbox_size, abuse_pause, rpc_owner_mismatch, inbound_lost, ses_object_lost from the inbound consumer, and ses_sending_paused and ses_rule_missing from the 15-minute SES platform check, which reads GetAccount and the receipt rule set): the reporting code writes an alert.fired audit row itself.
  2. Read the current state: for each alert key, the latest audit_log row with action IN ('alert.fired', 'alert.resolved') AND target_id = <alert key>.
  3. Transition, with pure rules in core::slo: a true condition on a key that is not firing writes alert.fired (details_json = { "severity", "values" }), logs alert_fired and writes one alert_fired point. A firing key re-notifies (another alert_fired point, no audit row) every 6 hours. A firing key whose condition has been false on two consecutive runs writes alert.resolved.
  4. Alert rows use tenant_id of the affected tenant, or NULL for deployment-level alerts. They contain IDs and numbers only.

6. Dashboards

The SQL for each panel is kept in deploy/observability/dashboards/ and runs against the Analytics SQL API, so it works from any tool that can call that API.

DashboardPanels
OverviewRequests and 5xx rate by route; SLO compliance per objective (section 4) with remaining error budget; firing alerts (from alert_fired)
InboundAccepted, staged, rejected and temp-failed per hour; quarantine reasons; verdict mix; inbound_ingest_ms and webhook_delivery_latency_ms{event_class=inbound} p50/p95/p99 by first_attempt; backscatter and throttling; drops by reason and source; SES: SNS rejections, auth disagreements, lost objects
Outbound and deliverySends by transport and outcome; bounce and complaint rate per domain; uncertain and reconciled; fallback sends per domain; provider quota errors; delivery orphans
WebhooksAttempts by result and error code; dead deliveries; endpoints over 10 consecutive failures; attempt latency
Identity and notificationsSigning calls by kind and result; identity-key operations; JWKS and directory requests by status class; notifications sent, failed and deferred by kind and reason; unsubscribes by result; usage alerts by feature and threshold
Search and AILatency per mode and scope; degraded and partial share; agentic status mix; AI call failures by purpose; index job failures
Jobs and privacyErasure durations and status; retention purges per store; DLQ open items per queue
DomainsDomains per state and per method; transitions; check outcomes per resolver; SES identities against the 10,000 per Region

7. Health checks

7.1 GET /health

  • No authentication, no dependency calls, no tenant data, no alert state.
  • 200 {"status": "ok", "version": "1.0.0", "commit": "abc1234", "env": "production"} when the isolate’s configuration is valid. env is PM_ENV, which Configuration says is shown here. When SES is configured, the body also has "ses_region": "eu-west-2" (the value of PM_SES_REGION), so anyone can see where AWS processes mail (Domains on any DNS host §11).
  • 200 {"status": "degraded", …} when the Worker runs with a feature off because its configuration is incomplete: "ses": "sns_topic_missing" (SES credentials and region set, PM_SES_SNS_TOPIC_ARN missing: the SES transport is off) or "billing": "stripe_secrets_missing" (PM_BILLING=stripe without its secrets: billing is not started) (Rust workspace › Startup rules).
  • 503 with the error envelope (unavailable) when the configuration is invalid (a required variable or secret is missing or malformed, or an optional variable is malformed; config_invalid is logged with the variable’s name, and the body’s details.config_invalid names it). Every other handler returns the same.
  • It is a liveness check. Dependency health is pmail doctor’s job.

7.2 pmail doctor

doctor runs every check, prints one line per check with pass, warn or fail, and a fix for each failure (FR-OPS-3). Output format and exit codes are defined in CLI and setup. The checks:

CheckFails whenFix printed
dns.platformMX, SPF, DKIM or DMARC for the platform domain missing or different from the provider API’s expected records, on either resolverThe exact record to add
routing.catch_allThe platform domain’s catch-all rule does not target the WorkerThe API call or dashboard step
sending.domainsA sending domain is not onboarded, or has preview_enabled = true (Privacy)Onboarding step; PATCH … {"preview_enabled": false}
sending.event_subscriptionsA sending domain has no event subscription to pm-delivery-events (delivery_events: "manual", the spike S9 fallback), or a subscription is left over from a removed domainpmail domains subscribe <domain>; for a left-over one, wrangler queues subscription delete <id>
bindingsA resource in the generated wrangler.toml is missing; D1 or R2 jurisdiction differs from PM_JURISDICTION; a queue lacks its dead-letter consumer; a bound Vectorize index (VECTORS, and VECTORS_NEXT during a re-embed) is not cosine with the eight metadata indexes, or its dimensions differ from those of the model named in its description (embed_model=…): 1,024 for @cf/baai/bge-m3, otherwise the length of a probe embedding from that model, as the re-embed step does ( Search §7.3); METRICS missingThe resource to create or pmail setup
secretsA required secret is missing (names only; values are never read). warn when PM_MASTER_KEY_NEXT is present (an unfinished master-key rotation)The secret to set, or the rotation step to finish
observabilityinvocation_logs is not false, or traces are enabled with PM_ENV = productionThe wrangler.toml lines
worker.versionThe deployed version differs from the CLI’spmail upgrade
healthGET /health is not 200, or reports degraded (a warn that names the feature that is off)Section 7.1; the missing variable or secret
alertsAny state alert is firing (GET /v1/audit-events?action=alert.fired, minus later alert.resolved)The alert’s runbook
dlqOpen dlq_items existpmail dlq list (GET /v1/platform/dlq)
quotaProvider quota errors in the last 24 hours (Analytics Engine SQL API); warn when PM_DAILY_SEND_QUOTA is unset, and warn (never fail) when the operator’s token lacks Account Analytics · Read, so the errors cannot be countedQuota exhausted; the permission in Deploy › step 2
web_bot_auth (when PM_WEB_BOT_AUTH = "on"; otherwise skip)GET /.well-known/http-message-signatures-directory does not answer 200 with Content-Type: application/http-message-signatures-directory+json, lists no key or more than three, or lacks a valid http-message-signatures-directory signature for each listed key (Agent signing keys §3.2)pmail keys rotate web_bot_auth; Deploy › Signed HTTP requests
ses (when PM_SES_REGION is set)GetAccount: production access not enabled, or account sending paused; with SES receiving configured, the active receipt rule set is not PM_SES_RULE_SET or lacks pm-deliver; PM_SES_REGION cannot receive mail; the region holds 10,000 identities (new SES domains are refused with ses_identity_limit). warn at 9,000 or more identities (ses_identities_90pct), and when the region is outside the EU and the UK under PM_JURISDICTION=eu (CLI › Doctor)SES account and receiving
cloudflare.zonesNever fails. Prints the account’s zone count, and warns above 1,000, because the zone limit of a non-Enterprise account is not documented (Domains on any DNS host §3.2)Ask Cloudflare to confirm the account’s zone limit
security_txtPM_SECURITY_CONTACT unset (warn), or Expires within 30 daysSet the variable; upgrade
mail_test (--mail-test)A message from the platform domain to a platform address does not arrive within 120 s with verdict: passPrints the observed authserv-id for PM_TRUSTED_AUTHSERV_ID

8. Dead-letter queues

Every work queue has a dead-letter queue with a consumer (FR-OPS-4, Rust workspace): pm-inbound-dlq, pm-outbound-dlq, pm-delivery-events-dlq, pm-webhooks-dlq, pm-index-dlq, each with max_batch_size = 100.

8.1 Consumer

For each message:

  1. Insert into dlq_items (INSERT OR IGNORE on (queue, message_id), so a repeat is absorbed): a new dlq_ ID, the source queue, the Cloudflare message ID, the body as received, its SHA-256, the tenant_id and kind named by the body if any, and first_seen_at = now.
  2. Log dlq_item (queue, message ID, tenant, kind from the body; no body text) and count it.
  3. ack(). A failed insert leaves the message for the dead-letter queue’s own retry.

The alert evaluator turns open items into dlq_items metrics and the dlq:{queue} alert (J8).

8.2 Table

dlq_items is defined in Data model: a dlq_ ID, the queue, the Cloudflare message ID (unique per queue), the body as received, its SHA-256, the tenant and kind, and the redrive bookkeeping. Rows are deleted after 14 days by the global retention job (Privacy).

8.3 Listing and redriving

Dead-letter items are read and redriven through the platform API, with a platform key holding platform:ops (REST API › Platform operations). The CLI wraps it as pmail dlq list and pmail dlq redrive (CLI and setup).

  • GET /v1/platform/dlq lists items (filters queue, status, tenant_id). It returns the queue, kind, tenant and timestamps, never the stored body: pointers can carry envelope addresses.
  • POST /v1/platform/dlq/{dlq_id}/redrive first checks that SHA-256 of body_json still equals body_sha256. A mismatch (the row was edited or damaged) is refused with 500 internal_error, logged as dlq_body_mismatch and nothing is published. Otherwise it publishes the stored body back to its source queue through the Worker’s own producer binding (Q_INBOUND, Q_OUTBOUND, Q_DELIVERY, Q_WEBHOOKS or Q_INDEX), then sets redriven_at and increments redrive_count in the same request, and writes an audit_log row (dlq.redrive). No Cloudflare API token is involved. Every consumer is idempotent (Design conventions), so a redrive is safe to repeat.

9. Runbooks

RunbookTypical alert
Bounce spikebounce_rate:{domain_id}
Complaint spikecomplaint_rate:{domain_id}
Quota exhaustedprovider_quota, provider_quota_80, quota_warning
Email Sending outageuncertain_spike, delivery_orphaned, outbound burn rules, notification_send_failures
Domain failingdomain_failing:{domain_id}, inbound_reject_spike, notification_send_failures
SES account and receivingses_object_lost, ses_sending_paused, ses_rule_missing, ses_identities_90pct
DLQ growthdlq:{queue}, inbound_tempfail
Integrator API downwebhook_failing, webhook_disabled
Parser bugpanics, reports of mis-parsed mail
Compromised keyReport, unusual usage, rpc_owner_mismatch
Abusive identityabuse_pause, mailbox_size
Erasure failureerasure_failed, erasure_overdue
Restore from PITRData corruption, a bad migration, inbound_lost

Every runbook ends by recording what was done in the incident log and checking that the alert resolved.

Bounce spike

  1. Diagnose. Overview and Outbound dashboards: which domain, which identities. Read GET /v1/identities/{id}/messages?status=bounced for samples; group deliveries.smtp_code and bounce_type. Hard bounces from one recipient domain usually mean stale addresses; soft bounces with 4.7.x mean throttling or reputation.
  2. Mitigate. Pause the sending identities (PATCH /v1/identities/{id} {"status": "paused"}) if the integrator is sending to a bad list. Hard bounces already create suppressions (FR-DLV-2). If a recipient provider is throttling, lower identity_daily_send_cap for the affected tenant.
  3. Verify. The bounce rate falls below 2% over the next hour; resume identities.

Complaint spike

  1. Diagnose. Which identities and message kinds. Check that marketing mail carries consent and unsubscribe headers (FR-OUT-8) and that the AI disclosure policy is applied.
  2. Mitigate. Identities above 0.3% complaints over their last 1,000 sends are already paused (FR-DLV-3). Pause the rest of the affected identities; suspend the tenant if the content is abusive (PATCH /v1/tenants/{id} {"status": "suspended"}). Complaint suppressions are permanent.
  3. Verify. No new complaints for 24 hours before resuming, then watch the rate for a week.

Quota exhausted

  1. Diagnose. provider_quota_errors_total by provider_code: E_DAILY_LIMIT_EXCEEDED (daily quota) or E_RATE_LIMIT_EXCEEDED (rate). Cloudflare applies the daily quota per account and raises it automatically over time; it is not exposed to the Worker as a number (Email Service limits page, read 2026-10-09). With PM_DAILY_SEND_QUOTA set to the figure shown in the dashboard, provider_quota_80 warns at 80%; without it, provider_quota fires on the first quota error. Update the variable when Cloudflare raises the quota.
  2. Mitigate. Nothing is lost: definitely-not-sent messages stay queued and back off for up to 24 hours (G3). Request a higher limit from Cloudflare; move urgent domains to SES (Email Sending outage) if they are pre-verified there. For a tenant’s own cap (quota.warning, 429 daily_cap_reached), raise identity_daily_send_cap or tenant_daily_send_cap in the tenant policy.
  3. Verify. transport_outcomes_total{outcome=accepted} resumes; no message reaches failed: quota_exhausted.

Email Sending outage

  1. Diagnose. Rising uncertain and retry outcomes or E_INTERNAL_SERVER_ERROR across all domains; check Cloudflare’s status page. Uncertain messages are never resent automatically (FR-OUT-2).

  2. Mitigate (J5). For each affected domain that has a verified SES identity (domains.ses_identity set and its Easy DKIM records published), switch its transport with a platform key:

    curl -X PATCH https://mail.example.com/v1/domains/dom_01JA… \
      -H "Authorization: Bearer $PYLOTA_MAIL_KEY" -H "Content-Type: application/json" \
      -d '{"transport": "ses"}'
    

    The change is audit-logged and starts a health check, which re-checks alignment for the new transport. Domains without SES fall back per policy only if they are failing; otherwise their mail waits in the queue.

  3. Resolve uncertain sends. After the outage, message.reconciled events settle most of them (FR-DLV-4). For the rest, the integrator checks with the recipient or its own records and calls POST …/messages/{id}/resolve.

  4. Verify. Switch transports back ({"transport": "cloudflare"}) once Cloudflare reports recovery, and watch delivery_events_total.

Domain failing

  1. Diagnose. GET /v1/domains/{domain_id}/health lists issues with the record and fix; GET /v1/domains/{domain_id}/records shows expected versus observed per resolver. A single resolver disagreeing never changes state (H7).
  2. Mitigate. Sending already uses the identity’s platform address (sent_via_fallback, FR-DOM-6). Give the operator the exact record from issues[].fix. For suspended (nameservers, ownership TXT or registration changed), issue a new ownership value with POST /v1/domains/{id}/reprove. For inbound_reject_spike, check whether a domain’s routing points elsewhere or a sender is guessing addresses (dictionary attack); both are visible in inbound_received_total by result. For notification_send_failures, check the platform domain first: system mail (sign-in, invitations, notifications) has no fallback, so while it is failing notification sends fail (reason domain_failing_no_fallback) or are held back (notifications_deferred_total{reason=platform_domain}). The Notifier keeps the items and retries hourly for 24 hours (O25); fixing the platform domain within that time loses nothing. If the platform domain is healthy, follow Email Sending outage.
  3. Verify. POST /v1/domains/{id}/verify twice, a minute apart; the state returns to healthy and domain.recovered is emitted.

SES account and receiving

  1. Diagnose. pmail doctor --check ses shows production access, the sending status, the receipt rule set and the identity count. For ses_object_lost, the ses_ingest row with status lost names the object and recipient: both the SNS push and the SQS backstop failed to get the message ingested before the 14-day lifecycle rule deleted it. Look for pm-inbound dead-letter items and backstop cron errors in that period.
  2. Mitigate.
    • ses_sending_paused: every SES domain already sends through its identities’ platform addresses (FR-DOM-6). Follow AWS’s instructions in the SES console to have sending resumed (the steps are AWS’s; verify them at the time).
    • ses_rule_missing: run pmail setup ses again. It is idempotent, adds pm-deliver to the active rule set and never deactivates another set. Until then, mail to SES domains does not reach the Worker.
    • ses_identities_90pct: the limit of 10,000 identities per Region can be raised only through the AWS account manager (quotas, read 2026-10-09). Ask for it now, or point new customers at nameservers or cloudflare_zone. At 10,000, creating a domain that needs an SES identity fails with 422 transport_unavailable (details.reason = "ses_identity_limit").
    • ses_object_lost: the message cannot be recovered (the sender’s server got a success reply). Tell the affected tenant, then fix why ingestion stalled (DLQ growth, a failing backstop).
  3. Verify. pmail doctor --check ses passes and the alert resolves; for a lost object, new mail to the same recipient arrives.

DLQ growth

  1. Diagnose. pmail dlq list --queue <queue> (GET /v1/platform/dlq?queue=…): the kind and tenant of each item, and the matching error log lines (by request_id or message_id). inbound_tempfail means R2 writes are failing in email() (J1): senders are retrying, nothing is lost.
  2. Mitigate. Fix the cause (a bug: deploy the fix; a dependency outage: wait). Then pmail dlq redrive --queue <queue>, which calls POST /v1/platform/dlq/{dlq_id}/redrive for each open item. Items are kept for 14 days.
  3. Verify. Redriven items leave the open set; dlq:{queue} resolves; for inbound items, the messages appear in their mailboxes.

Integrator API down

  1. Diagnose. GET /v1/webhooks/{webhook_id}/deliveries?status=failed shows error codes (timeout, tls, dns, status_5xx, ssrf_blocked). Deliveries retry for about 72 hours (J4).

  2. Mitigate. Nothing to do while the endpoint is down. After 100 consecutive failures over at least 24 hours the endpoint is disabled. When the integrator is back: re-enable (PATCH /v1/webhooks/{id} {"enabled": true}) and replay what died:

    curl -X POST https://mail.example.com/v1/webhooks/whk_01J9…/replay \
      -H "Authorization: Bearer $PYLOTA_MAIL_KEY" -H "Content-Type: application/json" \
      -d '{"since": "2026-10-08T00:00:00Z", "until": "2026-10-09T00:00:00Z", "status": "dead"}'
    

    Replay reaches back 30 days from each event’s occurred_at (or the tenant’s retention.events_days, if shorter); older events cannot be replayed.

  3. Verify. Replayed deliveries succeed; consumers deduplicate on webhook-id.

Parser bug

  1. Diagnose. Reproduce with the raw message (GET …/messages/{id}/raw, while within raw_days) against crates/conformance; add the case to the corpus with addresses rewritten to RFC 2606 names.
  2. Fix. Release with the fix and an incremented parser_version.
  3. Re-parse (J3). Start a reparse job for each affected tenant with a platform key holding platform:ops: POST /v1/platform/jobs with {"kind": "reparse", "tenant_id": "ten_…", "after": "<first affected date>"} (REST API › Platform operations). It re-parses affected messages from raw with the new parser_version and re-emits their events with reprocessed: true (Inbound). Follow it with GET /v1/platform/jobs/{job_id}. Messages past raw_days cannot be re-parsed and are counted in the job’s result.
  4. Verify. Spot-check re-parsed messages and their message.received events with reprocessed.

Compromised key

  1. Contain (J6). Revoke at once: DELETE /v1/keys/{key_id}. Revoke its descendants too: list keys and revoke every key whose created_by_key_id chain leads to the compromised key (revocation does not cascade).
  2. Investigate. GET /v1/audit-events?actor_key_id=key_… lists the key’s administrative actions. Sends are not audit rows (each is recorded by its message, events and delivery log): list the outbound messages of the identities the key reaches, and query Workers Logs for key_id = <key> over the last 7 days (route, status, message_id). Check webhooks created by the key (URLs pointing somewhere unexpected) and keys it created. If the key held identities:sign, its signature_minted lines name the identities it signed as; an assertion lives at most 10 minutes and a signed request at most 5, and the identity’s private key was never exposed, so no identity key needs rotating for this alone.
  3. Remediate. Rotate integrator secrets that may have been read through the key (webhook secrets with rotate-secret). Cancel queued sends made by the key (POST …/cancel). If the key was a platform key, review every tenant. For rpc_owner_mismatch, treat it as a possible isolation bug: capture the logged IDs and open a private security advisory.
  4. Verify. Requests with the old key return 401 key_revoked.

Abusive identity

  1. Diagnose. identity.paused with reason: abuse_threshold carries the complaint and bounce metrics. Review recent outbound messages and recipients.
  2. Mitigate. Keep the identity paused (inbound continues). Resume only with a tenant or platform key after the cause is fixed (PATCH … {"status": "active"}, audit-logged). For a whole tenant, suspend it. For mailbox_size (above 70% of 10 GB), set retention.message_days for the tenant or split traffic across identities; raw MIME and attachments are already in R2.
  3. Verify. Rates stay below the thresholds for a week after resuming.

Erasure failure

  1. Diagnose. GET /v1/erasure-requests/{id} shows failed and the partial receipt; erasure.failed names the step and error. Find job_step and job_failed log lines by job_id.
  2. Mitigate. Fix the cause (for example a Vectorize or R2 outage), then submit the same erasure again (POST /v1/erasure-requests with the same scope and target). Erasure is idempotent; the new receipt shows what was still left. NFR-PRV-1 counts from the first request, so act within the 24-hour window.
  3. Verify. The new request is completed (or completed_with_holds) with zero probe hits.

Restore from PITR

D1 has Time Travel (30 days on Workers Paid) and SQLite-backed Durable Objects have point-in-time recovery (30 days). R2 has no point-in-time recovery, no object versioning and no bucket replication (PutBucketVersioning and PutBucketReplication are listed as not implemented on the R2 S3 API compatibility page, last updated 2026-07-31, read 2026-10-09). The only copy of a deleted blob is the optional backup bucket (Privacy › R2 backup copy).

  1. Scope. Decide what to restore: D1, one or more mailboxes, or both. Pick the target time T.

  2. Freeze. Suspend affected tenants (PATCH /v1/tenants/{id} {"status": "suspended"}): inbound gets a temporary failure, so senders retry and nothing is lost; sends are refused.

  3. Save what a restore would undo. Before restoring D1, export erasure requests and key revocations made after T: pmail erasure list --json and pmail keys list --json, filtered by time.

  4. Restore D1.

    npx --yes wrangler@4.139.0 d1 time-travel info pylota-mail --timestamp=2026-10-09T09:00:00Z
    npx --yes wrangler@4.139.0 d1 time-travel restore pylota-mail --bookmark=<bookmark>
    

    The restore is destructive and in place, cancels in-flight queries, and prints a bookmark that undoes it; record that bookmark (D1 Time Travel docs, read 2026-10-09).

  5. Restore a mailbox. Inside the object: ctx.storage.getBookmarkForTime(T), then ctx.storage.onNextSessionRestoreBookmark(bookmark) (which returns an undo bookmark), then abort the object so it restarts restored (Durable Objects SQLite storage API, read 2026-10-09). workers-rs 0.8.7 does not wrap these methods (docs.rs, read 2026-10-09), so the restore tooling (P1) adds externs on Storage::as_raw() behind a platform-key-only operator entry point. The PITR API is not available in local development, so drills run on staging.

  6. Reconcile. R2 is not rewound:

    • inbound messages received after T in a restored mailbox still have raw.eml; re-queue their pointers (ingest deduplicates on raw_sha256);
    • outbound messages sent after T lost their rows and their idempotency ledger. The restore tooling lists t/{ten}/i/{idn}/out/ objects uploaded after T whose message is missing from the restored mailbox, and for each re-inserts the message from the stored MIME with status uncertain and flag reprocessed, plus an idempotency row from the object’s idem_key_sha256, fingerprint and operation metadata, with response_json built from the re-inserted row. A retry with the same Idempotency-Key then replays instead of sending again, and a person resolves each uncertain message as usual;
    • re-apply the saved key revocations, then re-submit the saved erasure requests with reason reapply_after_restore:{era_id} (Privacy).
  7. Resume the tenants and run pmail doctor --mail-test.

RPO and RTO (NFR-OPS-2): D1 and Durable Object recovery is continuous, which meets the 1-minute RPO for indexes. R2 objects are written once, before the row that points to them, and deleted only by retention and erasure; R2’s durability covers infrastructure loss, which meets the 15-minute RPO for blobs. Against a bug that deletes objects, nothing protects blobs by default; with PM_BACKUP_BUCKET set, the nightly copy limits the loss to objects created since the last run (RPO 24 hours). The 4-hour RTO is rehearsed in the staging drill.

10. Tests

TestProvesCovers
it::ops::j8_dlq_consumerA message forced into each dead-letter queue is recorded in dlq_items, counted, alerts after 15 minutes of fake time, is listed by GET /v1/platform/dlq without its body, and is redriven by POST /v1/platform/dlq/{dlq_id}/redrive; a non-platform key gets 403J8, FR-OPS-4
it::ops::provider_quota_80With PM_DAILY_SEND_QUOTA set, the evaluator fires provider_quota_80 at 80% of the day’s sends; unset, only the first quota error fires provider_quotaG3
it::ops::restore_rebuilds_ledgerAfter a simulated mailbox restore, a send made after the restore point replays with its original key instead of sending againNFR-OPS-2
it::logs::i5_no_content_in_logsNo canary content or address in any captured log line, including metric linesI5, FR-PRV-6
it::ops::metrics_emittedEach catalogued metric with its labels appears as event = "metric" lines for the flows that emit it; no metric carries a message or identity ID as a labelsection 3
it::ops::alert_evaluator_transitionsFire on a true condition, one audit row, re-notify after 6 h, resolve after two false runssection 5.4
it::ops::health_semantics/health needs no key, touches no binding, returns 503 unavailable with an invalid configuration, and has ses_region exactly when PM_SES_REGION is setsection 7.1
it::ops::ses_alertsses_identities_90pct fires at 9,000 counted identities; ses_sending_paused and ses_rule_missing fire from a fake GetAccount and rule set; the ses doctor check reports the samesection 5.3
it::ses::object_lostA lifecycle-deleted object sets the ledger row to lost, increments ses_object_lost_total and fires ses_object_lostN4, NFR-REL-1
it::inbound::d4_backscatter_droppedbackscatter_total incrementsD4
it::inbound::d5_sender_throttleinbound_throttled_total increments, and 101 throttled messages in an hour meet the inbound_throttled alert conditionD5
it::send::g3_quota_backoffprovider_quota_errors_total increments and the alert condition is metG3
it::delivery::g8_racedelivery_orphaned_total increments after the retry scheduleG8
it::notify::platform_domain_failing_retriesWith the platform domain failing, notification items are kept and retried hourly for 24 hours, notifications_deferred_total{reason=platform_domain} or notifications_failed_total increments, and domain_failing:{domain_id} fires for the platform domainO25
it::notify::daily_capsThe 51st notification for a person in a day goes to the digest and increments notifications_deferred_total{reason=cap_person}O24
core::slo::burn_rate_targetsThe Custom Alert targets in section 5.2 follow from the formulasection 5.2
core::slo::alert_state_machineTransition rules are pure and deterministicsection 5.4
live::ops::metrics_reach_analytics_engineOn staging, metrics written by a send are queryable through the SQL APIsection 3
live::ops::restore_drillStaging drill: D1 Time Travel restore and a mailbox PITR restore complete within 4 hours with the reconcile stepsNFR-OPS-2
it::ops::slo_from_metricsEach SLO row of section 4 (NFR-REL-1 to NFR-REL-4, NFR-PERF-1 to NFR-PERF-6, NFR-PRV-1) is computed by the SLO evaluator from metric lines that a scripted flow emitted, with the expected good and total countssection 4
it::bench::send_api_p951,000 sends through the simulator in workerd: send_api_ms p95 ≤ 500 ms; reports the figure, CI warns aboveNFR-PERF-1
it::bench::queue_to_transport_p951,000 queued sends: outbound_queue_to_transport_ms p95 ≤ 60 sNFR-PERF-2
it::bench::hybrid_p95Hybrid search on the 50,000-message mailbox with the fake AI at the recorded Workers AI latencies: p95 ≤ 800 ms (the real figure comes from staging in M20)NFR-PERF-4
it::bench::tenant_fanout_p95Tenant search over 10 identities: p95 ≤ 1 sNFR-PERF-5
it::bench::agentic_p95Agentic search with the scripted model at recorded latencies: p95 ≤ 8 s, first evidence ≤ 1.5 sNFR-PERF-6
live::slo::inbound_to_webhookOn staging, Gmail and Outlook mail to a webhook endpoint over the live run: p95 ≤ 30 s, p99 ≤ 120 sNFR-REL-3
live::ops::idle_cost_reviewAfter a week of idling on staging, the Cloudflare usage report shows no compute beyond the cron and alarm invocations; recorded in the release notesNFR-COST-1