Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Metrics dictionary — the normative definitions

Every scoreboard field the API serves (GET /workflow/scoreboard) is defined here exactly once: formula, source (data lineage), window semantics, inclusion/exclusion rules, unit, tier availability, and the industry citation it follows. This file is pinned by the meta-test every_scoreboard_field_has_a_dictionary_entry — a scoreboard field cannot ship without its dictionary entry. All rates are integer ten-thousandths (10000 = 100%); per-thousand densities are hundredths; times are seconds; money is cents.

Machine-readable twin: metrics/metrics.json (schema-versioned, scorer_version-stamped) mirrors every scoreboard / report-cadence entry below with the full attribute set structured — name · formula · unit · source table.column · window · inclusion/exclusion · citation · tier availability. The 18 “Server telemetry series” rows further down are doc-only rows and have no JSON twin. Two meta-tests pin the twins together: every_scoreboard_field_has_a_dictionary_entry (code → docs and code → JSON, with a partial docs → code reverse check on _units rows and five allowlisted names) and every_entry_source_table_exists_in_schema (every lineage table exists in src/migration.rs). Benchmarks are quoted as reference points, never claims.

Tier availability: every metric here is available on every deployment tier (T1 solo → T4 global) — tiers are config, not forks; no metric is gated behind a tier.

Posture: documented measurement, not certification. Fields whose data source does not exist in this system are not emitted (no invented telephony/CRM numbers) — see “Deliberately absent” at the end.

Scoreboard fields

FieldDefinition / formulaSource (lineage)WindowCitation
fcr_unitsshare of scored runs with no repeat contact: runs_without_repeat / runs_scored. A recurrence recorded inside the FCR window marks its predecessor as not-first-contact-resolved.workflow_runs.state_json (repeat_contact, prev_contact_age_secs) + fail-closed audit linkage (audit_events)BRAIN_FCR_WINDOW_DAYS repeat-attribution window, default 7 daysSQM-class FCR repeat-window methodology; COPC R8.0 FCR discipline
repeat_contact_rate_unitscomplement of FCR: runs_with_repeat / runs_scored. The primary demand metric — deflection never trades against it.same as fcr_unitssame FCR windowCOPC R8.0; docs/kb-deflection.md
correctness_unitsshare of runs whose recorded findings contain no contradiction/incorrect marker.workflow_runs.state_json.findingslast 1000 runs (scored cohort)ISO 18295-1 process-clause posture; AI Act Art.12 traceability
override_rate_unitsshare of workflow steps where human guidance overrode the engine’s step output.workflow_steps rows derived into StepRowslast 1000 runsHITL law (COMPLIANCE.md); NIST AI RMF
gap_rate_unitsknowledge-gap rate. Currently pinned to 0 in the run-derived scorer — gaps derive from proposals, not runs alone; non-zero emission rides the flywheel release.reserved (proposals tables)—KCS v6 Solve-loop gap capture
abstention_rate_unitsshare of steps where the engine abstained rather than guessed. Higher is honest, not worse.workflow_steps (abstained)last 1000 runsAI Act Art.14 human-oversight posture
guidance_acceptance_unitsaccepted guidance over offered guidance: accepted / (accepted + rejected); SCALE when none offered.workflow_steps (guidance_accepted)last 1000 runsCOPC R8.0 QA calibration discipline
handoff_completeness_unitsshare of runs reaching completed status with an I-PASS-complete handover record.workflow_runs.status + handover packet predicates (src/workflow/relay.rs::packet_missing)last 1000 runsI-PASS handover research; COPC service-level management
justified_handoff_rate_unitsinteger per-mille of recorded control:soft_handoff rows that fired (fires:true) AND carry a non-empty justification; integer floor division, 0 when no rows exist (fail-closed).agent_session_events (control:soft_handoff payloads)all recorded rowsI-PASS handover research; R12 soft-handoff latch law (80% integer rule)
closed_without_closurecount of resolved runs (last 1000 by id) whose persisted case carries NO closure record; an unparsable resolved state counts as without (fail-closed). Post-1.32.7 the A8 gate (“no closure artifact — a case that was not closed with its customer does not close”) makes a new closure-less resolution structurally impossible; the count watches legacy rows and drift.workflow_runs.status + workflow_runs.state_json.closurelast 1000 runsNAM 2015 Improving Diagnosis step 6 (communication of the diagnosis); the closed_looks_good defect-class ban
open_return_contractscount of referral return obligations still open: latest state per contract key over the back_referral session-log rows with status:"open", ordered by deadline_epoch (soonest first — the follow-up queue). Past-deadline opens flip escalated (+ a HITL-queue task with an audited justification) and never auto-resolve; only an operator decision carrying the complete required report releases a contract.agent_session_events (back_referral payloads)all recorded rowsDutch gatekeeping continuity standard (“Closing the Referral Loop: Receipt of Specialist Report”); 1.32.7 Back-Referral law red_flag_handoff_never_blocks_on_back_referral
audit_greenboolean: every scored run references at least one workflow audit row (fail-closed — absence never counts green).audit_events linkage per runlast 1000 runsAI Act Art.12 logging; SOC 2 readiness
escalation_honored_unitsshare of runs where recorded escalation requests were honored (default true only when nothing was requested).workflow_runs.state_json.escalation_honoredlast 1000 runsISO 18295-1 customer-handling clauses
runs_scoredcount of runs in the scored cohort (most recent 1000 by id).workflow_runslast 1000 runs—
return_rate_unitsshare of runs that are return/RMA runs: return_runs / runs_scored.workflow_runs.kind = 'return'last 1000 runsreturns/warranty KPI set (ClaimLane canon)
warranty_claim_rate_unitsshare of runs that are warranty-claim runs: warranty_runs / runs_scored.workflow_runs.kind = 'warranty_claim'last 1000 runsreturns/warranty KPI set
ftfr_unitsFirst-time-fix rate for repair-field work: repair-field runs with no repeat inside the FCR window over all repair-field runs — FCR’s repeat-window method applied to first-VISIT resolution (BRAIN_FCR_WINDOW_DAYS, default 7). The headline field-service economics metric (~1.6 extra dispatches per missed first visit). Empty cohort scores 0 — absence is never dressed up as perfection.workflow_runs.kind='repair_field' + state_json.repeat_contact / prev_contact_age_secssame FCR windowSQM-class repeat-window methodology; FSM FTFR benchmarks
refund_cycle_time_median_secsmedian seconds from run creation to terminal resolution over resolved return/warranty runs; 0 when none resolved.workflow_runs.created_at/updated_at + terminal statuslast 1000 runsrefund cycle-time KPI set
returnless_share_unitsshare of RETURN runs disposed returnless-refund, over all return runs; 0 when the cohort is empty. Returnless refunds pair with fraud review (disposition ranking gates it).workflow_runs.state_json.returnless over kind='return'last 1000 runsreturnless-refund/fraud-detection pairing (2026 practice)
aftersales_fraud_flag_rate_unitsshare of RETURN runs whose deterministic fraud signals flagged them, over all return runs; 0 when the cohort is empty. Signals inform HITL — they never auto-deny.workflow_runs.state_json.fraud_flagged over kind='return'last 1000 runsfraud-signals-feed-HITL posture
goodwill_total_cents_30dsum of amount_cents over APPROVED remedy proposals in the trailing 30 days whose approval audit row verifies (fail-closed — an unaudited remedy never aggregates).proposals kind='complaint_remedy' status='approved' × audit_events target/detail hash linkagetrailing 30 daysISO 10002 remedy discipline; goodwill-ledger-visible posture
goodwill_entries_30dcount of audited approved remedies in the window.same as goodwill_total_cents_30dtrailing 30 days—
goodwill_unaudited_excluded_30dapproved remedies EXCLUDED for missing audit linkage — surfaced, never folded away.same scantrailing 30 daysfail-closed evidence law
voc_contacts_totalcount of workflow runs — the contact volume the VoC ratio denominates.workflow_runsrolling (all runs)ISO 10004 satisfaction monitoring as data
voc_complaints_totalcount of runs with kind = 'complaint'.workflow_runs.kindrolling (all runs)ISO 10002 register posture
voc_complaints_per_thousand_contacts_unitscomplaints * 100_000 / max(contacts, 1) — complaints per thousand contacts in hundredths (per-mille × 100). Zero contacts score 0. The CSAT/DSAT instruments stay CRM-side; this is the lineage-derived complaint-density twin.workflow_runs.kind countsrolling (all runs)ISO 10004 §complaint-per-thousand-contacts KPI canon

Report-cadence fields (same read, weekly report ride)

FieldDefinition / formulaSource (lineage)WindowCitation
calibration_report_emittedtrue when THIS read crossed the weekly boundary and landed a machine-generated CalibrationRecord on the audit chain.src/workflow/calibration.rsweekly cadencemonthly signed-recalibration posture (COMPLIANCE.md)
kcs_linkage_rate_unitsshare of published knowledge linked from closed-run evidence.src/workflow/kcs.rs::kcs_measuresrolling (all governed articles)KCS v6 Evolve loop
searched_found_rate_unitsshare of recall/search events that ended in a found article (SIR proxy).kcs_measuresrollingKCS v6 Solve loop (search-and-solve)
article_freshness_median_age_secsmedian age in seconds since last review across governed articles.kcs_measuresrollingKCS v6 article-health
self_service_deflection_unitsINDICATIVE deflection from on-page KB feedback (solved-proofs over total feedback). Never traded against repeat_contact_rate_units.kcs::kb_feedback_measuresrollingdocs/kb-deflection.md governs; KCS v6 self-service
kb_feedback_totaltotal on-page feedback events counted.kcs::kb_feedback_measuresrolling—
kb_hot_topicstop slugs by feedback count above KB_HOT_TOPIC_THRESHOLD.kcs::kb_hot_topicsrollingKCS v6 Evolve (content-defect queue)
reask_ratere-ask events (case/reask) ÷ closed cases, in hundredths. Three deterministic sources emit the event: crm_merge (Zendesk/Salesforce merges, Genesys reopens via the Bridges sync), marked (operator reask note / brain workflow note --reask), derived (duplicate-open heuristic — exact hashed-subject match within BRAIN_REASK_WINDOW_DAYS, default 3 days, HITL-gated as case_merge_suggested; the approved merge emits). No fuzzy matching; no surveys.outbox topic='case/reask', workflow_runs.statusrolling; window semantics per BRAIN_REASK_WINDOW_DAYSCXC customer-effort canon (effort-proxy dimension); Keystone v1.28.36

Approval-fatigue telemetry (ASI09, Attestation v1.28.62)

The reviewer’s own anti-rubber-stamp detector (the console’s calibration strip, client/src/panels/review.rs rubber_stamp()) computed SERVER-SIDE so the DPO sees the signal on the scoreboard, not only in one reviewer’s console. The window and the sample cap mirror the client’s fetch exactly (trailing 7 days on created_at, latest 200 per status), and the verdict is pinned against the client arithmetic by scoreboard_uniformity_matches_client_math — the scoreboard and the reviewer’s console can never disagree.

FieldDefinition / formulaSource (lineage)WindowCitation
review_independence_risk1 when approve_rate > 0.9 AND decisions >= 20 over the windowed sample (the client detector’s verdict, verbatim arithmetic); else 0. An empty window scores 0 — absence is never dressed up as either safety or risk.proposals.status, proposals.created_at (decided proposals only)trailing 7 days, latest 200 decisions per statusASI09 approval-fatigue posture; COPC R8.0 QA calibration discipline
approval_uniformity_ratioapproved ÷ (approved + rejected) over the same sample, integer ten-thousandths (truncating; 10000 = 100%). Shows HOW near uniform, not just the binary risk.same sample as review_independence_risksame windowASI09 (Attestation v1.28.62); parity-pinned to the client arithmetic
review_decisions_windowapproved + rejected in the uniformity sample — the denominator context that makes the two signals above interpretable.same samplesame windowASI09 (Attestation v1.28.62)

Derived proxy (planned scorer integration)

customer_effort_events — a deterministic CES proxy per case computed from the lineage: repeat contacts × channel switches × handovers × re-asks (case/reask, weighted like a repeat — see frontdesk::effort_proxy: score = repeats×2 + switches×1 + handovers×3 + re_asks×2). No survey instrument exists here (VoC surveys stay CRM-side per ISO 10004); this is the lineage-derived twin. It lands as a scored dimension in the next scorer version with gold-set families extended; until then it is defined here so the formula is fixed before any code emits it.

Metric versioning (the scorer_version discipline)

The dictionary is versioned with the scorer: SCORER_VERSION (in crates/brain-engine-sdk/src/pure/calibration.rs, re-exported as CALIBRATION_SCORER_VERSION) stamps every CalibrationRecord on the audit chain, every gold-pack case (crates/gold-sets — GoldCase::validate fail-closed rejects a mismatched pack), and metrics/metrics.json. A formula change bumps the version, this file, the JSON twin, and the gold-pack expectations together, in one PR — pinned by the meta-test formula_change_bumps_scorer_version.

Server telemetry series (/metrics + /health/db)

The Prometheus text surface (GET /metrics, Read-gated) and the /health/db JSON carry the server’s own telemetry. Every emitted series carries a dictionary row here — pinned by the meta-test metrics_series_have_dictionary_rows (a series cannot ship without a row, the scoreboard discipline applied to ops telemetry). Counters are process-local (single-process truth since process start; multi-site aggregation remains Parcels federation). Gauges are scrape-time snapshots.

SeriesTypeDefinitionSource
brain_rss_mibgaugeProcess resident set in MiB — the same measurement the /health/db capacity block reports (capacity.rss_mib; /health itself returns only {status, version}), NOT whole-host memoryhttp_limit::process_rss_mib
brain_pool_connectionsgaugeGlobal pool connection counts by state label (idle/busy)r2d2::Pool::state() at scrape
brain_pool_in_usegaugePer-domain pool connections in use (connections − idle) — the pool-saturation signal under the concurrent benchr2d2::Pool::state() per registered domain
brain_pool_idlegaugePer-domain pool idle connectionsr2d2::Pool::state() per registered domain
brain_pool_timeouts_totalcounterPool checkouts that timed out (r2d2 get() failure) — counted at the existing handler error seam (HandlerError::db_down) and the workflow lane’s checkout arm; zero cost on success pathsconcurrency::CONCURRENCY
brain_busy_errors_totalcounterSQLITE_BUSY-family errors observed at the governed-write BEGIN sites (WorkflowTx::begin + the workflow lane’s BEGIN IMMEDIATE) — write contention after the 5 s busy_timeout burn, counted where the error arm already propagatesconcurrency::CONCURRENCY
brain_wal_pages_pendinggaugeWAL frames not yet checkpointed, per domain (log − checkpointed from the PASSIVE checkpoint row). The PRAGMA runs ONLY inside /health/db (cold path); /metrics reports the last snapshot — absent domains have no snapshot yet/health/db WAL sweep → concurrency::CONCURRENCY
brain_delivery_intents_pendinggaugeDelivery-family outbox rows sitting pending, per domain — non-zero reads as “awaiting its crank”: the /due crank (POST /workflow/delivery/due) drains them in bounded batches, so a value that never falls between cranks is the alarm, not the value itself. It exists so a LOST intent is distinguishable from one not yet crankedconnector::delivery::pending_intent_census at scrape
brain_delivery_untrusted_rows_pendinggaugeDelivery-family outbox rows sitting pending whose idempotency key is NOT a kernel ddl-intent- mint, per domain. Unlike the intent gauge, a non-zero value is NOT expected: it means the reserved topic root was written without the mintersame census, classified through delivery_intents::intent_kind
brain_lock_wait_micros_p50gaugeBucket-quantile (lower edge, µs) of contended lock-acquire waits across the instrumented request-path Mutex/RwLock holders (token store, rate limiter, replay cache, audit chain keys, domain registry, embed/rerank/screen models, …). Only CONTENDED acquires are recorded (try_lock fast path costs nothing), so 0 = no contention observed, never “gauge wired off”. Honest scope: the workflow lane’s mutex is NOT wait-instrumented (a plain lock()), and neither are the mcp binary, the connector token cache, nor the scrape-path locks. The value is a histogram bucket lower edge over the fixed µs edges in concurrency::LOCK_WAIT_BUCKET_EDGES_US — a deterministic read, not an interpolated percentile; moving an edge is a dictionary-visible changeconcurrency::CONCURRENCY.lock_wait_histogram()
brain_lock_wait_micros_p95gaugeThe p95 twin of brain_lock_wait_micros_p50 — same histogram, same edges, same contended-only recordingconcurrency::CONCURRENCY.lock_wait_histogram()
brain_capacity_statusgaugeCapacity posture: 0=unknown (the capacity could not be measured, e.g. pool exhausted), 1=ok, 2=warning, 3=exceededcapacity::classify
brain_audit_chain_okgauge1 = every registered domain’s audit chain verifies; 0 = tamper detected (TTL-cached; authoritative answer on /audit/verify)audit::verify_chain
brain_db_busy_totalcounterSQLITE_BUSY events surfaced at the audit seam specifically (audit-tx settle failures after busy_timeout burn-through) — the narrower audit-seam twin of brain_busy_errors_totalaudit::busy_hits()
brain_jwt_azp_rejected_totalcounterAccess tokens REFUSED because their RFC 7519 §4.1.3 azp claim was absent or named a different application (the token-intent / confused-deputy class). Only ever non-zero when BRAIN_JWT_AZP is configured, so a rising series is a positive statement that the control is live and biting; 0 is ambiguous between “not configured” and “nothing refused”, and the boot line (P63.2 disclosure) is what states the posture, not this series. Mints from /auth/refresh carry the verified azp forward, so rotation cannot trip thisauth::jwt::azp_rejections()
brain_model_calls_totalcounterModel-surface operations since process start, labelled class — the closed DecisionClass census. open_generate = LLM provider streams, classify = injection-screen calls, encode = texts submitted to the embedder (not batches: embed cost scales with texts). The class is a property of the CALL SITE, never of the call’s content; the label is a total function of a fieldless enum, so it carries no model id, prompt, principal or domain. Every declared class emits a row, including classes at zero — a dashboard must never confuse “nothing happened” with “not instrumented”. Process-local: a restart zeroes itdecision_class::note_call
brain_model_tokens_totalcounterProvider-reported tokens (input + output) by class, folded at the SAME observation seam that updates the exchange budget’s enforced total — one path, one number, never two meters that can drift. classify and encode are always 0 and that is the honest reading: neither surface reports token usage, and this tree deliberately does not substitute a proxy. (Reporting embedding dimensions as tokens is a real defect elsewhere in the tree, reported not fixed — see the R53a evidence §7.) Process-localdecision_class::note_tokens
brain_model_incomplete_totalcounterCalls that started and ended without a MessageEnd, so their spend is UNKNOWN, not zero, by class. Ships because the plan did not ask for it: without it the call count silently under-counts, and a reader dividing by it later would be dividing by a denominator with invisible holes. A rising series is a positive statement that spend is going unaccounted — it is not a health signaldecision_class::note_incomplete

/health/db JSON additive keys (v1.28.58): concurrency.pool_timeouts_total, concurrency.busy_errors_total, and concurrency.wal_pages_pending (a {domain: frames} object) — the same numbers as the series above. /health/db additive keys (v1.28.59): durability.synchronous (full|normal), durability.wal_autocheckpoint_pages, and durability.capacity_target (desktop|jetson) — the static boot-time echo of the connection-init policy (envelope defaults ⊕ the fail-closed BRAIN_SYNCHRONOUS / BRAIN_WAL_AUTOCHECKPOINT overrides), never a per-request pragma read. /health/db additive keys (v1.28.80): authn.enabled (a token resolves), authn.required (BRAIN_REQUIRE_AUTH=1), and hardening.allow_policy_bypasses (ingests unscreened under INJECTION_POLICY=allow, monotonic).

Deliberately absent (scope guards)

  • AHT decomposition (talk + hold + ACW): appears only when CRM data provides the components; no telephony feed exists today.
  • Abandonment rate: requires a telephony feed; absent until one exists.
  • CSAT/VoC: instruments stay CRM-side (ISO 10004); only lineage-derived proxies live here.
  • No forecasting/scheduling/capacity metrics: WFM alignment is interop (GET/POST /ops/shifts, GET /ops/skills), not reimplementation.