Metrics dictionary — the normative definitions
Every scoreboard field the API serves (GET /workflow/scoreboard) is defined
here exactly once: formula, source (data lineage), window semantics,
inclusion/exclusion rules, unit, tier availability, and the industry citation
it follows. This file is pinned by the meta-test
every_scoreboard_field_has_a_dictionary_entry — a scoreboard field cannot ship
without its dictionary entry. All rates are integer ten-thousandths
(10000 = 100%); per-thousand densities are hundredths; times are seconds;
money is cents.
Machine-readable twin: metrics/metrics.json
(schema-versioned, scorer_version-stamped) mirrors every scoreboard /
report-cadence entry below with the full attribute set structured — name ·
formula · unit · source table.column · window · inclusion/exclusion · citation ·
tier availability. The 18 “Server telemetry series” rows further down are
doc-only rows and have no JSON twin. Two meta-tests pin the twins together:
every_scoreboard_field_has_a_dictionary_entry (code → docs and code → JSON,
with a partial docs → code reverse check on _units rows and five allowlisted
names) and every_entry_source_table_exists_in_schema (every lineage table
exists in src/migration.rs). Benchmarks are quoted as reference points, never
claims.
Tier availability: every metric here is available on every deployment tier (T1 solo → T4 global) — tiers are config, not forks; no metric is gated behind a tier.
Posture: documented measurement, not certification. Fields whose data source does not exist in this system are not emitted (no invented telephony/CRM numbers) — see “Deliberately absent” at the end.
Scoreboard fields
| Field | Definition / formula | Source (lineage) | Window | Citation |
|---|---|---|---|---|
fcr_units | share of scored runs with no repeat contact: runs_without_repeat / runs_scored. A recurrence recorded inside the FCR window marks its predecessor as not-first-contact-resolved. | workflow_runs.state_json (repeat_contact, prev_contact_age_secs) + fail-closed audit linkage (audit_events) | BRAIN_FCR_WINDOW_DAYS repeat-attribution window, default 7 days | SQM-class FCR repeat-window methodology; COPC R8.0 FCR discipline |
repeat_contact_rate_units | complement of FCR: runs_with_repeat / runs_scored. The primary demand metric — deflection never trades against it. | same as fcr_units | same FCR window | COPC R8.0; docs/kb-deflection.md |
correctness_units | share of runs whose recorded findings contain no contradiction/incorrect marker. | workflow_runs.state_json.findings | last 1000 runs (scored cohort) | ISO 18295-1 process-clause posture; AI Act Art.12 traceability |
override_rate_units | share of workflow steps where human guidance overrode the engine’s step output. | workflow_steps rows derived into StepRows | last 1000 runs | HITL law (COMPLIANCE.md); NIST AI RMF |
gap_rate_units | knowledge-gap rate. Currently pinned to 0 in the run-derived scorer — gaps derive from proposals, not runs alone; non-zero emission rides the flywheel release. | reserved (proposals tables) | — | KCS v6 Solve-loop gap capture |
abstention_rate_units | share of steps where the engine abstained rather than guessed. Higher is honest, not worse. | workflow_steps (abstained) | last 1000 runs | AI Act Art.14 human-oversight posture |
guidance_acceptance_units | accepted guidance over offered guidance: accepted / (accepted + rejected); SCALE when none offered. | workflow_steps (guidance_accepted) | last 1000 runs | COPC R8.0 QA calibration discipline |
handoff_completeness_units | share of runs reaching completed status with an I-PASS-complete handover record. | workflow_runs.status + handover packet predicates (src/workflow/relay.rs::packet_missing) | last 1000 runs | I-PASS handover research; COPC service-level management |
justified_handoff_rate_units | integer per-mille of recorded control:soft_handoff rows that fired (fires:true) AND carry a non-empty justification; integer floor division, 0 when no rows exist (fail-closed). | agent_session_events (control:soft_handoff payloads) | all recorded rows | I-PASS handover research; R12 soft-handoff latch law (80% integer rule) |
closed_without_closure | count of resolved runs (last 1000 by id) whose persisted case carries NO closure record; an unparsable resolved state counts as without (fail-closed). Post-1.32.7 the A8 gate (“no closure artifact — a case that was not closed with its customer does not close”) makes a new closure-less resolution structurally impossible; the count watches legacy rows and drift. | workflow_runs.status + workflow_runs.state_json.closure | last 1000 runs | NAM 2015 Improving Diagnosis step 6 (communication of the diagnosis); the closed_looks_good defect-class ban |
open_return_contracts | count of referral return obligations still open: latest state per contract key over the back_referral session-log rows with status:"open", ordered by deadline_epoch (soonest first — the follow-up queue). Past-deadline opens flip escalated (+ a HITL-queue task with an audited justification) and never auto-resolve; only an operator decision carrying the complete required report releases a contract. | agent_session_events (back_referral payloads) | all recorded rows | Dutch gatekeeping continuity standard (“Closing the Referral Loop: Receipt of Specialist Report”); 1.32.7 Back-Referral law red_flag_handoff_never_blocks_on_back_referral |
audit_green | boolean: every scored run references at least one workflow audit row (fail-closed — absence never counts green). | audit_events linkage per run | last 1000 runs | AI Act Art.12 logging; SOC 2 readiness |
escalation_honored_units | share of runs where recorded escalation requests were honored (default true only when nothing was requested). | workflow_runs.state_json.escalation_honored | last 1000 runs | ISO 18295-1 customer-handling clauses |
runs_scored | count of runs in the scored cohort (most recent 1000 by id). | workflow_runs | last 1000 runs | — |
return_rate_units | share of runs that are return/RMA runs: return_runs / runs_scored. | workflow_runs.kind = 'return' | last 1000 runs | returns/warranty KPI set (ClaimLane canon) |
warranty_claim_rate_units | share of runs that are warranty-claim runs: warranty_runs / runs_scored. | workflow_runs.kind = 'warranty_claim' | last 1000 runs | returns/warranty KPI set |
ftfr_units | First-time-fix rate for repair-field work: repair-field runs with no repeat inside the FCR window over all repair-field runs — FCR’s repeat-window method applied to first-VISIT resolution (BRAIN_FCR_WINDOW_DAYS, default 7). The headline field-service economics metric (~1.6 extra dispatches per missed first visit). Empty cohort scores 0 — absence is never dressed up as perfection. | workflow_runs.kind='repair_field' + state_json.repeat_contact / prev_contact_age_secs | same FCR window | SQM-class repeat-window methodology; FSM FTFR benchmarks |
refund_cycle_time_median_secs | median seconds from run creation to terminal resolution over resolved return/warranty runs; 0 when none resolved. | workflow_runs.created_at/updated_at + terminal status | last 1000 runs | refund cycle-time KPI set |
returnless_share_units | share of RETURN runs disposed returnless-refund, over all return runs; 0 when the cohort is empty. Returnless refunds pair with fraud review (disposition ranking gates it). | workflow_runs.state_json.returnless over kind='return' | last 1000 runs | returnless-refund/fraud-detection pairing (2026 practice) |
aftersales_fraud_flag_rate_units | share of RETURN runs whose deterministic fraud signals flagged them, over all return runs; 0 when the cohort is empty. Signals inform HITL — they never auto-deny. | workflow_runs.state_json.fraud_flagged over kind='return' | last 1000 runs | fraud-signals-feed-HITL posture |
goodwill_total_cents_30d | sum of amount_cents over APPROVED remedy proposals in the trailing 30 days whose approval audit row verifies (fail-closed — an unaudited remedy never aggregates). | proposals kind='complaint_remedy' status='approved' × audit_events target/detail hash linkage | trailing 30 days | ISO 10002 remedy discipline; goodwill-ledger-visible posture |
goodwill_entries_30d | count of audited approved remedies in the window. | same as goodwill_total_cents_30d | trailing 30 days | — |
goodwill_unaudited_excluded_30d | approved remedies EXCLUDED for missing audit linkage — surfaced, never folded away. | same scan | trailing 30 days | fail-closed evidence law |
voc_contacts_total | count of workflow runs — the contact volume the VoC ratio denominates. | workflow_runs | rolling (all runs) | ISO 10004 satisfaction monitoring as data |
voc_complaints_total | count of runs with kind = 'complaint'. | workflow_runs.kind | rolling (all runs) | ISO 10002 register posture |
voc_complaints_per_thousand_contacts_units | complaints * 100_000 / max(contacts, 1) — complaints per thousand contacts in hundredths (per-mille × 100). Zero contacts score 0. The CSAT/DSAT instruments stay CRM-side; this is the lineage-derived complaint-density twin. | workflow_runs.kind counts | rolling (all runs) | ISO 10004 §complaint-per-thousand-contacts KPI canon |
Report-cadence fields (same read, weekly report ride)
| Field | Definition / formula | Source (lineage) | Window | Citation |
|---|---|---|---|---|
calibration_report_emitted | true when THIS read crossed the weekly boundary and landed a machine-generated CalibrationRecord on the audit chain. | src/workflow/calibration.rs | weekly cadence | monthly signed-recalibration posture (COMPLIANCE.md) |
kcs_linkage_rate_units | share of published knowledge linked from closed-run evidence. | src/workflow/kcs.rs::kcs_measures | rolling (all governed articles) | KCS v6 Evolve loop |
searched_found_rate_units | share of recall/search events that ended in a found article (SIR proxy). | kcs_measures | rolling | KCS v6 Solve loop (search-and-solve) |
article_freshness_median_age_secs | median age in seconds since last review across governed articles. | kcs_measures | rolling | KCS v6 article-health |
self_service_deflection_units | INDICATIVE deflection from on-page KB feedback (solved-proofs over total feedback). Never traded against repeat_contact_rate_units. | kcs::kb_feedback_measures | rolling | docs/kb-deflection.md governs; KCS v6 self-service |
kb_feedback_total | total on-page feedback events counted. | kcs::kb_feedback_measures | rolling | — |
kb_hot_topics | top slugs by feedback count above KB_HOT_TOPIC_THRESHOLD. | kcs::kb_hot_topics | rolling | KCS v6 Evolve (content-defect queue) |
reask_rate | re-ask events (case/reask) ÷ closed cases, in hundredths. Three deterministic sources emit the event: crm_merge (Zendesk/Salesforce merges, Genesys reopens via the Bridges sync), marked (operator reask note / brain workflow note --reask), derived (duplicate-open heuristic — exact hashed-subject match within BRAIN_REASK_WINDOW_DAYS, default 3 days, HITL-gated as case_merge_suggested; the approved merge emits). No fuzzy matching; no surveys. | outbox topic='case/reask', workflow_runs.status | rolling; window semantics per BRAIN_REASK_WINDOW_DAYS | CXC customer-effort canon (effort-proxy dimension); Keystone v1.28.36 |
Approval-fatigue telemetry (ASI09, Attestation v1.28.62)
The reviewer’s own anti-rubber-stamp detector (the console’s calibration
strip, client/src/panels/review.rs rubber_stamp()) computed SERVER-SIDE
so the DPO sees the signal on the scoreboard, not only in one reviewer’s
console. The window and the sample cap mirror the client’s fetch exactly
(trailing 7 days on created_at, latest 200 per status), and the verdict is
pinned against the client arithmetic by
scoreboard_uniformity_matches_client_math — the scoreboard and the
reviewer’s console can never disagree.
| Field | Definition / formula | Source (lineage) | Window | Citation |
|---|---|---|---|---|
review_independence_risk | 1 when approve_rate > 0.9 AND decisions >= 20 over the windowed sample (the client detector’s verdict, verbatim arithmetic); else 0. An empty window scores 0 — absence is never dressed up as either safety or risk. | proposals.status, proposals.created_at (decided proposals only) | trailing 7 days, latest 200 decisions per status | ASI09 approval-fatigue posture; COPC R8.0 QA calibration discipline |
approval_uniformity_ratio | approved ÷ (approved + rejected) over the same sample, integer ten-thousandths (truncating; 10000 = 100%). Shows HOW near uniform, not just the binary risk. | same sample as review_independence_risk | same window | ASI09 (Attestation v1.28.62); parity-pinned to the client arithmetic |
review_decisions_window | approved + rejected in the uniformity sample — the denominator context that makes the two signals above interpretable. | same sample | same window | ASI09 (Attestation v1.28.62) |
Derived proxy (planned scorer integration)
customer_effort_events — a deterministic CES proxy per case computed
from the lineage: repeat contacts × channel switches × handovers ×
re-asks (case/reask, weighted like a repeat — see frontdesk::effort_proxy:
score = repeats×2 + switches×1 + handovers×3 + re_asks×2). No survey
instrument exists here (VoC surveys stay CRM-side per ISO 10004); this is the
lineage-derived twin. It lands as a scored dimension in the next scorer
version with gold-set families extended; until then it is defined here so the
formula is fixed before any code emits it.
Metric versioning (the scorer_version discipline)
The dictionary is versioned with the scorer: SCORER_VERSION (in
crates/brain-engine-sdk/src/pure/calibration.rs, re-exported as
CALIBRATION_SCORER_VERSION) stamps every CalibrationRecord on the audit
chain, every gold-pack case (crates/gold-sets — GoldCase::validate
fail-closed rejects a mismatched pack), and metrics/metrics.json. A
formula change bumps the version, this file, the JSON twin, and the gold-pack
expectations together, in one PR — pinned by the meta-test
formula_change_bumps_scorer_version.
Server telemetry series (/metrics + /health/db)
The Prometheus text surface (GET /metrics, Read-gated) and the
/health/db JSON carry the server’s own telemetry. Every emitted series
carries a dictionary row here — pinned by the meta-test
metrics_series_have_dictionary_rows (a series cannot ship without a
row, the scoreboard discipline applied to ops telemetry). Counters are
process-local (single-process truth since process start; multi-site
aggregation remains Parcels federation). Gauges are scrape-time snapshots.
| Series | Type | Definition | Source |
|---|---|---|---|
brain_rss_mib | gauge | Process resident set in MiB — the same measurement the /health/db capacity block reports (capacity.rss_mib; /health itself returns only {status, version}), NOT whole-host memory | http_limit::process_rss_mib |
brain_pool_connections | gauge | Global pool connection counts by state label (idle/busy) | r2d2::Pool::state() at scrape |
brain_pool_in_use | gauge | Per-domain pool connections in use (connections − idle) — the pool-saturation signal under the concurrent bench | r2d2::Pool::state() per registered domain |
brain_pool_idle | gauge | Per-domain pool idle connections | r2d2::Pool::state() per registered domain |
brain_pool_timeouts_total | counter | Pool checkouts that timed out (r2d2 get() failure) — counted at the existing handler error seam (HandlerError::db_down) and the workflow lane’s checkout arm; zero cost on success paths | concurrency::CONCURRENCY |
brain_busy_errors_total | counter | SQLITE_BUSY-family errors observed at the governed-write BEGIN sites (WorkflowTx::begin + the workflow lane’s BEGIN IMMEDIATE) — write contention after the 5 s busy_timeout burn, counted where the error arm already propagates | concurrency::CONCURRENCY |
brain_wal_pages_pending | gauge | WAL frames not yet checkpointed, per domain (log − checkpointed from the PASSIVE checkpoint row). The PRAGMA runs ONLY inside /health/db (cold path); /metrics reports the last snapshot — absent domains have no snapshot yet | /health/db WAL sweep → concurrency::CONCURRENCY |
brain_delivery_intents_pending | gauge | Delivery-family outbox rows sitting pending, per domain — non-zero reads as “awaiting its crank”: the /due crank (POST /workflow/delivery/due) drains them in bounded batches, so a value that never falls between cranks is the alarm, not the value itself. It exists so a LOST intent is distinguishable from one not yet cranked | connector::delivery::pending_intent_census at scrape |
brain_delivery_untrusted_rows_pending | gauge | Delivery-family outbox rows sitting pending whose idempotency key is NOT a kernel ddl-intent- mint, per domain. Unlike the intent gauge, a non-zero value is NOT expected: it means the reserved topic root was written without the minter | same census, classified through delivery_intents::intent_kind |
brain_lock_wait_micros_p50 | gauge | Bucket-quantile (lower edge, µs) of contended lock-acquire waits across the instrumented request-path Mutex/RwLock holders (token store, rate limiter, replay cache, audit chain keys, domain registry, embed/rerank/screen models, …). Only CONTENDED acquires are recorded (try_lock fast path costs nothing), so 0 = no contention observed, never “gauge wired off”. Honest scope: the workflow lane’s mutex is NOT wait-instrumented (a plain lock()), and neither are the mcp binary, the connector token cache, nor the scrape-path locks. The value is a histogram bucket lower edge over the fixed µs edges in concurrency::LOCK_WAIT_BUCKET_EDGES_US — a deterministic read, not an interpolated percentile; moving an edge is a dictionary-visible change | concurrency::CONCURRENCY.lock_wait_histogram() |
brain_lock_wait_micros_p95 | gauge | The p95 twin of brain_lock_wait_micros_p50 — same histogram, same edges, same contended-only recording | concurrency::CONCURRENCY.lock_wait_histogram() |
brain_capacity_status | gauge | Capacity posture: 0=unknown (the capacity could not be measured, e.g. pool exhausted), 1=ok, 2=warning, 3=exceeded | capacity::classify |
brain_audit_chain_ok | gauge | 1 = every registered domain’s audit chain verifies; 0 = tamper detected (TTL-cached; authoritative answer on /audit/verify) | audit::verify_chain |
brain_db_busy_total | counter | SQLITE_BUSY events surfaced at the audit seam specifically (audit-tx settle failures after busy_timeout burn-through) — the narrower audit-seam twin of brain_busy_errors_total | audit::busy_hits() |
brain_jwt_azp_rejected_total | counter | Access tokens REFUSED because their RFC 7519 §4.1.3 azp claim was absent or named a different application (the token-intent / confused-deputy class). Only ever non-zero when BRAIN_JWT_AZP is configured, so a rising series is a positive statement that the control is live and biting; 0 is ambiguous between “not configured” and “nothing refused”, and the boot line (P63.2 disclosure) is what states the posture, not this series. Mints from /auth/refresh carry the verified azp forward, so rotation cannot trip this | auth::jwt::azp_rejections() |
brain_model_calls_total | counter | Model-surface operations since process start, labelled class — the closed DecisionClass census. open_generate = LLM provider streams, classify = injection-screen calls, encode = texts submitted to the embedder (not batches: embed cost scales with texts). The class is a property of the CALL SITE, never of the call’s content; the label is a total function of a fieldless enum, so it carries no model id, prompt, principal or domain. Every declared class emits a row, including classes at zero — a dashboard must never confuse “nothing happened” with “not instrumented”. Process-local: a restart zeroes it | decision_class::note_call |
brain_model_tokens_total | counter | Provider-reported tokens (input + output) by class, folded at the SAME observation seam that updates the exchange budget’s enforced total — one path, one number, never two meters that can drift. classify and encode are always 0 and that is the honest reading: neither surface reports token usage, and this tree deliberately does not substitute a proxy. (Reporting embedding dimensions as tokens is a real defect elsewhere in the tree, reported not fixed — see the R53a evidence §7.) Process-local | decision_class::note_tokens |
brain_model_incomplete_total | counter | Calls that started and ended without a MessageEnd, so their spend is UNKNOWN, not zero, by class. Ships because the plan did not ask for it: without it the call count silently under-counts, and a reader dividing by it later would be dividing by a denominator with invisible holes. A rising series is a positive statement that spend is going unaccounted — it is not a health signal | decision_class::note_incomplete |
/health/db JSON additive keys (v1.28.58): concurrency.pool_timeouts_total,
concurrency.busy_errors_total, and concurrency.wal_pages_pending (a
{domain: frames} object) — the same numbers as the series above.
/health/db additive keys (v1.28.59): durability.synchronous
(full|normal), durability.wal_autocheckpoint_pages, and
durability.capacity_target (desktop|jetson) — the static boot-time
echo of the connection-init policy (envelope defaults ⊕ the fail-closed
BRAIN_SYNCHRONOUS / BRAIN_WAL_AUTOCHECKPOINT overrides), never a
per-request pragma read.
/health/db additive keys (v1.28.80): authn.enabled (a token resolves),
authn.required (BRAIN_REQUIRE_AUTH=1), and hardening.allow_policy_bypasses
(ingests unscreened under INJECTION_POLICY=allow, monotonic).
Deliberately absent (scope guards)
- AHT decomposition (talk + hold + ACW): appears only when CRM data provides the components; no telephony feed exists today.
- Abandonment rate: requires a telephony feed; absent until one exists.
- CSAT/VoC: instruments stay CRM-side (ISO 10004); only lineage-derived proxies live here.
- No forecasting/scheduling/capacity metrics: WFM alignment is interop
(
GET/POST /ops/shifts,GET /ops/skills), not reimplementation.