Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Deployment

Brain Server is designed to run as a persistent, self-managed service on a single host. This page covers installing it, configuring it, keeping it healthy, and backing it up.


Service install (macOS)

scripts/install-service.sh builds the release binaries, installs them to ~/.local/bin, relocates the auth token from the launchd plist into a 0600 secret file, restarts the service, and waits for /health. It is idempotent.

scripts/install-service.sh

This installs:

  • brain-server — the server (launchd-managed, KeepAlive=true, RunAtLoad=true).
  • brain — the operator CLI (status, query, explain, ingest-dir, reconcile, resolve, backup, …).
  • mcp — the MCP bridge (search/recall/ingest as MCP tools).
  • bench — the latency/recall harness.
  • brain-migrate-rehearse — migration rehearsal / recovery.
  • brain-connector-stub (and brain-connector-gh when the feature is enabled).

Optional: brain-connector-crm (feature connector-crm) is built best-effort by install-service.sh — present only when the feature was enabled for a prior build; the script compiles it on the first run that needs it and skips cleanly otherwise, same posture as brain-connector-gh. The cron recipes in CRM case intake below need it installed.

macOS note: newly copied executables can get a com.apple.provenance xattr that Gatekeeper uses to SIGKILL on first exec (exit 137). The install script strips it. A manual cp does not.


Configuration

Brain Server is configured through environment variables (all resolved in src/config.rs). The most important:

VariableDefaultDescription
BIND_HOST127.0.0.1Bind address. 0.0.0.0 without BIND_PUBLIC logs a loud warning and still binds (the opt-in is env presence — any value counts); an unparseable host without BIND_PUBLIC refuses boot; any non-loopback bind with no auth configured refuses boot
BIND_PORT8765Listen port. Fail-closed (R70/F8-10): a present-and-malformed value refuses boot with a message naming the key, the value and the range. It used to be .parse().unwrap_or(8765), so a typo silently bound the production port. 0 is refused specifically — it parses, but port 0 binds a kernel-chosen ephemeral port that changes every restart. Unset (or empty) still binds 8765
BRAIN_DB_PATH~/.openclaw/workspace/brain.dbSQLite database path
CORS_ORIGINShttp://localhost:3000,http://localhost:8080CORS allowlist (scheme included)
AUTH_TOKEN / AUTH_TOKEN_FILE—Opaque bearer token(s); newline-separated = live rotation; off if unset
BRAIN_REQUIRE_AUTHunset1 = refuse to boot when no token resolves (fail-closed; without it a token-less boot carries a loud warn — the single-user-loopback posture it implies). Recommended on ANY deployment with a token file present
BRAIN_JWT_ISSUER—Enables JWT mode when set + keys loaded
INJECTION_POLICYquarantinequarantine | reject | allow
BRAIN_AUDIT_READ_EVENTSon (JWT) / off (loopback)Read-event audit
BRAIN_AUDIT_RETENTION_DAYSunset = foreverAudit retention window
BRAIN_WEBHOOK_TIMESTAMP_REQUIREDoff/unset1 = require the Standard Webhooks header set on /webhooks/* and verify v1, HMAC-SHA256 over {id}.{timestamp}.{body} (v1.20.4) — an opt-in hard replay window for first-party senders. GitHub sends no such timestamp; its replay protection is x-github-delivery idempotency, so leaving this unset keeps the legacy sha256= path unchanged

See Configuration and src/config.rs for the full list, including the JWT key directory, PRF tuning, suggest kill-switch, and DSAR webhook.

GDL provider profile

GDL uses a server-owned provider profile; do not put provider destination, model, or secret fields in a launch request. Set these together in the service environment:

BRAIN_GDL_PROVIDER_BASE_URL=https://provider.example/v1/stream
BRAIN_GDL_PROVIDER_MODEL=operator-selected-model
BRAIN_GDL_PROVIDER_SECRET_FILE=provider.key
BRAIN_GDL_PROVIDER_SECRET_ROOT=/absolute/operator-owned/secret-root

Create the root with operator-only directory permissions and the bearer file with mode 0600. The file is confined beneath the configured root; symlinks, outside-root paths, multiline/control content, and oversized values are refused. Keep the bearer out of command arguments and logs.

The four variables must be complete. All absent is an explicit disabled GDL provider; a partial or invalid profile refuses bootstrap. A complete profile must use a safe HTTPS endpoint. The existing address screen and DNS pinning run at launch, redirects are refused, and no provider client is kept in AppState. The readiness body reports only gdl_provider: disabled|configured|invalid; invalid is NOT_READY.

Grant the least-privilege workflow-operator role through the public role API to JWT operators that must launch GDL. Do not add workflow to the agent preset. A provider failure after admission is terminal and non-retryable: expect HTTP 503/gdl_provider_failed on the first launch and HTTP 409 with the same code on a later launch, with no provider replay. The provider request has a 25-second total body deadline; slow-drip responses cannot extend it, and receiver cancellation drops the in-flight HTTP future. Raw provider bodies, bearer values, secret paths, and secret-bearing URLs are not emitted.


Security posture in deployment

  • Loopback-safe by default — binding 0.0.0.0 without BIND_PUBLIC logs a loud warning (the opt-in is env presence); an unparseable host without BIND_PUBLIC refuses boot. In addition (v1.20.29) the server fails closed on startup: a non-loopback bind with no auth configured (no bearer token, no JWT keys) refuses to start, so an unauthenticated superuser API is never exposed off the loopback.
  • Two auth modes:
    • Opaque bearer (default): AUTH_TOKEN / AUTH_TOKEN_FILE, constant-time compare, multiple tokens for rotation.
    • JWT/JWS (opt-in): set BRAIN_JWT_ISSUER + generate keys with brain key generate. RS256/RS384/RS512/ES256/ES384/EdDSA only; revocation + refresh-chain reuse detection; per-route AuthZ.
  • Auth token file is 0600. The install script relocates any plaintext token out of the launchd plist into the secret file.

See Security for the full model.


Health & operations

brain doctor          # health + readiness
brain status          # counts, model, version
brain check-consistency   # duplicates, conflicts, stale sources

The audit log is read via the HTTP API (GET /audit) or the client console, not the brain CLI (the CLI has no audit subcommand).

/health reports liveness plus a capacity object (docs / DB size / RSS) and a hardening object (unsafe blocks, panics caught). Writes are guarded by a capacity envelope — reads are never blocked.


Security operations runbook (v1.20.5)

Loopback posture (the one-line checklist)

A token-bearing deployment should say so in the boot posture: set BRAIN_REQUIRE_AUTH=1 in the service environment (the plist) so a missing, deleted, or mis-resolved token file REFUSES boot instead of degrading to an unauthenticated single-user server (v1.28.80’s fail-closed admission). The /health/db authn.required echo names the live posture — false means you are relying on the loud-warn default. Verify after any install: curl -s localhost:8765/health/db -H "Authorization: Bearer $(head -1 ~/.config/brain-server/auth-token)" | jq .authn should read {"enabled":true,"required":true}.

Token rotation

The v1.20.2 machine-identity pattern: agents are not shared service accounts. Give each agent principal its own token and rotate on a cadence (≤90d recommended).

# opaque bearer: rotate atomically — fresh 0600 temp, fsync, rename (v1.27.12)
brain token rotate
# (or, manually: write a new token into the 0600 file; file-watch hot-reloads it)
umask 077 && head -c 32 /dev/urandom | base64 > ~/.config/brain-server/auth-token
# JWT mode: mint a fresh key, let the old one drain, then prune
brain key generate
# …wait ≥ max token lifetime (24h refresh)…
brain key prune
scripts/install-service.sh   # reload the key set

brain token rotate refuses to replace a group/world-readable token file and the server fails closed at startup on wide secret modes (token file, JWT keys, webhook signing secret, UMP signing keys — v1.27.12). Restart the server after rotating (scripts/install-service.sh) to load the new token.

Incident response — suspected memory poisoning

If a recall result, review item, or audit row looks planted:

  1. Review the blast radius — brain check-consistency (near-dups + contradictions) + GET /decayed to see what is currently decayed.
  2. Propose the cleanup — GET /consolidate/propose surfaces the duplicate / conflicting / stale-source candidates; approve the resolutions you trust.
  3. Purge the planted rows — POST /purge by id/owner (hard, audited, tombstoned) or POST /dsar {subject, action: purge} for a subject-scoped sweep. Every purge leaves a tombstone + audit row.
  4. Re-verify the chain — GET /audit/verify → {"ok": true}; the audit is tamper-evident, so the purge itself is provable.
  5. Rotate tokens — steps above, so the planted session (if any) dies with the old credential.

Classifier operations (v1.20.3, layer 2)

The optional ONNX classifier is auto-on (v1.28.71 “Pores”): unset or on loads it when the default artifact resolves at ~/.config/brain-server/models/injection-classifier/{model.onnx,tokenizer.json} (absent artifact → absent, not an error); only an explicit BRAIN_INJECTION_CLASSIFIER=off disables it. When loaded:

  • FPR calibration — watch the quarantine rate (/audit quarantined rows; the client Security panel surfaces the flag count). Tune BRAIN_INJECTION_THRESHOLD_HIGH/LOW — policy + thresholds read per call, so a flip takes effect without a restart (only the model load is cached).
  • Retrain trigger — re-run adaptive evals on a threat-model shift (new obfuscation technique or delivery vector observed); the blocklist + quarantine stay the always-on defense while a retrain is pending.
  • Model artifact hash-pin — pin the model file with sha256sum in the deployment config and verify on boot; the model file is itself a supply-chain artifact (LLM04/ASI04), so it is trusted like a dependency, not like a blob.
# pin the model artifact (the gate in the feature's docs)
sha256sum /path/to/model.onnx >> models.sha256

Backup & restore

brain backup <out-path>    # AES-256-GCM encrypted, checksummed, excludes secrets (DB from BRAIN_DB_PATH/default)
brain restore <in-path>

Warm standby shipper (v1.28.61)

A warm standby = encrypted base + shipped WAL chunks + a rehearsed promote. The shipper is an operator-run process, never a server thread (a shipper inside the server it protects is a correlated failure). Runbook: runbooks.md — promote procedure, ceilings, and the dated drill record.

macOS launchd (~/Library/LaunchAgents/com.brain.server.standby.plist):

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
  "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0"><dict>
  <key>Label</key><string>com.brain.server.standby</string>
  <key>ProgramArguments</key><array>
    <string>/usr/local/bin/brain</string>
    <string>standby</string>
    <string>start</string>
    <string>--to</string><string>/Volumes/standby/brain-follower</string>
    <string>--interval-secs</string><string>30</string>
    <string>--passphrase-file</string><string>/usr/local/etc/brain-server/backup.pass</string>
  </array>
  <key>EnvironmentVariables</key><dict>
    <key>BRAIN_DB_PATH</key><string>/Users/you/.openclaw/workspace/brain.db</string>
  </dict>
  <key>KeepAlive</key><true/>
  <key>RunAtLoad</key><true/>
  <key>StandardOutPath</key><string>/usr/local/var/log/brain-standby.log</string>
  <key>StandardErrorPath</key><string>/usr/local/var/log/brain-standby.log</string>
</dict></plist>

Linux systemd (/etc/systemd/system/brain-standby.service):

[Unit]
Description=brain-server warm standby shipper
After=brain-server.service

[Service]
ExecStart=/usr/local/bin/brain standby start --to /srv/standby/brain-follower \
  --interval-secs 30 --passphrase-file /etc/brain-server/backup.pass
Environment=BRAIN_DB_PATH=/var/lib/brain-server/brain.db
Restart=always
RestartSec=15

[Install]
WantedBy=multi-user.target

Monitor with cron: brain standby status --to <dir> and alarm when last cycle age exceeds 2 × interval — that is the shipper being dead. Rehearse the promote with brain standby promote-check --from <dir> and record the timings in the runbook.


The client GUI

Two GUIs ship over the same API:

  • Dioxus control surface (client/) — runs as a web app served by the server at /app, and as a desktop app. This is the default bundle (BRAIN_CLIENT_DIST → client/dist).
  • SvelteKit + Tauri shell (shell/) — the active successor: a typed-wire SvelteKit SPA with a Tauri desktop core, generated against openapi.yaml.

The Dioxus client/ removal is frozen until the shell’s parity gates pass.

# In the client/ directory — build the web bundle, then deploy it
./deploy-web.sh

To serve the shell at the same seat instead, point the dist at its build output (pnpm build → shell/build/):

BRAIN_CLIENT_DIST=shell/build ./target/release/brain-server

⚠️ Caveat: the shell’s current build is root-absolute (/_app/...), while the /app seat serves under a /app prefix, so its assets do not resolve from that mount today. Serving it at /app needs a base-path build first; serving it as its own origin works as-is.

See Client GUI.


Edge deployment (Jetson Nano / Raspberry Pi)

  • Set BRAIN_WORKER_THREADS=2 to trim RSS and context-switch overhead.
  • The release profile is speed-optimized (opt-level = 2) and the memory ceiling is bounded and configurable (CAPACITY_MAX_RSS_MIB defaults 512 on Jetson, 1024 on desktop targets; RSS is an advisory soft signal, not a hard kill).
  • No GPU, no embedding API, no Docker stack required.

CRM case intake (v1.28.22 “Bridges”)

brain-connector-crm (feature connector-crm) pulls support cases from Zendesk, Salesforce, or Genesys Cloud into the universal loop — operator- cranked via cron, one loop per invocation. Case bodies enter as proposals under BRAIN_WRITE_POSTURE=review; envelopes open governed runs and post crm/case/updated / crm/case/closed events. Config: 0600 JSON in ~/.config/brain-server/connectors/ (zendesk-*.json = {subdomain, email, api_token_file}; salesforce-*.json = {instance_url, client_id, client_secret_file, api_version?}; genesys-*.json = {region, client_id, client_secret_file, worktype?, org_id?}).

# Zendesk — every 5 minutes (respects the ~10 req/min incremental cap)
*/5 * * * * brain-connector-crm --source zendesk \
  --config ~/.config/brain-server/connectors/zendesk-acme.json \
  --checkpoint ~/.openclaw/workspace/brain.db >> ~/Library/Logs/brain-crm.log 2>&1

# Salesforce — incremental by SystemModstamp
*/5 * * * * brain-connector-crm --source salesforce \
  --config ~/.config/brain-server/connectors/salesforce-acme.json \
  --checkpoint ~/.openclaw/workspace/brain.db >> ~/Library/Logs/brain-crm.log 2>&1

# Genesys Cloud — workitems by worktype
*/10 * * * * brain-connector-crm --source genesys \
  --config ~/.config/brain-server/connectors/genesys-acme.json \
  --checkpoint ~/.openclaw/workspace/brain.db >> ~/Library/Logs/brain-crm.log 2>&1

Cursors persist in crm-state-{source}-{org}.json beside each config file. Custom CRMs: see connector-crm-custom.md.


The personal assistant crank (v1.28.42 “Valet”)

The trinity holds: cron or socket, never a daemon in the kernel. A reminder is just a governed run whose SLA envelope came due; brain valet due is a request-scoped, idempotent crank (outbox key valet-{run}-{due_at} — a double cron never double-fires). The Signal bridge is a separate zero-dependency edge process (tools/valet-relay/relay.js) holding ONLY its own 0600 config: it receives the server’s signed alert envelopes and forwards valet/due pings; your replies flow back through /webhooks/signal (HMAC-verified, replay-capped, injection-screened — every inbound byte is untrusted).

# The scheduler IS the cron recipe — every 15 minutes, weekdays.
*/15 * * * 1-5 brain valet due >> ~/Library/Logs/brain-valet.log 2>&1

# The morning brief, once a day at 07:30.
30 7 * * * brain valet brief >> ~/Library/Logs/brain-valet.log 2>&1

Setup: brain valet consent grant (the one-subject Outreach-lite registry — without it, envelopes fire locally but nothing is sent), then run the relay under launchd/KeepAlive with BRAIN_ALERT_WEBHOOK_URL pointing at its /alert listener and BRAIN_SIGNAL_WEBHOOK_SECRET_FILE mirroring the relay secret. Content-plan import: scripts/import-content-plan.ts plan.csv [--dry-run] creates one valet/reminder run per planned post.


The WhatsApp governed edge (v1.28.44 “Caravel”)

WhatsApp is governance MAPPING, not invention — Meta enforces the discipline; the adapter translates platform law onto kernel law. The edge is a separate Rust process (tools/channel-bridge, config-off by default: absent config = channel dark) that owns the PUBLIC webhook surface so brain-server never does:

  • Handshake + signature. Meta’s subscription GET (hub.challenge) is answered BY THE EDGE — the kernel never sees a challenge. Every POST is verified against X-Hub-Signature-256 (raw-body HMAC-SHA256 with the app secret, length-checked, constant-time) BEFORE any parse; only then are payloads projected into normalized envelopes, signed Standard-Webhooks style, and forwarded to POST /webhooks/channel/whatsapp. Verified bytes are the ONLY thing the kernel receives.
  • The 24-hour window rides the kernel gate exactly. Free-form replies inside 24h of the customer’s last inbound; outside it ONLY template messages — and a template send is a PROPOSAL (channel/template): double-approved by construction (Meta’s registry AND ours; ours carries the content digest). Business-initiated contact needs ALL THREE gates every time: template + standing consent in the shared registry + approved digest-bound proposal.
  • Statuses become lineage. sent/delivered/read/failed receipts land as case/channel_status outbox events on the thread’s case — hashes and refs on the audit chain, bodies never.
  • Quality tiers throttle deterministically. The tier state lives in a 0600 file under the state dir; a FRESH state is the MOST RESTRICTIVE tier until a status webhook upgrades it (fail-closed). Downgrades alert the operator via the bus metadata-only (number alias + old/new tiers).
  • Media digests-and-quarantine. Attachments downloaded by the edge are SHA-256’d; bytes sit in the retention dir named by digest, never auto- opened, never proxied through brain-server to a browser. Only the hash rides inbound (recorded verbatim ON the case note).

Config ($BRAIN_CONNECTOR_CONFIG_DIR/channel-whatsapp-{tenant}.json, 0600)

The SAME substrate file both sides read (domain + webhook_secret for the kernel seam; the WhatsApp keys for the edge):

{
  "domain": "acme",
  "webhook_secret": "whsec-…",
  "verify_token": "…",
  "phone_number_id": "1234567890",
  "app_secret_path": "app_secret.txt",
  "access_token_path": "access_token.txt"
}

Secret files are 0600, referenced by path (relative resolves beside the config); upward traversal refuses. Optional graph_api_version pins the Cloud API (default v21.0) — re-verify the account-quality webhook taxonomy against the pinned version at deploy.

Running

# Build the edge.
cargo build --release -p channel-bridge --manifest-path tools/channel-bridge/Cargo.toml

# Run (TLS terminates at YOUR reverse proxy in front of the loopback port).
tools/channel-bridge/target/release/channel-bridge \
  --config $BRAIN_CONNECTOR_CONFIG_DIR/channel-whatsapp-acme.json \
  --port 8791 --brain-url http://127.0.0.1:8765 \
  --retention-dir /var/lib/brain-server/channel-media \
  --state-dir /var/lib/brain-server/channel-bridge-state \
  --tick-secs 5

Run it under launchd/systemd KeepAlive like any governed edge. No extra cron: outbound drain is an internal tick loop paced by the tier table (throttled rows defer to later ticks). Registration evidence posts at boot over the same HMAC seam (channel:whatsapp mount, config-digest recomputed server- side). Template sends use parameterless templates (parameterized components are a documented ceiling).


The Slack and Teams operator annexes (v1.28.45 “Herald”)

The channels operators already live in become the console’s ANNEXES: case rooms, Relay handover pings, and digest-bound approvals where the people are. Both adapters are edge processes in the SAME tools/channel-bridge binary (config-off by default: absent config = channel dark), and the kernel-side pieces they ride are the SAME two HMAC seams as WhatsApp plus ONE new console seam:

Slack (Socket Mode)

  • No inbound listener exists by construction. The bridge DIALS Slack over the Socket-Mode WebSocket (apps.connections.open → wss, reconnect with capped exponential backoff + jitter). The slack kind binds NOTHING — pinned by socket_mode_never_opens_an_inbound_listener.
  • message events in the config’s mapped_channels become screened case notes through the ordinary inbound seam (thread map or [case N]); the sender’s OPAQUE user id rides as actor_ref (display names are never read).
  • Approve-by-button: pending renderable proposals render as Slack Blocks with the content preview AND the digest in the block; Approve/Reject buttons carry that digest in their value. A click whose digest is missing or mismatched is refused BRIDGE-SIDE (logged, never relayed) — and the kernel re-verifies it server-side. Two independent enforcement points.
  • Slash commands /brain due, /brain crank <run>, /brain approve <id>, /brain pending [limit] relay over the console seam; the kernel maps the clicking user through the user map and role-checks there.
  • User map: a Slack user is NOBODY until an approved channel/user_map proposal maps their opaque id to a principal with explicit roles. There is no auto-trust path.
  • Presence: mapped operator activity feeds the Crew roster as the closed activity kind channel — activity KINDS only, never content, and only while the domain’s Crew DPO switch is on.

Teams (Bot Framework + Adaptive Cards)

  • The supported Bot Framework route ONLY: the bridge registers an Azure bot, exposes POST /messaging behind the operator’s TLS proxy, verifies every activity’s Bot Framework JWT (JWKS, iss/aud pinned) BEFORE any parse, and answers with Adaptive Cards. The deprecated O365-connector path is deliberately NOT implemented.
  • Activities in mapped conversations become screened case notes (same threading law); proposal cards carry the digest field and Action.Submit returns it — the same digest binding as Slack buttons.
  • Room mapping: channel-bridge --config channel-teams-acme.json --list-channels enumerates the bot’s teams/channels via Graph (read-only, operator-run) so the operator can copy ids into mapped_channels.

Relay handover pings

When a handover OFFER is created, ONE channel/ping outbox row is enqueued with the I-PASS completeness state (refs only). The bridge drain resolves the receiving operator’s mapped platform refs + the case room and posts the ping in-channel (the case’s room; else the config’s handover_channel); an unmapped principal is audited loud and consumed — the drain never wedges. Accept/decline stays on the console (the ping coaches; the human decides there).

The user map (kernel side)

POST /workflow/channel/user-map FILES a channel/user_map proposal ({action: add|remove, channel, tenant, platform_user_id, principal, roles[]}); approval is the ONLY writer of the channel_user_map table (schema 1.28.45, additive). Roles resolve against the role store at file AND apply time. The console seam denies any actor that is unmapped, unroled, or lacking the action’s capability — 403, audited.

Config examples (0600, same substrate law as WhatsApp)

// channel-slack-acme.json
{
  "domain": "acme",
  "webhook_secret": "whsec-…",
  "mapped_channels": ["C0123ABCD"],
  "handover_channel": "C09HANDOVER",
  "app_token_path": "slack_app_token.txt",
  "bot_token_path": "slack_bot_token.txt"
}

// channel-teams-acme.json
{
  "domain": "acme",
  "webhook_secret": "whsec-…",
  "mapped_channels": ["19:…@thread.tacv2"],
  "bot_app_id": "00000000-0000-0000-0000-000000000000",
  "bot_tenant_id": "00000000-0000-0000-0000-000000000000",
  "bot_password_path": "teams_bot_password.txt"
}

Least privilege at the workspace-app level: install the Slack app with access scoped to the mapped channels only, and the Teams bot to its team only; channel tokens grant nothing beyond their mapped channels. Tokens live in 0600 files referenced by path — the bridge holds NO brain token, ever (pinned house-wide by self-grep).

Running

cargo build --release -p channel-bridge --manifest-path tools/channel-bridge/Cargo.toml

# Slack: dials OUT; binds nothing.
tools/channel-bridge/target/release/channel-bridge \
  --config $BRAIN_CONNECTOR_CONFIG_DIR/channel-slack-acme.json \
  --brain-url http://127.0.0.1:8765 --tick-secs 5

# Teams: one loopback listener behind YOUR TLS proxy.
tools/channel-bridge/target/release/channel-bridge \
  --config $BRAIN_CONNECTOR_CONFIG_DIR/channel-teams-acme.json \
  --port 8792 --brain-url http://127.0.0.1:8765 --tick-secs 5

# Teams room mapping (operator-run, read-only):
tools/channel-bridge/target/release/channel-bridge \
  --config $BRAIN_CONNECTOR_CONFIG_DIR/channel-teams-acme.json --list-channels

Deployment tiers (ISO 18295-1 applicability: any size)

The standard applies to a centre of any size; so does this server. The same binary scales from one operator to a global BPO by configuration, not by forks. Pick the tier that matches the operation — every tier ships the full audit chain and fail-closed gates.

Tiers are config, not forks: each tier is a checked-in env profile — deploy/tiers/t1.env, deploy/tiers/t2.env, deploy/tiers/t3.env, deploy/tiers/t4.env — that CI boots as part of the tier-smoke matrix, and a meta-test (guide_and_profiles_never_drift) fails if a profile sets a key this guide does not document. (The reverse — a key this guide documents that no profile sets — is not covered by the meta-test; the matrix below marks such keys “unset”.) Copy the profile into your service environment and add only site-specific values (BRAIN_DB_PATH, BIND_PORT, auth material).

TierWhoShapeProfile
T1 soloOne operator / micro-centreloopback bind, single domain, single DB, no rolesdeploy/tiers/t1.env
T2 teamA small team (≤ ~25 agents)roles enabled, HITL proposal review queue on, crew presence visibledeploy/tiers/t2.env
T3 siteA site or BPO campaignmulti-domain/multi-DB, calibration + public KB feedback live, WFM feeds feeding the centre’s tooldeploy/tiers/t3.env
T4 globalMulti-site / multi-regionT3 plus knowledge parcels, residency stamps, follow-the-sun handover via the shift ringdeploy/tiers/t4.env

Per-tier config matrix

VariableT1 soloT2 teamT3 siteT4 globalWhy
BRAIN_WRITE_POSTUREopen (the operator IS the reviewer; proposals still audited)reviewreviewreviewagent writes become HITL proposals from T2 up
BIND_PUBLIC0000never expose without auth; the server refuses any non-loopback bind with none regardless of tier
BRAIN_AUDIT_READ_EVENTSoff (loopback default)onononshared surfaces get read-audited once more than one person uses them
BRAIN_MULTI_DBunsetunset11domain-per-campaign databases at site scale
BRAIN_MAX_DOMAIN_DBSunsetunset1664explicit cap under the bounds law; size to your domain count
BRAIN_WEBHOOK_TIMESTAMP_REQUIREDunset (off)unset (off)11replay-hard webhook intake for first-party senders at site scale (T1/T2 profiles leave the key unset rather than set it to 0 — same off behavior)
BRAIN_OTEL_ENABLEDunsetunsetoptional1instrumented decision cores for multi-region ops visibility
BRAIN_TRUST_PROXYunsetunsetoptional1set only when TLS terminates on a trusted proxy chain

Sizing guidance

SQLite WAL headroom is the sizing lever, not heroics: keep the WAL under a few hundred MB by running brain backup (which checkpoints) on the cadence below, and promote to BRAIN_MULTI_DB when a single DB’s write contention or backup window stops fitting the maintenance slot. On edge ARM hardware set BRAIN_WORKER_THREADS=2 and keep CAPACITY_MAX_RSS_MIB at its 512 default. No sizing promise beyond what you measure — bench against YOUR corpus before promoting a tier.

Cadences (cron recipes)

CadenceT1T2T3T4
CRM connector sync—dailyevery 5–10 min (see CRM case intake)every 5–10 min per site
brain backupweeklynightlynightly + pre-calibrationnightly per region
KB build / publish (kb build)ad hocweeklydaily + feedback-loop drivendaily per locale set
Human-signed calibration—quarterlymonthly (the signed register extract rides it, v1.28.37)monthly per site
Valet crank (brain valet due)——weekdays every 15 min (the personal-assistant heartbeat, v1.28.42)same, per operator
Token rotation≤90d≤90d≤90d≤90d (staggered per principal)

Upgrade path

Tier promotion is additive: nothing configured at T1 blocks T4 features later. Move up by merging the next profile’s keys into your environment, restarting, and re-running the smoke suite (brain doctor, brain check-consistency, GET /audit/verify). There is no downgrade migration either — drop back by removing keys, never by editing data. The deliberate ceilings (workload visibility is measured, never enforced; no forecasting/scheduling engines — WFM alignment is interop) hold at every tier.

Next steps

Enterprise pilot profile — openclaw.json hardening (2026-09-09)

The personal-use defaults in ~/.openclaw/openclaw.json are correct for a trusted loopback host but too permissive for a multi-tenant pilot. Apply this profile for any internet-facing or group pilot:

{
  "plugins": {
    "brain-server": {
      "agents": ["*"],
      "autoCapture": false,
      "autoRecall": true, "autoRecallTopK": 5, "recallMaxChars": 2500,
      "allowedChatTypes": ["direct", "explicit"],
      "baseUrl": "http://127.0.0.1:8765",
      "defaultDomain": "global",
      "minQueryLength": 5,
      "requestTimeoutMs": 8000,
      "strictDomain": true,
      "authToken": "${BRAIN_SERVER_AUTH_TOKEN}"
    }
  },
  "tools": {
    "fs": { "workspaceOnly": true }
  }
}

${BRAIN_SERVER_AUTH_TOKEN} is openclaw-host environment substitution for the plugin’s authToken field. Prefer the plugin’s own ladder where possible: BRAIN_TOKEN_FILE (0600 secret file) → BRAIN_TOKEN (env) → authToken — see the authToken row in the integration reference.

Server side for the same pilot:

BRAIN_WRITE_POSTURE=review
BRAIN_AUDIT_READ_EVENTS=on
BRAIN_AUDIT_RETENTION_DAYS=180

Tuned enterprise retrieval (neural profiles + tuned classifier). Requires a neural-embed + rerank-tier build; switching embedding dimensions on an existing DB fails closed with the --re-embed instruction:

MODEL_PROFILE=enterprise
BRAIN_RERANK_MODEL_DIR=models/mxbai-rerank-large-v1/
BRAIN_INJECTION_CLASSIFIER=/path/to/tuned-model.onnx
BRAIN_INJECTION_TOKENIZER=/path/to/tuned-tokenizer.json
BRAIN_INJECTION_THRESHOLD_HIGH=0.9
BRAIN_INJECTION_THRESHOLD_LOW=0.7

A nonexistent classifier path refuses boot; thresholds band reject vs quarantine without restart. /health/db echoes the resolved profile, rerank arming, and classifier state — verify there before piloting.

Why each change (see MEMORY_STACK_REPORT_2026-09-09.md §1):

SettingPersonal defaultPilot valueWhy
autoCapturefalsefalseKeep it off: every group/channel message would otherwise auto-queue as a proposal
allowedChatTypes["direct","explicit"] (group/channel excluded)["direct","explicit"]Plugin recall in groups = cross-tenant prompt-injection via query; keep the default exclusion
strictDomainfalsetrueFail-closed on unknown domain instead of global sink
autoRecallTopK / recallMaxChars3 / 10005 / 2500Raise the ceiling slightly for pilot context depth; still bounded per turn
tools.fs.workspaceOnlyfalse + alsoAllow ["*"]true + explicit allowlistMaximally permissive is personal-only

Linux appliance install (the clean cycle)

For a single-host deployment — a government office, a back office — that is turned off at the end of the day and back on in the morning:

sudo ./deploy/install.sh                       # binaries, unit, service user
sudo systemctl start brain-server
/usr/local/bin/brain-clean-cycle-check         # the morning check

install.sh refuses to overwrite an existing store and prints the upgrade sequence instead. uninstall.sh removes the service and never the data; if you intend to remove the data, it tells you to run brain shred first and then delete it by hand.

The full runbook — the evening stop, the morning check, the storage rules, the backup rules and the off-site approval — is clean-cycle.md. Read it before the first production copy.

One-shot backup shipping

brain standby start is an infinite loop and cannot be run by a scheduler. For a timer, a CronJob, or a monthly ritual:

brain standby ship --to /path/to/follower [--passphrase-file PATH] [--db PATH]

It runs exactly one ship cycle and exits with its status, producing the same encrypted base, WAL chunk and signed manifest the shipper produces, which brain standby promote-check then verifies.

Deployment reference architecture

The shape a larger deployment takes — two hosts on separate circuits, a per-node battery, a cold standby, and an off-site vault — is recorded, with its unmeasured parts labelled as such, in deployment-reference-architecture.md.

Pilot caveats (honest ceilings)

  • Residency panel today shows DB file + BRAIN_REGION stamp, not per-tenant key isolation — per-tenant keys (SQLCipher + KMS, BRAIN_TENANT_KEY_FILE per tenant) ship in v3.7 (Q1 2027). See COMPLIANCE.md §10.3 / THREAT_MODEL.md.
  • Rate limiting is loopback-scoped until v2.1. The shared loopback bucket (X-A10 / S2-40) is correct for loopback. Any internet-facing pilot requires per-principal rate buckets (carry until v2.1) — put the reverse proxy’s per-IP limit in front and note the gate in the deployment runbook.
  • “Enterprise-pilot-ready” is subject to operator attestation. ISO 42001 / SOC 2 Type II attestation remains an external operator audit; the repo provides the posture, not the certificate. See COMPLIANCE.md header + §6.1.