Overview
Brain Server is a deterministic decision and memory substrate for teams and their AI agents.
It is one Rust binary that stores what a team knows — past resolutions, runbooks, KB articles, decisions, customer context — and both recalls it and supports the structured decisions that depend on it the same way every time, on the operator’s own hardware: private, offline-capable, no per-query cost on the hot path, and a human gate on everything an agent writes into permanent state (direct operator/API writes are screened and audited, and can be gated too — see BRAIN_WRITE_POSTURE).
The core idea is simple: recall that never has to think, and decisions that leave a trace. Instead of asking a language model whether to recall, and instead of paying an embedding API on every read and write, Brain Server uses a static, local embedding model and a deterministic retrieval pipeline. Permanent writes and configuration changes remain under explicit human control, and every significant action lands on a tamper-evident audit chain.
This is not a toy or a “local RAG.” It is the compliance-grade substrate that enterprises deploy when both memory and the decisions that depend on it must be private, explainable, and under human control.
Why it exists
Cloud memory services (Zep, Mem0, Letta Cloud) are powerful but carry three structural costs that don’t fit every use case:
- Per-query cost — an LLM or embedding API is charged on every read and write.
- Data egress — the agent’s memory lives in someone else’s datacenter.
- Network latency — recall waits on a round-trip to the cloud.
Brain Server inverts all three: zero per-query cost, zero data egress on the
retrieval path, zero
network latency on recall. (Operator-configured egress exists and is pinned
at the boundary: webhook/DSAR sinks, OIDC/JWKS fetch, the loop engine’s
provider calls — all behind the SSRF-hardened egress policy; see
docs/architecture.md.) It is designed to run on a 4 GB ARM device (Jetson
Nano, Raspberry Pi 5, a small mini PC).
Who it is for
- Support, helpdesk & contact-center teams whose agents must give customers the same grounded answer every time — past resolutions and KB articles recalled deterministically, with the review queue turning every solved case into reviewed knowledge (the KCS loop, as data).
- Teams that share one brain — domains, roles, procedures, case rooms, and handovers, so knowledge lives in one governed place instead of ten inboxes; agents join the same store under the same rules.
- Edge / privacy-first agent builders — people who can’t or won’t use an embedding API, and want the memory to live on the device.
- Knowledge-workers who think in domains — health, business, code, and more as separate brains that cross-reference on a miss.
The full audience map — including BPOs, in-house contact & support centers, regulated enterprises (finance, healthcare, legal, government), edge/field deployments, and delivery partners — is in Who it’s for — target audiences, with every segment marked shipped vs. planned (multi-client tenancy is the v2.0 “Cortex” milestone).
The six differentiators
① Zero-token, deterministic recall — no LLM in the loop
Every turn, the agent calls one /recall and gets the evidence to inject. No LLM
decides whether to recall, and no LLM extracts memories on write. Token accounting:
0 decision tokens, 0 embedding tokens. Only the capped returned snippets cost
context.
② Local embeddings — offline, private, ~free on CPU
The default profile uses potion-retrieval-32M via model2vec — a static
model, no transformer forward pass, just token lookup — running in-process with
no GPU and no network. There is no embedding API dependency in any profile:
embeddings are always a local library call. Opt-in MODEL_PROFILE=enterprise
(BAAI/bge-m3, 1024-d) or MODEL_PROFILE=desktop (gte-base-en-v1.5, 768-d)
swap in larger local transformer embeddings — still zero-API, zero-egress — and
an optional cross-encoder rerank tier (mixedbread-ai/mxbai-rerank-large-v1,
fallback bge-reranker-v2-m3) fine-tunes the fused order on the profiles that
arm it.
③ Per-domain knowledge graphs with automatic routing
Memories live in scoped domains (health, business, code, …), each with its own entity/relationship graph. Routing between domains is automatic via per-domain centroids — no manual tagging on ingest or query — with cross-domain fallback on a miss.
④ Edge-first, memory-bounded, single binary
A single Rust binary with embedded SQLite + sqlite-vec. int8/binary vector
quantization (4–32× smaller), bounded connection pools, a configurable memory
ceiling (512 MiB on a 4 GB ARM device — the default jetson target; 1 024 MiB
on the desktop target; CAPACITY_MAX_RSS_MIB tightens either). No separate
vector-DB process, no Python runtime, no Docker stack.
⑤ Native OpenClaw memory plugin
Ships as a kind: "memory" plugin occupying the memory slot, with per-agent opt-in
and group/channel exclusions for data-leakage prevention.
⑥ Human-gated write-back — meaningful control, not a rubber stamp
Agent-captured fragments never become permanent memory unreviewed. A captured
fragment is scored, not
stored (POST /ingest/proposal), and enters the store only after a human approves it —
optionally superseding the chunk it contradicts. (Honest scope: the compiled
default of BRAIN_WRITE_POSTURE is open for direct operator/API writes —
screened and audited but not proposal-gated; review is the installer’s
new-install default and gates all six agent-facing write surfaces.) The control room (Review panel, Memory
Operations panel with live SLA clocks + gate health, Agent Memory Register) is built to
make the operator a critical evaluator: raw evidence, sourcing prompt, and screen
verdict on every card, with every decision written to a tamper-evident audit chain. See
Human in the loop.
One-line positioning
Brain Server is a deterministic, self-hosted knowledge server for teams and their AI agents — one binary that stores what your team knows, recalls it the same way every time, and never lets a write bypass a human.
What’s inside
- Hybrid retrieval — vector KNN + lexical FTS5 fused via Reciprocal Rank Fusion, with deterministic PRF expansion and full provenance.
- Temporal evidence — every ingest stamps
observed_at/valid_from/valid_to; point-in-time recall returns the revision active at a timestamp. - Knowledge graph — entities and relationships extracted from markdown, traversable and queryable, with faithful multi-hop explanations.
- Governance — append-only audit log, prompt-injection quarantine, write-back gating with human approval, GDPR export/purge/DSAR, and calibrated abstention.
See The memory lifecycle for the full end-to-end path a fact takes from capture to storage, retention, recall, and erasure — and Human in the loop for the review gate + erasure procedure.
One-line positioning: Brain Server is a local-first, governed decision and memory substrate for reproducible agent systems — deterministic retrieval, human-gated permanent state, and tamper-evident provenance.
Continue to the Quickstart to get running.