Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

The Cited Work: every source behind these mechanisms, with links

Scope: every external source cited across docs/research/ and docs/blog/, gathered into one place with a short summary and a link that resolves. The mechanism notes keep their own inline citations; this is the index into them.

Why this note exists. The mechanism notes cite accurately but sparsely: an arXiv ID in parentheses, an author and year in prose, occasionally a bare journal name. That is the right density for a note whose subject is the implementation, and the wrong density for a reader who wants to go read the paper. Fourteen arXiv identifiers were cited in this directory and none of them carried a resolvable link. This note is the fix.

Verification rule applied here. Every entry below was checked against the published record during authoring, not recalled. Where a source is a preprint, a standard, or a guideline rather than a peer-reviewed paper, it says so. Two discrepancies surfaced during that check and are corrected in place; both are noted below rather than quietly amended.

Retrieval and fusion

Reciprocal Rank Fusion

Cormack, Clarke & Büttcher (2009), SIGIR. Scores each document 1/(k + rank) and sums across result lists.

The problem it solves is the one that makes naive hybrid retrieval awkward: two retrievers return scores on incomparable scales. A cosine distance and a BM25 score cannot be added without normalizing them, and any normalization you pick is a tunable parameter you now own. RRF sidesteps this by ignoring scores entirely and using only ranks, which is why it is parameter-light and hard to get wrong.

The paper reports RRF almost invariably beating the best individual system, and beating Condorcet Fuse and CombMNZ, across TREC and LETOR. Brain Server uses RRF_K = 60 (src/search/mod.rs:31), the standard value from the paper. Used in 08-hybrid-fusion.

The Probabilistic Relevance Framework: BM25 and Beyond

Robertson & Zaragoza (2009), Foundations and Trends in Information Retrieval 3(4).

The reference treatment of BM25, deriving it from a probabilistic model rather than presenting it as a heuristic, and explaining why the saturating term exists: repeated terms should stop helping, because a document that says “audit” thirty times is not thirty times more relevant. Brain Server’s lexical leg is SQLite FTS5 with BM25 ranking, so this is the leg’s theoretical basis. Used in 08-hybrid-fusion.

  • Robertson, S. E., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 3(4). https://doi.org/10.1561/1500000019

Jégou, Douze & Schmid (2011), IEEE TPAMI 33(1).

Vector search has a space problem: a float32 embedding is large, and scanning millions of them is slow. Product quantization decomposes a vector into subvectors, quantizes each against a learned codebook, and represents the whole vector as a short code. Distances are then approximated from the codes.

Brain Server does not implement PQ. It stores vectors as int8 and binary quantized in a vec0 table (vec_quantize_int8(…, 'unit') plus vec_quantize_binary(…)), which is simpler scalar quantization in the same family. The claim in the docs is a storage and speed trade of 4× to 32×, against some recall at the margins. Citing PQ is citing the family, not claiming the same compression ratio. Used in 08-hybrid-fusion.

Pseudo-relevance feedback, the classic result

PRF takes the top-k results of a first pass, assumes they are relevant, and uses their terms to expand the query. The standard formulation is Lavrenko & Croft (2001), SIGIR, whose relevance-based language models give the RM1, RM2, and RM3 variants, with RM3 the one usually meant by “classic PRF”. The earlier lineage is Ponte & Croft (1998), which introduced the language-modeling approach to retrieval that PRF builds on.

This codebase uses neither formula directly. Its PRF is a gate: expansion fires only when the cross-retriever evidence agrees, so a confident single retriever cannot rewrite the query on its own. The citation is for the technique being gated, not for the gate. Used in 07-prf-evidence.

Correction worth recording: this entry previously attributed PRF to Cormack et al. 2008, which is wrong. Cormack is the RRF author; the PRF line is Lavrenko & Croft, with Ponte & Croft as its predecessor. Corrected here rather than quietly amended.

Chunking and RAG lineage

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Lewis, Perez, Piktus et al. (2020), NeurIPS.

The paper that made chunk, then retrieve, then generate the default shape for knowledge-intensive NLP. It is cited here for framing only: the chunk-then- retrieve unit it established is what a memory store is organized around. This server deliberately does the retrieval half deterministically and hands the result to a model rather than training an end-to-end retriever-generator, so the paper is lineage, not method.

Retrieval-Augmented Generation for Large Language Models: A Survey

Gao et al. (2023), arXiv:2312.10997.

A survey of the chunking strategies that grew out of RAG, including the fixed-size versus structure-aware trade-off. The mechanism note is honest that structure-aware chunking is an engineering practice rather than a single citable algorithm: the heading-aware CommonMark splitter in src/chunker.rs is this project’s own choice, benchmarked against fixed-size in that module’s tests. This survey is the closest thing to a citable survey of the trade-off. Used in 10-chunking.

GraphRAG: From Local to Global, a Graph RAG Approach

Edge et al. (2024), arXiv:2404.16130.

Microsoft’s approach to the question that plain vector retrieval answers badly: queries about a whole corpus rather than a document (“what themes recur here?”) need a summary of structure, not top-k nearest neighbours. GraphRAG builds an entity graph and community summaries so global questions have something to retrieve.

Related lineage, explicitly not the same thing: it summarizes and embeds clustered text rather than splitting markdown, which is why the mechanism note lists it as adjacent rather than as a source. Brain Server’s graph leg is Personalized PageRank, closer to 04-ppr-graph. Used in 10-chunking.

Agent memory and anticipation

Generative Agents: Interactive Simulacra of Human Behavior

Park et al. (2023), UIST.

The canonical “memory as a first-class agent component” architecture: a memory stream scored by recency, importance, and relevance, plus reflective memory that synthesizes higher-order abstractions. This is the ancestor of every agent-memory product, and it is cited for the scoring shape rather than for any claim of similarity. Used in 09-anticipation.

MemGPT: Towards LLMs as Operating Systems

Packer, Wooders, Lin et al. (2023), arXiv:2310.08560. Preprint.

The OS analogy: treat context as a virtual address space and page between a small main context and larger external memory, with the model deciding what to page. The relevant lesson for this codebase is narrow and stated as such in the note: anticipatory memory must be reviewable, nothing is silently injected.

Correction worth recording: this identifier is MemGPT, and earlier in this project’s notes it was associated with Generative Agents, which is arXiv:2304.03442. The two are different papers. The citation is now correct. Used in 09-anticipation.

Mem0

The feedback API shape (memory_id, feedback) is the interoperability surface cited in the anticipation note. Mem0 is a product, not a paper, so it is listed here without a canonical citation; its own documentation is the reference. This matters for a related reason: the repo’s own lock-in post argues from vendor documentation rather than marketing, so the same standard applies.

Calibration

On Calibration of Modern Neural Networks

Guo, Pleiss, Sun & Weinberger (2017), ICML.

The paper behind expected calibration error. Modern networks are overconfident: a 0.9 prediction is right about 72% of the time. The paper introduces temperature fitting as the fix, applied against held-out data, and frames it as a property you must measure rather than assume.

Directly load-bearing for the System-1 port. The implementation has a hand-computable ECE with a NaN-means-no-measure law, and the rollout consequence is conservative 0.85 thresholds with escalate-heavy behavior until a temperature fit exists. Auto-action stays behind a fine-tuned checkpoint with a pinned SHA plus ECE evidence. Used in 06-abstention-verify and 14-governed-diagnostic-loop.

Graph retrieval, 2026 wave

These four identifiers are cited in the mechanism notes for the 2026 graph-memory direction. They are preprints, and the notes’ own ceiling language is the right frame: the graph-memory design space is active and unsettled, and a citation is a pointer to a position, not an endorsement of a result.

Memory benchmarks

LoCoMo

Maharana et al. (2024), ACL, arXiv:2402.17753. Very long multi-session conversations evaluated with QA plus event summarization.

This is the reference benchmark shape for the field, and the reason the memory-benchmark note exists. Its own headline is contested: the note opens on two vendors publishing different scores for the same benchmark, one of them in a vendor blog, and the honest conclusion is that self-reported numbers on LoCoMo are not comparable without a stated protocol. Used in 13-benchmark-landscape-2026.

LongMemEval and BEAM are cited in the same note as the 2026 standard alongside LoCoMo. Both are listed there by name without a canonical citation; the note’s own planned deliverable is a public harness over them, which is the right place for their identifiers to land when that harness exists.

Clinical process shape

These are the sources for the governed diagnostic loop’s process layer, and the note is careful that they are cited as process shape, not as diagnostic instruments. The loop enforces that a closure happens with evidence. It does not practice medicine.

Improving Diagnosis in Health Care

National Academies of Sciences, Engineering, and Medicine (2015).

Diagnosis as a multi-step process with named failure points. Step 6 carries the closure discipline the loop gates as A8 and A9: no resolution without a law-clean closure artifact, and reflexive closure refused. Used in 14-governed-diagnostic-loop.

  • National Academies of Sciences, Engineering, and Medicine (2015). Improving Diagnosis in Health Care. National Academies Press. https://doi.org/10.17226/21894

Changes in Medical Errors after Implementation of a Handoff Program

Starmer et al. (2014), NEJM. The I-PASS study.

Sender-owned sections (illness severity, patient summary, action list, situation awareness, synthesis) assembled by the sender and never synthesized by the receiver. The loop’s ipass_facts renders sender-owned sections only, and an escalation lands exactly one pre-filled offer draft behind a human gate. Used in 14-governed-diagnostic-loop.

Emergency Severity Index, v4 (AHRQ) and Emergency Triage (Manchester)

Triage acuity bands with wait windows. The loop ports the shape (MTS-style bands plus ESI 1–5, at least one required at triage exit, closed sets) while keeping acuity a monitor beside the authoritative P-class SLA, which takes the tighter of the two and never the looser. Acuity is advisory by construction and never binds resourcing. Used in 14-governed-diagnostic-loop.

Software engineering research

These come from the docs-truth blog post, which argued that a repo can encode its own rules and have a machine check them. They are grouped here because they are the empirical backing for gates rather than for memory mechanics.

What this note is not

It is not a claim that these papers validate this system. A citation means the mechanism note drew a shape from the work. It does not mean the paper benchmarked our implementation, or that our numbers match, or that we reproduced the result. Where a figure is quoted, the note that quotes it carries its own ceiling, and this note does not upgrade it by restating it.

It is not complete. It covers the sources cited from docs/research/ and docs/blog/. Compliance and threat-model documents cite standards and regulations (OWASP, NIST AI RMF, ISO 42001, SOC 2, GDPR, CRA) that belong in a standards register rather than a papers bibliography; those are gathered, verified, and linked in The Standards Register.

Two entries remain deliberately unlinked. LongMemEval and BEAM are named in the benchmark note without canonical identifiers. Rather than guess, they are flagged here as owed, and the planned public harness is where their identifiers should land.

Identifiers drift. Two were corrected during this pass (MemGPT’s, and a local-calibration figure whose provenance turned out to be a video claim rather than a measurement). A bibliography is a claim about sources, so it is worth the same treatment as any other: verify before citing, and record the correction when one is found.