Connectors — supervised external backfill
Connectors let Brain Server backfill external sources into the existing source/revision pipeline, supervised by an operator — the same way you ingest markdown or memories, but from a live external system (today: GitHub).
This page is verified against src/connector/, src/bin/brain-connector-gh.rs,
and the connect/sync/connector-status commands in src/bin/brain.rs.
What a connector is
A connector is a supervised ingester. It fetches items from an external system
and feeds them through the same source + immutable-revision pipeline the
manual ingest paths use — so connector-loaded content carries full provenance,
participates in the knowledge graph and hybrid recall, and is reconciled like any
other source. The connector’s source_path (github://…) keys the source row,
and reconciliation sweeps it under kind github.
A supervisor process owns the lifecycle: register → authenticate → sync → reconcile → report. The operator sees and controls it; nothing runs autonomously.
Today’s connector: GitHub issues (App auth)
The shipped connector pulls GitHub issues for configured repositories, authenticating as a GitHub App (installation access token), not a personal token.
Prerequisites
- A GitHub App with an installation on the target org/repos.
- The App’s App ID and Installation ID.
- The App’s private key file (PEM) — used to mint the short-lived installation token.
- (Optional) a webhook secret file for the issue webhook path.
Register (authenticate)
brain connect github \
--app-id 123456 \
--install-id 9876543 \
--key-file ./github-app.pem \
--repo acme/widgets --repo acme/docs
The GitHub App flow is implemented in src/connector/auth/github_app.rs
(GitHubAppConfig / GitHubAppProvider) and the HTTP client in
src/connector/github/client.rs — an installation token is minted from the App
key and used for the fetch.
Sync (backfill)
# backfill the registered instance(s)
brain sync github --config PATH # explicit config file
brain sync github --instance NAME # a named registered instance
brain syncis backed bybrain-connector-gh, a separate feature-gated binary (--features connector-github) because it pulls in the GitHub HTTP client.- Backfill functions:
backfill_issues_for_repoandreconcile_github_sources(src/connector/github/).
Inspect
brain connector-status # id, kind, instance, state, last_sync_at
connector-status reads GET /connectors and prints the registered
connectors; if none are registered it prints the brain connect usage line.
The kind column currently shows github.
Feature gate
The brain-connector-gh binary is feature-gated:
cargo build --release --features connector-github --bin brain-connector-gh
The brain connect/sync/connector-status commands in the main brain
binary are always compiled (they delegate to the server / connector binary as
appropriate); only the standalone connector binary needs the feature.
How connector chunks are stamped (the honest version)
The GitHub connector posts to POST /ingest/markdown, and that handler stamps
the chunk’s source column as markdown — the chunk-level ingest-kind
vocabulary is memory | markdown | structured | manual | vault, and there is
no connector chunk kind. Connector provenance lives one level up, in the
sources row (kind github, keyed by the github://… source_path) and
the revision lineage. Chunks are stamped imported origin (per
gate::origin_for_source — anything that is not manual/memory is
imported). The confidence ×0.9 “unverified external source” discount
(gate::confidence) keys on the SOURCE STRING containing
connector/github/web — which a markdown-stamped chunk does not carry,
so connector chunks do not receive the ×0.9 factor under the current
wiring; the imported-origin label is what carries the trust signal today.
See Memory lifecycle for the origin mapping.
Reconciliation
Like file/markdown sources, connector sources can be reconciled — orphans from sources that were deleted are swept so the shared store doesn’t answer from dead material:
brain reconcile <path> [--kind vault]
# or over HTTP:
POST /sources/reconcile
Security model
- Auth is App-scoped, never a personal token — least privilege, revocable, short-lived installation tokens minted per sync.
- Connector config lives under
~/.config/brain-server/connectors/github-{instance}.json(mode-checked like other secrets; the server’s fail-closed secret-permission check applies to the configured key/secret files). - Sync is operator-initiated; there is no autonomous background fetch. The
connector surfaces its state (
state,last_sync_at) for operator review.
CRM case connectors (v1.28.22 “Bridges”)
brain-connector-crm (feature connector-crm) is one binary, three sources
(--source zendesk|salesforce|genesys), operator-cranked via cron — the same
discipline as GitHub: config-derived hosts only, redirects refused, bounded
timeouts, secrets in 0600 files, cursors in a connector-owned state file.
Case bodies enter through the UMP ingest path (proposals under
BRAIN_WRITE_POSTURE=review); case envelopes open governed runs and post
crm/case/updated / crm/case/closed outbox events; the crm_cases
table binds each stable case_ref to its run. Customer identity is stored
only as a salted SHA-256 subject ref. Cron recipes: deployment.
Custom CRMs (Freshdesk, ServiceNow, JSM): connector-crm-custom.md.
Honest ceiling
- Registry vs runnable binary. The connector-kind registry
(
CONNECTOR_KINDS:github,crm-salesforce,crm-hubspot,slack,email-imap,jira,linear,notion,hris-readonly,ehr-readonly) is open for registration (POST /connectors/register, profile-gated by family), but onlykind=githubhas a runnable binary — the CLI names the shipped kinds and points at the GitHub connector as the working backfill template rather than quoting a version. GitHub issues are the concrete backfill; the connector contract (src/connector/mod.rs+src/connector/supervisor.rs) is designed to be extensible to other kinds. The inbound channel-bridge sibling (tools/channel-bridge,/webhooks/channel/{kind}) is documented in deployment. - It pulls issues via App auth over the GitHub REST API; it does not sync arbitrary repository content, PRs, or code.
- The CRM connectors are pull-only intake — there is no CRM writeback (posting resolutions back to the vendor is a later, separately-gated release), no background supervisor sync (cron only), and the custom-CRM path is docs + pure mappers, deliberately no generic JSONPath runtime.
Next steps
- Source lifecycle — provenance (
source+ immutablerevision). - Memory lifecycle — origin tiers and the
connectorkind. - API reference —
GET /connectors,POST /sources/reconcile.