✅ fresh
last synced 2026-09-29T21:12:19.681221+00:00 · coverage 100% (
reindex)Validation by Beadloom
doc_sync— same source assync-check.
Reindex
Full and incremental reindex pipeline for rebuilding the architecture graph database.
Source: src/beadloom/application/reindex/ (package; decomposed by cohesion in BDL-059 S4 into models, rules_loader, indexing, enrichment, sync_state, change_detection, full, incremental; BDL-074 C1 added test_index; with the package __init__ re-exporting the stable public + back-compat surface)
Specification
Purpose
The reindex module orchestrates the complete data pipeline that transforms YAML graph definitions, Markdown documentation, and source code into a queryable SQLite database. It provides two modes: a full reindex that drops all tables and rebuilds from scratch, and an incremental reindex that processes only changed files. The incremental path uses SHA-256 file hashes stored in a file_index table to detect changes, and falls back to full reindex when graph YAML files change or no prior file index exists.
Data Structures
ReindexResult (dataclass)
| Field | Type | Default | Description |
|---|---|---|---|
nodes_loaded | int | 0 | Number of graph nodes loaded from YAML (full reindex) or live-DB total (incremental) |
edges_loaded | int | 0 | Number of graph edges loaded from YAML (full reindex) or live-DB total (incremental) |
docs_indexed | int | 0 | Number of Markdown documents indexed |
chunks_indexed | int | 0 | Number of document chunks created |
symbols_indexed | int | 0 | Number of code symbols extracted (full reindex) or live-DB total (incremental) |
imports_indexed | int | 0 | Number of code imports resolved |
rules_loaded | int | 0 | Number of architecture rules loaded from rules.yml |
test_files_indexed | int | 0 | Test files recorded in test_files (BDL-074 C1) |
test_files_unplaced | int | 0 | Of those, the files not under a kind folder, which bind to no node |
nothing_changed | bool | False | True when incremental reindex detects no file changes |
errors | list[str] | [] | Fatal errors encountered during reindex |
warnings | list[str] | [] | Non-fatal warnings (e.g., duplicate doc references) |
Constants
_TABLES_TO_DROP
Ordered list of tables dropped during full reindex. Order matters for foreign key constraints:
_TABLES_TO_DROP = [
"search_index", "test_imports", "test_files", "test_overrides",
"sync_state", "code_imports", "rules", "code_symbols", "chunks",
"docs", "declared_docs", "edges", "nodes", "meta",
]_CODE_EXTENSIONS
Frozen set of file extensions scanned for code symbols:
_CODE_EXTENSIONS = frozenset({
".py", ".ts", ".tsx", ".js", ".jsx", ".vue", ".go", ".rs",
".kt", ".kts", ".java", ".swift", ".m", ".mm", ".c", ".h", ".cpp", ".hpp",
})_EXT_TO_LANG
Mapping of file extensions to language labels for route extraction:
_EXT_TO_LANG: dict[str, str] = {
".py": "python", ".ts": "typescript", ".tsx": "typescript",
".js": "javascript", ".jsx": "javascript", ".go": "go",
".java": "java", ".kt": "kotlin", ".kts": "kotlin",
".graphql": "graphql", ".gql": "graphql", ".proto": "protobuf",
}_DEFAULT_SCAN_DIRS
Default source directories when config.yml has no scan_paths:
_DEFAULT_SCAN_DIRS = ("src", "lib", "app")Full Reindex Pipeline
reindex(project_root, *, docs_dir=None) executes the following steps in order:
| Step | Action | Module |
|---|---|---|
| 0 | Snapshot sync baselines (symbols_hash, file_symbols_hash, two-phase hashes and baseline PROVENANCE from sync_state) | _snapshot_sync_baselines |
| 1 | Drop all tables (_TABLES_TO_DROP) | _drop_all_tables |
| 2 | Create schema | infrastructure.db.create_schema |
| 3 | Load YAML graph from .beadloom/_graph/*.yml | graph.loader.load_graph |
| 3a | Copy each node's tests: declaration into test_overrides; a declaration that is not a list of path strings becomes a warning and binds nothing | test_index.record_declared_test_overrides |
| 3b | Store deep config in root node's extra | onboarding.config_reader.read_deep_config |
| 4 | Index Markdown documents from docs directory | doc_sync.doc_indexer.index_docs |
| 4b | Cache the DECLARED doc surface (every docs: entry, existing or not) | read_declared_docs / store_declared_docs |
| 4c | Index the TO-BE space in place, from the configured doc_roots | index_to_be_space → doc_sync.doc_indexer.index_space_documents |
Step 4c runs on BOTH paths. The incremental path rebuilds the TO-BE space wholesale rather than tracking it in file_index: the planning tree is small, and the two reindex paths disagreeing about what is in the index is a defect class this project has already paid for twice (BDL-UX #142, #146). | 5 | Extract and index code symbols from source files | context_oracle.code_indexer.extract_symbols | | 5b | Extract code imports and create depends_on edges | graph.import_resolver.index_imports | | 5c | Load architecture rules from .beadloom/_graph/rules.yml | graph.rule_engine.load_rules | | 5e | Analyze git activity and store in nodes.extra | _store_git_activity | | 5f | Extract API routes and store in nodes.extra | _extract_and_store_routes | | 5g | Populate file_index — BEFORE anything derives ownership from it | _populate_file_index | | 5h | Index test files into test_files / test_imports and rebuild every node's extra["tests"] from the binding | test_index.index_test_files | | 6 | Build sync_state with preserved symbol hashes for drift detection | _build_initial_sync_state | | 7 | Populate FTS5 search index | context_oracle.search.populate_search_index | | 8 | Clear bundle_cache, set meta, take health snapshot | Multiple internal functions | | 9 | Store parser fingerprint | _store_parser_fingerprint |
Step 5g moved ahead of step 6 in BDL-061.50, and the ordering is now load bearing rather than incidental. A node's owned files are read from file_index (so a module with no top-level symbol is still paired with its doc), and a full reindex drops every table first — so populating file_index at the END left the very first build of a fresh index with an empty table, and those pairs appeared only on the SECOND reindex. A checker whose input arrives after it runs reports clean because there was nothing there.
load_graph (step 3) also STATs each node's declared source: a directory written without a trailing slash is normalised to carry one, and a source that names no path on disk becomes a ReindexResult warning naming the ref_id — see the graph-loader component doc.
Test Index
test_index.py (BDL-074 C1) reads the project's test layout from the tests: block of .beadloom/config.yml (context_oracle.test_layout.load_test_layout, BDL-074 G2) and records every file whose project-relative PATH matches one of its patterns (default pytest, go_test, jest, junit, xctest; a pattern with a / matches the end of the path, so __tests__/** reads a whole folder, beadloom-2mj3.15) in three places: under each declared root (default tests/, test/, spec/ and a top-level __tests__/, each where it exists), under each build tool's test tree the project has (default src/test/java/, src/test/kotlin/, SwiftPM's Tests/), and — when the layout has beside_code: true, the default — among the code scan's files outside all of those. A test beside the code is taken from the code scan rather than a second walk. A root or test tree is read only when a folder of exactly that spelling exists, so a case-insensitive disk does not read Tests/ for tests or Spec/ for spec. A file anywhere else is not read, whatever its path: binding it would take a guess at its node, so ctx and the debt report say where a test file is read instead. __pycache__/, __snapshots__/ and node_modules/ are skipped. A walked folder is never a scan path, so a mutmut copy under mutants/ is never indexed unless a scan path covers it. Files under a root never enter code_symbols, code_imports or file_index: they must not become code. A test beside the code was already indexed as code by the code scan, and this step does not change that. Each file is bound by context_oracle.test_binding.bind_test_file (see the Test Mapping SPEC) and stored as:
test_files(path, kind, ref_id, placement, test_count, file_hash);test_imports(file_path, line_number, import_path, resolved_ref_id), each import resolved bygraph.import_resolver.resolve_import_to_node, memoised per import path. Imports are read from Python files only; a file in another language is counted by suffix and records none.
It runs after file_index is populated (step 5g), because a test's imports resolve through the same ownership as the code's. A file whose hash matches the recorded one is not parsed again. It then rebuilds nodes.extra["tests"] for every node with a source or a tests: declaration, in the four-key shape (framework, test_files, test_count, coverage_estimate). A parent's files are the union of its descendants' along part_of. framework is named from the pattern groups the node's bound files matched; a node with no bound file states the project's. The tests: declaration is read from the graph only at a full reindex (step 3a), because the rebuild overwrites the extra["tests"] it arrived in.
The layout is recorded in the index as meta.test_layout (TestLayout.recorded(), stored by RecordedTestLayout.encode(), each group's patterns included since beadloom-2mj3.15), so ctx, the debt report and the rule engine can state their counts against it without importing context_oracle. Since beadloom-2mj3.17 the record holds the roots that exist as roots and the roots in force that do not as absent_roots, both read from disk by _recorded_layout(), so a reader names only the roots a project has. is_test_index_current() compares the same record, so a root folder created or removed forces one test re-index. IndexedTestFiles.warnings carries a sentence for each unusable part of the tests: config block and one for each node's tests: prefix that covers no indexed test file, and both reindex paths add them to ReindexResult.warnings:
Node 'billing': `tests:` prefix 'tests/e2e' covers no indexed test file, so it binds nothing (a folder is declared with a trailing '/')beadloom reindex prints the placement counts after the other totals, on both the changed and the nothing_changed branch. On this repository, measured at 067df32a on 2026-09-29:
Tests: 623 files (278 bound to a node, 167 unplaced, 75 acceptance step, 103 self-check)The bound and unplaced counts are always printed. The bound count covers the mirror, override and beside_code placements, and names the last in parentheses when non-zero (N bound to a node (B beside the code), BDL-074 G2). W unowned appears only when non-zero. The other_kind files are named by their recorded kind, one entry per kind with its count (A acceptance step, S self-check; a kind without a label is named as recorded), because the two bind differently and one phrase over both was true of neither (BDL-074 F1). An index without the test tables prints no Tests: line. The line above is this repository's, measured by beadloom reindex on 2026-09-28.
Incremental Reindex Pipeline
incremental_reindex(project_root, *, docs_dir=None) follows this decision tree:
- Scan current project files and compute SHA-256 hashes.
- Read stored file hashes from
file_indextable. - Fallback to full reindex if:
file_indexis empty (first run or post-upgrade).- Parser fingerprint changed (new tree-sitter grammar installed).
- The index predates derived-edge provenance (
meta.import_edge_provenanceabsent or older). One rebuild is required because a deriveddepends_onedge is otherwise indistinguishable from a graph-declared one, so refreshing the first would delete the second. - The index predates the test tables (
meta.test_index_versionabsent or not1,needs_full_test_reindex). Only a full rebuild reads thetests:declarations. - Any graph YAML file changed, detected via
_graph_yaml_changed()which directly compares hashes for files withkind == "graph"(belt-and-suspenders check that catches changes even whenfile_indexis stale).
- Early return if no files changed and the test index matches the test files on disk (
is_test_index_current; test files under the roots are not infile_index, so a test-only change is detected by hashing the files under the roots and test trees, and a test layout that differs from the onemeta.test_layoutrecords is a change too) (setsnothing_changed=True, updates meta timestamp, takes health snapshot). - True incremental path:
- Snapshot
symbols_hashfromsync_statebefore modifications for drift preservation. - Delete old data for changed and deleted files (from
docs,code_symbols,sync_state). - Re-index changed and added files individually.
- Re-extract imports for the code files touched, forget those deleted, then rebuild the derived
depends_onedge set (reindex_file_imports). Without this stepcode_imports— and therefore everyforbid_import, cycle and layer rule — described the tree as it was at the last FULL rebuild, so the documentedreindex && lintloop reported a clean boundary over a real violation (BDL-UX #142). - Re-extract API routes and update
nodes.extra. - Rebuild
sync_statefrom scratch (full table delete + rebuild) with preservedsymbols_hash. - Rebuild FTS5 search index.
- Clear
bundle_cache(conservative invalidation). - Update
file_indexincrementally. - Rebuild the test index and
extra["tests"]wholesale (index_test_files): a code file added or removed can change what a mirror names. - Update meta timestamps and take health snapshot.
- Backfill result counts: Populate
nodes_loaded,edges_loaded, andsymbols_indexedwith live-DB totals (not per-run deltas), matching the behavior of thenothing_changedpath.
- Snapshot
Configuration
Configuration is read from .beadloom/config.yml:
| Key | Type | Default | Description |
|---|---|---|---|
docs_dir | str | "docs" | Relative path to documentation directory from project root |
scan_paths | list[str] | ["src", "lib", "app"] | Directories to scan for source code |
tests | mapping | see the Test Mapping SPEC | The test layout (BDL-074 G2): roots, kinds, patterns, mirrors, beside_code |
File Hashing
Files are classified into three kinds in the file_index:
| Kind | Source | Extensions |
|---|---|---|
"graph" | .beadloom/_graph/*.yml | .yml |
"doc" | <docs_dir>/**/*.md | .md |
"code" | <scan_paths>/**/* | .py, .ts, .tsx, .js, .jsx, .vue, .go, .rs, .kt, .kts, .java, .swift, .m, .mm, .c, .h, .cpp, .hpp |
Hashes are computed as: hashlib.sha256(file_bytes).hexdigest()
Doc-to-Node Reference Map
_build_doc_ref_map scans YAML graph files for nodes with docs lists and builds a {relative_doc_path: ref_id} mapping. When a doc path is referenced by multiple nodes, the first reference wins and a warning is emitted.
CLI Interface
beadloom reindex [--project DIR] [--docs-dir DIR] [--full]| Option | Type | Default | Description |
|---|---|---|---|
--project | Path | . | Path to the project root |
--docs-dir | Path | from config | Documentation directory |
--full | flag | False | Force full rebuild (drop all tables and re-create) |
By default, performs an incremental reindex (only changed files). Use --full to force a complete rebuild. When nothing_changed is detected, displays current DB totals instead of reindex counts. Warns about missing language parsers when symbols_indexed == 0.
API
Public Functions
def reindex(project_root: Path, *, docs_dir: Path | None = None) -> ReindexResultFull reindex: drop all tables, recreate schema, and reload everything from disk. Returns a ReindexResult with counts and diagnostics.
def incremental_reindex(project_root: Path, *, docs_dir: Path | None = None) -> ReindexResultIncremental reindex: only process files that changed since the last reindex. Falls back to reindex() when graph YAML changed or no prior file index exists. The returned ReindexResult has nodes_loaded, edges_loaded, and symbols_indexed populated with live-DB totals (not per-run deltas), ensuring accurate reporting even when the incremental path does not touch the graph.
def resolve_scan_paths(project_root: Path) -> list[str]Resolve source scan directories from .beadloom/config.yml. Returns ["src", "lib", "app"] when config is absent or has no scan_paths key.
Internal Functions
def _snapshot_sync_baselines(
conn: sqlite3.Connection,
) -> tuple[dict[str, str], dict[tuple[str, str], _SyncPairSnapshot]]Snapshot sync_state before the table drop. Returns ({ref_id: symbols_hash} for entries with a non-empty hash, {(doc_path, code_path): _SyncPairSnapshot}), or two empty dicts if the table does not exist yet (first run). Every pair is snapshotted, not only those carrying two-phase data: the snapshot also carries baseline_source, and a baseline that comes back without its provenance is indistinguishable from one that was earned (BDL-UX #175). It also carries file_symbols_hash, the pair's own file surface. _build_initial_sync_state carries it rather than recomputing it, and where there is none to carry it computes one only when _node_in_drift says the node's carried hash still matches the tree — a file fact invented against a node ALREADY in drift would contradict it and silently win, and recomputing a carried one would re-baseline against the tree just indexed, which is how integrating a parallel wave erased the drift it brought in (BDL-UX #133 / #175). Both reindex paths call this one function, so they cannot disagree about what survives a rebuild.
def _drop_all_tables(conn: sqlite3.Connection) -> NoneDrop all application tables to allow a clean re-create. Iterates _TABLES_TO_DROP in order.
def _resolve_docs_dir(project_root: Path) -> PathResolve docs directory from .beadloom/config.yml key docs_dir, defaulting to <project_root>/docs. Delegates to infrastructure.doc_roots.resolve_docs_dir, the single reader of the key: it was read in three places before, so a project keeping its documentation elsewhere had one reader looking where the others had not (beadloom-mr2l.75).
def _build_doc_ref_map(
graph_dir: Path,
project_root: Path,
docs_dir: Path,
) -> tuple[dict[str, str], list[str]]Build a mapping of {relative_doc_path: ref_id} from YAML graph nodes. Returns (ref_map, warnings). Built on read_declared_docs, so the doc→ref map and the declared surface are parsed once, from one place.
def read_declared_docs(
graph_dir: Path,
project_root: Path,
docs_dir: Path,
) -> list[tuple[str, str, str]]
def store_declared_docs(
conn: sqlite3.Connection,
declared: list[tuple[str, str, str]],
) -> NoneRead every doc a node DECLARES in its docs: list as (declared_path, doc_path, ref_id) — declared_path resolved project-relative (both spellings, docs/domains/x/README.md and domains/x/README.md, land on the same file) and doc_path docs-dir-relative, the key of the docs table — and cache them in declared_docs. The docs table only holds files found on disk, so without this a deleted doc simply stopped being indexed and the gate had nothing to miss (BDL-UX #174). Both reindex paths write it; a graph-YAML change already forces a full reindex. The graph directory is read through onboarding.graph_files.each_graph_file, which is where the skip policy is stated: BDL-069 measured that this body reads the directory for NODES, and it was the frame BDL-UX #220 measured init tracebacking in, twice — a hand-edited graph file that does not parse raised yaml.parser.ParserError here, and one holding a top-level list raised AttributeError on data.get, which no except yaml.YAMLError catches.
def _index_code_files(
project_root: Path,
conn: sqlite3.Connection,
seen_ref_ids: set[str],
) -> tuple[int, list[str]]Scan source files, extract symbols, insert into SQLite, and create touches_code edges for annotated symbols. Returns (symbols_indexed, warnings).
def _build_initial_sync_state(
conn: sqlite3.Connection,
*,
preserved_symbols: dict[str, str] | None = None,
preserved_pairs: dict[tuple[str, str], _SyncPairSnapshot] | None = None,
) -> NonePopulate sync_state table from docs and code_symbols with shared ref_ids. When preserved_symbols is provided, keeps old symbols_hash for drift detection; otherwise computes a fresh baseline. preserved_pairs carries the two-phase hashes and, through _baseline_provenance, each pair's baseline_source:
- no snapshot ->
index_build— this pair had no baseline before the build, so the hashes just written ARE the current tree and prove nothing on their own; - a snapshot -> its own value, verbatim.
index_buildthat survives a reindex is stillindex_build: a fabricated baseline does not become earned by being copied. A pre-provenance row ('') reads ascarried, since it genuinely came from an earlier generation.
check_sync corroborates an index_build pair against git rather than reporting it fresh.
def _load_rules_into_db(
rules_path: Path,
conn: sqlite3.Connection,
result: ReindexResult,
) -> NoneLoad architecture rules from rules.yml into the rules table. _serialize_rule covers every rule type the loader produces — deny, require, cycle, import-boundary, forbid-edge, layer, cardinality, unregistered-feature-candidate, module-coverage, scenario-coverage, doc-area-coherence, summary-facts, and since BDL-074 C3 the three suite rules — and raises TypeError on a type it does not know, so a rule type added to the loader without a serializer fails loudly instead of vanishing from the rules table. Each rule is stored WHOLE: a forbid_import exemption, a scenario_coverage.non_behavioural declaration and a doc_area_coherence threshold are all part of what the rule currently means, and a reader of the table must not see a stricter rule than the one that runs. summary_facts stores an empty definition because it has no configuration to store. A layer rule's exempt: entries are stored for the same reason, and since BDL-070 B4 there is a reader that needs them: the architecture view reads its layer rule from this table and asks that rule which edges to draw red, so an index without the entries would make the site flag crossings the Gate excuses. The key is written only when the rule declares entries, so the row of a project that excuses none is unchanged. The suite rules are stored whole by _serialize_suite_rule, with their exemptions for the same reason: test_binding as {files?, for?, exempt?}, where each exempt entry carries files or nodes with its reason and until; test_import_boundary as {from_glob, to_glob, of?, exempt?}, with forbid_import's {from, to, reason, until} entries; and scenario_binding as {features, exempt?}. A key marked ? is written only when the rule declares it.
Test Index Functions
Module src/beadloom/application/reindex/test_index.py:
record_declared_test_overrides(conn) -> list[str]-- copy each node'stests:list intotest_overrides; returns one warning per declaration that is not a non-empty list of path strings.discover_test_files(project_root, layout=None, *, code_files=()) -> dict[str, str]-- every test file the layout reads, by project-relative path, with its text: the files under the roots and the present test trees whose paths match a pattern and, when the layout reads tests beside the code, each of code_files outside those whose path matches one. A file anywhere else is not read. layout defaults to the one the project declares.present_mirror_roots(project_root, layout) -> tuple[str, ...]-- the build tools' test trees of the layout that exist, spelled exactly as declared (BDL-074 G2b).index_test_files(project_root, conn, *, code_files) -> IndexedTestFiles-- rebuildtest_files/test_importsand every node'sextra["tests"]; setsmeta.test_index_versionandmeta.test_layout.IndexedTestFiles.by_placementcounts files per placement, withtotalandunplacedproperties, andwarningsholds the layout problems and thetests:prefixes that bind nothing.is_test_index_current(project_root, conn) -> bool-- whether the recorded layout is the one the config declares andtest_filesholds exactly the files under the roots and test trees on disk, by hash. It is asked only when no code file changed, so a test beside the code is unchanged by construction.needs_full_test_reindex(conn) -> bool-- whether the index predates the test tables.placement_counts(conn) -> dict[str, int]-- files per placement, read fromtest_files. Since BDL-074 C2 it delegates toinfrastructure.repository.count_test_files_by_placement, becausectxand the debt report state the same counts and neither may import the reindex.kind_counts(conn) -> dict[str, int]--other_kindfiles per recorded kind (BDL-074 F1), delegating toinfrastructure.repository.count_other_kind_test_files.describe_placements(counts, kinds) -> str-- the text afterTests:on the reindex output; kinds iskind_counts, each named throughinfrastructure.repository.label_test_kind.
change_detection.code_paths(files) -> frozenset[str] gives the code paths of a _scan_project_files result, as POSIX paths, which the mirror resolves against.
_store_test_mappings was removed in BDL-074 C1, together with its entry in the package __all__: the heuristic mapper no longer writes extra["tests"].
def _update_node_extra(conn: sqlite3.Connection, ref_id: str, key: str, value: object) -> NoneMerge a key/value into a node's extra JSON column. Does nothing if ref_id does not exist.
def _drop_node_extra_key(conn: sqlite3.Connection, ref_id: str, key: str) -> NoneRemove a key from a node's extra JSON column. Writes nothing if the node does not exist or does not carry the key.
def _scan_routes(project_root: Path) -> list[dict[str, object]]Every API route under the scan directories, each with method, path, handler, the project-relative file, line and framework.
def _extract_and_store_routes(project_root: Path, conn: sqlite3.Connection) -> NoneScan source files for API routes using _EXT_TO_LANG for language detection and store aggregated results in nodes.extra["routes"]. A route is stored on every node whose source its file lies under, by path component, as infrastructure.node_source.NodeSource.holds answers: the file is the source, or continues it past a /. A node whose source is '' or absent is given no route. Until BDL-069 beadloom-rqma.4 the match was a string prefix, so src/ledger/ took the routes of src/ledger_archive/ and a root with source: '' took every route. The store is the whole answer, not a merge: a node that holds no route now loses its routes key, so a route deleted from the code, or attributed under the old rule, is withdrawn by the next reindex that runs. Before, a node was only ever written when it had routes, and measured on a foreign repository an incremental reindex after a code change left the prefix-attributed route in place. An incremental reindex that finds no changed file returns before this step, so an index built under the old rule keeps those routes until a file changes or reindex --full runs.
def _store_git_activity(conn: sqlite3.Connection, project_root: Path) -> NoneAnalyze git activity via analyze_git_activity() and store results in nodes.extra["activity"] (level, commits_30d, commits_90d, last_commit, top_contributors).
def _compute_file_hash(path: Path) -> strCompute SHA-256 hex digest of a file's contents.
def _scan_project_files(
project_root: Path,
docs_dir: Path,
) -> dict[str, tuple[str, str]]Scan all project files and return {relative_path: (sha256_hex, kind)}.
def _get_stored_file_index(conn: sqlite3.Connection) -> dict[str, tuple[str, str]]Read file_index from DB. Returns {path: (hash, kind)}. Filters out sentinel rows (paths starting with __).
def _diff_files(
current: dict[str, tuple[str, str]],
stored: dict[str, tuple[str, str]],
) -> tuple[set[str], set[str], set[str]]Compare current vs stored file index. Returns (changed, added, deleted) sets of relative paths.
def _graph_yaml_changed(
current_files: dict[str, tuple[str, str]],
stored_files: dict[str, tuple[str, str]],
) -> boolCheck whether any graph YAML file was added, removed, or modified by directly comparing hashes for files with kind == "graph". This belt-and-suspenders check catches changes even when file_index is stale.
def _populate_file_index(conn: sqlite3.Connection, current_files: dict[str, tuple[str, str]]) -> NoneReplace the entire file_index with current files (used after full reindex).
def _update_file_index(
conn: sqlite3.Connection,
current_files: dict[str, tuple[str, str]],
changed: set[str],
added: set[str],
deleted: set[str],
) -> NoneIncrementally update file_index for affected paths (used after incremental reindex).
def _index_single_doc(conn, md_path, docs_dir, ref_map) -> tuple[int, int]Index one doc file. Returns (docs_count, chunks_count).
def _index_single_code_file(conn, file_path, project_root, seen_ref_ids) -> intIndex one code file. Returns symbol count.
Public Classes
@dataclass
class ReindexResult:
nodes_loaded: int = 0
edges_loaded: int = 0
docs_indexed: int = 0
chunks_indexed: int = 0
symbols_indexed: int = 0
imports_indexed: int = 0
rules_loaded: int = 0
test_files_indexed: int = 0
test_files_unplaced: int = 0
nothing_changed: bool = False
errors: list[str] = field(default_factory=list)
warnings: list[str] = field(default_factory=list)Invariants
- Full reindex always snapshots
symbols_hashbaselines via_snapshot_sync_baselines()before dropping tables, preserving drift detection state. The incremental path takes the SAME snapshot (it used to re-derive one inline, which silently dropped baseline provenance). - A baseline's provenance is never promoted by a reindex: only an attestation (
sync-update) or an observed doc edit upgrades it. - Full reindex always drops ALL tables before recreating them (clean slate guarantee).
- WAL mode is enabled on every database connection opened by
open_db. - Foreign keys are enabled per-connection via
open_db. - File hashes are SHA-256 hex digests.
- Incremental reindex always rebuilds
sync_statefrom scratch (full delete + rebuild) even though only some files changed, using preservedsymbols_hashvalues. - Incremental reindex always clears
bundle_cache(conservative invalidation). - Incremental reindex re-extracts API routes after code changes.
- A route, a symbol handed to
docs polishand a commit counted as activity are attributed to a node by one rule,NodeSource.holds. None of those three readers compares a file path with a source by itself, and a test asserts it of each. Ownership, which node a file BELONGS to, is a different rule and lives inrepository.source_covers. - Incremental reindex re-extracts imports for changed/added code files and deletes those of removed files, then rebuilds the derived
depends_onedges, so the incremental import graph is identical to the one a full rebuild produces. Measured on Beadloom's own tree (67 nodes, 1255 symbols, 1322 imports): +29 ms for one changed file, +42 ms for five, against 755 ms for a full rebuild. - Only
depends_onedges carryingextra.derived = "imports"are deleted by that refresh; an edge declared in the graph YAML is never touched. - Incremental reindex backfills
nodes_loaded,edges_loaded, andsymbols_indexedwith live-DB totals (not per-run deltas), ensuring accurate reporting even when the incremental path does not touch the graph or code symbols. file_indexis fully replaced after full reindex and incrementally updated after incremental reindex.- Meta key
last_reindex_atis updated on every successful reindex (including no-change incremental runs). - Test files are never written to
code_symbols,code_importsorfile_index, and onlytests/is walked for them. - Both reindex paths rebuild
extra["tests"]from the binding, so every node with asourceor atests:declaration carries the four-key shape. _graph_yaml_changed()performs a direct hash comparison on graph files by kind, independent of_diff_files(), to catch changes even whenfile_indexis stale.
Constraints
- Full reindex is not atomic: it drops all tables then recreates them. A crash mid-reindex leaves the database in an incomplete state. Re-running reindex resolves this.
- Incremental reindex conservatively invalidates
sync_stateandbundle_cacheentirely, even when only a single file changed. - Any graph YAML change (
.beadloom/_graph/*.yml) forces a full reindex. There is no incremental graph update path. - The
file_indextable must exist and be populated for incremental reindex to work. An empty or missingfile_indextriggers automatic fallback to full reindex. _build_doc_ref_mapresolves doc path conflicts by keeping the first reference. Subsequent references to the same doc from different nodes emit warnings but do not overwrite.- Code symbol indexing depends on
tree-sitterbeing available for the target language. Missing parsers result in zero symbols for that file (not an error).
Testing
Test files bound to reindex (110 tests in 12 files, measured by beadloom reindex on 2026-09-28): tests/integration/application/reindex/ (test_reindex.py, test_reindex_config.py, test_reindex_tests.py, test_reindex_indexes_test_files_in_their_own_tables.py, test_reindex_activity.py, test_reindex_routes.py, test_an_incremental_reindex_refreshes_the_imports.py, test_the_sync_baseline_is_rebuilt_with_one_provenance.py, test_tests_are_found_beside_the_code_and_under_declared_roots.py and test_a_build_tools_test_tree_is_indexed_by_its_mirror.py, BDL-074 G2) and tests/unit/application/reindex/ (test_a_graph_change_is_detected_by_its_hashes.py, test_the_reindex_hub_keeps_its_exports.py). The CLI surface is tested in tests/test_cli_reindex.py, which binds to no node yet (placement unplaced). The Tests: line's kinds are tested in tests/integration/services/commands/index_ops/test_the_tests_line_names_each_kind.py, bound to cli-commands.
Tests should cover the following scenarios:
- Full reindex end-to-end: Verify that a project with YAML graph, docs, and source code produces a populated database with correct counts in
ReindexResult. - Sync baseline preservation: Verify
_snapshot_sync_baselines()capturessymbols_hashbefore drop and that_build_initial_sync_state()restores them. - Incremental no-change: Verify
nothing_changed=Truewhen no files have been modified since the last reindex. - Incremental doc change: Modify a Markdown file, run incremental reindex, verify the doc is re-indexed and chunks updated. Verify
nodes_loaded,edges_loaded, andsymbols_indexedare populated with live-DB totals. - Incremental code change: Modify a source file, run incremental reindex, verify symbols are re-indexed. Verify
symbols_indexedreflects the live-DB total. - Incremental file addition: Add a new file, verify it appears in results.
- Incremental file deletion: Delete a file, verify its data is removed from the database.
- Graph YAML change triggers full reindex: Modify a
.beadloom/_graph/*.ymlfile, verify incremental falls back to full reindex via_graph_yaml_changed(). - Parser fingerprint change triggers full reindex: Verify that a changed parser fingerprint causes incremental to fall back to full.
- Empty file_index triggers full reindex: On a fresh database, verify incremental falls back to full reindex.
- Config resolution: Verify
resolve_scan_pathsand_resolve_docs_dircorrectly read fromconfig.ymland fall back to defaults. - Doc ref map conflicts: Create YAML nodes referencing the same doc path, verify warnings are emitted and the first reference is kept.
_diff_files: Unit test with known current/stored dicts to verify correct changed/added/deleted sets.- Test index: Verify
index_test_files()records test files intest_files/test_imports, never incode_symbolsorfile_index, skipsmutants/, rebuildsnodes.extra["tests"]in the four-key shape, and that a test-only change is picked up by an incremental reindex. - Git activity: Verify
_store_git_activity()populatesnodes.extra["activity"]. - Route extraction: Verify
_extract_and_store_routes()populatesnodes.extra["routes"]. - Route attribution:
tests/integration/infrastructure/node_source/test_a_file_lies_under_a_source_by_one_rule.pyruns one table of source shapes against the route store,docs polishand git activity, and withdraws a stale route through a real incremental reindex.tests/acceptance/features/routes_under_source.featurerunsinit,reindexanddocs polishend to end.