Skip to content

✅ fresh

last synced 2026-09-29T21:12:19.681221+00:00 · coverage 100% (agent-prime)

Validation by Beadloom doc_sync — same source as sync-check.

Agent Prime ​

Cross-IDE context injection via a three-layer architecture.

Layers ​

  1. AGENTS.md (static) — .beadloom/AGENTS.md with MCP tool list, architecture rules, conventions. Regenerated by generate_agents_md(), preserves content below ## Custom marker.

  2. IDE adapters (pointers) — .cursorrules, .windsurfrules, .clinerules files that instruct agents to read AGENTS.md. Created by setup_rules_auto() on bootstrap.

  3. The ignore block (one write) — bootstrap_project() also calls ignore_block.ensure_ignore_block(), so the derived state bootstrap is about to create under .beadloom/ (the index, and later the guard firing record) does not arrive as untracked churn in the adopter's repository. It is appended once and never rewritten, the result is returned as ignore_added / ignore_skipped_reason, and beadloom init prints it. See the ignore-block component.

  4. The domain-parent post-condition (one invariant, two writers) — every node bootstrap_project() writes leaves the function with at least one outgoing part_of edge, to its classified parent where one exists and to the root service node otherwise. It reads no kind since BDL-067 .17: until then it said kind: domain, the kind today's generated rules require a parent for, while generate_rules also writes feature-needs-parent and the sibling statement in import_docs already covered every kind — two writers of one invariant disagreeing about its population (the review of .16, minor 2). One node is carved out and the sentence says so since BDL-067 .6: a node whose ref_id is the root's own gets no edge, because an edge from a node to itself is not a parent. That ref_id collision was reachable on src/<project>/ until BDL-069 and is not produced by either writer now — see Two nodes never share a ref_id below; the carve-out stays because this function also runs over a graph a hand edit can reach. The same call writes domain-needs-parent into the adopter's rules.yml whenever it writes a domain, so a domain without that edge makes init exit 0 over a graph the very next lint --strict rejects — measured as rc 0 then rc 1 on a flat TypeScript project (BDL-UX #192). parent_edges.missing_parent_edges(nodes, root_ref_id, parented) enforces it over the whole node list rather than inside the branch that was reported, and reads root_ref_id back from the root node instead of recomputing it, because cluster refs pass through _sanitize_ref_id and the root ref does not.

    import_docs() is the second writer of domain nodes and did not receive the invariant until BDL-067 .14: every document the classifier could not place became a domain with no parent, in imported.yml, in the same run that wrote domain-needs-parent at error severity — so init --yes --mode both exited 0 and the next lint --strict exited 1 on three nodes, BDL-UX #192's signature on a branch the epic had declared covered. It calls the SAME function bootstrap_project calls, and _existing_graph(graph_dir) finds the root by reading the graph on disk: the one node of kind service that no part_of edge leaves. It attaches every kind, not only the two the generated rules require a parent for, because a post-condition that tracked today's rule set would go stale the next time generate_rules gains a rule. With no single root — an import-only run on a virgin project, or a graph with two unparented services — no parent is named rather than one guessed.

    Two nodes never share a ref_id, and that is a writer's duty rather than the loader's. The graph identifies a node BY its ref_id, so a writer that emits one name twice does not write two nodes: it writes one, and the report it prints counts two. On the ordinary single-package src-layout — a project called myapp holding src/myapp/ — the root service takes its name from the manifest and the package takes its from the directory, and those are one string. Measured on the published 4.0.0 wheel: init --yes --mode bootstrap reported Graph: 2 nodes and beadloom status then reported Nodes: 1. The node dropped was the one carrying source: src/myapp/, so the package the project is named after was absent from every answer the graph gave, and domain-needs-parent went inert rather than red — a rule cannot fail over a node the graph does not hold (BDL-UX #214). The same shape reached import_docs through the document's file name: two documents called setup.md in two directories, or one named after the project, asked for a ref_id the graph already held.

    scanner/ref_ids.RefIdAllocator hands out every ref_id both writers use. take(preferred, qualifier=...) returns preferred when it is free — which is every node on every project the collision does not touch — then preferred-qualifier, then preferred-qualifier-2 and upwards. Callers pass the node's kind as the qualifier, so the second myapp is written as myapp-domain and reads as what it is. WHICH asker keeps the plain name is the caller's judgement and not the allocator's: bootstrap_project gives the root the project's name first, before any cluster asks, because that ref_id titles the architecture document and is what generate_rules names as the parent every domain must have. import_docs seeds the allocator with the ref_ids already on disk (_existing_graph(...).ref_ids), because it adds to a graph another writer produced; imported.yml is not among the files read, so a re-import hands out the ref_ids it handed out last time.

    A rename reaches the EDGES as well as the node. Three passes used to recompute a cluster's ref_id from its directory name — the manifest dependency loop, _quick_import_scan and the top-level attachment loop — and a rename that reached only the node would have traded a lost node for an edge naming nothing, which the loader drops just as quietly. All three read the ref_id the cluster was written under, and _quick_import_scan takes that mapping as a parameter rather than recomputing it. The attachment loop's old carve-out — skip the cluster whose sanitized name equals the project's — is gone with the collision it worked around.

    One post-condition, one implementation, since BDL-067 .21. Until then the two writers carried the same private name in two modules with the same loop body, differing in a str() call and in their parameter order, and the only thing binding them was a docstring naming a symbol .17 had renamed away. They had already drifted once and been repaired by editing both — one covered every kind, the other only domain (the review of .16, minor 2; the review of .20, major 3). scanner/parent_edges.py now holds missing_parent_edges(nodes, root_ref_id, parented) and parented_by(edges), and both writers import them. What stays each writer's own is the one thing that genuinely differs: where parented comes from. The bootstrap produces the whole graph and reads it off the edges it is about to write; the importer adds to a graph and reads it off the one on disk. A THIRD writer of nodes is caught rather than repaired afterwards: tests/test_one_parent_post_condition_over_every_writer.py derives the writers from the source — every function that reaches write_yaml_atomic and builds a payload holding nodes — and fails when that set changes.

    ONE POLICY FOR THE READING SIDE TOO, since BDL-067 .24. graph_files.each_graph_file(graph_dir, *, also_skip=frozenset()) states the skip policy once: a file whose name is not a graph file's is skipped, a file that will not read or will not parse is skipped, and a file that parses to anything other than a mapping is skipped — so a caller may read data["nodes"] without asking again whether it can. rules.yml belongs to the policy, because a rules file is not a graph file for any reader; also_skip is the one genuine difference between the callers and is passed at the call site, by doc_classify._existing_graph alone, naming imported.yml because the run that asks is about to replace it. There were four bodies with four policies until .24 — doc_generator._load_graph_from_yaml, doc_generator._patch_docs_field, doc_classify._existing_graph and setup._graph_file_of_each_node — and two carried no guard at all. .21 removed generate_skeletons' node-list parameter, so init --bootstrap began reading the tree through an unguarded one and a hand-edited .beadloom/_graph/legacy.yml that does not parse reached the adopter as a raw yaml.parser.ParserError traceback, while the same commit added exactly that guard to two of the siblings and listed it as delivered (the review of .23, major 3). The mapping guard is separate from the parse guard and is needed separately: a graph file holding a top-level list parses without complaint and then raises AttributeError on data.get. tests/test_graph_files_are_read_under_one_policy.py derives the readers from the source — a function that both LISTS a directory and PARSES YAML, by any name the standard library or PyYAML offers for either — and fails when a fifth body appears under onboarding/ or in setup.py. The detector was widened at BDL-067 .25: .24's asked for glob with the literal "*.yml" and for yaml.safe_load by name, which is the spelling each_graph_file happens to use rather than what makes a body a reader, and five bodies that read the directory were measured passing it.

    The policy covers the readers init's own modules hold and no others, which is a scope and not a claim about the command. init still ends in a Python traceback on a graph file it cannot handle, and BDL-067 did not close that (BDL-UX #220, open). MEASURED at .25 over init's own eight (entry point x mode) cells crossed with three shapes of a hand-edited .beadloom/_graph/legacy.yml: 24 runs, of which the 15 that reach the file traceback — --bootstrap, --import and all three wizard modes, on a file that does not parse, on a file whose top level is a list, and on a file carrying an unquoted date. Two frames, both outside onboarding: application/reindex/indexing.read_declared_docs and graph/loader.load_graph. The date shape is the one that shows a parse guard would not be enough — every reader init owns yields that file happily. --yes reaches none of them, and not because a guard works: non_interactive_init returns skipped when .beadloom/ already exists, and --force deletes the directory first. graph/diff.py, reindex/change_detection.py and services/commands/index_ops.py walk the directory too. The measurement is pinned in the same test file, which fails as soon as somebody closes it.

    "No single root" counts DISTINCT ref_ids since BDL-067 .17. The candidates were collected into a list and counted there, and bootstrap_project produces one root under two node entries on an ordinary project shape: it writes the root service node under the project name and its top-level attachment loop skips the cluster whose sanitized name equals that name. A repository named after one of its own source directories therefore read as two candidates, the import attached nothing, and init --yes --mode both exited 1 on every run — measured on a project named core holding src/core/ and src/orders/ (the review of .16, major 1). The graph identifies a node by its ref_id and the loader keeps one node per ref_id, so the population the uniqueness test ranges over is the distinct ref_ids.

    An import-only run on a virgin project leaves those domains unparented and is green, which rests on one fact and no longer on two: no rules.yml is on disk, so lint --strict evaluates nothing. The second reason — that --mode import reached no verdict at all — was withdrawn in .17. It held only until the next init on the same tree, because the wizard's re-init does not delete .beadloom/, so imported.yml survived into a later bootstrap that wrote the rule and met the nodes. Every branch of init that writes a file under .beadloom/_graph/ now takes the verdict.

  5. beadloom prime (dynamic) — CLI command and MCP tool that queries the DB for current project state: architecture summary, stale doc-code pairs, lint violations, domain list.

  6. One order for all three entry points (BDL-067 .18 and .21, BDL-UX #216) — init bootstraps, then imports, then generates the doc skeletons, whether the mode arrived through --yes --mode, through --bootstrap or through the wizard's prompt. non_interactive_init() used to generate the skeletons inside its bootstrap block and import after them, so under --mode both it classified the documents it had written seconds earlier. Measured on a project with src/orders/, src/catalog/ and one document of the adopter's own: --yes --mode both imported four documents — architecture, readme, readme, payments — where the wizard answering both imported one. Three of the four are Beadloom's own scaffolding, and the two readme nodes are a single ref_id, because the importer names a node after the file stem and the loader keeps one node per ref_id: that graph had already dropped a document it claimed to describe. The defect predates this epic, and .14 changed its character — the parent post-condition gives every imported node a part_of edge to the root, so the wrong graph became structurally valid and both runs exited 0.

    The ORDER fixes it rather than a filter over the import scan. An exclusion would have to name docs/architecture.md and docs/domains/*/README.md, and those are the ADOPTER's documents whenever the adopter wrote them first — generate_skeletons() never overwrites an existing file — so a filter would drop a real document from the graph and leave the two entry points disagreeing about a different population. generate_skeletons() reads every graph file on disk: rendering docs/architecture.md from the bootstrap's nodes alone produced a whole-tree document describing part of the tree, which is the same divergence one file further on.

    .18 closed that between --yes and the wizard by changing one call. There are three entry points, and --bootstrap still passed its node list: on a tree carrying a graph file an earlier run had left, init --bootstrap and the wizard answering bootstrap left different trees — only the wizard wrote docs/domains/ledger/README.md, only the wizard's docs/architecture.md named ledger, and only the wizard's run patched that node's docs: field back into the file holding it (the review of .20, major 1). BDL-067 .21 closed it by REMOVING the parameter rather than by editing the third call site: generate_skeletons(project_root) takes the project root and nothing else, so a document about the whole tree cannot be handed part of the tree by any caller, including one written later. That is an API change for anyone importing beadloom.onboarding.generate_skeletons.

  7. A path the render broke in half is a path nobody can copy. The wizard's edit answer names the graph file to open, and rich hard-wraps at the console width — 80 when the output is not a terminal — inserting a real newline wherever the line runs out, token or no token. Whether that lands inside .beadloom/_graph/services.yml depends on how long the project's own path happens to be, so the assertion over it was green on macOS for nine consecutive runs and red on all six CI legs, whose temporary prefix is 68 characters and leaves exactly 12 before the break. The message now prints with soft_wrap, so the emitted text carries no inserted newline and the terminal folds it visually instead; test_the_path_it_names_survives_the_render_whole asserts the tail contiguously, which is the only form of the claim a wrap cannot satisfy by accident.

API ​

prime_context(project_root, *, fmt="markdown") ​

Returns compact project context. Static layer (config, rules, AGENTS.md) always available. Dynamic layer (DB queries) degrades gracefully without DB.

  • fmt="markdown" — human-readable output (~1000-1500 tokens)
  • The finding lists are bounded (MAX_LISTED_FINDINGS = 10, BDL-061 S4). prime used to print one line per stale pair and per lint violation with no limit, which kept its size promise only while the lists were empty: opting this repository into scenario-coverage (68 findings) grew the output from 2.6 KB to 13.1 KB — five times the budget, in the artifact whose whole job is to fit in one. The count is never truncated, only the list, and the cut says how many are hidden and which command shows them (beadloom lint / beadloom sync-check)
  • The stale list counts and names PAIRS (BDL-069 beadloom-yn6i). A pair is a document AND a code file, so three code files of one package give three stale pairs over one document. The health line reads N stale pair(s), the section is ## Stale Pairs, and each line is - <doc> <-> <code> (<ref_id>), the pair as sync-check's text renders it. Measured before the change on a repository with one README over three stale pairs: Health: 3 stale docs above three identical lines naming the README alone. The cut note says stale pair(s) for the same reason. fmt="json" is unchanged: health.stale_docs already carried doc_path, code_path and ref_id per pair
  • The health line states what the violation count was taken over (BDL-070 A4). N lint violations says the same words whether the layer rule judged 357 of 365 live depends_on edges or all 365, so the line carries a clause per declared layer rule: Health: 0 stale pair(s), 55 lint violations, architecture-layers judged 357 of 365 live depends_on edge(s) | Last reindex: …, printed on this repository on 2026-09-13. The wording is graph/rules/layer_reach.py::population_phrase, shared with the Gate line and the four lint renderings, so one fact has one form. It is one clause per RULE and not per finding, so the bounded list above can grow without it growing; a project that declares no layer rule gets no clause. fmt="json" carries the same list under health.layer_populations. Both facts come from ONE lint run, held together on prime.LintSnapshot: reading them separately would lint twice, and a count and a denominator taken over two different indexes are worse than neither
  • fmt="json" — structured dict for programmatic use

setup_rules_auto(project_root) ​

Auto-detects IDEs by marker files and creates adapter files. Returns list of created file paths.

generate_agents_md(project_root) ​

Generates .beadloom/AGENTS.md with v2 template. Injects rules from rules.yml. Preserves user content below ## Custom.

Each injected rule is labelled by scanner/rules_gen._detect_rule_type(), which reads the authoring keys from graph.rules.loader.AUTHORING_KEYS rather than from a copy of its own (BDL-073 B3), so every key the loader accepts has a label and the word unknown is left for a rule that names none. check is labelled cardinality and forbid forbid_edge; every other key is its own label. A rule naming two keys, which the loader rejects, is labelled by the first one its author wrote.

refresh_claude_md(project_root, *, dry_run=False) ​

Regenerates the auto-managed regions of .claude/CLAUDE.md between <!-- beadloom:auto-start SECTION --> / <!-- beadloom:auto-end --> markers and returns the names of the regions whose content changed. Two regions are rendered:

  • project-info — stack, dependencies, tests, linter, typing, packages and version, every one of them read from the target project's own manifest, tree and flow.yml. A fact that cannot be read is omitted, never substituted, and a section with nothing to say prints what it looked for rather than nothing at all.

    This sentence was already in this SPEC while the code did otherwise, which is worth stating plainly. Until BDL-UX #183 the version bullet was application.doctor.get_actual_version() — Beadloom's __version__, correct about Beadloom (BDL-UX #92) and false for every adopter — the architecture line said DDD packages whatever the project declared, the stack line matched the project's manifest against Beadloom's own dependency names (sqlite, click, rich, tree-sitter) and stated our Python floor as theirs, and the package scan fell back to looking for src/beadloom/ inside the adopter's tree. Each read correct on this repository by coincidence. :mod:beadloom.onboarding.scanner.project_facts now owns every one of those reads, and tests/support/adopter_project.py renders non-Beadloom fixtures so a coincidence cannot pass for a measurement again.

  • doc-language — the "ALL documents MUST be written in …" sentence, derived from language: in .beadloom/flow.yml (default en). The scaffolded flow used to state English unconditionally, so a team documenting in another language had to override the shipped default in prose (BDL-UX #136).

project_facts — what the TARGET project declares about itself ​

One module, one responsibility: read a fact out of somebody else's project. Nothing in it may consult beadloom.__version__, our package layout, or anything else true of this repository — every value comes from a file under project_root.

FunctionReadsUnknown
detect_project_version(root)pyproject.toml ([project], [tool.poetry], then a dynamic version via [tool.hatch.version] / [tool.setuptools.dynamic]), package.json, Cargo.tomlNone — the bullet is not rendered
detect_source_packages(root)src/<pkg>/<child>/__init__.py under the targetempty set
detect_requires_python(root)requires-python, verbatimNone
detect_declared_dependencies(root)[project].dependencies or package.json dependencies, first six, declared orderempty tuple
manifest_text(root)every readable dependency manifest, concatenatedNone — "nothing declares this" and "we could not look" are different answers

A VCS tag is deliberately not consulted: a tag is a release marker on a commit rather than a statement the project makes about itself, and reading one would need the infrastructure layer onboarding is forbidden to import.

The same reads back beadloom doctor's audit of an adopter's CLAUDE.md. Before BDL-UX #183 that audit compared four claims — version, packages, stack, test framework — against Beadloom's state (get_actual_version(), a scan for src/beadloom/, the literal keyword set {python, sqlite} and the literal string pytest). A TypeScript adopter with a correct file was told their stack claim was missing keywords and their test framework was not pytest. Each check now reads the project, and a fact the project does not declare is reported INFO … not verified rather than as drift — unknown is not zero, and it is not a verdict either.

blank_auto_regions(text) ​

Replaces every auto-region BODY with a fixed token. config-check compares the composed part of CLAUDE.md against compose("claude", ...); the regions are generated per project and are supposed to move, so blanking them keeps the two checks from reporting each other's drift.

setup_mcp_auto(project_root) ​

Auto-detects editor (claude-code, cursor, windsurf) by marker files and creates MCP config. Returns editor name on success, or None if config already exists.

Data Types ​

ScanResult (TypedDict) ​

Result of scan_project() — discovered project structure.

python
class ScanResult(TypedDict):
    manifests: list[str]      # Found manifest files (pyproject.toml, package.json, etc.)
    source_dirs: list[str]    # Discovered source directories
    file_count: int           # Total code file count
    languages: list[str]      # Detected file extensions

ClusterEntry (TypedDict) ​

A two-level directory cluster produced by _cluster_with_children().

python
class ClusterEntry(TypedDict):
    files: list[str]                    # All code files in the cluster
    children: dict[str, list[str]]      # Child dir name → code files
    source_dir: str                     # Owning top-level source directory

IndexCounts / Reindexer (reindex_port.py) ​

What init needs from the layer above it, stated as a type onboarding owns.

python
class IndexCounts(Protocol):
    symbols_indexed: int    # code symbols the run indexed
    imports_indexed: int    # import statements the run indexed
    edges_loaded: int       # graph edges the run loaded
    docs_indexed: int       # documents the run indexed (the wizard reports this one)

Reindexer = Callable[[Path], IndexCounts]

interactive_init(project_root, *, reindex) and non_interactive_init(project_root, *, reindex, mode=..., force=...) take the re-index as a required keyword argument, and services/commands/setup.py supplies application.reindex.reindex.

Why it is handed in. The declared direction is services → application → domains → infrastructure; onboarding is a domain and the re-index is an application use case, so init_flow.py importing beadloom.application.reindex ran against it. It did, twice and function-locally, and that was the only reverse-direction edge of the 357 that inheritance through part_of brings into the layer check's scope on this repository once layer membership is inherited through part_of (BDL-070 beadloom-46am). The inversion moves only where the callable comes from: the call itself still runs after every block that writes a graph file, which is the ordering BDL-067 .14 and .18 established.

Why it is required rather than defaulted. A default meaning "do not re-index" would let a caller take init's verdict over an index the run never refreshed — rc 0 from init, rc 1 from the adopter's next lint --strict, which is exactly what .14 closed. A default that resolves the import lazily would be the layering violation unchanged, and importlib.import_module would remove the derived edge by hiding the import from the scanner, which is the BDL-059 S3 workaround tests/test_no_domain_package_imports_application.py exists to prevent. A required argument fails at the call site instead.

CLI ​

  • beadloom prime [--json] [--update] [--project PATH]
  • beadloom setup-rules [--tool cursor|windsurf|cline] [--project PATH]
  • beadloom setup-mcp [--tool claude-code|cursor|windsurf] [--remove] [--project PATH]

MCP ​

  • prime tool (no parameters) — returns JSON context for agent sessions

Source ​

  • src/beadloom/onboarding/scanner/ — cohesion-split package; prime.py (prime_context()), agents_md.py (setup_rules_auto(), generate_agents_md(), setup_mcp_auto()), types.py (ScanResult, ClusterEntry), plus bootstrap.py / init_flow.py / project_scan.py / summary.py / entry_points.py / import_scan.py / readme.py / doc_classify.py / rules_gen.py / claude_md.py / constants.py / reindex_port.py; the package __init__.py re-exports the full public surface
  • src/beadloom/services/commands/query.py — prime CLI command
  • src/beadloom/services/commands/setup.py — setup-rules and setup-mcp CLI commands
  • src/beadloom/services/mcp_server.py — prime MCP tool