Skip to content

✅ fresh

last synced 2026-09-29T21:12:19.681221+00:00 · coverage 100% (docs-audit)

Validation by Beadloom doc_sync — same source as sync-check.

Documentation Audit ​

Zero-config detection of stale facts in project markdown documentation.

Source: src/beadloom/doc_sync/audit.py, src/beadloom/doc_sync/scanner.py, src/beadloom/doc_sync/audit_coverage.py

Specification ​

Purpose ​

The documentation audit feature detects stale numeric facts in markdown documentation by comparing mentioned values against ground-truth data extracted from the project. It uses a two-pass architecture: first collecting facts from project infrastructure (manifest files, graph DB, code symbols, MCP tools, CLI commands), then scanning documentation for keyword-proximate numbers and version strings, and finally comparing the two to produce stale/fresh findings.

Architecture ​

The audit pipeline has three stages:

  1. FactRegistry (audit.py) -- Collects ground-truth facts from multiple project data sources: pyproject.toml (version), graph DB (node/edge/test/framework/rule counts), the MCP tool catalog (tool count), Click CLI introspection (command count), and code_symbols (language count). Extra facts can be injected via config.yml. A source that cannot produce a value records the reason instead of dropping the fact, and the two surface facts are gated by audit_self_surface.py so they are declared only for the project that provides them.

  2. DocScanner (scanner.py) -- Scans markdown files for numeric and version string mentions. A line is tokenized on whitespace first, and only a token whose whole core is a number is a candidate (Layer 0); each candidate is then associated with a fact type based on nearby keywords within a configurable proximity window. False positives (dates, hex colors, issue IDs, line references, version pins) are masked before extraction.

    scan_line is the seam scan_file is written in terms of, and the entry point for prose that is not in a file at all. A node summary in the architecture graph is one line of exactly this kind, and the graph-summary-facts lint rule reads it here (BDL-062 .1). It exists so this codebase holds ONE notion of "a version" and ONE keyword table: a second extractor built beside this one would agree with it on the day it was written and diverge on the first token-boundary or clause-scope repair that landed on only one of the two. origin is what the returned Mention records as its source and the only thing the low-confidence filename check reads, so a caller whose prose has no file passes the identifier a reader needs to find it.

  3. Comparator (compare_facts in audit.py) -- Matches mentions against facts and applies configurable tolerances. Version strings require exact match; numeric facts support percentage-based tolerances (e.g., +/-10% for growing metrics like node_count).

  4. Coverage (audit_coverage.py) -- Reports, for every declared fact, whether the run checked anything at all, and the scan surface it ran over. A count of findings describes what the audit FOUND; coverage describes what it COVERED, and the two were measured to differ by a factor of nine (see Coverage reporting).

Fact Types ​

Fact NameSourceDefault ToleranceDescription
versionpyproject.toml / package.json / Cargo.toml0.0 (exact)Project version string
node_countgraph DB nodes table0.10 (+/-10%)Total graph nodes
edge_countgraph DB edges table0.10 (+/-10%)Total graph edges
language_countcode_symbols file extensions0.0 (exact)Distinct programming languages
test_countnodes.extra JSON tests.test_count0.05 (+/-5%)Total test count
nodes_with_frameworknodes.extra JSON tests.framework0.0 (exact)Nodes with detected test frameworks
mcp_tool_countinfrastructure.mcp_tools.MCP_TOOL_CATALOG0.0 (exact)Number of MCP tools. Declared only for Beadloom itself -- see Facts about the running package
cli_command_countClick main group recursive traversal0.0 (exact)Number of CLI commands. Declared only for Beadloom itself -- see Facts about the running package
rule_type_countgraph DB rules table0.0 (exact)Number of architecture rules

Facts about the running package ​

Two of the nine facts are read out of the running Beadloom package, not out of the project being audited: mcp_tool_count comes from MCP_TOOL_CATALOG and cli_command_count from the live Click group. Both are true of Beadloom and false of everybody else, and until BDL-062 .3 both were collected unconditionally -- so in every adopter repository docs audit declared two facts about the tool as facts about their documentation, and counted them in the denominator of "N of 9 verified". Measured: a project named invoice-svc was told it had 18 MCP tools and 43 CLI commands.

doc_sync/audit_self_surface.py gates both. A surface fact is declared only when the project under audit is the distribution whose surfaces this process can introspect, decided from the name the project declares for itself (pyproject.toml, package.json, Cargo.toml) against the running package's name. There is deliberately no directory-name fallback: a project that names itself nowhere is unknown, and unknown must not resolve to a match.

A project with its own MCP server or CLI declares the count itself:

yaml
docs_audit:
  extra_facts:
    mcp_tool_count:
      value: 7
      source: "our MCP server"

An extra_facts value withdraws the decline -- the fact is audited like any other. This is the escape hatch every decline reason names in its own text.

Three populations ​

A fact that could not be computed used to be dropped without trace, so the denominator moved in silence: measured in-process on this repository, an unregistered CLI surface turned 3 of 9 declared fact(s) verified into 3 of 8 and nothing named the fact that left. Every collector now records why it declared nothing, and the audit reports three populations:

PopulationWhere it appearsMeaning
verifiedverified_facts, coverage[*].status == "verified"A document stated the fact and it was judged
not applicable to this projectnot_applicable[name].reasonThe audit declared no value here, and says why. Outside the denominator entirely
declared but unverifiedunverified_factsA value exists and nothing checked it. Named, never counted as fine

version on this repository sits in the first population with two mentions, measured 2026-09-14: docs/getting-started.md states the current release as a claim, and the example in docs/services/cli.md that quotes that claim is read the same way. Every other version literal the audit reads is a dependency pin, a token attributed to another product, or one of five docs_audit.ignore triples with a stated reason. Two of the five are the v2.2.0 example and the 3.0.0 past tense; the other three silence 4.0.0 in documents that record a measurement taken on that release. A triple matches a path, a fact and a value and no line, so it silences that value everywhere in its document -- .beadloom/config.yml states, beside the cli.md triple, which test catches a current claim that regresses under it.

Version attribution: whose version a version is ​

A semantic version is attributed to the nearest subject NAME to its left inside its own clause. Only a version whose nearest name is this project's -- or that has no name at all -- is compared against this project's version; everything else is that product's release, is never compared, and is reported with its subject under attributed_versions.

The rule replaced a suppression per document. _extract_versions used to match every \bv?\d+\.\d+\.\d+\b outside a pin and hand each one to an exact comparison, so "Measured on bd 1.0.4" was a Gate finding; ten docs_audit.ignore triples stood on this repository for that one sentence shape, three of them in user-facing guides (BDL-UX #253, and the foreign-subject face of #190). Eight went inert when the rule landed, measured with a real DocScanner over the audit's 68-document surface.

SentenceNearest nameRead as
Measured on bd 1.0.4bdthe tracker's release
Every verdict on CPython 3.13.7cpythonthe interpreter's release
The current release is 7.0.0nonethis project's version
bd 1.0.4 answers and beadloom 7.0.0 asksbd, then beadloomone each

The vocabulary is derived where a project already declares it and configured where it cannot be: every distribution in pyproject.toml, package.json or Cargo.toml; the interpreter families implied by requires-python, engines.node or rust-version; git when the project is a git repository; and docs_audit.subjects for a name no manifest carries. Each entry records where it came from, and the audit prints them.

A source that cannot be consulted answers neither yes nor no. git is confirmed by the environment rather than by a file the project ships, and the absent .git was read as the assertion that this project has nothing to do with git. A directory built by git archive HEAD -- every clean room beadloom clean-room builds -- carries no .git by construction, so git 2.49.0 lost its subject and was compared against this project's own version. Every clean-room Gate run on this repository was rc 1 for that one line, in docs/domains/application/components/active-table/DOC.md:227, from beadloom-0mdo.63 landing until this repair (BDL-UX #266).

Such a name is UNRESOLVED. It stays in the vocabulary and still wins the attribution walk, so the version beside it is not judged against this project; and the audit reports the token it declined rather than dropping it. The two exempt populations stay separate because they are exempt for different reasons: attributed_versions is a subject this project confirmed, unjudged_versions is a subject this DIRECTORY could not confirm. Merging them would hide a directory that cannot see its own environment behind a rule that works.

SurfaceWhat it carries
docs audit --jsonunjudged_versions, summary.unjudged_version_count, unresolved_version_subjects (name + reason)
docs auditN version token(s) the audit could not judge here: git x1 (no .git here ...)
beadloom ci docs-audit lineCOULD NOT JUDGE N version token(s) naming git — unconfirmed here

A project that names git under docs_audit.subjects has answered the question the marker could not, and the name resolves. The repair was NOT to add git to a declared list: the derivation exists so that no second vocabulary can drift from the first, and the shipped change is to what an absent source MEANS, not to what the vocabulary contains.

It is a vocabulary and not a silencer, and the difference is the failure mode. A name nobody declared still produces a finding, so an unknown subject fails LOUD. The alternative shape -- read any word beside a version as a subject unless it is a function word -- was measured against this repository's own prose and rejected: Phase 3.0.0, Implemented 3.0.0, Release 2.1.0, dated 3.0.0 and published 2.2.0 all put an ordinary English word beside this project's OWN version, so that rule trades a loud false positive for a silent false negative.

Two faces of the same sentence family are outside this rule, and are declared rather than assumed covered:

  • a version MENTIONED rather than used -- the example token v2.2.0 inside the sentence stating what the extractor must not read. No subject stands beside it, because the sentence is about the token itself (BDL-UX #190's example face).
  • this project's OWN past version -- "the alternative shipped in 3.0.0 was worse". There is no foreign subject to find, so no attribution rule reaches it; it needs a notion of past tense (BDL-UX #205).

graph-summary-facts reads node summaries through the same scan_line seam with an EMPTY vocabulary, so every version in a summary stays this project's claim. That rule declares itself pure of the filesystem and the vocabulary is derived from a manifest; the limit is stated here rather than worked around.

False-Positive Filtering ​

Extraction starts from a token boundary rule (Layer 0), then applies a 3-layer false-positive reduction pipeline that reduces FP rate from ~60% to ~11%:

Layer 0: Token Boundary ​

A number that is part of a larger token is an identifier, not a claim, and is never extracted. The bead reference BDL-061.33, the version v2.2.0, the language version Python 3.10, the reference PR #33, the location cli.py:645 and the ratio 33/40 all end in digits that mean nothing on their own. Scanning for digits near a keyword read those tails as facts and failed the Gate twice (BDL-UX #169).

The boundary is whitespace, and only whitespace. Whitespace is the one separator every prose convention agrees on; . - / : # and = are precisely the characters that hold identifiers together, so treating them as boundaries is the defect rather than the fix.

A token's core is the token with wrapping characters removed — markdown emphasis and bracket/quote punctuation at either end, plus sentence punctuation at the end only (33., 33,, 33:). Sentence punctuation is deliberately not stripped from the start: a leading . # or - is exactly what an identifier tail looks like once its prefix is masked (masking BDL-061 leaves .33), and stripping it would restore the bug.

A token becomes a fact candidate only when its whole core is a number — either all digits, or digits in thousands groups (6,390, read whole as 6390). Reading a grouped number by its tail is the audit's worst available outcome: 1,067 nodes used to extract as 067, compare equal to a project count of 67, and stamp a false claim verified.

This single rule subsumes the per-pattern skips it replaced (0xFF, >=0.80, limit=10, 20+, 33%, 0-100, L42, file.py:15) — none of those cores is a number.

Layer 1: Blocklist Modifiers ​

Numbers near modifier words or phrases are skipped as configuration parameters or thresholds, not factual claims. Checked within a +/-3 token window around the number, scoped to the number's own clause (see Clause scope).

Single-word modifiers: default, max, minimum, limit, cap, target, threshold, about, approximately, per, depth, days, hours, minutes, seconds.

Multi-word phrases: up to, at least, at most, no more than, capped at.

Regex-based modifiers: share patterns written with a space (N %). The no-space forms N%, N+ and key=N are not number tokens at all and are rejected by Layer 0.

Layer 2: Proximity Scoring ​

When multiple fact keywords appear near a number, the closest keyword wins. On ties, keywords appearing after the number are preferred (e.g., "63 edges") over those before it. Uses the same PROXIMITY_WINDOW = 5 but with distance-based ranking via _keyword_distance(), scoped to the number's own clause.

Clause scope ​

Both windows above stop at a phrase separator (, ; : and the em/en dash). A word on the far side of one neither modifies the number nor names what it counts, and a flat +/-N window was measured wrong in both directions (BDL-UX #173):

SentenceOld behaviourWith clause scope
The graph holds 316 edges, one per import.nothing extracted — per is in the windowedge_count = 316
exposes 18 tools: 14 over the graphtools bound the 14 too, so a breakdown read as a restatement of the totalonly the 18

Punctuation is a coarse proxy for syntax; it is also the only one available without a parser, and it is what distinguishes what this number counts from what the rest of the sentence talks about. The separator set was chosen by measurement rather than taste:

  • Parentheses are deliberately not separators. They would cost the true verification in MCP tools (18): — an appositive restates its noun — and a lost true verification is precisely the silent false negative this rule exists to remove.
  • Digit groups inside a number belong to the number: the comma in 6,390 is not a boundary between the number and the next word.

MEASURED repo-wide: 0 mentions gained, 5 lost, and all five were confirmed false positives that had needed a docs_audit.ignore entry to stay quiet. Three of those entries were retired with this change, because a suppression that matches nothing reads as coverage it does not have.

Layer 3: File-Type Heuristics ​

Files with lower-confidence names suppress count-type fact matching (versions still matched):

  • Low-confidence filenames: SPEC.md, CONTRIBUTING.md
  • Excluded glob patterns: _graph/features/*/SPEC.md, docs/**/features/*/SPEC.md, docs/**/features/**/SPEC.md

The heuristic is load-bearing and was re-measured before being kept: shadow-scanning the 33 files it hides on this repo yields 17 count/version matches, all of them local examples, historical figures ("edge count went from 51 to 146"), table rows or Python interpreter versions. What changed is that the files are no longer hidden silently -- every one of them is named on the scan surface with the pattern that excluded it (--verbose, or scan_surface in --json).

Pattern Masking ​

The DocScanner also masks the following patterns before number extraction to prevent false matches:

PatternExampleRegex
ISO dates2026-02-19\b\d{4}-\d{2}-\d{2}\b
Month-year datesFeb 2026Month name + 4-digit year
Issue IDs#123, BDL-021#\d+, [A-Z]+-\d+
Hex colors#FF0000#[0-9a-fA-F]{3,8}
Hex literals0xFF0x[0-9a-fA-F]+
Version pins>=0.80, ^1.2.3Operator + version
Line references:15, line 42, L42Various patterns

Numbers 0 and 1 are always skipped as too common and ambiguous.

Several masks above are now subsumed by Layer 0 (#123, 0xFF, >=0.80, :15, L42, 0-100 are not number tokens). They are retained as defence in depth, because masking also governs version extraction and the word positions used for proximity.

Declared blind spots (measured 2026-08-23, resolved 2026-08-24) ​

A false positive announces itself by failing the Gate; a false negative is silent. The three silent ones measured on this repo are all settled -- one by a fix, two by being declared in the output rather than left implicit. Nothing here is silenced by a tolerance or an ignore entry: those hide a true-positive channel to quiet a false one.

Blind spotResolution
A Layer 1 modifier word suppressed a number it did not modify (316 edges, one per import)Fixed -- windows are clause-scoped (see Clause scope)
Counts below MIN_READABLE_COUNT (10) are not extracted for count facts, and 0/1 are never extracted at allKept, and reported. Removing the floor was re-measured: binding a single digit to an immediately following keyword yields 14 extra mentions on this repo, 13 of them ordinals (5. Domain list), table cells (| 5 | tests |) and category breakdowns ((4 tools):) -- several of which would have failed the Gate. The floor stays; the facts it costs are now named unreadable in the coverage report, so nothing reads green about them
SPEC.md / CONTRIBUTING.md suppress count facts, and docs/**/features/*/SPEC.md is excluded outright (33 of 79 markdown files on this repo)Reported. Every skipped file is named on the scan surface with the reason it was skipped

The residual, and it is stated rather than hidden: a doc claim written below the floor (indexes 7 languages) is invisible whether it is right or wrong. The floor is a property of the extractor, so unreadable_reason() states it against the fact rather than leaving the reader to infer it from a zero.

Which facts the floor applies to is scanner.is_count_fact(), not the _count suffix. The suffix names most of them and decides by default; _COUNT_FACTS_WITHOUT_SUFFIX names the ones whose own name says what they count instead. Renaming framework_count to nodes_with_framework (BDL-UX #193) would otherwise have taken that fact out of the floor without anybody deciding it, and test_every_registered_count_fact_is_recognised fails the moment a registered fact other than version falls outside the predicate.

Coverage reporting ​

N mention(s) fresh counts what the audit FOUND. It says nothing about the facts nothing was found for, and on this repo the gap was the whole report: 9 facts declared, 13 verifications, all thirteen of the same fact. A green docs-audit meant "one fact of nine was checked" and printed as a clean bill of health (BDL-UX #173).

Every declared fact therefore carries a FactCoverage:

StatusMeaning
verifiedAt least one mention was compared against the fact. This is about being checked, not about being right -- a stale mention is coverage
not_coveredNo document states the fact. Nothing to check, said out loud
unreadableThe scanner cannot read a claim of this fact at all -- its value is below an extraction floor, or no keywords are registered for it. The fact is structurally unverifiable, and reason says why

Two consequences hold by construction:

  • A mention dropped by a docs_audit.ignore rule is not coverage. A suppression that hides the only mention of a fact leaves that fact unchecked.
  • A fact with zero judged mentions is never counted as passing. It is named -- in the Ground Truth block, in unverified_facts, and on the beadloom ci line.

Coverage does not fail or WARN the gate step, deliberately. Silence in the documentation about a fact is not a defect in the code, and a WARN that every project would carry on every run would spend the channel sync-check needs for a genuinely missing baseline. The number rides on the line everybody reads, and --fail-if unverified>N is there for a project that wants every declared fact stated somewhere.

Keyword-Proximity Matching ​

The DocScanner uses a sliding window of PROXIMITY_WINDOW = 5 word positions around each detected number. If any keyword associated with a fact type appears within this window, the number is classified as a mention of that fact type.

Each fact type has a list of associated keywords:

Fact TypeKeywords
language_countlanguage, lang, programming language
mcp_tool_countMCP, tool, server tool
cli_command_countcommand, CLI, subcommand
rule_type_countrule type, rule kind, rule
node_countnode, module, domain, component
edge_countedge, dependency, connection
test_counttest, spec, assertion
nodes_with_frameworknode with framework, node with a framework, node with a test framework, node declar a framework, node declar a test framework

Keywords use prefix matching (e.g., "language" matches "languages", and "declar" matches "declare", "declares" and "declaring"). A multi-word keyword matches only consecutive words. When two facts sit the same distance from a number, the longer keyword phrase wins: "84 nodes declare a test framework" is one word from both node and node declar a test framework, and the phrase that accounts for more of the sentence is what the sentence is about.

Tolerance System ​

Tolerances control how much a mentioned value may deviate from the ground truth before being flagged as stale:

  • Exact match (tolerance = 0.0): The mentioned integer must equal the ground truth exactly.
  • Percentage tolerance (tolerance > 0.0): The mentioned value must fall within [actual * (1 - t), actual * (1 + t)].
  • Version strings: Always exact string comparison (leading v prefix is stripped).
  • Special case: When the ground truth is 0, only an exact match of 0 is accepted (regardless of tolerance).

Tolerances are merged in order: built-in defaults, then user overrides from config.yml.

Data Structures ​

Fact (frozen dataclass) ​

FieldTypeDescription
namestrFact identifier (e.g., "version", "node_count")
valuestr | intGround-truth value
sourcestrHuman-readable origin (e.g., "pyproject.toml", "graph DB")

Mention (frozen dataclass) ​

FieldTypeDescription
fact_namestrAssociated fact type
valuestr | intMentioned value
filePathSource markdown file
lineintLine number
contextstrStripped line content for display

AuditFinding (frozen dataclass) ​

FieldTypeDescription
mentionMentionThe documentation mention
factFactThe ground-truth fact it was compared against
statusstr"stale" or "fresh"
tolerancefloatApplied tolerance

AuditResult (frozen dataclass) ​

FieldTypeDescription
factsdict[str, Fact]All collected ground-truth facts
findingslist[AuditFinding]Findings for matched mentions
unmatchedlist[Mention]Mentions with no corresponding fact
coveragedict[str, FactCoverage]Per-fact coverage -- what the run checked
surfaceScanSurface | NoneDocuments read and skipped (None when built from mentions directly)
not_applicabledict[str, str]Fact name -> the reason no value was declared for it here

verified_facts and unverified_facts (properties) name the first and third populations, sorted; not_applicable carries the second with its reasons.

FactSet (frozen dataclass) ​

What FactRegistry.collect_set() returns: facts (dict[str, Fact]) and not_applicable (dict[str, str], fact name -> reason). The two are disjoint -- a name is in exactly one. FactRegistry.collect() returns facts alone and cannot tell an absent fact from a declined one, which is why a caller that reports coverage wants collect_set().

FactCoverage (frozen dataclass) ​

FieldTypeDescription
factFactThe declared fact
mentionsintMentions judged against it
statusstrverified / not_covered / unreadable
reasonstr | NoneFor unreadable, the scanner's statement of the limit

ScanSurface (frozen dataclass) ​

FieldTypeDescription
scannedtuple[Path, ...]Files read for mentions
excludedtuple[ExcludedDoc, ...]Files never opened, each with path + reason
count_suppressedtuple[Path, ...]Files read for versions only (subset of scanned)

CLI Interface ​

beadloom docs audit [--json] [--fail-if EXPR] [--stale-only] [--verbose] [--path GLOB] [--project DIR]
OptionTypeDefaultDescription
--jsonflagFalseOutput results as structured JSON
--fail-ifstrNoneCI gate expression (e.g., stale>0, stale>=5)
--stale-onlyflagFalseShow only stale findings
--verboseflagFalseInclude extra detail (unmatched mentions, fact sources)
--pathstr (multiple)NoneOverride default scan paths with custom glob patterns
--projectPathcurrent directoryProject root

The --fail-if expression supports the stale and unverified metrics with > and >= operators. When the condition is met, the command exits with code 1.

  • stale>N -- mentions that disagree with ground truth.
  • unverified>N -- declared facts the run checked nothing for (not_covered + unreadable). Opt-in: a project that wants every fact it declares to be stated somewhere enforces it here.

--verbose additionally names the documents that were not read and those whose counts were suppressed. --json carries coverage, verified_facts, unverified_facts, not_applicable, scan_surface, and a summary with declared_fact_count / verified_fact_count / unverified_count / unreadable_count / not_applicable_count alongside the existing counts. verified_facts, not_applicable and summary.not_applicable_count were added by BDL-062 .3. Nothing was removed, so a consumer parsing the 3.0.0 payload keeps working.

Configuration ​

Tolerance overrides and extra facts are configured in .beadloom/config.yml:

yaml
docs_audit:
  tolerances:
    test_count: 0.10
    node_count: 0.05
  extra_facts:
    custom_metric:
      value: 42
      source: "manual config"
  • tolerances: Per-fact tolerance overrides merged on top of built-in defaults.
  • extra_facts: User-defined facts with a value (str or int) and optional source label. A project declares its own mcp_tool_count or cli_command_count here; the value withdraws the decline described under Facts about the running package.

Debt Report Integration ​

The docs audit contributes to the debt report under the meta_doc_staleness category. Stale findings from the audit increase the architecture debt score.

Default Scan Paths ​

The DocScanner resolves markdown files using these default glob patterns:

  • *.md -- Root-level markdown files
  • docs/**/*.md -- All markdown files under docs/
  • .beadloom/*.md -- Beadloom configuration markdown files

CHANGELOG.md is always excluded. Directories .git, __pycache__, .venv, venv, and node_modules are also excluded.

API ​

Public Functions ​

python
def run_audit(
    project_root: Path,
    db: sqlite3.Connection,
    *,
    scan_paths: list[str] | None = None,
) -> AuditResult

Full audit facade: collect facts, scan docs, compare. Loads tolerance overrides from config if present.

python
def compare_facts(
    facts: dict[str, Fact],
    mentions: list[Mention],
    tolerances: dict[str, float] | None = None,
    ignore: list[IgnoreRule] | None = None,
    not_applicable: dict[str, str] | None = None,
) -> AuditResult

Compare mentions against ground-truth facts with configurable tolerances. not_applicable is carried through to the result unchanged, so the report can name the facts no value was declared for.

python
def foreign_project_reason(project_root: Path) -> str | None

Why Beadloom's own surfaces do not describe project_root -- None when the project under audit is this distribution. declared_project_name(project_root) is the manifest read it rests on, and returns None rather than a directory-name fallback.

python
def parse_fail_condition(expr: str) -> tuple[str, str, int]

Parse a --fail-if expression. Returns (metric, operator, threshold). Raises click.BadParameter on invalid input.

python
def fail_condition_triggered(
    condition: tuple[str, str, int], *, stale_count: int, unverified_count: int
) -> bool

Whether the run crosses the condition's threshold. One place decides what each metric means, so the reported number and the exit code cannot disagree.

python
def assess_coverage(
    facts: dict[str, Fact], findings: list[AuditFinding]
) -> dict[str, FactCoverage]

Per-fact coverage: what the run checked, as opposed to what it found (audit_coverage.py).

python
def unreadable_reason(fact_name: str, value: str | int) -> str | None

The scanner's own statement of why no document could state this fact readably -- or None when one could (scanner.py).

Public Classes ​

python
class FactRegistry:
    def collect_set(self, project_root: Path, db: sqlite3.Connection) -> FactSet: ...
    def collect(self, project_root: Path, db: sqlite3.Connection) -> dict[str, Fact]: ...

class DocScanner:
    def scan(self, paths: list[Path]) -> list[Mention]: ...
    def scan_file(self, file_path: Path) -> list[Mention]: ...
    def scan_line(self, line: str, *, origin: Path, line_number: int = 1) -> list[Mention]: ...
    def resolve_paths(self, project_root: Path, scan_globs: list[str] | None = None) -> list[Path]: ...
    def resolve_surface(self, project_root: Path, scan_globs: list[str] | None = None) -> ScanSurface: ...

@dataclass(frozen=True)
class Fact: ...

@dataclass(frozen=True)
class FactSet: ...

@dataclass(frozen=True)
class Mention: ...

@dataclass(frozen=True)
class AuditFinding: ...

@dataclass(frozen=True)
class AuditResult: ...

@dataclass(frozen=True)
class FactCoverage: ...

@dataclass(frozen=True)
class ScanSurface: ...

@dataclass(frozen=True)
class ExcludedDoc: ...

Invariants ​

  • FactRegistry.collect_set never raises; each data source is wrapped in try/except, and a source that fails records its reason in not_applicable rather than dropping the fact.
  • A fact is in facts or in not_applicable, never in both and never in neither.
  • A fact read from the running Beadloom package is declared only when the project under audit is that distribution. Being correct about Beadloom does not excuse stating it about somebody else.
  • Version extraction uses a priority fallback: pyproject.toml > package.json > Cargo.toml (first match wins).
  • The DocScanner skips code blocks (lines between triple-backtick fences).
  • False-positive masking replaces matched patterns with spaces of equal length to preserve character positions.
  • Each number in a line is matched to at most one fact type (first keyword match wins).
  • A modifier or keyword on the far side of a phrase separator is never matched to the number.
  • A fact with zero judged mentions is never counted as verified, and is named in the output.
  • A mention suppressed by docs_audit.ignore is not coverage.
  • Tolerance merging order: built-in DEFAULT_TOLERANCES < user overrides from config.
  • When ground truth is 0 and tolerance > 0, only an exact mention of 0 is accepted.

Constraints ​

  • Requires a populated SQLite database. Running the audit before beadloom reindex will produce no findings (no facts to collect from DB).
  • Keyword-proximity matching is heuristic; it may produce false positives for numbers near unrelated keywords.
  • The scanner only processes .md files; other documentation formats are not supported.
  • Version detection relies on regex, not full TOML/JSON parsing, which may miss edge cases.
  • The --fail-if expression only supports the stale and unverified metrics with > and >= operators.
  • Coverage answers "was this fact checked", never "is this fact stated correctly everywhere" -- the scanner cannot know about a claim it did not extract, which is why the extraction floors are reported as unreadable rather than inferred from a zero.

Testing ​

Test files: tests/integration/doc_sync/audit/test_docs_audit_cli.py, tests/integration/doc_sync/test_doc_scanner.py, tests/integration/doc_sync/test_doc_scanner_tokenization.py, tests/test_docs_audit_coverage.py, tests/integration/doc_sync/audit/test_audit_ignore.py

Key scenarios:

  • Fact collection: Verify facts are collected from pyproject.toml, graph DB, MCP tools, CLI commands.
  • Version extraction: Verify semantic versions are detected and version pins are ignored.
  • Number extraction: Verify keyword-proximity matching for each fact type.
  • False-positive masking: Verify dates, issue IDs, hex colors, line refs are masked.
  • Tolerance comparison: Verify exact match, percentage tolerance, and zero-value special case.
  • Code block skipping: Verify numbers inside code fences are ignored.
  • Config loading: Verify tolerance overrides and extra facts from config.yml.
  • Full audit pipeline: Verify run_audit end-to-end with stale and fresh findings.
  • Fail condition parsing: Verify valid and invalid --fail-if expressions.
  • CLI integration: Verify beadloom docs audit command options, output formats (JSON/Rich), and CI gate behavior.
  • Clause scope: Verify a modifier in another clause does not suppress a genuine count, that one in the same clause still does, and that a breakdown after a separator is not bound to the total's noun.
  • Coverage: Verify a stated fact reads verified, an unstated one not_covered, one below the extraction floor unreadable, that a stale mention still counts as coverage, and that an ignored mention does not.
  • Scan surface: Verify excluded and count-suppressed documents are named with their reason in both the Rich and JSON output.