✅ fresh
last synced 2026-09-29T21:12:19.681221+00:00 · coverage 100% (
debt-report)Validation by Beadloom
doc_sync— same source assync-check.
Debt Report
Architecture debt aggregation, scoring, and trend tracking.
Source: src/beadloom/application/debt_report/ (package — models, config, collect, scoring, trend, render; __init__.py re-exports the stable public surface). Decomposed by cohesion in BDL-059 S4; import paths unchanged.
Specification
Purpose
The debt report module aggregates all architecture health signals (lint violations, documentation gaps, complexity smells, test coverage gaps) into a single quantified debt score (0-100) with category breakdown, per-node top offenders, and trend tracking against graph snapshots. It provides a unified answer to "how healthy is our architecture?" and supports CI gating via the --fail-if flag on the CLI.
Debt Score Formula
debt_score = min(100, sum(category_scores))
category_scores:
rule_violations = (error_count * rule_error) + (warning_count * rule_warning)
doc_gaps = (undocumented * undocumented_node) + (stale * stale_doc)
+ (untracked * untracked_file)
complexity = (oversized * oversized_domain) + (high_fan_out * high_fan_out)
+ (dormant * dormant_domain)
test_gaps = (untested * untested_domain)Default weights (configurable via config.yml debt_report section):
| Weight | Default | Description |
|---|---|---|
rule_error | 3.0 | Per lint error |
rule_warning | 1.0 | Per lint warning |
undocumented_node | 2.0 | Per node without docs |
stale_doc | 1.0 | Per stale doc-code pair |
untracked_file | 0.5 | Per untracked source file |
oversized_domain | 2.0 | Per oversized domain |
high_fan_out | 1.0 | Per high fan-out node |
dormant_domain | 0.5 | Per dormant domain |
untested_domain | 1.0 | Per untested node (the name predates the test binding) |
Default thresholds:
| Threshold | Default | Description |
|---|---|---|
oversized_symbols | 200 | Symbol count above which a domain is oversized |
high_fan_out | 10 | Edge count above which a node has high fan-out |
dormant_months | 3 | Months without git activity for dormant classification |
Severity Classification
| Score Range | Severity | Indicator |
|---|---|---|
| 0 | clean | check mark (green) |
| 1-10 | low | filled circle (yellow) |
| 11-25 | medium | triangle (yellow) |
| 26-50 | high | diamond (red) |
| 51-100 | critical | X mark (red bold) |
Data Structures
DebtWeights (frozen dataclass)
Per-item weights and thresholds for debt score computation.
| Field | Type | Default | Description |
|---|---|---|---|
rule_error | float | 3.0 | Weight for lint errors |
rule_warning | float | 1.0 | Weight for lint warnings |
undocumented_node | float | 2.0 | Weight for undocumented nodes |
stale_doc | float | 1.0 | Weight for stale docs |
untracked_file | float | 0.5 | Weight for untracked files |
oversized_domain | float | 2.0 | Weight for oversized domains |
high_fan_out | float | 1.0 | Weight for high fan-out nodes |
dormant_domain | float | 0.5 | Weight for dormant domains |
untested_domain | float | 1.0 | Weight per untested node |
oversized_symbols | int | 200 | Oversized threshold |
high_fan_out_threshold | int | 10 | Fan-out threshold |
dormant_months | int | 3 | Dormant threshold (months) |
DebtData (frozen dataclass)
Raw counts aggregated from all data sources.
| Field | Type | Description |
|---|---|---|
error_count | int | Lint rule errors |
warning_count | int | Lint rule warnings |
undocumented_count | int | Nodes without docs |
stale_count | int | Stale sync pairs |
untracked_count | int | Untracked source files |
oversized_count | int | Oversized domains |
high_fan_out_count | int | High fan-out nodes |
dormant_count | int | Dormant domains |
untested_count | int | Nodes the test binding covers with no bound test file; 0 while it is withheld |
node_issues | dict[str, list[str]] | Per-node issue tracking for top offenders |
meta_doc_stale_count | int | Stale fact mentions in project documents |
layer_populations | list[str] | One clause per declared layer rule — how much of its edge set it judged |
test_population | str | What untested_count was counted over, or why it was withheld (BDL-074 C2); empty when not collected |
CategoryScore (frozen dataclass)
| Field | Type | Description |
|---|---|---|
name | str | Category: rule_violations, doc_gaps, complexity, test_gaps |
score | float | Weighted score for this category |
details | dict[str, int | float] | Per-item breakdown (e.g. {"errors": 2, "warnings": 1}) |
NodeDebt (frozen dataclass)
| Field | Type | Description |
|---|---|---|
ref_id | str | Graph node reference ID |
score | float | Debt contribution for this node |
reasons | list[str] | Issue reasons (e.g. ["undocumented", "stale_doc"]) |
DebtTrend (frozen dataclass)
| Field | Type | Description |
|---|---|---|
previous_snapshot | str | Display string for the previous snapshot (ISO date + optional label) |
previous_score | float | Debt score from previous snapshot |
delta | float | Change in overall debt score |
category_deltas | dict[str, float] | Per-category score changes |
DebtReport (frozen dataclass)
| Field | Type | Description |
|---|---|---|
debt_score | float | Overall debt score, 0-100 |
severity | str | Severity label: clean/low/medium/high/critical |
categories | list[CategoryScore] | Four category scores |
top_offenders | list[NodeDebt] | Top 10 nodes ranked by debt contribution |
trend | DebtTrend | None | Trend vs last snapshot, or None |
layer_populations | list[str] | Carried through from DebtData, unweighted |
test_population | str | Carried through from DebtData, unweighted |
What the rule-violation count was counted over (BDL-070 A4)
error_count and warning_count are counts over whatever set the rules could look at. A layer rule looked only at edges whose ends carry a declared layer tag — 16 of 363 on this repository on 2026-09-12, because it read a node's OWN tags — and since BDL-070 B3 an end takes its layer from the nearest part_of container that declares one, so a node carrying no tag at all is judged: 357 of 365 here, measured 2026-09-13. This collector is one of the two surfaces that call evaluate_all without ever building a LintResult, so the population reaches it here or it reaches nobody.
_count_violations reads the reaches with layer_rule_reaches over the rules it has already loaded, from the same reach_of the evaluator uses, rather than parsing the finding's prose back into integers. The clauses travel on DebtData.layer_populations → DebtReport.layer_populations, and appear under Rule Violations in the Rich report as counted over: <rule> judged <evaluated> of <total> live <edge_kind> edge(s), and under layer_populations in format_debt_json — on a project where the collector finds the rules file, which this repository is not (below).
On this repository, measured 2026-09-13, neither the clause nor a violation count appears at all. _count_violations looks for the rules file at <root>/rules.yml and then at <root>/.beadloom/rules.yml, and this project declares its rules in .beadloom/_graph/rules.yml, so the function returns before loading anything: beadloom status --debt-report prints Rule Violations 0 pts, errors: 0 and warnings: 0 while beadloom lint reports 0 error(s) and 55 warning(s) over the same graph, and DebtReport.layer_populations is empty, so the population has nothing to qualify. A zero that means "the rules were never read" is printed in the shape a clean project prints, which is this epic's thesis about a count with no population, one layer up.
The rules file this collector reads is not the one the rest of the product writes (BDL-UX #291). Every other reader, among them lint, reindex, the TUI and the MCP server, resolves <root>/.beadloom/_graph/rules.yml, so a project with the standard layout scores zero rule violations however many it has. It is recorded rather than repaired because the repair moves this repository's raw rule-violations score from 0 to the warning count lint reports, at the default rule_warning weight of 1.0. tests/integration/application/debt_report/test_the_debt_report_states_the_layer_population.py holds the current behaviour, so a repair fails there first.
They are carried UNWEIGHTED. A statement of how much of the graph a count covers is not itself debt, and scoring it would put a number in the score that measures the check rather than the code. Nothing else about the count changed: the population advisory is still counted among the warnings, exactly as BDL-070 A2 left it.
Data Collection Sources
| Category | Source | Module |
|---|---|---|
| Rule violations | evaluate_all(conn, rules) | graph/rules/ (re-exported via the graph.rule_engine shim) |
| Doc gaps -- undocumented | Nodes without docs (LEFT JOIN) | application/debt_report/collect.py |
| Doc gaps -- stale | sync_state entries with status='stale' | application/debt_report/collect.py |
| Doc gaps -- untracked | Nodes with source but no sync_state | application/debt_report/collect.py |
| Complexity -- oversized | Symbol count per node vs threshold | application/debt_report/collect.py |
| Complexity -- fan-out | Edge count per node vs threshold | application/debt_report/collect.py |
| Complexity -- dormant | analyze_git_activity() with dormant level | infrastructure/git_activity.py |
| Test gaps | nodes.extra["tests"] with an empty test_files, from the test binding | application/debt_report/collect.py |
What the untested count was counted over (BDL-074 C2)
_count_untested(conn) returns (count, ref_ids, population). It reads the binding the reindex wrote into each node's extra["tests"] (test mapping): the population is every node that carries that key, and a node whose test_files is empty is untested. The name-guessing mapper it replaced (test_mapper.map_tests, deleted in the same change) counted a node only when its coverage_estimate was none.
While any test file is unplaced — not under a mirrored kind folder or a build tool's test tree, and not inside a node's source — the count is WITHHELD: untested_count is 0 and no node is marked untested. An unplaced file binds to no node, so a node with no bound test may still be tested by one, and counting it would charge a project for its layout rather than its tests. The population then reads not counted: <describe_unplaced sentence>, so a node with no bound test may still be tested, the sentence ctx prints under its Tests: line. Since BDL-074 G2 that sentence is stated against the test layout the reindex recorded (infrastructure.repository.read_test_layout), so it names the project's own folders. Once every test file is placed the count is live and the population reads counted over N node(s) the test binding covers, all M test file(s) placed. Whenever a layout is recorded, withheld and counted alike, the population ENDS with ; and what a test file is read by, from describe_test_file_recognition() (beadloom-2mj3.15; before it, only when the index held no test file). A file outside every root, test tree and node source is not read, so "all M test file(s) placed" means all M files those patterns matched. On this repository, measured at 067df32a on 2026-09-29: not counted: 167 of 623 test file(s) are unplaced (not under tests/integration/ or tests/unit/) and bind to no node, so a node with no bound test may still be tested; a test file is read when its path matches a pattern of pytest (test_*.py, *_test.py) under the root tests. Under the default layout the clause names each group with its patterns and the roots and test trees the project has, since the index records only the roots that exist (beadloom-2mj3.17). With none of them, and tests beside the code read, it ends and it lies beside a node's code, since none of the roots tests, test, spec, __tests__ exists (beadloom-2mj3.19). So a project whose tests match no pattern is charged for every covered node and the report says why. The placement counts come from infrastructure.repository.count_test_files_by_placement.
The binding reads a test file by the project's layout — its roots, its patterns, the build tools' test trees and tests beside the code, each with a default (see the Test Mapping SPEC) — so an adopter whose tests are not tests/**/test_*.py scores what the retired mapper scored. Measured by beadloom-2mj3.11 and beadloom-2mj3.13 on 2026-09-28, against main at db5c3f28: a Go module with a test beside each package scores untested: 0 as on main, where the binding before G2 read 3; a Python project whose tests are mirrored under a declared test/ root, or sit beside the code, scores as on main; a Maven, a Gradle-Kotlin and a SwiftPM project score untested: 0 as on main. beadloom-2mj3.15 added the conventions that were still worse than main — a Jest project with __tests__/ folders, and JVM and SwiftPM test-tree files whose names carry no test affix — each now scoring as on main. A Python project whose tests sit flat under test/, a TypeScript project with flat spec/*.spec.ts files, or a Jest project with a top-level __tests__/ folder (beadloom-2mj3.17), has them read under the default roots tests, test, spec and __tests__ and unplaced, so its count is withheld (0, as on main). tests/integration/application/debt_report/test_an_adopter_scores_what_it_scored_before.py runs those layouts, one project per convention.
Three conventions of the retired mapper are non-goals, each stated in the Test Mapping SPEC. Two of them change the score, because the file they concern is not read or names no framework, so nothing withholds the count:
- NG2: an Xcode
ShopTests/target is outside every root, test tree and node source, so it is not read untiltests.mirrorsdeclares it, and every covered node counts untested where main counted 0. - NG4: a marker file without a test file (
conftest.py,jest.config.*) names no framework, so a project with no test file has every covered node untested where main counted 0.
NG3 does not change the score. A test bound on main only by its imports or by a folder named after a node, such as tests/billing/test_x.py, is read under its root and counted unplaced, because no kind folder, mirror or tests: list places it. The count is withheld: 0, as on main (TestANameAFolderOrAnImportUnderARootIsNotAGuessAtItsNode in tests/integration/application/reindex/test_a_convention_main_reached_by_a_guess_is_stated_not_guessed.py).
The population is carried UNWEIGHTED, like layer_populations: under Test Gaps in the Rich report, and as test_population in format_debt_json.
Top Offenders
compute_top_offenders() ranks individual graph nodes by their weighted debt contribution. Each node's score is computed from its issue list in DebtData.node_issues:
"violation:error:<rule>"-- weighted byrule_error"violation:warning:<rule>"-- weighted byrule_warning- Issue keywords (
undocumented,stale_doc,oversized,high_fan_out,dormant,untested) -- weighted via_ISSUE_WEIGHT_MAP
Nodes are sorted by descending score (ties broken alphabetically by ref_id). Default limit: 10 nodes.
Trend Tracking
compute_debt_trend() compares the current debt report against the most recent graph snapshot. Snapshot data contains structural information (nodes, edges, symbol count) but does not include dynamic data (rules, docs, tests). Therefore:
- The complexity category is recomputed from snapshot edges (high fan-out).
- The rule_violations, doc_gaps, and test_gaps categories are set to 0 for the snapshot (not computable from snapshot data).
Trend output shows per-category directional arrows: improved (down arrow), regressed (up arrow), or unchanged (equals sign).
Config Loading
load_debt_weights() reads the debt_report section from config.yml at the project root. The section has two subsections: weights (per-item multipliers) and thresholds. Missing keys fall back to defaults. Missing file or invalid YAML also falls back to defaults.
Example config.yml:
debt_report:
weights:
rule_error: 3
rule_warning: 1
undocumented_node: 2
stale_doc: 1
thresholds:
oversized_symbols: 200
high_fan_out: 10
dormant_months: 3Category Short Names
The --category CLI flag and MCP category argument accept short names mapped to internal names:
| Short Name | Internal Name |
|---|---|
rules | rule_violations |
docs | doc_gaps |
complexity | complexity |
tests | test_gaps |
CLI Interface
beadloom status --debt-report [--json] [--fail-if=EXPR] [--category=NAME] [--project DIR]--debt-report: Show debt report instead of standard status.--json: Output as structured JSON.--fail-if=EXPR: CI gate. Expressions:score>N,errors>N. Exits with code 1 if condition is met.--category=NAME: Filter to one category.
MCP Interface
Tool: get_debt_report
Arguments:
trend(bool, default false): Include trend vs last snapshot.category(string, optional): Filter to a specific category.
Returns JSON with: debt_score, severity, categories, top_offenders, trend, layer_populations, test_population. With trend the MCP handler attaches the trend by dataclasses.replace, and status --debt-report --category narrows the categories the same way, so neither drops a field the report carries.
Output Formats
Rich (human-readable):
- Header panel: "Architecture Debt Report"
- Score line with severity indicator and label
- Category breakdown with per-item detail lines (tree-style prefixes); the
layer_populationsclauses under Rule Violations andtest_populationunder Test Gaps - Top offenders table (rank, node, score, reasons)
The declared text in those lines — the layer_populations phrases, test_population (which names the test-file patterns, such as Jest's __tests__/**/*.[jt]s) and each offender's ref_id and reasons — is passed through rich.markup.escape before Rich reads it as markup. Unescaped, the Jest pattern printed as __tests__/**/*.s, and a declared pattern holding [/x] raised a MarkupError (beadloom-2mj3.19). The JSON form carries the same text unchanged.
JSON (machine-readable):
debt_score: floatseverity: stringcategories: list of{name, score, details}top_offenders: list of{ref_id, score, reasons}trend: null or{previous_snapshot, previous_score, delta, category_deltas}layer_populations: list of strings (additive, BDL-070 A4)test_population: string (additive, BDL-074 C2)
API
Public Functions
def load_debt_weights(project_root: Path) -> DebtWeightsLoad debt weights from config.yml debt_report section, falling back to defaults.
def collect_debt_data(
conn: sqlite3.Connection,
project_root: Path,
weights: DebtWeights | None = None,
) -> DebtDataAggregate raw counts from all data sources (lint, sync, doctor, git activity, the test binding).
def compute_debt_score(
data: DebtData,
weights: DebtWeights | None = None,
) -> DebtReportApply the weighted formula to produce a complete debt report. Caps score at 100.
def compute_top_offenders(
data: DebtData,
weights: DebtWeights,
limit: int = 10,
) -> list[NodeDebt]Rank nodes by their debt contribution and return the top N.
def compute_debt_trend(
conn: sqlite3.Connection,
current_report: DebtReport,
project_root: Path,
weights: DebtWeights | None = None,
) -> DebtTrend | NoneCompare current debt against the last snapshot. Returns None if no snapshot exists.
def format_debt_report(report: DebtReport) -> strRender a debt report as Rich-formatted terminal output.
def format_trend_section(trend: DebtTrend | None) -> strRender trend data as plain text with directional arrows.
def format_debt_json(
report: DebtReport,
category: str | None = None,
) -> dict[str, Any]Serialize a debt report to a JSON-safe dict with optional category filter.
def format_top_offenders_json(
offenders: list[NodeDebt],
) -> list[dict[str, object]]Serialize a list of NodeDebt to JSON-safe dicts.
Public Classes
@dataclass(frozen=True)
class DebtWeights: ...
@dataclass(frozen=True)
class DebtData: ...
@dataclass(frozen=True)
class CategoryScore: ...
@dataclass(frozen=True)
class NodeDebt: ...
@dataclass(frozen=True)
class DebtTrend: ...
@dataclass(frozen=True)
class DebtReport: ...Invariants
- The debt score is always clamped to the range [0, 100].
- All four categories (rule_violations, doc_gaps, complexity, test_gaps) are always present in
DebtReport.categories, even when their score is 0. compute_debt_scorenever modifies the database. All data collection is read-only.load_debt_weightsalways returns a validDebtWeightsinstance, even with missing or malformed config files.- Top offenders are sorted by descending score with alphabetical ref_id tiebreaking for deterministic output.
- Trend computation is based on structural snapshot data only; categories not stored in snapshots (rules, docs, tests) have 0 as the previous value.
- Each private data collection helper (
_count_undocumented,_count_stale, etc.) gracefully handles import failures or missing tables, returning zero counts.
Constraints
- Requires a populated SQLite database. Running the debt report before
beadloom reindexwill produce a zero score (no data to aggregate). - Trend tracking depends on at least one graph snapshot existing (created by
beadloom snapshot saveor automatically during reindex). - Trend comparisons for non-structural categories (rules, docs, tests) always show zero for the previous snapshot since these are not captured in snapshot data.
- The
--fail-ifCI gate only supports two expressions:score>Nanderrors>N. Other expressions produce an error. - Weight configuration lives in
config.ymlunder thedebt_reportkey. The module readsconfig.ymlfrom the project root, not from.beadloom/.
Testing
Test files: tests/integration/application/debt_report/test_debt_report.py, tests/integration/application/debt_report/test_the_debt_report_reads_the_test_binding.py (the untested count read from the binding, withheld while files are unplaced, and the --category report keeping its population clauses), tests/integration/application/debt_report/test_an_adopter_scores_what_it_scored_before.py (the count on a Go module, three Python layouts and the Maven, Gradle-Kotlin and SwiftPM layouts, and the population of a project with no test file), and tests/acceptance/features/ctx_and_debt_report_read_the_test_binding.feature.
Tests should cover the following scenarios:
- Zero debt: Verify that an empty/healthy graph produces a debt score of 0 with severity
clean. - Severity boundaries: Verify correct severity labels at boundaries (0, 1, 10, 11, 25, 26, 50, 51).
- Category scoring: Verify each category independently contributes the correct weighted score.
- Score capping: Verify that extreme values are capped at 100.
- Config loading: Verify loading weights from
config.yml, including missing file, missing section, and partial overrides. - Top offenders: Verify ranking by score, tiebreaking by ref_id, and limit enforcement.
- Per-node issue tracking: Verify that
violation:error:*,violation:warning:*, and issue keywords are correctly weighted. - Trend computation: Verify delta calculation when a snapshot exists, and
Nonereturn when no snapshot exists. - Format JSON: Verify JSON serialization with and without category filter.
- Format Rich: Verify that Rich output contains expected sections (header, score, categories, offenders).
- Declared text as written: The Rich output carries
__tests__/**/*.[jt]s, a layer phrase and an offender's bracketed reason verbatim, and a pattern holding[/x]does not raise (tests/unit/application/debt_report/test_the_rich_report_prints_declared_text_as_written.py). - Category short names: Verify that short names (
rules,docs,tests) map correctly to internal names.