✅ fresh
last synced 2026-09-29T21:12:19.681221+00:00 · coverage 100% (
import-resolver)Validation by Beadloom
doc_sync— same source assync-check.
Import Resolver
Import analysis and depends_on edge generation via tree-sitter-based source code parsing.
Source: src/beadloom/graph/import_resolver.py
Specification
Purpose
Extract import statements from source files using tree-sitter grammars, resolve each import to an architecture graph node, store the results in the code_imports table, and generate depends_on edges between graph nodes. This forms the foundation for automated dependency detection and architectural rule enforcement.
Supported Languages
| Language | File Extensions | Import Syntax Handled | Skipped Imports |
|---|---|---|---|
| Python | .py | import X, from X import Y | Relative imports (from . import, from ..) |
| TypeScript/JavaScript | .ts, .tsx, .js, .jsx | import ... from 'path' | Relative (./, ../), npm packages |
| Go | .go | import "path", import (...) blocks | Standard library (no / in path) |
| Rust | .rs | use path::to::module | Built-in crates (std, core, alloc), self, super |
Constants
_RUST_BUILTIN_CRATES: frozenset[str] = frozenset({"std", "core", "alloc"})
_TS_ALIAS_MAP: dict[str, str] = {
"@/": "src/",
"~/": "src/",
}Data Structures
ImportInfo
Frozen dataclass representing a single extracted import.
| Field | Type | Description |
|---|---|---|
file_path | str | Path to the source file containing the import. |
line_number | int | 1-based line number of the import statement. |
import_path | str | Raw import path (e.g. "beadloom.auth.tokens"). |
resolved_ref_id | str | None | Resolved graph node ref_id, or None. |
Import Extraction
def extract_imports(file_path: Path) -> list[ImportInfo]- Detect language via file extension using
get_lang_config(suffix). Return empty list if unsupported. - Read file content as UTF-8. Return empty list on
OSError,UnicodeDecodeError, or empty content. - Parse content with
tree_sitter.Parserusing the detected language grammar. - Dispatch to language-specific extractor based on extension.
AST traversal
Every extractor iterates _walk(root) — the whole tree in document order — not just the root's children. Extraction used to consider only a file's top-level statements, so an import inside a function, a class body, an if TYPE_CHECKING: guard or a try: block was invisible to the dependency graph. Those are exactly the places an import is put to defer cost or to break a cycle, so the graph was blind to the very edges the cycle and boundary rules exist to judge (BDL-UX #159). Measured on Beadloom when the walk was introduced: 460 nested imports, 231 of them first-party — about a third of all imports; the depends_on edge count went from 51 to 146, and six real boundary violations surfaced that had been hidden behind nested imports.
The statement types the extractors match never nest inside themselves, so a full walk cannot double-count a single statement.
Language-Specific Extractors
Python (_extract_python_imports):
- Walks every
import_statement/import_from_statementin the tree. import_statement: extracts thedotted_namechild as the import path.import_from_statement: checks forrelative_importchild; if present, skips. Otherwise extracts the firstdotted_nameas the module path.
TypeScript/JavaScript (_extract_ts_imports):
- Walks every
import_statementnode in the tree. - Extracts the string source via
_get_ts_import_source(looks forstring->string_fragmentchildren). - Skips imports starting with
"."or".."(relative).
Go (_extract_go_imports):
- Walks
import_declarationnodes. - Handles both single
import_specand groupedimport_spec_list. - Extracts
interpreted_string_literal_contentfrom each spec. - Skips standard library packages (heuristic: no
/in the path).
Rust (_extract_rust_imports):
- Walks
use_declarationnodes. - Extracts the path via
_get_rust_use_path, handlingscoped_identifier,identifier,scoped_use_list, anduse_wildcardnode types. - Determines root crate from the first
::segment. - Skips built-in crates (
std,core,alloc) and relative imports (self,super). - Emits at most one
ImportInfoperuse_declaration.
Import Resolution
def resolve_import_to_node(
import_path: str,
file_path: Path,
conn: sqlite3.Connection,
scan_paths: list[str] | None = None,
*,
is_ts: bool = False,
) -> str | None| Parameter | Type | Default | Description |
|---|---|---|---|
import_path | str | required | Raw import path to resolve. |
file_path | Path | required | Path of the file containing the import. |
conn | sqlite3.Connection | required | Database connection. |
scan_paths | list[str] | None | None (defaults to ["src", "lib", "app"]) | Source directories to search. |
is_ts | bool | False | Whether the import is from a TS/JS file. |
Resolution strategies (tried in order):
Strategy 1 -- Code-symbols annotation lookup:
- Convert the import path to candidate file paths via
_import_path_to_file_paths(replaces.with/, prepends each scan_path prefix, generates both.pyand__init__.pyvariants). - For each candidate, query
code_symbolsforannotationsJSON. - Parse the annotations and look for keys
domain,service, orfeaturewhose values match anodes.ref_id(constructed as"{kind}:{value}"). - Return the first matching
ref_id.
Strategy 2 -- Hierarchical source-prefix matching:
- For TypeScript/JavaScript (
is_ts=True): normalize the import path via_normalize_ts_import. ReturnsNonefor npm packages (non-aliased, non-relative paths), terminating resolution. - For other languages: convert the dotted path to a directory path (replace
.with/). - Call
_find_node_by_source_prefix(dir_path, scan_paths, conn):- Prepend each scan_path prefix (plus bare path).
- Split into path segments, walk from deepest to shallowest.
- For each segment level, query
nodes.sourcewith and without trailing/. - Return the first matching
ref_id.
Internal Resolution Helpers
| Function | Description |
|---|---|
_import_path_to_file_paths | Convert dotted import path to candidate file paths with scan_path prefixes. Generates .py and __init__.py variants. |
_normalize_ts_import | Resolve @/ and ~/ aliases to src/. Returns None for npm packages. |
_find_node_by_source_prefix | Walk path hierarchy from deepest to shallowest, query nodes.source with and without trailing /. |
_find_node_for_file | The node that OWNS a file — delegates to infrastructure/repository.get_owning_ref_id (most specific source wins). Used by create_import_edges. Previously walked up from the file's PARENT directory, so a node whose source IS a file never owned that file and its imports were credited to the enclosing directory's node. |
_walk | Pre-order traversal of the whole AST in document order; every extractor iterates it so imports below the top level are seen. |
_part_of_ancestors | Each node mapped to the set of nodes it is transitively part_of, for the containment skip in create_import_edges. Reads the direct part_of edges and delegates the climb to graph/rules/layers.py::part_of_ancestors, which is the one ancestry walk in the codebase (BDL-070 A1). |
Edge Generation
def create_import_edges(conn: sqlite3.Connection) -> int- Query all distinct
(file_path, resolved_ref_id)fromcode_importswhereresolved_ref_id IS NOT NULL. - For each row, determine the source node via
_find_node_for_file(rel_path, conn)(file ownership). - Skip if no source node is found or if
source_ref_id == target_ref_id(self-reference). - Skip containment in one direction only — an edge from a node to a node it is
part_of(a container depending on its own part). A package façade re-exporting its children says nothing thepart_ofedge did not, and paired with a child's ordinary upward import it manufactures a node-level cycle with no module-level counterpart. The reverse is KEPT: a child reaching into shared code that lives in its container but belongs to no other node is a real dependency, and dropping it made such a node report "depends on nothing" — false, and worse than a coarse answer. - Deduplicate
(source, target)pairs via aseenset. - Insert
depends_onedge withINSERT OR IGNORE, stampingextra = {"derived": "imports"}. The marker is the provenance that lets an incremental refresh delete the derived set without touching a graph-declared edge;INSERT OR IGNOREmeans a pair also declared in YAML keeps the YAML row and stays unmarked. - Commit and return the count of edges created.
Derived-Edge Refresh
def delete_derived_import_edges(conn: sqlite3.Connection) -> int
def refresh_import_edges(conn: sqlite3.Connection) -> intThe derived edge set is a pure function of code_imports, so a refresh is delete-then-recreate. Doing only the recreate half kept a dependency edge alive after its import was removed, for the cycle and layer rules to trip over.
Incremental Indexing
def reindex_file_imports(
project_root: Path,
conn: sqlite3.Connection,
*,
touched: Sequence[str],
removed: Sequence[str],
) -> intDeletes code_imports rows for the touched and removed paths, re-extracts the touched ones, then calls refresh_import_edges. This is what an incremental reindex calls; without it every import rule read an index frozen at the last FULL rebuild, so reindex && lint passed a real boundary break (BDL-UX #142).
Full Indexing Pipeline
def index_imports(project_root: Path, conn: sqlite3.Connection) -> int- Resolve scan paths via
resolve_scan_paths(project_root)from config. - Collect source files via
_collect_source_files(project_root), which usesresolve_scan_pathsandsupported_extensions()to enumerate files under each scan directory. - For each file: a. Call
extract_imports(file_path). Skip if empty. b. Read file content, compute SHA-256 hash, compute relative path. c. Determineis_tsflag from file extension (.ts,.tsx,.js,.jsx,.vue). d. For eachImportInfo, callresolve_import_to_nodeto resolve it. e. Upsert intocode_importswithON CONFLICT(file_path, line_number, import_path) DO UPDATE SET resolved_ref_id, file_hash. - Commit.
- Call
refresh_import_edges(conn)to regeneratedepends_onedges. - Return the total count of imports indexed.
Configuration
Scan paths are configurable via .beadloom/config.yml:
scan_paths:
- src
- lib
- appDefault: ["src", "lib", "app"].
API
Public Functions
def extract_imports(file_path: Path) -> list[ImportInfo]: ...
def resolve_import_to_node(
import_path: str,
file_path: Path,
conn: sqlite3.Connection,
scan_paths: list[str] | None = None,
*,
is_ts: bool = False,
) -> str | None: ...
def create_import_edges(conn: sqlite3.Connection) -> int: ...
def delete_derived_import_edges(conn: sqlite3.Connection) -> int: ...
def refresh_import_edges(conn: sqlite3.Connection) -> int: ...
def index_imports(project_root: Path, conn: sqlite3.Connection) -> int: ...
def reindex_file_imports(
project_root: Path,
conn: sqlite3.Connection,
*,
touched: Sequence[str],
removed: Sequence[str],
) -> int: ...Public Classes
@dataclass(frozen=True)
class ImportInfo:
file_path: str
line_number: int
import_path: str
resolved_ref_id: str | NoneInvariants
- Self-references (source node == target node) never generate
depends_onedges. - Each
(source_ref_id, target_ref_id)pair generates at most onedepends_onedge (deduplicated viaseenset increate_import_edgesandINSERT OR IGNORE). - Imports are upserted with
ON CONFLICT(file_path, line_number, import_path) DO UPDATE, ensuring idempotent reindexing. extract_importsreturns an empty list (never raises) for unsupported languages, unreadable files, or empty files.- Resolution strategies are tried in strict order: annotation lookup first, then source-prefix matching.
_import_path_to_file_pathsalways includes the bare (no-prefix) variant as the last set of candidates.
Constraints
- Requires tree-sitter grammar packages for each supported language (e.g.
tree-sitter-python,tree-sitter-typescript). Returns empty list if the grammar is not installed. - Only processes files located under directories listed in
scan_paths. - Relative imports are always skipped (language-specific detection):
- Python:
relative_importAST node presence. - TypeScript/JavaScript: path starts with
"."or"..". - Go: no
/in path (stdlib heuristic). - Rust: root identifier is
selforsuper.
- Python:
- npm packages (non-aliased, non-relative TypeScript/JavaScript imports) are skipped by
_normalize_ts_importreturningNone. - The
code_symbolstable must be populated for annotation-based resolution to work (Strategy 1). - The
nodestable must be populated for source-prefix resolution to work (Strategy 2). - File content is read as UTF-8; files that raise
UnicodeDecodeErrorare silently skipped.
Testing
Extraction Tests
- Python imports. Parse a file with
import foo,from bar import baz, andfrom . import relative. Assert the first two yieldImportInfoentries; the relative import is skipped. - TypeScript imports. Parse
import X from '@/components/Button'andimport Y from './local'andimport Z from 'react'. Assert only the aliased import is extracted; relative and npm are skipped. - Go imports. Parse
import ("fmt"; "github.com/org/pkg"). Assert only the non-stdlib import is extracted. - Rust imports. Parse
use std::io; use my_crate::module; use super::sibling;. Assert onlymy_crate::moduleis extracted. - Unsupported extension. Pass a
.txtfile. Assert empty list returned. - Empty file. Assert empty list returned.
- Unreadable file. Assert empty list returned without exception.
Resolution Tests
- Annotation lookup hit. Insert a
code_symbolsrow withannotations={"domain": "billing"}for a candidate file path. Insert a node withref_id="domain:billing". Assert resolution returns"domain:billing". - Source-prefix matching. Insert a node with
source="src/beadloom/auth/". Resolve import pathbeadloom.auth.tokenswithscan_paths=["src"]. Assert the correctref_idis returned. - TS alias resolution. Resolve
@/shared/utilswithis_ts=True. Assert it maps tosrc/shared/utilsand matches the appropriate node. - TS npm package skip. Resolve
reactwithis_ts=True. AssertNoneis returned. - No match. Resolve an import path with no corresponding annotation or node. Assert
None.
Edge Generation Tests
- Edges created. Insert two nodes and a resolved code_import. Call
create_import_edges. Assert onedepends_onedge is created. - Self-reference skipped. Import where source and target resolve to the same node. Assert zero edges.
- Deduplication. Multiple imports from the same source to the same target. Assert exactly one edge.
Pipeline Tests
index_importsend-to-end. Set up a project with source files, nodes, and code_symbols. Callindex_imports. Assert:code_importstable is populated with correctfile_path,line_number,import_path,resolved_ref_id.depends_onedges are created in theedgestable.- Return count matches the number of imports processed.
- Idempotent reindex. Call
index_importstwice. Assert the same results with no duplicates (upsert behavior).