Skip to content

Latest commit

 

History

History
218 lines (178 loc) · 9.6 KB

File metadata and controls

218 lines (178 loc) · 9.6 KB
estado Completed

Analyze

Run all inference detectors over a directory and produce a structured report of schema and content patterns. The report feeds schema apply; it is not a document-repair proposal report.

Usage

rootline analyze [directory]           # defaults to .
rootline analyze docs/ -o json
rootline analyze docs/ --incremental
rootline analyze docs/ --threshold 0.75

Flags

Flag Description
--incremental Report only inferences not covered by existing .stem files
--threshold <0.0-1.0> Section-pattern detection threshold (default 0.60)

Global flags --output json|table and --field <path> also apply.

Detectors

Fourteen detectors run per invocation — twelve data detectors and two governance detectors.

Markdown is parsed into an AST for every record before the detectors run. This is required by the section-pattern, invariant, and formal-dependency detectors; their output is therefore part of the normal command contract rather than an optional parsing mode.

Data: field types, required fields, enum values, constant fields, link types, back references, cross references, section patterns, invariants, formal dependencies, traceability links, structural rules.

Governance: schema coverage (directories without a .stem) and validation gaps. Validation-gap reporting is de-duplicated against data-detector coverage: untyped fields are reported directly, enum-without-values gaps surface when no record value already produced enum-value coverage, sequence gaps mean missing prefix/digits in the .stem declaration rather than skipped observed numbers, and required-understatement is only reachable at the narrow boundary not already covered by required-field inference.

Data Detectors

Detector Description
Field Type Inference Infer the data type of each field from record values
Required Field Detection Identify fields present in >80% of records
Enum Value Detection Extract discrete value sets for enum fields
Constant Field Detection Find fields with only one value across all records
Link Type Validation Validate wiki-links against link schema rules
Back Reference Detection Identify missing reciprocal link declarations
Cross Reference Detection Detect broken document-internal cross-references
Body Section Patterns Identify section headings in markdown bodies
Invariant Extraction Extract INV\d+ identifiers from bodies
Formal Dependency Extraction Extract semantic wiki-link dependencies with the built-in prefixes blocks, relates, and extends
Traceability Link Extraction Identify traceability field claims (Contribuye a, Cubre, Satisface)
Structural Rule Detection Infer directory naming patterns and hierarchy rules

Governance Detectors

Detector Description
Schema Coverage Directories without an owning .stem file
Validation Gaps De-duplicated governance gaps: untyped fields, reachable enum-without-values cases, incomplete sequence declarations (prefix/digits), and the required-understatement boundary case

JSON Output

{
  "version": 1,
  "kind": "rootline/analyze",
  "path": "docs/roadmap/O16-autoupdate-integration",
  "categories": [
    {
      "id": "field_types",
      "name": "Field Type Inference",
      "inference_count": 2,
      "inferences": [
        {
          "type": "field_type",
          "field": "tipo",
          "value": "enum",
          "message": "field \"tipo\" inferred as enum (4/4 records)",
          "requires_agent": false
        },
        {
          "type": "field_type",
          "field": "estado",
          "value": "enum",
          "message": "field \"estado\" inferred as enum (3/4 records)",
          "requires_agent": false
        }
      ]
    },
    {
      "id": "required_fields",
      "name": "Required Field Detection",
      "inference_count": 1,
      "inferences": [
        {
          "type": "required_field",
          "field": "tipo",
          "message": "field \"tipo\" appears in >80% of records — required",
          "requires_agent": false
        }
      ]
    },
    {
      "id": "enum_values",
      "name": "Enum Value Detection",
      "inference_count": 2,
      "inferences": [
        {
          "type": "enum_values",
          "field": "tipo",
          "value": "[outcome task]",
          "message": "field \"tipo\" has enum values: [outcome task]",
          "requires_agent": false
        },
        {
          "type": "enum_values",
          "field": "estado",
          "value": "[Completed]",
          "message": "field \"estado\" has enum values: [Completed]",
          "requires_agent": false
        }
      ]
    }
  ],
  "summary": {
    "total_inferences": 24,
    "agent_required": 3,
    "engine_resolved": 21
  }
}

For identical inputs and flags, analyze -o json emits each category's inferences[] in a deterministic order. Repeated runs are byte-stable, while the category sequence, inference membership, and summary counts remain unchanged.

Output Fields

  • version — contract version.
  • kind — rootline/analyze (distinguishes from other report kinds).
  • path — the scanned directory (as specified on the command line).
  • categories[] — one per detector: id, name, inference_count, inferences[].
    • Each inference carries type, message, and requires_agent; field, value, and source are type-dependent and omitted when not meaningful.
    • requires_agent: true marks inferences needing human or agent disambiguation (not automatically applied by fix/schema commands).
  • summary:
    • total_inferences — count of all inferences across all detectors.
    • agent_required — count of inferences with requires_agent: true.
    • engine_resolved — count of inferences ready for automatic application.

Consuming the Report

Analyze generates schema and diagnostic inferences. Feed supported schema inference types to schema apply to update .stem files. Analyze-derived changes are planned in memory and pass the same prospective hierarchy gate as schema proposal reports before dry-run actions or writes are published. A malformed governing .stem above the report root appears in stem_health with a relative path and blocks both dry-run and write mode before any apply action is published. Document repairs use the versioned rootline/proposals report produced by fix --all --dry-run.

Workflow

# Generate analyze report
rootline analyze docs/ -o json > analyze.json

# Preview schema changes
rootline schema apply --report analyze.json --dry-run

# Apply schema changes (--incremental to skip already-covered inferences)
rootline schema apply --report analyze.json

Inferences with requires_agent: true are logged in the report but skipped by schema apply. Human or agent review is needed to resolve them; they remain actionable for future tooling.

Canonical Section Inferences

Section candidates carry a real type plus a canonical source binding:

notes:
  type: string
  source: body.section["## Notes"]

Each selector component contains one to six # characters, one space, and the exact parsed heading text. The # characters encode the heading level, not the original Markdown form. Thus, an ATX heading ## Notes ## produces body.section["## Notes"]. The selector body.section["## Notes ##"] identifies literal parsed text Notes ##. Setext headings use the same form. A multiline Setext heading uses an escaped \n in the quoted text, such as body.section["## First\nSecond"]. Inference preserves the parsed level and text. It does not preserve whether the source used ATX or Setext syntax.

Every record contributes to the denominator. Rootline calculates the frequency of each exact heading family. Rootline discards each family that does not reach the threshold. A family that reaches the threshold is optional unless the heading occurs in every record. For a hierarchical section, analysis can emit the shortest common selector, such as body.section["## Parent"]["### Notes"]. The components are contiguous, and the selector matches a contiguous suffix of the heading path. The shortest common selector can be simple, so analysis does not always emit a qualified selector. After the threshold filter, Rootline resolves selectors and checks logical-name collisions only among families that reach the threshold. A discarded family does not produce a collision. If families that reach the threshold collide, Rootline fails inference and reports each colliding heading. Rootline requires explicit logical names and does not invent logical names. Analyze, schema proposal, and schema application preserve the canonical source identity.

Filtering with --incremental

By default, analyze reports all inferences. Use --incremental to report only inferences not already covered by existing .stem files:

rootline analyze docs/ --incremental -o json

This is useful in iterative schema refinement: each run shows only the new patterns not yet encoded in .stem files, avoiding re-reporting known patterns.

Threshold Tuning

Section-pattern detection sensitivity is controlled by --threshold (default 0.60, range 0.0-1.0). Higher threshold = fewer pattern proposals:

rootline analyze docs/ --threshold 0.80   # Conservative
rootline analyze docs/ --threshold 0.40   # Aggressive

Structural naming analysis scores directory names and Markdown record-file stems as separate populations. An unrelated directory therefore cannot become an outlier merely because the files beside it follow a record naming pattern.