This document is the per-variable reference for the three most-tuned JSON files in priv/config/:
scoring.json— heatmap, change_risk, god_modules, unprotected_hubs scoring constantsdispatch_patterns.json— runtime-dispatch patterns the AST walker can't statically seescan_defaults.json— universal source-root list for Mix projects
priv/config/ also holds enrichment_sources.json, dispatch_invariants.json, otp_checks.json and relevance.json, which have no per-variable section here yet — see ARCHITECTURE.md §5 for what each controls.
All are loaded once at daemon startup and cached in :persistent_term for free reads. All except relevance.json are tracked by Knowledge.CodeDigest so edits invalidate the L2 cache automatically; relevance.json filters responses at read time and needs no invalidation.
Each config file has a dedicated loader module that mirrors the same pattern:
| File | Loader |
|---|---|
scoring.json |
Giulia.Knowledge.ScoringConfig |
dispatch_patterns.json |
Giulia.Knowledge.DispatchPatterns |
dispatch_invariants.json |
Giulia.Config.DispatchInvariants |
relevance.json |
Giulia.Config.Relevance |
scan_defaults.json |
Giulia.Context.ScanConfig |
enrichment_sources.json |
Giulia.Enrichment.Registry |
otp_checks.json |
Giulia.Config.OtpChecks |
On first call, the loader reads the JSON, parses with atom keys, and caches the result in :persistent_term. Subsequent reads are free. None of the loaders watch the file for changes — pick up edits with a daemon restart.
Knowledge.CodeDigest hashes the content of every config file in the table above except relevance.json — six files: the three documented here plus enrichment_sources.json, dispatch_invariants.json and otp_checks.json — alongside the BEAM md5s of eleven code-tier modules. The persisted L2 caches (graph + metrics) are tagged with the digest at write time. On daemon startup, warm-restore compares stored vs current digest:
- Match → load cache as-is
- Mismatch → log
"Code digest changed (X -> Y) — invalidating … cache", drop the cache, force a rebuild on next scan
So the operator workflow for tuning any config is:
- Edit the JSON
- Restart the daemon (
docker compose restart giulia-worker) - Trigger a scan or wait for the next warm-restore — caches will rebuild with the new values automatically
The AST cache is not invalidated by config edits — only graph + metric caches. AST extraction is the expensive step (~10s on 580+ files) and shouldn't run on every config tweak. If a config edit needs full re-extraction (rare), use ?force=true on /api/index/scan.
Scoring constants for the four metric families. Defaults are calibrated for typical Elixir application/library shapes and must remain valid for every codebase per the universal-defaults principle (no per-project override file).
The composite score is a weighted sum of four normalized factors. Each factor is normalized to 0–100 by dividing by its cap (saturated at 100), then weighted, then summed and truncated.
score = trunc(
norm_centrality * weights.centrality +
norm_complexity * weights.complexity +
norm_test * weights.test_coverage +
norm_coupling * weights.coupling
)
How each factor contributes to the composite. Should sum to 1.0.
| Field | Default | Meaning |
|---|---|---|
centrality |
0.30 | How much fan-in (incoming module-level edges) drives the score |
complexity |
0.25 | How much AST complexity (control-flow node count) drives the score |
test_coverage |
0.25 | How much missing tests drives the score (penalty when has_test = false) |
coupling |
0.20 | How much max coupling to any single peer drives the score |
Saturation caps. Values above the cap don't increase the factor — every cap value is treated as "this is already worst-case."
| Field | Default | Meaning |
|---|---|---|
centrality_cap |
15 | In-degree at which centrality factor saturates at 100. Increase for very large codebases where 15 dependents isn't unusual |
complexity_cap |
200 | Module-level AST complexity at which the factor saturates. Increase for codebases with intentionally large modules |
coupling_cap |
50 | Max call-count to a single peer at which the factor saturates |
missing_test_factor |
100 | Raw factor value when has_test = false. With the default 0.25 test-coverage weight, this contributes +25 to the score |
Score thresholds for the red/yellow/green classification.
| Field | Default | Meaning |
|---|---|---|
red_min |
60 | Score at or above this is red zone |
yellow_min |
30 | Score at or above this (but below red_min) is yellow zone |
Below yellow_min is green.
Multiplicative composite score. The base captures intrinsic complexity + outward dependency surface; the multiplier amplifies by inward dependency count.
api_penalty = trunc(api_ratio * total_funcs)
base =
(complexity * weights.complexity) +
(fan_out * weights.fan_out) +
(max_coupling * weights.max_coupling) +
api_penalty +
(total_funcs * weights.total_funcs)
multiplier = 1 + (centrality / centrality_divisor)
score = trunc(base * multiplier)
| Field | Default | Meaning |
|---|---|---|
complexity |
2 | Coefficient on AST complexity in base |
fan_out |
2 | Coefficient on outgoing dependency count |
max_coupling |
2 | Coefficient on max calls to any single peer |
total_funcs |
1 | Coefficient on total function count (def + defp) |
| Field | Default | Meaning |
|---|---|---|
centrality_divisor |
2 | Divides centrality (in-degree) when computing multiplier. A value of 2 means each incoming dependent adds 50% to the base score |
| Field | Default | Meaning |
|---|---|---|
top_n |
20 | Number of top-scored modules to return from the endpoint |
Additive score weighted by three factors. No multiplicative term — large breadth alone is enough.
score =
(func_count * weights.func_count) +
(complexity * weights.complexity) +
(centrality * weights.centrality)
| Field | Default | Meaning |
|---|---|---|
func_count |
1 | Coefficient on total function count |
complexity |
2 | Coefficient on AST complexity |
centrality |
3 | Coefficient on in-degree (heaviest weight — a god module with many dependents is the highest-priority refactor target) |
| Field | Default | Meaning |
|---|---|---|
top_n |
20 | Number of top-scored modules to return |
A "hub" is a module with sufficient in-degree to make low spec/doc coverage risky.
| Field | Default | Meaning |
|---|---|---|
default_hub_threshold |
3 | Minimum in-degree to qualify as a hub. Overridable per-call via ?hub_threshold=N |
| Field | Default | Meaning |
|---|---|---|
red_max |
0.5 | Modules with spec_count / public_count below this are red severity |
yellow_max |
0.8 | Modules between red_max and yellow_max are yellow severity. Above is green and excluded from the report |
Runtime-dispatch patterns that AST analysis can't statically see. Consumed by Giulia.Knowledge.DispatchPatterns at startup. Used by dead_code to exempt functions that look unreachable in source but ARE called via runtime mechanisms.
Three types are currently supported.
File-content regex match. For dispatch patterns living outside the AST graph (e.g., shell scripts).
| Field | Meaning |
|---|---|
id |
Stable identifier for logs and reporting |
type |
Always "text_match" |
description |
Free-text rationale |
file_glob |
Glob pattern (relative to project root) of files to scan |
call_regex |
Regex with capture groups; non-matching files are skipped |
arity |
Arity to assume for the matched function (no AST to count from) |
capture |
{module: <group_index>, function: <group_index>} mapping the regex captures |
Example: Mix Release shell overlays.
{
"id": "mix_release_overlays",
"type": "text_match",
"description": "Mix Release overlay shell scripts invoke Elixir functions via `<app> eval Module.function`. Callers live outside the AST graph (shell is not Elixir).",
"file_glob": "rel/overlays/*.sh",
"call_regex": "eval\\s+([A-Z][A-Za-z0-9_.]+)\\.([a-z_][A-Za-z0-9_?!]*)\\s*$",
"arity": 0,
"capture": { "module": 1, "function": 2 }
}Find modules that use one of a list of behaviours and exempt their functions matching a name regex. Universal naming-convention dispatch.
| Field | Meaning |
|---|---|
id |
Stable identifier |
type |
Always "use_based_function_regex" |
description |
Free-text rationale |
behaviours |
List of module-name strings — modules that use any of these qualify |
function_regex |
Regex matched against function names |
arity |
Arity required to match (functions of other arities are not exempted) |
Example: ExMachina factories.
{
"id": "ex_machina_factories",
"type": "use_based_function_regex",
"description": "ExMachina testing factories — modules that `use ExMachina` expose `<name>_factory/0` functions invoked via runtime name-dispatch.",
"behaviours": ["ExMachina", "ExMachina.Ecto"],
"function_regex": "^[a-z_][A-Za-z0-9_]*_factory$",
"arity": 0
}AST-level detection of the defmacro __using__/1 do quote do: apply(__MODULE__, arg, []) end end idiom. Universal across Phoenix-style helper modules — mix phx.new generates this pattern in MyAppWeb.
| Field | Meaning |
|---|---|
id |
Stable identifier |
type |
Always "meta_macro_using_apply" |
description |
Free-text rationale |
enabled |
true/false toggle (use to disable temporarily without removing the entry) |
The mechanism has no other tunables — detection is purely structural.
{
"id": "elixir_meta_macro_using_apply",
"type": "meta_macro_using_apply",
"description": "The Phoenix-style `use MyAppWeb, :shape` idiom: defmacro __using__(arg) do apply(__MODULE__, arg, []) end. Every `use Mod, :shape` caller invokes Mod.shape/0 at compile time but the call is inside the macro's quoted body.",
"enabled": true
}- Edit
dispatch_patterns.json— add a new entry to thepatternsarray. - Restart the daemon. CodeDigest detects the change and invalidates L2 metric caches.
- Trigger a scan (or wait for warm-restore + first metric query).
If the pattern requires a NEW pattern type (not one of the three above), code changes in Knowledge.DispatchPatterns are required. Only the existing three types are runtime-loadable from JSON.
Source-root list for Mix projects. Walked by the indexer to collect Elixir source files for AST extraction.
A list. Entries may be directories (walked recursively for *.ex and *.exs) or individual files.
"source_roots": [
"lib",
"test/support",
"test/test_helper.exs"
]| Default entry | Rationale |
|---|---|
lib |
Universal Mix convention — every project has it |
test/support |
ExUnit convention for shared test helpers; modules here are typically called by every test |
test/test_helper.exs |
Universal ExUnit bootstrap; setup hooks here exercise project code at compile time |
Missing paths are skipped silently — a project without test/support doesn't error.
In addition to source_roots, the indexer parses the target project's mix.exs for def/defp elixirc_paths clauses and unions all string literals across them with the configured source_roots. This catches project-declared non-standard compile dirs (e.g., Plausible's extra/lib) without per-codebase opt-in.
The mechanism is in Giulia.Context.ScanConfig.mix_exs_roots/1 — entirely automatic, no JSON tuning required.
If you find that another path is a universal Elixir project convention (e.g., a future common test layout), add it to source_roots. Per the universal-defaults principle, only add entries that are valid for every Mix project — don't add entries that only apply to a specific codebase.
Allowlist of directories from which POST /api/index/enrichment will accept a payload_path. Anything outside is rejected with HTTP 422 (prevents the endpoint from becoming an arbitrary-file-read primitive).
"enrichment_payload_roots": [
"/tmp",
"/var/tmp",
"tmp",
"_build"
]| Default entry | Resolution | Rationale |
|---|---|---|
/tmp |
Absolute | Universal scratch dir for CI artifacts |
/var/tmp |
Absolute | Persistent scratch on Linux |
tmp |
Project-relative | Phoenix / Mix convention; Credo + Dialyzer often write here |
_build |
Project-relative | Compile output; PLT files for Dialyzer live here |
Project-relative entries are resolved against the caller's project value before the validation runs. Symlink resolution is intentionally NOT performed — the allowlist applies to the raw caller-supplied path, so a malicious symlink in /tmp cannot smuggle in /etc/passwd.
Number of historical builds to retain in ArcadeDB before pruning by Giulia.Storage.Arcade.Consolidator. The Consolidator runs on a 30-min timer plus on-demand via POST /api/index/compact?include=arcade, and deletes CALLS + DEPENDS_ON edges where build_id < (max - retention).
"arcade_history_builds": 10| Aspect | Behavior |
|---|---|
| Default | 10 |
| Minimum | Clamped to 3 even if config sets it lower |
| Why ≥3 | Drift / coupling / hotspot detectors require ≥3 builds of history per module to detect monotonic trends |
Effect on verify_l3 |
Without retention, count_parity.status == "l3_exceeds_l1" accumulates forever as scans repeat. With N=10, count_parity remains match regardless of how many prior builds exist |
| Effect on disk | Each retained build keeps its CALLS edge set in ArcadeDB; on a 1500-edge project with N=10, that's ~15k edges |
Read by Giulia.Context.ScanConfig.arcade_history_builds/0.
| Symptom | Where to look |
|---|---|
| Heatmap reports too many red modules | scoring.json → heatmap.zones.red_min (raise) or heatmap.weights.test_coverage (lower) |
| Change_risk top-10 is dominated by tiny modules | scoring.json → change_risk.weights.complexity or centrality_divisor |
| God_modules list ignores high-fan-in modules | scoring.json → god_modules.weights.centrality (raise) |
| Public function called via runtime mechanism flagged dead | Add a dispatch_patterns.json entry |
Mix-style project lays code in non-lib/ dir |
Usually auto-detected via mix.exs; otherwise add to scan_defaults.json source_roots |
For changes that should affect cached metrics: edit JSON → restart daemon → trigger a scan or wait for the next metric warming.
Thresholds and MFA lists for the OTP deep-analysis checks behind
GET /api/knowledge/otp_risks. Which calls count as blocking is policy that
varies by codebase and by year, so it lives here rather than in .ex source;
only the matching mechanics are compiled in.
Calls that block inside init/1, which runs inside the supervisor's start
sequence — a blocking call serialises boot and turns a down dependency into a
restart-intensity cascade.
| Key | Meaning |
|---|---|
error_mfas |
Error tier: network, DB drivers, cross-process calls, sleeps |
warning_mfas |
Warning tier: File.*, since reading config at boot is legitimate and file size is statically unknowable |
repo_convention |
When true, a call to any module whose final segment is exactly Repo counts as an error-tier DB call |
Pattern forms: "Req.*" matches any function on that module; "Finch.request"
matches exactly; ":httpc.*" and ":gen_tcp.connect" are the Erlang
equivalents. Error tier wins when both would match.
repo_convention exists because enumerating every project's repo module is
impossible. Ecto's naming convention is the only portable signal, and it matches
MyApp.Repo.all and MyApp.Tenant.Repo.one while correctly ignoring
RepoHelper.all.
| Key | Default | Meaning |
|---|---|---|
sync_chain.max_depth |
2 |
Chains deeper than this are flagged; every hop carries its own timeout, so a 3-hop chain is a 15s worst-case budget |
singleton_bottleneck.fan_in_threshold |
8 |
Distinct modules calling a singleton's synchronous API before it is suspected |
singleton_bottleneck.queue_len_threshold |
100 |
Observed message_queue_len at which a static suspicion is escalated to a runtime-confirmed error |
one_for_all.max_children |
5 |
:one_for_all supervisors with more children than this are flagged — one crash restarts them all |
missing_catch_all_handle_info.severity is the only other key and has no
tunable threshold: the check is precise by construction, firing only when
handle_info/2 clauses exist and none is an unguarded catch-all.
Per the universal-defaults principle these must produce sane output on any Elixir codebase with zero tuning — a default that is wrong somewhere is release-gating, not something to push onto the user.