A configuration-driven Python pipeline that answers one question:
How does City National Bank compare to its peer group, metric by metric, quarter by quarter?
It reads quarterly SNL Excel exports off disk, computes 208 configured metrics across 13 sections for 12 banks (City National Bank plus 11 peers), builds a peer median that excludes CNB itself, renders two charts per metric, and writes a PowerPoint deck and an Excel workbook you can hand to a Treasurer, a CFO, or a board committee without further editing.
Almost everything about it — which banks, which metrics, which formulas, which colors — lives in two YAML files. You should not need to write Python to repoint this at a different peer group or a different analysis.
- What you get
- Data source — read this first
- Requirements and setup
- Quickstart on the bundled sample data
- Running against real SNL data
- CLI flags
- Environment variables
- "Where do I change X?" routing table
- What the pipeline computes
- How it works
- Run characteristics
- Things that will bite you
- Repository layout
- Documentation index
- Tests
One run produces, in the output folder:
| File | What it is |
|---|---|
CNB_Enhanced_Peer_Analysis_<timestamp>_widescreen.pptx |
16:9 deck — 222 slides on the current config: title, peer group, table of contents, a 3-slide executive dashboard, then a divider plus one two-chart slide per metric for each of the 13 sections. |
CNB_Enhanced_Peer_Analysis_<timestamp>_standard.pptx |
Identical content in 4:3, for older projectors and printed handouts. |
CNB_Enhanced_Peer_Analysis_<timestamp>.xlsx |
Four sheets: Summary (wide pivot of every field by bank and quarter), Appendix Deltas (Latest / QoQ Δ / YoY Δ / vs Peer), Charts (both PNGs embedded per metric), Raw Data (the full long-format frame). |
SNL_Data_Stacked.xlsx |
Intermediate — the raw extract, before any math. Useful for debugging a missing field. |
SNL_Data_Transformed.xlsx |
Intermediate — after derived metrics and medians. |
run_manifest_<id>.json |
Audit record: run id, timestamps, runtime per step, row counts, resolved paths, peer count, metric count, git commit, and the list of files created. A new file per run; they are never overwritten. |
logs/pipeline_<YYYY-MM-DD>.log |
Appended log of every run that day. |
Every metric slide carries a line chart (CNB in blue, peer median dashed orange, a shaded peer range
band, optionally a dotted purple GSIB median) and a horizontal ranking bar chart of the latest
quarter with a Rank: #n/N badge — plus a one-line "Key Insight" whose color reflects whether CNB
is on the favorable side of the peer median for that metric.
This pipeline reads SNL (S&P Global Market Intelligence) Excel exports that already exist on disk. It does not call any API.
- The underlying figures are Call Report data as repackaged by SNL, not raw FFIEC filings.
- The pipeline never touches the network (unless you opt into
--ai-narratives, which calls the Anthropic API — see CLI flags). - Input files must be named exactly
SNL DATA VALUES MM.DD.YY.xlsxand must sit directly in one folder. The file-discovery glob is case-sensitive and non-recursive. - The SNL data is licensed third-party content. Do not commit real SNL exports to this or any
other repository. The
samples/folder in this repo contains synthetic, fabricated data with the same structure and clearly fake values.
There is a separate, unrelated FFIEC-based project at ~/Downloads/cnb_peer_analysis_ffiec
which pulls the same kind of analysis directly from the FFIEC. If you need API-sourced Call Report
data rather than SNL exports, start there instead. The stub FFIEC_MIGRATION_SPEC.md in this repo
just points at it.
- Python 3.10 or newer (developed and verified on 3.14).
- macOS or Linux. Windows is untested; paths in the code use
pathlib, so it should work, but the documented commands below are POSIX. - No database, no API key, no network access required for a normal run.
On macOS with Homebrew Python, a bare pip install will fail:
error: externally-managed-environment
× This environment is externally managed
That is Homebrew protecting its Python install, not a problem with this repo. Use a venv:
cd enhanced_snl_peer_analysis # wherever you cloned it
python3 -m venv .venv
.venv/bin/pip install --upgrade pip
.venv/bin/pip install -r requirements.txtVerify:
.venv/bin/python run_pipeline.py --version
# enhanced-snl-pipeline 2.0.0You can source .venv/bin/activate if you prefer, but every command in this README uses the
explicit .venv/bin/python form so it works whether or not you activated anything.
requirements.txt installs pandas, numpy, openpyxl, xlsxwriter, matplotlib, python-pptx, pyyaml,
pydantic, pydantic-settings, scipy, and anthropic, plus pytest for the test suite. The anthropic
package is required even if you never use --ai-narratives, because the test suite imports it.
samples/input/ holds ten synthetic quarterly files — same sheets, same header layout, same
entity list, entirely fabricated values — so you can get a complete, working run without any access
to licensed SNL data:
cd enhanced_snl_peer_analysis # wherever you cloned it
.venv/bin/python run_pipeline.py \
--values-folder samples/input \
--output /tmp/esnl_demoRoughly 30-45 seconds (28.6 s on the machine this was written on), exit 0, ending in
PIPELINE COMPLETED SUCCESSFULLY! and the list of files written. You should see these lines go by:
Selected 10 files from 2023-12-31 through 2026-03-31
Filtered to 12 companies, 30,360 rows
Derived metrics: 72 computed, 0 skipped. Total columns: 325
Computed Peer Median from 11 banks
Transform complete: 40,507 records, 13 entities
/tmp/esnl_demo/ then contains a 227-slide widescreen deck (~18 MB), the matching 4:3 deck, a
~19 MB workbook, the two intermediates and the run manifest. Open the _widescreen.pptx and page
through it — that is exactly what a real run produces, with fake numbers. All 208 metrics render on
the sample data, which makes it a good regression baseline: on real SNL data five metrics drop out
because their source fields are empty.
Send the demo run to a scratch directory like /tmp/esnl_demo, not to samples/output — that
folder holds the committed reference copy, and overwriting it puts a large binary diff in your next
commit.
A faster smoke test that skips chart rendering and both exports — the quickest way to check that a config edit parses and that every bank still matches:
.venv/bin/python run_pipeline.py \
--values-folder samples/input \
--output /tmp/esnl_demo \
--no-charts --no-excel --no-ppt --verboseThe sample entity list deliberately includes the two traps you will meet in the real data — both
BMO Bank National Association and the low-coverage Bank of Montreal - Chicago Branch, and a
Comerica Bank that reports nothing — so you can practise a peer swap safely.
samples/output/ holds a pre-generated reference output, so you can look at the deliverables
without running anything. See samples/README.md for what each sample file is,
how it was generated (tools/make_sample_data.py for the inputs, tools/build_sample_outputs.sh
for the outputs), and how it differs from real SNL data.
Point --values-folder at the folder holding the SNL DATA VALUES MM.DD.YY.xlsx files, and
--output at wherever the deliverables should land:
.venv/bin/python run_pipeline.py \
--values-folder ~/Downloads/VALUES \
--output ~/Downloads/peer_analysis_outputWith the GSIB overlay (adds JPM, BAC, C, GS, MS, WFC as a dotted purple "GSIB Median" line; it does not change the peer median or CNB's rank):
.venv/bin/python run_pipeline.py \
--values-folder ~/Downloads/VALUES \
--output ~/Downloads/peer_analysis_output \
--include-gsibIf you would rather not type the paths every time, export them instead — see Environment variables:
export ESNL_VALUES_FOLDER=~/Downloads/VALUES
export ESNL_OUTPUT_FOLDER=~/Downloads/peer_analysis_output
.venv/bin/python run_pipeline.pyCaveat that catches everyone: the built-in default paths in
src/enhanced_snl_peer_analysis/config/settings.pypoint at~/Downloads/Resolution Planning Research/SNL DATA/VALUES, which does not exist on a fresh machine. A barerun_pipeline.pywith no flags and no env vars dies withFileNotFoundError: VALUES folder not found. Always pass--values-folder, setESNL_VALUES_FOLDER, or edit the default insettings.pyonce for your machine.
The pipeline takes the most recent 13 quarters by filename date (oldest first). If the folder
holds fewer than 13 files it silently uses all of them — check the Selected N files from X through Y line in the log.
Invoke as .venv/bin/python run_pipeline.py [flags]. This is the complete list; there are no other
flags.
| Flag | Default | What it does |
|---|---|---|
--include-gsib |
off | Load the 6 GSIB banks and add a dotted purple GSIB Median series to line and bar charts, plus a GSIB column on the peer-group slide. GSIB banks never affect the peer median or CNB's rank. |
--config PATH, -c PATH |
packaged config/ |
Directory containing metrics.yaml and companies.yaml. Use this to keep an alternate peer group or metric set outside the package. See the warning below. |
--values-folder PATH |
settings.py default |
Folder of SNL DATA VALUES MM.DD.YY.xlsx files. |
--output PATH, -o PATH |
settings.py default |
Where all outputs, intermediates, logs and temp charts are written. Created if missing. |
--no-charts |
off | Skip chart rendering. Also silently suppresses both PowerPoint decks and disables AI narratives, because the deck builder needs the chart objects. The Excel workbook is still written, with an empty Charts sheet. |
--no-ppt |
off | Skip PowerPoint export only. |
--no-excel |
off | Skip Excel export only. |
--ai-narratives |
off | Replace the deterministic "Key Insight" lines in the deck with Claude-written analysis. Requires ANTHROPIC_API_KEY in the environment. Adds roughly 230 sequential API calls — expect many extra minutes of runtime and real API cost. If the key or the SDK is missing, the step logs AI narratives skipped: … and the run completes normally with the deterministic text. Note the Excel workbook always keeps the deterministic wording. |
--verbose, -v |
off | DEBUG-level logging. Needed to see Metric not in data: … lines, which are the only signal that a configured metric was dropped. |
--version |
— | Prints enhanced-snl-pipeline 2.0.0. |
Warning about --config: if the directory you point at is missing companies.yaml, the
pipeline does not error — it falls back to a stale hardcoded peer list baked into
pipeline.py (which still contains Comerica and Webster, both since acquired). You get a
plausible-looking deck built on the wrong peer group. If it is missing metrics.yaml, the run
crashes after extract with a confusing AttributeError. Copy both files together.
Two other invocation forms work identically if you have installed the package
(.venv/bin/pip install -e .):
.venv/bin/enhanced-snl-pipeline --values-folder ~/Downloads/VALUES
.venv/bin/python -m enhanced_snl_peer_analysis --values-folder ~/Downloads/VALUESSettings are read by pydantic-settings at import time, so export them before launching
Python. There is no .env file support.
| Variable | Default | Effect |
|---|---|---|
ESNL_VALUES_FOLDER |
~/Downloads/Resolution Planning Research/SNL DATA/VALUES |
Input folder. Overridden by --values-folder. |
ESNL_OUTPUT_FOLDER |
~/Downloads/Resolution Planning Research/SNL DATA |
Output folder. Overridden by --output. |
ESNL_TEMPLATE_PPT |
none | Path to a branded .pptx/.potx used as the deck template. There is no --template CLI flag; this env var (or editing settings.py) is the only way. The builder looks for a slide layout whose name contains "blank"; if none exists it falls back to the template's last layout (export_ppt.py:1004-1010), so a template without a "Blank" layout works fine and needs no editing. Note that supplying a template also skips the slide_width/slide_height assignment (export_ppt.py:996-1001), so both decks inherit the template's own dimensions — the "standard" output will match the template's aspect ratio, not 10×7.5in. |
ESNL_DATE_LOOKBACK_QUARTERS |
13 |
How many quarterly files to load, most recent first. No CLI flag exists for this. |
ESNL_EXPORT_TIMESTAMP_FILENAMES |
true |
Set false for stable output filenames (overwrites previous run). |
ESNL_EXPORT_CLEANUP_TEMP_FILES |
true |
Set false to keep the chart PNGs in <output>/_temp_charts_enhanced/. This is the only supported way to get the PNGs as standalone files. |
ESNL_VIS_CHART_DPI |
150 |
Chart resolution. 150 is calibrated for PowerPoint; raising it mostly just grows file size. |
ESNL_VIS_CNB_PRIMARY |
#0055A4 |
Subject-bank line/heading color. |
ESNL_VIS_PEER_MEDIAN |
#F57C00 |
Peer median color. |
ESNL_VIS_GSIB_MEDIAN |
#7B1FA2 |
GSIB median color. |
ESNL_VIS_POSITIVE_DELTA / ESNL_VIS_NEGATIVE_DELTA |
#2E7D32 / #C62828 |
Favorable / unfavorable text and count colors. |
ESNL_VIS_BAND_ALPHA |
0.3 |
Opacity of the shaded peer-range band. Raise it if the band is invisible in print. |
ESNL_VIS_* (28 fields total) |
see settings.py |
Every field on VisualSettings is overridable with the ESNL_VIS_ prefix and the field name upper-cased — line widths, marker size, font sizes, grid color, band colors. |
Precedence is CLI flag → environment variable → default in settings.py.
Not every setting is wired up. ESNL_EXPORT_GENERATE_EXCEL and ESNL_EXPORT_GENERATE_POWERPOINT
exist in settings.py but are never read — use --no-excel / --no-ppt. Likewise
ESNL_VIS_SLIDE_WIDTH_INCHES / SLIDE_HEIGHT_INCHES do nothing; slide geometry is hardcoded in
export_ppt.py (13.333×7.5in widescreen, 10×7.5in standard) — except when ESNL_TEMPLATE_PPT is
set, in which case both decks take the template's dimensions instead.
This is the table to bookmark. Almost every change you want to make is a YAML edit.
| I want to… | Edit this | Detail in |
|---|---|---|
| Swap a peer bank in or out | src/enhanced_snl_peer_analysis/config/companies.yaml → peer_banks |
docs/CONFIGURING_PEERS.md |
| Change which banks vote in the peer median | same file → excluded_from_median (uses full names, not abbreviations) |
docs/CONFIGURING_PEERS.md |
| Turn the GSIB overlay on, or change who is in it | --include-gsib; membership in companies.yaml → gsib_banks |
docs/CONFIGURING_PEERS.md |
| Change the subject bank away from CNB | companies.yaml and several hardcoded "CNB" strings in Python — this is the one change that is not config-only |
docs/CONFIGURING_PEERS.md |
| Add / remove / reorder a metric | src/enhanced_snl_peer_analysis/config/metrics.yaml |
docs/CONFIGURING_METRICS.md |
| Add a new derived ratio (numerator ÷ denominator etc.) | metrics.yaml, is_derived: true + one of 7 calculation: types. No Python needed. |
docs/CONFIGURING_METRICS.md |
| Add / rename / reorder a whole section | metrics.yaml top-level sections: list (name, order, description) |
docs/CONFIGURING_METRICS.md |
| Cut the deck down for a shorter audience | a trimmed metrics.yaml in a separate directory, used with --config |
docs/CONFIGURING_METRICS.md |
| Change brand colors, line widths, fonts, DPI | src/enhanced_snl_peer_analysis/config/settings.py → VisualSettings, or ESNL_VIS_* env vars |
Environment variables below |
| Change how many quarters are analyzed | ESNL_DATE_LOOKBACK_QUARTERS, or DateSettings.lookback_quarters in settings.py |
Environment variables below |
| Change the input or output folder | --values-folder / --output, or ESNL_VALUES_FOLDER / ESNL_OUTPUT_FOLDER |
CLI flags above |
| Apply a branded PowerPoint template | ESNL_TEMPLATE_PPT env var (no CLI flag exists); note it also overrides both decks' slide dimensions |
Environment variables below |
| Understand what the input Excel files must look like | — | docs/DATA_CONTRACT.md |
| Produce a conforming input file from S&P Capital IQ Pro | — | docs/DATA_CONTRACT.md |
| Work out why a metric is missing from the deck | — | docs/TROUBLESHOOTING.md |
| Run the quarterly refresh end to end | — | docs/HANDOFF.md |
| Understand the code before changing it | — | docs/ARCHITECTURE.md |
| Change Python and open a PR | — | CONTRIBUTING.md |
Peer banks are matched by exact string equality against column A of the SNL sheet. There is no fuzzy matching, no normalization, no alias table. A typo means the bank is silently missing and the peer median is quietly computed from fewer banks — the run still exits 0.
# src/enhanced_snl_peer_analysis/config/companies.yaml
peer_banks:
- full_name: "BMO Bank National Association" # must match the SNL sheet byte-for-byte
abbreviation: "BMO" # used on charts, tables and the rank badgeTwo real traps from this exact peer group, both worth internalizing:
- The SNL data contains both
BMO Bank National Association(the real U.S. bank, with the great majority of fields populated) andBank of Montreal - Chicago Branch(a foreign branch, under 10% of fields populated). Both strings match cleanly; picking the wrong one gives you near-empty charts rather than an error. Comerica Bankwas still listed as an entity in the 03.31.26 file after being acquired, but every field was blank — 0% coverage. Extract cannot tell "bank dropped out of the data" from "bank has no value for this field."
So: after any peer change, check field coverage, don't just check that the name matched. The repo ships a read-only inspector that answers both questions — the exact name string and how much data that entity actually has:
# every entity whose name contains "montreal", with its field coverage
.venv/bin/python tools/inspect_values_file.py \
"samples/input/SNL DATA VALUES 03.31.26.xlsx" --search montreal
# does a field exist, and which sheet carries it?
.venv/bin/python tools/inspect_values_file.py \
"samples/input/SNL DATA VALUES 03.31.26.xlsx" --field "Total Deposits"It takes any SNL DATA VALUES *.xlsx path, so it works on the synthetic samples and on a real
export. --min-coverage PCT filters to entities with at least that much data.
docs/CONFIGURING_PEERS.md walks the whole swap end to end.
# src/enhanced_snl_peer_analysis/config/metrics.yaml, inside a section's `metrics:` list
- name: Securities / Assets
format: '{:.1f}%'
higher_is_better: null # true, false, or null (null = no favorable/unfavorable coloring)
is_derived: true
calculation: ratio # numerator / denominator * 100
numerator: Total Securities
denominator: Total Assets
definition: Total investment securities as a percentage of total assets.The seven calculation: types are ratio, sum_ratio, sum_over_sum, difference_ratio,
qoq_pct_change, rename, and insured_deposits. Each takes a different set of field keys.
docs/CONFIGURING_METRICS.md has a copy-pasteable template for every one of them.
For a metric that already exists as a raw SNL field, omit is_derived entirely and set name to
the field's exact description text.
13 sections, 208 configured metrics, rendered in order:
| Order | Section | Metrics | Covers |
|---|---|---|---|
| 1 | Treasury Metrics | 21 | Core balance sheet and funding metrics monitored by Treasury |
| 3 | Profitability & Efficiency | 13 | Earnings generation and operational cost management |
| 4 | Capital Adequacy | 11 | Regulatory and tangible capital strength |
| 5 | Asset Quality & Credit Risk | 18 | Loan portfolio health and loss absorption capacity |
| 6 | Loan Composition | 22 | Portfolio allocation across loan types |
| 7 | Securities Portfolio Composition | 17 | Investment portfolio structure and classification |
| 8 | Deposit Composition & Funding | 20 | Deposit base structure and funding stability |
| 9 | Granular Deposit Mix | 18 | Detailed deposit type breakdowns |
| 10 | Liquidity | 16 | Cash reserves, funding buffers, and wholesale reliance |
| 11 | Off-Balance Sheet Exposure | 11 | Contingent commitments and derivatives notional |
| 12 | Interest Rate Risk & Yield/Cost Detail | 24 | Asset yields, funding costs, and rate sensitivity |
| 13 | Noninterest Income Breakdown | 11 | Fee income diversification by revenue source |
| 14 | Growth Trends | 6 | Balance sheet and portfolio growth trajectories |
order: 2 is deliberately unused — the deck inserts its three-slide Executive Dashboard in that
position (overview table, positive trends by section, areas requiring attention). That dashboard is
generated by the deck builder, not by metrics.yaml.
On the current real dataset 203 of the 208 metrics render; five are dropped because their underlying SNL fields are entirely empty for this peer group. That is expected, not a failure — see docs/TROUBLESHOOTING.md.
The peer group is 12 banks. CNB is the subject; the other 11 form the median. These full_name
strings are matched exactly against the SNL data, so they are reproduced here verbatim:
| Abbrev | full_name in companies.yaml |
|---|---|
CNB |
City National Bank — subject bank, excluded from the median |
KEY |
KeyBank National Association |
ZION |
Zions Bancorporation, National Association |
BMO |
BMO Bank National Association |
CFG |
Citizens Bank, National Association |
HBAN |
The Huntington National Bank |
FITB |
Fifth Third Bank, National Association |
FHN |
First Horizon Bank |
EWBC |
East West Bank |
WAL |
Western Alliance Bank |
FCNC.A |
First-Citizens Bank & Trust Company |
MTB |
Manufacturers and Traders Trust Company |
GSIB overlay, loaded only with --include-gsib: JPM, BAC, C, GS, MS, WFC.
Four steps, run in order by pipeline.py. Each is logged with row counts and a runtime, and each is
recorded in the run manifest.
1 · Extract (extract.py) — Finds the most recent N quarterly files, opens each with openpyxl,
reads field names from row 2, keeps only columns I onward, filters rows to the configured banks, and
stacks everything into one long frame of Company Name | Field Description | Date | Value | Source Sheet. A field that appears on more than one sheet is resolved by a fixed sheet-priority list.
$000 fields are divided by 1000 and relabeled ($M). The NA / NM / NULL text placeholders
become NaN — the "… non-numeric values coerced to NaN" warning you will see (a six-figure count
on real data) is normal and expected, not an error.
2 · Transform (transform.py) — Pivots to a wide frame, computes derived metrics from
metrics.yaml via a seven-handler dispatch table, then computes the Peer Median (median of each
peer's own value per quarter, with excluded_from_median banks left out) and, if requested, the
GSIB Median. Both are inserted as synthetic "banks." A second pass builds the delta table
(Latest, QoQ Δ, YoY Δ, vs Peer). The three deltas are plain level differences — for a
percentage-formatted metric they are in percentage points, not percent change. YoY Δ compares to
four quarters back and is blank when a metric has fewer than five quarters of data.
3 · Visualize (visualize.py) — Two matplotlib charts per metric at 150 DPI, 960 × 780 px, saved
to <output>/_temp_charts_enhanced/ and deleted at the end of the run.
4 · Export (export_excel.py, export_ppt.py) — One workbook and two decks, built from the same
in-memory chart objects.
Full call graph, function-by-function behavior, and the exact contracts between steps are in docs/ARCHITECTURE.md.
Measured on 13 quarters × 12 banks against the real SNL data, on an Apple-silicon Mac:
| Wall clock, full run | ~70–85 seconds |
| Extract → Transform | ~365,000 stacked rows → ~261,000 transformed rows |
| Metrics rendered | 203 of 208 configured |
| Slides per deck | 222 |
| Chart PNGs rendered | ~404 (two per metric) |
*_widescreen.pptx |
~19–20 MB |
*_standard.pptx |
~19–20 MB |
*.xlsx |
~25–27 MB |
SNL_Data_Stacked.xlsx |
~11 MB |
SNL_Data_Transformed.xlsx |
~5.6 MB |
run_manifest_*.json |
~50 KB |
--no-charts --no-excel --no-ppt cuts a real run to roughly 45–50 seconds (extract and transform
only); on the sample data it is a few seconds. --ai-narratives adds several minutes.
Short list. The full catalogue with fixes is in docs/TROUBLESHOOTING.md.
- A bare
run_pipeline.pyfails on a fresh machine. The default input path does not exist. Pass--values-folder. - Almost nothing is a hard error. A misspelled bank, a metric with a missing source field, a bad
calculation:type, an unknown YAML key — all of these log a warning (or nothing at all) and the run still exits 0. Exit code 0 does not mean the deck is complete. Read the log, and checkFiltered to N companiesandComputed Peer Median from N banksagainst what you expect. --no-chartssilently kills the PowerPoint too. No warning; the run reports success.- Excel's
Appendix Deltascolors by sign, not by favorability. A rising NPA ratio shows green there, while the deck correctly shows it red. - Two metrics collide on chart filenames (both truncate to
Cost_of_Interest_Bearing_Depos), so one metric's slide shows the other's charts. Known bug, documented in docs/TROUBLESHOOTING.md. - The Table of Contents slide numbers are 3 too high. Known off-by-three; the sections themselves are in the right order.
- If an output file is open in Excel, the write fails with a logged error and the run continues. Close the workbook before rerunning.
enhanced_snl_peer_analysis/
├── README.md ← you are here
├── CONTRIBUTING.md ← workflow for people changing the Python
├── run_pipeline.py ← entry point: .venv/bin/python run_pipeline.py
├── requirements.txt
├── pyproject.toml ← packaging, pytest, ruff, black, mypy config
│
├── docs/ ← handoff documentation (see index below)
│ ├── HANDOFF.md ← start here on day one
│ ├── CONFIGURING_PEERS.md ← change WHO is compared
│ ├── CONFIGURING_METRICS.md ← change WHAT is measured
│ ├── DATA_CONTRACT.md ← what the input files must look like
│ ├── ARCHITECTURE.md ← how the code fits together
│ └── TROUBLESHOOTING.md ← indexed by error message
│
├── samples/
│ ├── README.md ← what the sample data is, and how it differs from real SNL data
│ ├── input/ ← 10 SYNTHETIC quarterly files (safe to publish)
│ └── output/ ← a reference run's deck, workbook and manifest. Every file
│ is stamped SYNTHETIC SAMPLE DATA; none is a deliverable.
│
├── tools/
│ ├── inspect_values_file.py ← read-only workbook inspector. Exact entity name strings,
│ │ per-entity field coverage, and which sheet owns a field.
│ │ Run it before every companies.yaml / metrics.yaml change.
│ ├── make_sample_data.py ← regenerates samples/input/
│ ├── build_sample_outputs.sh ← regenerates samples/output/ from samples/input/
│ └── stamp_synthetic_notice.py ← marks generated artifacts as fabricated, so a deck or
│ workbook forwarded on its own still says so.
│ --check audits a folder and exits 1 if anything is bare.
│
├── src/enhanced_snl_peer_analysis/
│ ├── config/
│ │ ├── companies.yaml ← WHO is compared (peers, GSIBs, median exclusions)
│ │ ├── metrics.yaml ← WHAT is measured (13 sections, 208 metrics)
│ │ └── settings.py ← paths, lookback, colors, fonts, export toggles
│ ├── cli.py ← argparse surface, exit codes
│ ├── pipeline.py ← orchestrator: the 4 steps, rank/QoQ, chart data assembly
│ ├── extract.py ← step 1: read quarterly workbooks → long frame
│ ├── transform.py ← step 2: derived metrics, medians, deltas
│ ├── visualize.py ← step 3: matplotlib line + bar charts
│ ├── export_ppt.py ← step 4a: both PowerPoint decks
│ ├── export_excel.py ← step 4b: the workbook
│ ├── narratives.py ← optional Claude-written insight text
│ ├── manifest.py ← per-run JSON audit record
│ └── utils.py ← shared helpers (logging, safe_divide, date parsing)
│
├── tests/ ← pytest suite (75 tests)
├── templates/ ← drop a branded .pptx here; point ESNL_TEMPLATE_PPT at it
├── .gitignore ← ignores generated .xlsx/.pptx/.png and logs/, with negations
│ that keep samples/ committable
├── FFIEC_MIGRATION_SPEC.md ← stub; points at the separate FFIEC project
└── cnb_peer_analysis_q4_2025.html ← legacy standalone report, not produced by this pipeline
| Document | Read it when |
|---|---|
| docs/HANDOFF.md | Start here on day one. Orientation, your first 30 minutes, your first real run, the quarterly operating procedure, pre-flight and post-run checklists, known limitations, and how this relates to the FFIEC project. |
| docs/CONFIGURING_PEERS.md | You are changing who is compared: peers, the subject bank, median exclusions, the GSIB overlay. Includes ready-made scripts to dump every entity name in a data file and to check a candidate bank's field coverage before you commit to it, plus a full worked example of the Comerica/Webster → BMO/East West swap. |
| docs/CONFIGURING_METRICS.md | You are changing what is measured: metrics, sections, formats, higher_is_better. Complete key reference, all seven calculation handlers with copy-pasteable templates, how to find the exact SNL field name, and why metrics silently disappear. |
| docs/DATA_CONTRACT.md | You need the input file spec: filename rules, sheet inventory, the row/column map, cross-sheet de-duplication, unit conventions, how to produce a conforming export, and a validator script. |
| docs/ARCHITECTURE.md | You are about to change Python, not YAML. Flow, module reference, the data shape at every boundary, the config-driven dispatch, the manifest, and walkthroughs for adding an export format or a chart type. |
| docs/TROUBLESHOOTING.md | Something is missing, wrong, or silently empty. Indexed by the exact error or warning text, plus a description of what a healthy run looks like. |
| CONTRIBUTING.md | You are modifying the code: environment setup, running the tests, type checking, conventions, adding a calculation handler, and the PR checklist. |
| samples/README.md | You want to know what the sample data is, how it was generated, and how it differs from real SNL data. |
The package is not installed into the venv by default, and there is no conftest.py, so the test
suite needs PYTHONPATH=src:
cd enhanced_snl_peer_analysis # wherever you cloned it
PYTHONPATH=src .venv/bin/python -m pytest -qExpected: 74 passed, 1 skipped in about 5 seconds. The skip is the end-to-end test, which needs a real folder of quarterly files and skips itself when it cannot find one.
# skip the slow end-to-end test explicitly
PYTHONPATH=src .venv/bin/python -m pytest -q -m "not slow"
# with coverage (currently ~46% overall; the export modules are thinly covered)
PYTHONPATH=src .venv/bin/python -m pytest --cov=src/enhanced_snl_peer_analysisAlternatively .venv/bin/pip install -e . once, after which plain pytest works.
Two honest caveats: the anthropic package must be installed for the test suite to pass even though
no test makes a network call, and mypy is configured in pyproject.toml but is not installed
by requirements.txt and does not currently pass — do not treat type checking as a gate.
MIT (see pyproject.toml). The code is MIT; the SNL data it consumes is not — that remains
licensed content from S&P Global Market Intelligence and must not be redistributed.