Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Enhanced SNL Peer Analysis Pipeline

A configuration-driven Python pipeline that answers one question:

How does City National Bank compare to its peer group, metric by metric, quarter by quarter?

It reads quarterly SNL Excel exports off disk, computes 208 configured metrics across 13 sections for 12 banks (City National Bank plus 11 peers), builds a peer median that excludes CNB itself, renders two charts per metric, and writes a PowerPoint deck and an Excel workbook you can hand to a Treasurer, a CFO, or a board committee without further editing.

Almost everything about it — which banks, which metrics, which formulas, which colors — lives in two YAML files. You should not need to write Python to repoint this at a different peer group or a different analysis.


Contents


What you get

One run produces, in the output folder:

File What it is
CNB_Enhanced_Peer_Analysis_<timestamp>_widescreen.pptx 16:9 deck — 222 slides on the current config: title, peer group, table of contents, a 3-slide executive dashboard, then a divider plus one two-chart slide per metric for each of the 13 sections.
CNB_Enhanced_Peer_Analysis_<timestamp>_standard.pptx Identical content in 4:3, for older projectors and printed handouts.
CNB_Enhanced_Peer_Analysis_<timestamp>.xlsx Four sheets: Summary (wide pivot of every field by bank and quarter), Appendix Deltas (Latest / QoQ Δ / YoY Δ / vs Peer), Charts (both PNGs embedded per metric), Raw Data (the full long-format frame).
SNL_Data_Stacked.xlsx Intermediate — the raw extract, before any math. Useful for debugging a missing field.
SNL_Data_Transformed.xlsx Intermediate — after derived metrics and medians.
run_manifest_<id>.json Audit record: run id, timestamps, runtime per step, row counts, resolved paths, peer count, metric count, git commit, and the list of files created. A new file per run; they are never overwritten.
logs/pipeline_<YYYY-MM-DD>.log Appended log of every run that day.

Every metric slide carries a line chart (CNB in blue, peer median dashed orange, a shaded peer range band, optionally a dotted purple GSIB median) and a horizontal ranking bar chart of the latest quarter with a Rank: #n/N badge — plus a one-line "Key Insight" whose color reflects whether CNB is on the favorable side of the peer median for that metric.


Data source — read this first

This pipeline reads SNL (S&P Global Market Intelligence) Excel exports that already exist on disk. It does not call any API.

  • The underlying figures are Call Report data as repackaged by SNL, not raw FFIEC filings.
  • The pipeline never touches the network (unless you opt into --ai-narratives, which calls the Anthropic API — see CLI flags).
  • Input files must be named exactly SNL DATA VALUES MM.DD.YY.xlsx and must sit directly in one folder. The file-discovery glob is case-sensitive and non-recursive.
  • The SNL data is licensed third-party content. Do not commit real SNL exports to this or any other repository. The samples/ folder in this repo contains synthetic, fabricated data with the same structure and clearly fake values.

There is a separate, unrelated FFIEC-based project at ~/Downloads/cnb_peer_analysis_ffiec which pulls the same kind of analysis directly from the FFIEC. If you need API-sourced Call Report data rather than SNL exports, start there instead. The stub FFIEC_MIGRATION_SPEC.md in this repo just points at it.


Requirements and setup

  • Python 3.10 or newer (developed and verified on 3.14).
  • macOS or Linux. Windows is untested; paths in the code use pathlib, so it should work, but the documented commands below are POSIX.
  • No database, no API key, no network access required for a normal run.

Create a virtual environment (do this, don't skip it)

On macOS with Homebrew Python, a bare pip install will fail:

error: externally-managed-environment
× This environment is externally managed

That is Homebrew protecting its Python install, not a problem with this repo. Use a venv:

cd enhanced_snl_peer_analysis      # wherever you cloned it
python3 -m venv .venv
.venv/bin/pip install --upgrade pip
.venv/bin/pip install -r requirements.txt

Verify:

.venv/bin/python run_pipeline.py --version
# enhanced-snl-pipeline 2.0.0

You can source .venv/bin/activate if you prefer, but every command in this README uses the explicit .venv/bin/python form so it works whether or not you activated anything.

requirements.txt installs pandas, numpy, openpyxl, xlsxwriter, matplotlib, python-pptx, pyyaml, pydantic, pydantic-settings, scipy, and anthropic, plus pytest for the test suite. The anthropic package is required even if you never use --ai-narratives, because the test suite imports it.


Quickstart on the bundled sample data

samples/input/ holds ten synthetic quarterly files — same sheets, same header layout, same entity list, entirely fabricated values — so you can get a complete, working run without any access to licensed SNL data:

cd enhanced_snl_peer_analysis      # wherever you cloned it

.venv/bin/python run_pipeline.py \
  --values-folder samples/input \
  --output /tmp/esnl_demo

Roughly 30-45 seconds (28.6 s on the machine this was written on), exit 0, ending in PIPELINE COMPLETED SUCCESSFULLY! and the list of files written. You should see these lines go by:

Selected 10 files from 2023-12-31 through 2026-03-31
Filtered to 12 companies, 30,360 rows
Derived metrics: 72 computed, 0 skipped. Total columns: 325
Computed Peer Median from 11 banks
Transform complete: 40,507 records, 13 entities

/tmp/esnl_demo/ then contains a 227-slide widescreen deck (~18 MB), the matching 4:3 deck, a ~19 MB workbook, the two intermediates and the run manifest. Open the _widescreen.pptx and page through it — that is exactly what a real run produces, with fake numbers. All 208 metrics render on the sample data, which makes it a good regression baseline: on real SNL data five metrics drop out because their source fields are empty.

Send the demo run to a scratch directory like /tmp/esnl_demo, not to samples/output — that folder holds the committed reference copy, and overwriting it puts a large binary diff in your next commit.

A faster smoke test that skips chart rendering and both exports — the quickest way to check that a config edit parses and that every bank still matches:

.venv/bin/python run_pipeline.py \
  --values-folder samples/input \
  --output /tmp/esnl_demo \
  --no-charts --no-excel --no-ppt --verbose

The sample entity list deliberately includes the two traps you will meet in the real data — both BMO Bank National Association and the low-coverage Bank of Montreal - Chicago Branch, and a Comerica Bank that reports nothing — so you can practise a peer swap safely.

samples/output/ holds a pre-generated reference output, so you can look at the deliverables without running anything. See samples/README.md for what each sample file is, how it was generated (tools/make_sample_data.py for the inputs, tools/build_sample_outputs.sh for the outputs), and how it differs from real SNL data.


Running against real SNL data

Point --values-folder at the folder holding the SNL DATA VALUES MM.DD.YY.xlsx files, and --output at wherever the deliverables should land:

.venv/bin/python run_pipeline.py \
  --values-folder ~/Downloads/VALUES \
  --output ~/Downloads/peer_analysis_output

With the GSIB overlay (adds JPM, BAC, C, GS, MS, WFC as a dotted purple "GSIB Median" line; it does not change the peer median or CNB's rank):

.venv/bin/python run_pipeline.py \
  --values-folder ~/Downloads/VALUES \
  --output ~/Downloads/peer_analysis_output \
  --include-gsib

If you would rather not type the paths every time, export them instead — see Environment variables:

export ESNL_VALUES_FOLDER=~/Downloads/VALUES
export ESNL_OUTPUT_FOLDER=~/Downloads/peer_analysis_output
.venv/bin/python run_pipeline.py

Caveat that catches everyone: the built-in default paths in src/enhanced_snl_peer_analysis/config/settings.py point at ~/Downloads/Resolution Planning Research/SNL DATA/VALUES, which does not exist on a fresh machine. A bare run_pipeline.py with no flags and no env vars dies with FileNotFoundError: VALUES folder not found. Always pass --values-folder, set ESNL_VALUES_FOLDER, or edit the default in settings.py once for your machine.

The pipeline takes the most recent 13 quarters by filename date (oldest first). If the folder holds fewer than 13 files it silently uses all of them — check the Selected N files from X through Y line in the log.


CLI flags

Invoke as .venv/bin/python run_pipeline.py [flags]. This is the complete list; there are no other flags.

Flag Default What it does
--include-gsib off Load the 6 GSIB banks and add a dotted purple GSIB Median series to line and bar charts, plus a GSIB column on the peer-group slide. GSIB banks never affect the peer median or CNB's rank.
--config PATH, -c PATH packaged config/ Directory containing metrics.yaml and companies.yaml. Use this to keep an alternate peer group or metric set outside the package. See the warning below.
--values-folder PATH settings.py default Folder of SNL DATA VALUES MM.DD.YY.xlsx files.
--output PATH, -o PATH settings.py default Where all outputs, intermediates, logs and temp charts are written. Created if missing.
--no-charts off Skip chart rendering. Also silently suppresses both PowerPoint decks and disables AI narratives, because the deck builder needs the chart objects. The Excel workbook is still written, with an empty Charts sheet.
--no-ppt off Skip PowerPoint export only.
--no-excel off Skip Excel export only.
--ai-narratives off Replace the deterministic "Key Insight" lines in the deck with Claude-written analysis. Requires ANTHROPIC_API_KEY in the environment. Adds roughly 230 sequential API calls — expect many extra minutes of runtime and real API cost. If the key or the SDK is missing, the step logs AI narratives skipped: … and the run completes normally with the deterministic text. Note the Excel workbook always keeps the deterministic wording.
--verbose, -v off DEBUG-level logging. Needed to see Metric not in data: … lines, which are the only signal that a configured metric was dropped.
--version — Prints enhanced-snl-pipeline 2.0.0.

Warning about --config: if the directory you point at is missing companies.yaml, the pipeline does not error — it falls back to a stale hardcoded peer list baked into pipeline.py (which still contains Comerica and Webster, both since acquired). You get a plausible-looking deck built on the wrong peer group. If it is missing metrics.yaml, the run crashes after extract with a confusing AttributeError. Copy both files together.

Two other invocation forms work identically if you have installed the package (.venv/bin/pip install -e .):

.venv/bin/enhanced-snl-pipeline --values-folder ~/Downloads/VALUES
.venv/bin/python -m enhanced_snl_peer_analysis --values-folder ~/Downloads/VALUES

Environment variables

Settings are read by pydantic-settings at import time, so export them before launching Python. There is no .env file support.

Variable Default Effect
ESNL_VALUES_FOLDER ~/Downloads/Resolution Planning Research/SNL DATA/VALUES Input folder. Overridden by --values-folder.
ESNL_OUTPUT_FOLDER ~/Downloads/Resolution Planning Research/SNL DATA Output folder. Overridden by --output.
ESNL_TEMPLATE_PPT none Path to a branded .pptx/.potx used as the deck template. There is no --template CLI flag; this env var (or editing settings.py) is the only way. The builder looks for a slide layout whose name contains "blank"; if none exists it falls back to the template's last layout (export_ppt.py:1004-1010), so a template without a "Blank" layout works fine and needs no editing. Note that supplying a template also skips the slide_width/slide_height assignment (export_ppt.py:996-1001), so both decks inherit the template's own dimensions — the "standard" output will match the template's aspect ratio, not 10×7.5in.
ESNL_DATE_LOOKBACK_QUARTERS 13 How many quarterly files to load, most recent first. No CLI flag exists for this.
ESNL_EXPORT_TIMESTAMP_FILENAMES true Set false for stable output filenames (overwrites previous run).
ESNL_EXPORT_CLEANUP_TEMP_FILES true Set false to keep the chart PNGs in <output>/_temp_charts_enhanced/. This is the only supported way to get the PNGs as standalone files.
ESNL_VIS_CHART_DPI 150 Chart resolution. 150 is calibrated for PowerPoint; raising it mostly just grows file size.
ESNL_VIS_CNB_PRIMARY #0055A4 Subject-bank line/heading color.
ESNL_VIS_PEER_MEDIAN #F57C00 Peer median color.
ESNL_VIS_GSIB_MEDIAN #7B1FA2 GSIB median color.
ESNL_VIS_POSITIVE_DELTA / ESNL_VIS_NEGATIVE_DELTA #2E7D32 / #C62828 Favorable / unfavorable text and count colors.
ESNL_VIS_BAND_ALPHA 0.3 Opacity of the shaded peer-range band. Raise it if the band is invisible in print.
ESNL_VIS_* (28 fields total) see settings.py Every field on VisualSettings is overridable with the ESNL_VIS_ prefix and the field name upper-cased — line widths, marker size, font sizes, grid color, band colors.

Precedence is CLI flag → environment variable → default in settings.py.

Not every setting is wired up. ESNL_EXPORT_GENERATE_EXCEL and ESNL_EXPORT_GENERATE_POWERPOINT exist in settings.py but are never read — use --no-excel / --no-ppt. Likewise ESNL_VIS_SLIDE_WIDTH_INCHES / SLIDE_HEIGHT_INCHES do nothing; slide geometry is hardcoded in export_ppt.py (13.333×7.5in widescreen, 10×7.5in standard) — except when ESNL_TEMPLATE_PPT is set, in which case both decks take the template's dimensions instead.


"Where do I change X?" routing table

This is the table to bookmark. Almost every change you want to make is a YAML edit.

I want to… Edit this Detail in
Swap a peer bank in or out src/enhanced_snl_peer_analysis/config/companies.yaml → peer_banks docs/CONFIGURING_PEERS.md
Change which banks vote in the peer median same file → excluded_from_median (uses full names, not abbreviations) docs/CONFIGURING_PEERS.md
Turn the GSIB overlay on, or change who is in it --include-gsib; membership in companies.yaml → gsib_banks docs/CONFIGURING_PEERS.md
Change the subject bank away from CNB companies.yaml and several hardcoded "CNB" strings in Python — this is the one change that is not config-only docs/CONFIGURING_PEERS.md
Add / remove / reorder a metric src/enhanced_snl_peer_analysis/config/metrics.yaml docs/CONFIGURING_METRICS.md
Add a new derived ratio (numerator ÷ denominator etc.) metrics.yaml, is_derived: true + one of 7 calculation: types. No Python needed. docs/CONFIGURING_METRICS.md
Add / rename / reorder a whole section metrics.yaml top-level sections: list (name, order, description) docs/CONFIGURING_METRICS.md
Cut the deck down for a shorter audience a trimmed metrics.yaml in a separate directory, used with --config docs/CONFIGURING_METRICS.md
Change brand colors, line widths, fonts, DPI src/enhanced_snl_peer_analysis/config/settings.py → VisualSettings, or ESNL_VIS_* env vars Environment variables below
Change how many quarters are analyzed ESNL_DATE_LOOKBACK_QUARTERS, or DateSettings.lookback_quarters in settings.py Environment variables below
Change the input or output folder --values-folder / --output, or ESNL_VALUES_FOLDER / ESNL_OUTPUT_FOLDER CLI flags above
Apply a branded PowerPoint template ESNL_TEMPLATE_PPT env var (no CLI flag exists); note it also overrides both decks' slide dimensions Environment variables below
Understand what the input Excel files must look like — docs/DATA_CONTRACT.md
Produce a conforming input file from S&P Capital IQ Pro — docs/DATA_CONTRACT.md
Work out why a metric is missing from the deck — docs/TROUBLESHOOTING.md
Run the quarterly refresh end to end — docs/HANDOFF.md
Understand the code before changing it — docs/ARCHITECTURE.md
Change Python and open a PR — CONTRIBUTING.md

The one-minute version: swapping a peer

Peer banks are matched by exact string equality against column A of the SNL sheet. There is no fuzzy matching, no normalization, no alias table. A typo means the bank is silently missing and the peer median is quietly computed from fewer banks — the run still exits 0.

# src/enhanced_snl_peer_analysis/config/companies.yaml
peer_banks:
  - full_name: "BMO Bank National Association"   # must match the SNL sheet byte-for-byte
    abbreviation: "BMO"                          # used on charts, tables and the rank badge

Two real traps from this exact peer group, both worth internalizing:

  • The SNL data contains both BMO Bank National Association (the real U.S. bank, with the great majority of fields populated) and Bank of Montreal - Chicago Branch (a foreign branch, under 10% of fields populated). Both strings match cleanly; picking the wrong one gives you near-empty charts rather than an error.
  • Comerica Bank was still listed as an entity in the 03.31.26 file after being acquired, but every field was blank — 0% coverage. Extract cannot tell "bank dropped out of the data" from "bank has no value for this field."

So: after any peer change, check field coverage, don't just check that the name matched. The repo ships a read-only inspector that answers both questions — the exact name string and how much data that entity actually has:

# every entity whose name contains "montreal", with its field coverage
.venv/bin/python tools/inspect_values_file.py \
  "samples/input/SNL DATA VALUES 03.31.26.xlsx" --search montreal

# does a field exist, and which sheet carries it?
.venv/bin/python tools/inspect_values_file.py \
  "samples/input/SNL DATA VALUES 03.31.26.xlsx" --field "Total Deposits"

It takes any SNL DATA VALUES *.xlsx path, so it works on the synthetic samples and on a real export. --min-coverage PCT filters to entities with at least that much data. docs/CONFIGURING_PEERS.md walks the whole swap end to end.

The one-minute version: adding a metric

# src/enhanced_snl_peer_analysis/config/metrics.yaml, inside a section's `metrics:` list
  - name: Securities / Assets
    format: '{:.1f}%'
    higher_is_better: null        # true, false, or null (null = no favorable/unfavorable coloring)
    is_derived: true
    calculation: ratio            # numerator / denominator * 100
    numerator: Total Securities
    denominator: Total Assets
    definition: Total investment securities as a percentage of total assets.

The seven calculation: types are ratio, sum_ratio, sum_over_sum, difference_ratio, qoq_pct_change, rename, and insured_deposits. Each takes a different set of field keys. docs/CONFIGURING_METRICS.md has a copy-pasteable template for every one of them.

For a metric that already exists as a raw SNL field, omit is_derived entirely and set name to the field's exact description text.


What the pipeline computes

13 sections, 208 configured metrics, rendered in order:

Order Section Metrics Covers
1 Treasury Metrics 21 Core balance sheet and funding metrics monitored by Treasury
3 Profitability & Efficiency 13 Earnings generation and operational cost management
4 Capital Adequacy 11 Regulatory and tangible capital strength
5 Asset Quality & Credit Risk 18 Loan portfolio health and loss absorption capacity
6 Loan Composition 22 Portfolio allocation across loan types
7 Securities Portfolio Composition 17 Investment portfolio structure and classification
8 Deposit Composition & Funding 20 Deposit base structure and funding stability
9 Granular Deposit Mix 18 Detailed deposit type breakdowns
10 Liquidity 16 Cash reserves, funding buffers, and wholesale reliance
11 Off-Balance Sheet Exposure 11 Contingent commitments and derivatives notional
12 Interest Rate Risk & Yield/Cost Detail 24 Asset yields, funding costs, and rate sensitivity
13 Noninterest Income Breakdown 11 Fee income diversification by revenue source
14 Growth Trends 6 Balance sheet and portfolio growth trajectories

order: 2 is deliberately unused — the deck inserts its three-slide Executive Dashboard in that position (overview table, positive trends by section, areas requiring attention). That dashboard is generated by the deck builder, not by metrics.yaml.

On the current real dataset 203 of the 208 metrics render; five are dropped because their underlying SNL fields are entirely empty for this peer group. That is expected, not a failure — see docs/TROUBLESHOOTING.md.

The peer group is 12 banks. CNB is the subject; the other 11 form the median. These full_name strings are matched exactly against the SNL data, so they are reproduced here verbatim:

Abbrev full_name in companies.yaml
CNB City National Bank — subject bank, excluded from the median
KEY KeyBank National Association
ZION Zions Bancorporation, National Association
BMO BMO Bank National Association
CFG Citizens Bank, National Association
HBAN The Huntington National Bank
FITB Fifth Third Bank, National Association
FHN First Horizon Bank
EWBC East West Bank
WAL Western Alliance Bank
FCNC.A First-Citizens Bank & Trust Company
MTB Manufacturers and Traders Trust Company

GSIB overlay, loaded only with --include-gsib: JPM, BAC, C, GS, MS, WFC.


How it works

Four steps, run in order by pipeline.py. Each is logged with row counts and a runtime, and each is recorded in the run manifest.

1 · Extract (extract.py) — Finds the most recent N quarterly files, opens each with openpyxl, reads field names from row 2, keeps only columns I onward, filters rows to the configured banks, and stacks everything into one long frame of Company Name | Field Description | Date | Value | Source Sheet. A field that appears on more than one sheet is resolved by a fixed sheet-priority list. $000 fields are divided by 1000 and relabeled ($M). The NA / NM / NULL text placeholders become NaN — the "… non-numeric values coerced to NaN" warning you will see (a six-figure count on real data) is normal and expected, not an error.

2 · Transform (transform.py) — Pivots to a wide frame, computes derived metrics from metrics.yaml via a seven-handler dispatch table, then computes the Peer Median (median of each peer's own value per quarter, with excluded_from_median banks left out) and, if requested, the GSIB Median. Both are inserted as synthetic "banks." A second pass builds the delta table (Latest, QoQ Δ, YoY Δ, vs Peer). The three deltas are plain level differences — for a percentage-formatted metric they are in percentage points, not percent change. YoY Δ compares to four quarters back and is blank when a metric has fewer than five quarters of data.

3 · Visualize (visualize.py) — Two matplotlib charts per metric at 150 DPI, 960 × 780 px, saved to <output>/_temp_charts_enhanced/ and deleted at the end of the run.

4 · Export (export_excel.py, export_ppt.py) — One workbook and two decks, built from the same in-memory chart objects.

Full call graph, function-by-function behavior, and the exact contracts between steps are in docs/ARCHITECTURE.md.


Run characteristics

Measured on 13 quarters × 12 banks against the real SNL data, on an Apple-silicon Mac:

Wall clock, full run ~70–85 seconds
Extract → Transform ~365,000 stacked rows → ~261,000 transformed rows
Metrics rendered 203 of 208 configured
Slides per deck 222
Chart PNGs rendered ~404 (two per metric)
*_widescreen.pptx ~19–20 MB
*_standard.pptx ~19–20 MB
*.xlsx ~25–27 MB
SNL_Data_Stacked.xlsx ~11 MB
SNL_Data_Transformed.xlsx ~5.6 MB
run_manifest_*.json ~50 KB

--no-charts --no-excel --no-ppt cuts a real run to roughly 45–50 seconds (extract and transform only); on the sample data it is a few seconds. --ai-narratives adds several minutes.


Things that will bite you

Short list. The full catalogue with fixes is in docs/TROUBLESHOOTING.md.

  • A bare run_pipeline.py fails on a fresh machine. The default input path does not exist. Pass --values-folder.
  • Almost nothing is a hard error. A misspelled bank, a metric with a missing source field, a bad calculation: type, an unknown YAML key — all of these log a warning (or nothing at all) and the run still exits 0. Exit code 0 does not mean the deck is complete. Read the log, and check Filtered to N companies and Computed Peer Median from N banks against what you expect.
  • --no-charts silently kills the PowerPoint too. No warning; the run reports success.
  • Excel's Appendix Deltas colors by sign, not by favorability. A rising NPA ratio shows green there, while the deck correctly shows it red.
  • Two metrics collide on chart filenames (both truncate to Cost_of_Interest_Bearing_Depos), so one metric's slide shows the other's charts. Known bug, documented in docs/TROUBLESHOOTING.md.
  • The Table of Contents slide numbers are 3 too high. Known off-by-three; the sections themselves are in the right order.
  • If an output file is open in Excel, the write fails with a logged error and the run continues. Close the workbook before rerunning.

Repository layout

enhanced_snl_peer_analysis/
├── README.md                     ← you are here
├── CONTRIBUTING.md               ← workflow for people changing the Python
├── run_pipeline.py               ← entry point: .venv/bin/python run_pipeline.py
├── requirements.txt
├── pyproject.toml                ← packaging, pytest, ruff, black, mypy config
│
├── docs/                         ← handoff documentation (see index below)
│   ├── HANDOFF.md                ← start here on day one
│   ├── CONFIGURING_PEERS.md      ← change WHO is compared
│   ├── CONFIGURING_METRICS.md    ← change WHAT is measured
│   ├── DATA_CONTRACT.md          ← what the input files must look like
│   ├── ARCHITECTURE.md           ← how the code fits together
│   └── TROUBLESHOOTING.md        ← indexed by error message
│
├── samples/
│   ├── README.md                 ← what the sample data is, and how it differs from real SNL data
│   ├── input/                    ← 10 SYNTHETIC quarterly files (safe to publish)
│   └── output/                   ← a reference run's deck, workbook and manifest. Every file
│                                   is stamped SYNTHETIC SAMPLE DATA; none is a deliverable.
│
├── tools/
│   ├── inspect_values_file.py    ← read-only workbook inspector. Exact entity name strings,
│   │                               per-entity field coverage, and which sheet owns a field.
│   │                               Run it before every companies.yaml / metrics.yaml change.
│   ├── make_sample_data.py       ← regenerates samples/input/
│   ├── build_sample_outputs.sh   ← regenerates samples/output/ from samples/input/
│   └── stamp_synthetic_notice.py ← marks generated artifacts as fabricated, so a deck or
│                                   workbook forwarded on its own still says so.
│                                   --check audits a folder and exits 1 if anything is bare.
│
├── src/enhanced_snl_peer_analysis/
│   ├── config/
│   │   ├── companies.yaml        ← WHO is compared (peers, GSIBs, median exclusions)
│   │   ├── metrics.yaml          ← WHAT is measured (13 sections, 208 metrics)
│   │   └── settings.py           ← paths, lookback, colors, fonts, export toggles
│   ├── cli.py                    ← argparse surface, exit codes
│   ├── pipeline.py               ← orchestrator: the 4 steps, rank/QoQ, chart data assembly
│   ├── extract.py                ← step 1: read quarterly workbooks → long frame
│   ├── transform.py              ← step 2: derived metrics, medians, deltas
│   ├── visualize.py              ← step 3: matplotlib line + bar charts
│   ├── export_ppt.py             ← step 4a: both PowerPoint decks
│   ├── export_excel.py           ← step 4b: the workbook
│   ├── narratives.py             ← optional Claude-written insight text
│   ├── manifest.py               ← per-run JSON audit record
│   └── utils.py                  ← shared helpers (logging, safe_divide, date parsing)
│
├── tests/                        ← pytest suite (75 tests)
├── templates/                    ← drop a branded .pptx here; point ESNL_TEMPLATE_PPT at it
├── .gitignore                    ← ignores generated .xlsx/.pptx/.png and logs/, with negations
│                                   that keep samples/ committable
├── FFIEC_MIGRATION_SPEC.md       ← stub; points at the separate FFIEC project
└── cnb_peer_analysis_q4_2025.html ← legacy standalone report, not produced by this pipeline

Documentation index

Document Read it when
docs/HANDOFF.md Start here on day one. Orientation, your first 30 minutes, your first real run, the quarterly operating procedure, pre-flight and post-run checklists, known limitations, and how this relates to the FFIEC project.
docs/CONFIGURING_PEERS.md You are changing who is compared: peers, the subject bank, median exclusions, the GSIB overlay. Includes ready-made scripts to dump every entity name in a data file and to check a candidate bank's field coverage before you commit to it, plus a full worked example of the Comerica/Webster → BMO/East West swap.
docs/CONFIGURING_METRICS.md You are changing what is measured: metrics, sections, formats, higher_is_better. Complete key reference, all seven calculation handlers with copy-pasteable templates, how to find the exact SNL field name, and why metrics silently disappear.
docs/DATA_CONTRACT.md You need the input file spec: filename rules, sheet inventory, the row/column map, cross-sheet de-duplication, unit conventions, how to produce a conforming export, and a validator script.
docs/ARCHITECTURE.md You are about to change Python, not YAML. Flow, module reference, the data shape at every boundary, the config-driven dispatch, the manifest, and walkthroughs for adding an export format or a chart type.
docs/TROUBLESHOOTING.md Something is missing, wrong, or silently empty. Indexed by the exact error or warning text, plus a description of what a healthy run looks like.
CONTRIBUTING.md You are modifying the code: environment setup, running the tests, type checking, conventions, adding a calculation handler, and the PR checklist.
samples/README.md You want to know what the sample data is, how it was generated, and how it differs from real SNL data.

Tests

The package is not installed into the venv by default, and there is no conftest.py, so the test suite needs PYTHONPATH=src:

cd enhanced_snl_peer_analysis      # wherever you cloned it
PYTHONPATH=src .venv/bin/python -m pytest -q

Expected: 74 passed, 1 skipped in about 5 seconds. The skip is the end-to-end test, which needs a real folder of quarterly files and skips itself when it cannot find one.

# skip the slow end-to-end test explicitly
PYTHONPATH=src .venv/bin/python -m pytest -q -m "not slow"

# with coverage (currently ~46% overall; the export modules are thinly covered)
PYTHONPATH=src .venv/bin/python -m pytest --cov=src/enhanced_snl_peer_analysis

Alternatively .venv/bin/pip install -e . once, after which plain pytest works.

Two honest caveats: the anthropic package must be installed for the test suite to pass even though no test makes a network call, and mypy is configured in pyproject.toml but is not installed by requirements.txt and does not currently pass — do not treat type checking as a gate.


License

MIT (see pyproject.toml). The code is MIT; the SNL data it consumes is not — that remains licensed content from S&P Global Market Intelligence and must not be redistributed.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages