Skip to content

Commit 6d4396b

Browse files
committed
Add reproducible method benchmark and methodology spec
benchmarks/masking_benchmark.py quantifies, on synthetic bills whose anomalies are known by construction (fixed seed, no private data), the advantage of the robust method over the textbook mean+stddev z-score: end-to-end F1 0.667 -> 1.000. Isolated scenarios attribute the gain to the MAD scale (masking) and the day-of-week baseline (seasonality) separately. --check runs in CI and fails on regression. METHODOLOGY.md documents the derivation and cites Iglewicz & Hoaglin (1993); architecture.md replaces the legacy HPMCE description with the real module layout.
1 parent f5bddb3 commit 6d4396b

4 files changed

Lines changed: 680 additions & 49 deletions

File tree

‎METHODOLOGY.md‎

Lines changed: 130 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,130 @@
1+
# Methodology
2+
3+
How `cloudsealed-jit` decides that a day of cloud spend is anomalous, why the
4+
method is built the way it is, and a reproducible benchmark quantifying its
5+
advantage over the textbook approach.
6+
7+
## The problem with the textbook approach
8+
9+
Cost anomaly detection is usually done by comparing each day against the period
10+
mean and flagging anything beyond two or three standard deviations. On cloud
11+
billing data that method fails in two specific, measurable ways.
12+
13+
**Standard deviation is inflated by the very spikes you are looking for.** A
14+
handful of large anomalies raises σ enough to pull themselves back inside the
15+
threshold and to hide every smaller anomaly with them. This is the *masking
16+
effect*, and it gets worse as the anomalies get bigger — exactly the regime
17+
that matters for cost control.
18+
19+
**A flat mean ignores the weekly cycle.** Most cloud bills have a pronounced
20+
weekday/weekend shape. Measured against a flat mean, ordinary Mondays look like
21+
overspend and ordinary Sundays look like savings, burying real signals under
22+
false positives.
23+
24+
## The method
25+
26+
### Baseline
27+
28+
Expected spend for a day is a *level* term times a *weekday* term:
29+
30+
```
31+
expected[i] = rolling_median(cost, 7)[i] × dow_factor[weekday(i)]
32+
```
33+
34+
The centred 7-day rolling median tracks organic growth and step changes without
35+
being dragged by spikes. The weekday factor is the median ratio of observed
36+
spend to the level term for that weekday; it is only estimated once at least
37+
two full weeks are present (below that every factor is 1.0). Factors are
38+
renormalised so seasonality reshapes the baseline without shifting its overall
39+
level. (`analysis.py:_baseline`, `_weekday_factors`)
40+
41+
### Scoring
42+
43+
Residuals against the baseline are scored with a modified z-score built on the
44+
median absolute deviation (MAD):
45+
46+
```
47+
z = 0.6745 × (x − baseline) / MAD
48+
```
49+
50+
The 0.6745 constant makes the MAD a consistent estimator of σ for normally
51+
distributed data, so the score keeps the familiar "number of deviations"
52+
reading while tolerating contamination in up to ~50% of the sample. Days at or
53+
above |z| = 3.5 are reported — the threshold recommended by Iglewicz & Hoaglin
54+
(1993) for the modified z-score. (`kernels.py:modified_zscores`,
55+
`analysis.py:Z_THRESHOLD`)
56+
57+
Where the MAD is exactly zero (a series that is constant apart from a few
58+
spikes), the scale falls back to a standard deviation; if both are zero the
59+
series is flat and every score is zero.
60+
61+
### Waste
62+
63+
Only *positive* excess counts. Waste percentage is the share of total spend
64+
sitting above the baseline on anomalous days, which converts directly to
65+
currency instead of being a count of unusual days. (`analysis.py:analyze`)
66+
67+
### Recommendations
68+
69+
Every recommendation carries a figure derived from the series itself,
70+
normalised to 30 days, and states its assumption in the description. Estimates
71+
that depend on facts the analyser cannot observe — whether a workload is
72+
production, whether a commitment is acceptable — are labelled conditional
73+
rather than presented as findings.
74+
75+
## Benchmark
76+
77+
`benchmarks/masking_benchmark.py` builds synthetic billing series whose true
78+
anomalies are known by construction (fixed seed, no private data) and reports
79+
precision / recall / F1 for each method. Reproduce with:
80+
81+
```
82+
python benchmarks/masking_benchmark.py
83+
```
84+
85+
### Result 1 — the scale estimator (masking)
86+
87+
Six spikes on an otherwise clean 90-day series: three very large (~+7000) and
88+
three moderate (~+1200). Both detectors use the *same* rolling-median baseline,
89+
so the only variable is MAD vs standard deviation.
90+
91+
| scale estimator | precision | recall | F1 |
92+
|---|---|---|---|
93+
| textbook standard deviation | 1.000 | 0.500 | 0.667 |
94+
| **robust MAD** | 0.857 | **1.000** | **0.923** |
95+
96+
The three large spikes inflate the standard deviation until its 3.5-σ threshold
97+
climbs *above* the three moderate spikes, masking them (recall 0.500). The MAD
98+
is unmoved by the large spikes and recovers all six (recall 1.000).
99+
100+
### Result 2 — the baseline (seasonality)
101+
102+
A strong weekday/weekend cycle (weekends 45% cheaper) with two genuine spikes,
103+
comparing a flat mean against the day-of-week baseline.
104+
105+
| baseline | precision | recall | F1 |
106+
|---|---|---|---|
107+
| textbook flat mean + stddev | 1.000 | 0.500 | 0.667 |
108+
| **robust day-of-week + MAD** | 1.000 | **1.000** | **1.000** |
109+
110+
### Result 3 — end to end (headline)
111+
112+
A realistic bill: organic growth (~+68% over the quarter), a weekly cycle,
113+
noise, and six spikes of graded size. Full shipped method vs full textbook
114+
method.
115+
116+
| method | precision | recall | F1 |
117+
|---|---|---|---|
118+
| textbook flat mean + stddev | 1.000 | 0.500 | 0.667 |
119+
| **robust (rolling median + DOW + MAD)** | 1.000 | **1.000** | **1.000** |
120+
121+
**The full method's F1 exceeds the textbook method's by 0.333** on this
122+
scenario. `benchmarks/masking_benchmark.py --check` re-runs this comparison and
123+
fails CI if the advantage ever drops below 0.25, so the claim cannot silently
124+
regress.
125+
126+
## Reference
127+
128+
Iglewicz, B. and Hoaglin, D. C. (1993). *How to Detect and Handle Outliers.*
129+
ASQC Quality Press. (Origin of the modified z-score and the |z| = 3.5
130+
threshold.)

‎README.md‎

Lines changed: 169 additions & 32 deletions
Original file line numberDiff line numberDiff line change
@@ -1,51 +1,188 @@
1-
# JIT-Optimization-Engine
1+
# cloudsealed-jit
22

3-
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://opensource.org/licenses/MIT)
4-
[![Python 3.9+](https://img.shields.io/badge/python-3.9+-green.svg)](https://www.python.org/downloads/)
5-
[![Performance: JIT-Compiled](https://img.shields.io/badge/Performance-JIT--Compiled-red.svg)]()
3+
Detects structural waste in cloud billing exports.
64

7-
## 🚀 Overview
5+
Given a billing export from AWS, GCP or Azure, it models what each day *should*
6+
have cost, reports the days that did not match, and turns the excess into a
7+
monthly figure. It is a library, a CLI and an HTTP service.
88

9-
**JIT-Optimization-Engine** is a high-performance data processing core designed for **analytical diagnostics** and stochastic optimization. At its heart, the project leverages **LLVM-based Just-In-Time (JIT) compilation** (via Numba) to achieve low-level execution speeds, allowing for the analysis of massive datasets in fractions of a second.
10-
11-
This engine was engineered to serve as a **Technical Audit and Simulation** layer, capable of processing hundreds of thousands of telemetry records and time-series data to identify computational inefficiencies and latency bottlenecks.
9+
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
10+
[![Python](https://img.shields.io/badge/python-3.10%2B-blue.svg)](https://www.python.org/downloads/)
1211

1312
---
1413

15-
## 🛠️ Technical Architecture & Key Pillars
14+
## The problem
1615

17-
The engine is built upon four pillars of advanced software engineering:
16+
Cloud cost anomaly detection is usually done by comparing each day against the
17+
period average and flagging anything beyond two or three standard deviations.
18+
On billing data that method fails in two specific ways.
1819

19-
1. **JIT Compilation (Numba/LLVM):** Transforms complex Python functions into native machine code. This allows the engine to perform mathematical and logical calculations with performance comparable to C++, which is essential for processing infrastructure logs without the overhead of the standard Python interpreter.
20-
2. **Massive Parallel Processing:** Utilizes `ProcessPoolExecutor` to distribute the analytical workload across multiple CPU cores, enabling the simultaneous processing of data from high-throughput databases such as **QuestDB**.
21-
3. **Stochastic Simulation Engine:** Implements specialized algorithms for calculating **Z-Score**, **Sharpe Ratio**, and **Expectancy**. In an engineering context, these metrics validate the stability and predictability of the analyzed datasets.
22-
4. **Micro-latency Diagnostics:** Designed for environments where milliseconds matter, capturing performance variations (jitter) that standard monitoring tools often overlook.
20+
**Standard deviation is inflated by the very spikes you are looking for.** A
21+
handful of large anomalies raises σ enough to pull themselves back inside the
22+
threshold, and to hide every smaller anomaly with them. This is the masking
23+
effect, and it gets worse as the anomalies get bigger.
2324

24-
---
25+
**A flat average ignores the weekly cycle.** Most cloud bills have a pronounced
26+
weekday/weekend shape. Measured against a flat mean, ordinary Mondays look like
27+
overspend and ordinary Sundays look like savings.
2528

26-
## 📈 Application in FinOps & Engineering (CloudSealed)
29+
## The method
2730

28-
This script serves as the technological foundation for **Advanced FinOps** diagnostics. While it does not automate refactoring, it provides the **data intelligence** required for:
31+
**Baseline.** Expected spend for a day is a level term times a weekday term:
2932

30-
* **Waste Auditing:** Analyzing CPU and Memory consumption logs to prove where legacy code is causing excessive cloud costs.
31-
* **Performance Validation:** Acting as the "benchmark" that compares system efficiency before and after senior-level code refactoring interventions.
32-
* **ROI Simulation:** Accurately quantifying the potential reduction in Cloud Spend when transitioning to high-performance architectures.
33+
```
34+
expected[i] = rolling_median(cost, 7)[i] × dow_factor[weekday(i)]
35+
```
3336

34-
---
37+
The rolling median follows growth and step changes without being dragged by
38+
spikes. The weekday factor is the median ratio of observed spend to the level
39+
term for that weekday. It is only estimated with at least two full weeks of
40+
data; below that every factor is 1.0.
41+
42+
**Scoring.** Residuals are scored with a modified z-score built on the median
43+
absolute deviation:
44+
45+
```
46+
z = 0.6745 × (x − baseline) / MAD
47+
```
48+
49+
The 0.6745 constant makes MAD a consistent estimator of σ for normal data, so
50+
the score keeps the familiar "number of deviations" reading while tolerating
51+
contamination in roughly half the sample. Days at or above |z| = 3.5 are
52+
reported — the threshold recommended by Iglewicz & Hoaglin (1993).
53+
54+
**Waste.** Only positive excess counts. Waste percentage is the share of total
55+
spend sitting above the baseline on anomalous days, which converts directly to
56+
currency instead of being a count of unusual days.
57+
58+
**Recommendations.** Each carries a figure derived from the series itself,
59+
normalised to 30 days, and states its assumption in the description. Estimates
60+
that depend on facts the analyser cannot observe — whether a workload is
61+
production, whether a commitment is acceptable — are labelled conditional
62+
rather than presented as findings.
3563

36-
## ⚡ Quick Start
64+
## Does it actually work better?
3765

38-
### Prerequisites
39-
* Python 3.9+
40-
* Libraries: `pandas`, `numpy`, `numba`, `requests`, `pytz`
66+
Yes, and it is measured, not asserted. `benchmarks/masking_benchmark.py` builds
67+
synthetic bills whose anomalies are known by construction and scores this
68+
method against the textbook mean+standard-deviation approach:
69+
70+
| scenario | textbook F1 | this method F1 |
71+
|---|---|---|
72+
| masking (scale estimator) | 0.667 | **0.923** |
73+
| seasonality (baseline) | 0.667 | **1.000** |
74+
| end-to-end | 0.667 | **1.000** |
75+
76+
Full derivation and reproduction steps in [METHODOLOGY.md](METHODOLOGY.md); the
77+
design of the codebase is in [architecture.md](architecture.md). The benchmark
78+
runs in CI (`--check`) and fails the build if the advantage ever regresses.
79+
80+
## Install
4181

42-
### Installation & Execution
4382
```bash
44-
# Clone the repository
45-
git clone [https://github.com/cloudsealed/JIT-Optimization-Engine.git](https://github.com/cloudsealed/JIT-Optimization-Engine.git)
83+
pip install cloudsealed-jit # library + CLI
84+
pip install "cloudsealed-jit[jit]" # + numba-compiled kernels
85+
pip install "cloudsealed-jit[jit,api]" # + HTTP service
86+
```
87+
88+
`numba` is optional. Without it the kernels run on pure NumPy and the results
89+
are identical; only large inputs get slower.
90+
91+
## Use
92+
93+
### CLI
94+
95+
```bash
96+
cloudsealed-jit billing-export.csv
97+
cloudsealed-jit billing-export.csv --json > findings.json
98+
cat export.csv | cloudsealed-jit - --type cost-forecast
99+
```
100+
101+
### Library
102+
103+
```python
104+
from cloudsealed_jit import parse_billing_csv, analyze
105+
106+
series = parse_billing_csv(open("export.csv").read())
107+
result = analyze(series)
108+
109+
print(result.metrics.wastePercentage)
110+
for r in result.recommendations:
111+
print(r.title, r.potentialSavings)
112+
```
113+
114+
### HTTP service
115+
116+
```bash
117+
docker run -p 8091:8091 cloudsealed/jit-optimization-engine
118+
```
119+
120+
```
121+
GET /health
122+
POST /v1/analyze-billing
123+
```
124+
125+
```bash
126+
curl -X POST localhost:8091/v1/analyze-billing \
127+
-H 'Content-Type: application/json' \
128+
-d '{"companyName":"Acme","csvContent":"date,cost\n2026-01-01,100\n..."}'
129+
```
130+
131+
Set `JIT_OPTIMIZATION_API_KEY` to require an `X-Api-Key` header. Set
132+
`JIT_MAX_CSV_BYTES` to change the 64 MB upload ceiling.
133+
134+
Response shape:
135+
136+
```jsonc
137+
{
138+
"anomalies": [
139+
{ "date": "2026-01-31", "expectedCost": 99.0, "actualCost": 500.0,
140+
"deviation": 405.05, "zScore": 7.82, "severity": "CRITICAL",
141+
"description": "Spend above the day-of-week baseline by USD 401.00 (405.1%)." }
142+
],
143+
"metrics": {
144+
"averageDailyCost": 106.32,
145+
"stdDeviation": 51.69,
146+
"sharpeRatio": 2.06, // spend stability: mean / stddev of daily cost
147+
"wastePercentage": 6.29 // share of total spend above the baseline
148+
},
149+
"recommendations": [
150+
{ "title": "...", "description": "...", "potentialSavings": 200.5, "effort": "MEDIUM" }
151+
],
152+
"summary": "..."
153+
}
154+
```
155+
156+
`sharpeRatio` is a **spend stability ratio** — mean daily cost divided by its
157+
standard deviation, the reciprocal of the coefficient of variation. Higher
158+
means more predictable spend. It is named for the field in the consuming API
159+
contract; it is not a risk-adjusted return.
160+
161+
## Supported exports
162+
163+
| Provider | Date column | Cost column |
164+
|---|---|---|
165+
| AWS Cost and Usage Report | `lineItem/UsageStartDate` | `lineItem/UnblendedCost` |
166+
| GCP billing export | `usage_start_time` | `cost` |
167+
| Azure cost export | `Date`, `UsageDateTime` | `Cost`, `CostInBillingCurrency` |
168+
| Generic | heuristic | heuristic |
169+
170+
Line items are aggregated to calendar days. Days with no line items are
171+
inserted as zero-spend days rather than skipped. Rows that cannot be parsed are
172+
counted and reported in the summary rather than dropped silently.
173+
174+
## Development
175+
176+
```bash
177+
pip install -e ".[jit,api,dev]"
178+
pytest
179+
```
180+
181+
The test suite builds synthetic exports whose correct answer is known in
182+
advance — a known spike at a known date, a known weekend-idle service, a stable
183+
series that must produce no findings — so the assertions test behaviour rather
184+
than the current output.
46185

47-
# Install dependencies
48-
pip install pandas numpy numba requests pytz
186+
## License
49187

50-
# Run the diagnostic engine
51-
python main.py
188+
MIT. See [LICENSE](LICENSE).

0 commit comments

Comments
 (0)