Skip to content

Port themefinder#148: switch evals to LiteLLM gateway, add dynamic model discovery - #1548

Merged
saashanair merged 8 commits into
mainfrom
saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow
Aug 18, 2026
Merged

saashanair merged 8 commits into
mainfrom
saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow

Conversation

@saashanair

Copy link
Copy Markdown
Contributor

Context

Model access should be allowed only through the central i.AI LLM gateway (LiteLLM). This PR makes that true everywhere in themefinder/ — evals, synthetic data generation, and the benchmark CLI — and replaces the hardcoded model list with one that updates automatically as the gateway's model list changes.

This ports i-dot-ai/themefinder#148, raised against the old standalone themefinder . That PR was raised without realising that themefinder had been subtree-merged into consult — this replays that PR's diff under the themefinder/ prefix. File paths are the only thing that changed; the logic below is unmodified from the original PR.

Stacked on #1546 (re-adding the missing themes.json ground-truth files) — needed for evals/tests/ and the live eval CI job to actually pass here. This PR will auto-retarget to main once that one merges.

Changes proposed in this pull request

Dynamic model discovery (new)
evals/utils_gateway.py fetches the gateway's /model_group/info and /health/latest and turns them into a list of chat-capable models, each with vendor family and health resolved. Filtering is composable (split_unhealthy, filter_by_family, select_by_name) so callers decide what to do with unhealthy or unknown matches.

benchmark.py — no more hardcoded model list
The old MODEL_REGISTRY (Azure/Vertex/locai) is gone, along with the Vertex code path (it never worked — create_llm() unconditionally raised NotImplementedError for it). Relevant models are now found dynamically by querying the gateway.

Before After
--provider {azure,vertex,locai,all} --models <name>... / --family {gpt,claude,gemini,locai}... / --all (exactly one required)
reasoning_effort hardcoded per model name; no way to sweep effort levels Creates a separate config per effort level by sweeping --reasoning-effort {low,medium,high}... across all models whose supports_reasoning flag, read from the gateway, is true

Model selection (--models vs. --family/--all) is two small functions, _select_named_models/_select_healthy_models, each returning only the fields relevant to its own mode — --models warns on unhealthy matches but still runs them (explicit request by name); --family/--all drops unhealthy matches and reports what was dropped.

Synthetic data generation — off direct Azure
evals/synthetic/* now talks to the gateway (openai.AsyncOpenAI(base_url=LLM_GATEWAY_URL, ...)) instead of constructing an AsyncAzureOpenAI client directly. The two models it uses are now named constants (DRAFTING_MODEL, RESPONSE_GENERATION_MODEL) instead of repeated string literals.

Gateway credentials — one source of truth
LLM_GATEWAY_URL/CONSULT_EVAL_LITELLM_API_KEY were each read directly via os.getenv in 10 separate places across benchmark.py, metrics.py, all four eval_*.py files, synthetic/cli.py, and generate_synthetic.py — only utils_gateway.py's own client validated them before use; everywhere else silently passed None/None into OpenAILLM/AsyncOpenAI on a misconfigured environment. Added utils_gateway.gateway_credentials() as the single source of truth; every call site now gets the same fail-fast validation.

Config cleanup
.env.example and the eval workflow no longer list AZURE_OPENAI_*/GOOGLE_CLOUD_*/LOCAI_* — confirmed unused anywhere in the repo post-migration. LLM_GATEWAY_URL/CONSULT_EVAL_LITELLM_API_KEY are documented (previously required but missing from .env.example). AUTO_EVAL_4_1_SWEDEN_DEPLOYMENT (read by the eval judge/scoring path) is now prefilled with a known-working value plus a TODO, instead of being left blank and failing with a confusing 400 partway through a run.

Tests
35 tests (evals/tests/test_utils_gateway.py, test_benchmark.py) covering the discovery and selection logic against mocked gateway responses.

Guidance to review

Verified in the original PR, against the old standalone repo:

  • generate_synthetic.py runs end-to-end (1 question, 10 responses) — verified real output on disk
  • All 35 unit tests pass
  • benchmark.py --models <unhealthy-name> through a full run — the warn-but-still-run behavior is unit-tested; no currently-unhealthy model was available on the live gateway to exercise it end-to-end

Re-verified directly in consult, after porting:

  • All 35 unit tests pass unchanged in this repo (uv run pytest evals/tests/ -v)
  • discover_chat_models() runs live against the real gateway, 57 chat-capable models discovered with real health status
  • benchmark.py --quick runs end-to-end here too; mapping/condensation/refinement produce real scores. generation currently errors on an unrelated pre-existing bug in src/themefinder's batch-processing retry path (not introduced by this PR — happy to open a separate issue)
  • benchmark.py with no selector exits 2 with the expected argparse error, matching test_no_selector_is_invalid
  • benchmark.py --models gpt-4.1-nano-sweden --evals mapping runs end-to-end, real scores (f1: 0.479 / 0.503 across the two question parts)
  • benchmark.py --family locai --evals mapping — model discovered, selected by family filter, and received live API calls through the gateway. Reproduces the exact same known, pre-existing locai structured-output failure the original PR documented (malformed/non-JSON responses to the mapping prompt) — not a regression from this PR

Things to check

  • I have added any new ENV vars in all deployed environments and updated the .env.test files in the repo

n/a — no new env vars. Required a routing/permissions change on the gateway side instead: i-dot-ai/core-llm-gateway#232 which fixed the reading of available models from the gateway, alongside updating the routes that CI key (CONSULT_EVAL_LITELLM_API_KEY) can access.

@saashanair
saashanair requested a review from a team as a code owner August 12, 2026 09:08
@saashanair
saashanair requested review from 252afh and removed request for a team August 12, 2026 09:08
@linear

linear Bot commented Aug 12, 2026

Copy link
Copy Markdown

PRO-618

@github-actions

github-actions Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 2m 23s —— View job


PR Review

Clean port with solid test coverage. Two issues worth addressing:


🔴 DO — benchmark.py lines 989–990 and 1004–1006: silent success on config failures

Both early-exit paths use bare return, so the process exits with code 0 even when no models were found or matched. CI will pass silently on a misconfigured environment.

# line 989
if missing:
    console.print(...)
    return          # ← sys.exit(1) needed here

# line 1004
if not selected:
    console.print(...)
    return          # ← and here

Fix this →%20after%20%60console.print(f%22%5Bred%5DNo%20matching%20models...%22)%60%20around%20line%20989%2C%20and%202)%20after%20%60console.print(%22%5Bred%5DNo%20models%20matched...%22)%60%20around%20line%201005.%20Both%20are%20error%20conditions%20that%20should%20exit%20non-zero%20for%20CI%20to%20catch%20them.&repo=i-dot-ai/consult)


🟡 TRY — utils_gateway.py line 183: direct key access on gateway response

item["supports_reasoning"] will raise KeyError if the gateway ever omits this field. filter_chat_models already uses .get() for mode — worth being consistent:

# current
supports_reasoning=item["supports_reasoning"],

# safer
supports_reasoning=item.get("supports_reasoning", False),

@saashanair
saashanair force-pushed the saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow branch from 6d7d3bd to ec266f2 Compare August 12, 2026 09:18
@github-actions

github-actions Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 1m 2s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

Clean port with solid test coverage. Two issues from the previous review remain unaddressed:


🔴 DO — benchmark.py lines 989 and 1004: silent success on config failures

Both early-exit paths use bare return, so the process exits with code 0 even when no models were found or matched. CI will pass silently on a misconfigured environment.

# line 989
if missing:
    console.print(...)
    return          # ← sys.exit(1) needed here

# line 1004
if not selected:
    console.print(...)
    return          # ← and here

Fix this →%20after%20%60console.print(f%22%5Bred%5DNo%20matching%20models...%22)%60%20around%20line%20989%2C%20and%202)%20after%20%60console.print(%22%5Bred%5DNo%20models%20matched...%22)%60%20around%20line%201005.%20Both%20are%20error%20conditions%20that%20should%20exit%20non-zero%20for%20CI%20to%20catch%20them.&repo=i-dot-ai/consult)


🟡 TRY — utils_gateway.py line 184: direct key access on gateway response

item["supports_reasoning"] will raise KeyError if the gateway ever omits this field. filter_chat_models already uses .get() for mode — worth being consistent:

# current
supports_reasoning=item["supports_reasoning"],

# safer
supports_reasoning=item.get("supports_reasoning", False),

@github-actions

github-actions Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 47s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

Both issues from previous reviews are now fixed — sys.exit(1) at both early-exit paths in benchmark.py, and item.get("supports_reasoning", False) in utils_gateway.py. No further issues found. ✅

@saashanair
saashanair force-pushed the saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow branch from 057f6a8 to c5a4503 Compare August 12, 2026 11:22
@github-actions

github-actions Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 1m 9s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

Both previous issues are resolved and no new ones found. Clean port. ✅

@saashanair
saashanair force-pushed the saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow branch from c5a4503 to 0a5f7cc Compare August 13, 2026 07:51
@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 2m 22s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

Clean port. Both previous DO/TRY issues (sys.exit(1) at early-exit paths and .get("supports_reasoning", False)) are confirmed fixed. ✅

One small nit:

⚪️ NIT — generate_synthetic.py line 58: carried-over comment references "medium reasoning"

The comment survived from the Azure port but "medium reasoning" isn't an API parameter being set — it describes an assumed default of the gpt-5-nano-sweden model. Worth removing or replacing with the constant name to stay consistent with the pattern introduced by this PR:

# before
# Initialise LLM for response generation (gpt-5-nano with medium reasoning)
# Medium reasoning ≈ o1 performance, 2x faster throughput than mini/low

# after (or just drop the comment — the constant is self-documenting)
# Initialise LLM for response generation

Base automatically changed from chore/re-add-missing-themefinder-json-files to main August 13, 2026 08:01
@saashanair
saashanair force-pushed the saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow branch from 0a5f7cc to d87685b Compare August 13, 2026 08:01
@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 1m 4s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

All previous DO/TRY issues (sys.exit(1) at both early-exit paths, .get("supports_reasoning", False)) remain fixed. ✅

One carry-over from the last review:

⚪️ NIT — generate_synthetic.py lines 58–59: stale comment

The two-line comment still describes Azure-specific behaviour that no longer applies:

# Initialise LLM for response generation (gpt-5-nano with medium reasoning)
# Medium reasoning ≈ o1 performance, 2x faster throughput than mini/low

The constant RESPONSE_GENERATION_MODEL is already self-documenting — this comment can just be dropped. Fix this →


Comment thread .github/workflows/themefinder-eval.yml
Comment thread themefinder/evals/utils_gateway.py
Comment thread themefinder/evals/synthetic/cli.py
Comment thread themefinder/evals/tests/conftest.py
Comment thread themefinder/.env.example
@github-actions

github-actions Bot commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 1m 9s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

All previously flagged issues are resolved: sys.exit(1) at both early-exit paths, .get("supports_reasoning", False), and the stale Azure comment in generate_synthetic.py. TODO for hardcoded model-family matching is now in place. ✅

One open thread from 252afh's review:

⚪️ NIT — utils_gateway.py line 25: _GPT_MARKERS missing "o5-"

"gpt" catches gpt-5* models, but o5- (and o2-) aren't covered by any existing marker. The TODO at lines 20–23 acknowledges this. Fix this →

saashanair and others added 8 commits August 18, 2026 10:12
Ports the code from i-dot-ai/themefinder#148 ("Switch evals/benchmark to
LiteLLM gateway, add dynamic model discovery"), raised against the old
standalone themefinder repo before it was subtree-merged into consult.
Replays that PR's diff under the themefinder/ prefix; file paths are the
only thing that changed, logic is unmodified.

Replaces the static, hand-maintained model list in evals/benchmark.py
with live discovery via the gateway's /model_group/info and
/health/latest endpoints (evals/utils_gateway.py), so evals stop
targeting deployment names that have quietly gone stale on the gateway.

Adds httpx as an explicit themefinder dependency (utils_gateway.py's
gateway calls) and regenerates uv.lock accordingly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
themefinder-eval.yml: drop AZURE_OPENAI_ENDPOINT, AZURE_OPENAI_API_KEY,
and OPENAI_API_VERSION secrets - nothing in evals/ reads them anymore
after the LiteLLM gateway migration.

themefinder-ci.yml: scope the coverage-gated pytest run to tests/ only.
It was running bare `pytest -v -s`, which would now also sweep up the
new evals/tests/ into the 95% coverage gate. Added as a separate,
uncounted step instead, matching how the original themefinder repo's
CI handled the same evals/tests/ addition.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both early-exit paths (no gateway match for a named model, nothing
selected at all) used a bare `return`, so main() completed normally
and the process exited 0 even though nothing ran. CI review flagged
this: on a misconfigured environment, a run that did nothing would
still show green.

Kept the existing fail-fast behavior for a missing named model
(abort rather than silently run a smaller comparison than requested)
but fixed the exit code, and added a comment explaining why that one
stays strict rather than warn-and-continue like the unhealthy case.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
item["supports_reasoning"] and the per-row check["model_name"]/
["checked_at"]/["status"] accesses were all direct dict lookups, so
one malformed entry from the gateway would KeyError and take down
model discovery entirely - CI review flagged the first one, the same
class of bug was present in latest_health_by_model too.

supports_reasoning now falls back to False (matches GatewayModel's own
default; defaulting True risks sending a reasoning_effort param to a
model that doesn't support it, defaulting False just skips an optional
sweep). A malformed health-check row is now skipped rather than
crashing the whole call - the model still gets discovered, it just
falls back to the existing "unknown" health state, same as a stale or
missing check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two failure modes we hit debugging this port surfaced as opaque
errors deep in discover_chat_models():

- A key without access to /model_group/info or /health/latest failed
  with a raw httpx.HTTPStatusError - now caught around the gather()
  call and re-raised as a RuntimeError naming the exact route and
  pointing at the gateway UI's allowed-paths setting for the key.
- A key whose model grant is a wildcard (e.g. "all-team-models") can
  get /model_group/info back unexpanded - a single "*" row instead of
  real model names - which would otherwise silently become a fake,
  unusable "*"-named model selectable via --all. Now detected and
  raised explicitly before any filtering happens.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Five new cases, none previously covered:
- supports_reasoning missing from a gateway entry defaults to False
- a malformed health-check row (missing model_name/checked_at/status)
  is skipped rather than crashing latest_health_by_model
- a 401/403 from either gateway route raises RuntimeError
- a non-401/403 error (e.g. 500) still propagates unchanged, rather
  than getting misconverted into the permissions-error message
- an unexpanded '*' model_group entry raises RuntimeError

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_FAMILY_SUBSTRINGS/_GPT_MARKERS require manually adding new model name
patterns by hand (e.g. a future o-series release) rather than deriving
family from real metadata. Flagged as a TODO rather than fixed now,
since the gateway doesn't currently expose family as a field.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…plies

The two-line comment at generate_synthetic.py's LLM init duplicated
RESPONSE_GENERATION_MODEL (self-documenting) and asserted a stale,
unverifiable performance claim disconnected from any actual parameter
in that code. Dropped per review NIT.

Moved the substance of that comment (why gpt-5-nano uses medium
reasoning) to response_generator.py, right where reasoning_effort is
actually set, instead of losing it entirely.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@saashanair
saashanair force-pushed the saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow branch from 3aa3cc2 to e420574 Compare August 18, 2026 09:12
@github-actions

github-actions Bot commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @saashanair's task in 1m 19s —— View job


PR Review

  • Gather context and read changed files
  • Analyze code quality, patterns, and potential issues
  • Post review feedback

All previously flagged DO/TRY issues remain fixed (sys.exit(1) at both early-exit paths, .get("supports_reasoning", False), stale Azure comment dropped). ✅

One carry-over NIT still open:

⚪️ NIT — utils_gateway.py line 25: _GPT_MARKERS still missing "o5-" and "o2-"

The TODO noting the limitation is in place, but the markers themselves haven't been extended. "gpt" doesn't catch o5-* or o2-* model names, so they'll get family=None and be unselectable via --family gpt. Fix this →

@saashanair
saashanair merged commit 640ac70 into main Aug 18, 2026
22 checks passed
@saashanair
saashanair deleted the saashanair/pro-618-switch-to-using-litellm-throughout-evalspipeline-and-allow branch August 18, 2026 09:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants