Add execution_environment fix + evals as code - #8
Open
quinnsebso wants to merge 1 commit into
Open
Conversation
- eval_dataset_versioned_test: 30-row eval dataset as a dbt table model - cortex_agent_evaluation materialization + 3 helpers: writes run-config YAML to a stage, optional opt-in eval run (run_on_build=false by default) - eval_versioned_test: run-config, ref()s agent + dataset so eval re-deploys with the agent - README: flag that analyst execution_environment.warehouse must be non-empty (blank causes 399504) - gitignore .claude/ and .DS_Store
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the analyst tool's missing warehouse (execution_environment was blank, causing 399504 on every eval query), and adds evals-as-code: the eval dataset and run-config as version-controlled dbt models plus a new cortex_agent_evaluation materialization, so evals live in git and re-deploy with the agent. Running an eval stays opt-in (run_on_build=false) since it costs credits.
Tested: versioned deploy preserves the eval run (ALTER, same object); --full-refresh drops it (CREATE OR REPLACE, new object). README updated to flag the non-empty warehouse requirement.