Rendered Surface Evolver examples: simulated liquid in green, solid constraints in orange.
How good are large language models at writing complex physical simulations in a custom data format?
Surface Evolver Bench asks models to write complete
Surface Evolver .fe datafiles
for liquid surfaces shaped by geometry, constraints, contact angles, gravity,
and volume/area terms. The generated file is not judged by vibes or by another
model: it has to load in Evolver, run through hidden scripts, and match measured
physical quantities from reference solutions.
Surface Evolver is a tool released in 1992 (!) for modeling liquid surfaces. It is useful for tasks such as studying solder deposition on chips, modeling liquid fuel tanks, or designing lab-on-a-chip networks. A simulation is defined using Evolver's custom datafile syntax: vertices, edges, faces, bodies, constraints, energies, and boundary integrals.
Surface Evolver is powerful, terse, and niche. A correct answer is not just a plausible-looking text file: the model has to choose the right datafile sections, orient elements consistently, define bodies and constraints, and produce a file that Evolver can actually load, evolve, and measure. This eval therefore tests tool use, domain-specific syntax, numerical validation, and debugging under feedback.
Each task asks a model to create a complete Surface Evolver datafile. The model is given a small tool environment:
get_evolver_doc: retrieve curated local Surface Evolver documentation fromtools/docs.run_surface_evolver: run a candidate.fefile against a command script and inspect Evolver output.submit_fe_file: submit the final raw.fefile through a structured tool call.
The point is to evaluate whether a model can use documentation and execution feedback to produce a real artifact. The final submission is not scraped from Markdown or prose; it must arrive as a structured tool call containing the raw file content.
Submitted .fe files should describe the starting geometry and may define
reusable Evolver commands such as gogo := { ... }. They must not execute
commands while loading: no quit, bare gogo, refine, g, print, or
similar top-level commands. A bare read marker may be used before command
definitions, but not to execute another command file. Public and hidden command
scripts decide when to run refinement or measurement steps, and the runner
appends quit automatically.
For each model and task, se_eval.run_eval performs two stages:
generate: send the task prompt to the model, expose the tools above, allow up to--max-roundsrepair rounds, and writesubmission.fewhen the model callssubmit_fe_file.grade: readsubmission.fe, run deterministic static checks, run hidden Surface Evolver checks, and writeresult.json.
A run directory contains:
submission.fe: the model's submitted datafile.transcript.json: model messages, tool calls, and tool results.generation.json: model id, rounds used, submission status, and token usage when available.result.json: score, pass/fail status, static check details, hidden Evolver output, and metric measurements.run_error.json: written when generation or grading errors out, with a coarse category such asmalformed_provider_response,connectivity,quota_or_credit,rate_limit,no_submission, orgrader_or_evolver.visual.offandvisual.svg: optional Evolver-derived visual artifacts when grading with--visual.
The current grader uses a simple weighted score:
0.3: static datafile checks pass.0.3: the hidden Surface Evolver run completes successfully.0.4: hidden metric checks pass within their tolerances.
passed is true only when the static checks and all dynamic checks pass. Static
checks catch obvious structural omissions, such as too few vertices, missing
sections, missing body ids, or required syntax fragments. Dynamic checks are the
authority: they load the submitted file in Surface Evolver, run task-specific
commands, print SE_METRIC values, and compare those values with expected
targets.
Docker is the easiest way to run the eval because the image installs Surface Evolver and the Python dependencies.
docker build -t se-llm-eval .
docker run --rm \
-e OPENROUTER_API_KEY="$OPENROUTER_API_KEY" \
-v "$PWD/runs:/app/runs" \
se-llm-evalRun a specific task and model:
docker run --rm \
-e OPENROUTER_API_KEY="$OPENROUTER_API_KEY" \
-v "$PWD/runs:/app/runs" \
se-llm-eval python -m se_eval.run_eval \
--task-visibility public \
--task two_bubbles_2d \
--baseline deepseek-v4-proRun all configured baselines for a task:
docker run --rm \
-e OPENROUTER_API_KEY="$OPENROUTER_API_KEY" \
-v "$PWD/runs:/app/runs" \
se-llm-eval python -m se_eval.run_eval \
--task-visibility public \
--task two_bubbles_2d \
--all-baselinesThe helper script builds the image and runs the default task:
./run.shIt uses OpenRouter by default and reads OPENROUTER_API_KEY or
.openrouter_key. Set SE_EVAL_API_PROVIDER=openai to use OPENAI_API_KEY or
.openai_key; when only OpenAI credentials are present, the script selects
OpenAI automatically.
Local runs require Python 3.12+ and a Surface Evolver executable. The runner
looks for EVOLVER_BIN, ./evolver, evolver, evolver-nox-d, and
evolver-nox, in that order.
Install Python dependencies:
uv syncRun a smoke check that does not call a model:
PYTHONPATH=src uv run python -m se_eval.run_eval --list-baselinesRun a full eval locally with the default private task directory:
OPENROUTER_API_KEY=... \
PYTHONPATH=src uv run python -m se_eval.run_eval \
--baseline gpt-5.5Run a public task locally:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--task-visibility public \
--task two_bubbles_2d \
--baseline deepseek-v4-proThe runner supports OpenRouter through its OpenAI-compatible Chat Completions
endpoint and native OpenAI through the Responses API. The native path preserves
reasoning items across tool calls. OpenRouter remains the default for backwards
compatibility. Select native OpenAI with --api-provider openai and configure
authentication with OPENAI_API_KEY (or a local .openai_key file):
OPENAI_API_KEY=... \
PYTHONPATH=src uv run python -m se_eval.run_eval \
--api-provider openai \
--model gpt-5.5 \
--task-visibility public \
--task two_bubbles_2dOpenRouter-style openai/ model prefixes are stripped automatically when the
native OpenAI provider is selected, so an existing OpenAI baseline can also be
overridden without changing its model id:
OPENAI_API_KEY=... \
PYTHONPATH=src uv run python -m se_eval.run_eval \
--api-provider openai \
--baseline gpt-5.5To make native OpenAI persistent for a configured model, set api_provider in
eval_config.json. The provider object is specifically OpenRouter routing and
must be omitted for native providers:
{
"name": "gpt-5.5-native",
"model": "gpt-5.5",
"api_provider": "openai",
"reasoning_effort": "high"
}List named baselines:
PYTHONPATH=src uv run python -m se_eval.run_eval --list-baselinesRun one named baseline:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--task-visibility public \
--task two_bubbles_2d \
--baseline deepseek-v4-proOverride the model id directly:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--task-visibility public \
--task two_bubbles_2d \
--model openai/gpt-5.5For models that support reasoning controls:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--task-visibility public \
--task two_bubbles_2d \
--baseline gpt-5.5 \
--reasoning-effort highNamed models in eval_config.json can also declare a default
reasoning_effort:
{
"name": "gpt-5.5-high",
"model": "openai/gpt-5.5",
"reasoning_effort": "high"
}When no reasoning option is passed, the runner uses the configured model's
reasoning_effort. If neither the CLI nor config sets one, the request omits
reasoning controls and the API default applies. Both APIs receive a nested
reasoning object with the selected effort.
Native OpenAI Responses do not include a dollar cost in their usage object, so
the runner estimates standard-tier cost from input, cached-input, cache-write,
and output token counts. New runs store the estimate per round in
generation.json; chart generation also derives it for existing native OpenAI
runs that have token usage but no recorded cost. Unknown models and non-default
service tiers are left without a cost rather than assigned a guessed price.
These are estimates and do not include regional-processing uplifts.
Supported environment variables:
OPENROUTER_API_KEYOPENAI_API_KEYOPENAI_BASE_URLOPENAI_MODELSE_EVAL_API_PROVIDERSE_EVAL_REASONING_EFFORTOPENROUTER_MODELOPENROUTER_BASELINEOPENROUTER_REASONING_EFFORTOPENROUTER_BASE_URLOPENROUTER_HTTP_REFEREROPENROUTER_APP_TITLEEVOLVER_BINSE_EVAL_TASK_VISIBILITYSE_EVAL_TASK_DIRSE_EVOLVER_DOC_DIR
Generate only:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--stage generate \
--task-visibility public \
--task two_bubbles_2d \
--baseline deepseek-v4-proRe-grade an existing run without calling a model:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--stage grade \
--task-visibility private \
--task <task_id> \
runs/deepseek-v4-pro/private_<task_id>Grade a direct .fe file and print the result without writing JSON:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--stage grade \
--task-visibility public \
--task two_bubbles_2d \
--no-write \
path/to/submission.feWrite Evolver-derived visual artifacts during grading:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--stage grade \
--task-visibility public \
--task two_bubbles_2d \
--visual \
path/to/submission.feUse se_eval.run_matrix to run selected models across the public and private
task sets, grade each submission, and consolidate compact outcomes for later
plotting.
Preview the full matrix without calling models:
PYTHONPATH=src uv run python -m se_eval.run_matrix \
--baseline deepseek-v4-pro \
--task-visibility all \
--dry-runRun selected baselines on all public and private tasks:
OPENROUTER_API_KEY=... \
PYTHONPATH=src uv run python -m se_eval.run_matrix \
--baseline deepseek-v4-pro \
--baseline mistral-medium-3-5 \
--task-visibility allMatrix runs now use stable, incrementally fillable paths:
runs/<model-name>/<visibility>_<task-id>/
For example, DeepSeek's configured run for the public two-bubble task lives at
runs/deepseek-v4-pro/public_two_bubbles_2d/. The run directories are the
canonical source of truth. Compact aggregate files are derived from those
artifacts:
summary.json: aggregate pass rates and mean scores by model and by task, rewritten after each completed run.runs/plots/merged_outcomes.jsonlandruns/plots/merged_outcomes.csv: optional plot/export artifacts suitable for pandas.
Useful options include --model <openrouter/model-id> for exact model ids,
--all-baselines, repeated --task <task_id> filters, --skip-existing for
incrementally filling directories that lack both result.json and
run_error.json, and --summary-file for a custom derived summary path.
Configured reasoning defaults keep the configured model name in run directories.
Explicit --reasoning-effort matrix overrides use _reasoning-<effort> or
_reasoning-na suffixes so comparisons do not collide; derived plot/export rows
include reasoning_effort and model_run_label.
Create the model/task directory grid without calling models:
PYTHONPATH=src uv run python -m se_eval.sync_runsThe model list comes from eval_config.json. Use full configured model names
such as deepseek-v4-pro; older short aliases like deepseek still resolve for
now but are deprecated.
To compare multiple reasoning efforts in one matrix, either add separate config
entries such as gpt-5.5-medium and gpt-5.5-high, or pass a comma-separated
list / repeat the option. Use na for a run that omits reasoning controls:
PYTHONPATH=src uv run python -m se_eval.run_matrix \
--baseline gpt-5.5 \
--task-visibility all \
--reasoning-effort high,na,lowUse se_eval.plot_results to join one or more run roots and create
plot-ready merged data plus SVG charts. If an icons/ directory is present, the
plotter embeds provider PNG icons in the SVG charts and colors model bars by
provider. Use --icon-dir <path> to point at another icon directory, or
--no-icons to disable embedding.
Plot the default stable runs root:
PYTHONPATH=src uv run python -m se_eval.plot_results \
--output-dir runs/plotsPlot a specific run root:
PYTHONPATH=src uv run python -m se_eval.plot_results \
runs \
--output-dir runs/joined_plotsThe plotting script writes:
index.html: a small local report embedding the charts.by_model.svg: pass rate, mean score, total tokens, and recorded cost by model.model_tool_usage.svg: average assistant turns and stacked tool calls by model.score_vs_total_tokens.svg: mean score versus total tokens by model.score_vs_total_cost.svg: mean score versus recorded inference cost by model, with a stepped efficient frontier.pass_rate_vs_total_tokens.svg: pass rate versus total tokens by model.pass_rate_vs_total_cost.svg: pass rate versus recorded inference cost by model, with a stepped efficient frontier.by_task.svg: pass rate and mean score by task.task_model_heatmap.svg: mean score for each task/model pair.merged_outcomes.jsonlandmerged_outcomes.csv: joined row-level data, including reasoning labels, token/cost fields, and interaction counts whengeneration.jsonandtranscript.jsonare available.aggregates.json: aggregate statistics by model and by task, including total/mean token and recorded-cost values.
The public web page lives in docs/ and is deployed by
.github/workflows/pages.yml. The charts are copied from the generated plot
artifacts, and docs/data/aggregates.json is published as the compact aggregate
snapshot. Row-level CSV and JSONL files are not published for now.
After adding tasks, models, or run results, refresh the site assets with:
PYTHONPATH=src uv run python tools/update_pages_site.py --replotThat command regenerates runs/plots/ from the current run artifacts under
runs/, then copies the public chart and aggregate artifacts into:
docs/assets/charts/docs/data/
To publish from a specific run root instead of the default runs/, pass it
after --replot:
PYTHONPATH=src uv run python tools/update_pages_site.py --replot \
runs/some_matrixCommit the changed files in docs/ and push to main; the Pages workflow will
deploy the updated static site.
When grading criteria change, use se_eval.rescore_results to re-run only the
deterministic grader on saved submission.fe files. This does not call models or
spend LLM credits.
Preview a rescore for one task/model pair:
PYTHONPATH=src uv run python -m se_eval.rescore_results \
runs \
--task two_bubbles_2d \
--model gpt-5.5 \
--dry-runRescore all runs under runs/:
PYTHONPATH=src uv run python -m se_eval.rescore_resultsThe rescore command overwrites each submitted run's result.json and refreshes
the derived summary.json. By default it creates a .bak-<timestamp> copy of
summary.json before rewriting it. After rescoring, run se_eval.plot_results
again to regenerate joined plots.
Tasks are JSON files in tasks_public/ and tasks_private/. Use
--task-visibility public or --task-visibility private to select which set to
load. Private is the default. --task-dir or SE_EVAL_TASK_DIR can still point
at an explicit directory and takes precedence over task visibility.
A task defines:
id,title, and the user-facingprompt.public_command_script, whichrun_surface_evolveruses when a tool call leavescommand_scriptempty or omitted.hidden_command_script, which the grader runs before metric checks.static_checks, such as minimum element counts, required substrings, and command hygiene for submitted datafiles.dynamic_checks.evolver_metrics, which are Evolver expressions with expected values and tolerances.
Example task families currently include:
two_bubbles_2d: build a two-bubble 2D string model with two prescribed enclosed areas and one shared internal edge.- droplet tasks with contact-angle energy and content integrals.
- bridge tasks with liquid surfaces spanning constrained solid supports.
- groove and well tasks with liquid surfaces shaped by channel-like constraints.
To add a task, create tasks_public/<task_id>.json or
tasks_private/<task_id>.json, make the public validation script useful but not
exhaustive, and put the real acceptance criteria in hidden Evolver metric
checks. Then run:
PYTHONPATH=src uv run python -m se_eval.run_eval \
--task-visibility public \
--task <task_id> \
--baseline deepseek-v4-proThe documentation tool reads curated pages from tools/docs. The supported
topics are:
datafilesyntaxelementscommandssingletoggleconstraintsquantitiesenergiesmound
The eval-time tool does not parse HTML. To derive curated markdown-ish pages from a local Surface Evolver manual checkout, use:
python tools/convert_evolver_docs.py /path/to/Evolver/doc \
--topic datafile \
--topic syntax \
--topic elements \
--topic commandsThe eval intentionally avoids asking the model to paste a .fe file into a
chat answer. Tool calls make the artifact boundary explicit:
- the model can retrieve only the docs exposed by the benchmark,
- candidate execution is captured in the transcript,
- the final submission is raw file content,
- grading does not depend on brittle Markdown extraction.
This also makes the benchmark closer to real agentic engineering work: read domain docs, write a candidate, run the simulator, repair mistakes, and submit a file that passes deterministic checks.
Apache-2.0. See LICENSE.



