Skip to content

GB10 (sm_121) + Strix Halo (gfx1151) edge build/run notes#3

Open
AzeezIsh wants to merge 1 commit into
Palanivelg:feat/bfcl-v4-combinedfrom
AzeezIsh:azeez/edge-build-notes
Open

GB10 (sm_121) + Strix Halo (gfx1151) edge build/run notes#3
AzeezIsh wants to merge 1 commit into
Palanivelg:feat/bfcl-v4-combinedfrom
AzeezIsh:azeez/edge-build-notes

Conversation

@AzeezIsh

@AzeezIsh AzeezIsh commented Jun 9, 2026

Copy link
Copy Markdown

Adds GB10 (sm_121, CUDA) and Strix Halo (gfx1151, native ROCm) build and run
notes to the BFCL v4 example README, which currently documents only Thor and
x86 dGPU. Both configs were validated end to end (single-turn and multi-turn)
against the Thor reference.

Also adds a reproducibility note recommending a pinned llama.cpp commit: at
temperature 0, single-turn accuracy can drift past the 1% gate between build
versions with otherwise identical config, model, and seed.

Targeting feat/bfcl-v4-combined so it can fold into mlcommons#346.

Palanivelg pushed a commit that referenced this pull request Jun 26, 2026
* feat: add client-side vLLM profiling trigger

Adds an optional client-side trigger that fires POST /start_profile
at the performance phase start and /stop_profile at run end, so a
profiled run can be driven from a YAML/CLI flag without coupling
endpoints to any vendor harness.

Schema: ProfilerEngine enum (currently {vllm}) and ProfilingConfig
hung off Settings. URLs are auto-derived per entry in
endpoint_config.endpoints (strip /v1, append engine-specific path).
Default-off; warn-don't-fail throughout.

Report.txt gets a Profiling section and a sibling profiling.json is
written next to result_summary.json when the trigger is enabled.

* fix: close unclosed Field() for metrics_tokenizer_workers

The profiler-trigger commit (a4fe30b) left the Field( call for
metrics_tokenizer_workers unterminated, so config/schema.py raised
SyntaxError and the inference-endpoint CLI could not import. Add the
missing ) so the module compiles.

* feat: allow separate profiling endpoint override

Add an optional profiling.endpoints (CLI --profile-endpoints) field so the
profiler start/stop triggers can target a different host than the inference
endpoint. When unset, URLs are still derived from endpoint_config.endpoints;
when set, derivation runs over the override list using the same engine-specific
protocol. Adds a scheme validator mirroring EndpointConfig and a matching
--profile-endpoints override on the from-config subcommand.

* refactor: drop --profile/--profile-urls overrides from from-config

Keep the from-config CLI surface minimal per review feedback: profiling
is configured via the YAML settings.profiling block for from-config runs.
This also removes the model_copy(update=...) path that bypassed
ProfilingConfig URL-scheme validation. offline/online keep the
schema-generated --profile/--profile-urls flags, which validate normally.

* test: cover profiling trigger config and helpers

Adds unit tests for the client-side profiling trigger (review finding #3):
- TestProfilingConfig (test_schema.py): defaults, engine enum coercion, and
  URL-scheme validation on both the direct-construction (offline/online) and
  model_validate (from-config YAML) paths.
- TestProfilingHelpers (test_benchmark.py): _derive_profile_urls /v1 stripping
  and empty-endpoints ValueError, _post_profile 200/404/connection-failure via
  mocked urlopen, _render_profile_status, and _write_profiling_section output
  plus profiling.json serializability.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant