GB10 (sm_121) + Strix Halo (gfx1151) edge build/run notes#3
Open
AzeezIsh wants to merge 1 commit into
Open
Conversation
Merged
9 tasks
Palanivelg
pushed a commit
that referenced
this pull request
Jun 26, 2026
* feat: add client-side vLLM profiling trigger
Adds an optional client-side trigger that fires POST /start_profile
at the performance phase start and /stop_profile at run end, so a
profiled run can be driven from a YAML/CLI flag without coupling
endpoints to any vendor harness.
Schema: ProfilerEngine enum (currently {vllm}) and ProfilingConfig
hung off Settings. URLs are auto-derived per entry in
endpoint_config.endpoints (strip /v1, append engine-specific path).
Default-off; warn-don't-fail throughout.
Report.txt gets a Profiling section and a sibling profiling.json is
written next to result_summary.json when the trigger is enabled.
* fix: close unclosed Field() for metrics_tokenizer_workers
The profiler-trigger commit (a4fe30b) left the Field( call for
metrics_tokenizer_workers unterminated, so config/schema.py raised
SyntaxError and the inference-endpoint CLI could not import. Add the
missing ) so the module compiles.
* feat: allow separate profiling endpoint override
Add an optional profiling.endpoints (CLI --profile-endpoints) field so the
profiler start/stop triggers can target a different host than the inference
endpoint. When unset, URLs are still derived from endpoint_config.endpoints;
when set, derivation runs over the override list using the same engine-specific
protocol. Adds a scheme validator mirroring EndpointConfig and a matching
--profile-endpoints override on the from-config subcommand.
* refactor: drop --profile/--profile-urls overrides from from-config
Keep the from-config CLI surface minimal per review feedback: profiling
is configured via the YAML settings.profiling block for from-config runs.
This also removes the model_copy(update=...) path that bypassed
ProfilingConfig URL-scheme validation. offline/online keep the
schema-generated --profile/--profile-urls flags, which validate normally.
* test: cover profiling trigger config and helpers
Adds unit tests for the client-side profiling trigger (review finding #3):
- TestProfilingConfig (test_schema.py): defaults, engine enum coercion, and
URL-scheme validation on both the direct-construction (offline/online) and
model_validate (from-config YAML) paths.
- TestProfilingHelpers (test_benchmark.py): _derive_profile_urls /v1 stripping
and empty-endpoints ValueError, _post_profile 200/404/connection-failure via
mocked urlopen, _render_profile_status, and _write_profiling_section output
plus profiling.json serializability.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds GB10 (sm_121, CUDA) and Strix Halo (gfx1151, native ROCm) build and run
notes to the BFCL v4 example README, which currently documents only Thor and
x86 dGPU. Both configs were validated end to end (single-turn and multi-turn)
against the Thor reference.
Also adds a reproducibility note recommending a pinned llama.cpp commit: at
temperature 0, single-turn accuracy can drift past the 1% gate between build
versions with otherwise identical config, model, and seed.
Targeting feat/bfcl-v4-combined so it can fold into mlcommons#346.