Skip to content

ci: use pull_request_target for PR triage workflow - #2237

Merged
fehiepsi merged 1 commit into
masterfrom
fix-pr-triage
Aug 6, 2026
Merged

fehiepsi merged 1 commit into
masterfrom
fix-pr-triage

Conversation

@Qazalbash

Copy link
Copy Markdown
Collaborator

Fork PRs get a read-only GITHUB_TOKEN under pull_request, which breaks requestReviewers (needs pull-requests: write).

Fork PRs get a read-only GITHUB_TOKEN under pull_request, which
breaks requestReviewers (needs pull-requests: write).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@Qazalbash Qazalbash self-assigned this Aug 5, 2026
@Qazalbash
Qazalbash requested a review from juanitorduz August 5, 2026 16:10
@github-actions

github-actions Bot commented Aug 5, 2026

Copy link
Copy Markdown

Benchmark report

this PR fix-pr-triage at 744fbcde vs baseline master at a47333ee

+ run time:     1 faster
  compile time: unchanged across 32 benchmarks

Significant changes (1)

                         ──────── run time ───────     ───── compile time ─────
  benchmark              baseline  this PR       Δ     baseline  this PR      Δ
───────────────────────────────────────────────────────────────────────────────
+ categorical_log_prob     2.5 ms   2.2 ms  -10.7%      75.2 ms  74.2 ms  -1.3%

Red is slower, green is faster; a row is coloured by the worse of its two columns. A delta in parentheses cleared the threshold on a measurement below the resolution floor, so it is shown without being called a change. † marks a benchmark that could not be compared — see below.

Full results

distributions

                                 ───────── run time ────────     ────── compile time ──────
  benchmark                      baseline  this PR         Δ     baseline   this PR       Δ
───────────────────────────────────────────────────────────────────────────────────────────
  biject_to_constraints            4.0 ms   4.1 ms     +1.6%     394.3 ms  369.6 ms   -6.2%
+ categorical_log_prob             2.5 ms   2.2 ms    -10.7%      75.2 ms   74.2 ms   -1.3%
  dirichlet_log_prob               707 µs   720 µs     +1.9%     540.3 ms  417.8 ms  -22.7%
  dirichlet_sample                45.6 ms  45.1 ms     -1.1%     883.6 ms  868.2 ms   -1.7%
  gamma_log_prob                   2.2 ms   2.2 ms     +1.1%       2.55 s    2.26 s  -11.2%
  gamma_sample                    19.9 ms  20.7 ms     +3.7%     891.2 ms  827.1 ms   -7.2%
  lkj_cholesky_sample              5.5 ms   5.3 ms     -3.7%       1.27 s    1.20 s   -5.0%
  mixture_same_family_log_prob     2.1 ms   2.1 ms     -2.6%     105.1 ms  107.9 ms   +2.7%
  multivariate_normal_log_prob     291 µs   343 µs  (+17.8%)     157.0 ms  165.5 ms   +5.4%
  normal_log_prob                  696 µs   694 µs     -0.2%      69.5 ms   58.5 ms  -15.8%
  normal_sample                   22.5 ms  21.7 ms     -3.7%     203.6 ms  207.6 ms   +2.0%
  stick_breaking_transform         6.4 ms   6.3 ms     -2.4%     226.2 ms  216.5 ms   -4.3%
  student_t_log_prob               3.1 ms   3.1 ms     -0.4%      79.0 ms   78.8 ms   -0.3%
  truncated_normal_log_prob        854 µs   778 µs   (-8.9%)      53.6 ms   53.0 ms   -1.1%

handlers

                                  ───────── run time ─────────     ────── compile time ──────
  benchmark                       baseline   this PR         Δ     baseline   this PR       Δ
─────────────────────────────────────────────────────────────────────────────────────────────
  initialize_model_hierarchical    40.4 ms   39.6 ms     -2.0%       3.96 s    3.85 s   -2.9%
  log_density_hierarchical          3.8 ms    3.7 ms     -1.2%       1.25 s    1.26 s   +1.0%
  nested_handler_stack              1.4 ms    1.4 ms     +2.5%       527 µs    467 µs  -11.3%
  potential_energy_and_grad          21 µs     25 µs  (+20.2%)     100.7 ms   96.4 ms   -4.2%
  predictive_forward_sampling     728.5 ms  726.5 ms     -0.3%     165.5 ms  180.0 ms   +8.7%
  trace_seeded_model                845 µs    829 µs     -1.8%     567.2 ms  574.4 ms   +1.3%

mcmc

                             ──────── run time ───────     ───── compile time ─────
  benchmark                  baseline   this PR      Δ     baseline  this PR      Δ
───────────────────────────────────────────────────────────────────────────────────
  hmc_logistic_regression    739.7 ms  727.4 ms  -1.7%       3.34 s   3.11 s  -7.0%
  nuts_dense_mass_funnel       1.20 s    1.18 s  -1.6%       2.88 s   2.68 s  -6.8%
  nuts_eight_schools           1.18 s    1.18 s  -0.5%       2.72 s   2.52 s  -7.4%
  nuts_hierarchical_glm        4.98 s    4.91 s  -1.5%       5.14 s   4.94 s  -3.8%
  nuts_logistic_regression     1.10 s    1.08 s  -2.2%       3.27 s   3.52 s  +7.7%
  nuts_vectorized_chains       2.48 s    2.51 s  +1.4%       3.00 s   2.95 s  -1.6%

svi

                                             ──────── run time ───────     ───── compile time ─────
  benchmark                                  baseline   this PR      Δ     baseline  this PR      Δ
───────────────────────────────────────────────────────────────────────────────────────────────────
  svi_autodelta_map_logistic                 302.9 ms  305.9 ms  +1.0%       3.08 s   2.97 s  -3.7%
  svi_autodiagonalnormal_hierarchical          1.07 s    1.04 s  -2.5%       5.29 s   5.20 s  -1.7%
  svi_automultivariatenormal_eight_schools   772.8 ms  738.4 ms  -4.5%       4.26 s   4.10 s  -3.8%
  svi_autonormal_logistic                    757.8 ms  756.5 ms  -0.2%       3.39 s   3.23 s  -4.6%
  svi_multi_particle_elbo                      1.52 s    1.52 s  -0.1%       3.49 s   3.29 s  -5.6%
  svi_trace_mean_field_elbo                    1.32 s    1.31 s  -0.5%       5.75 s   5.43 s  -5.6%
Methodology and environment

Each benchmark is set up untimed, then called once with the JAX caches cleared and several more times warm. Run is the fastest warm call; compile is the first call minus that, i.e. the tracing, lowering and XLA compilation the warm calls did not have to pay for.

Both refs were measured on the same runner over 2 interleaved round(s), taking the best observation per benchmark. A result is called neutral when it moves less than ±5% (run) or ±25% (compile), or when the measurement itself is under 1 ms (run) / 50 ms (compile) — a shared CI runner cannot resolve changes below that. Compile time gets the looser band because it is measured once per round rather than best-of-N, and swings by roughly 20% even between two runs of identical code. A delta shown in parentheses did clear its threshold, but on a measurement below the resolution floor, so it is reported without being called a change.

baseline this PR
ref master fix-pr-triage
commit a47333ee 744fbcde
numpyro 0.21.0 0.21.0
jax 0.11.0 0.11.0
backend cpu cpu
python 3.14.6 3.14.6

Runner: Linux-6.17.0-1020-azure-x86_64-with-glibc2.39, 4 CPUs.

Produced by this benchmark run.

@fehiepsi
fehiepsi merged commit 0aef50c into master Aug 6, 2026
10 checks passed
@fehiepsi
fehiepsi deleted the fix-pr-triage branch August 6, 2026 00:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants