Skip to content

Switch to CompilerCaching.jl.#961

Merged
luraess merged 3 commits into
mainfrom
tb/compilercaching
Jul 22, 2026
Merged

Switch to CompilerCaching.jl.#961
luraess merged 3 commits into
mainfrom
tb/compilercaching

Conversation

@maleadt

@maleadt maleadt commented Jul 6, 2026

Copy link
Copy Markdown
Member

No description provided.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMDGPU.jl Benchmarks

Details
Benchmark suite Current: 4915289 Previous: 240107f Ratio
amdgpu/synchronization/context/device 610 ns 610 ns 1
amdgpu/synchronization/stream/blocking 250 ns 250 ns 1
amdgpu/synchronization/stream/nonblocking 330 ns 350 ns 0.94
array/accumulate/Float32/1d 86291 ns 84951 ns 1.02
array/accumulate/Float32/dims=1 355675 ns 342295 ns 1.04
array/accumulate/Float32/dims=1L 135692 ns 137362 ns 0.99
array/accumulate/Float32/dims=2 133702 ns 126782 ns 1.05
array/accumulate/Float32/dims=2L 2828710 ns 2825642 ns 1.00
array/accumulate/Int64/1d 97981 ns 98651 ns 0.99
array/accumulate/Int64/dims=1 339875 ns 329425 ns 1.03
array/accumulate/Int64/dims=1L 167512 ns 168173 ns 1.00
array/accumulate/Int64/dims=2 127022 ns 125522 ns 1.01
array/accumulate/Int64/dims=2L 3004263 ns 3003954 ns 1.00
array/broadcast 101982 ns 92211 ns 1.11
array/construct 1770 ns 1760 ns 1.01
array/copy 38751 ns 37590 ns 1.03
array/copyto!/cpu_to_gpu 183723 ns 114062 ns 1.61
array/copyto!/gpu_to_cpu 161533 ns 183553 ns 0.88
array/copyto!/gpu_to_gpu 66871 ns 60371 ns 1.11
array/iteration/findall/bool 182373 ns 181042 ns 1.01
array/iteration/findall/int 201543 ns 188513 ns 1.07
array/iteration/findfirst/bool 128432 ns 118562 ns 1.08
array/iteration/findfirst/int 118671 ns 118282 ns 1.00
array/iteration/findmin/1d 169972 ns 170152 ns 1.00
array/iteration/findmin/2d 156433 ns 156083 ns 1.00
array/iteration/logical 348615 ns 348075 ns 1.00
array/iteration/scalar 294094 ns 303894 ns 0.97
array/permutedims/2d 74711 ns 73381 ns 1.02
array/permutedims/3d 72981 ns 73131 ns 1.00
array/permutedims/4d 76651 ns 76332 ns 1.00
array/random/rand/Float32 51450 ns 50961 ns 1.01
array/random/rand/Int64 57430 ns 57361 ns 1.00
array/random/rand!/Float32 89911 ns 87231 ns 1.03
array/random/rand!/Int64 89632 ns 96922 ns 0.92
array/random/randn/Float32 90981 ns 86641 ns 1.05
array/random/randn!/Float32 106741 ns 80272 ns 1.33
array/reductions/mapreduce/Float32/1d 134082 ns 133782 ns 1.00
array/reductions/mapreduce/Float32/dims=1 96071 ns 95681 ns 1.00
array/reductions/mapreduce/Float32/dims=1L 776461 ns 778072 ns 1.00
array/reductions/mapreduce/Float32/dims=2 98132 ns 97781 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 296604 ns 294594 ns 1.01
array/reductions/mapreduce/Int64/1d 135682 ns 134242 ns 1.01
array/reductions/mapreduce/Int64/dims=1 96552 ns 95802 ns 1.01
array/reductions/mapreduce/Int64/dims=1L 782951 ns 782372 ns 1.00
array/reductions/mapreduce/Int64/dims=2 98242 ns 96722 ns 1.02
array/reductions/mapreduce/Int64/dims=2L 297564 ns 296194 ns 1.00
array/reductions/reduce/Float32/1d 133532 ns 133413 ns 1.00
array/reductions/reduce/Float32/dims=1 96391 ns 95362 ns 1.01
array/reductions/reduce/Float32/dims=1L 770962 ns 772321 ns 1.00
array/reductions/reduce/Float32/dims=2 98442 ns 95681 ns 1.03
array/reductions/reduce/Float32/dims=2L 296314 ns 293734 ns 1.01
array/reductions/reduce/Int64/1d 133962 ns 134022 ns 1.00
array/reductions/reduce/Int64/dims=1 97791 ns 95612 ns 1.02
array/reductions/reduce/Int64/dims=1L 781442 ns 783112 ns 1.00
array/reductions/reduce/Int64/dims=2 98811 ns 96961 ns 1.02
array/reductions/reduce/Int64/dims=2L 298285 ns 296764 ns 1.01
array/reverse/1d 44371 ns 44210 ns 1.00
array/reverse/1dL 76171 ns 75981 ns 1.00
array/reverse/1dL_inplace 103922 ns 95681 ns 1.09
array/reverse/1d_inplace 71101 ns 77711 ns 0.91
array/reverse/2d 52091 ns 51820 ns 1.01
array/reverse/2dL 102171 ns 100941 ns 1.01
array/reverse/2dL_inplace 175052 ns 86952 ns 2.01
array/reverse/2d_inplace 84571 ns 132582 ns 0.64
array/sorting/1d 345175 ns 341995 ns 1.01
integration/byval/reference 39350 ns 39641 ns 0.99
integration/byval/slices=1 40230 ns 40761 ns 0.99
integration/byval/slices=2 160222 ns 148282 ns 1.08
integration/byval/slices=3 235113 ns 234153 ns 1.00
integration/volumerhs 5024212 ns 5028074 ns 1.00
kernel/indexing 65371 ns 129912 ns 0.50
kernel/indexing_checked 73611 ns 131002 ns 0.56
kernel/launch 1560 ns 1350 ns 1.16
kernel/rand 193932 ns 195323 ns 0.99
latency/import 1709436321 ns 1601453816 ns 1.07
latency/precompile 37339939945 ns 36585492651 ns 1.02
latency/ttfp 5731761660 ns 2182646656 ns 2.63

This comment was automatically generated by workflow using github-action-benchmark.

@maleadt
maleadt force-pushed the tb/compilercaching branch 2 times, most recently from 57f2026 to 0f18147 Compare July 10, 2026 14:28
@maleadt
maleadt force-pushed the tb/compilercaching branch from 0f18147 to bdb2ee9 Compare July 13, 2026 09:09
@maleadt
maleadt marked this pull request as ready for review July 13, 2026 10:04
@maleadt

maleadt commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

Done on my end. I'll leave it to AMDGPU.jl maintainers to merge this.

@luraess

luraess commented Jul 13, 2026

Copy link
Copy Markdown
Member

@wsmoses we have this now failing for Enzyme tests as this PR requires now GPUCompiler v2 and Enzyme is still at 1.6.2 (see https://buildkite.com/julialang/amdgpu-dot-jl/builds/3886#019f5b1f-0fd3-4bc7-8c10-878504fd0dc9/L669). Could we bump compat and patch Enzyme?

@wsmoses

wsmoses commented Jul 13, 2026

Copy link
Copy Markdown
Member

@vchuravy has a pr starting to look into it.

That said, @maleadt as the changes here don't look huge, is it possible to conditionally gate this with something like

if gpucompilerver >= 2
a()
else
b()
end

that way downstream packages that haven't adapted yet [like enzyme/reactant] can still run ci with?

And the same for cuda/metal?

@maleadt

maleadt commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

I'd rather not. I've added enough cruft already in GPUCompiler just to keep supporting 1.10 for Enzyme, and I don't want that leak into the back-ends too.

@wsmoses

wsmoses commented Jul 13, 2026

Copy link
Copy Markdown
Member

Fair enough, in which case would you be able to help us adapt enzyme/reactant to gpucompiler 2?

@maleadt

maleadt commented Jul 14, 2026

Copy link
Copy Markdown
Member Author

Valentin created PRs already. The changes are not tricky; point an LLM at my changes and let it update those PRs.

@luraess

luraess commented Jul 15, 2026

Copy link
Copy Markdown
Member

@wsmoses let me know if you can things rolling - as ideally I'd like to have Enzyme CI passing before merging this.

@wsmoses

wsmoses commented Jul 15, 2026

Copy link
Copy Markdown
Member

there's a draft pr open for enzyme by @vchuravy and no pr at all for reactant.

without some assistance from gpucompiler folks, there's likely going to be non-trivial latency to adapt to the api churn as we're still working through some higher priority blockers atm.

@christiangnrd

Copy link
Copy Markdown
Member

@wsmoses I've opened EnzymeAD/Enzyme.jl#3354 targeting the draft PR which keeps support for GPUCompiler v1. This doesn't seem like too much to maintain at least until most things have transitioned to GPUCompiler 2.

I gave Reactant.jl a quick look and I don't think any of the GPUCompiler methods used by Reactant.jl need to adjust to a need signature/name. Obviously behaviour may have changed and require fixing but it's possible it'll just be a compat bump

@luraess

luraess commented Jul 21, 2026

Copy link
Copy Markdown
Member

What's the status of EnzymeAD/Enzyme.jl#3354 ? In the meantime, I am thinking of allowing for soft fail (as CUDA does) and merge this PR.

@luraess
luraess merged commit 48774a6 into main Jul 22, 2026
6 of 7 checks passed
@luraess
luraess deleted the tb/compilercaching branch July 22, 2026 10:05
@luraess

luraess commented Jul 22, 2026

Copy link
Copy Markdown
Member

Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants