Switch to CompilerCaching.jl.#961
Conversation
There was a problem hiding this comment.
AMDGPU.jl Benchmarks
Details
| Benchmark suite | Current: 4915289 | Previous: 240107f | Ratio |
|---|---|---|---|
amdgpu/synchronization/context/device |
610 ns |
610 ns |
1 |
amdgpu/synchronization/stream/blocking |
250 ns |
250 ns |
1 |
amdgpu/synchronization/stream/nonblocking |
330 ns |
350 ns |
0.94 |
array/accumulate/Float32/1d |
86291 ns |
84951 ns |
1.02 |
array/accumulate/Float32/dims=1 |
355675 ns |
342295 ns |
1.04 |
array/accumulate/Float32/dims=1L |
135692 ns |
137362 ns |
0.99 |
array/accumulate/Float32/dims=2 |
133702 ns |
126782 ns |
1.05 |
array/accumulate/Float32/dims=2L |
2828710 ns |
2825642 ns |
1.00 |
array/accumulate/Int64/1d |
97981 ns |
98651 ns |
0.99 |
array/accumulate/Int64/dims=1 |
339875 ns |
329425 ns |
1.03 |
array/accumulate/Int64/dims=1L |
167512 ns |
168173 ns |
1.00 |
array/accumulate/Int64/dims=2 |
127022 ns |
125522 ns |
1.01 |
array/accumulate/Int64/dims=2L |
3004263 ns |
3003954 ns |
1.00 |
array/broadcast |
101982 ns |
92211 ns |
1.11 |
array/construct |
1770 ns |
1760 ns |
1.01 |
array/copy |
38751 ns |
37590 ns |
1.03 |
array/copyto!/cpu_to_gpu |
183723 ns |
114062 ns |
1.61 |
array/copyto!/gpu_to_cpu |
161533 ns |
183553 ns |
0.88 |
array/copyto!/gpu_to_gpu |
66871 ns |
60371 ns |
1.11 |
array/iteration/findall/bool |
182373 ns |
181042 ns |
1.01 |
array/iteration/findall/int |
201543 ns |
188513 ns |
1.07 |
array/iteration/findfirst/bool |
128432 ns |
118562 ns |
1.08 |
array/iteration/findfirst/int |
118671 ns |
118282 ns |
1.00 |
array/iteration/findmin/1d |
169972 ns |
170152 ns |
1.00 |
array/iteration/findmin/2d |
156433 ns |
156083 ns |
1.00 |
array/iteration/logical |
348615 ns |
348075 ns |
1.00 |
array/iteration/scalar |
294094 ns |
303894 ns |
0.97 |
array/permutedims/2d |
74711 ns |
73381 ns |
1.02 |
array/permutedims/3d |
72981 ns |
73131 ns |
1.00 |
array/permutedims/4d |
76651 ns |
76332 ns |
1.00 |
array/random/rand/Float32 |
51450 ns |
50961 ns |
1.01 |
array/random/rand/Int64 |
57430 ns |
57361 ns |
1.00 |
array/random/rand!/Float32 |
89911 ns |
87231 ns |
1.03 |
array/random/rand!/Int64 |
89632 ns |
96922 ns |
0.92 |
array/random/randn/Float32 |
90981 ns |
86641 ns |
1.05 |
array/random/randn!/Float32 |
106741 ns |
80272 ns |
1.33 |
array/reductions/mapreduce/Float32/1d |
134082 ns |
133782 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=1 |
96071 ns |
95681 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=1L |
776461 ns |
778072 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2 |
98132 ns |
97781 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2L |
296604 ns |
294594 ns |
1.01 |
array/reductions/mapreduce/Int64/1d |
135682 ns |
134242 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=1 |
96552 ns |
95802 ns |
1.01 |
array/reductions/mapreduce/Int64/dims=1L |
782951 ns |
782372 ns |
1.00 |
array/reductions/mapreduce/Int64/dims=2 |
98242 ns |
96722 ns |
1.02 |
array/reductions/mapreduce/Int64/dims=2L |
297564 ns |
296194 ns |
1.00 |
array/reductions/reduce/Float32/1d |
133532 ns |
133413 ns |
1.00 |
array/reductions/reduce/Float32/dims=1 |
96391 ns |
95362 ns |
1.01 |
array/reductions/reduce/Float32/dims=1L |
770962 ns |
772321 ns |
1.00 |
array/reductions/reduce/Float32/dims=2 |
98442 ns |
95681 ns |
1.03 |
array/reductions/reduce/Float32/dims=2L |
296314 ns |
293734 ns |
1.01 |
array/reductions/reduce/Int64/1d |
133962 ns |
134022 ns |
1.00 |
array/reductions/reduce/Int64/dims=1 |
97791 ns |
95612 ns |
1.02 |
array/reductions/reduce/Int64/dims=1L |
781442 ns |
783112 ns |
1.00 |
array/reductions/reduce/Int64/dims=2 |
98811 ns |
96961 ns |
1.02 |
array/reductions/reduce/Int64/dims=2L |
298285 ns |
296764 ns |
1.01 |
array/reverse/1d |
44371 ns |
44210 ns |
1.00 |
array/reverse/1dL |
76171 ns |
75981 ns |
1.00 |
array/reverse/1dL_inplace |
103922 ns |
95681 ns |
1.09 |
array/reverse/1d_inplace |
71101 ns |
77711 ns |
0.91 |
array/reverse/2d |
52091 ns |
51820 ns |
1.01 |
array/reverse/2dL |
102171 ns |
100941 ns |
1.01 |
array/reverse/2dL_inplace |
175052 ns |
86952 ns |
2.01 |
array/reverse/2d_inplace |
84571 ns |
132582 ns |
0.64 |
array/sorting/1d |
345175 ns |
341995 ns |
1.01 |
integration/byval/reference |
39350 ns |
39641 ns |
0.99 |
integration/byval/slices=1 |
40230 ns |
40761 ns |
0.99 |
integration/byval/slices=2 |
160222 ns |
148282 ns |
1.08 |
integration/byval/slices=3 |
235113 ns |
234153 ns |
1.00 |
integration/volumerhs |
5024212 ns |
5028074 ns |
1.00 |
kernel/indexing |
65371 ns |
129912 ns |
0.50 |
kernel/indexing_checked |
73611 ns |
131002 ns |
0.56 |
kernel/launch |
1560 ns |
1350 ns |
1.16 |
kernel/rand |
193932 ns |
195323 ns |
0.99 |
latency/import |
1709436321 ns |
1601453816 ns |
1.07 |
latency/precompile |
37339939945 ns |
36585492651 ns |
1.02 |
latency/ttfp |
5731761660 ns |
2182646656 ns |
2.63 |
This comment was automatically generated by workflow using github-action-benchmark.
57f2026 to
0f18147
Compare
0f18147 to
bdb2ee9
Compare
|
Done on my end. I'll leave it to AMDGPU.jl maintainers to merge this. |
|
@wsmoses we have this now failing for Enzyme tests as this PR requires now GPUCompiler v2 and Enzyme is still at 1.6.2 (see https://buildkite.com/julialang/amdgpu-dot-jl/builds/3886#019f5b1f-0fd3-4bc7-8c10-878504fd0dc9/L669). Could we bump compat and patch Enzyme? |
|
@vchuravy has a pr starting to look into it. That said, @maleadt as the changes here don't look huge, is it possible to conditionally gate this with something like if gpucompilerver >= 2 that way downstream packages that haven't adapted yet [like enzyme/reactant] can still run ci with? And the same for cuda/metal? |
|
I'd rather not. I've added enough cruft already in GPUCompiler just to keep supporting 1.10 for Enzyme, and I don't want that leak into the back-ends too. |
|
Fair enough, in which case would you be able to help us adapt enzyme/reactant to gpucompiler 2? |
|
Valentin created PRs already. The changes are not tricky; point an LLM at my changes and let it update those PRs. |
|
@wsmoses let me know if you can things rolling - as ideally I'd like to have Enzyme CI passing before merging this. |
|
there's a draft pr open for enzyme by @vchuravy and no pr at all for reactant. without some assistance from gpucompiler folks, there's likely going to be non-trivial latency to adapt to the api churn as we're still working through some higher priority blockers atm. |
|
@wsmoses I've opened EnzymeAD/Enzyme.jl#3354 targeting the draft PR which keeps support for GPUCompiler v1. This doesn't seem like too much to maintain at least until most things have transitioned to GPUCompiler 2. I gave Reactant.jl a quick look and I don't think any of the GPUCompiler methods used by Reactant.jl need to adjust to a need signature/name. Obviously behaviour may have changed and require fixing but it's possible it'll just be a compat bump |
|
What's the status of EnzymeAD/Enzyme.jl#3354 ? In the meantime, I am thinking of allowing for soft fail (as CUDA does) and merge this PR. |
|
Thanks! |
No description provided.