Index-based mlebabeclap_adotx kernel - #5619
Merged
WeiqunZhang merged 1 commit intoSep 2, 2026
Merged
Conversation
ankithadas
force-pushed
the
MLEBABecLap-Adotx-Index
branch
from
August 19, 2026 04:18
4dcff73 to
ed9709c
Compare
WeiqunZhang
approved these changes
Sep 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Rewrites
mlebabeclap_adotx(2D and 3D) from aBox-based kernel that loops internally(
amrex::Loop(box, ncomp, ...)) into a per-cell(i,j,k,n)kernel, and switches its EBgeometry arguments from ten separate
Array4s (flag,vfrc,apx/apy/apz,fcx/fcy/fcz,ba,bc) to the existingEBDataview, matching whatmlebabeclap_gsrbalready does.MLEBABecLap::Fapplynow calls it throughAMREX_HOST_DEVICE_PARALLEL_FOR_4Dinstead ofAMREX_LAUNCH_HOST_DEVICE_LAMBDA, and gets the geometry viafactory->getEBData(mfi).The arithmetic inside the kernel is untouched: the loop wrapper is removed and the body is
only re-indented, and the EB arrays are re-bound to identically named local references
(
auto const& apx = ebdata.get<EBData_t::apx>();etc.), so the expressions are characterfor character the same. Results are bitwise unchanged.
Additional background
This is the first of several small PRs that split up the stale WIP #4922 ("GPU specific
kernels for MLEBABecLap"), as requested there. It supersedes the
mlebabeclap_adotxpartof #4922. The templating on
Tthat #4922 also introduced has been dropped, per thereview comment on that PR. The
(i,j,k,n)form is the prerequisite for a later PR thatadds a fused (
MultiFab-wideParallelFor) GPU path forFapply; that follow-up is notincluded here.
Testing (all commands run from the repo root unless noted):
cmake -S . -B build-split -DAMReX_SPACEDIM=3 -DAMReX_EB=ON -DAMReX_LINEAR_SOLVERS_EM=OFF -DAMReX_ENABLE_TESTS=ON -DAMReX_TEST_TYPE=Small -DAMReX_MPI=ON -DCMAKE_BUILD_TYPE=Releasethen
cmake --build build-split -j8andctest --test-dir build-split --output-on-failure→ builds clean, 9/9 tests pass.
Tests/LinearSolvers/CellEB,make -j8 COMP=llvm USE_MPI=FALSE DIM=3andDIM=2,run over seven configurations (
sphere,sphere+eb_is_dirichlet=1,rotated_box,two_spheres,flower, two-levelsphere, periodicsphere;n_cell=64,verbose=2).The MLMG/BiCGStab residual histories and the initial/final max, 1- and 2-norm residuals
are bitwise identical to
developmentin both 2D and 3D.CUDA build on an NVIDIA RTX A5000 (CUDA 13.2, gcc 11.4):
Tests/LinearSolvers/CellEB,make -j8 COMP=gnu USE_MPI=FALSE USE_CUDA=TRUE CUDA_ARCH=86 DIM=3,same seven configurations. MLMG iteration counts are identical to
development(9, 12, 11, 3, 11, 28, 11); the residual values differ only at the level of the
run-to-run nondeterminism of the unmodified binary (GPU reductions are not
bit-reproducible), and all solves converge to
resid/resid0 ~ 1e-13.The same CellEB matrix was re-run on Linux/gcc 11.4 (
make -j8 COMP=gnu USE_MPI=FALSE,DIM=3andDIM=2) against adevelopmentbuild in a sibling worktree: residualhistories again bitwise identical in both dimensions.
Performance
Neither timing changed measurably; this PR is a refactor, not an optimisation.
CPU,
Tests/LinearSolvers/CellEBmain3d.gnu.TEST.ex(
make -j8 COMP=gnu USE_MPI=FALSE DIM=3, gcc 11.4, single rank),inputs n_cell=128 eb2.geom_type=sphere eb_is_dirichlet=1 verbose=1.All binaries were built first and then timed interleaved in the same session
(3 reps each) on a shared machine, so the absolute numbers are inflated but the
comparison is fair. Best of 3, MLMG
Timers: Solve[s]:developmentmax_grid_size=32max_grid_size=64GPU (NVIDIA RTX A5000, CUDA 13.2,
make -j8 COMP=gnu USE_MPI=FALSE USE_CUDA=TRUE CUDA_ARCH=86 DIM=3), same test atn_cell=256, again built up front and timed interleaved, best of 3:developmentmax_grid_size=32(512 boxes)max_grid_size=64(64 boxes)MLMG converged in the same number of iterations in every one of these runs.
Checklist
The proposed changes:
P.S Generated using Claude Code