Skip to content

Index-based mlebabeclap_adotx kernel - #5619

Merged
WeiqunZhang merged 1 commit into
AMReX-Codes:developmentfrom
ankithadas:MLEBABecLap-Adotx-Index
Sep 2, 2026
Merged

Index-based mlebabeclap_adotx kernel#5619
WeiqunZhang merged 1 commit into
AMReX-Codes:developmentfrom
ankithadas:MLEBABecLap-Adotx-Index

Conversation

@ankithadas

@ankithadas ankithadas commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Summary

Rewrites mlebabeclap_adotx (2D and 3D) from a Box-based kernel that loops internally
(amrex::Loop(box, ncomp, ...)) into a per-cell (i,j,k,n) kernel, and switches its EB
geometry arguments from ten separate Array4s (flag, vfrc, apx/apy/apz,
fcx/fcy/fcz, ba, bc) to the existing EBData view, matching what
mlebabeclap_gsrb already does.

MLEBABecLap::Fapply now calls it through AMREX_HOST_DEVICE_PARALLEL_FOR_4D instead of
AMREX_LAUNCH_HOST_DEVICE_LAMBDA, and gets the geometry via factory->getEBData(mfi).

The arithmetic inside the kernel is untouched: the loop wrapper is removed and the body is
only re-indented, and the EB arrays are re-bound to identically named local references
(auto const& apx = ebdata.get<EBData_t::apx>(); etc.), so the expressions are character
for character the same. Results are bitwise unchanged.

Additional background

This is the first of several small PRs that split up the stale WIP #4922 ("GPU specific
kernels for MLEBABecLap"), as requested there. It supersedes the mlebabeclap_adotx part
of #4922. The templating on T that #4922 also introduced has been dropped, per the
review comment on that PR. The (i,j,k,n) form is the prerequisite for a later PR that
adds a fused (MultiFab-wide ParallelFor) GPU path for Fapply; that follow-up is not
included here.

Testing (all commands run from the repo root unless noted):

  • cmake -S . -B build-split -DAMReX_SPACEDIM=3 -DAMReX_EB=ON -DAMReX_LINEAR_SOLVERS_EM=OFF -DAMReX_ENABLE_TESTS=ON -DAMReX_TEST_TYPE=Small -DAMReX_MPI=ON -DCMAKE_BUILD_TYPE=Release
    then cmake --build build-split -j8 and ctest --test-dir build-split --output-on-failure
    → builds clean, 9/9 tests pass.

  • Tests/LinearSolvers/CellEB, make -j8 COMP=llvm USE_MPI=FALSE DIM=3 and DIM=2,
    run over seven configurations (sphere, sphere + eb_is_dirichlet=1, rotated_box,
    two_spheres, flower, two-level sphere, periodic sphere; n_cell=64, verbose=2).
    The MLMG/BiCGStab residual histories and the initial/final max, 1- and 2-norm residuals
    are bitwise identical to development in both 2D and 3D.

  • CUDA build on an NVIDIA RTX A5000 (CUDA 13.2, gcc 11.4):
    Tests/LinearSolvers/CellEB, make -j8 COMP=gnu USE_MPI=FALSE USE_CUDA=TRUE CUDA_ARCH=86 DIM=3,
    same seven configurations. MLMG iteration counts are identical to development
    (9, 12, 11, 3, 11, 28, 11); the residual values differ only at the level of the
    run-to-run nondeterminism of the unmodified binary (GPU reductions are not
    bit-reproducible), and all solves converge to resid/resid0 ~ 1e-13.

  • The same CellEB matrix was re-run on Linux/gcc 11.4 (make -j8 COMP=gnu USE_MPI=FALSE,
    DIM=3 and DIM=2) against a development build in a sibling worktree: residual
    histories again bitwise identical in both dimensions.

Performance

Neither timing changed measurably; this PR is a refactor, not an optimisation.

CPU, Tests/LinearSolvers/CellEB main3d.gnu.TEST.ex
(make -j8 COMP=gnu USE_MPI=FALSE DIM=3, gcc 11.4, single rank),
inputs n_cell=128 eb2.geom_type=sphere eb_is_dirichlet=1 verbose=1.
All binaries were built first and then timed interleaved in the same session
(3 reps each) on a shared machine, so the absolute numbers are inflated but the
comparison is fair. Best of 3, MLMG Timers: Solve [s]:

grids development this PR
max_grid_size=32 4.284 4.328
max_grid_size=64 4.534 4.554

GPU (NVIDIA RTX A5000, CUDA 13.2,
make -j8 COMP=gnu USE_MPI=FALSE USE_CUDA=TRUE CUDA_ARCH=86 DIM=3), same test at
n_cell=256, again built up front and timed interleaved, best of 3:

grids development this PR
max_grid_size=32 (512 boxes) 1.085 1.083
max_grid_size=64 (64 boxes) 0.834 0.833

MLMG converged in the same number of iterations in every one of these runs.

Checklist

The proposed changes:

  • fix a bug or incorrect behavior in AMReX
  • add new capabilities to AMReX
  • changes answers in the test suite to more than roundoff level
  • are likely to significantly affect the results of downstream AMReX users
  • include documentation in the code and/or rst files, if appropriate

P.S Generated using Claude Code

@ankithadas
ankithadas force-pushed the MLEBABecLap-Adotx-Index branch from 4dcff73 to ed9709c Compare August 19, 2026 04:18
@WeiqunZhang
WeiqunZhang merged commit f191389 into AMReX-Codes:development Sep 2, 2026
75 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants