Skip to content

Index-based mlebabeclap_adotx_centroid kernel - #5621

Open
ankithadas wants to merge 2 commits into
AMReX-Codes:developmentfrom
ankithadas:MLEBABecLap-AdotxCentroid-Index
Open

ankithadas wants to merge 2 commits into
AMReX-Codes:developmentfrom
ankithadas:MLEBABecLap-AdotxCentroid-Index

Conversation

@ankithadas

@ankithadas ankithadas commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Summary

Rewrites mlebabeclap_adotx_centroid (2D and 3D) from a Box-based kernel that loops
internally (amrex::Loop(box, ncomp, ...)) into a per-cell (i,j,k,n) kernel, and
replaces its eleven separate EB Array4 arguments (flag, vfrc, apx/apy/apz,
fcx/fcy/fcz, ccent, ba, bcent) with the existing EBData view.

MLEBABecLap::Fapply now calls it through AMREX_HOST_DEVICE_PARALLEL_FOR_4D and gets
all EB geometry from factory->getEBData(mfi), which lets the getVolFrac() /
getAreaFrac() / getFaceCent() / getBndryArea() / getBndryCent() / getCentroid()
lookups at the top of Fapply be dropped entirely.

The arithmetic inside the kernel is untouched: the loop wrapper is removed and the body is
only re-indented, and the EB arrays are re-bound to identically named local references
(auto const& ccent = ebdata.get<EBData_t::centroid>(); etc.), so the expressions are
character for character the same. Results are bitwise unchanged.

Additional background

This is part of the ongoing split of the stale WIP #4922 ("GPU specific kernels for
MLEBABecLap"), as requested there; it supersedes the mlebabeclap_adotx_centroid part of
#4922. The templating on T that #4922 also introduced has been dropped, per the review
comment on that PR.

This PR is stacked on top of the mlebabeclap_adotx PR (branch
ankithadas:MLEBABecLap-Adotx-Index) because both edit the same block of
MLEBABecLap::Fapply; please merge that one first. The diff shown here against
development therefore contains that commit as well.

Testing (all commands run from the repo root unless noted):

  • cmake -S . -B build-split -DAMReX_SPACEDIM=3 -DAMReX_EB=ON -DAMReX_LINEAR_SOLVERS_EM=OFF -DAMReX_ENABLE_TESTS=ON -DAMReX_TEST_TYPE=Small -DAMReX_MPI=ON -DCMAKE_BUILD_TYPE=Release
    then cmake --build build-split -j8 and ctest --test-dir build-split --output-on-failure
    → builds clean, 9/9 tests pass. A 2D library build
    (-DAMReX_SPACEDIM=2 -DAMReX_EB=ON) also builds clean.

  • Tests/LinearSolvers/LeastSquares is the test that actually exercises this kernel: it
    calls MLEBABecLap::setPhiOnCentroid() and then MLMG::apply(). Built with
    make -j8 COMP=llvm USE_MPI=FALSE DIM=2 and DIM=3, and run on
    inputs.2d.base, inputs.2d.askew-x, inputs.2d.fullyrotated,
    inputs.2d.trianglewave, inputs.3d.poiseuille.askew-all and
    inputs.3d.poiseuille.aligned.xy-x. The resulting plotfiles are byte for byte
    identical
    to those produced by development (diff -r over all six).

  • Tests/LinearSolvers/CellEB, make -j8 COMP=llvm USE_MPI=FALSE DIM=3 and DIM=2,
    seven configurations (sphere, sphere + eb_is_dirichlet=1, rotated_box,
    two_spheres, flower, two-level sphere, periodic sphere; n_cell=64,
    verbose=2): MLMG/BiCGStab residual histories and final residual norms are bitwise
    identical to development.

  • CUDA build on an NVIDIA RTX A5000 (CUDA 13.2, gcc 11.4):
    Tests/LinearSolvers/CellEB, make -j8 COMP=gnu USE_MPI=FALSE USE_CUDA=TRUE CUDA_ARCH=86 DIM=3.
    MLMG iteration counts are identical to development (9, 12, 11, 3, 11, 28, 11); the
    residual values differ only at the level of the run-to-run nondeterminism of the
    unmodified binary (GPU reductions are not bit-reproducible).

  • The same CellEB matrix was re-run on Linux/gcc 11.4 (make -j8 COMP=gnu USE_MPI=FALSE,
    DIM=3 and DIM=2) against a development build in a sibling worktree: residual
    histories again bitwise identical in both dimensions.

Performance

Neither timing changed measurably; this PR is a refactor, not an optimisation.

CPU, Tests/LinearSolvers/CellEB main3d.gnu.TEST.ex
(make -j8 COMP=gnu USE_MPI=FALSE DIM=3, gcc 11.4, single rank),
inputs n_cell=128 eb2.geom_type=sphere eb_is_dirichlet=1 verbose=1.
All binaries were built first and then timed interleaved in the same session
(3 reps each) on a shared machine, so the absolute numbers are inflated but the
comparison is fair. Best of 3, MLMG Timers: Solve [s]:

grids development this PR
max_grid_size=32 4.284 4.325
max_grid_size=64 4.534 4.553

GPU (NVIDIA RTX A5000, CUDA 13.2,
make -j8 COMP=gnu USE_MPI=FALSE USE_CUDA=TRUE CUDA_ARCH=86 DIM=3), same test at
n_cell=256, again built up front and timed interleaved, best of 3:

grids development this PR
max_grid_size=32 (512 boxes) 1.085 1.083
max_grid_size=64 (64 boxes) 0.834 0.824

MLMG converged in the same number of iterations in every one of these runs.

Checklist

The proposed changes:

  • fix a bug or incorrect behavior in AMReX
  • add new capabilities to AMReX
  • changes answers in the test suite to more than roundoff level
  • are likely to significantly affect the results of downstream AMReX users
  • include documentation in the code and/or rst files, if appropriate

P.S Generated using Claude Code

@ankithadas
ankithadas force-pushed the MLEBABecLap-AdotxCentroid-Index branch from 0ce7913 to e28c136 Compare August 19, 2026 04:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant