Repository navigation
feat(core/mpi/nccl) Add an error checking mechanism - #281
rbourgeois33 wants to merge 26 commits into
Conversation
It is my preference too.
I am ok, but it will need some thoughts about
Do you mean
This is the really hard thing with distributed computing. I think it is good to focus on local error handling but we have to be contious that we can deadlock or fail. |
@dssgabriel, is that enforced somewhere ? |
|
Added stronger error checking on the nccl backend: Core idea: If we catch a cuda error, we return an unexpected In the case where the nccl error is In Which requires to have access to the nccl communicator, now a member of nccl |
|
By choosing the default move ctor for nccl As a result ~Request() noexcept {
if (request_ != nullptr) {
KC_CUDA_CHECK(cudaEventDestroy(request_));
}To fix it, I rewrote the move ctor: Request(Request&& other)
: request_(std::move(other.request_)),
callbacks_(std::move(other.callbacks_)),
status_(std::move(other.status_)),
comm_(std::move(other.comm_)) {
other.request_ = nullptr;
}; |
@rbourgeois33, is this what I reported in #264? I am planning on pushing the code fixing that over the weekend 👍 |
Description
Improves error handling with
expected.Related:
Technical description
This PR proposes an error checking mechanism and applies it to
send/recvas well asbroadcast. I will extend it to other primitives once we converged on technical choices. It is based onexpected. I did not wrapRequestintoexpectedbut rather added astatus_member toRequest, so that.valuein the calls (send(...).value().wait()). Crucially, this PR does not breaks the current API.Requestobject, despite failures that can happen both at the communication creationauto req= KokkosComm::send(..), and thewait()callreq.wait().Blocking call that don't return aBlocking calls are untouched.Requestcan now return anexpectedtype (see e.g.mpi/send).Some examples (from the tests):
Some questions to answer together:
expected? It seems that this was ~agreed upon (Error handling, how to do it? #29 (comment), Error handling, how to do it? #29 (reply in thread))tl::expected, to avoid imposing ac++23compiler ? (as suggested Error handling, how to do it? #29 (reply in thread))KOKKOSCOMM_ABORT_ON_ERROR=OFFand only works because I setMPI_ERRORS_RETURNby hand. Should we handle all combinations ? In particular ifMPI_ERRORS_ARE_FATALis set, the code can abort even ifKOKKOSCOMM_ABORT_ON_ERROR=OFF*. Should there be constraints between the two ?KokkosComm::Abort, but I am not sure how to implement it. (Error handling, how to do it? #29 (comment), Implement Teuchos MPI operations #10)TODO:
Changes
N.B. the following list is AI generated
Error,ErrorCodeandstatus_type(tl::expected<void, Error>) inerror.hpptl-expecteddependency (CMake + package config)KC_MPI_*/KC_NCCL_*/KC_CUDA_CHECK_REQcheck macros that return a failedRequestinstead of abortingKOKKOSCOMM_ABORT_ON_ERRORmakes all check macros abort instead of returningRequest: store error status, addfailed(),has_error(),error_code(),backend_error_code()Request::wait(): record backend errors instead of aborting, skip callbacks on errorRequeston error:isend,irecv,ibroadcast,iallgather,iallreduce,ialltoall,ireducesend,recv,broadcast,allgather,allreduce,alltoall(incl. pre-2.28 grouped path),reduceNotSupportedinstead of abortingmpi::functions,Channel,test,wait_all,wait_anyare unchanged (not used by the core API)fail_iftompi::deprecated/nccl::deprecated, only kept for the unchanged code aboveKC_CUDA_CHECK/KC_NCCL_CHECKlibrary macrosScopedRegionfor profiling regionsMPI_ERRORS_RETURNintest_main, error checks in send/recv and broadcast testsNotSupportedtests (broadcast, allgather, allreduce, alltoall)Checklist