Skip to content

4.4: Multi-GPU And Distributed Processing Research #1827

Description

@kmeirlaen

Included survey bullets:

  • Multi-GPU training support.
  • Pool VRAM across GPUs in one machine.
  • Network/distributed processing with queue and dispatcher, similar to Metashape.
  • Remote or cluster processing interest appears in adjacent requests.

Suggested solution:

Keep this as a research track until block training proves the partition/merge model. Multi-GPU and distributed processing are larger architecture commitments and should not block v0.6's first bigger-scene win.

Actionable scope:

  • Document trainer assumptions that prevent multi-GPU.
  • Identify whether data parallel, model parallel, or block parallel is most realistic.
  • Evaluate using block training as the basis for distributed execution.
  • Define what would be shared between machines: source data, trained blocks, previews, progress.

Acceptance signal:

  • A technical design exists for either local multi-GPU or distributed block processing, with known risks and implementation phases.

Dependencies:

  • Topic 4.1 block training.
  • Runtime job system.
  • Project/session serialization.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions