Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 9 additions & 3 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,13 +16,19 @@ after its public API and release process are established.

### Added

- Metal Gaussian compute/image tests now consume a backend-private persistent
attribute store. Camera and metadata changes reuse buffers; matching revision
ranges upload only edited attributes, with GPU version copies protecting
in-flight readers. Transactional updates, source/handle identity, SH layout
changes and a live-byte budget have Apple GPU coverage. Renderer scheduling
and telemetry integration remain open; GPU mode support is unchanged.

- Metal Gaussian gather packs GPU-sorted records into the scalar raster stream
and writes indirect draw arguments, including empty-frame resets. The Apple
GPU harness runs preparation through indirect raster in one submission and
compares color/depth/IDs with CPU-prepared direct draws; ABI, artifact identity
and install checks include the new kernel. Renderer integration and persistent
Gaussian attribute residency remain open; `prefer` still falls back and
`require` rejects.
and install checks include the new kernel. Renderer integration remains open;
`prefer` still falls back and `require` rejects.

- Shared Slang Gaussian projection, covariance, SH, culling/compaction and
deterministic radix-sort kernels now compile into the Metal Gaussian library.
Expand Down
1 change: 1 addition & 0 deletions backend/merlin-metal/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@ find_library(MERLIN_COREGRAPHICS_FRAMEWORK CoreGraphics REQUIRED)

add_library(merlin-metal STATIC
src/backend.mm
src/gaussian_residency.mm
src/resource_table.cpp
)
merlin_target_defaults(merlin-metal)
Expand Down
93 changes: 93 additions & 0 deletions backend/merlin-metal/src/gaussian_residency.hpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
#pragma once

#import <Metal/Metal.h>

#include <merlin/extraction/frame_snapshot.hpp>

#include <atomic>
#include <memory>
#include <vector>

namespace merlin::metal {

// Backend-private immutable attribute versions. The byte budget includes old
// versions and staging retained by unfinished command buffers, not just the
// current scene. Calls are externally serialized on one Metal command queue.
class GaussianResidency {
struct Budget {
std::atomic<std::uint64_t> live{};
std::uint64_t limit{};
};

public:
struct Buffer {
Buffer() = default;
Buffer(const Buffer&) = delete;
Buffer& operator=(const Buffer&) = delete;
id<MTLBuffer> metal = nil;
std::shared_ptr<Budget> budget;
~Buffer();
};
using BufferPtr = std::shared_ptr<const Buffer>;

struct Resource {
extraction::GaussianRecord record;
BufferPtr positions;
BufferPtr covariances;
BufferPtr opacities;
BufferPtr radiance;
};

struct Update {
// Ordered by the complete generation-bearing resource handle for stable
// frame-wide tie breaking. Metadata belongs to this immutable snapshot.
std::vector<Resource> resources;
std::uint64_t upload_bytes{};
std::uint64_t upload_range_count{};
std::uint64_t device_copy_bytes{};
std::uint64_t allocation_count{};

private:
friend class GaussianResidency;
struct Copy {
BufferPtr source;
BufferPtr destination;
std::uint64_t destination_offset{};
std::uint64_t bytes{};
};
std::vector<Copy> copies;
std::shared_ptr<Budget> owner;
std::uint64_t epoch{};
std::uint64_t source_id{};
bool encoded{};
bool committed{};
};

GaussianResidency(id<MTLDevice> device, std::uint64_t byte_budget);
GaussianResidency(const GaussianResidency&) = delete;
GaussianResidency& operator=(const GaussianResidency&) = delete;
// Preparation is transactional: allocation/validation failure, or dropping
// an unsubmitted update, leaves the resident scene untouched.
std::shared_ptr<Update> Prepare(const extraction::FrameSnapshot& snapshot);
// Encode before preparation kernels; automatically retain the update and
// all its buffers until completion, even if the store/frame is destroyed.
void Encode(const std::shared_ptr<Update>& update, id<MTLCommandBuffer> command);
// Call only after committing that command to the same serial queue. An
// upload failure invalidates dependent submissions: the owner must Reset
// and report the GPU failure before using residency again.
void Commit(const std::shared_ptr<Update>& update);
void Reset();
[[nodiscard]] std::uint64_t live_bytes() const noexcept;

private:
BufferPtr Allocate(std::uint64_t bytes, MTLResourceOptions options, Update& update);
void ValidateUpdate(const std::shared_ptr<Update>& update) const;

id<MTLDevice> device_;
std::shared_ptr<Budget> budget_;
std::vector<Resource> resident_;
std::uint64_t source_id_{};
std::uint64_t epoch_{};
};

} // namespace merlin::metal
192 changes: 192 additions & 0 deletions backend/merlin-metal/src/gaussian_residency.mm
Original file line number Diff line number Diff line change
@@ -0,0 +1,192 @@
#include "gaussian_residency.hpp"

#include <merlin/render/backend.hpp>

#include <algorithm>
#include <cstring>
#include <limits>
#include <type_traits>
#include <unordered_map>

namespace merlin::metal {
namespace {
[[noreturn]] void Fail(render::RendererErrorCode code, const char* message) {
throw render::RendererError(code, "synchronize Metal Gaussian attributes", message);
}

void Validate(const extraction::GaussianRecord& record) {
if (!record.positions || !record.covariances || !record.opacities ||
!record.spherical_harmonics_coefficients || record.spherical_harmonics_degree > 3 ||
record.positions->size() > std::numeric_limits<std::uint32_t>::max() ||
record.positions->size() != record.covariances->size() ||
record.positions->size() != record.opacities->size()) {
Fail(render::RendererErrorCode::InvalidRequest, "Gaussian attribute payload is malformed");
}
const auto coefficients = (record.spherical_harmonics_degree + 1U) *
(record.spherical_harmonics_degree + 1U);
if (record.spherical_harmonics_coefficients->size() != record.positions->size() * coefficients) {
Fail(render::RendererErrorCode::InvalidRequest, "Gaussian SH payload size is inconsistent");
}
for (const auto& range : record.particle_ranges) {
if (range.first > record.positions->size() ||
range.count > record.positions->size() - range.first) {
Fail(render::RendererErrorCode::InvalidRequest, "Gaussian changed range is out of bounds");
}
}
}
} // namespace

GaussianResidency::Buffer::~Buffer() {
if (metal != nil) budget->live.fetch_sub(metal.length, std::memory_order_relaxed);
}

GaussianResidency::GaussianResidency(id<MTLDevice> device, std::uint64_t byte_budget)
: device_(device), budget_(std::make_shared<Budget>()) {
if (!device) Fail(render::RendererErrorCode::InvalidRequest, "Metal device is null");
budget_->limit = byte_budget;
}

std::uint64_t GaussianResidency::live_bytes() const noexcept {
return budget_->live.load(std::memory_order_relaxed);
}

GaussianResidency::BufferPtr GaussianResidency::Allocate(
std::uint64_t bytes, MTLResourceOptions options, Update& update) {
// Even empty attributes have a legal binding; byte-addressed shaders use
// 32-bit offsets. Reject before allocation or narrowing to NSUInteger.
bytes = std::max(bytes, std::uint64_t{16});
const auto live = live_bytes();
if (bytes > std::numeric_limits<std::uint32_t>::max() || bytes > device_.maxBufferLength)
Fail(render::RendererErrorCode::Unsupported, "Gaussian attribute exceeds the Metal buffer or shader address limit");
if (live > budget_->limit || bytes > budget_->limit - live)
Fail(render::RendererErrorCode::ResourceExhausted, "Gaussian residency exceeds the live-byte budget");
auto result = std::make_shared<Buffer>();
result->budget = budget_;
result->metal = [device_ newBufferWithLength:bytes options:options];
if (!result->metal) Fail(render::RendererErrorCode::BackendFailure, "Metal attribute buffer allocation failed");
budget_->live.fetch_add(result->metal.length, std::memory_order_relaxed);
++update.allocation_count;
return result;
}

std::shared_ptr<GaussianResidency::Update> GaussianResidency::Prepare(
const extraction::FrameSnapshot& snapshot) {
auto update = std::make_shared<Update>();
update->owner = budget_;
update->epoch = epoch_;
update->source_id = snapshot.source_id;
std::unordered_map<std::uint64_t, const Resource*> previous;
if (snapshot.source_id == source_id_) {
for (const auto& resource : resident_) previous.emplace(resource.record.gaussian, &resource);
}
update->resources.reserve(snapshot.gaussians.size());
for (const auto& record : snapshot.gaussians) {
Validate(record);
const auto found = previous.find(record.gaussian);
const auto* old = found == previous.end() ? nullptr : found->second;
const auto coefficients = (record.spherical_harmonics_degree + 1U) *
(record.spherical_harmonics_degree + 1U);
const bool partial = old && snapshot.source_id != 0 &&
old->record.revision == record.particle_base_revision &&
old->record.positions->size() == record.positions->size() &&
old->record.spherical_harmonics_degree == record.spherical_harmonics_degree &&
std::any_of(record.particle_ranges.begin(), record.particle_ranges.end(),
[](const auto& range) { return range.count != 0; });
const auto attribute = [&](const auto& payload, const auto& previous_payload,
std::uint64_t revision, std::uint64_t previous_revision,
BufferPtr previous_buffer, std::uint32_t multiplier) -> BufferPtr {
if (previous_buffer && revision == previous_revision && payload == previous_payload)
return previous_buffer;
using Element = typename std::decay_t<decltype(*payload)>::value_type;
const auto bytes = std::uint64_t{payload->size()} * sizeof(Element);
auto destination = Allocate(bytes, MTLResourceStorageModePrivate, *update);
const auto stage = [&](std::uint64_t first, std::uint64_t count) {
const auto size = count * sizeof(Element);
if (!size) return;
auto staging = Allocate(size, MTLResourceStorageModeShared, *update);
std::memcpy(staging->metal.contents, payload->data() + first, size);
update->copies.push_back({staging, destination, first * sizeof(Element), size});
update->upload_bytes += size;
++update->upload_range_count;
};
if (partial && previous_buffer && bytes) {
// Version on the GPU before patching: never overwrite a buffer that an
// earlier command may still read, and never copy unchanged data on CPU.
update->copies.push_back({previous_buffer, destination, 0, bytes});
update->device_copy_bytes += bytes;
for (const auto& range : record.particle_ranges)
stage(std::uint64_t{range.first} * multiplier, std::uint64_t{range.count} * multiplier);
} else {
stage(0, payload->size());
}
return destination;
};
Resource next;
next.record = record;
next.positions = attribute(record.positions, old ? old->record.positions : nullptr,
record.positions_revision, old ? old->record.positions_revision : 0,
old ? old->positions : nullptr, 1);
next.covariances = attribute(record.covariances, old ? old->record.covariances : nullptr,
record.covariance_revision, old ? old->record.covariance_revision : 0,
old ? old->covariances : nullptr, 1);
next.opacities = attribute(record.opacities, old ? old->record.opacities : nullptr,
record.opacity_revision, old ? old->record.opacity_revision : 0,
old ? old->opacities : nullptr, 1);
next.radiance = attribute(record.spherical_harmonics_coefficients,
old ? old->record.spherical_harmonics_coefficients : nullptr,
record.radiance_revision, old ? old->record.radiance_revision : 0,
old ? old->radiance : nullptr, coefficients);
update->resources.push_back(std::move(next));
}
std::sort(update->resources.begin(), update->resources.end(),
[](const auto& a, const auto& b) { return a.record.gaussian < b.record.gaussian; });
for (std::size_t i = 1; i < update->resources.size(); ++i) {
if (update->resources[i - 1].record.gaussian == update->resources[i].record.gaussian)
Fail(render::RendererErrorCode::InvalidRequest, "Duplicate Gaussian resource handle");
}
return update;
}

void GaussianResidency::ValidateUpdate(const std::shared_ptr<Update>& update) const {
if (!update || update->owner != budget_ || update->epoch != epoch_ || update->committed)
Fail(render::RendererErrorCode::InvalidRequest, "Stale or foreign Gaussian residency update");
}

void GaussianResidency::Encode(const std::shared_ptr<Update>& update, id<MTLCommandBuffer> command) {
ValidateUpdate(update);
if (!command || command.device != device_ || update->encoded ||
command.status != MTLCommandBufferStatusNotEnqueued || !command.retainedReferences)
Fail(render::RendererErrorCode::InvalidRequest, "Expected an unsubmitted retaining Metal command buffer");
if (!update->copies.empty()) {
auto encoder = [command blitCommandEncoder];
if (!encoder) Fail(render::RendererErrorCode::BackendFailure, "Metal attribute blit encoder allocation failed");
for (const auto& copy : update->copies) {
[encoder copyFromBuffer:copy.source->metal sourceOffset:0
toBuffer:copy.destination->metal destinationOffset:copy.destination_offset size:copy.bytes];
}
[encoder endEncoding];
}
// Capturing the plan keeps allocation accounting and old versions alive as
// long as the GPU uses them; the plan itself does not retain the command.
__block auto retained = update;
[command addCompletedHandler:^(id<MTLCommandBuffer>) { retained.reset(); }];
update->encoded = true;
}

void GaussianResidency::Commit(const std::shared_ptr<Update>& update) {
ValidateUpdate(update);
if (!update->encoded) Fail(render::RendererErrorCode::InvalidRequest, "Gaussian update has not been encoded");
auto resources = update->resources;
resident_.swap(resources);
source_id_ = update->source_id;
update->committed = true;
++epoch_;
}

void GaussianResidency::Reset() {
resident_.clear();
source_id_ = 0;
++epoch_;
}

} // namespace merlin::metal
28 changes: 26 additions & 2 deletions docs/design/metal-gaussian-execution.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,8 +88,8 @@ projection/sort policies, transforms, multiple resources, hidden/all-rejected
input, workgroup boundaries, asymmetric alpha composition and a prefilled opaque
depth attachment. They do not establish general scene or host-presentation
parity. This remains a correctness harness, not renderer GPU execution:
persistent attributes, completion-safe frame scheduling and telemetry
integration remain Phase 2 work. Controlled performance captures
renderer integration of persistent attributes, completion-safe frame scheduling
and telemetry remain Phase 2 work. Controlled performance captures
and updated native viewport/HgiMetal comparisons remain unfinished; this is
not completion of the phase gate below.

Expand Down Expand Up @@ -130,6 +130,30 @@ update metadata. Replaced buffers and slots remain alive until their last GPU
consumer completes. Bound resident and scratch allocation and report failures
through the existing diagnostic/fallback contract.

The private Metal attribute store now supplies the compute/image harness with
immutable device-local position, covariance, opacity and SH buffers. It keys
reuse by source identity, the complete resource handle, attribute revisions and
shared payload identity. Camera, transform and visibility changes reuse those
buffers. Matching particle-base revisions and unchanged layouts permit staged
range updates; a GPU copy creates the new attribute version before patching it,
so earlier submissions can continue reading the old version. Revision gaps,
source changes, count changes and SH layout changes upload complete affected
attributes. Unchanged attributes retain their original buffers.

Preparation is transactional, and a separate commit publishes the resident
scene only after submission. Command completion retains staging and old
versions, including their contribution to an explicit live-byte budget.
Allocation failure leaves the previous scene usable. The caller must invalidate
residency and report a failed upload command before scheduling dependent work.
The harness checks actual bytes, partial SH ranges, removal/generation reuse,
abandoned/invalid updates, budget exhaustion and multiple blocked submissions.
Continuous image comparisons check static, camera, transform, localized edits,
visibility, removal and reintroduction with the existing color/depth/ID tolerance.
This is a backend-private building block: renderer scheduling, scratch reuse,
error recovery and public telemetry integration are still unfinished. It scans
resource metadata, and partial updates currently copy the full changed attribute
on the GPU; it does not yet implement a resource-delta fast path or an arena.

Connect the complete frame path on the GPU:

```text
Expand Down
Loading
Loading