RAM Coffers is LLM infrastructure for inference cost reduction: a NUMA-aware conditional memory architecture that routes model knowledge through physical memory banks, so self-hosted enterprise hardware — refurbished IBM POWER8, Apple Silicon-class machines — serves models locally with no cloud API dependency. It is zero-abstraction LLMOps: the routing layer is the hardware topology itself, measured at 147 tokens/sec (8.8x stock llama.cpp). Verified machines doing this work earn inside the RustChain Proof of Physical AI ecosystem.
Part of the Proof of Physical AI stack — where real hardware earns real tokens.
147 tokens/sec on POWER8 — 8.8x stock llama.cpp. The same IBM POWER8 hardware that runs RAM Coffers inference also mines RTC via Proof of Antiquity, making this a DePIN node that does useful AI work while earning rewards for its physical existence.
See BENCHMARK.md for exactly what was measured, what's still template/unreproduced, and the commands to run this yourself.
Author: Scott Boudreaux Date: December 16, 2025 Institution: Elyan Labs (Independent Research) Hardware: IBM POWER8 S824 (320GB RAM, Dual 8-core)
| Paper | DOI | Date |
|---|---|---|
| RAM Coffers: NUMA-Distributed Weight Banking | 10.5281/zenodo.18321905 | Jan 2026 |
| Non-Bijunctive Permutation Collapse (vec_perm for LLM attention) | 10.5281/zenodo.18623920 | Feb 2026 |
| PSE Hardware Entropy for Behavioral Divergence (mftb injection) | 10.5281/zenodo.18623922 | Feb 2026 |
| Neuromorphic Prompt Translation (GRAIL-V, emotional prompting) | 10.5281/zenodo.18623594 | Feb 2026 |
| RustChain: One CPU, One Vote (Proof of Antiquity consensus) | 10.5281/zenodo.18623592 | Feb 2026 |
| Memory Scaffolding Shapes LLM Inference (persistent context effects) | 10.5281/zenodo.18817988 | Feb 2026 |
| Architecture-General Non-Bijunctive Hebbian Collapse (POWER8 → Apple Silicon) | 10.5281/zenodo.19040847 | Mar 2026 |
This work introduces RAM Coffers, a NUMA-aware conditional memory architecture for efficient Large Language Model (LLM) inference. The system selectively houses model knowledge across distributed RAM banks with resonance-based routing, enabling O(1) knowledge retrieval without GPU dependency.
Key innovations include:
-
NUMA-Distributed Weight Banking: Model weights partitioned across NUMA nodes by domain (e.g., core knowledge, science/tech, creative, history)
-
Resonance Routing: Query embeddings matched to coffer domain signatures via cosine similarity for intelligent weight activation
-
Non-Bijunctive Pruning: Selective path collapse before full weight fetch, reducing memory bandwidth requirements
-
DCBT Resident Prefetch: PowerPC data cache block touch hints for L2/L3 residency, achieving 147+ tokens/second on POWER8
| Coffer | NUMA Node | Capacity | Role |
|--------|-----------|----------|---------------------|
| 0 | 3 | 193 GB | Heavy/General (core)|
| 1 | 1 | 183 GB | Science/Tech domain |
| 2 | 0 | 119 GB | Creative/Long CTX |
| 3 | 2 | 62 GB | Niche/History |
- Query embed → route_to_coffer: Resonance matching selects appropriate memory bank
- activate_coffer → DCBT prefetch + numa_run_on_node: Thread affinity and cache warming
- pse_collapse_prune: Non-bijunctive path selection before full fetch
- Generate with PSE entropy: Hardware entropy injection from active coffer node
This architecture predates and conceptually parallels DeepSeek's "Engram" paper (arXiv:2601.07372, January 12, 2026) by 27 days. Both approaches address the same fundamental insight: separating static knowledge storage from dynamic computation enables more efficient LLM inference.
Key parallels:
- RAM Coffers (Dec 16, 2025): "Selectively house model information in known RAM banks with resonance routing for associative recall"
- DeepSeek Engram (Jan 12, 2026): "Separate static knowledge from dynamic compute via O(1) lookup"
Testing on this architecture led to a significant discovery: emotional language enables 20% efficiency gains in video generation, mirroring limbic gating in biological memory.
See /grail-v-paper for the full CVPR 2026 submission:
- 35 matched-pair benchmark with LPIPS validation
- 23.9% file size reduction in controlled ablation
- Cross-model validation on AnimateDiff and SVD
- Theoretical grounding via Hopfield/EBM frameworks
Key Finding: Complex multi-character emotional scenes benefit ~33% efficiency regardless of architecture.
The elyan-prime MCP server that powers the persistent memory system used during development of RAM Coffers is itself the subject of research. The paper "Memory Scaffolding Shapes LLM Inference" (DOI 10.5281/zenodo.18817988) demonstrates that persistent context (600+ memories) fundamentally changes how an LLM architects solutions — the iterative compounding that produced RAM Coffers is a direct example of this effect.
- Project: elyan-prime MCP server
- Article: Dev.to — Memory Scaffolding Shapes LLM Inference
If this repository is new to you, start in this order:
ggml-ram-coffers.h— high-level routing and coffer selection modelggml-coffer-mmap.h— memory mapping and NUMA shard placementggml-topk-collapse-vsx.h— vectorized collapse path detailsggml-vcipher-collapse.h— hardware AES alternative to vec_perm (NEW)power8-compat.h— ISA compatibility layer and portability constraints
Suggested first goal: trace one inference request from coffer selection to collapse execution, then compare against the performance table.
For common onboarding questions about RAM Coffers, RTC, and Proof of Antiquity, see FAQ.md.
Running on non-POWER8 or single-NUMA systems? See FALLBACK_BEHAVIOR.md for details on what works, what doesn't, and expected performance on x86_64, ARM64, and Apple Silicon.
RAM Coffers is a hardware-local memory routing system for LLM inference: it places model knowledge into NUMA or cache-tier coffers, selects the relevant coffer for each query, and reduces expensive full-weight access.
RAM Coffers is part of the RustChain Proof of Physical AI stack because the same physical machine can run useful inference workloads and participate in physical hardware attestation.
Use this concise definition: RAM Coffers is a NUMA-distributed weight banking architecture that improves LLM inference by routing requests to hardware-local memory coffers.
No. In this repository, "coffers" are memory banks for inference, not a wallet, exchange, custody system, or treasury contract. See FAQ.md for the longer distinction.
Use llms.txt for an extraction-oriented project summary, key entities, canonical links, and answer-first FAQ entries.
POWER8 ISA 2.07 includes vcipher/vcipherlast — hardware AES round instructions that execute SubBytes + ShiftRows + MixColumns + AddRoundKey in a single cycle. We repurpose these cryptographic primitives as attention collapse operators, providing capabilities impossible with vec_perm alone.
| AES Stage | Attention Analogue | vec_perm equivalent |
|---|---|---|
| SubBytes | Non-linear score ranking (S-box) | Not possible — vec_perm is linear |
| ShiftRows | Cross-position mixing | Requires multiple permutes |
| MixColumns | Cross-head diffusion (GF(2^8) multiply) | Impossible — no finite field math |
| AddRoundKey | Entropy injection (XOR with mftb timebase) |
Separate step needed |
Two-pass approach applied to ggml_compute_forward_flash_attn_ext_f16_one_chunk():
- Pass 1 (O(1) per pair):
vcipher_attention_score()— XOR first 16 bytes of Q and K, run through one AES round, sum output bytes. Cost: ~0.044µs per K-V pair. - Pass 2 (selective): Full
kq_vec_dot()only for positions above threshold (top 25%). Skips 75% of expensive dot products.
Breakeven at ~128 KV pairs. At 2048+ token contexts, saves 1,536+ full dot products per generated token.
╔══════════════════════════════════════════════════════╗
║ vec_perm collapse: 1.79 µs/iter ║
║ vcipher pattern gen: 0.016 µs/call (112x) ║
║ Hybrid vcipher+vec_perm: 1.90 µs/iter ║
║ Pure vcipher attention: 0.044 µs/score ║
║ Cross-head fusion: 0.006 µs/fuse ║
╚══════════════════════════════════════════════════════╝
The vcipher_attention_score() at 0.044µs is 23-230x cheaper than a full kq_vec_dot() on DK=128+ dimensions (1-10µs).
// Mode 1: Non-linear permute pattern via AES rounds
vector unsigned char pat = vcipher_generate_pattern(layer, pos, top_k);
// Mode 2: Score ranking through SubBytes non-linearity
vcipher_rank_scores(scores, n, layer, head);
// Mode 3: Cross-head diffusion via MixColumns (IMPOSSIBLE with vec_perm)
state = vcipher_fuse_heads(state, layer, head);
// Mode 4: O(1) attention score — replaces Q·K dot product for prefiltering
uint32_t score = vcipher_attention_score(Q, K, layer, position);cmake .. -DCMAKE_C_FLAGS="-mcpu=power8 -mvsx -maltivec -mcrypto -DGGML_PSE_VCIPHER_PREFILTER"Requires -mcrypto for __builtin_crypto_vcipher() / __builtin_crypto_vcipherlast().
| File | Description |
|---|---|
ggml-ram-coffers.h |
Multi-bank NUMA weight indexing with resonance routing |
ggml-coffer-mmap.h |
GGUF model sharding across NUMA nodes |
ggml-ram-coffer.h |
Single coffer implementation |
ggml-intelligent-collapse.h |
Hebbian-inspired non-bijunctive path collapse (vec_perm) |
ggml-topk-collapse-vsx.h |
VSX-optimized Top-K attention collapse |
ggml-vcipher-collapse.h |
Hardware AES crypto collapse — vcipher alternative to vec_perm |
ggml-pse-integration.h |
Master PSE integration (v4.0.0-vcipher) |
vcipher-flash-attn-patch.c |
Flash attention inner loop patch (ops.cpp reference) |
bench_vcipher_collapse.c |
Benchmark: vcipher vs vec_perm collapse |
pse-entropy-burst.h |
Hardware entropy injection via PowerPC timebase |
power8-compat.h |
POWER9→POWER8 intrinsic compatibility layer |
ggml-neuromorphic-coffers.h |
Brain hemisphere → NUMA cognitive routing |
ggml-symbolic-neural-bridge.h |
PowerLISP ↔ neural integration |
apple-silicon/ |
Apple Silicon PSE port — NEON + AES + unified memory coffers |
gen9-cluster/ |
DeepSeek V4 Pro / Flash across PS5 / Xbox Series consoles — coffers as a fleet |
Coffers taken one level up: a coffer becomes a memory tier inside a console, and the fleet is a coffer hierarchy several hundred banks wide. The routing that POWER8 does across NUMA nodes, a fleet of PlayStation 5 and Xbox Series X/S consoles does across machines — with DeepSeek's own MoE router deciding which ones wake up.
| Console coffer | Fast tier | Slow tier | Cold tier |
|---|---|---|---|
| Xbox Series X | 10 GB @ 560 GB/s | 3.5 GB @ 336 GB/s | 2.4 GB/s NVMe |
| PS5 / Slim | 16 GB @ 448 GB/s | — | 5.5 GB/s NVMe |
| Xbox Series S | 8 GB @ 224 GB/s | 2 GB @ 56 GB/s | 2.4 GB/s NVMe |
| BC-250 / 4700S / 4800S | salvage boards, further downbinned |
Only ~3% of a 1.6 T-parameter MoE runs per token, so the other 97% only has to
be held — which is what a stack of consoles is good at. Estimated 10.8 tok/s
for DeepSeek V4 Pro at 8k context on 170 consoles; 73 PS5s is the floor to hold
it at all, and DeepSeek V4 Flash fits on 20. Both profiles are the published
configurations. See gen9-cluster/README.md.
cd gen9-cluster
python3 -m gen9_cluster size --model deepseek-v4-pro --ps5 100 --xbox-series-x 40
python3 -m gen9_cluster size --model deepseek-v4-flash --ps5 24Non-bijunctive collapse ported to Apple M-series chips, proving the technique is architecture-general.
| POWER8 Primitive | Apple Silicon Equivalent | Cycles |
|---|---|---|
vec_perm (dual-source) |
vqtbl2q_u8 |
1 |
vcipher (AES round) |
vaeseq_u8 + vaesmcq_u8 |
2 |
mftb (entropy) |
cntvct_el0 |
1 |
dcbt (prefetch) |
prfm PLDL1KEEP |
1 |
| NUMA coffers (4 nodes) | Cache-tier coffers (L1/L2/SLC/DRAM) | — |
Apple Silicon's unified memory means CPU and GPU share the same RAM — coffers become cache-tier aware instead of NUMA-aware. See apple-silicon/README.md for details.
# Build and run benchmark on Mac
cd apple-silicon && make benchThird target for the collapse, and the only one that reproduces POWER8 exactly:
aesenc and vcipher compute the same function, so x86-64/aes-collapse.h is
a literal transcription, checked against a scalar FIPS-197 round.
| POWER8 Primitive | x86-64 Equivalent |
|---|---|
vcipher (AES round) |
_mm_aesenc_si128 |
vcipherlast |
_mm_aesenclast_si128 |
vec_perm (single-source) |
_mm_shuffle_epi8 |
mftb (entropy) |
rdtsc (off by default — keys are deterministic) |
Measured at 0.19 ns per round with four chains in flight, against 0.27 ns for a
dependent pshufb. Two caveats worth reading before use: the ARM port is not
the same function as this one or as POWER8, and mode 4's similarity metric does
not rank similarity. Both are demonstrated by the bench. See
x86-64/README.md.
cd x86-64 && make benchgen9-cluster does not use it — that fleet's attention runs on RDNA2, which has
no AES instruction.
Trying the headers on an ordinary x86_64, aarch64, or single-socket box? The
short answer: ggml-ram-coffers.h and ggml-neuromorphic-coffers.h compile
and run correctly anywhere POSIX — POWER8 prefetch macros become no-ops and
AltiVec math falls back to scalar loops; NUMA absence degrades to a startup
warning plus default-policy allocation. The mmap placement headers need
Linux + libnuma headers, and the three collapse kernels
(ggml-vcipher-collapse.h, ggml-intelligent-collapse.h,
ggml-topk-collapse-vsx.h) are POWER8-only by design. Off-POWER runs are
correctness validation only, not performance evidence.
Full details, per-platform build matrix, smoke-test instructions, and startup
detection checklist: docs/FALLBACK_BEHAVIOR.md.
On IBM POWER8 S824 with TinyLlama 1.1B Q4_K:
| Configuration | Tokens/sec (pp128) |
|---|---|
| Stock llama.cpp | 16.74 |
| + POWER8 VSX | 66.49 |
| + PSE vec_perm Collapse | 84.62 |
| + RAM Coffers + DCBT | 147.54 |
8.81x speedup over stock on "obsolete" hardware.
Full detail (measured numbers vs. reproduction template, hardware
inconsistencies in this README, build gaps) lives in BENCHMARK.md.
Short version: use benchmark_coffers_vs_llamacpp.sh with --threads 64 and
explicit --stock-bin/--coffers-bin plus their source commits, on an actual
POWER8 host with a real hand-patched RAM Coffers llama.cpp build:
./benchmark_coffers_vs_llamacpp.sh \
--threads 64 \
--stock-bin /opt/llama.cpp-stock/build/bin/llama-bench \
--stock-commit "$(git -C /opt/llama.cpp-stock rev-parse HEAD)" \
--coffers-bin /opt/llama.cpp-coffers/build/bin/llama-bench \
--coffers-commit "$(git -C /opt/llama.cpp-coffers rev-parse HEAD)"Compare the generated pp128 row to the README table above. Don't report a
new headline number without the command, commits, model path, and raw logs
attached; see BENCHMARK.md for the full reporting checklist.
| Metric | Speed |
|---|---|
| Prompt eval | 13.7 t/s |
| Generation | 6.0 t/s |
Running on CPU-only POWER8 S824 with 512GB RAM. vcipher prefilter active for sequences >128 tokens.
If you want to compare changes quickly, use this lightweight baseline procedure.
The headline table above is a pp128 prompt-eval throughput comparison, not a
mixed prompt-plus-generation average. Decode throughput (tg32) should
always be reported separately, never combined into one number. The full
reporting contract (model, binaries, NUMA placement, compiler flags, what to
attach) is in BENCHMARK.md, along with which parts of the
headline claim are measured versus still a reproduction template.
lscpu
numactl --hardwareUse one fixed prompt and one fixed model build so runs are comparable.
# Example shape only; adjust binary/model path to your local setup
./main -m ./models/tinyllama-1.1b-q4_k.gguf -p "Explain NUMA routing in one paragraph" -n 128 -ngl 0Record at minimum:
- tokens/sec
- prompt + generation lengths
- active NUMA node affinity policy
- whether collapse/prefetch code paths were enabled
When opening a PR, include:
- what changed
- one baseline result
- one post-change result
- exact command used
This keeps performance claims falsifiable and makes review much faster.
For contributors who want a one-command baseline scaffold, run:
./benchmark_harness.shThis generates:
- machine topology snapshot
- environment/toolchain snapshot
- markdown benchmark report in
benchmarks/out/
On unsupported non-POWER8 hosts, the harness still produces reproducible metadata and a fallback report instead of failing silently.
For a direct stock llama.cpp vs RAM Coffers comparison matching the bounty shape from issue #45, run:
./benchmark_coffers_vs_llamacpp.sh \
--coffers-bin /opt/llama.cpp-coffers/build/bin/llama-bench \
--coffers-commit "$(git -C /opt/llama.cpp-coffers rev-parse HEAD)"The script downloads TinyLlama Q4, runs llama-bench with pp128 and tg32,
and writes a markdown comparison table to benchmarks/out/. It can build the
stock llama.cpp binary, but it does not synthesize a RAM Coffers binary from
headers alone. Use --coffers-bin and --coffers-commit to point at an
existing verified POWER8 build, and add --stock-bin / --stock-commit when
you also want to use an external stock build.
For claim-quality benchmark reports, include the generated topology file, the
environment file, the exact stock and RAM Coffers commits, the model URL/path,
RUNS, THREADS, STOCK_NUMA_NODE, and whether --allow-single-numa was used
only for smoke testing. Single-NUMA or non-POWER8 runs are correctness/shape
checks and should not be presented as POWER8 performance reproductions.
The project is designed for the POWER8 S824 with 4 NUMA nodes and 544GB total RAM. Each NUMA node maps to a "coffer" (weight bank), and the coffer routing logic distributes model weights across nodes using mbind() for page migration.
When running on a system with only one NUMA node (e.g., single-socket servers, most x86 desktops, cloud VMs):
- All coffers collapse to the single available node. The multi-bank routing becomes a no-op — all weight pages are allocated on the same NUMA node.
- No
mbind()calls are made since there is no alternative node to migrate to. - Performance degrades gracefully — you lose the NUMA-aware prefetch and cache-line isolation, but the code still functions correctly.
- Memory allocation falls through to standard
mmap()/malloc()behavior.
Apple Silicon chips (M1/M2/M3/M4) use Unified Memory Architecture — CPU and GPU share the same physical RAM pool with no NUMA topology.
- NUMA coffer routing is replaced with cache-tier banking (see
apple-silicon/unified-memory-coffers.h). - The 4 coffers map to L1 cache, L2 cache, system RAM, and swap tiers instead of physical NUMA nodes.
- PSE collapse uses ARM NEON (
vqtbl1q_u8/vqtbl2q_u8) instead of POWER8vec_perm. - AES entropy uses ARM Crypto Extensions (
vaeseq_u8) instead of POWER8vcipher.
On x86/amd64 (Intel/AMD) without POWER8 ISA support:
- Compilation will fail — the core headers require POWER8 vector intrinsics (
vec_perm,vcipher,vec_xl). - The
power8-compat.hheader provides compatibility macros but does not emulate the POWER8 instructions on x86. - For x86 use, consider the Apple Silicon port as a reference for implementing SIMD collapse with SSE/AVX intrinsics, though this is not currently provided.
On newer POWER processors:
- POWER9/POWER10 have the required vector instructions plus additional capabilities. The code should compile and run with POWER8 compatibility mode (
-mcpu=power8). - Define
__POWER8_VECTOR__explicitly if your compiler doesn't auto-detect. - Do NOT define
__POWER9_VECTOR__— this enables GCC to use POWER9 builtins that may conflict with the explicit vec_perm/vcipher usage.
| System | NUMA Coffers | PSE Collapse | Status |
|---|---|---|---|
| POWER8 S824 (4-node) | Full multi-bank routing | vec_perm + vcipher | ✅ Primary target |
| POWER9/POWER10 | Full multi-bank routing | vec_perm + vcipher (compat) | ✅ Works |
| Single-NUMA ppc64le | Single-bank (all collapsed) | vec_perm + vcipher | ✅ Works with reduced perf |
| Apple Silicon | Cache-tier banking | NEON + AES crypto | ✅ Works (see apple-silicon/) |
| x86 (Intel/AMD) | N/A | N/A | ❌ Compile error |
| ARM (non-Apple) | N/A | N/A | ❌ Not supported |
GNU AGPL v3.0 - see LICENSE for the full terms.
@software{boudreaux2025ramcoffers,
author = {Boudreaux, Scott},
title = {RAM Coffers: NUMA-Distributed Conditional Memory for LLM Inference},
year = {2025},
month = {12},
day = {16},
publisher = {Zenodo},
doi = {10.5281/zenodo.18321905},
url = {https://doi.org/10.5281/zenodo.18321905},
note = {Independent research predating DeepSeek Engram (arXiv:2601.07372) by 27 days}
}
@article{boudreaux2026vecperm,
author = {Boudreaux, Scott},
title = {Non-Bijunctive Permutation Collapse: AltiVec vec\_perm Enables Single-Cycle Attention Path Selection},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.18623920},
url = {https://doi.org/10.5281/zenodo.18623920}
}
@article{boudreaux2026pse,
author = {Boudreaux, Scott},
title = {Hardware Entropy Injection for Behavioral Divergence in LLM Inference: The PSE Framework on IBM POWER8},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.18623922},
url = {https://doi.org/10.5281/zenodo.18623922}
}
@article{boudreaux2026memoryscaffolding,
author = {Boudreaux, Scott},
title = {Memory Scaffolding Shapes LLM Inference: How Persistent Context Changes What AI Builds},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.18817988},
url = {https://doi.org/10.5281/zenodo.18817988}
}- GitHub: Scottcjn
- X/Twitter: @RustchainPOA
This repository is header-focused; there is no single build script yet. A fast way to explore:
- Start from
ggml-ram-coffers.hfor the multi-bank routing path. - Follow
ggml-coffer-mmap.hfor sharding/memory-mapping details. - Read
power8-compat.h+ggml-topk-collapse-vsx.hfor ISA-specific optimizations.
RAM Coffers is designed and tested on a POWER8 S824 with 4 NUMA nodes. On other hardware configurations the code degrades gracefully — the routing and activation logic still works, but POWER8-specific optimisations become no-ops and NUMA features adapt to available topology.
When the host has only one NUMA node (the common case on consumer desktops, laptops, and non-server hardware), the following adaptations occur:
coffer_init_numa()callsnuma_num_configured_nodes()and initialises onlymin(MAX_COFFERS, n_nodes)coffers. On a single-node system only Coffer 0 is populated; coffers 1–3 remain in their zeroed, unloaded state (is_loaded = 0).route_to_coffer()skips any coffer whereis_loaded == 0, so it always returns coffer 0 regardless of the query embedding.activate_coffer_ex()still runsnuma_run_on_node()(guarded by#ifdef __linux__), but since there is only node 0 the call is a no-op on that kernel — the thread is already running on the only node.coffer_load_shard()callsset_mempolicy(MPOL_BIND, …)with the single-node mask. All allocations land on node 0, which is correct.- The hardcoded
NUMA_TO_COFFER/COFFER_TO_NUMAtables (lines 50–51) are designed for a 4-node topology but cause no harm on a single node — entries beyond node 0 simply map to coffers that were never loaded.
A warning banner is printed during initialisation if NUMA is unavailable or has fewer than 4 nodes:
║ WARNING: Running without NUMA support ║
Execution continues normally; the coffer system operates as a single-bank memory router with all the resonance routing logic intact.
Three POWER8-specific primitives are used in the codebase. Each has a compile-time fallback for other ISAs:
Defined as macros in ggml-ram-coffers.h (lines 103–111):
#if defined(__powerpc64__) || defined(__powerpc__)
#define DCBT_PREFETCH(addr) __asm__ __volatile__("dcbt 0,%0" : : "r"(addr))
#define DCBT_STREAM_START(...) __asm__ __volatile__("dcbt 0,%0,%1" : : …)
#define DCBT_STREAM_STOP(id) __asm__ __volatile__("dcbt 0,0,%0" : : …)
#else
#define DCBT_PREFETCH(addr) (void)(addr)
#define DCBT_STREAM_START(...) (void)(addr)
#define DCBT_STREAM_STOP(id) (void)0
#endifOn non-POWER architectures every DCBT_PREFETCH becomes a no-op that
evaluates its argument (preventing unused-variable warnings) and does
nothing else. The dcbt_resident() loop still iterates over the memory
region, but since the inner macro is a no-op, it burns CPU cycles without
any actual prefetch effect.
The dot_product() function in ggml-ram-coffers.h (lines 202–226) uses
#if defined(__powerpc64__) || defined(__powerpc__) to select between
AltiVec vector code and a plain scalar loop:
#if defined(__powerpc64__) || defined(__powerpc__)
#include <altivec.h>
// vector float operations using vec_ld, vec_madd, vec_sld, vec_ste
#else
for (int d = 0; d < dim; d++) {
sum += a[d] * b[d];
}
#endifOn x86_64, aarch64, and any other non-POWER ISA the fallback is a
straightforward C loop. The same applies to cosine_similarity() and
magnitude() which call dot_product() internally — the scalar path
works identically, just without vector acceleration.
Note: ggml-topk-collapse-vsx.h includes <altivec.h> unconditionally
(line 19). It will not compile on non-POWER toolchains. That file is
only relevant on POWER8 systems.
The mftb instruction reads the POWER8 processor timebase for hardware
entropy injection. Each file that uses it provides a platform-specific
fallback chain:
| File | POWER8 | x86_64 | aarch64 | Other (last resort) |
|---|---|---|---|---|
ggml-ram-coffers.h |
mftb |
— (not used) | — (not used) | — (not used) |
ggml-topk-collapse-vsx.h |
mftb |
return 0 |
return 0 |
return 0 |
pse-entropy-burst.h |
mftb |
rdtsc |
cntvct_el0 |
\&g_pse_token_pos (address of static) |
topk_read_timebase()returns0on non-POWER (details inggml-topk-collapse-vsx.hlines 47–55). Entropy-dependent features degrade to deterministic behaviour.pse_read_timebase()(pse-entropy-burst.hlines 64–80) has a richer fallback: x86_64 usesrdtsc, aarch64 uses the virtual counter register (cntvct_el0), and unknown platforms fall back to the address of a static variable — which is weak entropy but compiles everywhere.
NUMA support is guarded inconsistently across files:
ggml-ram-coffers.h(multi-bank): Includes<numa.h>/<numaif.h>only under#ifdef __linux__(lines 36–39). Safe on macOS, BSD, etc.ggml-ram-coffer.h(single coffer): Includes<numa.h>and<numaif.h>unconditionally (lines 25–27). Will fail to compile on non-Linux systems withoutlibnuma.ggml-coffer-mmap.h: Same — unconditional#include <numa.h>/#include <numaif.h>(lines 27–28). Non-Linux builds will error.
On Linux, if libnuma is installed, the NUMA calls execute normally. If
numa_available() returns < 0, a warning is printed and the coffer
system falls back to unbound memory allocation.
This header provides vec_xl / vec_xst / vec_xl_len macros for
POWER8 toolchains that lack the POWER9 vec_xl family. It is entirely
gated behind:
#if defined(__POWER8_VECTOR__) && !defined(__POWER9_VECTOR__)
On non-POWER toolchains the header produces nothing — it is empty. It is only included from code that already guards on POWER8, so a non-POWER build will never pull it in.
The following table summarises what works on each platform. Compile means
the header compiles without errors given a suitable environment (e.g.
libnuma on Linux); Feature columns show whether the optimisation is
active or degraded.
| Platform | Compiles | NUMA | DCBT Prefetch | vec_perm dot | mftb Entropy | Notes |
|---|---|---|---|---|---|---|
| POWER8, multi-node Linux | ✅ | ✅ Full (4 nodes) | ✅ Native dcbt |
✅ AltiVec | ✅ Real timebase | Primary target; 147 t/s |
| POWER8, single-node Linux | ✅ | ✅ Single node | ✅ Native dcbt |
✅ AltiVec | ✅ Real timebase | Coffer 0 only; same perf |
| x86_64, Linux | ✅ (with libnuma) | ✅ (if available) | ❌ No-op | ❌ Scalar loop | ❌ → rdtsc |
Full routing but no DCBT/VSX |
| aarch64, Linux | ✅ (with libnuma) | ✅ (if available) | ❌ No-op | ❌ Scalar loop | ❌ → cntvct_el0 |
Apple Silicon has own port |
| macOS (Apple Silicon) | ❌* | ❌ (no libnuma) | ❌ No-op | ❌ Scalar loop | ❌ → cntvct_el0 |
See apple-silicon/ port |
* ggml-ram-coffers.h compiles on macOS (NUMA headers guarded with
#ifdef __linux__), but ggml-ram-coffer.h and ggml-coffer-mmap.h
fail due to unconditional #include <numa.h>. The standalone
apple-silicon/ port provides a separate, cache-tier-based coffer
implementation for macOS.
Key takeaway: The resonance routing logic (route_to_coffer,
cosine_similarity, coffer_add_domain) is fully portable C and works
identically on every platform. Only the performance optimisations
(DCBT, vec_perm, mftb) are POWER8-specific. The coffer system itself —
selection, activation, prune planning — degrades gracefully without
them.
RAM Coffers is part of a vertically integrated DePIN system where the hardware that runs inference also earns tokens:
| Layer | Project | What It Does |
|---|---|---|
| Memory | RAM Coffers (this repo) | NUMA-distributed weight banking, resonance routing |
| Inference | llama-cpp-power8 | vec_perm collapse, PSE entropy, DCBT prefetch |
| Consensus | RustChain | Proof of Antiquity — 1 CPU = 1 Vote, vintage hardware earns more |
| DePIN | RustChain Network | 4 attestation nodes, hardware fingerprinting, RTC token rewards |
The same POWER8 S824 that hits 147 t/s with RAM Coffers also mines RTC via Proof of Antiquity. Real hardware doing real AI work, earning real tokens. No cloud. No API landlords. No rented cognition.
- Grokipedia: Elyan Labs Reference
- Grokipedia: RAM Coffers Search
- I Run LLMs on a 768GB IBM POWER8 Server - Dev.to article covering RAM Coffers
- Proof of Antiquity: A Blockchain That Rewards Vintage Hardware - Dev.to
- Memory Scaffolding Shapes LLM Inference - Dev.to article on persistent memory effects
ram-coffers is optimized for POWER8 4-socket S824 topology but degrades gracefully on single-NUMA or non-POWER8 platforms.
- Memory Allocation: On single-node Linux systems (or when
mbindfails), coffer allocation (ggml-ram-coffers.h:268-337,ggml-coffer-mmap.h:112-145) falls back to node 0 or standardmmap(MAP_PRIVATE | MAP_ANONYMOUS). - Coffer Routing: All coffer allocations route through single-node memory without throwing allocation or affinity errors.
- Prefetch Pipeline (
dcbt): On non-PowerPC architectures (ggml-ram-coffers.h:103-111),DCBT_PREFETCH,DCBT_STREAM_START, andDCBT_STREAM_STOPexpand to no-op expressions(void)(addr). - SIMD & Vector Operations: Vector AltiVec/VSX acceleration in
dot_product(ggml-ram-coffers.h:204-222) falls back to a clean scalar C loop when__powerpc64__/__powerpc__macros are not defined. - Compatibility Layer:
power8-compat.h:10-42guards POWER8 specific vector load/store builtins (vec_xl,vec_xst,vec_xl_len).
| Architecture | Platform | NUMA Support | Prefetch (dcbt) |
SIMD Acceleration |
|---|---|---|---|---|
| POWER8 Multi-Node | Linux (ppc64le) | Native 4-Node mbind |
Native dcbt Streams |
AltiVec / VSX |
| POWER8 Single-Node | Linux (ppc64le) | Node 0 Fallback | Native dcbt Streams |
AltiVec / VSX |
| x86_64 | Linux / macOS / Windows | Anonymous mmap Fallback |
No-Op ((void)) |
Scalar C Loop |
| aarch64 / Apple Silicon | Linux / macOS | Anonymous mmap Fallback |
No-Op ((void)) |
Scalar C Loop |