Research into unified cross-modal representation learning for arbitrary- dimensional signals. Part of the Stratum intelligence substrate research track at Mihok Labs.
# Install alloy package
pip install -e alloy/
# Quick training test (100 steps)
alloy train --experiment 004 --steps 50
# Run smoke tests
alloy test
# Evaluate a checkpoint
alloy evaluate --checkpoint runs/checkpoint.ptPhase 1 — Modality Bridging (phases/phase1/)
Interleaved alignment blocks with gradient separation. Best result: mixing
ratio 0.140, unification ratio 1.184. 14 experiments.
Phase 2 — Coordinate-Value Tokenization (phases/phase2/)
Unified tokenization where every signal is a function over coordinates.
No modality distinction at the representation level. Active.
| ID | Name | Description |
|---|---|---|
| 001 | coordinate_baseline |
Reconstruction-only over coordinate-value tokens |
| 002 | patch_as_function |
Perceiver with coordinate-contrastive loss |
| 003 | perceiver_capacity |
Higher-capacity Perceiver (32 latents, 6 layers) |
| 004 | coordinate_aware |
4TB+3AB, gradient separation, local value-gated InfoNCE + VICReg |
| 005 | mock_validation |
Test experiment for CI and pipeline validation |
To list registered experiments at any time:
alloy --list-experimentsThe primary entry point is the alloy CLI, installed as a script via pip install -e alloy/.
| Command | Description |
|---|---|
alloy train |
Train an experiment with staged pipeline |
alloy evaluate |
Evaluate a checkpoint and generate HTML report |
alloy test |
Run smoke tests (CPU, <30 seconds) |
alloy compare |
Compare two experiments against Phase 1 baselines |
alloy sweep |
Grid search hyperparameter sweep |
alloy --list-experiments |
List registered experiments |
alloy --version |
Print alloy version (v0.1.0) |
See Alloy Documentation for full architecture and Alloy ADRs for design decisions.
alloy train --experiment 004Flags:
| Flag | Type | Default | Description |
|---|---|---|---|
--experiment, -e |
str | "001" |
Experiment ID to run (001, 002, 003, 004, 005) |
--steps |
int | config value | Override training steps from config |
--config, -c |
path | config built-in | Path to JSON or YAML configuration file |
--resume |
flag | off | Resume from the latest checkpoint in results_dir |
--runpod-mode |
flag | off | Enable RunPod monitoring, signal handlers, and S3 sync |
--auto-memory |
flag | off | Auto-detect GPU memory and apply training preset |
Examples:
# Train experiment 004 with default config (35000 steps)
alloy train --experiment 004
# Quick test with 100 steps
alloy train --experiment 004 --steps 100
# Resume from last checkpoint
alloy train --experiment 004 --resume
# Train with auto memory detection (GPU preset)
alloy train --experiment 004 --auto-memory
# Train with custom YAML config
alloy train --experiment 004 --config my_config.yaml
# RunPod mode with S3 checkpoint sync
alloy train --experiment 004 --runpod-mode --resumeExpected output (first few lines):
============================================================
Alloy Phase 2 — Experiment 004: coordinate_aware
============================================================
device: mps
d_model: 128
num_heads: 8
num_transformer_layers: 4
num_alignment_layers: 3
num_fourier_freq: 32
rho: 0.1
local_sim_threshold: 0.5
batch_size: 16
num_steps: 100
mask_ratio: 0.3
tau_anneal: 0.5 -> 0.05
============================================================
Starting training...
VALIDATION STAGE — 100 steps
...
Training complete. Final recon loss: 0.042156
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004Flags:
| Flag | Type | Default | Description |
|---|---|---|---|
--checkpoint, -c |
path | required | Path to checkpoint.pt file |
--experiment, -e |
str | required | Experiment ID to build model/tokenizer architecture |
--baselines |
str (repeatable) | exp_011 |
Baseline experiment names to compare against |
--output, -o |
path | none | Directory to write evaluation report (JSON + HTML) |
Examples:
# Evaluate a checkpoint with baseline comparison
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004
# Generate HTML report
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004 --output reports/
# Compare against multiple baselines
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004 \
--baselines exp_011 --baselines exp_003Expected output:
Loading checkpoint: runs/checkpoint.pt
Step: 34999 Loss: 0.042156
============================================================
Evaluation
============================================================
============================================================
Key Metrics
============================================================
Recon MSE (1D): 0.0312
Recon MSE (2D): 0.0523
Recon MSE (combined): 0.0418
Semantic Alignment: 0.326
Probing Accuracy: 0.9840 (avg)
Mixing Ratio: 0.1400 (avg)
Unification Ratio: 1.1840 (avg)
Outcome: A (Strong unification)
============================================================
Baseline Comparison
============================================================
vs exp_011:
Better: probing_accuracy (+0.0520), mixing_ratio (+0.0100).
Worse: recon_mse (-0.0060)
Reports written to: reports/
Done.
alloy testFlags:
| Flag | Type | Default | Description |
|---|---|---|---|
--category, -k |
choice | smoke |
Test category: smoke, slow, gpu, all, integration |
--verbose, -v |
flag | off | Increase pytest verbosity |
Examples:
# Run smoke tests (fast, <5s each)
alloy test
# Run smoke tests with verbose output
alloy test --verbose
# Run slow tests (full training runs, >30s)
alloy test --category slow
# Run GPU tests (requires CUDA or MPS)
alloy test --category gpu
# Run integration tests (multi-module)
alloy test --category integration
# Run all tests
alloy test --category allDirect pytest equivalents:
pytest -m smoke # smoke tests
pytest -m slow # slow tests
pytest -m gpu # GPU tests
pytest -m integration # integration tests
pytest # allMarker reference:
| Marker | Description | Typical runtime |
|---|---|---|
smoke |
Fast sanity checks | <5s each |
slow |
Full training runs | >30s each |
gpu |
Requires CUDA or MPS | varies |
integration |
Multi-module tests | <30s each |
Expected output:
Running tests: pytest /path/to/alloy/tests -x -m smoke --timeout 30
==================================== test session starts =====================================
platform darwin -- Python 3.10.x, pytest-8.x.x, pluggy-1.x.x
rootdir: /path/to/alloy
configfile: pytest.ini
plugins: timeout-2.x.x
collected 23 items / 18 deselected / 5 selected
alloy/tests/test_cli.py ..... [100%]
===================================== 5 passed in 2.30s =====================================
alloy compare --exp-a runs/exp_a_checkpoint.pt --exp-b runs/exp_b_checkpoint.pt --experiment 004Flags:
| Flag | Type | Default | Description |
|---|---|---|---|
--exp-a, -a |
path | required | Path to first checkpoint.pt file |
--exp-b, -b |
path | required | Path to second checkpoint.pt file |
--experiment, -e |
str | required | Experiment ID to resolve model architecture |
--output, -o |
path | none | Optional JSON file to write full comparison data |
Example:
alloy compare \
--exp-a experiments/exp_004/results/checkpoint.pt \
--exp-b experiments/exp_004_ablated/results/checkpoint.pt \
--experiment 004 \
--output comparison.jsonExpected output:
Loading checkpoint A: checkpoint.pt
Loading checkpoint B: checkpoint.pt
============================================================
Comparison
============================================================
Metric A B Delta
──────────────────────── ── ── ───
recon_mse 0.0418 0.0472 +0.0054
probing_accuracy 0.9840 0.9720 -0.0120
mixing_ratio 0.1400 0.1280 -0.0120
unification_ratio 1.1840 1.2100 +0.0260
semantic_alignment 0.3260 0.2980 -0.0280
A step: 34999 B step: 34999
Summary (A vs B): Better: recon_mse (+0.0054), unification_ratio (+0.0260).
Worse: probing_accuracy (-0.0120), mixing_ratio (-0.0120),
semantic_alignment (-0.0280)
Comparison JSON written to: comparison.json
alloy sweep --param rho --values "0.05,0.1,0.15,0.2" --experiment 004 --steps 500Flags:
| Flag | Type | Default | Description |
|---|---|---|---|
--param |
str | required | Parameter name to sweep (e.g. rho, d_model, training__lr_transformer) |
--values |
str | required | Comma-separated values (e.g. 0.05,0.1,0.15) |
--experiment, -e |
str | required | Experiment ID to use as base configuration |
--steps |
int | config value | Training steps per sweep trial |
The sweep engine runs each value as a separate trial sequentially. Short parameter names
(e.g. rho) are automatically resolved to dotted config keys (model__rho). Full
dotted keys with __ (e.g. training__lr_transformer) are used as-is.
Each trial gets its own results directory under results_dir/sweep_{param}_{value}/
to avoid checkpoint conflicts. A single trial failure does not abort the sweep.
Example:
# Sweep the rho hyperparameter across 3 values, 500 steps each
alloy sweep --param rho --values "0.05,0.1,0.15" --experiment 004 --steps 500Expected output:
============================================================
Alloy Parameter Sweep
============================================================
Experiment: 004
Parameter: rho
Values: 0.05, 0.1, 0.15
Steps/run: 500
Total runs: 3
============================================================
─── Trial 1/3: model__rho = 0.05 ───
...
─── Trial 2/3: model__rho = 0.1 ───
...
─── Trial 3/3: model__rho = 0.15 ───
...
Results saved to /path/to/results/sweep_results.json
Sweep results JSON contains per-trial metrics (final_loss, loss_curve, config_summary) plus
aggregate statistics (best, worst, mean_loss, std_loss, per-param-value breakdown).
Alloy auto-detects hardware on startup via detect_hardware() from alloy.hardware.
The detection priority is: CUDA → MPS → CPU. The first available accelerator is selected.
from alloy.hardware import detect_hardware
hw = detect_hardware()
print(f"Device: {hw.device}")
print(f"Name: {hw.name}")
print(f"Memory: {hw.memory_gb:.1f} GB")
print(f"FP16: {hw.supports_fp16}")
print(f"BF16: {hw.supports_bf16}")
print(f"Batch size: {hw.max_batch_size}")
print(f"Strategy: {hw.recommended_strategy}")CUDA (NVIDIA GPU)
Device: cuda
Name: NVIDIA RTX 4090
Memory: 24.0 GB
FP16: True
BF16: True
Batch size: 64
Strategy: full
FP16 support requires compute capability 5.0+ (Pascal and newer). BF16 requires 8.0+ (Ampere and newer). VRAM tiers for batch sizing: >=24GB → 64, >=16GB → 32, >=8GB → 16, else → 8.
MPS (Apple Silicon)
Device: mps
Name: Apple MPS
Memory: 36.0 GB
FP16: True
BF16: False
Batch size: 24
Strategy: cpu_fallback
Apple Silicon supports float16 but not bfloat16 natively. The MPS backend has incomplete
operator coverage, so the recommended strategy is cpu_fallback. When an MPS operation
fails with RuntimeError, the fallback decorator catches it, runs on CPU, and moves the
result back to MPS automatically.
CPU
Device: cpu
Name: Apple M2 Pro
Memory: 36.0 GB
FP16: False
BF16: False
Batch size: 8
Strategy: tiled
The --auto-memory flag on alloy train detects GPU memory and applies an optimal
training preset before building model components. It configures batch size, gradient
accumulation steps, tiling, mixed precision, and gradient checkpointing based on
available VRAM or system memory:
| Memory | Batch size | Grad accum steps | Mixed precision | Tiled |
|---|---|---|---|---|
| >= 70 GB | 64 | 1 | yes | no |
| >= 35 GB | 32 | 1 | yes | yes |
| >= 24 GB | 24 | 1 | yes | yes |
| >= 16 GB | 16 | 4 | yes | yes |
| >= 8 GB | 8 | 8 | conditional | yes |
| < 8 GB | 4 | 16 | conditional | yes |
For MPS (Apple Silicon), the preset uses unified memory tiers with conservative batch sizes (since model activations and macOS share the same memory pool). For CPU, it falls back to batch_size=4 with no tiling or mixed precision.
Training runs through four sequential stages, each checkpointed on transition:
ValidationStage (100 steps)
→ Sanity-checks losses for NaN/Inf values
→ Checks gradient norms against threshold (default 10.0)
→ Halts early if either check fails
↓
WarmupStage (500 steps)
→ Linear LR interpolation: 1e-5 → target_lr
→ Logs LR at first, midpoint, and last step
↓
TrainingStage (num_steps from config)
→ Cosine decay from target_lr → min_lr (1% of target)
→ Early stopping: monitors recon_loss plateau (patience: 5000)
→ Periodic logging and checkpointing at config intervals
→ Graceful exit via SIGINT/SIGTERM signal handler
↓
EvaluationStage (single-shot)
→ model.eval() + tokenizer.eval()
→ Runs full 10-metric evaluation suite
→ Pretty-prints nested metric dicts
All stages run through TrainingHarness.run_staged(), which handles batch generation,
dual optimizer management (transformer + alignment), checkpoint persistence, S3 sync,
RunPod monitoring, and memory profiling.
Each checkpoint saved to results_dir/checkpoint.pt contains:
| Key | Description |
|---|---|
step |
Global training step counter |
loss / recon_loss |
Reconstruction loss at checkpoint time |
model |
Model state dict |
tokenizer |
Tokenizer state dict |
opt_t |
Transformer optimizer state (Adam) |
opt_a |
Alignment optimizer state (Adam) |
config |
Full experiment configuration dict |
rng_state |
Random number generator states |
config_hash |
Config fingerprint for resumability |
comparator |
LocalValueComparator state (exp_004 only) |
# Resume from the latest checkpoint in results_dir
alloy train --experiment 004 --resume
# Resume with RunPod mode
alloy train --experiment 004 --resume --runpod-modeWhen --resume is set, the harness finds the latest checkpoint in results_dir and
restores model, tokenizer, and dual optimizer states before entering the training loop.
The RNG state is also restored for deterministic continuation.
alloy evaluate --checkpoint results/checkpoint.pt --experiment 004This builds the model and tokenizer architecture from the registered experiment, loads saved weights, and runs the full evaluation suite. The 10-metric suite covers:
- Reconstruction (MSE per modality + combined)
- Semantic alignment (cross-modal representational similarity)
- Probing accuracy (linear probe on learned representations, per-layer)
- Mixing ratio (how intermixed 1D and 2D tokens are in latent space)
- Unification ratio (whether the model treats both modalities as one)
- Coordinate sensitivity (how much token representations vary with coordinate position)
- Coordinate neighborhood (smoothness of coordinate-adjacent token representations)
- PCA overlap grids (per-layer overlap between 1D and 2D token projections)
- Latent slot specialization (per-slot activation statistics)
- Outcome classification (A/B/C letter grade based on thresholds)
By default, evaluation compares against Phase 1's exp_011 (the best Phase 1 result
with mixing ratio 0.140 and unification ratio 1.184). Additional baselines can be
specified with --baselines:
alloy evaluate \
--checkpoint results/checkpoint.pt \
--experiment 004 \
--baselines exp_011 \
--baselines exp_001Use --output to generate standalone dark-themed HTML reports with inline CSS
(no CDN, no JS required). Each report includes: key metrics table, per-layer
analysis tables, coordinate sensitivity plots, PCA overlap grids, latent slot
specialization bars, and outcome classification.
alloy evaluate --checkpoint results/checkpoint.pt --experiment 004 --output reports/
# Writes: reports/evaluation_report.html + reports/evaluation_report.jsonalloy compare \
--exp-a experiments/exp_004/results/checkpoint.pt \
--exp-b experiments/exp_004_v2/results/checkpoint.pt \
--experiment 004Prints a side-by-side metric comparison with per-metric deltas and a human-readable
summary of improvements and regressions. Use --output for JSON export.
The GridSearch engine varies one hyperparameter across a list of values while
holding the base config constant:
# Sweep rho across 4 values, 500 steps per trial
alloy sweep --param rho --values "0.05,0.1,0.15,0.2" --experiment 004 --steps 500Each trial constructs a modified config via ExperimentConfig.override(), creates a
trial-specific results directory, and runs the full training pipeline. Failures set
final_loss=inf and continue to remaining trials.
from alloy.sweep import GridSearch
from alloy.configs import PROD_CONFIG
sweep = GridSearch(
base_config=PROD_CONFIG,
param_name="rho",
values=[0.05, 0.1, 0.15, 0.2],
)
results = sweep.run(experiment_name="004", steps_per_trial=500)
# Aggregate statistics
stats = sweep.aggregate(results)
print(f"Best: rho={stats['best']['param_value']} loss={stats['best']['final_loss']:.6f}")
# Save results
sweep.save("sweep_rho_results.json")Short parameter names (e.g. rho) are auto-resolved to model__rho. Full dotted keys
with __ (e.g. training__lr_transformer) are used as-is.
alloy test # Smoke tests (<5s each)alloy test --category slow # Full training runs (>30s)
alloy test --category gpu # GPU-required tests
alloy test --category integration # Multi-module tests
alloy test --category all # Everythingpytest alloy/tests/ -v -m smoke
pytest alloy/tests/ -v -m slow
pytest alloy/tests/ -v -m gpu
pytest alloy/tests/ -v -m integration
pytest alloy/tests/ -v # All testsTest files:
| File | Coverage |
|---|---|
test_cli.py |
CLI help, list-experiments, train/evaluate help, version |
test_training.py |
Harness imports, stage construction, early stopping, LR scheduling, minimal training run |
test_hardware.py |
Hardware detection, profile fields, MPS fallback, optimal config, CPU fallback |
test_evaluation.py |
Evaluation reports, baseline registry, ASCII visualizations |
alloy/
├── cli/ # Click-based CLI (train, evaluate, test, compare, sweep)
│ ├── main.py # CLI entry point, --list-experiments, --version
│ ├── train.py # alloy train command
│ ├── evaluate.py # alloy evaluate command
│ ├── test.py # alloy test command (pytest wrapper)
│ ├── compare.py # alloy compare command
│ └── sweep.py # alloy sweep command
├── core/ # Core abstractions
│ ├── base.py # BaseExperiment ABC
│ ├── config.py # ExperimentConfig, ModelConfig, TrainingConfig, etc.
│ ├── harness.py # TrainingHarness (training loop + run_staged)
│ ├── registry.py # ExperimentRegistry (auto-discovery)
│ └── logging.py # MetricsLogger (JSONL logging)
├── model/ # Model architectures
│ ├── coordinate_alloy.py # CoordinateAlloy (4TB + 3AB interleaved)
│ ├── bidirectional.py # BidirectionalTransformer
│ ├── tokenizer.py # CoordinateTokenizer
│ ├── perceiver.py # Perceiver (cross-attention bottleneck)
│ └── fourier.py # Fourier feature coordinate encoding
├── training/ # Training pipeline and scheduling
│ ├── pipeline.py # TrainingPipeline (4-stage orchestration)
│ ├── stage.py # Stage, ValidationStage, WarmupStage, TrainingStage, EvaluationStage
│ └── scheduling.py # LRSchedule (cosine), EarlyStopping
├── evaluation/ # Evaluation suite
│ ├── report.py # EvaluationReport (HTML + JSON + WandB)
│ ├── baselines.py # BaselineRegistry (Phase 1 comparisons)
│ └── visualizations.py # ASCII visualizations (PCA grids, loss curves, metric comparisons)
├── experiments/ # All registered experiments
│ ├── exp_001_coordinate_baseline/
│ ├── exp_002_patch_as_function/
│ ├── exp_003_perceiver_capacity/
│ ├── exp_004_coordinate_aware/
│ └── exp_005_mock_validation/
├── hardware/ # Hardware detection and capability probing
│ ├── detector.py # detect_hardware() — CUDA → MPS → CPU
│ ├── profiles.py # HardwareProfile, get_optimal_config()
│ └── fallbacks.py # fallback_decorator() — MPS → CPU retry
├── sweep/ # Hyperparameter sweeping
│ └── engine.py # GridSearch — single-param grid sweep
├── deployment/ # Production infrastructure
│ ├── checkpoint_manager.py # Checkpoint save/load/S3 sync
│ ├── signal_handlers.py # GracefulExitHandler (SIGINT/SIGTERM)
│ ├── monitoring.py # RunPodMonitor (heartbeat, HTTP)
│ ├── memory.py # MemoryManager (GPU profiling)
│ └── resume.py # Checkpoint resume loading
├── configs/ # Built-in config presets
│ ├── prod.py # PROD_CONFIG (35000 steps, d_model=128)
│ ├── debug.py # DEBUG_CONFIG (100 steps, d_model=32)
│ └── sweep_base.py # SWEEP_BASE_CONFIG (500 steps)
├── compat/ # Phase 1 compatibility layer
│ └── phase1.py # Phase1CheckpointLoader (baseline metrics)
├── utils/ # Shared utilities
│ ├── signals.py # RichSignal1D, RichSignal2D generators
│ ├── patching.py # Patch extraction and reconstruction
│ ├── similarity.py # Cross-modal similarity functions
│ └── evaluation.py # Shared evaluation helpers
├── tests/ # Test suite
│ ├── conftest.py # Pytest fixtures and marker registration
│ ├── test_cli.py # CLI smoke tests
│ ├── test_training.py # Training pipeline tests
│ ├── test_hardware.py # Hardware detection tests
│ └── test_evaluation.py # Evaluation report tests
├── docs/ # Architecture and design docs
│ ├── architecture.md # Full Phase 2 architecture overview
│ └── adr/ # Architecture Decision Records
│ ├── 0001-coordinate-value-tokenization.md
│ ├── 0002-gradient-separation-pattern.md
│ └── 0003-interleaved-architecture.md
├── notebooks/ # Jupyter notebooks (exploratory analysis)
├── pyproject.toml # Package config, dependencies, CLI entry point
├── pytest.ini # Pytest config: markers, timeout, paths
├── requirements.txt # Runtime dependencies
└── README.md # This file
Alloy Phase 2 inverts the premise of Phase 1. Rather than building alignment machinery on top of modality-specific embeddings, it asks whether coordinate-value tokenization prevents the modality gap from forming at all. Every signal, regardless of original dimensionality, enters the model as a set of (coordinate, value) pairs. The tokenizer doesn't know whether a patch came from a 1D or 2D signal. It knows the coordinate where the patch was sampled and the values observed there.
-
Coordinate-value tokenization: Both 1D and 2D patches are encoded as
[coord_embed || value_embed]tokens. Fourier feature encoding on coordinates. Shared coordinate projection across dimensionalities. Separate value projections per dimensionality (necessary because raw patch sizes differ). -
Gradient separation: Transformer and alignment parameters have structurally separated gradient paths. The transformer optimizer (opt_t) only steps on reconstruction gradients. The alignment optimizer (opt_a) only steps on InfoNCE and VICReg gradients. This prevents alignment losses from distorting the reconstruction backbone.
-
Interleaved architecture (exp_004): 4 transformer blocks interleaved with 3 alignment blocks. Alignment blocks use cross-modal cross-attention. No modality label is ever injected into the token stream.
-
Staged training: 4 stages (validation → warmup → training → evaluation), each with own lifecycle hooks. Validation catches NaN/exploding gradients before they waste compute. Warmup does linear LR interpolation. Training uses cosine decay with early stopping. Evaluation runs the full suite once.
-
Hardware auto-detection: CUDA, MPS, and CPU with capability-specific presets. Mixed precision where available. Graceful MPS fallback for unsupported operations.
For the full architecture document, see alloy/docs/architecture.md.
docker build -t alloy:latest .
docker run -it alloy:latest alloy train --experiment 004 --steps 100# Build
docker build -f Dockerfile.runpod -t alloy-exp004:latest .
# Run on RunPod (GPU required)
docker run --gpus all \
-e S3_BUCKET=my-bucket \
-e S3_PREFIX=exp004 \
alloy-exp004:latest# .github/workflows/test.yml
# Runs on push and PR to main
# Python 3.10 + 3.11, smoke tests, 10 min timeout
# pip install -e "alloy/[dev]", then pytest -v -m smokeRUNPOD_MODE=1 # Enable RunPod optimizations (heartbeat, S3 sync, signal handlers)
S3_BUCKET=my-bucket # S3 bucket for checkpoint persistence
S3_PREFIX=exp004 # S3 key prefix for organizing checkpoints
MONITOR_PORT=8080 # Enable HTTP monitoring on this port (with RunPod mode)When RUNPOD_MODE=1 is set (or --runpod-mode is passed), the training harness:
- Starts a
RunPodMonitorthat writes heartbeat files for the RunPod daemon - Optionally starts an HTTP server on
MONITOR_PORTfor health/metrics endpoints - Registers SIGINT/SIGTERM handlers for graceful checkpoint-on-exit
- Syncs checkpoints to S3 after each save (via
CheckpointManager.s3_sync())
The phases/phase2/ entry point still works for backwards compatibility:
cd phases/phase2
# Train
python main.py --experiment 004
# Evaluate
python main.py --experiment 004 --eval-only
# Resume
python main.py --experiment 004 --resume auto
# RunPod mode (monitoring, signal handlers, S3 sync)
python main.py --experiment 004 --resume auto --runpod-modeNote: This is the legacy interface. For new workflows, use the alloy CLI documented above.
Symptom: recon_loss becomes NaN or inf during validation or training.
Causes and fixes:
- Learning rate too high: Reduce
lr_transformerandlr_alignmentin your config. - Gradient explosion: Enable gradient clipping (
grad_clipin training config). Check gradient norms in the validation stage output. - Mask ratio too high: With
mask_ratio > 0.5, there may be too few visible tokens to reconstruct. Trymask_ratio=0.3or lower. - Numerical instability in tau: Set
tau_startlower (e.g.0.5instead of1.0). Extremely high tau can cause InfoNCE overflow. - Run with
--auto-memoryto apply conservative presets.
Symptom: All contrastive pair labels are negative (no anchors find nearby positives).
Causes and fixes:
local_sim_thresholdtoo high: The similarity threshold for considering two tokens "close" may be too strict. Trylocal_sim_threshold=0.3or0.1.rhotoo small: The InfoNCE temperature mixing ratio controls how much local vs global structure is considered. Tryrho=0.2orrho=0.3.- Patch size mismatch: If 1D and 2D patch sizes differ dramatically, their value embeddings may never be comparable. Ensure
value_embed_dimis the same for both modalities.
Symptom: Training on Apple Silicon produces warnings like MPS operation not implemented, falling back to CPU.
Cause: The MPS backend has incomplete operator coverage. Some operations (e.g. certain indexing patterns, torch.unique, some reductions) don't have MPS kernels.
What happens: The fallback_decorator in alloy/hardware/fallbacks.py automatically catches RuntimeError on MPS, runs the operation on CPU, and moves the result back to MPS. A warning is logged each time. This is normal behavior on Apple Silicon and doesn't affect correctness.
Symptom: ModuleNotFoundError: No module named 'alloy' or ImportError.
Fixes:
- Ensure alloy is installed:
pip install -e alloy/ - Verify
.venvis activated:source .venv/bin/activate - If using legacy
phases/phase2/entry point, run from thephases/phase2/directory
"No experiments registered":
alloy --list-experiments
# Expected: 001, 002, 003, 004, 005If experiments are missing, check that alloy/experiments/exp_*/experiment.py files exist and are importable.
CUDA out of memory:
Use --auto-memory to apply a memory-appropriate preset, or reduce batch_size in your config.
Checkpoint not found on resume:
The --resume flag looks for the latest checkpoint in results_dir (default: alloy/results/). Verify the path exists and contains checkpoint.pt files.
- Architecture Overview — Full Phase 2 architecture, component diagrams, data flow
- ADR 0001: Coordinate-Value Tokenization — Why tokenize as functions instead of embeddings
- ADR 0002: Gradient Separation Pattern — Dual-optimizer design rationale
- ADR 0003: Interleaved Architecture — Why interleave alignment blocks in the transformer stack
- User Guide — Configuration reference, evaluation metrics, RunPod deployment guide
- Phase 1 Documentation — Original modality bridging experiments (exp_001 through exp_014)
- CLAUDE.md — Claude handoff reference (project structure, running experiments)