Skip to content

Repository files navigation

Alloy

Research into unified cross-modal representation learning for arbitrary- dimensional signals. Part of the Stratum intelligence substrate research track at Mihok Labs.

Quick Start

Local

# Install alloy package
pip install -e alloy/

# Quick training test (100 steps)
alloy train --experiment 004 --steps 50

# Run smoke tests
alloy test

# Evaluate a checkpoint
alloy evaluate --checkpoint runs/checkpoint.pt

Phases

Phase 1 — Modality Bridging (phases/phase1/) Interleaved alignment blocks with gradient separation. Best result: mixing ratio 0.140, unification ratio 1.184. 14 experiments.

Phase 2 — Coordinate-Value Tokenization (phases/phase2/) Unified tokenization where every signal is a function over coordinates. No modality distinction at the representation level. Active.

Phase 2 Experiments

ID Name Description
001 coordinate_baseline Reconstruction-only over coordinate-value tokens
002 patch_as_function Perceiver with coordinate-contrastive loss
003 perceiver_capacity Higher-capacity Perceiver (32 latents, 6 layers)
004 coordinate_aware 4TB+3AB, gradient separation, local value-gated InfoNCE + VICReg
005 mock_validation Test experiment for CI and pipeline validation

To list registered experiments at any time:

alloy --list-experiments

CLI Reference

The primary entry point is the alloy CLI, installed as a script via pip install -e alloy/.

Command Description
alloy train Train an experiment with staged pipeline
alloy evaluate Evaluate a checkpoint and generate HTML report
alloy test Run smoke tests (CPU, <30 seconds)
alloy compare Compare two experiments against Phase 1 baselines
alloy sweep Grid search hyperparameter sweep
alloy --list-experiments List registered experiments
alloy --version Print alloy version (v0.1.0)

See Alloy Documentation for full architecture and Alloy ADRs for design decisions.

alloy train

alloy train --experiment 004

Flags:

Flag Type Default Description
--experiment, -e str "001" Experiment ID to run (001, 002, 003, 004, 005)
--steps int config value Override training steps from config
--config, -c path config built-in Path to JSON or YAML configuration file
--resume flag off Resume from the latest checkpoint in results_dir
--runpod-mode flag off Enable RunPod monitoring, signal handlers, and S3 sync
--auto-memory flag off Auto-detect GPU memory and apply training preset

Examples:

# Train experiment 004 with default config (35000 steps)
alloy train --experiment 004

# Quick test with 100 steps
alloy train --experiment 004 --steps 100

# Resume from last checkpoint
alloy train --experiment 004 --resume

# Train with auto memory detection (GPU preset)
alloy train --experiment 004 --auto-memory

# Train with custom YAML config
alloy train --experiment 004 --config my_config.yaml

# RunPod mode with S3 checkpoint sync
alloy train --experiment 004 --runpod-mode --resume

Expected output (first few lines):

============================================================
Alloy Phase 2 — Experiment 004: coordinate_aware
============================================================
device:              mps
d_model:             128
num_heads:           8
num_transformer_layers:   4
num_alignment_layers:     3
num_fourier_freq:         32
rho:                      0.1
local_sim_threshold:      0.5
batch_size:          16
num_steps:           100
mask_ratio:          0.3
tau_anneal:           0.5 -> 0.05
============================================================

Starting training...

  VALIDATION STAGE — 100 steps
  ...

Training complete.  Final recon loss: 0.042156

alloy evaluate

alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004

Flags:

Flag Type Default Description
--checkpoint, -c path required Path to checkpoint.pt file
--experiment, -e str required Experiment ID to build model/tokenizer architecture
--baselines str (repeatable) exp_011 Baseline experiment names to compare against
--output, -o path none Directory to write evaluation report (JSON + HTML)

Examples:

# Evaluate a checkpoint with baseline comparison
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004

# Generate HTML report
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004 --output reports/

# Compare against multiple baselines
alloy evaluate --checkpoint runs/checkpoint.pt --experiment 004 \
    --baselines exp_011 --baselines exp_003

Expected output:

Loading checkpoint: runs/checkpoint.pt
  Step: 34999  Loss: 0.042156

============================================================
Evaluation
============================================================

============================================================
Key Metrics
============================================================
  Recon MSE (1D):       0.0312
  Recon MSE (2D):       0.0523
  Recon MSE (combined): 0.0418
  Semantic Alignment:   0.326
  Probing Accuracy:     0.9840 (avg)
  Mixing Ratio:         0.1400 (avg)
  Unification Ratio:    1.1840 (avg)
  Outcome:              A (Strong unification)

============================================================
Baseline Comparison
============================================================

  vs exp_011:
    Better: probing_accuracy (+0.0520), mixing_ratio (+0.0100).
    Worse: recon_mse (-0.0060)

Reports written to: reports/

Done.

alloy test

alloy test

Flags:

Flag Type Default Description
--category, -k choice smoke Test category: smoke, slow, gpu, all, integration
--verbose, -v flag off Increase pytest verbosity

Examples:

# Run smoke tests (fast, <5s each)
alloy test

# Run smoke tests with verbose output
alloy test --verbose

# Run slow tests (full training runs, >30s)
alloy test --category slow

# Run GPU tests (requires CUDA or MPS)
alloy test --category gpu

# Run integration tests (multi-module)
alloy test --category integration

# Run all tests
alloy test --category all

Direct pytest equivalents:

pytest -m smoke          # smoke tests
pytest -m slow           # slow tests
pytest -m gpu            # GPU tests
pytest -m integration    # integration tests
pytest                   # all

Marker reference:

Marker Description Typical runtime
smoke Fast sanity checks <5s each
slow Full training runs >30s each
gpu Requires CUDA or MPS varies
integration Multi-module tests <30s each

Expected output:

Running tests: pytest /path/to/alloy/tests -x -m smoke --timeout 30
==================================== test session starts =====================================
platform darwin -- Python 3.10.x, pytest-8.x.x, pluggy-1.x.x
rootdir: /path/to/alloy
configfile: pytest.ini
plugins: timeout-2.x.x
collected 23 items / 18 deselected / 5 selected

alloy/tests/test_cli.py .....                                                          [100%]

===================================== 5 passed in 2.30s =====================================

alloy compare

alloy compare --exp-a runs/exp_a_checkpoint.pt --exp-b runs/exp_b_checkpoint.pt --experiment 004

Flags:

Flag Type Default Description
--exp-a, -a path required Path to first checkpoint.pt file
--exp-b, -b path required Path to second checkpoint.pt file
--experiment, -e str required Experiment ID to resolve model architecture
--output, -o path none Optional JSON file to write full comparison data

Example:

alloy compare \
    --exp-a experiments/exp_004/results/checkpoint.pt \
    --exp-b experiments/exp_004_ablated/results/checkpoint.pt \
    --experiment 004 \
    --output comparison.json

Expected output:

Loading checkpoint A: checkpoint.pt
Loading checkpoint B: checkpoint.pt

============================================================
Comparison
============================================================

  Metric                         A          B       Delta
  ────────────────────────        ──         ──       ───
  recon_mse                0.0418     0.0472    +0.0054
  probing_accuracy         0.9840     0.9720    -0.0120
  mixing_ratio             0.1400     0.1280    -0.0120
  unification_ratio        1.1840     1.2100    +0.0260
  semantic_alignment       0.3260     0.2980    -0.0280

  A step: 34999  B step: 34999

  Summary (A vs B): Better: recon_mse (+0.0054), unification_ratio (+0.0260).
                    Worse: probing_accuracy (-0.0120), mixing_ratio (-0.0120),
                    semantic_alignment (-0.0280)

Comparison JSON written to: comparison.json

alloy sweep

alloy sweep --param rho --values "0.05,0.1,0.15,0.2" --experiment 004 --steps 500

Flags:

Flag Type Default Description
--param str required Parameter name to sweep (e.g. rho, d_model, training__lr_transformer)
--values str required Comma-separated values (e.g. 0.05,0.1,0.15)
--experiment, -e str required Experiment ID to use as base configuration
--steps int config value Training steps per sweep trial

The sweep engine runs each value as a separate trial sequentially. Short parameter names (e.g. rho) are automatically resolved to dotted config keys (model__rho). Full dotted keys with __ (e.g. training__lr_transformer) are used as-is.

Each trial gets its own results directory under results_dir/sweep_{param}_{value}/ to avoid checkpoint conflicts. A single trial failure does not abort the sweep.

Example:

# Sweep the rho hyperparameter across 3 values, 500 steps each
alloy sweep --param rho --values "0.05,0.1,0.15" --experiment 004 --steps 500

Expected output:

============================================================
Alloy Parameter Sweep
============================================================
  Experiment: 004
  Parameter:  rho
  Values:     0.05, 0.1, 0.15
  Steps/run:  500
  Total runs: 3
============================================================

─── Trial 1/3: model__rho = 0.05 ───
...
─── Trial 2/3: model__rho = 0.1 ───
...
─── Trial 3/3: model__rho = 0.15 ───
...

Results saved to /path/to/results/sweep_results.json

Sweep results JSON contains per-trial metrics (final_loss, loss_curve, config_summary) plus aggregate statistics (best, worst, mean_loss, std_loss, per-param-value breakdown).

Hardware Setup

Alloy auto-detects hardware on startup via detect_hardware() from alloy.hardware. The detection priority is: CUDA → MPS → CPU. The first available accelerator is selected.

Detection by platform

from alloy.hardware import detect_hardware

hw = detect_hardware()
print(f"Device:     {hw.device}")
print(f"Name:       {hw.name}")
print(f"Memory:     {hw.memory_gb:.1f} GB")
print(f"FP16:       {hw.supports_fp16}")
print(f"BF16:       {hw.supports_bf16}")
print(f"Batch size: {hw.max_batch_size}")
print(f"Strategy:   {hw.recommended_strategy}")

CUDA (NVIDIA GPU)

Device:     cuda
Name:       NVIDIA RTX 4090
Memory:     24.0 GB
FP16:       True
BF16:       True
Batch size: 64
Strategy:   full

FP16 support requires compute capability 5.0+ (Pascal and newer). BF16 requires 8.0+ (Ampere and newer). VRAM tiers for batch sizing: >=24GB → 64, >=16GB → 32, >=8GB → 16, else → 8.

MPS (Apple Silicon)

Device:     mps
Name:       Apple MPS
Memory:     36.0 GB
FP16:       True
BF16:       False
Batch size: 24
Strategy:   cpu_fallback

Apple Silicon supports float16 but not bfloat16 natively. The MPS backend has incomplete operator coverage, so the recommended strategy is cpu_fallback. When an MPS operation fails with RuntimeError, the fallback decorator catches it, runs on CPU, and moves the result back to MPS automatically.

CPU

Device:     cpu
Name:       Apple M2 Pro
Memory:     36.0 GB
FP16:       False
BF16:       False
Batch size: 8
Strategy:   tiled

Auto-memory mode

The --auto-memory flag on alloy train detects GPU memory and applies an optimal training preset before building model components. It configures batch size, gradient accumulation steps, tiling, mixed precision, and gradient checkpointing based on available VRAM or system memory:

Memory Batch size Grad accum steps Mixed precision Tiled
>= 70 GB 64 1 yes no
>= 35 GB 32 1 yes yes
>= 24 GB 24 1 yes yes
>= 16 GB 16 4 yes yes
>= 8 GB 8 8 conditional yes
< 8 GB 4 16 conditional yes

For MPS (Apple Silicon), the preset uses unified memory tiers with conservative batch sizes (since model activations and macOS share the same memory pool). For CPU, it falls back to batch_size=4 with no tiling or mixed precision.

Training pipeline stages

Training runs through four sequential stages, each checkpointed on transition:

ValidationStage (100 steps)
  → Sanity-checks losses for NaN/Inf values
  → Checks gradient norms against threshold (default 10.0)
  → Halts early if either check fails
  ↓
WarmupStage (500 steps)
  → Linear LR interpolation: 1e-5 → target_lr
  → Logs LR at first, midpoint, and last step
  ↓
TrainingStage (num_steps from config)
  → Cosine decay from target_lr → min_lr (1% of target)
  → Early stopping: monitors recon_loss plateau (patience: 5000)
  → Periodic logging and checkpointing at config intervals
  → Graceful exit via SIGINT/SIGTERM signal handler
  ↓
EvaluationStage (single-shot)
  → model.eval() + tokenizer.eval()
  → Runs full 10-metric evaluation suite
  → Pretty-prints nested metric dicts

All stages run through TrainingHarness.run_staged(), which handles batch generation, dual optimizer management (transformer + alignment), checkpoint persistence, S3 sync, RunPod monitoring, and memory profiling.

Checkpoint format

Each checkpoint saved to results_dir/checkpoint.pt contains:

Key Description
step Global training step counter
loss / recon_loss Reconstruction loss at checkpoint time
model Model state dict
tokenizer Tokenizer state dict
opt_t Transformer optimizer state (Adam)
opt_a Alignment optimizer state (Adam)
config Full experiment configuration dict
rng_state Random number generator states
config_hash Config fingerprint for resumability
comparator LocalValueComparator state (exp_004 only)

Resume from checkpoint

# Resume from the latest checkpoint in results_dir
alloy train --experiment 004 --resume

# Resume with RunPod mode
alloy train --experiment 004 --resume --runpod-mode

When --resume is set, the harness finds the latest checkpoint in results_dir and restores model, tokenizer, and dual optimizer states before entering the training loop. The RNG state is also restored for deterministic continuation.

Evaluation

Checkpoint evaluation

alloy evaluate --checkpoint results/checkpoint.pt --experiment 004

This builds the model and tokenizer architecture from the registered experiment, loads saved weights, and runs the full evaluation suite. The 10-metric suite covers:

  • Reconstruction (MSE per modality + combined)
  • Semantic alignment (cross-modal representational similarity)
  • Probing accuracy (linear probe on learned representations, per-layer)
  • Mixing ratio (how intermixed 1D and 2D tokens are in latent space)
  • Unification ratio (whether the model treats both modalities as one)
  • Coordinate sensitivity (how much token representations vary with coordinate position)
  • Coordinate neighborhood (smoothness of coordinate-adjacent token representations)
  • PCA overlap grids (per-layer overlap between 1D and 2D token projections)
  • Latent slot specialization (per-slot activation statistics)
  • Outcome classification (A/B/C letter grade based on thresholds)

Baseline comparison

By default, evaluation compares against Phase 1's exp_011 (the best Phase 1 result with mixing ratio 0.140 and unification ratio 1.184). Additional baselines can be specified with --baselines:

alloy evaluate \
    --checkpoint results/checkpoint.pt \
    --experiment 004 \
    --baselines exp_011 \
    --baselines exp_001

HTML reports

Use --output to generate standalone dark-themed HTML reports with inline CSS (no CDN, no JS required). Each report includes: key metrics table, per-layer analysis tables, coordinate sensitivity plots, PCA overlap grids, latent slot specialization bars, and outcome classification.

alloy evaluate --checkpoint results/checkpoint.pt --experiment 004 --output reports/
# Writes: reports/evaluation_report.html + reports/evaluation_report.json

Comparing two checkpoints

alloy compare \
    --exp-a experiments/exp_004/results/checkpoint.pt \
    --exp-b experiments/exp_004_v2/results/checkpoint.pt \
    --experiment 004

Prints a side-by-side metric comparison with per-metric deltas and a human-readable summary of improvements and regressions. Use --output for JSON export.

Sweeps

Grid search over a single parameter

The GridSearch engine varies one hyperparameter across a list of values while holding the base config constant:

# Sweep rho across 4 values, 500 steps per trial
alloy sweep --param rho --values "0.05,0.1,0.15,0.2" --experiment 004 --steps 500

Each trial constructs a modified config via ExperimentConfig.override(), creates a trial-specific results directory, and runs the full training pipeline. Failures set final_loss=inf and continue to remaining trials.

Programmatic API

from alloy.sweep import GridSearch
from alloy.configs import PROD_CONFIG

sweep = GridSearch(
    base_config=PROD_CONFIG,
    param_name="rho",
    values=[0.05, 0.1, 0.15, 0.2],
)
results = sweep.run(experiment_name="004", steps_per_trial=500)

# Aggregate statistics
stats = sweep.aggregate(results)
print(f"Best: rho={stats['best']['param_value']} loss={stats['best']['final_loss']:.6f}")

# Save results
sweep.save("sweep_rho_results.json")

Short parameter names (e.g. rho) are auto-resolved to model__rho. Full dotted keys with __ (e.g. training__lr_transformer) are used as-is.

Testing

Quick smoke tests

alloy test                    # Smoke tests (<5s each)

Full test categories

alloy test --category slow         # Full training runs (>30s)
alloy test --category gpu          # GPU-required tests
alloy test --category integration  # Multi-module tests
alloy test --category all          # Everything

Direct pytest

pytest alloy/tests/ -v -m smoke
pytest alloy/tests/ -v -m slow
pytest alloy/tests/ -v -m gpu
pytest alloy/tests/ -v -m integration
pytest alloy/tests/ -v              # All tests

Test files:

File Coverage
test_cli.py CLI help, list-experiments, train/evaluate help, version
test_training.py Harness imports, stage construction, early stopping, LR scheduling, minimal training run
test_hardware.py Hardware detection, profile fields, MPS fallback, optimal config, CPU fallback
test_evaluation.py Evaluation reports, baseline registry, ASCII visualizations

Project Structure

alloy/
├── cli/                       # Click-based CLI (train, evaluate, test, compare, sweep)
│   ├── main.py                # CLI entry point, --list-experiments, --version
│   ├── train.py               # alloy train command
│   ├── evaluate.py            # alloy evaluate command
│   ├── test.py                # alloy test command (pytest wrapper)
│   ├── compare.py             # alloy compare command
│   └── sweep.py               # alloy sweep command
├── core/                      # Core abstractions
│   ├── base.py                # BaseExperiment ABC
│   ├── config.py              # ExperimentConfig, ModelConfig, TrainingConfig, etc.
│   ├── harness.py             # TrainingHarness (training loop + run_staged)
│   ├── registry.py            # ExperimentRegistry (auto-discovery)
│   └── logging.py             # MetricsLogger (JSONL logging)
├── model/                     # Model architectures
│   ├── coordinate_alloy.py    # CoordinateAlloy (4TB + 3AB interleaved)
│   ├── bidirectional.py       # BidirectionalTransformer
│   ├── tokenizer.py           # CoordinateTokenizer
│   ├── perceiver.py           # Perceiver (cross-attention bottleneck)
│   └── fourier.py             # Fourier feature coordinate encoding
├── training/                  # Training pipeline and scheduling
│   ├── pipeline.py            # TrainingPipeline (4-stage orchestration)
│   ├── stage.py               # Stage, ValidationStage, WarmupStage, TrainingStage, EvaluationStage
│   └── scheduling.py          # LRSchedule (cosine), EarlyStopping
├── evaluation/                # Evaluation suite
│   ├── report.py              # EvaluationReport (HTML + JSON + WandB)
│   ├── baselines.py           # BaselineRegistry (Phase 1 comparisons)
│   └── visualizations.py      # ASCII visualizations (PCA grids, loss curves, metric comparisons)
├── experiments/               # All registered experiments
│   ├── exp_001_coordinate_baseline/
│   ├── exp_002_patch_as_function/
│   ├── exp_003_perceiver_capacity/
│   ├── exp_004_coordinate_aware/
│   └── exp_005_mock_validation/
├── hardware/                  # Hardware detection and capability probing
│   ├── detector.py            # detect_hardware() — CUDA → MPS → CPU
│   ├── profiles.py            # HardwareProfile, get_optimal_config()
│   └── fallbacks.py           # fallback_decorator() — MPS → CPU retry
├── sweep/                     # Hyperparameter sweeping
│   └── engine.py              # GridSearch — single-param grid sweep
├── deployment/                # Production infrastructure
│   ├── checkpoint_manager.py  # Checkpoint save/load/S3 sync
│   ├── signal_handlers.py     # GracefulExitHandler (SIGINT/SIGTERM)
│   ├── monitoring.py          # RunPodMonitor (heartbeat, HTTP)
│   ├── memory.py              # MemoryManager (GPU profiling)
│   └── resume.py              # Checkpoint resume loading
├── configs/                   # Built-in config presets
│   ├── prod.py                # PROD_CONFIG (35000 steps, d_model=128)
│   ├── debug.py               # DEBUG_CONFIG (100 steps, d_model=32)
│   └── sweep_base.py          # SWEEP_BASE_CONFIG (500 steps)
├── compat/                    # Phase 1 compatibility layer
│   └── phase1.py              # Phase1CheckpointLoader (baseline metrics)
├── utils/                     # Shared utilities
│   ├── signals.py             # RichSignal1D, RichSignal2D generators
│   ├── patching.py            # Patch extraction and reconstruction
│   ├── similarity.py          # Cross-modal similarity functions
│   └── evaluation.py          # Shared evaluation helpers
├── tests/                     # Test suite
│   ├── conftest.py            # Pytest fixtures and marker registration
│   ├── test_cli.py            # CLI smoke tests
│   ├── test_training.py       # Training pipeline tests
│   ├── test_hardware.py       # Hardware detection tests
│   └── test_evaluation.py     # Evaluation report tests
├── docs/                      # Architecture and design docs
│   ├── architecture.md        # Full Phase 2 architecture overview
│   └── adr/                   # Architecture Decision Records
│       ├── 0001-coordinate-value-tokenization.md
│       ├── 0002-gradient-separation-pattern.md
│       └── 0003-interleaved-architecture.md
├── notebooks/                 # Jupyter notebooks (exploratory analysis)
├── pyproject.toml             # Package config, dependencies, CLI entry point
├── pytest.ini                 # Pytest config: markers, timeout, paths
├── requirements.txt           # Runtime dependencies
└── README.md                  # This file

Architecture Overview

Alloy Phase 2 inverts the premise of Phase 1. Rather than building alignment machinery on top of modality-specific embeddings, it asks whether coordinate-value tokenization prevents the modality gap from forming at all. Every signal, regardless of original dimensionality, enters the model as a set of (coordinate, value) pairs. The tokenizer doesn't know whether a patch came from a 1D or 2D signal. It knows the coordinate where the patch was sampled and the values observed there.

Key architectural decisions

  1. Coordinate-value tokenization: Both 1D and 2D patches are encoded as [coord_embed || value_embed] tokens. Fourier feature encoding on coordinates. Shared coordinate projection across dimensionalities. Separate value projections per dimensionality (necessary because raw patch sizes differ).

  2. Gradient separation: Transformer and alignment parameters have structurally separated gradient paths. The transformer optimizer (opt_t) only steps on reconstruction gradients. The alignment optimizer (opt_a) only steps on InfoNCE and VICReg gradients. This prevents alignment losses from distorting the reconstruction backbone.

  3. Interleaved architecture (exp_004): 4 transformer blocks interleaved with 3 alignment blocks. Alignment blocks use cross-modal cross-attention. No modality label is ever injected into the token stream.

  4. Staged training: 4 stages (validation → warmup → training → evaluation), each with own lifecycle hooks. Validation catches NaN/exploding gradients before they waste compute. Warmup does linear LR interpolation. Training uses cosine decay with early stopping. Evaluation runs the full suite once.

  5. Hardware auto-detection: CUDA, MPS, and CPU with capability-specific presets. Mixed precision where available. Graceful MPS fallback for unsupported operations.

For the full architecture document, see alloy/docs/architecture.md.

Docker / CI

Standard Docker

docker build -t alloy:latest .
docker run -it alloy:latest alloy train --experiment 004 --steps 100

RunPod Docker

# Build
docker build -f Dockerfile.runpod -t alloy-exp004:latest .

# Run on RunPod (GPU required)
docker run --gpus all \
    -e S3_BUCKET=my-bucket \
    -e S3_PREFIX=exp004 \
    alloy-exp004:latest

CI (GitHub Actions)

# .github/workflows/test.yml
# Runs on push and PR to main
# Python 3.10 + 3.11, smoke tests, 10 min timeout
# pip install -e "alloy/[dev]", then pytest -v -m smoke

Environment Variables

RUNPOD_MODE=1       # Enable RunPod optimizations (heartbeat, S3 sync, signal handlers)
S3_BUCKET=my-bucket # S3 bucket for checkpoint persistence
S3_PREFIX=exp004    # S3 key prefix for organizing checkpoints
MONITOR_PORT=8080   # Enable HTTP monitoring on this port (with RunPod mode)

When RUNPOD_MODE=1 is set (or --runpod-mode is passed), the training harness:

  • Starts a RunPodMonitor that writes heartbeat files for the RunPod daemon
  • Optionally starts an HTTP server on MONITOR_PORT for health/metrics endpoints
  • Registers SIGINT/SIGTERM handlers for graceful checkpoint-on-exit
  • Syncs checkpoints to S3 after each save (via CheckpointManager.s3_sync())

Legacy CLI (Phase 2)

The phases/phase2/ entry point still works for backwards compatibility:

cd phases/phase2

# Train
python main.py --experiment 004

# Evaluate
python main.py --experiment 004 --eval-only

# Resume
python main.py --experiment 004 --resume auto

# RunPod mode (monitoring, signal handlers, S3 sync)
python main.py --experiment 004 --resume auto --runpod-mode

Note: This is the legacy interface. For new workflows, use the alloy CLI documented above.

Troubleshooting

NaN loss

Symptom: recon_loss becomes NaN or inf during validation or training.

Causes and fixes:

  • Learning rate too high: Reduce lr_transformer and lr_alignment in your config.
  • Gradient explosion: Enable gradient clipping (grad_clip in training config). Check gradient norms in the validation stage output.
  • Mask ratio too high: With mask_ratio > 0.5, there may be too few visible tokens to reconstruct. Try mask_ratio=0.3 or lower.
  • Numerical instability in tau: Set tau_start lower (e.g. 0.5 instead of 1.0). Extremely high tau can cause InfoNCE overflow.
  • Run with --auto-memory to apply conservative presets.

Zero positive pairs

Symptom: All contrastive pair labels are negative (no anchors find nearby positives).

Causes and fixes:

  • local_sim_threshold too high: The similarity threshold for considering two tokens "close" may be too strict. Try local_sim_threshold=0.3 or 0.1.
  • rho too small: The InfoNCE temperature mixing ratio controls how much local vs global structure is considered. Try rho=0.2 or rho=0.3.
  • Patch size mismatch: If 1D and 2D patch sizes differ dramatically, their value embeddings may never be comparable. Ensure value_embed_dim is the same for both modalities.

MPS fallback

Symptom: Training on Apple Silicon produces warnings like MPS operation not implemented, falling back to CPU.

Cause: The MPS backend has incomplete operator coverage. Some operations (e.g. certain indexing patterns, torch.unique, some reductions) don't have MPS kernels.

What happens: The fallback_decorator in alloy/hardware/fallbacks.py automatically catches RuntimeError on MPS, runs the operation on CPU, and moves the result back to MPS. A warning is logged each time. This is normal behavior on Apple Silicon and doesn't affect correctness.

Import errors

Symptom: ModuleNotFoundError: No module named 'alloy' or ImportError.

Fixes:

  • Ensure alloy is installed: pip install -e alloy/
  • Verify .venv is activated: source .venv/bin/activate
  • If using legacy phases/phase2/ entry point, run from the phases/phase2/ directory

Other common issues

"No experiments registered":

alloy --list-experiments
# Expected: 001, 002, 003, 004, 005

If experiments are missing, check that alloy/experiments/exp_*/experiment.py files exist and are importable.

CUDA out of memory: Use --auto-memory to apply a memory-appropriate preset, or reduce batch_size in your config.

Checkpoint not found on resume: The --resume flag looks for the latest checkpoint in results_dir (default: alloy/results/). Verify the path exists and contains checkpoint.pt files.

Documentation Links

About

Unified cross-modal representation learning — one transformer, one embedding space, any signal dimensionality.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Used by

Contributors

Languages