Skip to content

Enhancements: Simulations and Actor Critic Approach #2

Description

@ChinarCypher

Enhanced JEPA

Add EMA target encoder (V-JEPA style) with update_target_encoder()
Add Predictor head (residual MLP) for self-supervised JEPA loss
Add encode_batch() for efficient transition training

Realistic Simulation

Nomoto 1st-order heading model (TAU_PSI=3.0) to replace the toy heading += yaw * 5
Ocean current + weather modes (calm/choppy/storm)
36-ray LIDAR simulation and AIS contact list returned in info dict every step
COLREGs encounter classification (head-on / crossing / overtaking) in info dict
Rich reward shaping: progress reward + success bonus + collision penalty + step cost

LIDAR/SONAR Alternative

The new module that makes MarlinNet viable as a LIDAR replacement:
LidarEncoder — 1D CNN over 36 LIDAR rays
AISEncoder — permutation-invariant Set Transformer for variable-N vessel contacts
CrossModalFusion — multi-head cross-attention fuses all 4 modalities into 387-dim latent
Gracefully degrades to camera-only if LIDAR/AIS unavailable

Full COLREGs

COLREGsEngine to implement Rules 13–17 with urgent/normal distance tiers
VelocityObstacleAvoidance to sample 64 candidate velocities and picks the safest one nearest the goal direction

Actor-Critic

Use full 387-dim latent (not just 384)
Squashed Gaussian actor for SAC-compatible stochastic sampling
Critic head for PPO value estimation
Orthogonal weight init throughout

End-to-End Training Loop

TransitionBuffer circular replay buffer with episode boundary markers
TransitionTrainer.run() alternates collect/train with cosine LR schedule, gradient clipping, and validation logging

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or requesthelp wantedExtra attention is needed

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions