Problem / motivation
The navigation-input design defers encoder pretraining. Pretraining only the map encoder would leave the rest of the encoder stack under a different initialization strategy.
Proposed solution
Evaluate coordinated pretraining for the camera, semantic-map, route, temporal-history, and world-model encoders. Compare the current initialization with frozen, partially frozen, and end-to-end fine-tuned pretrained variants, and report convergence and downstream KITScenes metrics.
Relevant references: MAE and I-JEPA.
Problem / motivation
The navigation-input design defers encoder pretraining. Pretraining only the map encoder would leave the rest of the encoder stack under a different initialization strategy.
Proposed solution
Evaluate coordinated pretraining for the camera, semantic-map, route, temporal-history, and world-model encoders. Compare the current initialization with frozen, partially frozen, and end-to-end fine-tuned pretrained variants, and report convergence and downstream KITScenes metrics.
Relevant references: MAE and I-JEPA.