Hybrid transformer architecture replacing discrete layers with Neural ODE blocks, enabling continuous semantic control (sentiment, style) at inference time via learned steering signals
-
Updated
Jan 30, 2026 - Jupyter Notebook
Hybrid transformer architecture replacing discrete layers with Neural ODE blocks, enabling continuous semantic control (sentiment, style) at inference time via learned steering signals
Neural ODE Transformer in Julia, trained via adjoint sensitivity methods. Matched-architecture vs. discrete Transformer on Penn Treebank: 119.9 val perplexity vs 113.7 discrete. Single-seed result; discrete wins at this scale as expected, continuous-depth advantage expected to emerge at larger scale.
Add a description, image, and links to the continuous-depth topic page so that developers can more easily learn about it.
To associate your repository with the continuous-depth topic, visit your repo's landing page and select "manage topics."