Skip to content
#

continuous-depth

Here are 2 public repositories matching this topic...

Language: All
Filter by language

Neural ODE Transformer in Julia, trained via adjoint sensitivity methods. Matched-architecture vs. discrete Transformer on Penn Treebank: 119.9 val perplexity vs 113.7 discrete. Single-seed result; discrete wins at this scale as expected, continuous-depth advantage expected to emerge at larger scale.

  • Updated Jun 30, 2026
  • Julia

Improve this page

Add a description, image, and links to the continuous-depth topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the continuous-depth topic, visit your repo's landing page and select "manage topics."

Learn more