Hi LION team,
Thanks for the really interesting paper.
I see in your repo you have shared the code for MLM and image classification. Could you also share the code for LRA as well? I'd like to take a look at your implementation as I don't quite understand how you got the performance bump over vanilla transformer or vanilla linear attention. Both saturate at ~40% on listops whereas LION-S works extremely well for some cases and not so much on others (ref. Table 6 in your manuscript).
Thanks,
Vedant
Hi LION team,
Thanks for the really interesting paper.
I see in your repo you have shared the code for MLM and image classification. Could you also share the code for LRA as well? I'd like to take a look at your implementation as I don't quite understand how you got the performance bump over vanilla transformer or vanilla linear attention. Both saturate at ~40% on listops whereas LION-S works extremely well for some cases and not so much on others (ref. Table 6 in your manuscript).
Thanks,
Vedant