Unsloth does not support context parallel for Qwen3.5 GatedDeltaNet, and it does not shard that hybrid model with FSDP2.
Qwen3.5-35B-A3B has 40 decoder layers, and 30 of them are GatedDeltaNet. The convolution and the gated delta rule are recurrent in time, so a rank that only holds a slice of the sequence cannot run them. Ring-style rotation of K and V cannot run that recurrence either. The other 10 layers are full attention. On one 96GB GPU the bf16 weights already occupy about 65 GiB, and the first LoRA step runs out of memory at sequence length 32768. The same thing happens at 65536.
Unsloth does not support context parallel for Qwen3.5 GatedDeltaNet, and it does not shard that hybrid model with FSDP2.
Qwen3.5-35B-A3B has 40 decoder layers, and 30 of them are GatedDeltaNet. The convolution and the gated delta rule are recurrent in time, so a rank that only holds a slice of the sequence cannot run them. Ring-style rotation of K and V cannot run that recurrence either. The other 10 layers are full attention. On one 96GB GPU the bf16 weights already occupy about 65 GiB, and the first LoRA step runs out of memory at sequence length 32768. The same thing happens at 65536.