Thank you so much for your amazing work! When running the code, we noticed that two VAR models of different scales are directly employed as the drafter and refiner models without additional training. This raises the question: why can these two models be used together seamlessly without further fine-tuning?
Thank you so much for your amazing work! When running the code, we noticed that two VAR models of different scales are directly employed as the drafter and refiner models without additional training. This raises the question: why can these two models be used together seamlessly without further fine-tuning?