The author previously mentioned in the issue that your current open-source code is the third row in the table.
However, I used your pretrain model and conducted sft on this basis, and the resulting effect is quite different from the data in your table.
I would like to know if during the sft process, drivelm was first used for vqa training for 12 epochs, and then nuscenes was used for trajectory planning training.
Has anyone reproduced the results in the table? Could we have a detailed discussion? Thank you!

The author previously mentioned in the issue that your current open-source code is the third row in the table.
However, I used your pretrain model and conducted sft on this basis, and the resulting effect is quite different from the data in your table.
I would like to know if during the sft process, drivelm was first used for vqa training for 12 epochs, and then nuscenes was used for trajectory planning training.
Has anyone reproduced the results in the table? Could we have a detailed discussion? Thank you!