I can train or finetune the model to predict text from a silent video. Can I somehow get the duration of each predicted text token in the input video?
I can train or finetune the model to predict text from a silent video. Can I somehow get the duration of each predicted text token in the input video?