Hi tin2tin — disclosure first: I work on Sonilo (licensed video-to-music), so this suggestion involves our product.
I saw the remote-backend adapter work land in June (Custom URL adapters, the fal.ai connector). That opens a door for a kind of audio Pallaidium doesn't have yet: music generated from the rendered strip itself, rather than from a text prompt the way MusicGen and AudioLDM2 work. You render the VSE sequence, the model watches it, and the returned track follows the cuts and the mood of what's on screen. Length matches the sequence automatically. Tracks are licensed and safe for commercial use (terms apply).
Given your line about editing to "the emotional weight of what you see and hear," music that was generated from the picture felt like a natural fit for the toolbox.
I'd be glad to build the adapter and test it against a real project file. General recipe for the pattern: https://github.com/cindyxu1030/sonilo-video-to-music-cookbook
Hi tin2tin — disclosure first: I work on Sonilo (licensed video-to-music), so this suggestion involves our product.
I saw the remote-backend adapter work land in June (Custom URL adapters, the fal.ai connector). That opens a door for a kind of audio Pallaidium doesn't have yet: music generated from the rendered strip itself, rather than from a text prompt the way MusicGen and AudioLDM2 work. You render the VSE sequence, the model watches it, and the returned track follows the cuts and the mood of what's on screen. Length matches the sequence automatically. Tracks are licensed and safe for commercial use (terms apply).
Given your line about editing to "the emotional weight of what you see and hear," music that was generated from the picture felt like a natural fit for the toolbox.
I'd be glad to build the adapter and test it against a real project file. General recipe for the pattern: https://github.com/cindyxu1030/sonilo-video-to-music-cookbook