Best prior for large cyclic/lipopeptide generation & validity drop during TL #319
|
Hi REINVENT team, Thanks for the great work. I’m working on generating large cyclic peptides (1,000–1,300 Da) with a hydrophobic lipid tail, similar to polymyxins/octapeptins, and I’m unsure which prior is most suitable for this chemical space. The standard ChEMBL prior and Mol2Mol prior both seem far from this domain, and when I try transfer learning on my peptide dataset, the SMILES validity drops dramatically (sometimes <1%), with the model often collapsing to repetitive structures even with gentle training. My molecules have long SMILES strings, ring closures, multiple amide bonds, and branching, so I suspect the chemical space is outside the original prior distribution and that the default filters/tokenization might not handle large macrocycles well. Could you advise on (1) which prior is best for macrocyclic/lipopeptide structures and (2) why validity decreases so much during TL, and whether there are recommended settings or workflows for handling long cyclic peptides in REINVENT? Thanks! |
Have you had a look into PepINVENT (the prior is on Zenodo) to see if this works for you? Macrocyclic peptides were part of the training set, I believe.