Hi Sam,
I'm reading up on the tool suite of PPI predictors you developed over the last years.
I was wondering on which PPIs exactly you trained (and evaluated) D-SCRIPT, Topsy-Turvy and TT3D.
In the TT paper you reference the D-SCRIPT paper. In the D-SCRIPT paper you write
we use data from the STRING DB (version 11) Szklarczyk et al., 2019. [...] In order to select only high-confidence physical protein interactions, we limited our positive examples to binding interactions associated with a positive experimental-evidence score.
STRING v11 does not distinguish between functional associations and physical interactions. I assume you nevertheless mean v11.
In TT3D you write
We evaluate TT3D in the same cross-species setting where D-SCRIPT and Topsy-Turvy were originally tested. Following (Sledzieski et al. 2021), TT3D was trained and validated on known human PPI from the STRING database (Szklarczyk et al. 2021), filtered for experimentally determined physical binding interactions.
Then, the best model trained on human PPIs was tested on known interactions from other model organisms such as mouse (Mus musculus), fly (Drosophila melanogaster), roundworm (Caenorhabditis elegans), Escherichia coli, and brewer’s yeast (Saccharomyces cerevisiae), also from STRING.
In this case, you refer to STRING v11.5. Did you use the physical interaction subset of STRING or all of STRING (combination of functional associations and physical interactions)?
I'm wondering if you re-trained D-SCRIPT and TT on STRING v11.5 data for the comparison with TT3D?
Lastly, did you include data from homology transfer in STRING? I.e., did you use the "experiments" column in 9606.protein.links.full.v11.0.txt.gz (127.6 Mb) or the "experimental" column in 9606.protein.links.detailed.v11.0.txt.gz (110.1 Mb) to filter for PPIs with experimental evidence? Same question for v11.5.
You have some files in https://github.com/samsledje/D-SCRIPT/tree/main/data/pairs but I'm not sure what exactly they are.
Same question goes for the MINT preprint if you have any insight. Looks like an exciting approach!
Many thanks,
Henrietta
Hi Sam,
I'm reading up on the tool suite of PPI predictors you developed over the last years.
I was wondering on which PPIs exactly you trained (and evaluated) D-SCRIPT, Topsy-Turvy and TT3D.
In the TT paper you reference the D-SCRIPT paper. In the D-SCRIPT paper you write
STRING v11 does not distinguish between functional associations and physical interactions. I assume you nevertheless mean v11.
In TT3D you write
In this case, you refer to STRING v11.5. Did you use the physical interaction subset of STRING or all of STRING (combination of functional associations and physical interactions)?
I'm wondering if you re-trained D-SCRIPT and TT on STRING v11.5 data for the comparison with TT3D?
Lastly, did you include data from homology transfer in STRING? I.e., did you use the "experiments" column in 9606.protein.links.full.v11.0.txt.gz (127.6 Mb) or the "experimental" column in 9606.protein.links.detailed.v11.0.txt.gz (110.1 Mb) to filter for PPIs with experimental evidence? Same question for v11.5.
You have some files in https://github.com/samsledje/D-SCRIPT/tree/main/data/pairs but I'm not sure what exactly they are.
Same question goes for the MINT preprint if you have any insight. Looks like an exciting approach!
Many thanks,
Henrietta