Feature request
Currently, CSV samplesheets only support local filesystem paths in path_prefix. Users working with cloud storage must use the JSON samplesheet format instead. It would be convenient to support cloud URIs (gs://, s3://, etc.) directly in CSV samplesheets.
Motivation
- CSV samplesheets are simpler to generate and more widely used
- Many production pipelines store genotype data in cloud buckets and generate samplesheets dynamically
- The JSON format works as a workaround but requires a different structure (
geno/pheno/variants fields instead of path_prefix)
Current behavior
SamplesheetParser.resolvePath() treats any path not starting with / as relative and resolves it against the samplesheet's parent directory. A path_prefix like gs://bucket/path/sample becomes /local/dir/gs:/bucket/path/sample.
Suggested fix
Detect URI schemes in resolvePath() and return them as-is, skipping relative path resolution:
private def resolvePath(path) {
if (path.matches('^[a-zA-Z][a-zA-Z0-9+.-]*://.*')) {
return path
}
// existing logic...
}
Happy to contribute a PR for this if it's welcome.
Feature request
Currently, CSV samplesheets only support local filesystem paths in
path_prefix. Users working with cloud storage must use the JSON samplesheet format instead. It would be convenient to support cloud URIs (gs://,s3://, etc.) directly in CSV samplesheets.Motivation
geno/pheno/variantsfields instead ofpath_prefix)Current behavior
SamplesheetParser.resolvePath()treats any path not starting with/as relative and resolves it against the samplesheet's parent directory. Apath_prefixlikegs://bucket/path/samplebecomes/local/dir/gs:/bucket/path/sample.Suggested fix
Detect URI schemes in
resolvePath()and return them as-is, skipping relative path resolution:Happy to contribute a PR for this if it's welcome.