Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions .github/workflows/build-book-release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -136,8 +136,10 @@ jobs:
remotes::install_github("HenrikBengtsson/matrixStats", ref="develop")
BiocManager::install("Cairo")
BiocManager::install("biclust")
BiocManager::install("cobiclust")
BiocManager::install("cobiclust")
BiocManager::install("ecodist")
BiocManager::install("HDF5Array")
BiocManager::install("kableExtra")
BiocManager::install("microbiome/microbiomeDataSets")
BiocManager::install("fionarhuang/TreeSummarizedExperiment")
# Remove when available in docker
Expand Down Expand Up @@ -167,7 +169,7 @@ jobs:
run: |
git config --global user.email "action@github.com"
git config --global user.name "GitHub Action"
git config --global --add safe.directory /__w/course_2022_oulu/course_2022_oulu
git config --global --add safe.directory /__w/course_2022_radboud/course_2022_radboud
git fetch --all
commit=$(git --work-tree=../ rev-parse --verify --short HEAD)
git worktree add --track -B gh-pages docs origin/gh-pages
Expand Down
105 changes: 0 additions & 105 deletions 01-program.Rmd
Original file line number Diff line number Diff line change
@@ -1,105 +0,0 @@
# Program

The course takes place daily from 9am – 5pm (CEST), including
coffee and lunch breaks.

We expect that participants will prepare for the course in advance, see section
\@ref(start). Online support is available.


The material follows open online book created by the course teachers,
Orchestrating Microbiome Analysis
https://microbiome.github.io/OMA. This is R/Bioconductor framework for
multi-omic data science.


<img src="fig.png" alt="ML4microbiome" width="50%"/>
<p style="font-size:12px">Figure source: Moreno-Indias _et al_. (2021) [Statistical and Machine Learning Techniques in Human Microbiome Studies: Contemporary Challenges and Solutions](https://doi.org/10.3389/fmicb.2021.635781). Frontiers in Microbiology 12:11.</p>


## Day 1 - Open data science

**Morning session**

9-10 Coffee, Welcome & Practicalities

10-11 Lecture: Open & reproducible workflows

11-12 Demo & hands-on: Introduction to CSC RStudio notebook

12-13 Lunch break


**Afternoon hands-on session**

13-15 Demo: Data science framework

15-17 Hands-on: microbiome data summaries & exploration

17-18 Presentations & Discussion


----------------------------------------------------------------

## Day 2 - Tabular data

**Morning session**

9-10 Lecture: Analysis & visualization of _tabular data_

10-12 Demo & hands-on: Univariate methods

12-13 Lunch break

**Afternoon hands-on session**

13-14 Demo: Multivariate data analysis & visualization

14-17 Hands-on: Multivariate data analysis & visualization

17-18 Presentations & Discussion

----------------------------------------------------------------

## Day 3 - Multi-assay data

**Morning session**

9-10 Lecture: multi-omic data integration

10-12 Demo & hands-on: multi-assay data container

12-13 Lunch break


**Afternoon hands-on session**

13-15: Demo & hands-on: association analysis

13-17: Demo & hands-on: machine learning

17-18 Presentations & Discussion


-----------------------------------------------------------------

## Day 4 - Advanced topics

**Morning session**

9-10 Summary of the learning material

10-12 Demo & hands-on: custom data & advanced tools

12-13 Q & A session


**Afternoon session**

13-14 Lunch

14-16 Wrap-up




68 changes: 2 additions & 66 deletions 01.1-codeofconduct.Rmd
Original file line number Diff line number Diff line change
@@ -1,11 +1,5 @@

# Project-wide Code of Conduct statement for Bioconductor
[link to code of conduct](https://bioconductor.github.io/bioc_coc_multilingual/)


(Adapted from the BioC 2020 Code of Conduct)

## Code of Conduct -Version 1.0.2 (July 27, 2021)
# Code of Conduct

The Bioconductor community values an open approach to science that promotes the

Expand All @@ -15,64 +9,6 @@ The Bioconductor community values an open approach to science that promotes the
- a kind and welcoming environment
- community contributions

In line with these values, Bioconductor is dedicated to providing a welcoming, supportive, collegial experience free of harassment, intimidation, and bullying regardless of:

- identity: gender, gender identity and expression, sexual orientation, disability, physical appearance, ethnicity, body size, race, age, religion, language etc.
- intellectual position: approaches to data analysis, software preferences, coding style, scientific perspective, stage of career, etc.

By participating in this community, you agree not to engage in behavior contrary to these values at any Bioconductor-sponsored event (in person or virtual, including but not limited to talks, workshops, poster sessions, social activities) or electronic communication channel (including but not limited to community-bioc Slack, the support site, online forums, package review site and social media communications). Furthermore, we require all participants to have identifiable accounts in Bioconductor online forums. Accounts that do not adhere to this after request to de-anonymise may be deleted.

We do not tolerate harassment, intimidation, or bullying of community members. Sexual language and imagery are not appropriate in presentations, communications or in online venues, including chats.

Any person/s violating the Code of Conduct may be sanctioned or expelled temporarily or permanently from an electronic platform or event at the discretion of the Code of Conduct committee.

**Examples of unacceptable harassment, intimidation, and bullying behavior**

Harassment includes, but is not limited to:

- Making comments in chats, to an audience or personally, that belittle or demean another person
- Sharing sexual images online
- Harassing photography or recording
- Sustained disruption of talks or other events
- Unwelcome sexual attention
- Advocating for, or encouraging, any of the above behavior

Intimidation and bullying include, but are not limited to:

- Aggressive or browbeating behavior
- Mocking or insulting another person’s intellect, work, perspective, or question/comment
- Making reference to someone’s gender, gender identity and expression, sexual orientation, disability, physical appearance, body size, race, age, religion, or other personal attributes in the context of a scientific discussion
- Deliberately making someone feel unwelcome
- Trolling behaviour (deliberately inflammatory or offensive messages)
- Sustained off-topic posts

## Enforcement

Anyone asked to stop harassing or intimidating behavior are expected to comply immediately.

If a person/s contravene the Code of Conduct the Code of Conduct committee retains the right to take any action that ensures a welcoming environment for all community members. This includes warning the alleged offender or temporary/permanent expulsion from the event and/or electronic platforms under Bioconductor’s control.

The Code of Conduct committee may take action to redress anything designed to, or with the clear impact of, disrupting an event or electronic communication platform or making the environment hostile for any community member.

We expect everyone in the Bioconductor community to comply with the Code of Conduct when participating in Bioconductor events and online communication platforms.

## Reporting

If someone makes you or anyone else feel unsafe or unwelcome, please report it as soon as possible. You can make a report either anonymously or personally. All reports will be reviewed by the Code of Conduct Committee and will be kept confidential.

*Electronically*
You can make an anonymous or non-anonymous report via the following link: https://forms.gle/gEWHBWnXvZbEdFsq5. It is a free-form text box that will be forwarded to the Code of Conduct Committee. Alternatively you can email the Code of Conduct Committee (code-of-conduct@bioconductor.org). If you are uncomfortable reporting to the Code of Conduct committee as a group, you can contact any individual committee member via email or a direct message on the community-bioc Slack channel. Please include screenshots/copies of all relevant electronic conversations whenever possible (you don’t need to compromise your anonymity!).

We can’t follow up an anonymous report with you directly, but we will fully investigate it and take whatever action is necessary to prevent a recurrence.

**Personal Report (for any Bioconductor events: in-person or virtual)**

You can make a personal report to any member of the event Code of Conduct committee present at an event.

When taking a personal report, we will ensure you are safe and cannot be overheard. We may involve other event staff to ensure your report is managed properly. Once safe, we’ll ask you to tell us about what happened. This can be upsetting, but we’ll handle it as respectfully as possible, and you can bring someone to support you. You won’t be asked to confront anyone, and we won’t tell anyone who you are.

Our team will be happy to help you get the relevant support (e.g. help contacting hotel/venue security, local law enforcement, local support services, provide escorts, or otherwise assist you to feel safe for the duration of the event).

We value your attendance and participation at Bioconductor events and in our community.


For the full version, enforcement, and reporting instructions, see the [Bioconductor code of conduct](https://bioconductor.github.io/bioc_coc_multilingual/).
187 changes: 187 additions & 0 deletions 02-data.Rmd
Original file line number Diff line number Diff line change
@@ -0,0 +1,187 @@

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo=TRUE, message=FALSE, warning=FALSE)
```


# Importing microbiome data

This section demonstrates how to import taxonomic profiling data in R.


## Data access

*ADHD-associated changes in gut microbiota and brain in a mouse model*

Tengeler AC _et
al._ (2020) [**Gut microbiota from persons with
attention-deficit/hyperactivity disorder affects the brain in
mice**](https://doi.org/10.1186/s40168-020-00816-x). Microbiome
8:44.

In this study, mice are colonized with microbiota from participants
with ADHD (attention deficit hyperactivity disorder) and healthy
participants. The aim of the study was to assess whether the mice
display ADHD behaviors after being inoculated with ADHD microbiota,
suggesting a role of the microbiome in ADHD pathology.

Download the data from
[data](https://github.com/microbiome/course_2022_radboud/tree/main/data)
subfolder.


## Importing microbiome data in R

**Import example data** by modifying the examples in the online book
section on [data exploration and
manipulation](https://microbiome.github.io/OMA/data-introduction.html#loading-experimental-microbiome-data).

The data files in our example are in _biom_ container, which is a
standard file format for microbiome data. Other file formats exist as
well, and import details vary by platform. Here, we import _biom_ data
files into a specific data container (structure) in R,
_TreeSummarizedExperiment_ (TreeSE) [Huang et
al. (2020)](https://f1000research.com/articles/9-1246).

The data container provides the basis for downstream data analysis in
the _miaverse_ data science framework.


<img src="https://raw.githubusercontent.com/FelixErnst/TreeSummarizedExperiment
/2293440c6e70ae4d6e978b6fdf2c42fdea7fb36a/vignettes/tse2.png" width="100%"/>

**Figure sources:**

**Original article**
- Huang R _et al_. (2021) [TreeSummarizedExperiment: a S4 class
for data with hierarchical structure](https://doi.org/10.12688/
f1000research.26669.2). F1000Research 9:1246.

**Reference Sequence slot extension**
- Lahti L _et al_. (2020) [Upgrading the R/Bioconductor ecosystem for microbiome
research](https://doi.org/10.7490/
f1000research.1118447.1) F1000Research 9:1464 (slides).




### Example solution

Let us first import the biom file into R / TreeSE container.

```{r import, message=FALSE, warning=FALSE}
# Load the mia R package
library(mia)

# Defining file paths
## Biom file (taxonomic profiles)
biom_file_path <- "data/Tengeler2020/Aggregated_humanization2.biom"

# Import the data into SummarizedExperiment container
se <- loadFromBiom(biom_file_path)

# Convert this data to TreeSE container (no direct importer exists)
tse <- as(se, "TreeSummarizedExperiment")
```


Check and clean up rowData (information on the taxonomic features).


```{r rowdata, message=FALSE, warning=FALSE}
# Investigate the rowData of this data object
print(head(rowData(tse)))

# We notice that the rowData fields do not have descriptibve names.
# Hence, let us rename the columns in rowData
names(rowData(tse)) <- c("Kingdom", "Phylum", "Class", "Order",
"Family", "Genus")

# We also notice that the taxa names are of form "c__Bacteroidia" etc.
# Goes through the whole DataFrame. Removes '.*[kpcofg]__' from strings, where [kpcofg]
# is any character from listed ones, and .* any character.
rowdata_modified <- BiocParallel::bplapply(rowData(tse),
FUN = stringr::str_remove,
pattern = '.*[kpcofg]__')

# Genus level has additional '\"', so let's delete that also
rowdata_modified <- BiocParallel::bplapply(rowdata_modified,
FUN = stringr::str_remove,
pattern = '\"')

# rowdata_modified is a list, so convert this back to DataFrame format.
# and assign the cleaned data back to the TSE rowData
rowData(tse) <- DataFrame(rowdata_modified)

# Recheck rowData after the modifications
print(head(rowData(tse)))
```


Next let us add sample information (colData) to our TreeSE object.


```{r coldata, message=FALSE, warning=FALSE}
## Sample phenodata
sample_meta_file_path <- "data/Tengeler2020/Mapping_file_ADHD_aggregated.csv"

# Check what type of data it is
# read.table(sample_meta_file_path)

# It seems like a comma separated file and it does not include headers
# Let us read the file; note that sample names are in the first column
sample_meta <- read.table(sample_meta_file_path, sep=",", header=FALSE, row.names=1)

# Check the data
print(head(sample_meta))

# Add headers for the columns (as they seem to be missing)
colnames(sample_meta) <- c("patient_status", "cohort",
"patient_status_vs_cohort", "sample_name")

# Add this sample data to colData of the taxonomic data object
# Note that the data must be given in a DataFrame format (required for our purposes)
colData(tse) <- DataFrame(sample_meta)

# Check the colData after modifications
print(head(colData(tse)))
```


Add phylogenetic tree (rowTree).

```{r rowtree, message=FALSE, warning=FALSE}
## Phylogenetic tree
tree_file_path <- "data/Tengeler2020/Data_humanization_phylo_aggregation.tre"

# Read the tree file
tree <- ape::read.tree(tree_file_path)

# Add tree to rowTree
rowTree(tse) <- tree
```

We can save the data object as follows.

```{r save, message=FALSE, warning=FALSE}
saveRDS(tse, file="tse.rds")
```


### Loading readily processed data

Alternatively, you can just load the already processed data set.

By using this readily processed data set we can skip the data import
step, and assume that the data has already been appropriately
preprocessed and available in the TreeSE container.


```{r, message=FALSE, warning=FALSE, eval=FALSE}
tse <- readRDS("data/Tengeler2020/tse.rds")
```

In addition to our example data, further demonstration data sets are
available in the TreeSE data container: see [OMA demo data
sets](https://microbiome.github.io/OMA/containers.html#example-data).

Loading