-
Notifications
You must be signed in to change notification settings - Fork 4
Expand file tree
/
Copy pathREADME.Rmd
More file actions
188 lines (124 loc) · 12 KB
/
Copy pathREADME.Rmd
File metadata and controls
188 lines (124 loc) · 12 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
---
output: github_document
---
<!-- README.md is generated from README.Rmd. Please edit that file -->
```{r, include = FALSE}
knitr::opts_chunk$set(
collapse = TRUE,
comment = "#>",
fig.path = "man/figures/README-",
out.width = "100%"
)
```
# MPLNClust
Finite Mixtures of Multivariate Poisson-Log Normal Model for Clustering Count Data
<!-- badges: start -->
<!-- https://www.codefactor.io/repository/github/anjalisilva/MPLNClust/issues -->
<!-- [](https://www.codefactor.io/repository/github/anjalisilva/mplnclust)-->
[](https://github.com/anjalisilva/MPLNClust/issues) [](./LICENSE)  
<!-- https://shields.io/category/license -->
<!-- badges: end -->
## Description
`MPLNClust` is an R package for performing clustering using finite mixtures of multivariate Poisson-log normal (MPLN) distribution proposed by [Silva et al., 2019](https://pubmed.ncbi.nlm.nih.gov/31311497/). It was developed for count data, with clustering of RNA sequencing data as a motivation. However, the clustering method may be applied to other types of count data. The package provides functions for functions for parameter estimation via 1) an MCMC-EM framework by [Silva et al., 2019](https://pubmed.ncbi.nlm.nih.gov/31311497/) and 2) a variational Gaussian approximation with EM algorithm by [Subedi and Browne, 2020](https://doi.org/10.1002/sta4.310). Information criteria (AIC, BIC, AIC3 and ICL) and slope heuristics (Djump and DDSE, if more than 10 models are considered) are offered for model selection. Also included are functions for simulating data from this model and visualization.
## Installation
To install the latest version of the package:
``` r
require("devtools")
devtools::install_github("anjalisilva/MPLNClust", build_vignettes = TRUE)
library("MPLNClust")
```
To run the Shiny app:
``` r
MPLNClust::runMPLNClust()
```
## Overview
To list all functions available in the package:
``` r
ls("package:MPLNClust")
```
`MPLNClust` contains 14 functions.
1. __*mplnVariational*__ for carrying out clustering of count data using mixtures of MPLN via variational expectation-maximization
2. __*mplnMCMCParallel*__ for carrying out clustering of count data using mixtures of MPLN via a Markov chain Monte Carlo expectation-maximization algorithm (MCMC-EM) with parallelization
3. __*mplnMCMCNonParallel*__ for carrying out clustering of count data using mixtures of MPLN via a Markov chain Monte Carlo expectation-maximization algorithm (MCMC-EM) with no parallelization
4. __*mplnDataGenerator*__ for the purpose of generating simlulation data via mixtures of MPLN
5. __*mplnVisualizeAlluvial*__ for visualizing clustering results as Alluvial plots
6. __*mplnVisualizeBar*__ for visualizing clustering results as bar plots
7. __*mplnVisualizeHeatmap*__ for visualizing clustering results as heatmaps
8. __*mplnVisualizeLine*__ for visualizing clustering results as line plots
9. __*AICFunction*__ for model selection
10. __*AIC3Function*__ for model selection
11. __*BICFunction*__ for model selection
12. __*ICLFunction*__ for model selection
13. __*runMPLNClust*__ is the shiny implementation of *mplnVariational*
14. __*mplnVarClassification*__ is an implementation for classification is currently under construction
Framework of __*mplnVariational*__ makes it computationally efficient and faster compared to __*mplnMCMCParallel*__ or __*mplnMCMCNonParallel*__. Therefore, __*mplnVariational*__ may perform better for large datasets. For more information, see details section below. An overview of the package is illustrated below:
<div style="text-align:center"><img src="inst/extdata/Overview_MPLNClust.png" width="800" height="450"/>
<div style="text-align:left">
<div style="text-align:left">
## Details
The MPLN distribution ([Aitchison and Ho, 1989](https://www.jstor.org/stable/2336624?seq=1)) is a multivariate log normal mixture of independent Poisson distributions. The hidden layer of the MPLN distribution is a multivariate Gaussian distribution, which allows for the specification of a covariance structure. Further, the MPLN distribution can account for overdispersion in count data. Additionally, the MPLN distribution supports negative and positive correlations.
A mixture of MPLN distributions is introduced for clustering count data by [Silva et al., 2019](https://pubmed.ncbi.nlm.nih.gov/31311497/). Here, applicability is illustrated using RNA sequencing data. To this date, two frameworks have been proposed for parameter estimation: 1) an MCMC-EM framework by [Silva et al., 2019](https://pubmed.ncbi.nlm.nih.gov/31311497/) and 2) a variational Gaussian approximation with EM algorithm by [Subedi and Browne, 2020](https://doi.org/10.1002/sta4.310).
### MCMC-EM Framework for Parameter Estimation
[Silva et al., 2019](https://pubmed.ncbi.nlm.nih.gov/31311497/) used an MCMC-EM framework via Stan for parameter estimation. This method is employed in functions __*mplnMCMCParallel*__ and __*mplnMCMCNonParallel*__.
Coarse grain parallelization is employed in *mplnMCMCParallel*, such that when a range of components/clusters (g = 1,...,G) are considered, each component/cluster size is run on a different processor. This can be performed because each component/cluster size is independent from another. All components/clusters in the range to be tested have been parallelized to run on a separate core using the *parallel* R package. The number of cores used for clustering is calculated using *parallel::detectCores() - 1*. No internal parallelization is performed for *mplnMCMCNonParallel*.
To check the convergence of MCMC chains, the potential scale reduction factor and the effective number of samples are used. The Heidelberger and Welch’s convergence diagnostic (Heidelberger and Welch, 1983) is used to check the convergence of the MCMC-EM algorithm. Starting values (argument: *initMethod*) and the number of iterations for each chain (argument: *nIterations*) play an important role for the successful operation of this algorithm.
### Variational-EM Framework for Parameter Estimation
[Subedi and Browne, 2020](https://doi.org/10.1002/sta4.310) proposed a variational Gaussian approximation that alleviates challenges of MCMC-EM algorithm. Here the posterior distribution is approximated by minimizing the Kullback-Leibler (KL) divergence between the true and the approximating densities. A variational-EM based framework is used for parameter estimation. This algorithm is implemented in the function __*mplnVariational*__. The parsimonious family of models implemented by considering eigen-decomposition of covariance matrix in [Subedi and Browne, 2020](https://doi.org/10.1002/sta4.310) is not yet available with this package.
## Model Selection and Other Details
Four model selection criteria are offered, which include the Akaike information criterion (AIC; Akaike, 1973), the Bayesian information
criterion (BIC; Schwarz, 1978), a variation of the AIC used by Bozdogan (1994) called AIC3, and the integrated completed likelihood (ICL; Biernacki et al., 2000). Slope heuristics (Djump and DDSE; Arlot et al., 2016) could be used for model selection if more than 10 models are considered.
Starting values (argument: *initMethod*) and the number of iterations for each chain (argument: *nInitIterations*) play an important role to
the successful operation of this algorithm. There maybe issues with singularity, in which case altering starting values or initialization
method may help.
## Shiny App
The Shiny app employing __*mplnVariational*__ could be run and results could be visualized:
``` r
MPLNClust::runMPLNClust()
```
<div style="text-align:center"><img src="inst/extdata/ShinyAppMPLNClust1.png" alt="ShinyApp1" width="650" height="400"/>
<div style="text-align:left">
<div style="text-align:left">
In simple, the __*runMPLNClust*__ is a web applications available with `MPLNClust`.
## Tutorials
For tutorials and plot interpretation, refer to the vignette:
``` r
browseVignettes("MPLNClust")
```
## Citation for Package
``` r
citation("MPLNClust")
```
Silva, A., S. J. Rothstein, P. D. McNicholas, and S. Subedi (2019). A multivariate Poisson-log normal mixture model for clustering transcriptome sequencing data. *BMC Bioinformatics*. 20(1):394.
``` r
A BibTeX entry for LaTeX users is
@Article{,
title = {A multivariate Poisson-log normal mixture model for clustering transcriptome sequencing data},
author = {A. Silva and S. J. Rothstein and P. D. McNicholas and S. Subedi},
journal = {BMC Bioinformatics},
year = {2019},
volume = {20},
number = {1},
pages = {394},
url = {https://pubmed.ncbi.nlm.nih.gov/31311497/},
}
```
## Package References
* [Silva, A., S. J. Rothstein, P. D. McNicholas, and S. Subedi (2019). A multivariate Poisson-log normal mixture model for clustering transcriptome sequencing data. *BMC Bioinformatics.*](https://pubmed.ncbi.nlm.nih.gov/31311497/)
* [Subedi, S., R.P. Browne (2020). A family of parsimonious mixtures of multivariate Poisson-lognormal distributions for clustering multivariate count data. *Stat.* 9:e310.](https://doi.org/10.1002/sta4.310)
## Other References
- [Aitchison, J. and C. H. Ho (1989). The multivariate Poisson-log normal distribution. *Biometrika.*](https://www.jstor.org/stable/2336624?seq=1)
- [Akaike, H. (1973). Information theory and an extension of the maximum likelihood principle. In *Second International Symposium on Information Theory*, New York, NY, USA, pp. 267–281. Springer Verlag.](https://link.springer.com/chapter/10.1007/978-1-4612-1694-0_15)
- [Arlot, S., Brault, V., Baudry, J., Maugis, C., Michel, B. (2016). _capushe: CAlibrating Penalities Using Slope HEuristics_. R package version 1.1.1](https://CRAN.R-project.org/package=capushe)
- [Biernacki, C., G. Celeux, and G. Govaert (2000). Assessing a mixture model for clustering with the integrated classification likelihood. *IEEE Transactions on Pattern Analysis and Machine Intelligence* 22.](https://hal.inria.fr/inria-00073163/document)
- [Bozdogan, H. (1994). Mixture-model cluster analysis using model selection criteria and a new informational measure of complexity. In *Proceedings of the First US/Japan Conference on the Frontiers of Statistical Modeling: An Informational Approach: Volume 2 Multivariate Statistical Modeling*, pp.69–113. Dordrecht: Springer Netherlands.](https://link.springer.com/chapter/10.1007/978-94-011-0800-3_3)
- [Ghahramani, Z. and Beal, M. (1999). Variational inference for bayesian mixtures of factor analysers. *Advances in neural information processing systems* 12.](https://cse.buffalo.edu/faculty/mbeal/papers/nips99.pdf)
- [Robinson, M.D., and Oshlack, A. (2010). A scaling normalization method for differential expression analysis of RNA-seq data. *Genome Biology* 11, R25.](https://genomebiology.biomedcentral.com/articles/10.1186/gb-2010-11-3-r25)
- [Schwarz, G. (1978). Estimating the dimension of a model. *The Annals of Statistics* 6.](https://www.jstor.org/stable/2958889?seq=1)
- [Wainwright, M. J. and Jordan, M. I. (2008). Graphical models, exponential families, and variational inference. *Foundations and Trends® in Machine Learning* 1.](https://onlinelibrary.wiley.com/doi/abs/10.1002/sta4.310)
## Maintainer
* Anjali Silva (anjali@alumni.uoguelph.ca).
## Contributions
`MPLNClust` welcomes issues, enhancement requests, and other contributions. To submit an issue, use the [GitHub issues](https://github.com/anjalisilva/MPLNClust/issues).
## Acknowledgments
* Dr. Marcelo Ponce, SciNet HPC Consortium, University of Toronto, ON, Canada for all the computational support.
* This work was funded by Natural Sciences and Engineering Research Council of Canada, Queen Elizabeth II Graduate Scholarship, and Arthur Richmond Memorial Scholarship.