|
15 | 15 | #' @param power power to raise test stat to |
16 | 16 | #' @param keep.boots Should the bootstrap values be saved in the output? |
17 | 17 | #' @param keep.samples Should the samples be saved in the output? |
| 18 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 19 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 20 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
18 | 21 | #' @return Output is a length 2 Vector with test stat and p-value in that order. That vector has 3 attributes -- the sample sizes of each sample, and the number of bootstraps performed for the pvalue. |
19 | 22 | #' @details The KS test compares two ECDFs by looking at the maximum difference between them. Formally -- if E is the ECDF of sample 1 and F is the ECDF of sample 2, then \deqn{KS = max |E(x)-F(x)|^p} for values of x in the joint sample. The test p-value is calculated by randomly resampling two samples of the same size using the combined sample. |
20 | 23 | #' |
|
52 | 55 | #' @param power power to raise test stat to |
53 | 56 | #' @param keep.boots Should the bootstrap values be saved in the output? |
54 | 57 | #' @param keep.samples Should the samples be saved in the output? |
| 58 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 59 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 60 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
55 | 61 | #' @return Output is a length 2 Vector with test stat and p-value in that order. That vector has 3 attributes -- the sample sizes of each sample, and the number of bootstraps performed for the pvalue. |
56 | 62 | #' @details The Kuiper test compares two ECDFs by looking at the maximum positive and negative difference between them. Formally -- if E is the ECDF of sample 1 and F is the ECDF of sample 2, then \deqn{KUIPER = |max_x E(x)-F(x)|^p + |max_x F(x)-E(x)|^p}. The test p-value is calculated by randomly resampling two samples of the same size using the combined sample. |
57 | 63 | #' |
|
90 | 96 | #' @param power power to raise test stat to |
91 | 97 | #' @param keep.boots Should the bootstrap values be saved in the output? |
92 | 98 | #' @param keep.samples Should the samples be saved in the output? |
| 99 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 100 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 101 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
93 | 102 | #' @return Output is a length 2 Vector with test stat and p-value in that order. That vector has 3 attributes -- the sample sizes of each sample, and the number of bootstraps performed for the pvalue. |
94 | 103 | #' @details The CVM test compares two ECDFs by looking at the sum of the squared differences between them -- evaluated at each point in the joint sample. Formally -- if E is the ECDF of sample 1 and F is the ECDF of sample 2, then \deqn{CVM = \sum_{x\in k}|E(x)-F(x)|^p}{CVM = SUM_(x in k) |E(x)-F(x)|^p} where k is the joint sample. The test p-value is calculated by randomly resampling two samples of the same size using the combined sample. Intuitively the CVM test improves on KS by using the full joint sample, rather than just the maximum distance -- this gives it greater power against shifts in higher moments, like variance changes. |
95 | 104 | #' |
|
128 | 137 | #' @param power power to raise test stat to |
129 | 138 | #' @param keep.boots Should the bootstrap values be saved in the output? |
130 | 139 | #' @param keep.samples Should the samples be saved in the output? |
| 140 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 141 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 142 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
131 | 143 | #' @return Output is a length 2 Vector with test stat and p-value in that order. That vector has 3 attributes -- the sample sizes of each sample, and the number of bootstraps performed for the pvalue. |
132 | 144 | #' @details The AD test compares two ECDFs by looking at the weighted sum of the squared differences between them -- evaluated at each point in the joint sample. The weights are determined by the variance of the joint ECDF at that point, which peaks in the middle of the joint distribution (see figure below). Formally -- if E is the ECDF of sample 1, F is the ECDF of sample 2, and G is the ECDF of the joint sample then \deqn{AD = \sum_{x \in k} \left({|E(x)-F(x)| \over \sqrt{2G(x)(1-G(x))/n} }\right)^p }{AD = SUM_(x in k) (|E(x)-F(x)|/sqrt(2G(x)*(1-G(x)))/n)^p} where k is the joint sample. The test p-value is calculated by randomly resampling two samples of the same size using the combined sample. Intuitively the AD test improves on the CVM test by giving lower weight to noisy observations. |
133 | 145 | #' |
|
169 | 181 | #' @param power power to raise test stat to |
170 | 182 | #' @param keep.boots Should the bootstrap values be saved in the output? |
171 | 183 | #' @param keep.samples Should the samples be saved in the output? |
| 184 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 185 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 186 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
172 | 187 | #' @return Output is a length 2 Vector with test stat and p-value in that order. That vector has 3 attributes -- the sample sizes of each sample, and the number of bootstraps performed for the pvalue. |
173 | 188 | #' @details The Wasserstein test compares two ECDFs by looking at the Wasserstein distance between the two. This is of course the area between the two ECDFs. Formally -- if E is the ECDF of sample 1 and F is the ECDF of sample 2, then \deqn{WASS = \int_{x \in R} |E(x)-F(x)|^p}{WASS = Integral |E(x)-F(x)|^p} across all x. The test p-value is calculated by randomly resampling two samples of the same size using the combined sample. Intuitively the Wasserstein test improves on CVM by allowing more extreme observations to carry more weight. At a higher level -- CVM/AD/KS/etc only require ordinal data. Wasserstein gains its power because it takes advantages of the properties of interval data -- i.e. the distances have some meaning. |
174 | 189 | #' |
|
207 | 222 | #' @param power also the power to raise the test stat to |
208 | 223 | #' @param keep.boots Should the bootstrap values be saved in the output? |
209 | 224 | #' @param keep.samples Should the samples be saved in the output? |
| 225 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 226 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 227 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
210 | 228 | #' @return Output is a length 2 Vector with test stat and p-value in that order. That vector has 3 attributes -- the sample sizes of each sample, and the number of bootstraps performed for the pvalue. |
211 | 229 | #' @details The DTS test compares two ECDFs by looking at the reweighted Wasserstein distance between the two. See the companion paper at [arXiv:2007.01360](https://arxiv.org/abs/2007.01360) or <https://codowd.com/public/DTS.pdf> for details of this test statistic, and non-standard uses of the package (parallel for big N, weighted observations, one sample tests, etc). |
212 | 230 | #' |
|
245 | 263 | #' @description (**Warning!** This function has changed substantially between v1.2.0 and v2.0.0) This function takes a two-sample test statistic and produces a function which performs randomization tests (sampling with replacement) using that test stat. This is an internal function of the `twosamples` package. |
246 | 264 | #' @param test_stat_function a function of the joint vector and a label vector producing a positive number, intended as the test-statistic to be used. |
247 | 265 | #' @param default.p This allows for some introduction of defaults and parameters. Typically used to control the power functions raise something to. |
| 266 | +#' @param weights.a Weights for observations in sample a. Not currently implemented -- reserved for future use. |
| 267 | +#' @param weights.b Weights for observations in sample b. Not currently implemented -- reserved for future use. |
| 268 | +#' @param paired Logical. If TRUE, performs a paired test where samples are assumed to be in corresponding order. Samples must have equal length. |
248 | 269 | #' @return This function returns a function which will perform permutation tests on given test stat. |
249 | 270 | #' @details test_stat_function must be structured to take two vectors -- the first a combined sample vector and the second a logical vector indicating which sample each value came from, as well as a third and fourth value. i.e. (fun = function(jointvec,labelvec,val1,val2) ...). See examples. |
250 | 271 | #' |
|
0 commit comments