Skip to content

Commit 7751976

Browse files
committed
V0.11.0
The validation summary now carries the predictor into the result so relative CI widths are computed against the predictor's 10th–90th percentile spread instead of a self-referential bootstrap proxy. Stability tiers are assigned consistently and exposed via a machine-readable `stability` list, and package metadata plus the bilirubin/CRC vignettes were updated to reflect the revised reporting and wording.
1 parent 1f9a96d commit 7751976

71 files changed

Lines changed: 3687 additions & 3422 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎DESCRIPTION‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,7 @@
11
Package: OptSurvCutR
22
Type: Package
33
Title: Optimal Survival Cut-Point Discovery for Time-to-Event Analysis with 'OptSurvCutR'
4-
Version: 0.10.1
4+
Version: 0.11.0
55
Authors@R:
66
c(person(given = "Payton",
77
family = "Yau",
@@ -11,7 +11,8 @@ Authors@R:
1111
person(given = "Suhirthakumar",
1212
family = "Puvanendran",
1313
role = "aut",
14-
email = "kumar.puvan@uwl.ac.uk"))
14+
email = "kumar.puvan@uwl.ac.uk",
15+
comment = c(ORCID = "0000-0003-2346-0483")))
1516
Description: Provides a robust workflow for optimal cut-point analysis in
1617
time-to-event ('survival') data. Functions determine the optimal number
1718
of cut-points via find_cutpoint_number(), find their precise locations

‎NEWS.md‎

Lines changed: 103 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,106 @@
1+
# OptSurvCutR v0.11.0 (2026-09-08)
2+
3+
The validation summary now carries the predictor into the result, so relative confidence
4+
interval widths are computed against the predictor's 10th–90th percentile spread rather
5+
than a self-referential bootstrap proxy. Stability tiers are assigned consistently and
6+
exposed through a machine-readable `stability` list. Both case study vignettes and the
7+
package documentation have been updated to reflect the revised reporting.
8+
9+
## Bug Fixes (Important)
10+
11+
* **Stability Metric Denominator:** `summary.validate_cutpoint_result()` computed the
12+
relative confidence interval width against the bootstrap distribution of the first
13+
cut-point rather than the 10th–90th percentile range of the predictor. The cause was
14+
that `validate_cutpoint()` did not carry the `userdata` element into its returned
15+
object, so the predictor values were unavailable to the summary method and an internal
16+
fallback was silently used instead. Reported relative widths were consequently inflated
17+
by roughly an order of magnitude, and **Stability Tier 1 was unreachable for any
18+
input**. Stability tiers reported by version 0.10.1 and earlier should be recomputed.
19+
20+
* **Removal of Silent Fallbacks:** The two fallback denominators
21+
(`bootstrap_distribution[, 1]` and `max(abs(medians))`) have been removed. Where the
22+
predictor values are unavailable, or where the 10th and 90th percentiles of a highly
23+
discrete predictor coincide, the relative width is now reported as undefined with a
24+
diagnostic message rather than substituted with an unrelated quantity.
25+
26+
* **Single Cut-point Classification:** Models with one cut-point have no adjacent
27+
interval, so interval separation is undefined and `has_distinct_separation` could never
28+
evaluate to `TRUE`. Such models could therefore only be assigned Tier 1 or Tier 4.
29+
Single-threshold models are now graded on relative width alone (< 30% Tier 1, 30–60%
30+
Tier 2, > 60% Tier 4) and can no longer be assigned Tier 3, which requires overlap.
31+
32+
## API Changes
33+
34+
* **Machine-readable Stability Output:** `summary.validate_cutpoint_result()` now attaches
35+
a `stability` list to the returned object, containing `tier`, `tier_label`, `percent`,
36+
`max_relative_width`, `relative_widths` (per cut-point), `data_spread`, `separated` and
37+
`worst_cut`. Tier assignment and interval widths can now be accessed programmatically
38+
instead of parsed from console output, which supports batch screening of multiple
39+
predictors.
40+
41+
* **Predictor Values Retained:** `validate_cutpoint()` now returns the analysis data in
42+
the `userdata` element, consistent with `find_cutpoint()`.
43+
44+
## Documentation
45+
46+
* **Tier Classification Documentation:** Clarified that the four tiers summarise two
47+
independent diagnostics — interval separation and relative width — and do not constitute
48+
an ordinal scale; whether Tier 2 or Tier 3 is preferable depends on whether separation
49+
or boundary precision matters for the intended application. Documented that the 30% and
50+
60% boundaries are pragmatic conventions adopted for interpretability rather than
51+
calibrated thresholds.
52+
53+
* **Diagnostic Messaging:** Where stability cannot be calculated, the message now
54+
identifies the required input (`userdata$factor`) and the conditions under which the
55+
metric is undefined.
56+
57+
* **Console Messages:** Corrected spelling in three user-facing messages in the systematic
58+
search engine, and generalised a hard-coded message to report the number of candidate
59+
positions and cut-points actually being searched.
60+
61+
* **Standards Compliance:** Expanded `@srrstats {G1.1}` to document the methodological
62+
origin of the approach (maximally selected rank statistics, information-theoretic model
63+
selection, multivariable Cox optimisation, and bootstrap resampling), and to clarify
64+
that the multi-threshold problem lies outside the stated scope of `maxstat`, `survminer`
65+
and `cutpointr` rather than representing a deficiency in those packages. Narrowed
66+
`@srrstats {G1.5}` to describe the comparison actually provided.
67+
68+
* **Vignettes:** Both case study vignettes were re-run against the corrected metric and
69+
the reported stability tiers updated accordingly. Guidance on replicate counts,
70+
permutation counts and the interpretation of overlapping confidence intervals has been
71+
expanded, and clinical background added for both cohorts.
72+
73+
* **README:** Replaced the simulated example, in which the information criterion selected
74+
zero cut-points, with the Mayo Clinic PBC analysis reported in the manuscript. Revised
75+
language describing what the bootstrap assessment establishes, and reserved claims of
76+
rOpenSci compliance until review has concluded.
77+
78+
## Testing
79+
80+
* **Regression Coverage:** Added `tests/testthat/test-stability-metric.R`, verifying that
81+
`userdata` survives into the validation object, that the relative width is computed
82+
against the predictor's percentile range and not the bootstrap distribution, that
83+
single-cut models are never assigned Tier 3, that undefined cases return no tier rather
84+
than an incorrect one, and that `summary()` returns the documented `stability` fields.
85+
86+
* **Reachability Test:** Added an explicit test that Tier 1 is attainable for a
87+
well-separated threshold in a large sample. The defect above was characterised by an
88+
entire classification category being unreachable, which value-comparison tests alone
89+
would not have detected.
90+
91+
* **Workflow Integration Tests:** Added `tests/testthat/test-workflow-integration.R`,
92+
which runs the three core functions in sequence and asserts that state is preserved
93+
across them: the number of thresholds located matches the number selected, one
94+
confidence interval and one relative width is reported per threshold, the predictor is
95+
identical at every stage, and covariates used during the search are reused during
96+
validation. Both defects fixed in this release were handoff failures between functions
97+
rather than faults within them, which unit tests cannot detect.
98+
99+
## Dependencies
100+
101+
* No changes to package dependencies in this release.
102+
103+
1104
# OptSurvCutR v0.10.1 (2026-04-09)
2105

3106
## Documentation & Standards Compliance

‎R/engine-systematic.R‎

Lines changed: 6 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@
66
#' Internal helper: Systematic Grid Search over Regulared Space
77
#'
88
#' @description
9-
#' Implements an exhaustive grid search over a regulared percentile space
9+
#' Implements an exhaustive grid search over a regularised percentile space
1010
#' to evaluate possible thresholds for 1 or 2 cut-points, respecting the minimum
1111
#' group size constraints.
1212
#'
@@ -30,7 +30,7 @@
3030
.systematic_search <- function(userdata, num_cuts, criterion,
3131
covariates, nmin, predictor_name,
3232
quiet, candidate_cuts = NULL, ...) {
33-
if (!quiet) cli::cli_alert_info("Running regulared systematic search for {num_cuts} cut-point(s)...")
33+
if (!quiet) cli::cli_alert_info("Running regularised systematic search for {num_cuts} cut-point(s)...")
3434
userdata <- userdata[order(userdata$factor), ]
3535

3636
cov_part <- if (!is.null(covariates)) paste(" +", paste(covariates, collapse = " + ")) else ""
@@ -66,7 +66,7 @@
6666
best_cut_val <- rep(NA_real_, num_cuts)
6767
all_stats_df <- NULL
6868

69-
# Establish the core regulared grid structure
69+
# Establish the core regularised grid structure
7070
if (!is.null(candidate_cuts)) {
7171
search_grid <- candidate_cuts
7272
} else {
@@ -95,7 +95,7 @@
9595
best_stat <- stats_per_cut[best_idx]
9696
all_stats_df <- data.frame(cut1 = search_grid, stat = stats_per_cut)
9797
} else { # num_cuts == 2
98-
if (!quiet) cli::cli_alert_info("Searching for 2 cuts over regulared coordinate space...")
98+
if (!quiet) cli::cli_alert_info("Searching over {length(search_grid)} candidate positions for {num_cuts} cut(s)...")
9999
foreach::registerDoSEQ()
100100

101101
if (length(search_grid) < 2) {
@@ -137,7 +137,7 @@
137137
all_stats_df <- results_list
138138
}
139139

140-
if (!quiet) cli::cli_alert_success("Systematic grid optimation complete.")
140+
if (!quiet) cli::cli_alert_success("Systematic grid optimisation complete.")
141141
return(list(
142142
optimal_cuts = best_cut_val,
143143
optimal_stat = best_stat,
@@ -225,7 +225,7 @@
225225
#' Internal helper: Systematic Model Selection and Delta IC Range Mining
226226
#'
227227
#' @description
228-
#' Evaluates models from 1 to `max_cuts` using the regulared systematic
228+
#' Evaluates models from 1 to `max_cuts` using the regularised systematic
229229
#' sweeper, computing Information Criteria (AIC/BIC) and profiling the permissible range.
230230
#'
231231
#' @param userdata Cleaned survival data frame.

‎R/find_cutpoint.R‎

Lines changed: 82 additions & 30 deletions
Original file line numberDiff line numberDiff line change
@@ -13,19 +13,52 @@
1313
#'
1414
#' @section srrstats compliance:
1515
#' .
16-
#' @srrstats {G1.1} Grounded in information-theoretic model selection (AIC/BIC)
17-
#' and multivariable survival optimization (Cox/log-rank) via exhaustive grid
18-
#' search (k <= 2) and genetic algorithms via `rgenoud` (k > 2). Whereas tools
19-
#' like `survminer` and `cutpointr` focus strictly on univariate single splits
20-
#' (k = 1), OptSurvCutR optimizes multiple thresholds simultaneously under active
21-
#' covariate adjustment with bootstrap stability validation. Accelerated using
22-
#' base `Rcpp` (`src/matrix_factory.cpp`) for discrete index binning without external
23-
#' linear algebra libraries.
16+
#' @srrstats {G1.1} This package implements an established approach — outcome-oriented
17+
#' cut-point selection for censored survival data — and extends it in four key
18+
#' directions. The methodological origin is the maximally selected rank statistic
19+
#' (Miller & Siegmund 1982; Lausen & Schumacher 1992), in which a log-rank statistic
20+
#' is maximised over candidate thresholds and the resulting p-value corrected for the
21+
#' selection. Faraggi & Simon (1996) established by simulation that uncorrected
22+
#' selection inflates Type I error and biases effect estimates, and Rota et al. (2015)
23+
#' compared correction strategies for censored outcomes. These references define the
24+
#' single-threshold problem that `maxstat`, `survminer::surv_cutpoint()`, and
25+
#' `cutpointr` address.
26+
#'
27+
#' Those packages are not deficient implementations of a multi-threshold method; the
28+
#' multi-threshold problem is outside their stated scope. `cutpointr` is designed for
29+
#' binary classification metrics and optimises a single cut-point over measures such as
30+
#' the Youden index; `maxstat` computes asymptotic approximations and exact null
31+
#' distributions of maximally selected rank statistics for a single split; `survminer`
32+
#' provides a plotting-oriented interface to the latter. Where a single unadjusted
33+
#' threshold answers the question, these remain appropriate tools.
34+
#'
35+
#' `OptSurvCutR` addresses four methodological problems that fall outside that scope:
36+
#' 1. Number of thresholds: Treated as a model selection problem rather than fixed a
37+
#' priori, using information criteria (Akaike 1974; Schwarz 1978) across candidate
38+
#' complexities.
39+
#' 2. Simultaneous optimisation: Multiple thresholds are optimised simultaneously rather
40+
#' than sequentially, using exhaustive enumeration for k <= 2 and an evolutionary
41+
#' genetic algorithm (Mebane & Sekhon 2011, via `rgenoud`) for k > 2, avoiding the
42+
#' path-dependence of hierarchical splitting.
43+
#' 3. Confounder control: Thresholds are selected within a multivariable Cox model
44+
#' (Cox 1972), evaluating candidate boundaries conditionally on clinical covariates
45+
#' rather than marginally.
46+
#' 4. Threshold reproducibility: Non-parametric bootstrap resampling (Efron 1979)
47+
#' re-runs the search across resampled cohorts to quantify spatial stability. While
48+
#' a permutation-corrected p-value establishes that an observed separation is unlikely
49+
#' under the null, the bootstrap CI reports whether the threshold coordinate itself
50+
#' is reproducible across comparable patient samples.
51+
#'
52+
#' Numerical implementation uses base `Rcpp` (`src/matrix_factory.cpp`) for discrete
53+
#' index binning, without external linear algebra dependencies.
2454
#' @srrstats {G1.0} References provided for Cox, log-rank, genetic optimisation.
2555
#' @srrstats {G1.3} Systematic grid search (1–2 cuts) and `rgenoud` global
2656
#' optimisation documented.
27-
#' @srrstats {G1.5} Compared with `cutpointr` and `survminer` in package
28-
#' vignette.
57+
#' @srrstats {G1.5} A feature-level comparison against `maxstat`, `survminer`,
58+
#' `CutpointsOEHR`, Evaluate Cutpoints, X-tile and Cutoff Finder is provided in the
59+
#' accompanying manuscript, and against `maxstat`/`survminer` in the bilirubin
60+
#' vignette. Numerical comparison of results across implementations is not attempted,
61+
#' as the packages address different estimands.
2962
#' @srrstats {G1.6} Numerical stability via `survival::coxph` and `rgenoud`;
3063
#' edge cases return `NA`.
3164
#' @srrstats {G2.3a} Uses `match.arg()` to validate `method` and `criterion`
@@ -63,30 +96,49 @@
6396
#' @srrstats {RE2.4a} Checks for collinearity among predictors.
6497
#' @srrstats {RE2.4b} Checks for collinearity between X and Y.
6598
#'
66-
#' @details
67-
#' `method = "systematic"`: grid search respecting `nmin`. Optimised via internal quantiles.
68-
#' `method = "genetic"`: `rgenoud` global optimisation.
69-
#' Systematic search is slow for `num_cuts > 2`; use `genetic`.
70-
#' Core vector partitions are calculated in compiled C++ via `Rcpp` for optimal performance.
71-
#'
7299
#' @references
73-
#' Altman, D. G., Lausen, B., Sauerbrei, W., & Schumacher,
74-
#' M. (1994). Dangers of Using “Optimal” Cutpoints in the Evaluation of
75-
#' Prognostic Factors. *JNCI: Journal of the National Cancer Institute*,
76-
#' 86(11), 829–835. \doi{10.1093/jnci/86.11.829}
100+
#' Akaike, H. (1974). A new look at the statistical model identification.
101+
#' *IEEE Transactions on Automatic Control*, 19(6), 716–723.
102+
#' \doi{10.1109/TAC.1974.1100705}
103+
#'
104+
#' Altman, D. G., Lausen, B., Sauerbrei, W., & Schumacher, M. (1994). Dangers
105+
#' of using "optimal" cutpoints in the evaluation of prognostic factors.
106+
#' *JNCI: Journal of the National Cancer Institute*, 86(11), 829–835.
107+
#' \doi{10.1093/jnci/86.11.829}
108+
#'
109+
#' Cox, D. R. (1972). Regression models and life-tables. *Journal of the
110+
#' Royal Statistical Society: Series B (Methodological)*, 34(2), 187–202.
111+
#' \doi{10.1111/j.2517-6161.1972.tb00899.x}
112+
#'
113+
#' Efron, B. (1979). Bootstrap methods: Another look at the jackknife.
114+
#' *The Annals of Statistics*, 7(1), 1–26. \doi{10.1214/aos/1176344552}
115+
#'
116+
#' Faraggi, D., & Simon, R. (1996). A simulation study of cross-validation for
117+
#' selecting an optimal cutpoint in univariate survival analysis.
118+
#' *Statistics in Medicine*, 15(20), 2203–2213.
119+
#' \doi{10.1002/(SICI)1097-0258(19961030)15:20<2203::AID-SIM357>3.0.CO;2-G}
120+
#'
121+
#' Lausen, B., & Schumacher, M. (1992). Maximally selected rank statistics.
122+
#' *Biometrics*, 48(1), 73–85. \doi{10.2307/2532740}
123+
#'
124+
#' Mantel, N. (1966). Evaluation of survival data and two new rank order
125+
#' statistics arising in its consideration. *Cancer Chemotherapy Reports*,
126+
#' 50(3), 163–170.
127+
#'
128+
#' Mebane Jr, W. R., & Sekhon, J. S. (2011). Genetic optimization using
129+
#' derivatives: The rgenoud package for R. *Journal of Statistical Software*,
130+
#' 42(11), 1–26. \doi{10.18637/jss.v042.i11}
77131
#'
78-
#' Cox, D. R. (1972). Regression Models and Life-Tables. *Journal
79-
#' of the Royal Statistical Society: Series B (Methodological)*, 34(2),
80-
#' 187–202. \doi{10.1111/j.2517-6161.1972.tb00899.x}
132+
#' Miller, R., & Siegmund, D. (1982). Maximally selected chi square statistics.
133+
#' *Biometrics*, 38(4), 1011–1016. \doi{10.2307/2529881}
81134
#'
82-
#' Mantel, N. (1966). Evaluation of survival data and two new
83-
#' rank order statistics arising in its consideration. *Cancer
84-
#' Chemotherapy Reports*, 50(3).
135+
#' Rota, M., Antolini, L., & Valsecchi, M. G. (2015). Optimal cut-point
136+
#' definition in biomarkers: The case of censored failure time outcome.
137+
#' *BMC Medical Research Methodology*, 15(1), 24.
138+
#' \doi{10.1186/s12874-015-0009-y}
85139
#'
86-
#' Mebane Jr, W. R., & Sekhon, J. S. (2011). Genetic
87-
#' Optimisation Using Derivatives: The rgenoud Package for R.
88-
#' *Journal of Statistical Software*, 42, 1–26.
89-
#' \doi{10.18637/jss.v042.i11}
140+
#' Schwarz, G. (1978). Estimating the dimension of a model. *The Annals of
141+
#' Statistics*, 6(2), 461–464. \doi{10.1214/aos/1176344136}
90142
#'
91143
#' @param data A data frame containing the analysis variables.
92144
#' @param predictor The continuous predictor variable name (character).

‎R/validate_cutpoint.R‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -351,6 +351,10 @@ validate_cutpoint <- function(cutpoint_result, num_replicates = 500,
351351
confidence_intervals = ci_df,
352352
bootstrap_distribution = bootstrap_df,
353353
boot_summary = boot_summary,
354+
# Carry the analysis data forward so that summary() can compute the
355+
# relative CI width against the spread of the PREDICTOR. Without this,
356+
# summary() falls back to an incorrect denominator.
357+
userdata = cutpoint_result$userdata,
354358
parameters = list(
355359
num_replicates = num_replicates,
356360
successful_reps = successful_reps,

0 commit comments

Comments
 (0)