01. SuperSurv with Ensemble

Introduction

The core feature of the SuperSurv package is its ability to combine multiple base survival learners using cross-validated ensemble weights. This tutorial walks through preparing data, defining a library of models, fitting the Super Learner, and generating predictions for new patients.

1. Load Data & Prepare Matrices

We will use the built-in metabric dataset, extracting the covariates and performing a standard 80/20 train-test split.

For a quick teaching example, we sample 200 patients and use three cross-validation folds and 20-point fitting grids. These settings illustrate both fitting objectives and the prediction workflow, rather than a full-data performance comparison.

library(SuperSurv)
library(survival)

# Load built-in METABRIC data
data("metabric", package = "SuperSurv")

# Use a fixed teaching subset so the package vignette remains quick to rebuild.
# The full-data comparison used in the manuscript is provided separately in the
# replication materials.
set.seed(42)
example_idx <- sample(seq_len(nrow(metabric)), 200)
metabric_example <- metabric[example_idx, ]
n_total <- nrow(metabric_example)
train_idx <- sample(1:n_total, 0.8 * n_total)

train <- metabric_example[train_idx, ]
test  <- metabric_example[-train_idx, ]

# Extract just the X covariates (assuming they are named x0, x1, etc.)
x_cols <- grep("^x", names(metabric), value = TRUE)
X_tr <- train[, x_cols]
X_te <- test[, x_cols]

# Define the prediction time grid (e.g., survival at 50, 100, 150, 200 months)
new.times <- c(50, 100, 150, 200)

common_control <- list(
  saveFitLibrary = TRUE,
  event.t.grid = seq(0, max(train$duration[train$event == 1]), length.out = 20),
  cens.t.grid = seq(0, max(train$duration[train$event == 0]), length.out = 20)
)

2. Define the Ensemble Library

We define a small library of parametric and tree-based survival models for a reproducible demonstration.

my_library <- c("surv.coxph", "surv.weibull")
if (has_rpart) {
  my_library <- c(my_library, "surv.rpart")
}

3. Train the SuperSurv Metalearner

Before we run the main SuperSurv engine, it is important to understand its key parameters. The Super Learner algorithm relies on cross-validation to assign weights to the base models, and it must model both the event and the censoring mechanism to avoid biased evaluations.

Here is the complete guide to the arguments you need to pass:

The Data Inputs

The Model Libraries

The Meta-Learner & Tuning

Let’s fit two models with verbose = FALSE to see how the meta-learner choice affects the final ensemble weights.

# Fit 1: Least Squares Meta-learner
set.seed(2026)
fit_ls <- SuperSurv(
  time = train$duration,
  event = train$event,
  X = X_tr,
  newdata = X_te,                 # Predict on the test set immediately
  new.times = new.times,       # Our evaluation time grid
  event.library = my_library,
  cens.library = my_library,
  metalearner = "brier", 
  control = common_control,
  verbose = FALSE,
  selection = "ensemble",
  nFolds = 3
)

# Fit 2: Negative Log-Likelihood Meta-learner
set.seed(2026) # Reuse the same cross-validation folds for a fair comparison.
fit_nll <- SuperSurv(
  time = train$duration,
  event = train$event,
  X = X_tr,
  newdata = X_te,
  new.times = new.times,
  event.library = my_library,
  cens.library = my_library,
  metalearner = "logloss",       # Swap to nloglik
  control = common_control,
  verbose = FALSE,
  selection = "ensemble",
  nFolds = 3
)

4. Package Object Interface

SuperSurv fits behave like ordinary R model objects. Use the print and summary methods for a quick overview, and use accessors for common fitted-model details.

fit_ls
#> SuperSurv fit
#>   Selection: ensemble 
#>   Event learners: 3 
#>   Censoring learners: 3 
#>   Predictions: 40 observations x 4 times
#>   Evaluation times: 4 values from 50 to 200 
#>   Nonzero event weights:
#> surv.coxph_screen.all surv.rpart_screen.all 
#>                0.6132                0.3868

summary(fit_ls)
#> Summary of SuperSurv fit
#>   Selection: ensemble 
#> 
#> Call:
#> SuperSurv(time = train$duration, event = train$event, X = X_tr, 
#>     newdata = X_te, new.times = new.times, event.library = my_library, 
#>     cens.library = my_library, verbose = FALSE, control = common_control, 
#>     metalearner = "brier", selection = "ensemble", nFolds = 3)
#> 
#> Event ensemble:
#>                  learner weight   risk status
#>    surv.coxph_screen.all 0.6132 3.6578     ok
#>    surv.rpart_screen.all 0.3868 3.6663     ok
#>  surv.weibull_screen.all 0.0000 3.6678     ok
#> 
#> Censoring ensemble:
#>                  learner weight   risk status
#>    surv.coxph_screen.all 0.4881 0.5579     ok
#>  surv.weibull_screen.all 0.2618 0.5673     ok
#>    surv.rpart_screen.all 0.2501 0.5707     ok
#> 
#> Predictions: 40 observations x 4 times
#> Evaluation times: 4 values from 50 to 200 
#> Elapsed time (seconds):
#> everything      train    predict 
#>      0.395      0.370      0.025

event_weights(fit_ls)
#>   surv.coxph_screen.all surv.weibull_screen.all   surv.rpart_screen.all 
#>               0.6132128               0.0000000               0.3867872

learner_names(fit_ls)
#> [1] "surv.coxph_screen.all"   "surv.weibull_screen.all"
#> [3] "surv.rpart_screen.all"

eval_times(fit_ls)
#> [1]  50 100 150 200

selected_variables(fit_ls, learner = 1)
#> [1] "x0" "x1" "x2" "x3" "x4" "x5" "x6" "x7" "x8"

5. Compare the Fitting Objectives

With selection = "ensemble", the meta-learner estimates a convex combination of the candidate predictions under the selected cross-validated objective. The two objectives can produce different weights.

First inspect the fitted weights and algorithm summaries.

cat("\n--- LEAST SQUARES METALEARNER ---\n")
#> 
#> --- LEAST SQUARES METALEARNER ---
summary(fit_ls)
#> Summary of SuperSurv fit
#>   Selection: ensemble 
#> 
#> Call:
#> SuperSurv(time = train$duration, event = train$event, X = X_tr, 
#>     newdata = X_te, new.times = new.times, event.library = my_library, 
#>     cens.library = my_library, verbose = FALSE, control = common_control, 
#>     metalearner = "brier", selection = "ensemble", nFolds = 3)
#> 
#> Event ensemble:
#>                  learner weight   risk status
#>    surv.coxph_screen.all 0.6132 3.6578     ok
#>    surv.rpart_screen.all 0.3868 3.6663     ok
#>  surv.weibull_screen.all 0.0000 3.6678     ok
#> 
#> Censoring ensemble:
#>                  learner weight   risk status
#>    surv.coxph_screen.all 0.4881 0.5579     ok
#>  surv.weibull_screen.all 0.2618 0.5673     ok
#>    surv.rpart_screen.all 0.2501 0.5707     ok
#> 
#> Predictions: 40 observations x 4 times
#> Evaluation times: 4 values from 50 to 200 
#> Elapsed time (seconds):
#> everything      train    predict 
#>      0.395      0.370      0.025

cat("\n--- NLOGLIK METALEARNER ---\n")
#> 
#> --- NLOGLIK METALEARNER ---
summary(fit_nll)
#> Summary of SuperSurv fit
#>   Selection: ensemble 
#> 
#> Call:
#> SuperSurv(time = train$duration, event = train$event, X = X_tr, 
#>     newdata = X_te, new.times = new.times, event.library = my_library, 
#>     cens.library = my_library, verbose = FALSE, control = common_control, 
#>     metalearner = "logloss", selection = "ensemble", nFolds = 3)
#> 
#> Event ensemble:
#>                  learner weight   risk status
#>    surv.coxph_screen.all      1 0.7015     ok
#>  surv.weibull_screen.all      0 0.7286     ok
#>    surv.rpart_screen.all      0 0.7899     ok
#> 
#> Censoring ensemble:
#>                  learner weight   risk status
#>    surv.coxph_screen.all 0.5131 0.5741     ok
#>  surv.weibull_screen.all 0.3991 0.5782     ok
#>    surv.rpart_screen.all 0.0878 0.5954     ok
#> 
#> Predictions: 40 observations x 4 times
#> Evaluation times: 4 values from 50 to 200 
#> Elapsed time (seconds):
#> everything      train    predict 
#>      0.363      0.337      0.025

Then evaluate both fitted ensembles on the same held-out observations using both IPCW Brier score and IPCW log-loss. This comparison illustrates the objectives without assuming in advance that one must perform better.

benchmark_ls <- eval_benchmark(
  fit_ls, X_te, test$duration, test$event, new.times
)
benchmark_nll <- eval_benchmark(
  fit_nll, X_te, test$duration, test$event, new.times
)

objective_comparison <- rbind(
  transform(
    benchmark_ls$summary[benchmark_ls$summary$Model == "SuperSurv_Ensemble", ],
    Fitting_Objective = "Brier"
  ),
  transform(
    benchmark_nll$summary[benchmark_nll$summary$Model == "SuperSurv_Ensemble", ],
    Fitting_Objective = "Log-loss"
  )
)
objective_comparison[, c("Fitting_Objective", "IBS", "IPCW_LogLoss", "Uno_C", "iAUC")]
#>   Fitting_Objective       IBS IPCW_LogLoss     Uno_C      iAUC
#> 1             Brier 0.2233944    0.6339592 0.5917061 0.6062454
#> 2          Log-loss 0.1967823    0.5738174 0.6671108 0.7134136

event_weights(fit) reports each learner’s contribution to the final convex combination. A zero weight means that the learner does not contribute to that fitted ensemble. The held-out table, rather than the training objective alone, should guide interpretation of predictive performance.

6. Generating Predictions on New Data

If you passed newdata during the training phase, SuperSurv already calculated the predictions for your test set. However, in a real-world clinical setting, you will often train the model once and then predict on brand new patients months later.

Because we set control = list(saveFitLibrary = TRUE) during training, we can use the standard R predict() method.

# Select 3 brand new patients from our test set
new_patients <- X_te[1:6, ]

# Generate predictions using the Least Squares ensemble
ensemble_preds <- predict(
  object = fit_ls, 
  newdata = new_patients, 
  new.times = new.times,
  type = "event"
)

cat("\n--- PREDICTED SURVIVAL PROBABILITIES ---\n")
#> 
#> --- PREDICTED SURVIVAL PROBABILITIES ---
final_matrix <- ensemble_preds
colnames(final_matrix) <- paste0("Time_", new.times)
rownames(final_matrix) <- paste0("Patient_", 1:6)

print(round(final_matrix, 4))
#>           Time_50 Time_100 Time_150 Time_200
#> Patient_1  0.8321   0.6422   0.5388   0.4367
#> Patient_2  0.8450   0.6958   0.6267   0.5663
#> Patient_3  0.8976   0.7698   0.6934   0.6117
#> Patient_4  0.9023   0.7817   0.7096   0.6325
#> Patient_5  0.8841   0.7423   0.6592   0.5723
#> Patient_6  0.9017   0.7784   0.7043   0.6247

Understanding the Output Matrix:

With type = "event", predict() returns the event-survival probability matrix directly. * Rows (\(N\)): Represent individual patients. * Columns (\(T\)): Represent the specific time points we defined in new.times. * Values: The estimated probability that the patient will survive past that specific time point. As time increases (moving left to right across a row), the survival probability naturally decreases.

7. Visualizing Patient-Specific Predictions

While raw probability matrices (\(N \times T\)) are perfect for downstream coding and performance benchmarking, they are difficult to interpret clinically. Doctors and researchers need to see the actual survival trajectories.

SuperSurv includes a built-in plot_predict() function to effortlessly translate this matrix into publication-ready survival curves for individual patients.

This plotting example runs only when the optional ggplot2 package is installed; fitting and numerical prediction do not require it.

# Plot the predicted survival curves for our 3 new patients (Rows 1, 2, and 3)
plot_predict(
  preds = ensemble_preds,
  eval_times = new.times,
  patient_idx = 1:6
)

8. Next Steps

You now know how to prepare data, define a model library, choose a meta-learner, and generate patient-specific survival curves.

Before applying a model, evaluate the ensemble and its component learners on appropriately held-out data. Head to Tutorial 2: Model Performance & Benchmarking for numerical and graphical examples using time-dependent Brier score, AUC, and Uno’s C-index.