Real datasets arrive with column names that are missing, misleading,
in the wrong language, or simply wrong: a column called
category_code holding continuous lab values, a
gender column that is actually a free numeric measurement,
an outcome buried under an opaque v7. Any tool that decides
a column’s statistical role from its name
inherits every one of those lies.
rolescry decides roles from the data
signature instead – the guiding principle is Data inspice,
non nomen (“inspect the data, not the name”). Renaming every column
to col_1, col_2, ... does not change a single role
assignment. This is the turnusol (litmus) invariant,
and it is the package’s keystone test.
set.seed(1)
pre <- rnorm(100, 10, 2) # a measurement, before ...
d <- data.frame(
arm = rep(c("control", "treatment"), each = 50), # a balanced 2-level grouping
pre = pre,
post = pre + rnorm(100, 1, 1) # ... and again after (paired with pre)
)
res <- detect_roles(d)
res
#> <role_detection> 100 observations x 3 variables
#> group_var arm pct=100.0
#> paired_pairs pre, post pct=77.3
#> agreement_pairs pre, post pct=79.1
#> covariate pre, post, arm pct=50.0
summary(res)
#> role found columns pct
#> 1 group_var TRUE arm 100.0
#> 2 paired_pairs TRUE pre,post 77.3
#> 3 agreement_pairs TRUE pre,post 79.1
#> 4 time_variable FALSE 0.0
#> 5 event_variable FALSE 0.0
#> 6 subject_id FALSE 0.0
#> 7 repeated_measures FALSE 0.0
#> 8 scale_items FALSE 0.0
#> 9 outcome_continuous FALSE 0.0
#> 10 outcome_binary FALSE 0.0
#> 11 covariate TRUE pre,post,arm 50.0The same call on the name-stripped twin yields the same roles by position:
Each column is first typed by value
(continuous, binary, categorical,
ID), never by name. Candidate roles are then scored by
signatures that capture the statistical shape a role implies –
correlation and distributional overlap for paired measurements,
Bland-Altman bias and intraclass correlation for agreement, event-rate
and right-skew for survival, inter-item correlation and a Cronbach-alpha
proxy for scale items, and so on. Every score is a transparent sum of
named components you can inspect:
res$roles$paired_pairs$components[[1]]
#> $name
#> [1] "Correlation"
#>
#> $score
#> [1] 20
#>
#> $max
#> [1] 20
#>
#> $detail
#> [1] "r=0.88"For a categorical column with level proportions \(p_1, \dots, p_k\), the normalized Shannon entropy
\[ H_{\text{norm}} = \frac{-\sum_i p_i \log_2 p_i}{\log_2 k} \in [0, 1] \]
measures how balanced the levels are. A grouping variable (treatment vs control) has high entropy (near-balanced); a near-constant flag has entropy near zero. Entropy drives both the value classifier and the group-balance signal.
To ask – name-blind – whether a candidate grouping actually
carries information about an outcome, rolescry
uses normalized mutual information:
\[ \text{NMI}(X, Y) = \frac{I(X; Y)}{\min\{H(X),\, H(Y)\}} \in [0, 1], \]
which is 0 for independent variables and 1 for a deterministic association, and is comparable across variables with different numbers of levels. It is exposed directly:
Names are not useless – they are just untrustworthy. When
you do trust them, pass a keyword dictionary via
name_bonus. Names then act only as a small,
capped tie-breaker (at most a +10 point nudge,
i.e. <= 10% of the selection score); the mathematical signature still
dominates (>= 90%), the relationship enforced by
score_gap_ok().
Here two columns both look like groupings; the data alone picks the
perfectly balanced site, but a keyword dictionary nudges
the choice to the intended treatment arm:
set.seed(4)
clin <- data.frame(
site = rep(c("north", "south"), length.out = 160), # perfectly balanced (the math pick)
treated = sample(c("no", "yes"), 160, replace = TRUE, prob = c(0.62, 0.38))
)
detect_roles(clin)$roles$group_var$columns # data alone -> "site"
#> [1] "site"
detect_roles(clin, name_bonus = rolescry_default_name_bonus())$roles$group_var$columns # name nudge -> "treated"
#> [1] "treated"read_data() reads a file with the header row found by
the same information-theoretic scorer (detect_header()), so
messy exports with title rows or merged cells still load with sensible
column names. Delimited text works with base R; spreadsheet and
statistical formats use optional packages and degrade gracefully if they
are not installed.
rolescry is derived from Boynukara, C. (2026).
MDStatR (v2.1.0 Veritas). Zenodo. https://doi.org/10.5281/zenodo.20707791. Run
citation("rolescry") to cite the package and its parent
engine.