Why de-identify?

Sharing a linelist externally — with collaborators, for CRAN examples, in a research supplement — requires that directly identifying information is removed. molting() automates this using a cryptographic hash: each row gets a unique fingerprint derived from its identifiers, the identifiers are stripped, and the fingerprint is stored in a lookup table that only authorised personnel hold.

The name comes from the natural process of moulting: a bird sheds its distinctive, identifiable plumage and temporarily becomes more uniform. The old plumage is not destroyed — it re-grows from the same follicles. The lookup table is those follicles.


What gets removed

molting() uses regular-expression pattern matching to detect PII columns. The default patterns cover:

  • Names: name, surname, firstname, lastname
  • Dates: dob, birth
  • Identifiers: mrn, urn, medicare, patient_id, subject_id, _id$
  • Contact: address, street, phone, email

Age category variables (age2cat, age5cat, age10cat, etc.) are automatically preserved — they are not directly identifying.

patient_data <- data.frame(
  patient_name = c("John Doe", "Jane Smith"),
  dob          = as.Date(c("1980-01-01", "1975-05-15")),
  mrn          = c("12345", "67890"),
  age5cat      = factor(c("18-64", "18-64")),   # preserved automatically
  diagnosis    = c("Condition A", "Condition B"),
  lab_value    = c(120, 95)
)

result <- suppressMessages(molting(patient_data))
names(result$deidentified)   # hash + retained columns
#> [1] "row_hash"  "age5cat"   "diagnosis" "lab_value"
names(result$lookup)         # hash + removed columns
#> [1] "row_hash"     "patient_name" "dob"          "mrn"

The output list

molting() returns a named list with two elements when return_lookup = TRUE (the default):

  • $deidentified — the de-identified data frame, with row_hash as the first column
  • $lookup — the lookup table, with row_hash plus all removed identifier columns
str(result, max.level = 1)
#> List of 2
#>  $ deidentified: tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
#>  $ lookup      : tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
head(result$lookup)
#> # A tibble: 2 × 4
#>   row_hash                                              patient_name dob   mrn  
#>   <chr>                                                 <chr>        <chr> <chr>
#> 1 89573bbf928ef324ba95e8d04fd1701dfc5c15265ab21b3dbbb2… John Doe     1980… 12345
#> 2 7b2536bf2d008eb3555f9404d1762dcc529e1a7d5e6880c66341… Jane Smith   1975… 67890

Store $lookup securely and separately from $deidentified. Consider encrypting the lookup file before archiving. In a Queensland Health context, the lookup table should remain within the Health Service network.


Hash algorithm selection

The default is SHA-256, which provides strong collision resistance for typical surveillance dataset sizes (tens of thousands of rows). For very large datasets where speed matters more than collision resistance, SHA-1 or MD5 are faster but should not be used where linkage integrity is critical.

# SHA-256 (default, recommended)
result_256 <- molting(patient_data, hash_method = "sha256")

# MD5 (shorter hash, faster, lower collision resistance)
result_md5 <- molting(patient_data, hash_method = "md5")

# Blake3 (fast and cryptographically strong — good for large datasets)
result_b3 <- molting(patient_data, hash_method = "blake3")

Controlling which columns are hashed

By default, molting() hashes all detected PII columns. Supply id_cols to override this — useful when you want a shorter, more stable hash based only on a true unique identifier.

# Hash only on MRN and DOB — more stable if name variations exist
result_ids <- suppressMessages(
  molting(patient_data, id_cols = c("mrn", "dob"))
)
result_ids$lookup
#> # A tibble: 2 × 4
#>   row_hash                                              patient_name dob   mrn  
#>   <chr>                                                 <chr>        <chr> <chr>
#> 1 255ad8c4452e80121bdee96a35de7cea0c25475a5e9e6f73088c… John Doe     1980… 12345
#> 2 107510d61f22148ec63b636654c41efea411429ace78b9c371f0… Jane Smith   1975… 67890

Adding columns to the removal list

Use additional_pii_cols for dataset-specific identifiers that don’t match the default patterns.

patient_data2 <- patient_data
patient_data2$study_code <- c("SC-001","SC-002")

result2 <- suppressMessages(
  molting(patient_data2, additional_pii_cols = "study_code")
)
names(result2$deidentified)
#> [1] "row_hash"  "age5cat"   "diagnosis" "lab_value"

Irreversible de-identification

If you genuinely do not need to relink (e.g. producing a public-use file), set return_lookup = FALSE. This is irreversible — there is no way to recover the original identifiers.

deidentified_only <- suppressMessages(
  molting(patient_data, return_lookup = FALSE)
)
class(deidentified_only)   # a data frame, not a list
#> [1] "tbl_df"     "tbl"        "data.frame"

Hash collisions

If two rows produce the same hash (extremely rare with SHA-256 for realistic dataset sizes but possible with very short hashes like CRC32), molting() warns you. If a collision is detected, switch to a stronger algorithm or add more columns to id_cols.


What comes next

Use [homing()] to relink the de-identified data when authorised (see vignette("homing")).

For aggregated outputs — monthly counts from roost() — de-identification may not be necessary at all if the counts are not small enough to be re-identifying. The ABS cell-suppression threshold of 5 is a useful rule of thumb.