---
title: "From marginal study summaries to synthetic patients"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{From marginal study summaries to synthetic patients}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

`summary2joint` estimates a joint latent Gaussian distribution from marginal
summaries in repeated independent studies of a common population. It does not
identify arbitrary dependence from one set of margins and does not recover the
original patient records.

## Input

Use a named list specifying continuous, binary, and ordinal variables. For each
study, supply its size and a named list of means/sample SDs, binary event counts,
and ordered category counts. All summaries in a study refer to the same people.
Unreported variables can be omitted, but every pair needs repeated joint reporting.

```{r}
library(summary2joint)
set.seed(31)
variables <- list(age = list(type = "continuous"),
                  response = list(type = "binary"),
                  severity = list(type = "ordinal", levels = 3L))
studies <- lapply(seq_len(80), function(i) {
  z <- matrix(rnorm(300), 100, 3)
  z[, 2] <- 0.4 * z[, 1] + sqrt(0.84) * z[, 2]
  age <- 55 + 8 * z[, 1]
  response <- as.integer(z[, 2] > 0)
  severity <- findInterval(z[, 3], c(-Inf, -0.5, 0.5, Inf))
  list(n = 100L, summaries = list(
    age = list(mean = mean(age), sd = sd(age)),
    response = list(events = sum(response)),
    severity = list(counts = tabulate(severity, 3))))
})
fit <- fit_summary_copula(studies, variables)
fit
fit$diagnostics[c("converged", "boundary", "inference_ok")]
```

## Probability and generation

```{r}
joint_probability(fit, lower = c(age = 60, response = 1, severity = 2))
head(simulate_summary_copula(fit, n = 100, seed = 42))
confint(fit)
```

The correlation matrix is on the latent normal scale, not generally the Pearson
correlation of the observed discrete variables. Probability intervals propagate
uncertainty in both the margins and dependence. Fits are ordinary serializable
R objects; save them using `saveRDS()`.

## Independent groups and limitations

If scientific knowledge supports independent groups, supply a partition to
`fit_summary_copula_groups()`. Each group must contain at least two variables.
Use `joint_probability_groups()` for the fitted grouped distribution;
`simulate_summary_copula()` works for both classes. Grouping is a modeling
assumption, not an automatic selection procedure.

Always inspect convergence, boundaries, and inference diagnostics. A full-rank
sandwich requires more studies than parameters, but that condition alone does not
ensure accurate intervals. Rare events and few studies can cause undercoverage.
Between-study heterogeneity can confound within-patient dependence. This package
assumes a common population, aligned definitions, independent participants across
studies, and reporting independent of measurements. It does not model missing
patient data or rounded counts. Details and numerical controls are in
`?fit_summary_copula`.
