Package {summary2joint}


Type: Package
Title: Joint Distribution Estimation from Marginal Study Summaries
Version: 0.1.0
Description: Estimates latent Gaussian joint distributions for normal continuous, binary, and ordinal variables using marginal summaries from independent studies of a common population. Fits a pairwise Gaussian working criterion using exact summary moments, with study-level sandwich uncertainty. Supports prespecified independent groups, joint event probabilities, and synthetic patient generation. Identification requires repeated joint reporting of variable pairs; heterogeneous populations, rare categories, and small study collections require caution.
License: GPL-3
Encoding: UTF-8
Depends: R (≥ 4.1.0)
Imports: Rcpp, mvtnorm, numDeriv, stats, utils
LinkingTo: Rcpp, RcppArmadillo
Suggests: knitr, rmarkdown
VignetteBuilder: knitr
NeedsCompilation: yes
Packaged: 2026-09-20 03:13:52 UTC; xzhang
Author: Xuekui Zhang ORCID iD [aut, cre]
Maintainer: Xuekui Zhang <ubcxzhang@gmail.com>
Repository: CRAN
Date/Publication: 2026-09-29 15:30:02 UTC

Joint Distribution Estimation from Marginal Study Summaries

Description

Reconstruct latent Gaussian distributions for normal continuous, binary, and ordinal variables from replicated marginal study summaries. See fit_summary_copula for the input contract and model assumptions. The package provides joint event probabilities and synthetic patient generation, not recovery of original patient records.

Author(s)

Xuekui Zhang ubcxzhang@gmail.com

See Also

fit_summary_copula_groups, joint_probability, simulate_summary_copula


Estimate a Joint Distribution from Marginal Study Summaries

Description

Fits a pairwise Gaussian working criterion with exact summary means and covariances. A Cholesky parameterization ensures a positive-definite latent correlation matrix. Continuous margins are normal; binary and ordinal variables are thresholded latent normal variables.

Usage

fit_summary_copula(studies, variables, control = list())

Arguments

studies

A list of independent studies. Each study has integer n >= 2 and a named summaries list. Continuous entries contain mean and sample sd (variance denominator n - 1); binary entries contain events or exact proportion; ordinal entries contain ordered counts or exact proportions. Counts must sum to n for ordinal variables. Omit entries for unreported variables; do not supply missing values. All summaries within a study must refer to the same participants.

variables

A uniquely named list of at least two variable specifications, each containing type = "continuous", "binary", or "ordinal". Ordinal specifications also contain integer levels >= 2.

control

Named list overriding the controls described below.

Details

Studies are independent samples from a common joint population. Reporting must be independent of measured values. Every variable pair must be jointly reported in at least one study to pass validation; informative estimation and asymptotic inference require repeated pair reporting across many studies. Two studies are a software minimum, not a recommendation for reliable estimation. Rounded proportions not implying integer counts are rejected.

Continuous summaries are centered and scaled before fitting. Optimization uses stats::optim with L-BFGS-B. Controls and defaults are:

starts = 2L

Number of deterministic optimization starts.

maxit = 500L, factr = 1e5, pgtol = 1e-6

Optimizer controls.

gradient_step = 1e-5, hessian_step = 2e-4

Relative central-difference steps.

min_probability = 1e-6, min_eigenvalue = 1e-4

Interior constraints.

inference = TRUE

Compute the study-level sandwich covariance.

keep_scores = FALSE

Retain study scores in the returned fit.

Parameter bounds are [-10,10] for standardized means, [-3,3] for log standard deviations, [-4,4] for binary and first ordinal thresholds, [-4,3] for log threshold increments, and [-5,5] for Cholesky coordinates.

Inference is withheld for failed or boundary fits, invalid sensitivity or variance estimates, or rank-deficient study scores. Full-rank inference requires more studies than parameters. This does not guarantee nominal coverage with few studies or rare categories. Check diagnostics before interpreting intervals.

Value

An object of class summary_copula, containing:

R, margins

Latent correlation matrix and original-scale marginal parameters. Latent correlations generally differ from observed Pearson correlations.

correlations

Pair estimates, sandwich standard errors, Fisher-scale 95 percent confidence limits, and counts of reporting studies.

coefficients, vcov, H, J

Internal parameter vector, its sandwich covariance, sensitivity, and study-score covariance. Coordinates are standardized means and log standard deviations, binary thresholds, ordinal first thresholds and log increments, then row-ordered Cholesky coordinates.

correlation_vcov, study_scores

Correlation covariance and optional study scores.

diagnostics

Convergence, boundary and inference flags, objective, gradient, eigenvalue, condition-number and rank diagnostics.

transform, control, variables, call

Scaling constants, numerical settings, variable definitions, and matched call.

n, K, pair_counts

Study sizes, study count, and joint reporting counts.

The fit contains no native pointer and can be serialized with saveRDS.

See Also

fit_summary_copula_groups, joint_probability

Examples

set.seed(17)
variables <- list(x = list(type = "continuous"), y = list(type = "binary"))
studies <- lapply(seq_len(60), function(i) {
  x <- rnorm(100)
  y <- as.integer(0.5 * x + sqrt(0.75) * rnorm(100) > 0)
  list(n = 100L, summaries = list(
    x = list(mean = mean(x), sd = sd(x)), y = list(events = sum(y))))
})
fit <- fit_summary_copula(studies, variables)
fit$R
fit$diagnostics
joint_probability(fit, lower = c(x = 0, y = 1))
head(simulate_summary_copula(fit, 10, seed = 42))

Fit Prespecified Independent Groups of Variables

Description

Fits each group using the original estimator and fixes between-group latent correlations to zero.

Usage

fit_summary_copula_groups(studies, variables, groups, control = list())

Arguments

studies, variables, control

As in fit_summary_copula.

groups

List of character vectors partitioning all variable names into at least two groups of at least two variables each. Membership is prespecified, not estimated.

Details

Studies reporting no variables from a group are omitted from that component fit and receive zero study scores for the group. The stacked sandwich covariance retains empirical cross-group score covariances. Group independence must be scientifically justified. It is not inferred from the supplied summaries. Full inference requires successful interior component fits and full rank of the stacked study scores.

Value

An object of class summary_copula_groups, with R, margins, variables, groups, component fits, study_ids, concatenated coefficients, parameter_indices, stacked vcov, study_scores, study_influences, within-group correlations, K, and diagnostics. Use joint_probability_groups for event probabilities and simulate_summary_copula for patient generation.


Joint Event Probabilities and Delta-Method Intervals

Description

Calculates the probability that all supplied bounds hold simultaneously, integrating over unspecified variables.

Usage

joint_probability(object, lower = numeric(), upper = numeric(),
                  inference = TRUE, level = .95)
joint_probability_groups(object, lower = numeric(), upper = numeric(),
                         inference = TRUE, level = .95)

Arguments

object

A summary_copula fit, or a summary_copula_groups fit for joint_probability_groups.

lower, upper

Named numeric vectors of original-scale bounds. Unspecified bounds are infinite. Binary values are 0/1 and ordinal values are 1 through the specified number of levels. Discrete endpoints are inclusive.

inference

Whether to calculate uncertainty when available in the fit.

level

Confidence level, conventionally 0.95.

Details

Deterministic multivariate normal integration supports at most six constrained variables per component. Larger events may be approximated from generated patients. Grouped probabilities factorize, but their uncertainty uses the complete stacked sandwich covariance. Intervals use the logit-scale delta method and include marginal-parameter uncertainty.

Value

A one-row data frame with estimate, se, lower, and upper. Uncertainty is NA if inference is disabled or unavailable, or the probability equals zero or one.

See Also

fit_summary_copula


Generate Synthetic Patients from a Fitted Joint Distribution

Description

Draws latent multivariate normal vectors and applies the fitted marginal observation maps. These are synthetic records, not reconstructed original patients.

Usage

simulate_summary_copula(object, n, seed = NULL)

Arguments

object

A fitted summary_copula or summary_copula_groups object, or a list with a valid latent correlation matrix R and named margins. Each margin contains type, with mean and sd for continuous variables, threshold for binary variables, or thresholds and levels for ordinal variables.

n

Positive integer number of synthetic patients.

seed

Optional integer random seed. When supplied, the caller's random state is restored after generation; otherwise the current random stream is used.

Value

A data frame with n rows and original variable names. Continuous variables retain measurement units, binary values are 0/1, and ordinal values are integers 1 through the number of levels.

See Also

fit_summary_copula


Inspect a Summary Copula Fit

Description

Displays diagnostic flags and latent correlations, extracts internal-scale parameters and covariance, or calculates correlation intervals.

Usage

## S3 method for class 'summary_copula'
print(x, ...)
## S3 method for class 'summary_copula'
coef(object, ...)
## S3 method for class 'summary_copula'
vcov(object, ...)
## S3 method for class 'summary_copula'
confint(object, parm, level = .95, ...)

Arguments

x, object

A fitted summary_copula object.

parm

Currently ignored; intervals for all latent correlations are returned.

level

Confidence level.

...

Further arguments, currently unused.

Value

print returns the fit invisibly; coef returns the internal parameter vector; vcov returns its covariance matrix; confint returns the correlation data frame with Fisher-transformed confidence limits. Intervals are unavailable when the fit's diagnostic checks fail.

See Also

fit_summary_copula