| Type: | Package |
| Title: | Joint Distribution Estimation from Marginal Study Summaries |
| Version: | 0.1.0 |
| Description: | Estimates latent Gaussian joint distributions for normal continuous, binary, and ordinal variables using marginal summaries from independent studies of a common population. Fits a pairwise Gaussian working criterion using exact summary moments, with study-level sandwich uncertainty. Supports prespecified independent groups, joint event probabilities, and synthetic patient generation. Identification requires repeated joint reporting of variable pairs; heterogeneous populations, rare categories, and small study collections require caution. |
| License: | GPL-3 |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.1.0) |
| Imports: | Rcpp, mvtnorm, numDeriv, stats, utils |
| LinkingTo: | Rcpp, RcppArmadillo |
| Suggests: | knitr, rmarkdown |
| VignetteBuilder: | knitr |
| NeedsCompilation: | yes |
| Packaged: | 2026-09-20 03:13:52 UTC; xzhang |
| Author: | Xuekui Zhang |
| Maintainer: | Xuekui Zhang <ubcxzhang@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-29 15:30:02 UTC |
Joint Distribution Estimation from Marginal Study Summaries
Description
Reconstruct latent Gaussian distributions for normal continuous, binary,
and ordinal variables from replicated marginal study summaries. See
fit_summary_copula for the input contract and model assumptions.
The package provides joint event probabilities and synthetic patient generation,
not recovery of original patient records.
Author(s)
Xuekui Zhang ubcxzhang@gmail.com
See Also
fit_summary_copula_groups, joint_probability,
simulate_summary_copula
Estimate a Joint Distribution from Marginal Study Summaries
Description
Fits a pairwise Gaussian working criterion with exact summary means and covariances. A Cholesky parameterization ensures a positive-definite latent correlation matrix. Continuous margins are normal; binary and ordinal variables are thresholded latent normal variables.
Usage
fit_summary_copula(studies, variables, control = list())
Arguments
studies |
A list of independent studies. Each study has integer |
variables |
A uniquely named list of at least two variable specifications,
each containing |
control |
Named list overriding the controls described below. |
Details
Studies are independent samples from a common joint population. Reporting must be independent of measured values. Every variable pair must be jointly reported in at least one study to pass validation; informative estimation and asymptotic inference require repeated pair reporting across many studies. Two studies are a software minimum, not a recommendation for reliable estimation. Rounded proportions not implying integer counts are rejected.
Continuous summaries are centered and scaled before fitting. Optimization uses
stats::optim with L-BFGS-B. Controls and defaults are:
starts = 2LNumber of deterministic optimization starts.
maxit = 500L,factr = 1e5,pgtol = 1e-6Optimizer controls.
gradient_step = 1e-5,hessian_step = 2e-4Relative central-difference steps.
min_probability = 1e-6,min_eigenvalue = 1e-4Interior constraints.
inference = TRUECompute the study-level sandwich covariance.
keep_scores = FALSERetain study scores in the returned fit.
Parameter bounds are [-10,10] for standardized means, [-3,3] for log standard deviations, [-4,4] for binary and first ordinal thresholds, [-4,3] for log threshold increments, and [-5,5] for Cholesky coordinates.
Inference is withheld for failed or boundary fits, invalid sensitivity or
variance estimates, or rank-deficient study scores. Full-rank inference requires
more studies than parameters. This does not guarantee nominal coverage with few
studies or rare categories. Check diagnostics before interpreting intervals.
Value
An object of class summary_copula, containing:
R,marginsLatent correlation matrix and original-scale marginal parameters. Latent correlations generally differ from observed Pearson correlations.
correlationsPair estimates, sandwich standard errors, Fisher-scale 95 percent confidence limits, and counts of reporting studies.
coefficients,vcov,H,JInternal parameter vector, its sandwich covariance, sensitivity, and study-score covariance. Coordinates are standardized means and log standard deviations, binary thresholds, ordinal first thresholds and log increments, then row-ordered Cholesky coordinates.
correlation_vcov,study_scoresCorrelation covariance and optional study scores.
diagnosticsConvergence, boundary and inference flags, objective, gradient, eigenvalue, condition-number and rank diagnostics.
transform,control,variables,callScaling constants, numerical settings, variable definitions, and matched call.
n,K,pair_countsStudy sizes, study count, and joint reporting counts.
The fit contains no native pointer and can be serialized with saveRDS.
See Also
fit_summary_copula_groups, joint_probability
Examples
set.seed(17)
variables <- list(x = list(type = "continuous"), y = list(type = "binary"))
studies <- lapply(seq_len(60), function(i) {
x <- rnorm(100)
y <- as.integer(0.5 * x + sqrt(0.75) * rnorm(100) > 0)
list(n = 100L, summaries = list(
x = list(mean = mean(x), sd = sd(x)), y = list(events = sum(y))))
})
fit <- fit_summary_copula(studies, variables)
fit$R
fit$diagnostics
joint_probability(fit, lower = c(x = 0, y = 1))
head(simulate_summary_copula(fit, 10, seed = 42))
Fit Prespecified Independent Groups of Variables
Description
Fits each group using the original estimator and fixes between-group latent correlations to zero.
Usage
fit_summary_copula_groups(studies, variables, groups, control = list())
Arguments
studies, variables, control |
As in |
groups |
List of character vectors partitioning all variable names into at least two groups of at least two variables each. Membership is prespecified, not estimated. |
Details
Studies reporting no variables from a group are omitted from that component fit and receive zero study scores for the group. The stacked sandwich covariance retains empirical cross-group score covariances. Group independence must be scientifically justified. It is not inferred from the supplied summaries. Full inference requires successful interior component fits and full rank of the stacked study scores.
Value
An object of class summary_copula_groups, with R,
margins, variables, groups, component fits,
study_ids, concatenated coefficients, parameter_indices,
stacked vcov, study_scores, study_influences,
within-group correlations, K, and diagnostics.
Use joint_probability_groups for event probabilities and
simulate_summary_copula for patient generation.
Joint Event Probabilities and Delta-Method Intervals
Description
Calculates the probability that all supplied bounds hold simultaneously, integrating over unspecified variables.
Usage
joint_probability(object, lower = numeric(), upper = numeric(),
inference = TRUE, level = .95)
joint_probability_groups(object, lower = numeric(), upper = numeric(),
inference = TRUE, level = .95)
Arguments
object |
A |
lower, upper |
Named numeric vectors of original-scale bounds. Unspecified bounds are infinite. Binary values are 0/1 and ordinal values are 1 through the specified number of levels. Discrete endpoints are inclusive. |
inference |
Whether to calculate uncertainty when available in the fit. |
level |
Confidence level, conventionally 0.95. |
Details
Deterministic multivariate normal integration supports at most six constrained variables per component. Larger events may be approximated from generated patients. Grouped probabilities factorize, but their uncertainty uses the complete stacked sandwich covariance. Intervals use the logit-scale delta method and include marginal-parameter uncertainty.
Value
A one-row data frame with estimate, se, lower, and
upper. Uncertainty is NA if inference is disabled or unavailable,
or the probability equals zero or one.
See Also
Generate Synthetic Patients from a Fitted Joint Distribution
Description
Draws latent multivariate normal vectors and applies the fitted marginal observation maps. These are synthetic records, not reconstructed original patients.
Usage
simulate_summary_copula(object, n, seed = NULL)
Arguments
object |
A fitted |
n |
Positive integer number of synthetic patients. |
seed |
Optional integer random seed. When supplied, the caller's random state is restored after generation; otherwise the current random stream is used. |
Value
A data frame with n rows and original variable names. Continuous
variables retain measurement units, binary values are 0/1, and ordinal values
are integers 1 through the number of levels.
See Also
Inspect a Summary Copula Fit
Description
Displays diagnostic flags and latent correlations, extracts internal-scale parameters and covariance, or calculates correlation intervals.
Usage
## S3 method for class 'summary_copula'
print(x, ...)
## S3 method for class 'summary_copula'
coef(object, ...)
## S3 method for class 'summary_copula'
vcov(object, ...)
## S3 method for class 'summary_copula'
confint(object, parm, level = .95, ...)
Arguments
x, object |
A fitted |
parm |
Currently ignored; intervals for all latent correlations are returned. |
level |
Confidence level. |
... |
Further arguments, currently unused. |
Value
print returns the fit invisibly; coef returns the internal
parameter vector; vcov returns its covariance matrix; confint
returns the correlation data frame with Fisher-transformed confidence limits.
Intervals are unavailable when the fit's diagnostic checks fail.