Compares an original dataset with its synthetic double across dataset-level
dimensions, numeric distributions, categorical distributions, and numeric
correlations. Returns a structured dataganger_comparison object.
Arguments
- original
The original data frame.
- synthetic
The synthetic data frame (from
synthesize_data()).- roles
Optional; a
dataganger_rolesobject fromdetect_roles().
Value
An S3 object of class dataganger_comparison, a list with
components dataset, numeric, categorical, relationship, interaction,
utility (a global pMSE-based utility diagnostic; see Details),
disclosure (role-aware disclosure risk diagnostics; see Details),
privacy_flags, and meta.
Details
utility reports the propensity-score pMSE / S_pMSE utility diagnostic of
Snoke, Raab, Nowok, Dibben & Slavkovic (2018), the same formula
synthpop::utility.gen() uses for its logistic-regression method:
S_pMSE near 1 means a model fit to distinguish original from synthetic
rows does no better than chance on the shared predictor columns (high
utility for that joint distribution); higher values mean the two datasets
are more detectably different. This is a utility measure, not a privacy
measure – it says nothing about disclosure risk.
disclosure evaluates role-aware disclosure risk diagnostics using
synthpop. Identity disclosure is assessed via replicated uniques
(synthpop::replicated.uniques()) across quasi-identifier keys identified in
roles: original uniques replicated in synthetic data indicate potential
re-identification risk. When sensitive target attributes are present,
attribute disclosure risk is evaluated via the DiSCO measure
(synthpop::disclosure()).
Examples
dat <- data.frame(x = 1:10, y = letters[1:10])
spec <- synth_spec(purpose = "demo")
syn <- synthesize_data(dat, spec)
#> Warning: Cannot guarantee level presence for columns: y (10 levels at k = 5 require
#> minimum n = 50). Largest minimum n = 50; output has n = 10. Restoring as many
#> levels as fit without removing another level's last copy.
compare_synthetic(dat, syn)
#> ℹ Not enough numeric columns (1) for correlation comparison.
#> Need at least 2 numeric columns with non-zero variance.
#>
#> ── DataGangeR Comparison ───────────────────────────────────────────────────────
#>
#> ── Dataset ──
#>
#> • Rows: 10 (original) -> 10 (synthetic)
#> • Columns: 2 (original) -> 2 (synthetic)
#> • Type match: 50%
#> • Missing: 0% (original) -> 0% (synthetic)
#>
#> ── Numeric -- top 3 by |standardized difference| ──
#>
#> • x: std diff = -0.429
#> Orig mean (SD): 5.5 (3.03)
#>
#> ── Categorical -- top 3 by distributional difference ──
#>
#> • y: p = 0.395, TVD = 1
#> Levels: 10 (orig) -> 10 (syn)
#>
#> ── Utility ──
#>
#> • S_pMSE = 2.11 (1.0 = the model could not tell original and synthetic rows
#> apart on 19 predictors; higher = more detectably different)
#> This is a utility measure, not a privacy measure -- a low score here is not
#> evidence the data are safe to release.
#>
#> ── Disclosure risk (synthpop) ──
#>
#> • Not computed: Roles must be provided to determine quasi-identifier keys and
#> sensitive targets.
