Aggregate subgroup information for the relaxed model
Source:vignettes/subgroup-identification.Rmd
subgroup-identification.RmdThe relaxed model gives the index and comparator treatments separate
covariate coefficients. The IPD identify beta_index; the
comparator coefficients are informed only through the aggregate
likelihood. Aggregate subgroup rows can add information, but a row count
alone cannot establish that all comparator coefficients are
identified.
This vignette separates three questions:
- Is the aggregate design structurally capable of separating the comparator intercept and coefficients?
- Does the nonlinear likelihood contain enough information at the fitted parameter values?
- Is the resulting target-population effect stable to the prior, integration approximation, and transport assumptions?
Only the first question has a simple exact answer, and only for the normal identity-link model.
What an aggregate row contributes
For the binomial, normal, and Poisson families, each mutually exclusive aggregate stratum contributes one scalar outcome summary. With K covariates, the relaxed comparator model has K + 1 coefficients: an intercept and K slopes. Therefore
S \geq K + 1
is a necessary counting condition. It is not sufficient.
For a normal identity-link model, stratum s has mean
\bar y_s = \mu_c + m_s^\mathsf{T}\beta_c,
where m_s is the declared covariate-mean vector. The exact aggregate design matrix is D=[\mathbf{1},M]. The comparator coefficients can be separated by the likelihood only when D has full column rank K+1. Equivalently, after centering the rows of M, the covariate-mean matrix must have rank K.
For a nonlinear mean model, each row instead contributes
\bar y_s = \mathbb{E}_{X\sim f_s}\left[g^{-1}(\mu_c+X^\mathsf{T}\beta_c)\right].
Its local Jacobian depends on the full within-stratum distribution
f_s, the link, and the parameter
values. Mean vectors alone do not determine that Jacobian. Differences
in variances, binary probabilities, or dependence can add information
that a mean-only rank calculation cannot see. Consequently,
check_identification() treats the mean-profile spectrum as
descriptive for binomial, Poisson, and normal-log models. When the row
count is adequate, flagged is NA, not
FALSE.
These distinctions follow the aggregate marginalization used by ML-NMR and ML-UMR (Phillippo et al. 2020; C. Chandler and Ishak 2026).
Rows must describe one population
Multiple set_agd() rows are assumed to be mutually
exclusive, collectively exhaustive strata of one comparator population.
They are not separate, overlapping subgroup analyses. Population
summaries are combined with outcome_n for binomial and
normal outcomes and with person-time exposure for Poisson outcomes. For
a normal comparator mean,
\bar y_c=\sum_s w_s\bar y_s,\qquad \operatorname{Var}(\bar y_c)=\sum_s w_s^2 SE_s^2, \qquad w_s=\frac{n_s}{\sum_j n_j},
assuming independent, mutually exclusive stratum summaries. Inverse-variance weights would describe a meta-analysis of estimates, not the mean of the comparator population.
If Poisson exposure varies with covariates, reported covariate summaries must match the person-time distribution relevant to the rate. Ordinary participant means are justified only under a fixed-exposure or suitable independence assumption (Remiro-Azocar et al. 2022).
Mean-profile geometry
check_identification() centers the declared subgroup
mean vectors, scales each covariate by its IPD SD, and examines the
singular values. Scaling prevents a covariate measured in large units
from dominating a binary proportion merely because of units. The
reported
\texttt{cond\_inv}=\frac{d_{\min}}{d_{\max}}
approaches zero as the mean profiles collapse onto a lower-dimensional set. For the identity-link model this summarizes the exact design matrix. For nonlinear models it describes only the spread of the declared means.
Consider one continuous and two binary covariates. The following designs all contain several rows, but they span different covariate directions.
ref_sd <- c(age = 1, sex = 0.5, prior = 0.49)
by_age_only <- cbind(
age = seq(-1.5, 1.5, length.out = 6),
sex = 0.5,
prior = 0.4
)
cross_tab <- rbind(
c(0, 0, 0), c(0, 1, 0), c(0, 0, 1), c(0, 1, 1)
)
jointly_defined <- rbind(
c(-1, 0, 0), c(-1, 1, 0), c(-1, 0, 1), c(-1, 1, 1),
c(1, 0, 0), c(1, 1, 1)
)
geometry <- function(M, ref_sd) {
M <- as.matrix(M)
K <- ncol(M)
M <- scale(M, center = TRUE, scale = FALSE)
M <- sweep(M, 2, ref_sd, "/")
d <- svd(M)$d
if (length(d) < K) d <- c(d, rep(0, K - length(d)))
c(cond_inv = min(d) / max(d),
eff_dim = sum(d^2)^2 / sum(d^4))
}
do.call(rbind, lapply(
c("by_age_only", "cross_tab", "jointly_defined"),
function(nm) c(design = nm, rows = nrow(get(nm)),
round(geometry(get(nm), ref_sd), 3))
))
#> design rows cond_inv eff_dim
#> [1,] "by_age_only" "6" "0" "1"
#> [2,] "cross_tab" "4" "0" "1.999"
#> [3,] "jointly_defined" "6" "0.707" "2.764"Age bands vary only in the age direction. The binary cross-tab varies in two directions but leaves age fixed. Only the jointly defined cells span all three covariates. More rows along an existing direction cannot replace a missing direction.
Using check_identification()
The function uses declared AgD means, so the identity-link result
does not depend on the realized QMC points and can be run before
add_integration().
set.seed(2026)
n_ipd <- 120
ipd_df <- data.frame(
trt = "A",
y = rnorm(n_ipd, 2, 1),
age = rnorm(n_ipd),
sex = rbinom(n_ipd, 1, 0.5),
prior = rbinom(n_ipd, 1, 0.4)
)
ipd <- set_ipd(ipd_df, treatment = "trt", outcome = "y",
covariates = c("age", "sex", "prior"), family = "normal")
make_design <- function(M) {
agd_df <- data.frame(
trt = "B",
n = rep(100, nrow(M)),
ybar = 2 + 0.1 * M[, 1] + 0.2 * M[, 2] - 0.1 * M[, 3],
yse = rep(0.1, nrow(M)),
age_mean = M[, 1],
age_sd = rep(1, nrow(M)),
sex_mean = M[, 2],
prior_mean = M[, 3]
)
agd <- set_agd(
agd_df, treatment = "trt", family = "normal",
outcome_mean = "ybar", outcome_se = "yse", outcome_n = "n",
cov_means = c("age_mean", "sex_mean", "prior_mean"),
cov_sds = c("age_sd", NA, NA),
cov_types = c("continuous", "binary", "binary")
)
combine_data(ipd, agd)
}
identity_age <- check_identification(
make_design(by_age_only), link = "identity", verbose = FALSE
)
identity_joint <- check_identification(
make_design(jointly_defined), link = "identity", verbose = FALSE
)
nonlinear_joint <- check_identification(
make_design(jointly_defined), link = "log", verbose = FALSE
)
data.frame(
design = c("age only", "joint identity", "joint log"),
scope = c(identity_age$diagnostic_scope,
identity_joint$diagnostic_scope,
nonlinear_joint$diagnostic_scope),
rows = c(identity_age$n_rows, identity_joint$n_rows,
nonlinear_joint$n_rows),
cond_inv = c(identity_age$cond_inv, identity_joint$cond_inv,
nonlinear_joint$cond_inv),
flagged = c(identity_age$flagged, identity_joint$flagged,
nonlinear_joint$flagged)
)
#> design scope rows cond_inv flagged
#> 1 age only identity 6 1.487234e-33 TRUE
#> 2 joint identity identity 6 7.070856e-01 FALSE
#> 3 joint log descriptive 6 7.070856e-01 NAThe package threshold cond_inv < 0.2 is an
exploratory warning rule for the identity-link design. It is not a
validated universal cutoff. The statistic is relative: it cannot see
subgroup sample sizes, outcome precision, overall between-stratum
magnitude, nonlinear link curvature, or model misspecification. A value
above 0.2 does not prove identification.
Recommended workflow
- Verify that rows are mutually exclusive and collectively exhaustive, and that their sample sizes or exposures represent the intended comparator population.
- Count the scalar outcome summaries. Fewer than K+1 rows cannot separate a relaxed comparator intercept and K slopes for the non-survival families.
- For normal identity models, inspect design rank and
cond_inv. Request jointly defined cells that vary across all covariates when the design is deficient. - For nonlinear models, treat the mean spectrum as descriptive. Check
the fitted coefficient posterior, sampler diagnostics, and
prior_sensitivity(). A prior sweep measures sensitivity; it does not assign a fraction of information to the data or the prior. - Check
check_integration()separately. Numerical fidelity of the QMC grid and statistical identification are different questions. - Examine overlap before transporting the comparator regression to the index or another target covariate distribution. A stable fit inside the comparator support does not validate extrapolation outside it (Phillippo et al. 2016; Conor Chandler and Ishak 2026).
- If the relaxed result is unstable, obtain more informative jointly defined strata, simplify the effect-modification structure, use SPFA with a clear assumption statement, and report prior and target-population sensitivity.
The marginal posterior-to-prior variance change printed for relaxed coefficients is descriptive. It can be negative because of posterior correlation or changed marginal geometry, and neither its sign nor magnitude is an identification certificate.
Distribution and dependence assumptions
Each nonlinear aggregate likelihood integrates over a reconstructed
within-stratum covariate distribution. Means, SDs, binary proportions,
and a copula do not uniquely identify that distribution. The same
supplied correlation matrix is used as a within-stratum dependence model
across rows. A pooled comparator correlation also contains
between-stratum covariance and is not generally the required
within-stratum correlation. For non-Gaussian continuous margins, supply
a defensible Spearman matrix with cor_adjust = "spearman"
or a calibrated latent Gaussian-copula matrix with
cor_adjust = "none"; an observed Pearson matrix is not
automatically either.
These are sensitivity assumptions, not details that subgroup row count can repair. Vary plausible dependence structures and marginal distributions when the estimand is sensitive to them.
Time-to-event outcomes
The counting argument does not apply to reconstructed survival data.
A comparator curve supplied through set_agd_surv()
contributes event and censoring likelihood terms across time, not one
scalar summary. Depending on the survival family and covariate
distribution, one marginal curve may identify some combinations of
comparator coefficients while leaving other directions weak.
check_identification() therefore refuses survival rather
than returning a misleading row-count result.
For a relaxed survival fit, inspect the comparator coefficient
posterior and run prior_sensitivity() with
prior_beta_comparator_scales. Also compare plausible
baseline assumptions, including aux_by = ".study" and
aux_by = "none" where scientifically defensible. RMST
differences are collapsible within a specified population, but this does
not make them automatically transportable to another population.
Simulation evidence
The sections above argue from geometry. This section measures whether
that geometry actually predicts anything. The script that produced it is
a development file, kept in the package repository rather than installed
with the package, and is read there: data-raw/simulate_subgroup_identification.R.
It writes a manifest beside its results recording the software that
produced them. The numbers below were read from a results file with MD5
4dc62e120a5b2ad21733c13d843ef9ec; a test compares that against the file
in the repository, so a regenerated result set that has not been written
back into this article fails the build rather than shipping quietly.
Six aggregate designs are crossed with the three non-survival families and 300 replications each, 5400 fits in total.
Every design is a tabulation of the same people. Each replication draws one index study and one comparator cohort at the individual level, then reports that single cohort six ways: as age bands, as a cross-tab of the two binary covariates, and as several designs that cut across all three. Each aggregate row carries the realized means and standard deviations of the members that fall in it, as a real report would, rather than a stratum mean pushed through the link.
That control is what makes the comparison mean anything, and holding the total sample size fixed is not enough on its own. Designs built around separate target profiles also differ in their pooled covariate distribution, in how far that distribution sits from the index population, and in their expected outcome level; a width difference could then come from any of those rather than from the geometry under study. Under a partition all of them are identical by construction, and the only thing that varies is where the cuts fall.
The designs are compared within a replication, not across replications. Because the six share one draw, the difference between two of them is a paired contrast: the replication-to-replication variation both designs see cancels, and what remains is the effect of the tabulation. The contrasts below are reported that way, with the Monte Carlo standard error computed from the paired differences, because that standard error describes the contrast rather than the two cell means separately. How much precision the pairing actually buys is reported with the contrasts, and it is not uniform.
Coverage is scored against the estimand the package
reports. The Stan models define delta_index as an
average over the realized index rows, so it is a property of the index
sample in hand and moves from one replication to the next. The truth is
therefore recomputed on each replication’s own covariates, which makes
it exact for that replication rather than approximate: it evaluates the
very same average the generated quantity estimates.
Non-convergence is an outcome, not an observation.
Interval width and coverage are summarized over fits that met every
retained diagnostic (split-Rhat at most 1.01, bulk and tail effective
sample size at least 400, no divergent transitions, no treedepth
saturation, and every chain finished); the rest are counted in their own
column. A posterior that has not converged is not a valid measurement of
what a design can achieve, and pooling it into a width average measures
the sampler as much as the design. The Failed column counts
fits that errored outright, and the study records the condition class
and message for each so a systematic failure can be told from a
transient one.
The comparator covariates are generated independently, and the integration is given that correlation structure rather than one estimated from the index sample. Left to itself the package estimates a matrix from the individual data and transports it to every aggregate row, which would add a second, replication-varying misspecification on top of the design being studied.
| Family | Design | Rows | Row n | cond_inv | Fitted | eff_dim | Width | Width CV | MCSE width | Coverage | MCSE cov | Failed | Not conv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| binomial | joint_8 | 8 | 39-115 | 0.78000 | 0.780 | 2.88 | 0.677 | 0.027 | 0.0010 | 0.953 | 0.012 | 0 | 0 |
| binomial | joint_6 | 6 | 39-190 | 0.75000 | 0.740 | 2.83 | 0.685 | 0.028 | 0.0011 | 0.953 | 0.012 | 0 | 0 |
| binomial | minimal_4 | 4 | 87-190 | 0.56000 | 0.560 | 2.46 | 0.711 | 0.037 | 0.0015 | 0.953 | 0.012 | 0 | 0 |
| binomial | crosstab_4 | 4 | 87-210 | 0.00043 | 0.025 | 2.00 | 1.872 | 0.156 | 0.0237 | 0.987 | 0.009 | 0 | 148 |
| binomial | ovat_age_6 | 6 | 100-100 | 0.00042 | 0.052 | 1.00 | 1.171 | 0.171 | 0.0115 | 0.993 | 0.005 | 0 | 0 |
| binomial | ovat_age_3 | 3 | 200-200 | 0.00000 | 0.000 | 1.00 | 1.251 | 0.171 | 0.0124 | 0.993 | 0.005 | 0 | 0 |
| normal | joint_8 | 8 | 42-113 | 0.78000 | 0.780 | 2.88 | 0.342 | 0.042 | 0.0008 | 0.940 | 0.014 | 0 | 0 |
| normal | joint_6 | 6 | 42-187 | 0.75000 | 0.740 | 2.83 | 0.345 | 0.042 | 0.0008 | 0.953 | 0.012 | 0 | 0 |
| normal | minimal_4 | 4 | 90-188 | 0.56000 | 0.560 | 2.46 | 0.362 | 0.046 | 0.0010 | 0.940 | 0.014 | 0 | 0 |
| normal | crosstab_4 | 4 | 90-218 | 0.00043 | 0.028 | 2.00 | 2.651 | 0.487 | 0.0745 | 0.987 | 0.007 | 0 | 0 |
| normal | ovat_age_6 | 6 | 100-100 | 0.00042 | 0.047 | 1.00 | 0.845 | 0.329 | 0.0160 | 0.970 | 0.010 | 0 | 0 |
| normal | ovat_age_3 | 3 | 200-200 | 0.00000 | 0.000 | 1.00 | 1.359 | 0.308 | 0.0242 | 0.990 | 0.006 | 0 | 0 |
| poisson | joint_8 | 8 | 41-114 | 0.78000 | 0.780 | 2.88 | 0.219 | 0.034 | 0.0004 | 0.957 | 0.012 | 0 | 0 |
| poisson | joint_6 | 6 | 41-188 | 0.75000 | 0.740 | 2.83 | 0.221 | 0.033 | 0.0004 | 0.953 | 0.012 | 0 | 0 |
| poisson | minimal_4 | 4 | 93-189 | 0.56000 | 0.560 | 2.46 | 0.229 | 0.041 | 0.0005 | 0.950 | 0.013 | 0 | 0 |
| poisson | crosstab_4 | 4 | 93-220 | 0.00043 | 0.027 | 2.00 | 1.543 | 0.645 | 0.0710 | 0.970 | 0.012 | 0 | 103 |
| poisson | ovat_age_6 | 6 | 100-100 | 0.00042 | 0.050 | 1.00 | 0.513 | 0.352 | 0.0105 | 0.939 | 0.014 | 0 | 6 |
| poisson | ovat_age_3 | 3 | 200-200 | 0.00000 | 0.000 | 1.00 | 0.719 | 0.287 | 0.0119 | 0.997 | 0.003 | 0 | 1 |
The cond_inv column is the geometry of the design as
requested: the same statistic evaluated at the population profiles.
It separates the six designs into two groups three orders of magnitude
apart. ovat_age_3, ovat_age_6 and
crosstab_4 fall below 0.001, because their requested
profiles span fewer than three directions; the other three sit above
0.5. None of them is exactly zero, because that statistic is computed on
a two-million-draw approximation of the population, so an exactly absent
direction reads as the floor of the approximation rather than as 0.
The Fitted column is the same quantity computed from the
rows the likelihood actually received, and it is not that number. A
stratum defined as holding sex at one half realizes a proportion near
one half but not exactly one half, so its profile matrix is numerically
full rank. Only ovat_age_3 is exactly rank deficient
whatever the draw, because three rows cannot span three directions after
centering.
What survives is a separation, not an exact zero. Across all 5,400 fits the fitted conditioning of the three deficient designs never rises above 0.12, and that of the three identified designs never falls below 0.47. The groups do not overlap, though the fitted gap is a factor of about 4, far narrower than the requested one: sampling noise does repay some of the missing direction. So “unidentified” below means a direction the design was never asked to supply and that noise repays only at a conditioning too small to be useful, not a matrix a computer would call singular.
Row count does not order the designs; geometry does
| Predictor | binomial | normal | poisson |
|---|---|---|---|
| row count | -0.77 | -0.77 | -0.77 |
| cond_inv (requested) | -0.83 | -0.83 | -0.83 |
| cond_inv (fitted) | -0.94 | -0.94 | -0.94 |
| eff_dim | -0.81 | -0.81 | -0.81 |
Six designs per family is a coarse basis for a rank correlation, and ties are extensive, so read the table as a direction rather than a calibration. The paired contrasts below carry the weight, because they compare designs on the SAME draw and so are not limited by having only six of them.
| Family | Design | Pairs | Width difference | MCSE | Wider in |
|---|---|---|---|---|---|
| binomial | crosstab_4 | 152 | 1.162 | 0.0227 | 100% |
| binomial | joint_6 | 300 | -0.027 | 0.0010 | 5% |
| binomial | joint_8 | 300 | -0.034 | 0.0011 | 1% |
| binomial | ovat_age_3 | 300 | 0.540 | 0.0118 | 100% |
| binomial | ovat_age_6 | 300 | 0.460 | 0.0110 | 100% |
| normal | crosstab_4 | 300 | 2.289 | 0.0742 | 100% |
| normal | joint_6 | 300 | -0.017 | 0.0006 | 4% |
| normal | joint_8 | 300 | -0.020 | 0.0005 | 2% |
| normal | ovat_age_3 | 300 | 0.997 | 0.0240 | 100% |
| normal | ovat_age_6 | 300 | 0.483 | 0.0159 | 100% |
| poisson | crosstab_4 | 197 | 1.313 | 0.0708 | 100% |
| poisson | joint_6 | 300 | -0.008 | 0.0003 | 9% |
| poisson | joint_8 | 300 | -0.010 | 0.0004 | 6% |
| poisson | ovat_age_3 | 299 | 0.490 | 0.0119 | 100% |
| poisson | ovat_age_6 | 294 | 0.284 | 0.0106 | 100% |
Each row is that design minus the four-row minimal_4,
over the replications where both converged. A positive difference is a
wider, worse interval.
What the pairing buys splits along the same line as everything else. For the identified designs the paired standard error is 42% to 59% smaller than differencing two independently estimated cell means, because both designs respond to the shared draw in much the same way and that response cancels. For the deficient designs it is 0% to 5% smaller, which is to say nothing at all: their width moves with the unidentified direction, which the reference design does not share, so there is little common variation left to cancel. Pairing is still the right standard error to report, and for these designs it is not a cheaper one.
Two things are worth separating here, because they carry different weight.
Ranked against the mean width, cond_inv orders these
designs better than the row count does, but the margin is modest and
rests on six designs per family, which is a thin basis for a
correlation. Read it as a direction, not as a calibration.
The inversion is the part that does not depend on how many designs were tried. The six-row one-variable-at-a-time design has more rows than the four-row minimal design and yields a wider, less stable interval, in every family. A reader counting rows would prefer the worse design. That is the concrete content of the claim that more rows along an existing direction cannot replace a missing direction, and one counterexample is enough to establish it.
The two failing designs lose opposite halves of the space
Both the one-variable-at-a-time bands and the bare cross-tab fall in
the deficient group, and neither recovers a usable third direction from
sampling noise, but they do not fail in the same way. The bands vary
only age, leaving eff_dim at 1 whether they have three rows
or six. The cross-tab varies both binaries and holds age fixed, giving
eff_dim = 2. One keeps the continuous covariate and loses
the binaries; the other keeps the binaries and loses the continuous
covariate.
An unidentified design is unpredictable, not merely imprecise
The failing designs are not just wider. Their widths vary far more from one replication to the next than the identified designs do, so a single analysis cannot be read as representative of what that design would have produced. The coefficient of variation of the width separates the two groups completely: every identified cell is at or below 0.05 and every deficient one at or above 0.16, reaching 0.64 in the worst cell. The widths themselves are far more separated than that; it is the stability of the width, not its size, that this measures.
The sampler notices, and only for the nonlinear families
Non-convergence tracks identification exactly in this study, across every diagnostic recorded. No fit in an identified design failed any of them, and all 258 fits that failed at least one belong to the deficient group.
The diagnostics disagree sharply about how many fits are affected: 130 exceed the split-Rhat threshold, 115 fall below the bulk effective sample size one, 14 below the tail one, 94 report a divergent transition, 0 saturate the treedepth, and 0 lost a chain. A study reporting Rhat alone would have understated the problem, which is why the table above filters on all six rather than on the one that is easiest to extract. So a rank-deficient aggregate design does not merely widen the interval, it also makes the posterior harder to sample, and the diagnostics one would run anyway will often say so.
The counts here come from the summaries the script retains rather than from what the sampler printed, because it suppresses fitting warnings; the two describe the same runs.
They do not always say so. Every fit failing a diagnostic here is binomial or Poisson; the normal family sampled cleanly in the unidentified designs too. A plausible reading is that with an identity link the unidentified direction is a flat ridge, which the sampler traverses without difficulty, while a logit or log link curves it. That is an explanation for the pattern rather than something this study tested, and the practical point does not depend on it: a clean set of convergence diagnostics is not evidence that the design is identified.
Coverage does not reveal any of this. The unidentified designs cover at or near the nominal rate, because a wide enough interval covers almost anything. An interval can be correct and useless at the same time, which is why the width and its variability, not the coverage, are what these designs should be judged on.
The shipped screen is exact where it speaks, and silent where it matters most
check_identification() ran on every design in every
replication and its verdict was kept, so the study also measures the
diagnostic this article recommends.
Under the identity link it is exact. Across the 1800 normal fits it flagged 900 and cleared 900, and that split is the deficient and identified split above, fit for fit: no identified design was flagged and no deficient one was missed.
Under the logit and log links it mostly declines. For 3000 of the
3600 binomial and Poisson fits it returns no verdict at all, which is
the refusal documented in ?check_identification: a
subgroup-mean screen is an exact statement about the design matrix only
when the link is the identity. The refusal is preserved here rather than
collapsed into “not flagged”, because reporting it as a clearance would
publish a verdict the package deliberately withholds.
What it still flags there is ovat_age_3, and on a
different ground: three rows is fewer than the K+1 = 4 that three covariates need, and a row
count is structural rather than link-dependent. So under a nonlinear
link the screen catches the design that is short of rows, and says
nothing about crosstab_4 or ovat_age_6, which
have rows enough and the wrong geometry.
The sampler does not reliably cover that gap. Of the 1200 deficient
binomial and Poisson fits where the screen abstains, 943 met every
diagnostic, which is 79% of them. The trouble is concentrated in
crosstab_4; ovat_age_6 samples cleanly in 99%
of its nonlinear fits while producing the third-widest intervals of the
six designs in its own family. A clean set of diagnostics is not
evidence that the design is identified, and here it is not even good
evidence.
That is the practical conclusion, and it is a negative one about the
two signals that arrive after fitting. Under a nonlinear link
the screen withholds a verdict, and the sampler mostly reports success,
for a design whose intervals are among the worst here. What separates
the designs in every family is the geometry, cond_inv and
eff_dim, and that is computable from the aggregate table
before anything is fitted. Compute it first; do not wait for the fit to
object.
What this vignette does not claim
There is no universal minimum number of rows beyond the necessary
K+1 count, no source-validated
cond_inv cutoff, and no guarantee that a non-flagged design
produces a practically precise target effect.
The simulation above is deliberately narrow. It uses synthetic strata rather than a real trial’s reported subgroups, six designs rather than a systematic sweep of the design space, three covariates, one effect size, one comparator size, and no survival family. Its correlations rest on six designs per family, which is a thin basis; read the ranking as a direction, not as a calibration. What it does establish is the inversion, and one counterexample is enough for that: a design with more rows can be strictly worse than a design with fewer.
The generator draws each cohort at the individual level and each
aggregate row reports the realized moments of its own members, so a
row’s members are exactly the people its summary describes. What the row
does not carry is the SHAPE of that membership: a design that bands on
age reports a mean and a standard deviation for a within-band age
distribution that is truncated, while the fit integrates a normal with
those two moments, and only crosstab_4 leaves the age
margin approximately Gaussian. This is a real misspecification, and it
falls unevenly across the designs, so it is worth bounding rather than
asserting away. Comparing each row’s true mean response against the one
the model integrates puts the discrepancy at most 4.2% of that row’s own
sampling standard error across every design and both nonlinear families,
and the age-banded designs sit lower than that, below the sampling noise
of the comparison itself. The width differences the table reports run to
factors of two and three, so the shape of the within-row age
distribution is not what produces them.
A study that justified a cond_inv threshold for a
specific application would need the actual within-stratum distributions
of that application, its own effect size and sample sizes, a design
sweep wide enough to place a cutoff, and the outcome family in use. The
guidance here remains structural and diagnostic, not a numerical
calibration.