13 Sources of invalid inference
The following are recurring errors in longitudinal modeling comparisons. Each is stated with the symptom, the cause, and the remedy.
13.1 Evaluating on level rather than change
Symptom: reported \(R^2\) above 0.9 for next-visit prediction. Cause: the target is dominated by the baseline value, which is known. Remedy: evaluate on change or log-ratio, against the no-change baseline (Chapter 12).
13.2 Comparing arms that are algebraically equivalent
Symptom: differences between the differential-equation and supervised arms that are small, unstable across folds, and without consistent sign. Cause: at \(T = 2\) with no pooling difference, the arms are the same model (Chapter 3). Remedy: design the comparison to exercise extrapolation, coupling, or latent time.
13.3 Interpreting prior-dominated subject-level parameters
Symptom: subject-specific rate estimates that are smooth functions of baseline covariates and have implausibly low variance. Cause: the random effects are shrunk to the population mean; the model is reporting its prior. Remedy: report the shrinkage factor; do not interpret parameters that are not data-determined.
13.4 Treating cross-sectional covariance as dynamical coupling
Symptom: a dense estimated coupling matrix with a structure resembling known anatomical or functional networks. Cause: regions covary cross-sectionally for many reasons — shared measurement error, global scaling, common covariates — none of which imply that one region’s change drives another’s. Remedy: at small \(T\), restrict coupling claims to models with an externally derived structural prior, and state that the coupling is assumed rather than inferred.
13.5 Ignoring the compositional constraint
Symptom: nearly all regions show correlated change of similar sign and magnitude. Cause: a single global scaling factor is being recovered repeatedly. Remedy: model on the log scale with a global term, or work in proportions (Chapter 2).
13.6 Cross-sectional processing of longitudinal data
Symptom: implausibly large within-subject variability; apparent volume increases in regions expected to decline. Cause: independent segmentation of repeated scans. Remedy: use an unbiased within-subject template pipeline (Reuter et al. 2012).
13.7 Template leakage into held-out visits
Symptom: held-out-visit performance that exceeds plausible limits and degrades when the pipeline is rerun. Cause: the longitudinal template was constructed using the held-out timepoint. Remedy: rebuild templates excluding held-out visits.
13.8 Scanner change confounded with time
Symptom: a step change in derived measures at a known date, or site-specific rate differences. Cause: software or hardware change during follow-up. Remedy: longitudinal harmonization, explicit version terms, and sensitivity analysis in subjects with no scanner change (Chapter 4).
13.9 Unaddressed regression toward the mean
Symptom: a strong negative coefficient on baseline value in the change model, approximately proportional to the measurement error fraction. Cause: measurement error in the predictor. Remedy: errors-in-variables modeling, or reporting with and without the same-region baseline feature (Chapter 6).
13.10 Informative attrition
Symptom: population rate estimates that flatten at later visits. Cause: subjects with faster progression are less likely to return. Remedy: report retention per visit; use a joint or weighted model; report sensitivity.
13.11 Unregularized neural components in hybrid models
Symptom: the mechanistic parameter of a hybrid model differs substantially from its value when the model is fitted without the neural term. Cause: the neural term has absorbed mechanistic signal, or noise (Philipps et al. 2025). Remedy: penalize the neural term toward zero, report the ablation, and report the mechanistic parameter under both fits (Rooij et al. 2025).
13.12 Single-start optimization
Symptom: results that change between runs or with initialization. Cause: non-convex likelihood surfaces, which are the norm in this class of model. Remedy: multi-start optimization with reporting of the distribution of optima (Philipps et al. 2025).
13.13 Aggregate reporting across regions of heterogeneous reliability
Symptom: modest aggregate performance with no method clearly preferred. Cause: regions where change is below the measurement floor dominate the average. Remedy: stratify by signal-to-noise ratio and define the primary analysis set in advance (Chapter 4).