15  Sources of invalid inference

The following are recurring errors in longitudinal modeling comparisons. Each is stated with the symptom, the cause, and the remedy.

15.1 Evaluating on level rather than change

Symptom: reported \(R^2\) above 0.9 for next-visit prediction. Cause: the target is dominated by the baseline value, which is known. Remedy: evaluate on change or log-ratio, against the no-change baseline (Chapter 14).

15.2 Comparing arms that are algebraically equivalent

Symptom: differences between the differential-equation and supervised arms that are small, unstable across folds, and without consistent sign. Cause: at \(T = 2\) with no pooling difference, the arms are the same model (Chapter 3). Remedy: design the comparison to exercise extrapolation, coupling, or latent time.

15.3 Interpreting prior-dominated subject-level parameters

Symptom: subject-specific rate estimates that are smooth functions of baseline covariates and have implausibly low variance. Cause: the random effects are shrunk to the population mean; the model is reporting its prior. Remedy: report the shrinkage factor; do not interpret parameters that are not data-determined.

15.4 Treating cross-sectional covariance as dynamical coupling

Symptom: a dense estimated coupling matrix with a structure resembling known anatomical or functional networks. Cause: regions covary cross-sectionally for many reasons — shared measurement error, global scaling, common covariates — none of which imply that one region’s change drives another’s. Remedy: at small \(T\), restrict coupling claims to models with an externally derived structural prior, and state that the coupling is assumed rather than inferred.

15.5 Imposing a decomposition the data do not support

Symptom: results organized along canonical networks, modules, or pathways, in a dataset whose change covariance shows no components exceeding a noise null. Cause: the structure was supplied by the constraint rather than recovered from the data. A structurally constrained model fits plausibly whether or not the structure is present, so its fit is not evidence for the structure. Remedy: estimate the effective dimension first; test the prior against alignment and null-structure controls; where nothing exceeds the null, report the measurement floor rather than adopting the prior (Chapter 5).

15.6 Reading dimensionality without a noise model

Symptom: an apparent effective rank that changes substantially with parcellation, cohort, or preprocessing. Cause: independent measurement noise is full rank and inflates apparent dimensionality, while shared artifact is low rank and manufactures it. The two act in opposite directions, and neither is interpretable without an estimate of the noise covariance. Remedy: whiten by a scan–rescan noise estimate before decomposing; confirm with cross-validated rank and split-half reproducibility of loadings (Chapter 5).

15.7 Ignoring the compositional constraint

Symptom: nearly all regions show correlated change of similar sign and magnitude. Cause: a single global scaling factor is being recovered repeatedly. Remedy: model on the log scale with a global term, or work in proportions (Chapter 2).

15.8 Cross-sectional processing of longitudinal data

Symptom: implausibly large within-subject variability; apparent volume increases in regions expected to decline. Cause: independent segmentation of repeated scans. Remedy: use an unbiased within-subject template pipeline (Reuter et al. 2012).

15.9 Template leakage into held-out visits

Symptom: held-out-visit performance that exceeds plausible limits and degrades when the pipeline is rerun. Cause: the longitudinal template was constructed using the held-out timepoint. Remedy: rebuild templates excluding held-out visits.

15.10 Scanner change confounded with time

Symptom: a step change in derived measures at a known date, or site-specific rate differences. Cause: software or hardware change during follow-up. Remedy: longitudinal harmonization, explicit version terms, and sensitivity analysis in subjects with no scanner change (Chapter 4).

15.11 Unaddressed regression toward the mean

Symptom: a strong negative coefficient on baseline value in the change model, approximately proportional to the measurement error fraction. Cause: measurement error in the predictor. Remedy: errors-in-variables modeling, or reporting with and without the same-region baseline feature (Chapter 8).

15.12 Informative attrition

Symptom: population rate estimates that flatten at later visits. Cause: subjects with faster progression are less likely to return. Remedy: report retention per visit; use a joint or weighted model; report sensitivity.

15.13 Unregularized neural components in hybrid models

Symptom: the mechanistic parameter of a hybrid model differs substantially from its value when the model is fitted without the neural term. Cause: the neural term has absorbed mechanistic signal, or noise (Philipps et al. 2025). Remedy: penalize the neural term toward zero, report the ablation, and report the mechanistic parameter under both fits (Rooij et al. 2025).

15.14 Single-start optimization

Symptom: results that change between runs or with initialization. Cause: non-convex likelihood surfaces, which are the norm in this class of model. Remedy: multi-start optimization with reporting of the distribution of optima (Philipps et al. 2025).

15.15 Aggregate reporting across regions of heterogeneous reliability

Symptom: modest aggregate performance with no method clearly preferred. Cause: regions where change is below the measurement floor dominate the average. Remedy: stratify by signal-to-noise ratio and define the primary analysis set in advance (Chapter 4).