5 Dimensionality
How many variables, and how many drivers.
The preceding chapters treat \(T\), the number of observations per subject. This chapter treats \(R\), the number of variables measured per observation, and the distinction between \(R\) and the number of underlying drivers.
5.1 Two counts
Let \(R\) be the number of measured variables — regions in a parcellation, genes in an assay, analytes in a panel — and let \(d\) be the dimension of the process generating their coordinated change. In most biological systems \(d \ll R\): a few hundred regional volumes are driven by a small number of spatially organized processes; tens of thousands of transcripts are driven by a much smaller number of regulatory programs.
Nearly every scientific claim in this setting is a claim about \(d\), not about \(R\). \(R\) is a property of the measurement instrument and the analyst’s choice of parcellation or panel; \(d\) is a property of the system. Confusing the two produces both overstated dimensionality — treating \(R\) correlated measurements as \(R\) pieces of evidence — and unwarranted reduction, imposing a \(d\) taken from prior knowledge without checking that the data support it.
The central methodological requirement of this chapter is that \(d\) and its structure be estimated and tested, not assumed.
5.2 How \(R\) enters the model
\(R\) acts through three channels which push in different directions.
Replication. For any parameter shared across variables, \(R\) multiplies effective sample size: a rate law whose shape is pooled across regions is estimated from \(N \times R\) arcs rather than \(N\). This is the channel through which large \(R\) helps.
Coupling dimension. For any model in which variables interact, parameter count grows as \(R^2\): 676 at \(R = 26\), 10{,}000 at \(R = 100\), \(10^6\) at \(R = 1000\). This is the channel through which large \(R\) is destructive. The choice of \(R = 26\) bilateral regions in Ziegler et al. (2017) was forced by this scaling, and even at that size inference required model comparison over restricted connectivity structures rather than estimation of a free coupling matrix.
Multiplicity. Absent at \(R = 1\); routine false-discovery control at \(R \approx 100\); at \(R \approx 1000\) the variables are spatially dependent, so standard procedures are miscalibrated in ways that depend on the parcellation, and per-variable model selection compounds the problem.
5.3 Effective dimensionality is not \(R\)
Two facts limit what large \(R\) delivers.
Finer partition subdivides existing degrees of freedom. Change in neighboring regions is strongly correlated. Increasing parcellation resolution resamples the same underlying variation more finely rather than adding independent information. The covariance of change therefore has an effective rank far below \(R\), frequently by one to two orders of magnitude.
Per-variable signal-to-noise falls as \(R\) rises. Smaller regions contain fewer voxels, suffer greater partial-volume contamination, and are more sensitive to registration error. The measurement floor of Chapter 4 worsens monotonically with \(R\). After applying an SNR threshold, the effective number of usable variables can be far smaller than the nominal \(R\).
Together these imply that the useful range of \(R\) saturates well below the maximum a given parcellation offers, and that the saturation point should be determined empirically rather than by convention.
5.4 Cost in \(R\) by method family
| Method | Parameter cost | Behavior as \(R\) grows |
|---|---|---|
| D1 uncoupled rate law | \(O(R)\); \(O(1)\) if shape pooled | Improves; \(R\) acts as replication |
| D2 coupled system, general | \(O(R^2)\) | Infeasible beyond \(R \approx 30\) |
| D2 low rank, \(\mathbf{A}=\mathbf{U}\mathbf{V}^\top\) | \(O(Rk)\) | Tractable at large \(R\) |
| D3 network diffusion | \(O(1)\) | Unaffected by \(R\) |
| D4 latent time | \(O(R)\) trajectories, \(O(N)\) shifts | Improves; more variables constrain each subject’s stage |
| ML per variable | \(R\) separate models | Degrades; splits the sample \(R\) ways |
| ML pooled or multi-task | \(O(1)\) plus embeddings | Improves |
| H3 universal DE | \(O(R)\) inputs and outputs | Degrades; the neural term absorbs noise faster |
D3 is the only coupled specification whose parameter count does not grow with \(R\). Because \(-\beta\mathbf{L}\) has fixed structure, \(\mathbf{L}\) is eigendecomposed once and \(\exp(-\beta\mathbf{L}t)\) evaluated cheaply for any \(\beta\), so the model is as tractable at \(R = 10^5\) as at \(R = 100\). This property is the reason structurally constrained models dominate high-\(R\) applications, and it is also the reason their apparent success requires the controls described below.
5.5 Estimating \(d\)
The following sequence establishes whether low-dimensional structure is present before any structural prior is applied.
Remove the global component first. The compositional constraint of Chapter 2 guarantees a large first component corresponding to global scaling. It is an artifact of normalization choice, not evidence of a driver. Model it explicitly — a total-volume term, or working in log-ratio coordinates — before examining the remaining spectrum.
Compare the eigenvalue spectrum against a random-matrix null. For a data matrix whose entries are pure noise, the eigenvalues of the sample covariance follow the Marchenko–Pastur distribution with a sharp upper edge determined by the noise level and the ratio of dimensions. Eigenvalues above that edge are candidate signal; everything below is consistent with noise. This is the standard basis for denoising and rank selection in both diffusion MRI (Veraart et al. 2016) and single-cell transcriptomics (Aparicio et al. 2020), and it applies directly to a matrix of regional change. Permutation-based parallel analysis is a distribution-free alternative that makes weaker assumptions about noise homogeneity.
Whiten by the measurement noise covariance. Where scan–rescan data exist, the difference field gives a direct estimate of the noise covariance, which is not isotropic: registration and segmentation errors are spatially structured. Solving the generalized eigenvalue problem of change covariance against noise covariance, rather than the ordinary eigenproblem, separates components by signal-to-noise rather than by raw variance. This is the most informative version of the diagnostic and should be used wherever the reliability data permit.
Select rank by cross-validation. Fit a \(k\)-component decomposition on training subjects and measure reconstruction error of held-out subjects as a function of \(k\). The minimum is an estimate of usable rank and is generally smaller than the count of components exceeding a null threshold.
Check reproducibility of loadings. Split the sample, decompose each half independently, and measure the alignment of the resulting component loadings. A component whose spatial pattern does not replicate across halves should not be interpreted, whatever its eigenvalue.
5.5.1 Two failure modes that push in opposite directions
Independent measurement noise is full rank and therefore inflates apparent dimensionality: a noisy dataset can appear high-dimensional simply because no structure rises above the noise.
Shared artifact — global scaling, a common registration target, a scanner change affecting all variables — is low rank and therefore manufactures apparent low-dimensional structure. A strong leading component is as easily an artifact signature as a biological driver.
Because these act in opposite directions, neither a high nor a low apparent rank can be interpreted without the noise model. The reproducibility and whitening steps above are what distinguish them.
5.6 Testing a prior decomposition rather than imposing one
Where a decomposition is already available from prior work — a canonical network partition, a structural covariance parcellation, a connectome, a pathway annotation — the temptation is to impose it. The difficulty is that a model constrained by a prior structure will produce a plausible and well-fitting result whether or not that structure is present in the data, because the structure is supplied by the constraint rather than recovered from the observations. A good fit is therefore not evidence for the prior.
Three controls make the prior testable. Their spatial-specific form, including the requirement that a null preserve spatial autocorrelation, is given in Chapter 6.
Alignment. Quantify the correspondence between the data-driven components estimated above and the prior partition — for example, the proportion of change variance the prior partition explains, compared against partitions matched to it in number and size but otherwise arbitrary.
Null-structure controls. Refit the constrained model with the prior structure replaced by null structures that preserve its nuisance properties while destroying the feature of interest. For a connectome, degree-preserving rewiring by edge swapping and spatial nulls that shuffle node positions while retaining connection profiles are standard; Zheng et al. (2019) used both, and reported that the empirical network outperformed each across all densities. The catalogue of appropriate nulls and their properties is reviewed by Váša and Mišić (2022). A structurally constrained model that does not outperform matched nulls has demonstrated nothing about the structure.
Nested comparison. Fit three specifications on identical folds: an unconstrained low-rank model with rank set by cross-validation, the prior-constrained model, and where feasible the free model. If the prior constraint costs predictive accuracy relative to the data-driven decomposition, the prior is not the structure operating in this dataset.
5.6.1 Interpreting the outcome
| Components above null | Aligned with prior | Interpretation |
|---|---|---|
| No | — | No recoverable structure. A structure-constrained fit would report the prior, not the data. Report the measurement floor and stop. |
| Yes | Yes | The prior is supported. The constrained model is justified and inherits its parsimony advantage. |
| Yes | No | Structure is present but the prior is wrong for this cohort, measure, or interval. Use the data-driven decomposition and report the discrepancy. |
The first row is the case that matters most in practice and is the one most often handled incorrectly. Noisy, small, or poorly harmonized data will not exhibit recoverable low-dimensional structure, and applying a canonical network partition to such data yields network-organized results by construction. The appropriate response is to report that the data do not support the decomposition, not to adopt it because it is well established elsewhere.
Prior knowledge that a small number of drivers exists in a system is a reason to test for them, not a licence to assume them in a particular dataset.
5.7 Behavior at three scales
\(R = 1\). A pure longitudinal problem. All \(N\) subjects support a single target, so a richly parameterized nonlinear mixed model is affordable. No coupling question, no multiplicity, no compositional constraint. Classical nonlinear mixed effects is strongest here and flexible supervised learning has least to add.
\(R \approx 100\). Coupling becomes a scientific question while remaining out of reach as a free parameter. Low-rank and structurally constrained specifications are the practical options; pooled and multi-task models deliver their largest gain; the compositional constraint is material. This is the scale assumed by the method ladder in Chapter 7 and Chapter 8.
\(R \gtrsim 1000\). The question changes character. Rather than asking which variables couple, one estimates \(d\), fits the dynamical model in the reduced space, and projects back. Spatial adjacency becomes exploitable: variables are no longer exchangeable, and a spatial prior is more appropriate than treating them as independent units. This is a different workflow, not a scaled version of the previous one.
5.8 Interaction with \(T\)
\(R\) and \(T\) are not substitutes.
- Coupling requires both. No value of \(R\) recovers directional coupling at \(T = 2\); cross-sectional covariance among variables is not dynamical coupling (Chapter 15). Equally, no value of \(T\) at \(R = 1\) says anything about coupling.
- \(R\) can partially substitute for \(T\) in estimating trajectory shape, under the assumption that variables share a common shape and differ only in rate and onset. The variables then act as replicates for the shape parameters, and shape becomes estimable at lower \(T\) than a single-variable analysis would allow. This is the mechanism disease-progression models exploit (Oxtoby 2023). It is an assumption and it fails where subsets follow genuinely different forms.
- \(T\) does not substitute for \(R\) in anything concerning spatial or structural organization.
Compactly: \(T\) constrains dynamical structure, \(R\) contributes replication and spatial structure, the product \(NRT\) governs shared parameters, and \(T\) alone governs anything subject-specific and dynamical.
5.9 Recommended sequence
- Estimate the measurement noise covariance (Chapter 4).
- Remove or model the global component.
- Estimate \(d\) by noise-whitened decomposition with a random-matrix or permutation null, confirmed by cross-validated rank and split-half reproducibility.
- Report the spectrum. If nothing exceeds the null, stop and report the floor.
- Test any prior decomposition against alignment and null-structure controls before adopting it.
- Fit the dynamical model in the reduced space, where a general coupling matrix is affordable again, and project back for reporting.