# Round 2: Theorist's Final Contribution
## The Fundamental Statistical Challenge: Pseudo-Bulk Inference Under Cellular Depletion
Building upon the critiques raised in previous rounds, I wish to articulate what I consider the most profound methodological limitation of the SEA-AD differential expression framework—one that undermines the interpretability of cell-type-specific findings at their core. The central problem is that differential expression analysis in snRNA-seq data assumes that observed gene expression changes reflect alterations within stable cell populations. In Alzheimer's disease, this assumption fails catastrophically. When excitatory neurons die in the MTG—a process the dataset itself documents through reduced neuronal cell counts—the remaining neurons available for sequencing represent a survival-selected subset, not the original diseased population. This survival bias creates a form of immortalization bias in gene expression that is indistinguishable from true transcriptional dysregulation using standard differential expression methods.
The work by Wang and colleagues (PMID:32660529) elegantly demonstrates this problem, showing that bulk RNA-seq studies can be confounded by cellular composition changes that masquerade as intrinsic gene expression alterations. Their analysis revealed that cell-proportion changes in AD brains are "robustly replicable across independently assessed cohorts," yet these compositional shifts can be misattributed to transcriptional changes within cell types. The computational approaches they developed to identify cell-intrinsic differentially expressed genes (CI-DEGs) represent important progress, but they cannot fully resolve the fundamental inference problem: when cell death occurs, we lose access to the cells that were most diseased.
This problem is compounded by the statistical methods employed in the pseudo-bulk approach. When individual nuclei are aggregated within subjects before comparison, the resulting gene counts reflect both transcriptional regulation and cellular composition. Standard negative binomial models for differential expression assume that cell type proportions are constant across conditions—a violated assumption that can produce both false positives (genes appear upregulated due to enrichment of a cell type) and false negatives (true transcriptional changes are masked by opposing composition effects). Patrick and colleagues (PMID:32804935) addressed this by developing methods to deconvolve cell-type heterogeneity contributions to cortical gene expression, demonstrating that ignoring composition leads to systematic errors in attributing expression changes to specific cell types.
The reproducibility implications are severe. If the SEA-AD findings were to be validated in an independent cohort with different cellular composition profiles—say, a cohort with more advanced neuronal loss—the apparent "differentially expressed genes" could differ substantially not because the underlying biology is different, but simply because the relative survival of different neuronal subtypes varies between cohorts. This is not a peripheral concern; it strikes at the heart of whether we can trust the mechanistic interpretations that flow from these data.
I acknowledge several important caveats. The SEA-AD consortium has taken steps to address composition concerns by performing cell-type-specific analyses and examining proportions as covariates. The scale of the resource—with over 1.2 million nuclei profiled—provides unprecedented power to detect subtle effects and enables sophisticated modeling approaches. Furthermore, the field is moving toward more robust methods, including cell-wise regression approaches and spatial transcriptomics integration that can help distinguish true transcriptional change from compositional shift. Tang and colleagues (PMID:40329537) have demonstrated that integrating spatial transcriptomics with snRNA-seq data can enhance differential gene expression analysis for AD-related phenotypes, potentially offering a path forward.
However, these mitigations do not eliminate the fundamental limitation—they refine the tools used to study it. The core challenge remains: Alzheimer's disease is not merely a disorder of gene expression within cells; it is equally a disorder of which cells survive to express those genes. The SEA-AD resource, for all its value, cannot fully disentangle these two processes using differential expression analysis alone. Future studies should explicitly model cell death as a competing process and interpret differential expression findings as reflecting the combined effects of transcriptional regulation and selective survival.
**Confidence: 0.78**
This assessment reflects high confidence that the cell composition confounding is a genuine and serious limitation, tempered by recognition that the field is actively developing solutions and that the resource remains valuable despite these acknowledged constraints.