Back

Reliability-weighted target prioritization in CD4+ T-cell Perturb-seq: a generalizability-theory decomposition

Cheng, C.

2026-07-15 bioinformatics
10.64898/2026.07.13.738312 bioRxiv
Show abstract

Genome-scale Perturb-seq screens prioritize candidate targets by the strength of a perturbations transcriptional effect. Effect strength does not answer a prior measurement question: is the readout dependable? A large effect estimated from a single guide, a single donor, or a pseudobulk of few cells need not survive replication, and for target prioritization each false lead costs a validation experiment. We treat each perturbation effect as a measurement in a crossed Target x Guide x Donor x Condition design and apply generalizability theory (Brennan, 2001; Cronbach et al., 1972) to separate the dependable part of an effect from facet-specific idiosyncrasy. Guides and donors enter as random facets; condition enters as a fixed facet and is analyzed within its levels. For each target we report a dependability profile over the facets and a joint generalizability coefficient over the two random facets, and we re-rank targets by effect magnitude weighted by that coefficient. On the released screen (Zhu et al., 2025), removing the measurement-error floor estimated from the non-targeting controls raises the number of genes with a dependable target-signal share above .10 from 40 to 7,674. Analyzed within activation states, dependability recovers the T-cell-receptor signaling module as reliably measurable only in activated cells, without recourse to gene annotation. A design study indicates that reliability is limited by the number of guides rather than the number of donors, so a future screen should add guides. Every methodological decision was recorded and adversarially reviewed, and all results regenerate from the released summary statistics.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Cell Systems
201 papers in training set
Top 0.1%
21.9%
2
eLife
5828 papers in training set
Top 7%
12.5%
3
Nature Communications
5641 papers in training set
Top 18%
9.8%
4
Bioinformatics
1204 papers in training set
Top 3%
7.9%
50% of probability mass above
5
PLOS Computational Biology
1863 papers in training set
Top 6%
6.3%
6
Cell Reports Methods
165 papers in training set
Top 0.6%
3.4%
7
Nature Methods
385 papers in training set
Top 3%
3.2%
8
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 22%
2.4%
9
Nature Genetics
286 papers in training set
Top 3%
2.1%
10
Genome Biology
637 papers in training set
Top 5%
2.1%
11
Nature Biotechnology
172 papers in training set
Top 3%
1.7%
12
Scientific Reports
3612 papers in training set
Top 56%
1.7%
13
PLOS Biology
486 papers in training set
Top 5%
1.7%
14
Cell Genomics
172 papers in training set
Top 2%
1.5%
15
Molecular Systems Biology
162 papers in training set
Top 2%
1.1%
16
Molecular Biology of the Cell
311 papers in training set
Top 3%
1.1%
17
Life Science Alliance
285 papers in training set
Top 5%
1.1%
18
Science Advances
1243 papers in training set
Top 25%
1.1%
19
Cell Reports
1498 papers in training set
Top 25%
1.0%
20
iScience
1154 papers in training set
Top 34%
0.8%
21
The American Journal of Human Genetics
234 papers in training set
Top 3%
0.8%
22
Bioinformatics Advances
203 papers in training set
Top 4%
0.8%
23
Patterns
78 papers in training set
Top 3%
0.8%
24
Science
477 papers in training set
Top 10%
0.6%