Back

Generative Models Validation via Manifold Recapitulation Analysis

Lazzaro, N.; Leonardi, G.; Marchesi, R.; Datres, M.; Saiani, A.; Tessadori, J.; Granados, A.; Henriksson, J.; Chierici, M.; Jurman, G.; Sales, G.; Tebaldi, T.

2024-10-26 bioinformatics Community evaluation
10.1101/2024.10.23.619602 bioRxiv
Show abstract

SummarySingle-cell transcriptomics increasingly relies on nonlinear models to harness the dimensionality and growing volume of data. However, most model validation focuses on local manifold fidelity (e.g., Mean Squared Error and other data likelihood metrics), with little attention to the global manifold topology these models should ideally be learning. To address this limitation, we have implemented a robust scoring pipeline aimed at validating a models ability to reproduce the entire reference manifold. The Python library Cytobench demonstrates this approach, along with Jupyter Notebooks and an example dataset to help users get started with the workflow. Manifold recapitulation analysis can be used to develop and assess models intended to learn the full network of cellular dynamics, as well as to validate their performance on external datasets. AvailabilityA Python library implementing the scoring pipeline has been made available via pip and can be inspected at GitHub alongside some Jupyter Notebooks demonstrating its application. Contactnlazzaro@fbk.eu or toma.tebaldi@unitn.it

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.