An end-to-end workflow for statistical analysis and inference of large-scale biomedical datasets
Heidari, E.; Sharifi-Zarchi, A.; Sadeghi, M. A.; Mirzaei, M.; Ahmadi, N.; Balazadeh-Meresht, V.; Sadr, M.
Show abstract
Throughout time, as medical and epidemiological studies have grown larger in scale, the challenges associated with extracting useful and relevant information from these data has mounted. General health surveys provide a good example for such studies as they usually cover large populations and are conducted throughout long periods in multiple locations. The challenges associated with interpreting the results of such studies include: the presence of both categorical and continuous variables and the need to compare them within a single statistical framework; the presence of variations in data resulting from the technical limitations in data collection; the danger of selection and information biases in hypothesis-directed study design and implementation; and the complete inadequacy of p values in identifying significant relationships. As a solution to these challenges, we propose an end-to-end analysis workflow using the MUltivariate analysis and VISualization (MUVIS) package within R statistical software. MUVIS consists of a comprehensive set of statistical tools that follow the basic tenet of unbiased exploration of associations within a dataset. We validate its performance by applying MUVIS to data from the Yazd Health Study (YaHS). YaHS is a prospective cohort study consisting of a general health survey of more than 30 health-related measurements and a questionnaire with over 300 questions acquired from 10050 participants. Given the nature of the YaHS dataset, most of the identified associations are corroborated by a large body of medical literature. Nevertheless, some more interesting and less investigated connections were also found which are presented here. We conclude that MUVIS provides a robust statistical framework for extraction of useful and relevant information from medical datasets and their visualization in easily comprehensible ways.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 92%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 91%
- COVID19-Global: A shiny application to perform a global comparative data visualization for the SARS-CoV-2 epidemic 91%
Similar papers in this journal
- %svy_logistic_regression: A generic SAS(R) macro for simple and multiple logistic regression and creating quality publication-ready tables using survey or non-survey data 93%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 93%
- BioDiscViz : a visualization support and consensus signature selector for BioDiscML results 92%
Similar papers in this journal
- Causal Analysis for Multivariate Integrated Clinical and Environmental Exposures Data 91%
- Combining symbolic regression with the Cox proportional hazards model improves prediction of heart failure deaths 91%
- An Ontology-based Approach to Guide and Document Variable and Data Source Selection and Data Integration Process to Support Integrative Data Analysis in Cancer Outcomes Research 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.