Back

A Roadmap to Artificial Intelligence (AI): Methods for Designing and Building AI ready Data for Womens Health Studies

Kidwai-Khan, F.; Wang, R.; Skanderson, M.; Brandt, C.; Jarad, S.; Womack, J.

2023-05-30 health informatics
10.1101/2023.05.25.23290399 medRxiv
Show abstract

ObjectivesEvaluating methods for building data frameworks for application of AI in large scale datasets for womens health studies. MethodsWe created methods for transforming raw data to a data framework for applying machine learning (ML) and natural language processing (NLP) techniques for predicting falls and fractures. ResultsPrediction of falls was higher in women compared to men. Information extracted from radiology reports was converted to a matrix for applying machine learning. For fractures, by applying specialized algorithms, we extracted snippets from dual x-ray absorptiometry (DXA) scans for meaningful terms usable for predicting fracture risk. DiscussionLife cycle of data from raw to analytic form includes data governance, cleaning, management, and analysis. For applying AI, data must be prepared optimally to reduce algorithmic bias. ConclusionAlgorithmic bias is harmful for research using AI methods. Building AI ready data frameworks that improve efficiency can be especially valuable for womens health. Lay SummaryWomens health studies are rare in large cohorts of women. The department of Veterans affairs (VA) has data for a large number of women in care. Prediction of falls and fractures are important areas of study related to womens health. Artificial Intelligence (AI) methods have been developed at the VA for predicting falls and fractures. In this paper we discuss data preparation for applying these AI methods. We discuss how data preparation can affect bias and reproducibility in AI outcomes.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.