Back

Generalization of the minimum covariance determinant algorithm for categorical and mixed data types

Beaton, D.; Sunderland, K. M.; ADNI, ; Levine, B.; Mandzia, J.; Masellis, M.; Swartz, R. H.; Troyer, A. K.; ONDRI, ; Binns, M. A.; Abdi, H.; Strother, S. C.

2020-03-31 bioinformatics
10.1101/333005 bioRxiv
Show abstract

The minimum covariance determinant (MCD) algorithm is one of the most common techniques to detect anomalous or outlying observations. The MCD algorithm depends on two features of multivariate data: the determinant of a matrix (i.e., geometric mean of the eigenvalues) and Mahalanobis distances (MD). While the MCD algorithm is commonly used, and has many extensions, the MCD is limited to analyses of quantitative data and more specifically data assumed to be continuous. One reason why the MCD does not extend to other data types such as categorical or ordinal data is because there is not a well-defined MD for data types other than continuous data. To address the lack of MCD-like techniques for categorical or mixed data we present a generalization of the MCD. To do so, we rely on a multivariate technique called correspondence analysis (CA). Through CA we can define MD via singular vectors and also compute the determinant from CAs eigenvalues. Here we define and illustrate a generalized MCD on categorical data and then show how our generalized MCD extends beyond categorical data to accommodate mixed data types (e.g., categorical, ordinal, and continuous). We illustrate this generalized MCD on data from two large scale projects: the Ontario Neurodegenerative Disease Research Initiative (ONDRI) and the Alzheimers Disease Neuroimaging Initiative (ADNI), with genetics (categorical), clinical instruments and surveys (categorical or ordinal), and neuroimaging (continuous) data. We also make R code and toy data available in order to illustrate our generalized MCD.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Biostatistics
24 papers in training set
Top 0.1%
18.6%
2
Bioinformatics
1204 papers in training set
Top 2%
15.1%
3
BMC Bioinformatics
457 papers in training set
Top 1%
6.8%
4
PLOS Computational Biology
1863 papers in training set
Top 6%
5.5%
5
PLOS ONE
5266 papers in training set
Top 30%
5.2%
50% of probability mass above
6
Biometrics
23 papers in training set
Top 0.1%
4.3%
7
The Annals of Applied Statistics
19 papers in training set
Top 0.1%
4.1%
8
Bioinformatics Advances
203 papers in training set
Top 1%
4.1%
9
Statistics in Medicine
40 papers in training set
Top 0.2%
3.4%
10
Journal of Computational Biology
48 papers in training set
Top 0.6%
1.7%
11
Scientific Reports
3612 papers in training set
Top 53%
1.7%
12
PeerJ
308 papers in training set
Top 5%
1.7%
13
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.3%
1.7%
14
Evolutionary Biology
14 papers in training set
Top 0.1%
1.3%
15
Patterns
78 papers in training set
Top 2%
1.3%
16
BioData Mining
22 papers in training set
Top 0.5%
1.1%
17
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
18
NeuroImage
903 papers in training set
Top 5%
1.1%
19
Briefings in Bioinformatics
354 papers in training set
Top 7%
0.8%
20
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
21
Genetic Epidemiology
55 papers in training set
Top 0.8%
0.6%
22
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
0.6%