Back

Finger Type Classification for Fingerprint Image Error Correction in Large Scale Biometric Databases

Sakif, T. I.; Dawson, J.; Nasrabadi, N.

2025-12-16 bioinformatics
10.64898/2025.12.14.694185 bioRxiv
Show abstract

Large-scale biometric systems, essential for national security and border management, increasingly rely on multimodal databases containing millions of identities. However, operational pressures and insufficient training lead to frequent image classification and labeling errors by human operators. These critical data integrity issues include the mislabeling of rolled vs. flat fingerprints, out-of-sequence captures, and the insertion of incorrect modalities. Such errors render enrollment records unreliable, compromising subsequent identity verification processes. Since manually sorting vast image archives is unfeasible, our study proposes an automated solution. The primary objective was to deploy a Siamese Network to classify fingerprints by their precise finger type and collection methodology (flat or rolled impressions). A secondary, but central, goal was to investigate the influence of varying embedding dimensions (64, 128, 256, 512) and similarity thresholds (0.5, 0.2, 0.1) on the networks performance metrics. Our most significant finding demonstrates a clear trade-off: a lower similarity threshold drastically increases conditional accuracy and precision (e.g., up to 98%) but simultaneously increases the proportion of images categorized as "uncertain" (up to 24%). In a practical, large-scale application, this necessitates balancing superior classification accuracy against a higher volume of images requiring costly manual inspection. This work provides a proof-of-concept tool capable of efficiently quantifying the percentage of images requiring human review across various modalities (fingerprints, face, iris). The eventual goal is a lightweight, efficient tool to establish standard preprocessing procedures for any large biometric dataset, dramatically reducing the time and cost associated with data integrity maintenance.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.