Back

TIPs-VF: An augmented vector-based representation for variable-length DNA fragments with sequence, length, and positional awareness

De los Santos, M.

2025-02-17 bioinformatics
10.1101/2025.02.15.637782 bioRxiv
Show abstract

The ability to accurately encode and represent genetic sequences in machine learning process is critical for advancements in biotechnology, specifically in genetic engineering and synthetic biology. Traditional sequence encoding method face significant limitations in handling sequence variability, maintaining reading frame integrity, and preserving biologically relevant features. This preliminary study presents TIPs-VF (Translator-Interpreter Pre-seeding for Variable-length Fragments), a simple and efficient encoding framework designed to address some of the key challenges in representing genetic sequences for machine learning. The results showed that TIPs- VF enables a variable-length sequence representation that retains biological context while ensuring the alignment of encodings with codon boundary, making it particularly suited for modular genetic construction. TIPs-VF demonstrated superior performance in truncation and fragmentation analysis, sequence homology detection, domain assessment, and splice junction identification. Unlike conventional methods that require fixed-length inputs, TIPs-VF dynamically adapts to sequence length variations, preserving essential features such as domain similarities and sequence motifs. Additionally, TIPs-VF improves open reading frame recognition and enhances the identification of vector parts and plasmid elements by unifying sequence embeddings with the three possible open reading frame. Overall, TIPs-VF offers a robust, biologically meaningful encoding framework that overcomes the constraints of traditional sequence representations by incorporating sequence, length, and positional awareness. The TIPs-VF encoding infrastructure is available at https://tips.logiacommunications.com.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.