DNAcrypt-AI combines genome-coordinate sampling, large-scale DNA reconstruction, and machine learning to explore a new bioinformatics approach to digital information concealment.
Rather than storing a password or cryptographic key as a conventional character sequence, DNAcrypt-AI distributes its representation across randomly selected locations in the human reference genome. These coordinates are reconstructed into DNA sequences and interpreted through machine learning to recover the original information.
The framework is described in the preprint βDNAcrypt-AI: A Genome-Scale Bioinformatics Framework for Coordinate-Based Cryptography Using Artificial Intelligence,β authored by ChordexBio scientists Marvin De los Santos, Chickie Lynn, and Vaughn Jhing. The research was posted on 16 June 2026 and is available under DOI 10.21203/rs.3.rs-10020713/v1
Looking at the genome as a computational resource
The human genome contains billions of positions distributed across chromosomes, strands, and sequence contexts. These positions have traditionally been studied to understand inheritance, evolution, biological function, and disease.
DNAcrypt-AI examines a different possibility: Can the immense coordinate space of the human genome also serve as a structured substrate for information processing?
The framework does not use a personβs genome or private genetic information. Instead, it operates on established human reference genome assembliesβhg19 and hg38βwhich function as large, reproducible genomic maps. During each encoding run, DNAcrypt-AI randomly samples locations from these reference assemblies. A coordinate may include the genome assembly, chromosome, strand orientation, and start and end positions. Together, these locations form a unique genome vocabulary representing the concealed information.
The researchers compare the concept to a genomic positioning system. A conventional GPS identifies a location through geographic coordinates; DNAcrypt-AI represents information through combinations of genomic coordinates.
High-entropy passwords and encryption keys
The ChordexBio scientists evaluated DNAcrypt-AI using two output classes:
-Passwords containing uppercase letters, lowercase letters, numbers, and selected symbols
-Encryption keys containing uppercase letters, lowercase letters, and numbers
The system was tested using sequence lengths ranging from 6 to 90 characters. Across five independent generations, 16-character passwords produced calculated Shannon entropy values of approximately 60 to 64 bits. Sixty-four-character encryption keys produced values of approximately 331 to 345 bits. The chart on page 15 of the preprint also shows that the longer keys reached approximately 87% to 91% of their calculated maximum uncertainty under the studyβs evaluation method.
These results show that stochastic genome-coordinate selection can generate diverse outputs across independent encoding events.
Reproducible recovery across common string lengths
DNAcrypt-AI was evaluated through five independent recovery attempts at lengths of 6, 12, 24, 48, 60, and 90 characters. Every attempt produced an exact match for sequences containing up to 60 characters. This range encompasses many practical password and key-generation configurations evaluated in the study.
At 90 characters, three of five attempts achieved complete recovery. Two attempts produced partial mismatches concentrated near positions 70 to 90. The researchers attributed this instability to limitations in the principal-component reduction stage used by the current Covary implementation. Additional executions subsequently recovered the complete sequences.
The recovery results are visualized on page 16 of the preprint, where successful positions are shown across the tested configurations and the mismatches are concentrated within the longest sequence class. These findings establish proof of concept while also identifying a clear engineering target: improving long-range embedding stability for larger encoded sequences.
Multiple layers of encoding variability
DNAcrypt-AI derives variability from three distinct processes:
- Genome-coordinate selection. Each run can sample a different combination of assemblies, chromosomes, strands, and genomic intervals.
- K-mer-to-character assignment. A k-mer dictionary determines how machine-learned sequence representations correspond to letters, numbers, and symbols.
- DNA hashing. A separate feature-level representation controls sequence-length recovery and contributes another layer of obfuscation.
The study found that minor changes to selected, non-critical genome-coordinate entries did not necessarily alter the recovered password. This suggests a degree of tolerance to limited incidental metadata variation.
Changing more than half of the k-mer-to-character assignments, however, produced substantially different decoded sequences while preserving the expected character length. The custom-dictionary experiment shown on page 18 demonstrates that mapping dictionaries can act as an independent source of output variability.
Repurposing genomic AI beyond conventional analysis
Machine learning in genomics is commonly developed to classify sequences, predict biological functions, identify mutations, or model evolutionary relationships.
DNAcrypt-AI explores another direction.
Instead of asking only what biological information can be extracted from the human genome, the framework asks how genomic structures and machine-learned sequence relationships can be repurposed as functional components of a computational system. The human reference genome is therefore not treated merely as a database to search. Its coordinate breadth, sequence diversity, and relational structure become active elements of the encoding and recovery process.
Future versions could potentially incorporate additional reference assemblies, pangenome resources, or cross-species coordinate systems. Alternative dimensionality-reduction and learned projection methods may also improve recovery stability for longer sequences.
Expanding the role of biological complexity
DNAcrypt-AI introduces a broader scientific idea: biological complexity does not need to remain solely an object of analysis. It can also become part of the computational machinery.
By connecting genome coordinates, reconstructed nucleotide sequences, k-mer representations, and machine-learning inference, ChordexBio scientists have demonstrated a new bioinformatics-adjacent use for the human reference genome.
The project establishes an early foundation for exploring how genomic data structures might contribute to future information-processing systemsβwhile opening a new interdisciplinary space between molecular biology, artificial intelligence, and cybersecurity.
The study is currently available as a preprint and has not yet undergone peer review: https://doi.org/10.21203/rs.3.rs-10020713/v1
β Hi, Iβm Chordie! Iβm the content management guru at ChordexBio, responsible for creating news, blogs, and resource materials. Iβm passionate about turning ideas into clear, informative contentβand Iβm always happy to write π