A new study at ChordexBio demonstrates how Covary can analyze hundreds of complete filovirus genomes in minutes, without conventional multiple sequence alignment.

The study, titled β€œRapid, Large-Scale and Multi-Species Phylogenomic Analysis of Orthoebolaviruses Using Covary,” evaluated whether biologically meaningful relationships among closely related filoviruses could be recovered using an alignment-free and translation-aware computational approach. The preprint was authored by ChordexBio scientists Bea Nicole S. Gacayan and Marvin De los Santos and was posted on 9 July 2026.

The work highlights Covary as an emerging genomic intelligence platform for rapidly comparing complete genomes using machine-learned sequence representations.

Moving beyond conventional genome alignment

Genomic surveillance allows scientists to monitor viral variation, investigate transmission patterns, and detect divergent or emerging lineages. However, the expanding number of available genome sequences creates a growing computational challenge.

Traditional phylogenomic workflows commonly begin by aligning sequences against one another. As datasets increase in size or contain genetically diverse organisms, multiple sequence alignment can become computationally demanding and may require substantial preprocessing.

Covary takes a different approach

Rather than aligning every nucleotide position, Covary converts genetic sequences into translation-aware numerical embeddings. These embeddings represent sequence characteristics in a form that machine-learning and statistical methods can compare directly.

Covary is powered by TIPs-VF and is designed to support alignment-free comparison, distance analysis, clustering, and visualization of biological sequences at scale. Its workflows can produce distance matrices and projections through PCA, t-SNE, and UMAP, together with hierarchical clustering and related exploratory outputs.

This allows scientists to examine relationships across entire genomes without first introducing alignment gaps or selecting a small set of marker genes.

Hundreds of filovirus genomes analyzed

The ChordexBio scientists assembled 750 complete genomes representing six recognized Orthoebolavirus species and Orthomarburgvirus marburgense, which was included as an external filovirus comparison.

After Covary automatically excluded sequences containing ambiguous nucleotide characters, 728 complete genomes proceeded to comparative analysis.The dataset included genomes from:, Zaire, Sudan, Reston, Bundibugyo, Bombali, TaΓ― Forest and Marburg virus

Each genome contains approximately 19,000 nucleotides. Collectively, the dataset represented millions of nucleotide positions that needed to be transformed, compared, and organized. The complete Covary workflow was performed in less than 10 minutes, according to the preprint.

Covary recovered recognizable species-level patterns

Most of the analyzed filoviruses formed distinct and taxonomically consistent groups. Reston, Sudan, Bundibugyo, and TaΓ― Forest ebolaviruses formed coherent species-level clusters. Marburg virus was also clearly separated from the ebolaviruses, while retaining internal substructure that may correspond to variation within the species.

One of the study’s most notable findings involved Orthoebolavirus zairense, or Zaire ebolavirus. Unlike most of the other species, Zaire ebolavirus genomes did not form one compact cluster. They were distributed across several areas of the embedding space and appeared in multiple subclusters within the dendrogram.

The finding therefore represents a basis for further evolutionary and epidemiological investigation, not a proposed revision of filovirus taxonomy.

A computational tool for genomic surveillance

The study demonstrates how machine-learned whole-genome representations could complement conventional phylogenetics. Because Covary does not require predefined marker genes, it may be used as an early exploratory layer for rapidly screening large sequence collections. Genomes with unusual distances, unexpected cluster placement, or divergent embedding patterns can be identified for more focused phylogenetic, epidemiological, or laboratory investigation.

This is particularly relevant to pathogen surveillance, where researchers may need to examine expanding genome datasets while outbreaks are still developing. Covary is designed for large-scale biological sequence analysis and supports applications in phylogenomics, taxonomic studies, pathogen screening, and genomic exploration. Its public web platform also enables researchers to perform and share whole-genome analyses through cloud-based workflows.

Building accessible genomic intelligence

For ChordexBio, the study represents more than an analysis of Ebola virus diversity. It demonstrates how independently developed machine-learning infrastructure can support internationally relevant questions in virology and genome science.

Covary was developed to make large-scale genomic analysis more accessible, particularly for researchers who may not have access to extensive local computing infrastructure. The framework combines sequence encoding, machine learning, distance-based analysis, and visual exploration within a reusable workflow.

By applying Covary to a complex, multi-species filovirus dataset, ChordexBio scientists showed that rapid computation does not have to come at the expense of biologically interpretable results.

The researchers nevertheless describe Covary as a complementary framework, rather than a replacement for conventional phylogenetic reconstruction. Embedding-derived relationships should be interpreted alongside established evolutionary methods, epidemiological records, ecological evidence, and balanced genomic sampling.

The study is currently available as a preprint and has not yet undergone peer review: https://doi.org/10.20944/preprints202607.0617.v1

C
Chordie

βœ‹ Hi, I’m Chordie! I’m the content management guru at ChordexBio, responsible for creating news, blogs, and resource materials. I’m passionate about turning ideas into clear, informative contentβ€”and I’m always happy to write πŸ™‚