AI Model Identified Hidden Genetic Elements in Bacteria
The research-stage framework reveals that most non-coding intergenic base-pairing remains currently unannotated.
Updated on Sept. 23, 2026 in Life Sciences

Live Poll
Should artificial intelligence be used more extensively to identify unknown elements in microbial genomes?
Researchers have developed Minerva, a language model capable of identifying non-coding genomic elements in prokaryotes. The tool found that 84.3% of predicted intergenic base-pairing in the 150 bacterial genomes analyzed currently lacks annotation.
Why it matters
Current genomic annotation is primarily protein-centric and relies on homology-based comparisons, often missing non-coding elements. This tool offers a way to map these hidden structural components, which are essential for understanding prokaryotic gene regulation.
Minerva identified that 84.3% of intergenic base-pairing across 150 bacterial genomes is currently unannotated. The research-stage model utilizes genome language models to generate two-dimensional maps of local interactions, applying categorical Jacobian fingerprinting to predict base-pairing.
The players
Minerva
A research-stage genome language model designed to map non-coding elements and ncRNA base-pairing in prokaryotic genomes.
Pseudomonas
A genus of bacteria used as a model organism to demonstrate the tool's capability in identifying secondary-structure extensions.
The details
Minerva operates by treating genomic sequences as language and predicting local molecular interactions through two-dimensional mapping. The framework employs interaction heads to isolate specific base-pairing motifs and categorical Jacobian fingerprinting—a mathematical method to track how specific inputs influence complex system outputs—to identify non-coding RNA (ncRNA) structures. Beyond basic mapping, the tool successfully detects open-reading-frame signatures at the DNA level and characterizes secondary-structure extensions in the TwoAYGGAY ncRNA family within Pseudomonas. It also identified that Unknown Group 27 (UG27) reverse transcriptase systems encode arrays of structurally conserved ncRNAs that template complementary DNA hairpin products.
Timeline
September 23, 2026: The research article was published on biorxiv.org.
The Tech Race
This development follows the precedent set by the human genome's dark matter annotation initiatives. It marks a departure from traditional homology-driven search methods by leveraging language models to map the previously invisible non-coding architecture of prokaryotic life.
This research-stage model provides a new tool for bioinformaticians and researchers to refine bacterial genome databases. It does not currently impact clinical workflows or diagnostic products, as the findings remain confined to computational research on prokaryotic structures.
The takeaway
Minerva proves that language models can effectively map the vast, unannotated structural components of bacterial DNA that traditional methods miss. Researchers should watch for subsequent validation studies that experimentally confirm the functional role of these newly identified ncRNA hairpins.
Further reading
For more context on how machine learning is changing genomic analysis, visit Life Sciences.
More information
Access the complete scientific research article for technical validation.
Source note: This article includes information reported by Biorxiv.
Live Poll
Should artificial intelligence be used more extensively to identify unknown elements in microbial genomes?






