This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: More than two decades after scientists first sequenced the entire human genome—all 3 billion base pairs, of DNA code—the meaning of much of this code remains a mystery. While an estimated 1% to 2% of human DNA codes for proteins, the rest is a mix of "junk DNA"—evolutionary holdovers that no longer code for anything—and regulatory elements that control when, where and how strongly genes are expressed.
These noncoding regions of the genome could hold the key to understanding a variety of inherited traits, including those that lead to diseases such as cancer, heart disease and autism. But first, scientists have to understand how variants in this DNA contribute to the multitude of traits that make each of us unique. Researchers at UC Berkeley have created a new genomic language AI model, called GPN-Star, that far outpaces its competitors at identifying the most important genetic variants that contribute to inherited traits, including those that lead to disease.
It is also far more computationally efficient than larger models, requiring only a fraction of the time and computing resources to train. "Our model excels in making predictions about the pathogenicity of genetic variants and identifying functional versus nonfunctional elements in the genome," said study senior author Yun Song, a professor of computer science and statistics at Berkeley and an investigator at the Innovative Genomics Institute. Along with the study, the researchers have published genome-wide predictions from their model, which highlight genetic variants that are likely to have the most influence on inherited traits.
Biologists can use these annotations to identify relevant genes and regulatory elements for further study. "We hope our work will help drive biological discovery," Song said. "People have developed really creative tools for assaying the impact of genetic variants, but they cannot experimentally test every single variant in the genome.
We believe our predictions will help prioritize the experiments that could have the greatest impact on human health." Song is also director of the Berkeley Center for Computational Biology and co-director of the UC Berkeley–UCSF Bakar Computational Biomedicine Initiative. The study was published in the journal Nature. Genomic language models work a little like chatbots for DNA, but instead of being trained on natural language, they are trained on vast troves of DNA sequences.
These models' advanced pattern-recognition skills can identify repeating patterns and sequences much faster than any human, allowing them to identify important elements of a genome that might otherwise be impossible to recognize. "Mathematically, a DNA sequence is just a string of letters—A, C, G and T. We don't know a priori which parts of the genome are functional elements, and a very small percentage of the genome is functional," Song said.
Extract — continue reading at the source.