A junk-filled genome
Norman Johnson · Evolution · 2023
“I would be quite proud to have served on the committee that designed the E. coli genome. There is, however, no way that I would admit to serving on a committee that designed the human genome. Not even a university committee could botch something that badly” David Penny, quoted in Graur et al. (2013). Bacterial genomes are sleek and no-nonsense. For example, Escherichia coli has 4,200 protein-coding genes, which contribute the vast majority (88%) of its 4 million base (megabase)-long genome (Blattner et al., 1997). Its genome also contains some well-characterized and well-conserved regulatory elements. There is not much repetitive DNA. I agree with David Penny; I would be proud to be a member of that committee. Even though I know these bacterial genomes were shaped by natural evolutionary processes, they look like the products of efficient and thrifty designers. They are “designoid” in the language of Dawkins (1996). By contrast, the human genome is anything but efficient and thrifty. Recall that eukaryotic genes have introns—the parts that are transcribed but spliced out before the final mRNAs are made—as well as the remaining exons. The DNA from the exons of our approximately 20,000 protein-coding genes total about 30 megabases, or roughly 1% of our total genome of 3.1 billion pairs. The introns—the parts that are thrown away—are much larger, totaling up to about 45% of genome. A smaller fraction of the genome consists of various classes of RNA-encoding genes, many of which appear to have regulatory roles. Another small fraction consists of numerous copies of short repeats, some at the telomeres and centromeres and some interspersed in the genome. But interestingly, somewhere between half and two thirds of the human genome comes from mobile elements or their remnants. This range is imprecise because ascertaining whether a stretch of DNA is derived from a mobile element can be challenging owing to the degradation of the sequence of transposable elements. If you are keeping score, you may have noticed that these figures total to somewhat more than 100%. That is because many of the mobile elements and their remnants lie in the introns. Thus, there may be some double counting. Regardless of the exact proportions of these categories, the human genome looks like a hot mess. I would not want to take credit for it either. Larry Moran at the University of Toronto puts his explanation of the composition of the human genome right in the title of his book: 90% of Your Genome Is Junk. By junk, Moran means parts that can be excised without any loss of functionality or reduction of fitness. So, only about a tenth of the human genome is needed; the rest is neutral or even slightly deleterious. That the bulk of the human genome is junk is not a novel proposition. In fact, for more than a half century—long before we had a human genome sequence—many leading evolutionary geneticists, biochemists, molecular biologists, and other researchers have expressed similar opinions. For instance, a 1972 paper by the Japanese geneticist Susumu Ohno has the title “So much junk in the genome” (Ohno, 1972). In another incidence, Orgel and Crick (1980, pp. 604–605) state that “…there is a large amount of evidence which suggests, but does not prove, that much DNA in higher organisms is little better than junk.” Putting aside equating of humans and similar organisms as “higher organisms,” most evolutionary biologists would not disagree with these statements. Controversy erupted in September 2012 when a consortium known as the Encyclopedia of DNA Elements (ENCODE) challenged this notion of a junk-filled human genome. ENCODE published a series of papers in high-profile journals supporting the conclusion that about 80% of the human genome is functional. In other words, the maximum amount of junk was 20% and could be lower still (ENCODE, 2012). ENCODE’s conclusion initially received widespread fanfare. For instance, writing in the journal Science, science writer Elizabeth Pennisi (2012, p. 1159) noted that the ENCODE findings “sound the death knell for the idea that our DNA is mostly littered with useless bases.” The ENCODE researchers did not restrict themselves to scientific journals. They also staged a multimedia campaign promoting their conclusions, including lavishly produced videos. In one of these videos, Ewan Birney, one of the leaders of ENCODE, states, “This metaphor about junk DNA has become, I think, very entrenched. It’s been entrenched publicly, entrenched scientifically. And ENCODE totally challenges that. We just don’t have big, blank, boring, bits of the genome; all the genome is alive at some level” (quoted in Moran, p. 245). Soon after, the pushback came. A series of papers, mostly from evolutionary geneticists, criticized ENCODE’s conclusion and presented several arguments in support of the junk-filled human genome (e.g., Doolittle, 2013; Graur et al., 2013; Palazzo & Gregory, 2014). These arguments have not been synthesized until now. Here, Moran provides a cogent synthesis of these arguments for the junk-filled genome hypothesis. Moran flatly states that he cannot prove this hypothesis but instead argues that it ought to be considered the null hypothesis. He then states, “What I can do is to show you that the concept of junk DNA is compatible with all the evidence, consistent with our understanding of evolution and population genetics, and possesses extraordinary explanatory power. It helps us make sense of biology” (p. 5). Part of the debate revolves around the meanings of words, but this is an important semantic debate. Both those who see the genome as mostly junk and those who see it as mainly functional agree that junk DNA is DNA that lacks function. Defining what functional DNA is has been more contentious. The ENCODE scientists implicitly defined functional DNA based on biochemical activity; to them, it includes DNA that is transcribed or has binding activity or has any other biochemical function. Moran argues that this expansive definition is flawed because many biochemically active DNA stretches may be meaningless to organismic function. After some philosophical discussion of function, Moran provides a one-sentence summary definition: “Functional DNA is any stretch of DNA whose deletion from the genome would reduce the fitness of the individual” (p. 98). As we will see later, Moran and other proponents of a junk-filled genome think that much of this junk may even be somewhat deleterious, but that selection may not be sufficiently powerful to get rid of it. In this case, the removal of excess DNA may even improve function and thus increase fitness. There are several ways to estimate the extent of junk in the genome. Ethics obviously prohibit experimentally deletions of DNA in humans. But these studies have been done in mice: megabases of DNA can be deleted without any obvious effect on survival or fecundity. Moreover, nature has done the experiment on humans. Because deletions occur naturally, one could (at least in theory) determine which parts of the genome can be deleted without large reductions in viability by scanning the genomes of lots of living people. The areas in which deletions are observed in some people are putative disposable regions. Such an endeavor requires a large database—like the United Kingdom’s Biobank (Halldorsson et al., 2022)—and much statistical analysis. Preliminary results of such investigations suggest much of the genome consists of disposable regions. Still, probably the best way to estimate the fraction of the genome that is junk is to look for the extent of sequence conservation. Regions of the genome that are functional are usually under strong purifying selection; thus, they evolve slower than disposable junk. Estimates of the fraction of the genome under such selective constraint generally fall under 10%. Because some parts of the genome are currently under selective constraint but have not been so for sufficiently long periods of time to show up as having low divergence, this figure is a lower bound. Still, 10% conserved (or 90% junk) is a reasonable estimate. Another line of inquiry that supports the junk-filled genome hypothesis comes from the immense variation in genome size, even among fairly closely related organisms. Moreover, the size of the genome does not correspond with complexity or number of protein-coding genes. For instance, the genome of an onion (Allium cepa) is about five times larger than the human genome. In addition, different species of onions within the genus Allium vary widely in genome content. Using these facts, the Guelph University genome researcher, Ryan Gregory, put forth a heuristic what has become known as “The Onion Test”: any claim that the genomes of humans and other large multicellular organisms are mostly functional needs to explain why onion genomes are so much larger than human genomes and why there is so much genome size variation among different onions (Palazzo & Gregory, 2014). While the onion test does not state that the bulk of noncoding DNA has to be junk, it provides a useful testing ground for functional hypotheses: can they explain onion genomes? Incidentally, onions are far from the most extreme genome giant: some plants and some salamanders have genomes in excess of 100 billion bases. In fact, Orgel and Crick (1980) brought up the large genomes of salamanders in their version of the onion test. Moreover, on the other side of the spectrum, a carnivorous plant, the humped bladderwort (Utricularia gibba) has a genome of only 82 megabases (2.5% of the size of the human genome and 0.5% of the size of the onion genome). Thus, bladderwort has 28,000 protein-coding genes—more than humans have—but only a small fraction (3%) of its genome comes from transposable elements. These huge differences in genome size present a major challenge for the idea that most of the human genome is functional. The junk-filled genome hypothesis posits that genome size variation among different species is a consequence of species varying in how rapidly they gain junk DNA, how rapidly they lose it, and how efficiently selection weeds out slightly deleterious DNA. This explanation is compatible with the neutral theory of molecular evolution and especially Tomoko Ohta’s extension, the nearly neutral theory (Ohta, 1973). This theory predicts that selection should be more efficient at purging slightly deleterious mutations when effective population sizes (N) are high. Selection efficacy correlates with N because the power of genetic drift to fix slightly deleterious variants decreases as N increases. Accordingly, if there is a very small cost to having individual chunks of junk DNA, we should expect species with low long-term N to accumulate junk DNA. Most large multicellular organisms have relatively low N compared with bacteria; hence, selection would not be expected to get rid of it. In other words, drift is a barrier to purging of junk DNA from large multicellular organisms (Lynch, 2007). Within large multicellular organisms, we would also expect a stronger drift barrier leading to more bloated genomes in lineages that have experienced long-term reductions in N. Ascertaining historical N has its challenges and only a few comparative studies have been done. Nevertheless, these studies generally support the predictions of the nearly neutral theory: lineages with low N generally have more bloated genomes. Note that ecology and physiology can also matter. Some organisms—such as carnivorous plants and hummingbirds—may face stronger selection to purge excess DNA. Consistent with the increased selection, these organisms typically have smaller, sleeker genomes. One factor affecting genome size (and variation in it) to which I wish Moran had devoted more attention is the cellular elimination of noncoding DNA. We have some evidence that species vary in how quickly cellular mechanisms dispose of excess noncoding DNA and that the elimination rate is negatively correlated with genome size. Drosophila flies eliminate DNA faster than Laupala crickets. Consistent with the faster elimination, Drosophila genomes are sleeker and smaller than Laupala genomes (Petrov et al., 2000). Montane grasshoppers have still more bloated genomes and a correspondingly slower rate of elimination than even Laupala crickets (Bensasson et al., 2001). More such studies need to be done. That most of the genome is junk does not invalidate the notion that some noncoding DNAs can have function. For instance, we know of many cases wherein transposable elements have been co-opted into functional roles. Interestingly, several of these are in mammal embryonic development (Cosby et al., 2019). Just as host-parasite systems can evolve into mutualistic relationships, we should expect some co-option of transposable elements for host function. Nevertheless, these co-opted transposable elements probably greatly outnumbered by remnant copies of transposable elements hanging around as junk. In addition to providing evidence and theory supporting the junk-filled genome hypothesis, Moran also directly criticizes the ENCODE results. He notes that much of the DNA stretches showing transcription in these experiments do so at low levels. Moran states that transcription is a noisy process. Low-level ectopic transcription is likely to be pervasive as transcription factors weakly bind to sequences that are one or two mutations away from a canonical binding site. These weak off-target binding and other processes generate low-abundance transcripts that lack actual function. In addition, the ENCODE researchers included cancer cell lines in their assays; hence, much of what they deemed functional may not be in a typical somatic cell. The subject matter in this book is not easy. Molecular biologists might well be challenged by the population genetics theory, while the biochemistry details may vex evolutionary biologists. Moran does an excellent job at presenting both of these aspects. I am also glad that he provided a historical perspective, showing that many of the current debates have a long history. In the Preface, Moran states that he was motivated to write this book in part due to what he views as failures in science communication regarding the nature of the genome. He reminds us about the importance of accuracy in science communication: “No matter how good your style, if the substance of what you are communicating is flawed, then you are not a good science communicator” (xiii). Narratives are useful in communicating science, but when they (or the hype) get in the way of telling the truth, the science, and the science communication suffer. The author declares no conflict of interest.