| Line 1: |
Line 1: |
| − | '''Pseudogenes''' are genes present in an organism's [[genome]] that have lost the ability to code for proteins due to mutation. Pseudogenes are often difficult to parse from the large amount of non-coding base pairs in the genome. Convention requires to two elements to be present to label a sequence a pseudogene. The first is [[homology]] which is the requirement that a sequence be demonstrated to descend from a functional copy of the gene and the second is non-functionality which is the requirement that the gene not code for a protein in the organism in question. | + | '''Pseudogenes''' are genes present in an organism's [[genome]] that have lost the ability to code for proteins due to mutation. <ref name=petrov>Petrov, D.A, Hartl, D.L. (2000). Pseudogene evolution and natural selection for a compact genome. The American Genetic Association 91:221-227. [http://www.stanford.edu/group/petrov/research/16.pdf]</ref> They were first identified and dubbed in the late 1970s when researchers began finding non-coding regions in some organisms that were similar to actual coding genes in other organisms. <ref name=sciam>Gerstein, M, Zheng, D. (2006). The real life of pseudogenes. Scientific American 95:48-55. [http://papers.gersteinlab.org/e-print/sciam2/preprint.pdf]</ref> So far an estimated 19,000 pseudogenes have been identified in the human genome, this is almost equal to the total number of coding genes (21,000). <ref name=sciam /> Pseudogenes have been identified in a wide range of organisms from bacteria to mice to humans, the total number of pseudogenes in a given genome is not predictable but specific pseudogenes are often compared across species to elucidate complex evolutionary relationships <ref name=sciam /> Humans have many pseudogenes including [[L-gulonolactone oxidase]] which is used to synthesize vitamin c. This gene was inactivated in the common ancestor of all [[simians]]. |
| | + | |
| | + | ==Finding pseudogenes== |
| | + | |
| | + | Pseudogenes are often difficult to parse from the large amount of non-coding base pairs in the genome. Convention requires two elements to be present to label a sequence a pseudogene. The first is [[homology]] which is the requirement that a sequence be demonstrated to descend from a functional copy of the gene and the second is non-functionality which is the requirement that the gene not code for a protein in the organism in question. <ref name=sciam /> |
| | + | |
| | + | Since all pseudogenes are descended from a functioning gene the first step is to find the parent gene that it descended from. This is done by using computer programs to compare sequences of DNA across species. <ref name=sciam /> This is a large computational problem but by keeping in mind the phylogenetic relationships between species the search time can be decreased by looking at species that share a more recent common ancestor.<ref name=mito>Bensasson, D., Zhang, D., Hartl, D., Hewitt, G. (2001). Mitochondrial pseudogense: evolution's misplaced witness. Trends in Ecology and Evolution 16: 314-321. [http://www.ioz.ac.cn/department/agripest/group/zhangdx/ZDX-s-pdf%5CTREE2001-Updated_numt.PDF]</ref> Once a functioning copy of a gene is detected its sequence is compared to the pseudogene. A high correlation in base pairs is used to assign homology. Non-functionality can be demonstrated by attempting to transcribe the sequence in-vitro. <ref name=sciam /> |
| | + | |
| | + | ==Pseudogenes and neutral selection theory== |
| | + | |
| | + | Because pseudogenes do not code for a function many scientist have hypothesized that the accumulation of mutations would not be constrained by selection pressures.<ref name=petrov /> This is known as neutral selection, and pseudogenes have been studied extensively to test various theories of neutral selection. <ref name=bustamante>Bustamante, C, Neilsen R, Hartl, D. (2002). A maximum likelihood method for analyzing pseudogene evolution: implications for silent site evolution in humans and rodents. 19:110-117. [http://www.mbe.oupjournals.org/cgi/content/abstract/19/1/110]</ref> It has been determined that mutations fixate in pseudogene regions at about 30 percent higher than in coding regions of DNA. <ref name=bustamante /> Some theorist have argued that there maybe some selection pressure on pseudogenes (such as on genome size in general) so conclusions should be tempered.<ref name=petrov /> Others have determined that base pair mutations are not completely random, favoring accumulation of [[guanin]] and [[cytosine]]. <ref name=bustamante /> Despite these findings research on pseudogenes still continues to be a productive avenue for exploring mutation and selection. |
| | | | |
| − | Since all pseudogenes are descended from a functioning gene the first step is to find a species that has a functioning copy of that gene. This is done by looking at [[phylogenetic tree]]s and testing organisms with a relatively recent [[common descent | common ancestor]] and working backwards until a copy is found. Once a functioning copy of a gene is detected its sequence is compared to the pseudogene. A high correlation in base pairs is used to assign homology. Non-functionality can be demonstrated by attempting to transcribe the sequence in-vitro. Humans have many pseudogenes including [[L-gulonolactone oxidase]] which is used to synthesize vitamin c. This gene was inactivated in the common ancestor of all [[simians]].
| |
| | | | |
| | ==See also== | | ==See also== |
| Line 9: |
Line 18: |
| | | | |
| | ==References== | | ==References== |
| − | | + | <references/> |
| − | *Vanin, E. F. (1985). "Processed pseudogenes: characteristics and evolution." Annu Rev Genet 19: 253-72.
| |