Genome annotation is the process of assigning biological roles to genomic sequences through systematic identification of genetic elements. This lecture covers structural-based annotation approaches that use rule-based methods to identify genomic features such as genes, exons, introns, and repetitive elements. The process involves three main steps: identifying non-coding regions (particularly repetitive elements), predicting functional elements like genes, and adding biological information to these elements. Structural annotations rely on intrinsic genomic features (sequence composition, codon usage, splice site patterns) and extrinsic evidence (homologous sequences, RNA-seq data) to predict gene models. Key challenges include high error rates in eukaryotic gene prediction (often exceeding 10%), the complexity of alternative splicing, and the need to distinguish true genes from spurious open reading frames. Tools like MAKER integrate multiple approaches by combining ab initio predictions with evidence-based methods to produce more accurate gene models. The lecture emphasizes that computational predictions should be treated as hypotheses requiring experimental validation, particularly for eukaryotic genomes where complete gene model prediction remains challenging despite advances in algorithmic approaches.
Genome Annotation in Bioinformatics: Gene Prediction Explained
Added:hi today for lecture 15 of the Bagram mattis class I'll be talking about annotation black shouldn't we talk about annotation for the next two classes but for this one we're specifically going to focus in things are learning the more structural based annotation of those routine predictions and other genetic element prediction so hopefully some of the learning outcomes you'll get from this is you on better understanding of the components of what a DNA genome annotation is what are some of the main processes by which this is accomplished how do we actually annotate different types of elements like repeats and genes and also how a little bit more about gene prediction programs and in class we'll spend some time actually running these programs so hopefully they'll be good understanding how they work both from a theoretical standpoint from a practical standpoint so annotation if we define it is just a note added by way of explanation or commentary so what does that mean when we talk about genome annotation that means that we're assigning a role or a note to a string of nucleotides right so we're adding in some piece of extra evidence and typically we think that we have a good handle on why this is we're not always showing direct evidence for this sometimes we're using multiple evidence lines of evidence to identify what exactly a specific region or part of a genome is possibly involved with or what its identity might be and you can think about this is basically metadata right this is data about data to do this process we often undertake something known as especially when we think about genes as gene prediction will also do repeat prediction but in this case what we're trying to do is we're trying to infer what the most likely gene model we can identify in this genome and you can use this with various levels of evidence so you can do with no evidence additional evidence you can do it with using a whole bunch of external data to help you do your prediction do you know annotation generally consists of three steps this is identifying the portions of the genome that do not code for proteins so repetitive elements and the non-coding regions of genome identifying the functional elements of the genome this is the gene prediction and then adding in some sort of biological information to these elements so adding in some type of additional annotations additional information to maybe identify what gene it is what function of my have that so I think about annotation there are sort of two types or structural based and functional annotations we're gonna mostly focus on the structural annotations here and these are these are rule-based approaches to doing genome annotation what you basically are doing you're looking for regions of interest in the genome and you can find them by kind of how they look so what are some of the characteristics of these regions things like how many exons introns or the isoforms well you know translated regions functional annotations a little bit different we're looking more at trying to identify what regions actually do so what are they code for and in this case you know we're often using homology based approaches to identify what some sequences might be similar to then through inference using things like the gene ontology databases try to figure out what the actual function of these things might be all will also use these sorts of things to do things like identify orthologs and pair logs but the whole goal of this whole process is to describe structures like we see here right when we have rare but identify where the exons the introns are where the open reading frames or the untranslated regions where the introns right where are the parts of the genome that can actually give us more information about the actual things that a genome is doing so this lecture we're going to mostly focus on structure based annotations next lecture we'll talk more about functional impatiens I'm not really going to talking about 3d 3d protein structures because I'm not a protein biologist and that's not a point of emphasis for me but just understand that there are structural based annotation approaches that you can also use for that so basically we're trying to do is go from this case where we just have DNA sequence to something that looks a little bit more like this it's a little deformative right well we can think oh I have a region of the genome and I can find where I can identify genes that are active so regions of genomes where we expect genes to be located but also understanding that there's regions of the genome where there is not there are not genes we can do things like characterize these regions based upon things like do you see content or we can identify where we see evidence of you know open reading frames as well to help us understand well we might predict their input genes located for a practical standpoint genome annotations are often presented in an interval file format so this is an example of one a commonly used one called a GFF file format or gene feature file format and in this case you've got about nine columns each of these columns codes for different things you know which sequence you're talking about where did the source of information come from what what's the type of annotation you're talking about it's an exon Natron mrna etc it's important here we call an interval file because it's identifying the interval on the sequence or the chromosome where we expect to find this feature sometimes there's a score so you can Frank you know how good of a identification you've been able find and then there is some information on strand well they're not you find with the positive or the negative strand and then usually in this last column there's some other additional information that can help you better understand what the things are actually looking at here so here you could see like you know this tells you oh yeah these are all the exons that are associated with this this mRNA in this gene so this is just a closer look at what some what these what one of these systems might look like there are a bunch of other types of file formats by the way that you can that are used that are very similar to GF f3 there are bed files of GTF file listing version 2 of the GF f files bigwigs other types of interval type structured files that can provide information about kinds of entities you might expect to find on a chromosome why is this is important well say we're doing something like differential gene expression right got two different conditions we're gonna map read so the genome we want to see whether or not if there's differences in gene expression between these two different conditions right knowing where the genes are is really important for us to know whether or not we should be looking at those reads or we shouldn't be looking at those reads especially if we do if we had said like some other genes and the region which is not uncommon so if we have two genes where we have overlapping reading frames their exons overlap here we can see you know there are a few reads where you could assign them to either one of these genes based upon this annotation it looks like it's pretty obvious that these reads are all associated with this gene and that this gene so it's important that we're able identify these types of components and how we might use them we're actually interpreting data one of the challenges about this process though in finding genes is its it's pretty complicated so if you just take any sort of length of sequence what you're gonna find is that typically there's a lot of open reading frames so when I say open the reading frame that means that I'm finding a segment of DNA in a larger sequence that code that that could be translated into protein sequence so you have a series of codons all in a row with no stop codons in between and in the middle of it and it's actually gonna effectively code for a protein sequence and so it's very easy to find this kind of protein sequence in any sequence this is looking at the six frames of one sequence so this is the starting at the first base starting at the second base starting near the third base of the positive strand and this is starting at the first the last base the second last base and the third the last base of the negative strand so looking it this way and so this will show you you know where you see evidence of protein coding sequence in these sequences he knows there's a bunch of them so how do we figure out which ones to be are the good ones and which ones are the bad ones it's difficult you know we can find open reading frames pretty easy but when you think about like how much complexity there actually is and how good we are at actually predicting this I was pretty challenging so here's some cases where we have like you know two genes here and we have the predicted and Minar mRNA sequences and you can see that a lot of these are there's a lot of variability first off like some of them are pretty consistent across resources but you can see that there's other weird predicted genes here and here that don't seem to match any of these gene models and these are the ones that were predicted by some of the gene prediction software so a lot of the direct evidence even when we have direct evidence it can be challenging for us to match what these predictions are with things that are actually presented with what the actual genes are doing so a good example here to I think is if you look at some of the gene prediction software I'm going to see how accurate is from a 2011 paper so vitold but you can see that in a lot of these gene prediction software's the accuracy is not particularly high even at ninety percent accurate right like that means that ten percent of the time you're wrong that's being wrong quite a bit so you you know it's these are challenging processes and you'll notice like you know some of these are actually quite good like these predictions right here actually match very much well with what the expectations are but over here maybe less so so one of the ways we deal with that is there's a lot of different approaches to doing things one of the most common pipelines that used and we'll talk about in this class is a program called maker maker one of its strengths is that it actually uses different approaches so it uses evidence-based approaches or extra extrinsic process approaches purchase they use data outside of itself to identify where gene annotations might be and it also uses ab initio gene fighters so rule-based gene prediction programs that use intrinsic data so these are predicting genes from first principles and what it does is it takes both of those a pieces of evidence which tries to combine them to try to minimize the amount of error that you might experience so one of things that i mentioned is that you know if we think about a genome annotation and the steps that it takes to get there the first step is actually there for the portions that don't code for proteins so often we think about annotation is being focused on the functional components but if you're actually annotating a genome the first thing you want to do is figure out where the repeats are and one of the reasons why is that you saw how much error was involved in this a genome annotation process you can imagine that removing a bunch of the repetitive element from the things that you're considering would be very helpful and identifying you know which where you should be looking for genes so if I want to get rid of a lot of those spurious open reading frame predictions finding where the repeats are so the areas I can ignore what I'm doing with the annotations actually really important plus it's just important that we know what the repetitive elements are and where they're located at genome so there are a bunch of different tools there's repeat master which is an extrinsic tool right using evidence specifically the rep based database and their intrinsic tools like repeat master te class that actually predict repeats from rules and when you're doing genome annotation you kind of have to use both of these because unless you're doing a model species there's probably very little information on the repeats that are actually in that organism so you're gonna have to do some intrinsic repeat prediction because rep base does not have the things that you're looking for in it so it won't be good for identifying repeats so why is identifying repeats importance let's talk about this little before but you know repeats are biologically meaningful that's the main thing right so there are drivers of evolution they're implicated in genome arrangements in drifts a new biological function and increased rates of evolution under stress they are they could both have negative and positive effects as you know so it's for us to understand genome function genome evolution it's really important that we understand repeats repeat masking as I said before is a very important partner in being able to identify and look comparatively at genes so basically from a practical standpoint we need to mask these genes before we perform Arctic analyses there's this great quote by Webb Miller every time you compare two species that are close to each other that then either is to humans we get nearly killed by unmastered this was communicated to a colleague by not to me personally so an example of this is to say I have these two genes right and I have these repeats that surround these genes it's gonna be really hard for me to do an alignment of this segment of DNA if I have this these repeats aligning to multiple places on this sequence right so it's gonna be hard to me to line them up to each other so it's important that if I if I mask them out in the tooling we keep Oscar now it becomes much easier for me to align the pieces the parts of the genome that I actually think are important to the kinds of questions I want to ask so yeah we can do this and we can mask these repeats right but repeat massacre is an extrinsic tool so one of the things is is we have to know what the repeats are what the repeat families are to be order to mask our genome and do this otherwise we're going to be stuck in this situation right which is much more complicated and difficult situation to interpret so how do we do repeat family an application we have to know the repeats well one of the ways is we do it manually and this has been done for a long time so we can look at the genomes of human and mouse and we can start to curate them manually and there are manually curated databases like rep base for instance and you know repeat master which is assist strings of Tula juice which uses red face basically uses a search algorithm similar to blast or a hidden Markov model searching tool called Hummer to identify matches from a database so it looks at either rec base or a fan which is the the hummer tool for the hidden Markov model searches and you know it'll find it'll take this library of repeats and it'll try to identify where those repeats are in a sequence if you look at a sample report from there you can see that this is what it looks like we need to look in rep base there's key words associated with it they'll tell you the type of repeat that it is it'll often give you a brief description of that repeat of what it does or what it's involved with giving its title its origin and I'll show you some links to download if you want to download one of the things that you might have noticed here is that when it does this where is it derived from is derived from consensus sequence typically when we think about brief identification we are identifying a repeat family and so how we're gonna do that is by creating a consensus sequence for that family so if we were to think about you know multiple repeats that may come from the same family this is actually how we would define the repeat family with this consensus sequence so we have these three sequences that we thought were from same repeats family right and they come from this so how you know all repeats aren't gonna be identical to each other these are not parts of the genome that are constrained by evolution right so they can mutate and there's no so consequence there's no protein sequence that they're disrupting necessarily or maybe not one that we would think could be as strong as it would be in protein sequences there are other reasons why they might be constrained but generally you think they're less constrained so they can change a lot more so within a group of sequences we tend to call a repeat family of a group that is you know similar to each other that can be identified as being part of a consensus sequence why consensus sequences well it's faster to compare one sequence to genome the many right so we can look for every single repeat the family independently it takes a lot of time but also consensus is can actually do a better job for identifying individual instances of a repeat especially when to try and discover copies new copies that may not be part of the family already that you might not know anything about so the example here is if we look at these three sequences right in their distance from the consensus sequence their individual distance to this to this new sequence right it looks like oh this is you know three base pairs difference this is about three bit this is three base pairs difference is four base pair of differences we might not think that that's actually very close at all to this new sequence but if we were to generate a consensus sequence of those we can notice that the consensus sequence is only a distance of two so if you're looking for new copies right often new copies are closer to consensus that they are to any other individual copy there are many different types of repeats we talked about this early in the semester there's interspersed element's values lines retrotransposons DNA transposons those kinds of things there are simple repeats there are many and microsatellites these are repetitive sequences there are non-coding RNAs as well we can also then it basically RNAs that don't code for protein sequences also long non-coding RNAs that aren't listed here and then there are often contaminant so we might think of being repeats in libraries too so some basics about how these tools work especially the extrinsic ones we think about repeat mascar uses a blast like tool to compare libraries to a query sequence it's got a couple different options there's cross-match full blast there's no RM blast and there's also ID fam each one of those does their search something differently some of them are faster than others and after you're done with the search right repeat master will report back in complete matches to the best of its ability and you'll find that certain types of repeats are typically incomplete as well but master is happy to avoid the factory so how one of the problems is these imp with these incomplete repeats is that we think about how repeats actually are they're often nested within each other so new people ask ur tries to detect this problem where you know if you have here you have just one repeat right and this we go through time like because this is an area that's not constrained by evolution you could have other repeats kind of build up around it or this one get nested inside of another one right and over time it might get nested inside of another one right and so what might have been an original repeat now so you can fragmented and smaller it's a smaller piece of a broader repeat family so to deal with that wreaking misaki trying to detect the different components of this there's a bunch of this is not a perfect process by any means so you know which library so what is the actual reference sequence you choose to do - - you use the mask your your genome as a big effect so for instance in de novo genome sequencing I have used you know the human repeat library to masks a horseshoe crabs right and if I do that I only end up with like 6 percent of the genome being identified is repetitive well that's because all the repeats that are in that are in horseshoe crabs are not in humans right so to mask those repeats I have to predict them and this all oh use another way to generate them and then I'll be able to mask you know 40 50 60 % does you know so you got to be careful about which library you pick make sure that it's correct for your target species some organisms that are commonly used have pre-selected libraries are rare for them but most don't and you should understand that when you do that masking you are often going to have incomplete masking whether it's because something's just not in your database or there are other things that can drive it repeats can be how they diverge from each other which means that they may not be easily identified and also there's a there's a process knowledge sometimes the ends can sometimes be broken off from the rest of the repeat they get fragmented and what you'll do is you'll look and you'll you'll end up finding some piece of sequence or tool which will blast the some other gene or structure or something like that and is this really a new feature or is it just part of what this fragmented repeat was previously so some things to think about so there are a bunch of different ways of identifying repeat families there's not just this one way on this extrinsic way and you know if we think about how you look good like all the different the different families identified in this a little sort of cartoon right how would we think that certain things belong to the same family and certain things don't belong to the same family right they would tend to have specific features about them that we used to define them you know there are definitely features about fishes that are very different from a lot of all the other organisms in these groups or birds for that matter right but a lot of them have four legs or four appendages so there are characteristics that we can use to identify repeat families consensus things and so we can think about I'm using it de novo tool like Rickey modeler and when I talk about a de novo repeat prediction tool so de novo meaning of new what's important is that you know as I was going here and I was looking these things I was coming up with rules right that's coming through oh this one's got feathers this one has scales this one has these these ones all have four legs you know like there are things that I can do to identify you know which ones have similar facial features you know all the things that I can do to make rules about how I'm gonna group things or identify things and so that's what de novo tool prediction tools are doing de novo to have a Miss your tools is there picking rule-based ways for identifying whatever thing they're trying to identify in this case it's repeats so here if we take a tool like remodeler what it actually does is a combination of using three different tools so three different algorithms to try to identify repeats without any evidence so it uses recon repeat scout and TRF and basically tandem repeats are sort of taken out of the equation but we're looking for it using repeat scout which is good for finding highly conserved append of elements and recon which is good for finding less highly conserved elements it's a nice quote by Eddy here or papers is the problem of automated repeat sequence classification is inherently messy and ill-defined it does not appear amenable to a clean algorithmic attack right so that's why we're using multiple programs in this case so if we think about again what is a repay family if it's a repeat family is a collection of similar sequences that appear many times in a genome we can think of a Lu elements it's a good example be family so these occur over a million times in the human genome make up a large portion of our repetitive elements it's about a 22 base pair consensus sequence and it's how we're able not so not all member aliments are gonna be 282 base pairs but it's gonna have a general consensus sequence that we'll be able to use it to identify that specific sequence right and there are consensus sequences of lots of other repeat families that we are gonna find in our genomes so his identifying once we identifying these repeat families an easy problem well if we take a whole bunch of say we had a DNA sequence right and we were trying to find the al elements within it it was broken up into pieces like a genome assembly where they could be in different scaffolds or contexts we might be able to find these used in the consensus and then I'll you know align these elements to each other one can build the ally consensus sequence for it right so this is a process by which we have type people identify these a lose these repetitive parts of this of these segments line them up and then create the family consensus sequence then we can use a tool like repeat massacre to try to identify where all these positions are not no sequences so some of the difficulties here are that regions containing repeat occurrences if we don't know where they are prior a priority how a priority how do we know in other words beforehand how do we know to look for them and where to look for them we don't know the boundaries so when we're trying to make this consensus sequence you know how do we know where to cut this thing off that's also a problem many repeats occur in partial copies because they've been mutations over time so you know when we build this consensus sequence is it okay to use sequences for we don't have the whole sequence and so there's a bunch of algorithms that you can use to do this says you'll notice a lot of them were first developed back in the early 2000s there still is some active worked on this but it's not as big of an issue for a lot of people because I think repeats kind of are ignored a lot of time but the tools work decently I would say so let's take a look at something like recon which i think is a so it's a pretty successful tool so there are some issues with trying to do pairwise similarity so if we say we were gonna look for all the alo elements for instance so if there's approximately a million al use right if we were going to try to do all the pairwise comparisons there we're linking it over 10 to the 12 pairwise alignments which is just too much computational time to try to find and do a pairwise approach or that's just too much data and there is again this issue with defining repeat boundaries and pairwise alignments so local sequence alignments generally don't associate with biological boundaries and it's very difficult to define elemental boundaries of all at the same time clustering elements into families it's a simultaneously difficult to do so repeat Scout takes an interesting approach to this it's one of the two tools used in acute mascar and so here you could have a stretch of a genome where you have all of your repeats spread across it and you can break those up into individual sequences with some flanking sequence right so you can kind of try to find where these repetitive regions are then align them to each other and when you do what you'll see you know we have different pieces of fragment some of them are more fragmented some of them are more distant from each other than other pieces right but they're all considerably part of the same family so how do we build a consensus sequence from this well what happens is we find the seed or the core of it which is this Kaymer seed which is the sick which is the same across all these sequences so once you find that same sequence then we start extending it in what's called a greedy fashion it's a specific specific type of algorithm but one base pair at a time so here we'll start extending it extending it and then you'll notice at some point we'll will have a section where one of these sequences will not align very well so it's it's a fragment of the whole family so we'll just discard that one sequence once it stops aligning to the consensus and we'll keep going and we'll be able to maintain an alignment and typically that alignment will depend on us keeping it above a certain score and so what we'll do is we'll stop aligning once we get to a point where the sequence is no longer really aligned to each other and you note that like if we were doing instead of consensus if we doing you know multiple sequence alignment consensus sequences here if we were doing pairwise alignment we would actually in what two of these cases would have kept going where at a lot of these other ones we wouldn't so just goes to show you that pairwise alignment is a poor way to establish the boundaries of these families and so first you do the right knee to the left and eventually you end up with what you can identify as the consensus sequence of the family based upon these results so that is how we would look at identifying portions that don't code for proteins next we're going to talk about the parts that do code for proteins and again we're going to take the same kind of approach they're both extrinsic and intrinsic tools once they use evidence and ones that don't and then we can take we can use a tool like maker to actually combine those results into one so we can use various types of evidence protein evidence RNA seek evidence homology whatever it is or we can use first principles this one where we think we see genes in a genome of unknown function so how do we do these gene predictions um well you know for things like protein coding sequences you know they're the same kind of rules that you guys have been studying for a long time you know where do we find open to reading frames where are the stop codons or the stop codons like what are the rules that we would associate with protein transcription and translation gene transcription and translation and so there are rules specifically you know the ones for protein coding genes are pretty clear for non-coding RNAs and for other regulatory elements they'd be less so but there are rules that do exist for some of these if you're doing a good job with your computational gene predictions for an OP for a new species or an unknown sequence genome sequence you would do these predictions and then you should at some point try to figure out what these things are based on some other evidence right so there are ways to confirm experimentally what you're looking at and the reason why this is important is eukaryotic gene predictions of really high error rates yeah before I showed you you know they could have up to 10% error rate or more so it's a pretty higher rate of error you can even see this shown in references like refseq where genes proteins and mRNAs that are experimentally confirmed are preceded by an n whereas those that are just computational predictions are preceded by an X so you might have different levels of confidence in the sequences joking on based upon that so the primary goal of computational gene prediction is basically the label the nucleotides of the genome sequence especially the ones where we can identify where we can label these genome sequences where there's a where there's a protein sequence so if we take any one sequence and we look at it we're trying to figure out which parts of that sequence are exon and which parts are spliced sites and which parts are introns right and we could look at different paths through the sequence based upon different rules and probability structures and the same way we could think about it kind of thinking that this P and now goes through those hidden markov models that we talked about early in in the class they're using the same sort of approach right here right where they're trying to figure it as I move through what's the path where I go from Incheon to splice site to exon and back to intron and back X I close the sides of things so they're just trying to describe what the different parts of the sequence might represent based upon rules and probabilities the basic properties of gene prediction basically you use models to do this right and so these models have to satisfy some biological constraint that's where the rules come from so codons coding regions a gene has to start with a start codon you know initial exons have to occur before splice sites and introns like basic stuff you know but are we gonna stop with the stop codon yeah that's these are all things that have to be kept to be identified they'll usually use some sort of algorithm to do this process an example is a finite state machine which is specific algorithmic approach and they'll often use some sort of species specific characteristics tune to improve gene prediction for instance different genomes have different base level GC content or codon bias or distributions of exon intron size and we're gonna predict like shouldn't my exon ended here or should I end over there from comparing two different paths well based upon the distribution of exons you can figure out if one is more probable or less probable in a genome and then you can also look at protein sequences from a closer related species to give you some sort of idea like what the evolutionary history might be there for them and how that might help you make a call between one Passover prokaryotic gene predictions are a little bit easier than eukaryotes because you don't have generally no splice sites right so a single open reading frame you have different types of start codons so that can be a little more challenging but there tend to be a little bit more sensitive and specific than eukaryotic gene predictions two commonly tools are like glimmer or G mark s or s 2 in this case those are all predictors you might use eukaryotic gene predictors Augustus is one of the big ones you can use it to be both ab initio or evidence-based gene prediction snap is also another good ab initio tool and then there are some ways to use some extrinsic methods using RNA seek and some combined approaches like maker where you can combine multiple tools in one the lines of baltimore results to generate one annotation additionally eukaryote gene protectors have high error rates sometimes as much as greater than 50% really high right there a lotta there tend to be lots of false positives in this case and a lot of that will depend on how well you've done in masking your repetitive elements as well there's a whole bunch of ways that gene models can be varied cause there to be all these complications you know you can have alternative splice sites isoforms you can have non canonical splice site so spy sites that don't match what you typically see you can have non canonical start codons so a place where you know binding is not like your actual start codon you can have stop codons that are just read through for various reasons you can have genes that are nested within each other you can have pseudogenes and you can even have genes that splice together from different parts of the core of the G of the chromosome from two parts of the genome so there's all these ways that gene miles can be really complicated and can really complicate your teen prediction so let's talk about gene prediction abacus you gene prediction is it's one of these rule-based approaches and so this is when you're going to be predicting genes based basically just on having DNA sequence alone and you're going to search for protein coding singles you know start start codons open reading frame splice sites things like that stop codons these are typically based on a probabilistic model you thought usually in hmm or and support vector machines which is kinda like a machine learning approach and a classic example this is a program called gene scan and how this works we've talked about hidden Markov models previously in this class this is a type of supervised machine learning algorithm in this case we're using Bayesian statistics and we'll use a training set to kind of create the profile the hidden Markov model that we're gonna want to build first we use these and all kinds of things they come from speech and gesture recognition applications but in bio from Attucks we use them for gene prediction sequence alignment chip C protein folding all kinds of stuff so if we think about just that gene scan each hmm model basically what it's gonna do right this we have like all the parts of a genome sort of showing parts of a gene excuse me shown here we've got the promoter or the utr right promoter these here the initial exon introns exons [Music] terminal X on the 3-prime utr poly a signal right all the different kinds of put components that we might think and we have to figure out to figure out what the best path is through this model and so you know there's going to be emission and transmission probabilities associated with all of these different spots on this path and so as we were to think if we were gonna move through it right we could start with our intergenic region we're gonna get it motor this is just one path that we could take through this through this map get then go to our initial exon introns internal exon go to another intron the terminal exon right then our three prime UTR and then to our poly a single right and this that's like an example of a path and that sort of path right ends up being described in this way and what we can do is we can use the hidden markov I'll describe the probability of moving through a path like that right and and we can sort of look across a DNA sequence in this we move along this DNA sequence we can be calculating the probabilities and whether or not it should move from one group on well then I should move from promoter to 5 prime UTR to exon right and that's what's going to determine where I establish all these things and we use sort of specific marks like start codons and stop codons and splice site models and things like that to predict where we would expect you'd have transitions between these entities evidence based gene predictions basically use sequence alignments to improve predictions so either usually usually either es Dec DNA so RNA seek or transcriptomics data or proteins have cosas David species to identify regions where we would have a much higher expectation of there being an exon in areas where you would expect to see introns so if you can use this type of information you really improve how good you are at identifying exons and that's what you see over here cases where you know as we're able to include more of this information we're able to improve our excellent specificity you can also improve predictions using comparative genomics so looking at whole genome alignments and looking at where genes were previous on other or closely related organisms or where genes are on your organism so these can programs like contrast when you look at these whole genome alignments they can help improve gene prediction again both you can see in this case contrast is helping both with exon specificity and exon sensitivity and a little and gene and sensitivity and specificity as well as compared to a tool like gene scan which is not only using that a j-jim model and you know the kinds of evidence that they might use in these cases again are either gonna be protein sequences so known amino acid sequences from other organisms typically or may be assembled RNA seek or download these these for protein sequences one thing to keep in mind is here we're saying that the more conserved two sequences the better it's going to be for annotation so it's gonna be more if the special sequence is more similar hasn't changed a whole lot your evolutionary time then it's gonna be better at identifying these genes there obviously is some bias here so you know if you're trying to do flies you're just plugging and do a lot better than if you're doing dragon flies because there's not you know like there's Drosophila up there a model system for flies but they're you know Oh de nada is a whole different group and there may not be the same sort of kind of reference to use in this case proteins can often be incomplete when you do this so they're gonna depend on these protein alignments but you get a fragments on either ends of the proteins and these pairwise alignments that may not line very well so your protein may look smaller than what you think it actually is usually how this happens is you'll align the proteins to the genome and then based upon that alignment you'll be able to identify the segments of the protein sequence and align to different positions on the genome so again you're aligning protein sequence here to DNA sequence so you're gonna have a trends translate it back from protein sequence to what might be a key in the alignment RNA seek works kind of how we expect it right when we're when we're doing RNA secret blue the the thing that we're looking at is mRNA sequence typically and so that would mean that when we generated this sequence what we actually have to look is that these regions will represent you know the exon structures that were originally in the genome so there is some expectation that you will have reads that will span splice sites right that will spin that this DNA sequencing read here you're gonna have to cut it in half or in some weight so that part of it aligns here and part of it aligns there but that's good that actually can help you identify where the splice sites are so if you can I would always suggest including RNA seek data annotation project because it's a very powerful way of validating your gene express your gene data by showing you genes that are truly expressed should use from the same organism if you use it from a different organism there's gonna be sequence bias based upon differences just and things like GC content opponents keep in mind that this is gonna be very noisy because you're gonna have alternative splice variants from different tissues so you probably want to use information from more than one tissue to compensate for that or more than one life stage you can also have preeminence in the sample which is complicated as well right because then you're trained it would look like a transcript but in fact it's gonna have the intron still in it usually you don't have those in high frequency so it's not as bad but just do something for where I have here avoid gonads but you know I'm pretty sure that the vertebrate genome project now says do brain and gonads so I would I would say you know just make sure using a variety of different tissues and again keeping in mind that you're gonna have Reed's that are gonna map the splice sites and you're gonna have to break them up and make sure that you map them that you map them so they can go to different parts of the genome and so there are specific mapping tools that you would use for that process and you can see when you align them this is what the pileups look like right you can see cases where there are splice sites right where there are gaps in the middle of reeds so we can identify where versus areas where there is no sequence things will be intronic regions right where some of these other ones look like they're more like splice sites so we can see the intergenic regions here we can see the splice sites between introns where the introns are exons splice site intron again here exon intron so we get a gap in these reads exon intron and notice here you know there are some examples we have some sequence in the middle this might be from having a pre mRNA sequence in our sequence or an alternative splice break so it is possible that you will get reads that go into this segment [Music] predicting where those splice junctions are such a really important part of gene prediction and that's another reason why RNA seek data is so important because it's going to be really hard to predict without it so just keeping in mind again as we go through this process you're gonna have certain reads that are gonna have like anybody broken up and those reads will have to be mapped just in fragment to different did you know two pieces of the genome that are somewhat distant from each other so it's a different computational problem but this process is important for helping us identify how we're splice junctions actually are this brings us to a second issue involved in genome assemblies whether or not you actually need to build a transcriptome assembly and so eventually we want to actually figure out what are the transcripts and protein sequences that are being produced in this annotation process and what you can do is you can build a reference based transcriptome assembly so you can take all of your RNA seek fragments that have a hard time mapping to your genome right you then find your best path through this alignment and what ends up happening is you're able to then compress this alignment down to identify the different possible isoforms that you can see and so this is one of the ways that we can take information on these assembled entities where we have reads that might be split into different places like I read them be split between these two places right we can find reads here that are split in three places one two three and as we build these pathways through we can find examples of you know what this blue pathway is here what this red pathway is here right what the yellow pathway is here the scene here and what this green pathway is there right we can see the different examples of isoforms that we can then generate from these two productions and so when we you can also do this process with a de novo transcriptome where we don't use a reference to help us drive this process we just assemble a transcriptome from scratch and this is an example of how that might look so after we're done with that de novo transcriptome where we just take our nice ink reads and assemble them into transcripts this would show this is an example of how that transcriptome and maps back and you can see that animal transcriptome actually did a pretty good job of generating transcripts and these these of actual reads and that in fact took the example are ultimately we want to do is generate consensus gene models so you know they all have different strengths and weaknesses all these different tools that we talked about and so a tool like glean or maker will then try to take all these different outcomes different tools and then combine them into one so we just take a look at clean arm which is used a lot more interests often systems here in this case we can have we can see that there are a bunch of different possible outcomes for various transcription here William mallanna gaster and Grimshaw and so we can look at at this region in Tri parallel what is the confirmed prediction in this case right and we can try to compress these entities together to remove variation after we're done with this type of process we can then further validate what we found by looking for biological evidence of our gene predictive model and I'll go into this in more detail next but basically yes all the ways we can cross validate our results so we can if you look at the gnomon NCBI line for instance like it's going to go through all these different parts this is mostly an evidence-based process right this is like identifying genes based upon other genes that already exist there's not going to be as much of the gene prediction components as some of the other pipelines and you know they if we're looking at same you know fifty to sixty percent sensitivity and specificity x' in other processes and just the rule based approaches alone the evidence-based approaches are also pretty good you know about seventy percent sensitivity and specificity however we could also say that it's thirty percent they're missing so it's not like just anyone approaches the answer to all this stuff a lot of times what they'll do is they'll use they'll often use RNA seek data from say like another species to help improve their gene predictions and then you know we have this gene that we predicted from this genome only predicted using melania gaster and scrimshaw and we can then apply it to something like a rector some common problems we have in these gene finding programs it'll split single genes into multiple predictions they just won't be good evidence that it's more than one could fuse neighboring genes together exons can be missing in some cases and sometimes it does it over predicts exons so it'll take a final open reading frame and it'll include it when it really shouldn't be and sometimes ice forms just are observed because they're not evidenced anywhere in the RNA secret protein data but ultimately you know using multiples of these approaches and combine them is probably the right way to get the best evidence for these gene predictions and maker is a is a good way of doing this how maker goes about doing this process is basically it'll identify Rapide to the line that you have seed in proteins it'll then do its AB of Nishio gene predictions and then it'll try to synthesize both the values of the ab initio prediction and the extrinsic approaches to gene prediction so goes through this whole process usually you'll need some sort of extrinsic data it could be protein data could be RNA seek data in various file formats different ways can be used and then for the intrinsic part of it there's often a training and Retraining part this were to run through the whole process one time then repeat that process in order to get better at it so there's a lot of different ways that we can go back doing this so as we get towards to kind of wrap sort to wrap this thing this idea of prediction up gene predictors are very important this whole process is very important for annotating the genome identifying what are the interesting features of a genome and so not just genes but also repeats as well predictions should be treated as hypotheses that need to be confirmed experimentally you know they're not it's not okay just to take computational data at face value in all cases sometimes it's very good other times you want to at some point go in and make sure that you're working from what what is to what is known to be true rather than what you just assumed it true eukaryotic gene predictors generally can accurately identify internal exons but you just get much lower specced sensitiveness message with specificities when predicting complete gene models so usually they do good in fragment but getting whole models is challenging and it's the end of the talk next time we'll talk a little bit more about some other annotation approaches that are more funky in their approach Thanks
Up Next

HPA Axis & Cortisol: The Stress Response Explained
@PhysiologyforHippies
70K views•2016-04-18

Circadian Metabolomics: Sleep, Food Timing & Human Clocks
@tscnlab
359 views•2022-11-10

Enteric Nervous System Explained: The Gut's Brain | Neurobiology Lecture
@alumniu6029
438 views•2018-09-12

Bacteriophages: Earth's Deadliest Killers and Future Antibiotics
@kurzgesagt
34.6M views•2018-05-13
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biology







































