Phage genome annotation requires manual curation of computational predictions (using tools like Glimmer, GeneMark, tRNA Scan SE, and HHpred) because phage genomes are highly diverse with many unique genes, and researchers must answer three fundamental questions for each gene: whether it is a gene, its start position, and its function, using comparative genomics and multiple lines of evidence to assign functions to the approximately 30% of phage genes that have known functions.
Phage Genome Annotation: Bioinformatics for Actinobacteriophages
Added:reading as we just were talking we realized the importance of genomic characterization of organisms and especially given the pages are specially adapted to their host and their lifestyle life cycle includes short genes and less capacity for redundant dna sequences and trna and other features but they're not any they're not too many specific tools for annotation of eight genomes and even to characterize them completely which includes their assembly and or predictions assigning functions and uh looking at the architecture of the genome etc so we iprc and phase directory thought that given the grave uh given the dire importance of this area and we understand that you know the genomic genomic characterization is very important to understand the phage biology and also to you know understand when we want to apply them for any application whether therapeutics or any other it's important to know their genomes so if this background ibrc and page directory are co-hosting a series of phage bioinformatics webinars and the first in the series is by dr deborah jacob serra from happy lab uh from hackford lab whom we had earlier in june last year we had him as a speaker she's from department of biological science university of pittsburgh usa and she's a veteran in the field actually i don't know how many people how many teachers and students she has trained today she's going to present on genomic annotation and comparative bioinformatic analysis of actinobacterophages she is developing actively developing the phage discovery and genomic platform for advancing science education she oversees program development for the fire program which is talk about in her talk and coordinates many aspects of c phages program as well debbie as we affectionately call everyone calls her debbie so i'll take that liberty to call her the same she is spearheading methods to use a variety of actinobacterial hosts for phage isolation which include mycobacterium arthrobacter microbacterium rotococcus and gordonia so with this background and with this introduction uh may uh maybe welcome uh you for this first webinar of the series and without taking more time uh floor is yours thank you very much for accepting this invitation and uh very excited me and jessica both are very excited to have you well it is very nice to be here um i um never met a phage i didn't like so for me to talk about them is always a fun thing to do i'm going to share my my screen and we'll go through a powerpoint and i am up for questions as we go um the trick is on zoom as you i'm sure you know is that it's hard to see the questions so jessica and army are going to help me and don't hesitate to slow me down or stop me with whatever you want to talk about okay the biggest question out there is why do we study them and how many phages are there and um we study them because they're there we started um i started with graham hatful over 15 years ago and i had come from a clinical background and teaching in schools and i was very much interested in how can we do things that help others and i come to graham hatfield's lab and he just wants to study phages because they're there what are they and how did they get to be that way and it really made me refocus how i think about the world um from a basic science from versus an applied science background and then lo and behold in the last couple years we added our applied sciences to this with phage therapy so there's lots of good reasons to look at these um and there's lots of good anecdotal data out there that suggest that people use the data that we learn from these phages in lots of different ways um one of the reasons we study them is because there's a lot of them and every phage biologist that starts this talk always starts with how some analogy is to kind of approximate how many phages there are graham likes to use um fluorescent a fluorescent dye picture of dna in seawater and approximates how many more phages there are than stars in the sky um the the late roger hendricks used to set used to give an analogy to liken phages to beetles and he used beetles because they're the most common numerous insect on the planet and you would when you approximate um if you make every phage a beetle we're talking about miles and miles of beetles that cover the earth but my favorite analogy actually comes from a phage biologist by the name acosta georgiopolis and and costa says let's try weighing them so if you can imagine a balanced beam and if you could put all the phages on one side what would you put on the other side to balance that that weight and he suggests um and this is where it works really well when we're all in the same room and you could look around is if you put the entire world population on the other side but only if we are um sumo wrestlers and so for some reason my ah there we go there we go okay so um there's my my silliness for the day okay um i work in the lab of dr graham hatful and he and we study actinobacteriophages and the the numbers of phages that we found i just looked up this morning is we have um over um 18 000 phages in our collection and we've sequenced um 3 700 of them and um almost all of them are by student researchers in um a program that's situated at the university of pittsburgh called fire and the one that's all across the globe mostly in the united states called see phages where students go out right now in a two semester course at university and they take a semester and they find the phages and characterize them on on the plates isolating purifying looking at restriction digest patterns electron microscope pictures and then sending dna off to us where we sequence between the two semesters and second semester they bioinformatically and analyze them and students love to work in a wet lab and then they like go well why am i doing this stuff at a computer and it takes them a while to realize that that's where graham does the science is at the computer analyzing and many many bench experiments start by graham hypothesizing what he sees in the genomes and um and then lab work is done to um either prove them right or prove them wrong and like any good group of lab um lab scientists who work for a pi that is extremely talented is it's fun to prove him right but it's also fun to prove him wrong so we have a good time at all of that um let me see where i'm at okay so the first thing you do that we do that's different from lots of other ways to think about phage biology is you can just go in and get an environmental sample put all the extract the dna out and you'll see a sampling of what's on the planet by the sample that you take um it's one way to look at the biodiversity we want to do more because once we identify and characterize the phage we then want to be able to go back and manipulate it because that's what we do as humans is we um understand science and we try to make sense of it right and then use it for our bidding and so we archive all our samples so that we can go back and use that phage in whatever way the science kind of dictates to us okay once we get the dna extracted we sequence right now the sequencing of choice is illumina because we can get good enough coverage and we have enough experience that we can finish a genome without additional um technology so it's i don't think we've done a sanger read to finish a phage now for over a year um so um the the coverage is really good um when the coverage is good and the quality is really good we've used we started at sanger sequencing when technology was at that point and so we'll move wherever the technology takes us but i will say that in the actinobacteriophage hatful approach to this we're really picky about how we sequence how we finish and the orientation of a genome and by by orientation i mean the sequencer has no orientation um if if it's a phage and it's circular it's going to start it wherever it feels like and just show you that the coverage is xyz and so it's the sequencing sequencer's job to figure out if in the phage um how that dna exists in a phage head and there's been lots studied by this and um there is a correlation between the terminus of a phage and what kind of um of um phage it is meaning how are the ends determined is there an overlap that goes that's a five prime overlap is there a three prime overlap does it have direct terminal repeats or is it circularly permuted meaning that no stage has the exact same um length of dna in it but the unique dna is all the same and so you don't know where it starts and where it ends and i wrote the quick the trick question of are phage genomes circular or linear because in the host if they're going to take over a host they're definitely going to be circular or they'll be chewed up but in the phage head they are a linear piece of dna and since we report phage dna we want that linear piece of dna to match what's in the phage head and then the order of the genes in the genome is there for human convenience right that if you go back to the beginning of the genomics of phages the best we could do is um look at a particular gene right you found the capsid gene on a gel and then you were able to sequence that and because it was out of context of the genome we wanted to be able to read it left to right and so it became very common that the structural genes read left to right and that we put those because they were the first things that we studied at the beginning of the genome so our orientation of um a genome is to read um the structural genes in the forward direction and so they end up on the left and in the forward direction and if there are defined ends in a phage you usually find them in a particular order and that particular order is something like terminase portal more capsid processing genes like scaffolding in the protease then the capsid genes the tail genes lysis cassette immunity cassette and then all the genes that mess with the host's dna and they can be primases and polymerases and reductases and we could go on and on and on of what you could find and none of those genes make sense because they're not needed by the phage um at all the the the phage just needs some dna binding sites genes and then the structural genes and it will work what's really nice about all of this part of sequencing is that my colleague at the university of pittsburgh dan russell who's in charge of our genome center has written some very nice tutorials on finishing an orientation and you can find them at our website which is phagesdb in the resources protocol pages okay the next step is annotation which is what i'm going to get to first um and then it is coupled with step four which is the genome comparison comparisons not going to be able to go there but the genome comparisons that we can do now actually inform how we do annotation um and it helps us to refine lots and lots of steps okay um in genome annotation the big overarching rules are is if we could do this by annotation through automation we would um we don't because the phages don't follow the rules and they're so small and their genes are so small it becomes difficult to assert the programming parameters and when we're done we're going to compare with everything else that we see as i just described sorry okay the skills of bioinformatics i believe are really skills that you've learned when you were we a we child and i contend that the skill that's needed for genome annotation is the skill of pattern recognition and i use this reading readiness test to test folks to say if you have this skill you can do genome annotation and um you can all figure out where in the fourth box you would put the the dot is that true i'm gonna take a yes okay um and so you all have the skills that you need to do this okay the second key concept is um that because we are in the bioinformatics world it is going to feel very uncomfortable um in in the true sense of biology if we could take every gene to the bench and study it characterize it find out what it does and how it interacts with all the other genes it would be a perfect world we don't want to forsake the good for the perfect and so we're going to use analogous homologous everything we can we're going to throw at these phages to infer as much as we can we are going to do what every good science does scientist does and make a claim look for some data some supporting data before we want to um make the claim right we're gonna we're gonna identify a question and then we're either gonna say yes or no and assign this as a gene and what its function is it's gonna feel very uncomfortable i work with very young people and this feels really um uncomfortable and they get really skittish and then they start to think what does it matter like it's it's make believe it's not make believe it's just putative okay and we're going to use supporting data to um to identify um those parts of the genome that we care about we're going to use glimmer and gene mark for their their predictive powers to predict coding potential they use interpolated hidden markov models they look at four nucleotides at a time um they even have a series where they can look at two nucleotides at a time but that gets pretty noisy and should be pretty much discounted um and that um both of those programs claim somewhere over 95 percent accurate unfortunately 95 accurate it's not good enough and so we actually curate these all by hand and no one really assigns functions in a consistent and meaningful way but all pieces of data are helpful as we move forward this is slide number nine i think in my powerpoint and it is the important slide if you take home anything this is the one that you wanna you wanna grab because it has all the links of all the programs that i'm gonna talk about first and foremost is our bioinformatics guide and um actually how you get down and do the work is all and how you use these programs is all written in this bioinfomatic guide and if you really want to learn how to annotate a phage from my perspective this is the place that you go and you get very familiar with it or you use it as a resource jump in and then go back and use it as a reference like i said we're going to use the prediction programs of gene mark and glimmer we're going to use aragorn and trna and they're both online and trna scan in particular can be downloaded on your own server and you can run it independently which might be smart because it's a little bit finicky we're going to use blasts both at the nucleotide level and the protein level in a lot of different ways there's a lot of information in blast data and the conserve domain database and there are ways you read the data for the question that you bring so if you keep in mind when you're annotating that you have a question is this a gene what does it start what is its function and you use the data to answer each one of those questions specifically you'll be very successful at that if you are in the actinobacteriophage data we also can have a database database that is blastable and so if you've used ncbi it takes um some time to do your blasting phagesdb is much quicker and you're going to match the same first 100 hits because we have so much data at ncbi one of the programs that we use that is essential and can be used by anyone else by contacting the creator of it is famirator and if you've ever looked at a sing any um paper by graham hatful um where we are describing a new set of phages we have these wonderful genome maps and they're really pretty and that's who makes them is the famirator program but more importantly it is a pairwise comparison at both the protein and nucleotide level and gives you lots of comparative genome data that you can then build upon and finely tune what you're looking at and then the key program that we use for genome annotation is a program written by university of pittsburgh researcher by the name of jeffrey lawrence and it's called dna master and dna master is a genome editor and most programs that let you interact with dna and genomes are actually viewers and so when you go in and you try to change things in them it gets really tricky dna master has this wonderful feature where it links everything you touch with the actual sequence so you can manipulate it and interact with it in a way that ends up being very productive and i wrote another favorite tool that's in dna master called genome comparison that will let you compare genomes very quickly and as you change them it'll show you the changes on the spot because all other programs are much more fixed okay our genome annotation is um described in the bioinformatics guide and um [Music] i'm having a heck of a time advancing my slides guys so it's it's strange sorry um and in that guide is our three beautiful pages of um what we do to um annotate um and they look like this and there's a lot of boxes and a lot of words and we do all of this um and i could not cover all of this in an hour but i want you to know that this is what we do and it is really all spelled out and gives you clear information on um how we come to the decisions that we make okay we start every annotation with a fasta file and i told you at the beginning the skill you need to be able to make sense of this is to find patterns and lo and behold many many decades ago people would look at this sequence and be able to find things in it if you looked really hard you could find starts you could find stops and then you would be able to make some sense of it thank goodness today we have lots of programs that can do that for us and we try and we do that by converting the nucleotide sequence to genes and we do that by codons and we use for the most part though not exclusively because there's phages that don't use the bacterial on plant plastic code but we do and it seems to work and what is unique to students is they know about the methionine start but they don't know about the gtg and the ttg start it takes me a while sometimes to remember what the stop codons are because they're represented as asterisks in my data and so um i really have to work sometimes at the stops um the space in between a start and a stop is called an open reading frame and then the programs work to find the coding potential in all of the orfs that are present know that for every sequence given there are um there's more than one way to translate it into codons and you know this um and so if we start with the letter a we get a series of codons if we start with the second letter we get another series we start with the third letter we get another set of amino acids in a codon we know that um if we start with the fourth letter we're back into the first row right we're back to this one we just skipped the atg right and we know that the dna in our fasta file is just one of two um sequences and so we know that there's a um an anti-parallel strand that also codes so there's three more ways to code it so there are six reading frame translations for every dna sequence and if we really want to look for it and study it we can get in and do this and dna master has a way of generating this exact document of a whole genome when graham annotates he looks at this all the time and he can readily seek see the promoter sequences and he's really good at finding like hairpin terminators and things like that um for the most part unless i know i want to look for them i really don't but dna master displays the sequence in a way that you can find where you are at all times the features found in actinobacteriophage genomes include our protein coding sequence which is the majority of the features we find we find trnas tmrnas we can identify at p sites um we can have identify terminators we don't normally find and annotate them in our genbank files that usually happens when we've really dug in deep with particular sets of genomes and we really want to know more about it and it's usually written in a paper one of the things that we routinely look for and annotate is the tail assembly chaperone program frameshift and it is one of the fun things i think that we annotate because it really shows an interesting feature of how genes are translated and we won't really deal with that on this talk because it's way down the road from the basics there are three questions we ask about every gene feature and is it a gene what is its start and what is its function so an average genome in the actinobacteriophages is about a 70 kb strand and there's usually around 100 genes so we're going to ask each of those questions a hundred times okay what is most critical is we make the claim of is it a gene what is its start and what is its function is that we have supporting data we're not making it up and we use the coding potential programs and a lot of comparative data to help us get to the answers dna master is excellent at helping us to add genes delete genes and to change the start with very little trouble okay um the coding potential is done and the most important thing about this slide is that you take your phage genome and you take it out to these to the website of gene mark glimmer is downloadable and glimmer presents you with a list of starts and stops genemark actually presents one of its outputs is a graphic output that's really good for when you're working with students to get them to see where the genes are and how they fit together you just have to remember at all times that both programs only use a sample of the genome so they it's a random sample and it looks for the biggest open reading frames and it looks for codon no not codon usage but display of four nucleotides at the time at a time it's easy to um sort of approximate that to codon usage even though it's not it looks at four nucleotides at a time and says what's the prevailing patterns that are there that must be where the genes are and that works really well unless the genes are from a different source than what you gave it to model after so if if they're newly acquired genes their coding pattern could be really different than the rest of that phage genome if you point to a host maybe the host you isolated on if it's in gene mark's repertoire a genome of the same genus you just have to remember that doesn't mean that that phage infects that host but does it give you information that you can use to make good decisions i'm sure it's always fun i think to go out to gene mark's website and use their program and compare the coding potential if it's a phage found on mycobacterium smegmatis find the coding potential compared to smeg versus e coli and you will quickly see garbage in garbage out you'll see that anything could be called a gene depending on the parameters that you use to evaluate it and it's important that you understand that these are predictive they're they're programs that help you predict what's there and that you can cipher through it all the next three slides are also in the guide and there are guiding principles of bacteriophage genome annotation and um they're there for you to read there's some highlights i want to pick up first is that in any part of the dna we only think there's one gene so they're not going to be on top of each other for the most part they also don't overlap each other very much um except for the quintessential beautiful three or four or one base pair overlap that actually the ribosome would prefer um and so this gets into the question of um where does translation transcription where does it all start right and in the simplest form a bacteriophage genome um has two jobs one is to go in and take over the host and the other is to make parts of a phage and assemble them and spit out phage and so you could think about a phage genome as being two operons to do that what we found in the limited bench data we have is that there's a couple more than that there can be promoters in the area of the repressor so if it's going to integrate this will help regulate how it regu how it does that and there can be promoters like around the capsid because it's going to make a whole lot more capsid genes then it's going to make some of the other genes that it makes so and there's no one way to do this right so there are a couple places and that's becoming more apparent as we do more rna-seq data but for now um you could simplify it and we do have genes in the arthrobacter i know and i'm pretty sure in the microbacteria that are like in the neighborhood of 15 kb and they're all structural genes with a little handful of dna binding genes at the end so that's exactly what they're doing they're they're grabbing on to the host machinery the host dna taking over its machinery and making phages so it's exactly what when i first began to study these phages what i thought all of phage needed however most phages have you know 40 genes that are all in dna regulation and what they're doing and how they help or hinder the phage or the bacteria is complicated and complex okay the gene density in phage genomes is that it's very high and this gets really tricky because if we just make it dense um we're doing a self-fulfilling prophecy and so we try really hard to make sure we have evidence for every g that we call but we do find that they are that there's very few gaps in the instances where there are gaps we aggressively look for what could be in that gap but don't aggressively fill it if there's no evidence for anything there so there are occasions where we have phages that have one or two kb which is huge in a phage genome where there's no genes or anything that we can recognize like a trna or whatever present doesn't mean there isn't something useful there we just don't know what it is um one of the strongest rules that we follow is that we're not going to cut off coding potential and this gets really interesting when we [Music] have students do it because the shine delgarno's and start start the ribosome binding site data is numerical and folks want to use big numbers but the only ribosome binding data that is really pertinent is one that includes all of the coding potential and you would be surprised when you compare the gene mark data trained on self to the gene mark data trained on a host how they can inform you and um it there's some nuances in there that with practice you begin to appreciate um most many phage genes are unique you won't find them anywhere in the literature and if you find them somewhere else they don't know what it is either so about 30 percent of our phage genes have functions which is way high from when i started um but just because you don't know what it does doesn't mean you don't call it and that's uncomfortable and uncomfortable for some more than others but i'm going to encourage you to call what is there um you also have to remember that there are genes and genomes that do not follow the patterns you begin to identify the notorious gene that comes to mind is an h endonuclease right it's capable of popping in and popping out of genomes what feels like willy-nilly and so it can have a way different coding potential than the rest of the genome because it's newly acquired and so you have to be willing to [Music] look for that search for it and find evidence to support that phages try their best to be efficient um remember they don't have the ability to produce energy so anything that requires energy is going to be tricky for a phage i think and so you want to be careful that you don't decorate a phage genome by calling a forward gene reverse gene a forward gene reverse gene a forward gene reverse gene it doesn't really work like that um because we're going to need the phage is going to need promoters and ways to get um the each of those genes started if they're like all by themselves going in a direction much more efficient that that the whole genome be in one direction and it could go in one swoop or in two directions and have two different um initiating machinery along the way um it calls the number eight is called into question all the time um that you know historically the smallest protein was about 40 amino acids i think we've broken that record um it were much smaller than that and we're finding that those little tiny genes are things that actually do do battle or provide um added utility to a bacteria um and so we've studied some of those things and found that they are critical in how well a phage survives with a host okay number 10 i now just learned recently that that can be broken but for the most part you have to have a stop codon the stop codon when you're learning to annotate is the only unique feature to identify a phage by lots of people call genes and then you start adding genes and subtracting genes so if you call them by name like as in numbers we typically identify our genes in each genome starting at one and going as many genes as we have but if you start to delete and add you can mess that number remember that the prediction programs can start with a random sample so they may not have and where they break down are in the small genes so what may be gene number 42 in one file isn't 42 in another file and when you're trying to do this in a classroom that becomes really treacherous and the simplest way to figure out what everybody's talking about is to identify the stop codon because it's the only thing that is unique in that particular open reading frame because you could all be talking about different starts um there are three start codons in our codon usage and they are atg gtg and ttg ttg is only used it's an infrequent usage one we've calculated which has been some time ago we're at about seven percent of the genes start with ttgs um they end up being evaluated a little bit more strongly because they're sort of forgotten in the programs that we use if you notice i have a number 12 there and i took it out of the main list and i'll tell you why in a second tnr trnas are [Music] are why are they in a phage genome why does it need a trna it shouldn't need one it's going to use the host and the host better have it or the host isn't going um to be present so um they're there we try to space special attention we have a graduate student in the lab who's working on getting some answers to trnas it seems to be a hot topic i've seen to have gotten a couple requests lately because people are starting to they're starting to gain traction in in investigation um and we're picky and we're trying to um call them as precisely as we can and so we spend a little bit of extra time in how we do that protein assignments are unbelievably rigorous we want data you will find in our data set that for a while we were calling a lot of things based on syntony for the last three years we've tried to revamp that and like kind of pull back a little bit the only gene gene now that i'm calling by syntony meaning where it's located in the genome are those big genes right after the tape measured gene um that have to be minor tails everything else i'm a little bit squeamish to call by where it's located in the genome if it doesn't have some sort of supporting data and that's supporting data most often comes from hh pred blast you have to be careful of because blast is like somebody said it was that so you're going to want to go into using the conservative mind database in the ncbi blast and make sure where that data came from i think that's basically the rule here is that you really want to make sure that you have some sort of data to back up whatever you want to claim and then um iteration the more you do it the more the better you get at it the more you're willing to make your claims and you know where to look for the data all the databases can be very intimidating and it takes a while till you get acclimated and um and get there this is number 12 and it's the longest of all of the um guiding principles and it's the least useful and the reason it's the least useful is it talks about shindogano sequences and getting a genome to be um hooked into the ribos or a transcript to get hooked into the ribosome so it can be translated and it is a very known um understood kind of phenomenon but a lot of the transcripts are leaderless um in our phages and so it's irrelevant and um and so i just don't want to go there my next picture is a picture of dna master with a phage open and if there's enough time i'll show you how i got here but basically i did what we call an auto annotation we took our phage genome out to glimmer and gene mark um and aragorn 1.0 and processed it and it came back and he said here's the predictive genes and it comes back in a list right which is right here in the feature table it tells me how it was called by glimmering gene mark and there's a shorthand for what this all means and i'm running out of time so i'm not going to explain that very well but it's in a list but this is the place this frames window is the place where dna master outperforms any other program i've tried to use because for every single row in this table there is a gene feature that's listed you can see there are six rows so all six translations are represented here and where these genes are in in those six frames are represented this lets you know where all the starts are and it knows it shows you very quickly the relationship of one gene to the other and if they're tightly packed you can see these genes are pretty tightly packed here's one to investigate because is this the right start or should be it be made bigger and we'll look at all starts but that's a a key to saying what would be better and dna master has the tools to help you evaluate that quite simply um the green or forwards the reverses are reds in my version there's a color chart you can make them any color you want and so that gets to be kind of fun but this is where the work is done it allows you to blast all the genomes the genes in the genome you can blast the whole genome at one time it's a rather tedious and long process because you're interfacing with ncbi dna master is a program that requires administrative access so you have to run this program as an administrator so that basically it can talk to other programs that are out on the web and that gets problematic depending on what kind of computer you have and what kind of privileges you have if you're in an institution with lots of firewalls you're going to have to get clearance of all those things there are preferences that are set in dna master dna masters easily download from jeffrey lawrence's website the install program installs one of the most early versions of dna master that no longer works the second thing you have to do with dna master is updated if you don't have administrative privileges you can't update it because you can't get your computer to talk to jeffrey's computers and so um that's your first clue as to how easy this is going to go in our guide we have you have to set up the preferences in a particular way or things don't work and nine times out of ten those are the two things that happen where we get messages from our student population and faculty population that says this isn't working why is that however once you get it working you begin to see how easily you can manipulate the genomic data and get your answers quickly and efficiently that particular visual representation mimics the gene mark output that again shows you six reading frames it shows you each open reading frame if you can see these lines that bisect each row so here's a good example of an open reading frame up ticks or starts down ticks or stops the coding potential is outlined with this black line the red lines are using an algorithm that looks at two nucleotides at a time and i just consider it noise i pay no mind to any of the red lines and then the black heavy line at the bottom is actually what gene mark has identified an area of interest may not even be any anything they really call because they may call other things and here's a really good example of some really what looks like crappy coding potential that should be investigated to see if you're looking for it my guess is neither glimmer or jean mark really called this guy and i'll put my money right now on that i'm going to want to call this guy because there's enough coding potential and it may be something significant and i won't know that until i investigate the familyraider map is found at farmraider.org again that link is on slide nine and it shows you this is two rows of two phages called zizzle and lamoula um the names are fun because um we allow the finders to find to name the phage um name the phages because when you find the phage you have to give it a name because if you don't with thousands of phages found you could never keep everything straight we believe in a non-systematic nomenclature meaning you do not anything about your phage in its relationship to other phages by where you found it when you found it um and what a plaque looks like all the plaques can look the same and different and there's no key component characterization that gives it something nameable in addition um all of the phages in the actinobacteriophage so far i'm gonna i'm gonna take that back most of the phages in our groups are um called a viral in nature so they're sipho verdes myovir days and podovir days um so you can't put all the siphos together they're they're too diverse um nor can you put all the myos or potos we just don't have as many of those and i don't have as much to say but this is two rows of these two phages and you will see a ruler which signifies the genome length and is a representation of the nucleotides and then there's boxes above and below that line the boxes above are genes that are predicted that um are going in the forward direction and the ones below are in the reverse direction okay the boxes are color coded and numbered and i'm sure you can't read the numbers but just because they're the same color doesn't mean they're the same thing because there's tens 15 000 different groups what we call fams of jeans and there's not enough colors to really tell them apart so always check your your numbers um but there's a number above the box that is the fam number and how many members are in it and i think the how many members is what's really important there fam numbers change every time we generate a new family or database and in our busy season we're doing that once a week however we can connect them all and so it isn't um the data is retrievable the other thing shown on this map are as as we finish a genome and little moolah is finished but zizzle is just a draft we add the functions in so when we're calling zizzle we know that this is the terminus large sum unit we're going to prove it by looking at all of our data but it gives us an idea of the orientation and what we can find the biggest gene is almost always a tape measure gene it's a gene that's involved in tail assembly chaperone its length is correlated to the length of the tail and so it's a good kind of benchmark of a phage genome and the other thing that this map shows is that it is a pairwise comparison and i've only given it a pair um but there's the colors in between um the two genomes and those colors are a representation of nucleotide similarity and the colors go the colors of the rainbow with violet being the most similar white being no similarity and in the case of um white you can sometimes see that things can still be in the same family and i don't really see a good example of that um but there's lots of things you can learn by using famirator and famirater is very interactive and you can move the genomes around and there's lots to be gained if you are studying phages outside of actinobacteriophages um steve croissan at james madison university is the owner of famrader and is willing to set up databases with just about anybody so he is a good person to have as a resource and then the bottom line comes into the first genome we find that's like unique we call them singletons at this stage of the game you have very little comparative data to which to rely on but as the comparative data comes in we get to be able to refine the data more and more and more we have a program and again it's um it's in you know free resource and you can set it up yourselves is it's a starter rater and it takes every gene that shares a fam and shows you where all the starts are so that you can really come to a consensus rather quickly when you use dna master know that this was written by jeffrey lawrence i think close to 20 years ago and that you should still consider a beta testing it's written in a language called delphi which is defunct now um and that it's very easy to corrupt and lose your work so you need to know that you have to run it as an administrator you have to save often every time you save you should save it as a new name never overwriting the last data that you know was good because if you overwrite corrupt with overwrite a good file with corrupt work you're going to have to redo it all and it works very nicely in a virtual machine and um we have our virtual machine of preference and you can use whatever you want is virtualbox because it's free and that's what i use um there's other choices out there there's also i don't know if i'm going to say this correctly but emulators like wine where you can run windows programs on um on macs and i would discourage you from doing that because it loses functionality and it runs so slowly it can be extremely painful and then i added the reminder is did you save your files again using a new name because if you lose your work you will cry and you can imagine that i have cried many many times as i have done this work and that's the way it goes so that's the overview of how you get started if you're interested in using our programs um like i said go to slide nine read the bioinformatics guide um our see phages program sponsored by the howard used medical institute um we also have another website called sephages.org and again there's a lot of public facing information there including forums that if you know your question you can the search function is likely to get you to other folks who have had the same problem and then you can get to an answer so i would say that like all the programs used for phages um we're trying really hard to be user-friendly [Music] but it's hard and the data is coming at us really really fast and trying to make sense of that data and then to get to the answers that we'd really like to get to isn't isn't easy and for those of you who are young enough and sturdy enough start programming because that answers the questions that you will be able to ask and get answers for are going to be found in big data sets and phages make good use of of that data and by all means try to do that um i think i might have a couple minutes for questions if you're uh if you'd like to ask them definitely debbie thank you so much there's so many questions raised hands and before i take questions uh there are there's there there's requests from several people if we can share the slides with them um yes i gave them to you will you share them is that okay okay i'll do that wonderful so that's a good news for all of you jessica you want to take the question i mean you want to ask the questions sure yep um are we going a little bit after the hour or what what's our time frame just before we get going you can go right outside right now sorry what's that we can go right now right away but how many how much time should we spend i think there are about uh 10 questions as they can see so hands are raised so baby they are very eager to talk so maybe we can have them speak first nitish wants to uh sorry okay you should be able to yeah hello yeah am i audible yeah so i have very uh basic question about what degree of exchanges in the genomes that you you find between fudges and their hosts and the follow-up question is um when you pair up fargis and host do you see any kind of a relationship with respect to their environment um i'm not sure exactly what you're asking and i'm not sure i heard the first question could you tell me the first question again sorry yeah uh so the first question is uh what degree of ex changes in the genome uh do you happen to see between fudges and the host um so um that's hard to say right because is um it's not easy to say um the place where you kind of see it is where um in the area around the integrase and so where the phage attaches to the host it can sometimes drag things that are very foreign and um not easy to understand from the host um it are the trnas i i don't know how to answer most of the question in that um is a polymerase from the host because the phage doesn't need a polymerase so it had to have come from the host um why does a phage bring its own polymerase um i don't know that answer either because it doesn't technically need it but does it make it more advantageous to the host no to the phage maybe being able to get in and out of a host quickly is you know maybe it's a time i have no idea um what you also see is that sometimes you find a phage on a on a host that it's the first time i don't know that is exactly the first time but it hasn't been in that host very long so it doesn't match anything about the host that we know a lot of um the same kinds of genes are consistently found even if we don't know what they are we can see them over and over again um and so you're asking me a question that would be great to be able to answer but i don't have an answer yeah i understand that yeah thank you there you should be able to okay i'm gonna go with rodrigo then next but nima if you if you want to after him you can so bear with us while we work on the special security measures i'm clicking ask to unmute but rodrigo yep there you go hello deborah thank you for the talk sure i was just wondering um the program can it identify phage satellites for example um well it would depend on what you gave it right um um no i i don't think um [Music] you would have to suspect strongly that you have satellites and then once you put it into dna master it would help you to easily identify what the genes are and then you would go source out what what genes are there and you would have to come to that conclusion um when you when you have a big data set like a bacterial genome or whatever comes in a whole genome shotgun sequencing project um you're still your best bet for finding phage-like items are the the programs like faster and things like that um which are clumsy and difficult and um we're working on that's actually one of the biggest problems at the moment because we know and we have sequences of phase satellites and sometimes we try to obviously identify what the coding regions are doing because these will carry other things some of them are picked up you know as ncbi and the other platforms but when you try to run them through faster for example faster will simply not give you an output because i think there's like a limitation with like the recognition of like conserved regions for example so normally they don't carry capsids and therefore faster doesn't give me anything despite of me already knowing this is clearly of age satellite because we've worked with it blah blah blah right no it so um one of the programs that's in our repertoire is a program that basically collects um all the database information in one place so it will take your gene mark and glimmer output right so um dna master could help you refine what you think are genes but then getting it blasted and hh predid and does it have any membranes pro protein you know whatever whatever you're interested in um [Music] is how do you collect that all so that you're not running your 100 things in a hundred different programs right um and um and so i think that helps um with that but right now the only way to get at that is to probably hh pred and set up your databases that you can blast against yeah we do all the all the stuff manually and right and it feels very manual and very difficult um right we need someone with a lot of money and a big computing system to start housing this data that is the most prolific data um of the planet what um kinds of phages are you looking at like what's the host where where are you at we normally look at e coli staphylococcus listeria enterococcus because we know there that you know we've studied these phage satellites there right but it's clearly what we've seen from other data sets is that there's more out there like many things have been recognized as you know defective prophages and these type of things that also a lot of these phage satellites were initially thought as oh these are defective and it turns out it's not the case right right so much so much to do yeah but thank you you're welcome okay so matt would you like to ask a question namaste has thanked basically his thanking his written a message so he's just congratulating the organizers and thanking debbie very much for an excellent talk and several of them have are doing this so maybe now we can have questions from the text or dr sanji are you there yeah uh thank you dr umi uh thank you dr devorah uh it was very nice uh informative talk i just want your comment on your or or your recommendation regarding the smart experimental verification of your annotation particularly with respect to say proteogenomics approach to verify your annotation um so um are you asking about our quality control procedures is that what you're really asking like how do we know what we're doing yeah experimental verification if i want to check whether my annotation how good that is so what are the smart experimental protocols or experimental uh experiments you would recommend um for example proteogenomics is becoming popular nowadays for annotating genes i'm i don't really have good recommendations because we're at the beginning of this um so i don't really know and that's my simplest answer there's a lot of insight into that um at this stage of the game um we are trying to document what we're doing in a very um human way um in our one of our databases called see phages we have forums and we try to write specific things like that um as a resource but we have not connected that to any of our reporting so we have not tried to use a system that would document that we know this is a capsid from an experimental um position um we just we just aren't there yet yeah okay thank you sure okay so should we ask some questions from the chat i'll ask the first one um so okay aren't the genes in t4 like phages arranged by their temporal expression during the life cycle um i don't know enough to know the answer to that um because i really work in the actinobacteria phage um so um we reference t4 a lot but i don't i don't know enough about t4 to answer that question well sorry um let's go with another one have you tried nanopore sequencing in phages we do have a nanopore sequencer and um it is not getting us to the finished state that the illumina is so right now we're sticking with alumina but we are using the nanopore for our bacterial bacterial sequencing and that seems to be working well too right the whole sequencing arena keeps changing and i'm going to say no matter what we say today like look for what we're going to do tomorrow which is the appropriate e-value for digging protein homology when using hh pred or other tools ah that's a that's a lovely question um i think we look at a 10 to the minus three to get some sort of answer um but that's out of context like there isn't a net like you won't find that number in our protocols because if you find something you know we can hit human proteins at ten to the minus three right um there are no human proteins in a phage there just aren't are there homologues that do something different for a phage that end up to be somehow related to it sure um the one that comes to mind is um every once in a while we hit a von the willow brands factor but it it's because it has a dna binding sequence that that is a homologous but it's not a clotting factor it has no clotting ability so um you just have to be careful of of it if you use hh pred to assign function i don't really use the e value when i'm looking at hh pred hh pred actually points you to their percent identity and how similar it is and we use a cutoff there of 90 again it has to make sense for the phage to do that and the other part of the signing function is definitely how much identity you have across or similarity you have across the whole thing you give it and there are times where it doesn't have to have a whole long significance a good example would be that it has some sort of peptidoglycan domain and it's the biggest gene in the genome and it's where the tape measure should be it's the tape measure even though it's the first time we've ever seen that particular kind of tape measure um and so one of the things i hope you'll all take away from this is if we could throw all this data in a in a computer and set up parameters and say if it if it hits you know um a 90 alignment with a 90 hit give it that function i'm pretty sure um we would do that but you can't like it just has to all be done in context and you know these pages are so diverse um it's all good stuff and the challenge is there and it's a puzzle and it's a really fun puzzle awesome um do we know what setup is used by dna master to execute june mark and glimmer in the background um so this is out of my area of expertise i just know the person who has it set up um it is set up the uh um dna master has to establish the link and then it's according to like glimmer and jean mark's downloadable executables right and it literally is there's a link to get to them um and i honestly don't know if you're going to do it independently um i assume that's an email to jeffrey lawrence and saying hey i want to use dna master you know this is the link to where my glimmer is this is where the link where my gene mark is and um i you know there there are windows that'll let you point to wherever it needs to be pointed to i'm not even sure if you'd need to ask him but that's what i would do cool how can one transform a fasta file to emdl file to make it easily readable with artemis software so dna master is excellent at um changing its formats um and so you start um with a fasta file but then dna master makes its own um database file called it i always call it a dylexic dyslexic dam file because it's dnam but i really read the word damn every time sorry um but it will um it can reformat into things i'm pretty sure there's an embl file but there's definitely fasta files and other files that um can go into other programs he's made it so that it will um you can take the data out just as easily as put it in right um can we find micro bacterial sage promoters with any tool available um sure so there's wonderful search functions in dna masters one of its um not introductory points that i just couldn't get to in an hour but there's a search function if you know what your promoter sequence that you want to point to it will search the whole genomes for those sequences it will look for promoters and score them um exactly how you know how far you want to believe it will depend on your usage of it i think and how you know when you compare that to things you really do know you'll have to establish your cutoffs of what scores work but yeah you can find your you can find you can easily find promoters in a sequence that you give it um jeffrey claims that you can import all of ncbi into dna master if you have a computer big enough to hold the data and it has what it calls a genome manager and will actually hold all the files and let you do a lot of comparative work so there's lots of things that can be done there's a scan function that graham is actually fond of that if you know a sequence that you really want to find it's better at finding it than most mostly for the simple fact is you can put it in the forward direction and it will find it in both directions which you know a lot of times we get tripped up because we're looking for something in a sequence that's really going in the opposite direction as us and that can be painful yeah it really is a really good program it's just um following it's not falling apart it's working fine you just have to be very careful know the limitations and you have to have a good computer to work with it and more than likely the newer versions of windows are going to become less and less compatible so you know us mac users we use it on a virtual machine but pc folks are also going to end up installing a virtual machine on your pc so that you can run an old operating system we say we're we're pointed to getting this updated um and the last time i checked with each of us all only having two hands it's it's it has to happen one day this isn't gonna work but right now we're holding on as best we can out of the six reading frame translations how do you know which is the correct one um so that's where coding potential comes in is that those programs use math right and they look at the codon usage i don't want to say codon uses the nucleotide patterns and figure out which ones are most likely based on the biggest open reading frames so the general concept is is that if you have a really big open reading frame i'll bet right now we probably know its function okay but even if we didn't it can't stay a big open reading frame it will degenerate if it's not used by the phage right and if it degenerates it will become smaller and smaller and smaller and be useless but if it stays big over time it has to be something that the phage makes and so it uses that to say if you are made like that things made then that's the simplest most efficient way for a phage to make all the things it makes does that make sense and so um that's basically the concept behind the math that both glimmer and gene mark use and then they also use um start start sequences and other things to refine it and that data comes back in less than a minute and you go i don't know how they came to it but it it truly is a math on a sample that includes the biggest dwarfs and what did you mean when you said gene mark is trained on stealth versus host okay so after it um looks at that sampling and has a pattern what does it compare it to so it can either compare it to itself right the biggest genes in the open reading frame or it could compare it to those four nucleotide patterns to a host and so it's comparing those patterns to what you know in the case of the mycobacteria to what smeg does or like more interestingly to what tb uses and does do those patterns fit that particular open reading frame and could that likely be a gene and when you do that you'll find that some of your things actually go to function so that gives you more um confidence that the bias that this introduces is legitimate um and it breaks down the smaller the gene gets and so then you have to evaluate and that evaluation is why we do what we do because if we could program it and say they all have to follow the same degree of bias i'm pretty sure programmers could write that code the problem is there's these 15 factors and they're not equally weighted across all the decision making and so you have to evaluate that and you have to keep yourself in check that you're not introducing any more bias or that as a program with thousands of annotators out there because that's what see phages is that we're all looking at stuff from the same bias and if we choose to change our bias we do what scientists do we document it and we make the claim provide our data and tomorrow we could change what we do because the date is better and i think that's part of why i love this whole arena is the concept of right and wrong is absent it's just the best you can do on any given day and the more you know the better you are at it and um when you were a child didn't you have like reading books where you didn't have to read a certain thing it was like you read at your own pace that's sort of what all this is like it's what science is it's how good you are and how far you want to go with it and what makes sense and if you want to start making claims you have to be able to support it and if you can support it then as a scientist you write it down present it to other scientists and they're either going to tell you you're full of bananas or this is sound scientific um protocol and we'll we'll advance the field by doing that and so we're at the beginning of it even though there's thousands of phages there's still how many more we haven't touched and we don't know and it makes it so exciting and i know how to say oh i see what you're saying i can now do that better i was going to say as a woman i could apologize for not doing as well as i could a minute ago but i'm trying to teach myself not to apologize for things like that i didn't know i did wrong is that good we're getting all the life lessons all in one all that fun right oh there's so many good lessons about life in phages um it's true another of my favorite i'm gonna give you one more is when students like find their phages you know they find a plaque that's so exciting when they find their first plaques and then they like purify and they amplify it but it's sort of like pregnancy right that when they go to electron microscope and see the actual particle for the first time they know a lot about the phage they've seen it they've seen how they've been tortured by it and all kinds of things but it's really sweet when they see that particle for the first time sort of like birth just this is a little bit and naming it is just like naming a kid you know like um so there's phages are life let's just leave it at that and we're happy love it army i see you've reappeared did you want to say something or yeah so i mean i think it's they're close to we have several several more questions so how long can we go on now uh debbie you have to tell us how much time do we have um how about we stop at 9 30. so a couple more a couple more minutes and um yeah the weather's really bad here so i have to sit tight so i'm i'm stuck so that's good then in that case we can take jessica i think you can go ahead okay i'll ask a few more um okay why are genes of phages in the opposite directions and what's the benefit of that um you have to ask the phage i think is the quick answer to that um because dna is dna i don't that there's no good answer um it's how the dna stacks up and where it goes um you probably it's probably useless to um for a for a phage genome to try to make phage particles if it can't really you know first it has to get in there and take over the genome um it has to do battle with the bacteria and so there's going to be promoters and ways for it to get in there and work in the genome and you have to remember the phage doesn't care whether it's forward or backwards like it just doesn't matter um it's just what we've seen there's a lot of patterns that we see and a very common one is all of um that there's two promoters actually at both ends of the genome and one's all the forward genes with all the structure and one's all the takeover genes and then those transcripts would meet in the middle um i don't know i've seen genomes that are all forward it is unlikely i will see a genome that's all reverse because by convention we put the terminus and the forward gene or the structural genes in the forward direction because they were the first studied and they always like studied them that way so the convention is know that there's a lot of people out there that send their sequencing to a service and then whatever comes back they accept in phage genomics that's not a good idea you should always orient your phage to its ends if it has them and if it doesn't we pick an arbitrary start as close to the terminus as we can considering that most terminases have a small terminate subunit ahead of them and then we try really hard not to cut a genome so that there's a gene that crosses the ends we're not always successful mostly because after we cut it we learned something um but genbank doesn't accept genes that cross the ends very well and so you like argue with them a little bit more than you have to so we look for a gap upstream of terminace for those genomes that are circularly permuted and it doesn't work if there's an overhang and so there's really defined ends and there's a gene that crosses the defined ends like it is that is what it is um but those are the kind of the things that we think about when we're trying to orient them there's nothing there's no harm in um to the phage if you orient your phage sequence like unconventionally like you'll get to all the right genes but if you want to compare it and compare it easily so that your brain can process the data if easily um follow the conventions um and i really did just review a paper recently where the sequencer they they purchased their sequence and the sequencer actually assembled it incorrectly so there was a gene that was in the right orientation and then the next gene was the next part of the dna was all backwards and i'm like if i wouldn't have had comparative genomics i might not have picked that up and it isn't that it couldn't happen but very unlikely that it did and so whatever sequence you look at you have to look at it holistically and we tend to drill down and try to look at genes independently and we always have to remember to pull back and get the context of the genome we just submitted a paper it's almost ready to come out where we've looked at the prophages of the obsessive phages that we have been working with and really looked at the attachment sites of these of these processes and where they go and how these phages are a whole new subset of mycobacteriophages like they're really different and um what are they doing to contribute to the pathogenicity and those are awesome questions out there and there's a lot of bacteria sequenced and if we could find [Music] good software that could find the prophages in the bacterial genomes that are already sequenced um we would get another good look at another whole component of of the phages that are in our world wow awesome thank you so much maybe one last little question because it came up in a couple of versions um how do you use dna master or formulator of pages that are not included in a page database um so the phage database really is the glue that holds us all together um famirator i'm gonna start with dna master actually has a genome manager so you could actually manage your a small sub you can manage the whole thing um in dna master um and then you would see where you want a database to be built around it famirator is available to any subset of phages that anyone wants to do you just go to famrader.org and there's a contact form there and tell them what you want and he'll actually get you started of your own database there's no point at this stage of the game to start putting you know mixing and matching the phages because the data sets are so big and um and wild um and they're not as comparable as fade should be a phage but they're all so different and it's so cool um so there's so dna masters portable it's whatever you do with it famirator's available to you i think every software program we've written for what we do is actually is freeware um but you know it's all written by lots of different people managed and you just have to have the right managers and and how to get that data going um um and so that that's the trick right is yeah we still need to be able to do this um you know in our own house kind of thing got it well that's we're 9 30 or 6 30 depending on where everyone is but thank you so much debra and yeah we're so happy to have you here for our first of this new series and we're going to make your materials available so join our slack channel page.directory slack we have a page bioinformatics channel set up in there and deborah i don't know if you're a slack person but you can also join and people might have more questions there or i can also share the other questions with you the chat log i usually share with speakers so okay great that would be great yeah um jeremy did you want to say anything else i want to thank debbie a lot very much for the presentation for sharing your slides and i hope that the questions which are not taken here because of the lack of time uh so then the community can discuss among themselves and if after that also if some questions remain answered unanswered then maybe we can send your way would that be okay sure i'd be happy i'm willing to answer any questions that i can and um it was a pleasure to be able to share what we do here yeah that is exciting wonderful very very nice thank you so much and as you can see on the chat a lot of people are thanking you profusely for a wonderful job so yeah very well thank you thank you for the opportunity and good luck with your series thank you again [Music] so are you okay yeah so i think we can close go ahead and thank all the participants also for excellent questions and for being here a very good attendance we had and so yeah thank you all very much for joining bye bye
Up Next

Photobioreactors Explained: Types, Design, and Mechanisms
@anasaldailami1464
767 views•2021-02-02

Algae Biofuels: Harnessing Microalgae for Renewable Energy
@LosAlamosNationalLab
623 views•2020-12-03

Microbial Degradation of Plastics: Biodegradation Pathways & Sustainability
@majeedhammad
2.9K views•2021-04-11

CRISPR and Genetic Engineering: How Gene Editing Works and Why It Matters
@kurzgesagt
30.5M views•2016-08-10
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biotechnology







































