Machine learning models can significantly improve CRISPR-Cas9 guide RNA design by predicting on-target efficiency (Azimuth model) and off-target effects (Elevation model). The Azimuth model uses boosted regression trees with features including position-dependent nucleotides and thermodynamic properties to predict guide effectiveness for gene knockout. The Elevation model addresses the combinatorial explosion of potential off-target sites by first creating a shortlist using heuristics, then applying a two-layer model that combines single-mismatch predictions with a second layer to account for multiple mismatches. These models outperform previous methods like CFD and have been validated through genome-wide pre-computation for the human genome, enabling researchers to select optimal guides by balancing on-target efficiency against off-target risk.
CRISPR Guide Design with Machine Learning | Jennifer Listgarten
Added:thanks Alex actually I'd say it started almost three three and a half years ago and actually it's so the work I'm gonna talk about is in two papers one of which came out in January 2016 and the other one was just accepted after five revisions of hell so alright so and actually I'm about to start a new chapter in my life in academia in January after more than a decade at Microsoft Research and I can't tell you where just yet but hopefully very soon oh you are allowed to say it's not wear it because I yes so I think I probably don't need to explain the intro to CRISPR but I feel like I should just to make sure I cover the basis so I'm sure everybody here knows what gene editing is and that there's a long history and that it's basically like cutting and pasting and these kinds of things right and so you know a few years ago this kind of thing was very very expensive and very very laborious and it wasn't until 2012 the CRISPR came to make these things much cheaper and much faster and more scalable and now people are doing all kinds of things with CRISPR but there's two sort of core tasks that people started to do with CRISPR and probably are the main functions of this technology right now and so one is sort of a classical edit where you change an A to a T and the other one is gene knockout and because all this work is joint work with John bench and many other people but john dunn chose in the screening platform and he is very interested in knockout and so the focus of all this work is knockout but of course the key task to both of these the key thing to both of these things is precisely cutting the DNA and so the modeling we've done has for the most part focused on that and so it stands to reason that much of what we've done is applicable for this first task but we haven't really established that right now so right so you know this CRISPR has sort of slowly made its way you know well first probably before the top-tier journals other journals and then slowly into the popular media New York Times CNN and then who saw this cover piece unwired from a few years ago this is everything that's written on this cover is actually in reference to crisper so apparently crisper means no hunger no pollution no disease and the end of life as we know it it's the genesis engine so that's you know I guess that speaks to the hype that is a little bit around crisper at the moment and so there's also some backlash as we see this this is now pretty old I haven't followed it I don't know if it's new but that was a direct causal result of the cover Wired magazine so but the reality as I'm sure everybody knows here is that there are already a lot of very promising proof of principle for therapeutic applications and it's already routinely used in the lab force drug screens and all kinds of other things so it is hyped but it's also obviously very very useful at the same time and so there are - there are two main problems and I'm gonna sort of abstract away what a biologist would say to a computer scientist mind so in my mind there's two main problems which are the on target and the off target effects and so again I'm gonna focus this talk on gene knockin so remember the goal is that for some gene let's say this one right here I want to prevent it from functioning and so the way CRISPR works to do that is you basically throw in these scissors which are the cast 9 for example and it's gonna make a cut to the double-stranded DNA and then it will repair itself if I don't give it a template it'll just kick in an innate repair mechanism will try to repair itself with but with some reasonable probability that repair will fail and that's what will cause the gene to not function anymore and so CRISPR has a few constraints of where you can apply these scissors and those that's determined by cast 9 or something like cast 9 which does the cutting and so cast 9 for example needs on one of the strands to have a GG and that's the Pam site and so basically it recognizes a pam site you can imagine GG is promiscuous and so you can cut in many many places and so now imagine like here I've just shown four places that I could in principle cut but there's actually hundreds of places for a single gene and it turns out that some of those work well to knock out the gene and some of them don't and in some genes it might be that many of them work well and in other genes it may be that very few of them work well and so already like now day to day people are trying to do these gene knockouts and they're not sure which thing to apply all they know is they cast nine needs a GG there but they don't know like which of these and so this is as you can imagine we can use machine learning to improve this kind of tasks and the flip side is the off-target effect so now imagine i've selected this particular site in this gene but i have the whole rest of the genome so what i haven't explained yet is it needs that GG yes yo see what's that there's kind of two ways that i could see that happening one is that the supervised signal is gonna be the result of an acid and that assay is a functional assay which tells you the genes not functioning so it requires both the cut and the lack of function but we in the off target stuff we sometimes model just the cutting so it kind of depends what's available to us as data people and from what we can see often these are highly correlated and that a model train on one works well on another and i'll get into a little bit about that later yeah and so for the off target though beyond the ng geom as I'm gonna explain there's a guide RNA that also needs to match like where you design you say you wanted to go here then you need to have an end eg here but then you also design something that's complementary to that to the DNA there and it may be somewhat complementary also elsewhere in the genome and it may have the capacity to also cut there and potentially disrupt and so these are the off target and obviously we want to ideally eliminate these or you know more practically speaking reduce them and also maybe understand where they occur because maybe some are worse than others so these are the two problems that I'm going to talk about are they clear to everyone okay so right so people will tackle this obviously in the laboratory improving the molecular biology and all kinds of things like that that I'm not going to talk about I'm a machine learning person so I'm going to tell you how we can use machine learning to also help us for these problems so I just really usually give it I mean probably everybody at the broad knows this history that it came from a yogurt company was the one who figured out what CRISPR actually does does does everyone here who put up your hand if you already know the company oh okay very few so okay super interesting right so for decades people sort of poorly funded biologists looking at bacteria and things like this we're studying CRISPR but no one knew what it was and that's why it was sort of obscure low-impact and poorly funded and Danone a yogurt company which obviously has and researchers at that company obviously discovered what it was and they did because it had a real impact on their bottom line and the impact was that they discovered it was associated with a naturally occurring immune defense mechanism in bacteria and that had practical consequences for them because some yogurts and their vats were being affected being infected and others weren't and so they had to throw out the ones that were infected basically and so they would pursue this and they're the ones who actually discovered that CRISPR is part of this bacterial immune defense mechanism and I'll explain that in just a second and then you can see a pretty short time later you know and I'm of course there's much other work that led up to these developments but in 2012 people realized how to take the core machinery of this natural defense mechanism and use it for general purpose gene editing and of course that has many implications for all kinds of things that humanity cares about so now let me go to the actually the system that Danone had figured or the researchers at Danone had figured out and so this is a bacteria and if it gets infected by a virus this virus has this RNA here and if the bacteria survives it what it will do or what it can do is it can take a cut out part of the viral RNA and it can put it in its own DNA down here and it sort of takes a snapshot so you can think of that as like a viral memory that it's storing in its own DNA and then if it gets infected again it's now kind of primed because it now has the capacity to recognize these things very quickly much as yours and my adaptive immune response can be primed by a vaccine I should I just had a vaccine this morning to recognize things very quickly and it has innate these sort of scissors in the system which do the cutting and so these scissors are kind of you know floating around and they are primed to recognize and the next time it gets infected it will cut it and it has therefore has the potential to fight off this virus so that's how it's a naturally occurring immune mechanism but as you can imagine there's a something here where you have a precise sort of thing that guides it to a particular piece of RNA and it makes a cut and so of course from everything I've just told you it's no surprise that people or maybe it as a big surprise in retrospect it's no surprise that people figured out how to use this for gene editing and so right now the system is in fact you can sort of take these bits and pieces and put them in almost any organism and they will still function and so what I used to call a sort of viral memory and one thing I forgot to mention is is how does the how does the bacteria sort of know that this region here is storing these memories and how does it delineate the endpoints so what it does is it puts little patterns here does anyone know what these patterns are these are the clustered regularly interspaced palindromic repeats so that's actually where the name comes from so there are sort of the delineation of these things in the native system right and so but now what we can do is we can say let's take some like maybe this is a human genome here and I know I want to either edit it or disabled in this region of the genome so what I do is instead of like waiting for bacteria to be infected and to fine part of a virus but it is I just synthetically generate this thing down here and we call that that's the guide RNA and we make it complimentary to the region in the genome that we want to cut and again we have the constraint that that needs to be if this is caste 9 there needs to also be a G G here otherwise even with the complementarity the caste 9 just won't work it won't go in there and it won't do its thing all right so does that people understand how the basic system works is there's a guide RNA it's about 20 to 30 nucleotides long and you throw in and some scissors and then it goes to the part of the gene in an ideal world goes to exactly where you want and cuts only there and does what you want of course it's not an ideal world and so right so now I'm going to tell you how we use machine learning to tackle the two problems I told you about in the beginning so one is to how do we get better on target efficiency using machine learning and we've dubbed this model as a move and then afterwards I'll tell you how we've reduced the off target effects which we call elevation I don't know so this is joint work with Niccolo who was speaking here last week and he I think he he liked these names that's how these names came to be right so what's the basic setup this is not actually super fancy machine learning I would say the off-target problems much harder the on target problems a little bit straightforward but just the due diligence we did we spend months and months making sure that we were doing things really robustly and solidly and so in any case the basic set up is as follows we it's a supervised machine learning problem which means we need some inputs and we need some examples of how well those inputs worked and then we're gonna try to build a model f of X which captures the generalities of this input output and then can be applied to new examples so that's just the basic setup of supervised machine learning and so the RNA guide here is going to be in its original formal to be like we as 30 nucleotides and then we have some measure that I'll explain that tells you how effective it was at knocking out the gene or not right and so again this is I think in the seminar series probably everybody here knows this kind of thing by now but oh thank you no way you can but you need not so the training that I'll talk about so yeah so this can be ya know you can think of this as anywhere you want and what we want is we want as much training data like this as possible and if we have more than one gene we're gonna have a much better training data set yeah there are issues that I might go into a little bit about dealing with the assay measurements and how they may or may not be consistent across different genes and how you deal with that but that's not a machine learning problem that's like an assay problem issue yes yes yes yes sorry that's the if that's the yes I didn't understand that was your exact question it's any guide and any any amount of effectiveness that's right they keep in mind that people design godlike so the training data is with guides that were designed with the hopes they would work and so people don't design promiscuous guides for example so in general these guides don't have a perfect match anywhere else in the genome so right so it's not like we're picking random guides in the training data and I think that might make it a bit weird yeah right so right so we have these we're gonna have to somehow take that 30 nucleotide sequence and convert it into features and we're gonna need some measure of effectiveness and then we're gonna have to pick some sort of model class and I mean these are just the core components of anything in machine learning and so right the I'd say luckily we have some very smart collaborators of the broad who know how to do this and so and what's so what this is is how did we get this supervised data so that's how did we get these labels over here for how effective each example was and so we did it in to meet John and and these guys at the Broad did it in two ways so in the first example of data we had this one I think was seven or eight they're each about seven or eight genes for a total of fifteen genes and flow cytometry was only cell surface proteins so proteins that are expressed on the surface of the cell because they're expressed on the surface of the cell you can attach fluorescence markers you can do sorting on them and things like this and so because of that ability you can then using ingenuity of biologists tease apart which ones got knocked out versus which ones did not get knocked out and back to yo C's question this is not cutting this is actually like a functional did like you know is that protein expressed functionally or not such that it actually gets tagged with the fluorescence and you can see it and the other assay is so and keep in mind that's very specific to just cell surface protein so you can't use that in general and the other one is also very specific to a different class of genes those are genes for which it's known that they have a particular kind of drug resistance and so I'm not going to go into the names of the drugs and all these things but the concept here is that it's known that for some genes when you apply a drug that if that gene is knocked out it will survive but not otherwise sorry into a cell so on applied to the cell line if the gene has been knocked out the cell will survive but not otherwise so again just from this prior knowledge of the drug sensitivity you can do a read out based on did the cell survive or not and get a sense of how much it was knocked out and so I'm emphasizing that these are very specialized because for a couple of reasons so one is why we're using machine learning if you can just measure these things in the lab well I think as most people here know like yeah we can measure them in a lab but only for like you know this is months and months of work to get this for 15 genes and a lot of money and of course we want to be able to do this genome-wide and so it's a scale problem as we can't easily do this for the whole genome and more importantly as I I don't know that all the developments but you can see these are very specialized so even if we had the time and money you can't even apply these asses genome-wide anyway so that's how we measure it and then the last right so now let me go a little bit into the future ization so there's a rich history in computational biology with machine learning on how to go from nucleotides to numeric values a lot of it came from the support vector machine community which actually when I was a grad student they were as hyped up as deep neural networks all right now and so they and you can so in support vector machines I guess the nuances here aren't important but they build kernels which have implicit feature spaces and you can back out of those and just get explicit feature spaces which is what we did so the idea is to have a slide yes so just a simple example would be make it say we want to know just is there an A in position 1 or a-team as you want to see like which letter is in position 1 then we can basically invert each letter this is a little standard one-hot encoding so you get there's 4 letters to each letter gets four digits and one of them is on and so now you can just encode in the first position this one had a T and the second position is the next four dishes it's had a G and so on so you get this really long string of 0 & 1 which is now suitable for any machine learning model that you'd like to apply and you can also do this for dinucleotides and you can the one this as I just showed you as I'm concatenated them it according to their position but you can also do things like count how many there are and you can do interactions of them and things like this so this is sort of one of the main features that provides a lot of the predictive power and so you can actually see here so some key you know people had been doing a little bit of machine learning before we started this and and they were using those features and some some of those features that you can see if you look at the future importance is here like those ones that tell you the position dependent nucleotide either order two means two nucleotides or order one these come out as actually the most important thing for this problem and you can see that when you add these other kinds of features that hadn't been there the one I guess actually those are those starred ones right people hadn't been using those than you boost the performance up on these various datasets so FC is that flow cytometry res is the drug resistance and this is their combination and so one thing that was kind of neat is John had and I guess other people had this hunch that the thermodynamics should play a role and you can see in this video actually Howcast 9 comes in and why that might be so cass 9 what it does is it goes in and actually it has to find the pam and then it has to pull apart the double-stranded DNA before it makes a cut and because it has to pull it apart that obviously requires energy and the energy it requires has to do with the thermodynamics of the bonds there and so if you can model that it's just you know it it's the reason that we could do better and in fact you can and we did a few things so john was aware of a crystallography paper that looked at more precisely how things were happening and said you know i think if you break it like we did the thermodynamics of the whole guide we also did it of parts of the guide according to those crystallography experiments and that also seemed to be useful right and so now in terms of modeling again like so this is just standard off-the-shelf stuff here we used yep what does it mean to turn thing is the thermodynamics you have some function that goes from the sequence to some notion of how much energy is required and separation a nonlinear function that someone else is learning yes it would be hard I do need features just throw it in a deep neural network right so we don't have anywhere near the amount of data yeah I know you know that but just to be clear for everyone that in principle you could probably learn this if you had enough data but we're like you know we tried deep neural networks just for kicks and you know if we'd been really like attention-seeking we would have just used them they work just as well and then it would have been on the front cover of science a deep learning Plus CRISPR right which you know has probably come since and and I get asked to review these kinds of things and they're usually crap and not needed but I did not mention any names that I know this is recorded you can't do this is not enough data we did this I think that it's starting to change and there may be enough data in some domain so there's many flavors of the CRISPR system also and things like this so I have seen some work that looks promising that's not public yet so I don't it's not all crop but there's certainly a large attempt by everybody to use deep learning for whatever without doing due diligence that it's actually needed and so it didn't in this case it didn't help any try to design more high-throughput assays or yes oh no so now I've explained how we we have for 15 genes we get these measurements of how the well the knock out works so you get some sort of efficiency like it's not just binary it's a sort of on average this is how much it knock things out and we're gonna use regression trees which probably after 40 of these mi a talk someone has explained regression trees but if not there's sort of a very nice model in the sense that they're they can be highly nonlinear and and you can control the complexity well enough pretty easily such that even if you don't have a lot of data you can still use them and so you basically pick one it picks one feature at a time and decides to split on it if it's a real value it figures out where to split if it's a binary or categorical it just chooses that and it basically in a greedy way kind of constructs a tree like this which has the effective if you had only two features of sort of partitioning the space and then within each part of the space it now can make like a regression prediction that could just be the average of the examples that fall on the leaf there or it could be building like a simple linear regression model with those examples that fall into that leaf question the negative examples or is there are so will be only if we were doing classification the word so I'd say one thing is that everybody in this field likes to do classification even though all the asses are real valued and they're throwing out a ton of information and we've shown that in two of our publications that you throw out a lot of information by doing that so if you don't do laying vacation there's an example there's this notion but in regression it's just it's a real-valued number so there's no I mean the low ones work poorly the high ones work well or maybe it's the other way around I don't know each assay is different but the good candidates to begin with so there is a bias in the training data set in that we're not sampling randomly at uniform from the input space those are designed by people who say I know I want it to be unique like in like an odd match elsewhere in things like this yeah that's right and so so that's how regression tree works and we actually use boosted regression trees and so again I don't know if people have gone over boosting in this but posting is really cool there's a very solid theoretical foundation behind it which I won't go into but there's also a very nice intuition and so what boosting does it takes like some sort of pretty weak model in our case we're gonna sort of force the regression trees to be pretty simple by keeping them very short mostly speaking so if you keep them short I can't use very many features and it can't really over fit and so when you have these sort of very low capacity models we call these weak learners and you could use a different kind of weak learner or anything that's just sort of not a super rich model and what boosting is intended to do is to start with a weak learner and then to add to keep relearning on a slightly tweaked version of the training data set each time it reweighed s-- the examples that effectively reweighed s-- them and learns another model and at the end what you have is you have a whole bunch a whole ensemble of this weak model in this case the weak model is the regression tree and then it gives you a rule on how to combine them so that's sort of the abstract concept and actually to give a bit more intuition but it actually does is and what you do is you take your original dany the original training data set and you fit in this case the regression tree like one that you force to be pretty small and then you look at how well it did on each example in sample not out-of-sample so the whole training data and then what you do is for those examples that it did badly on you up wake them proportional to how badly it did and then you learn the next one and so that's how you keep and you can keep going and then you wait each of those models by how well it did overall on the accuracy on that training data set and so I mean that's I think there's some intuition here that sounds very pleasing but as I said you like if you go and read about this there's a whole set of theoretical reasons that this is also very sensible and there's interpretations of of what of this model also in different ways that are quite interesting but I won't go and do here I just want to at least give you that intuition and that in general this can be a very nice way to get a very powerful model and it turns out that it is it was for us the best model and so actually what you can see here is on these data sets so when it says train on FC test on FC this is obviously with cross validation we're never testing on what we trained actually in all of these it's cross validation and in particular the flavor of cross validation that we do is so right across the ocean you normally like chunk up your data and just say ten folds you hold one fold out train on the remaining nine tests on the last one and here we're actually gonna make sure we do it by gene meaning that we're gonna test one gene will be one holdout and we're gonna train on the other ones and the reason we want specifically to do it that way is it sort of mimics the use case in the real world right like chances are people are not going to be using any of the genes we had in our training data they're gonna be using some totally different gene that we had no data on and so it's sort of the closest we can get to the use case and if you do it the other way you actually can get a higher accuracy because it sometimes it I don't know if we quantified that or not but anyway so john and his colleagues had just actually published a paper when i met him or we're in the middle of revision and this is their model and as the best it was the state of the art at the time it was a combination of SVM and logistic regression and that's the performance it has and then the model that we published and the models all the other models are actually just our own models that we tried out of curiosity so they're shaded a bit lighter because they're not sort of competing models but and you can see the the boosted regression tree is what did well but you can see that actually super simple models do well here and that's just because one like there's very little data and probably the other reason is often I think especially in biology that linear additive signal is very powerful and so you can get away with them and so as so even though this is kind of simple machine learning as I said we spend months and months sort of really trying to assure ourselves that this was the best thing and we didn't over fit because we knew if this was we did a good job people would really use this it is really satisfying the year after this came out that the two different individual or sets of people actually showed this was the state-of-the-art on you know data we'd never seen and so a lot of people did start using this this tool that we put out and so right in fact it's actually been adopted by two startups and it gets used all the time it's available through the broad website it's also the source codes available and the reason we yeah I so i glossed over that and I'll go into a bit more detail with the off-target but there's a real problem with the assay readout is not super comparable between say like different even different genes let alone the different assays and so what we do and I've glossed over here is we and we followed actually what John did so which is to rank normalize within a gym and so when yakura saying is there anything gene specific here that's the only sense in which something is gene specific and it's not quite the right thing to do because what that means is that every gene now has the same scale but the reality is is one gene could have many things that work and another gene might not so this is not ideal and if there's time at the end of this talk I'll tell you we actually have like a separate machine learning project we've been working on in parallel to address this but it's just like a we didn't have time to do it with everything else here let me yeah let me I didn't answer your I got sidetracked with my own answer so the more direct answer to your first questions we do Spearman correlation because at the end of the day how people use this is they want to rank they want to say I want to knock down this gene can you give me a rank order of the guides I should try and that's what they that's the sort of most prominent use case there may be other ones but that's sort of the real thing that people want and so what you want to evaluate is the rank order actually and you don't care about the precise value and therefore you don't care about bias and in general in machine no one cares at all about playas it's the statisticians that go crazy with bias because there's a bias-variance tradeoff and most machine learning people like want like things that work and they know if they don't have infinite data they'd rather have a bias and have smaller variance so for all of you bias obsessed statisticians right so okay so that was the first to have for the talk and I think that gives you a sense of you know how you can use machine learning it was kind of largely off-the-shelf here but this the second part is is was we had to be a lot more creative and I'll tell you about why in a second yep the way it works is once we've learned this model if someone goes to the server and says I want to knock out this gene what it does is it goes into it gets the sequence for the gene it finds out everywhere you could apply CRISPR with that and kept with cast line for example which is NDG and that might be like 400 different places then it takes each of those 400 places and it pipes it through the model and gives the user a rank order yeah we we assume the sequences know that's right and so yeah and sometimes people say well does the model work if you know you have snips and these kinds of things and yes it does we haven't pre computed these kinds of things because that's impossible but there's no reason this model wouldn't work for any sort of snip issue as well to be pretty good yeah there any worry that the model would not know they were picked to be pretty good in the full so I don't know all the design that went behind at the one part of the design I know is they were picked to be pretty good in the sense they didn't have a lot of off target activity but I'm modeling on target activity and I'm not sure how much they could have even put in bias with respect to that because I don't think they knew very much so and also they really what they did is they statically covered the whole gene so I mean those genes were chosen for exactly was it and is what they big things that so the busine algorithm itself gives a recipe for combining them that's part of the boosting algorithm and it basically says add the predictions together in a way that's proportional to how accurate each of the trees was roughly speaking yeah right so this problem is much harder for a couple of reasons so this is the off-target problem and so now remember that so far everything I told you about there was a perfect complementarity by design right I had a guide and it was gonna match exactly to where I wanted to cut and I would only measure there but with the off-target accidents can happen right I can target somewhere but it might be that all along the genome I have an almost perfect match maybe there's one mismatch maybe there's two and it might still be active and I need to model all of these because I need to basically scan so right so that's that's why there are two problems I need to scan the whole genome just for one guide that I want to use whereas before four on target I just say how well did it work here but for off target I don't say how well they work here and say oh my god like everywhere else did it actually like caused some sort of disruption and I also need to account for the fact that there might be these mismatches and so you get this combinatorial explosion that you need to model and you have very limited data and so that those two things together maybe this is a very very hard problem and at the same time we couldn't use very sophisticated models because we have so few data and so right so in the end this is how we decided to break up the problem so the problem is gonna be that for a given guide we want to know the off target effects so we've decided to use this guide that's targeting this place and so the first thing we're gonna do is we're going sort of deuce make a shortlist across the genome of things we think might be active and so I mentioned that it could be active with some mismatches it turns out even with up to six mismatches it can be active and so I need a pretty like good shortlist right in a perfect world the shortlist would be the whole genome but I can't afford to do that so we do something a bit smarter and we could actually use machine learning here but we didn't because we decided to focus on machine learning for the second two things which I'll tell you about but and the first thing what we do is we use sort of heuristics that have been floating around in the community that had to do with how many mismatches so we know there's activity up to six but we also know for example that the distal three nucleotides are four nucleotides away from the pam site are not super as relevant and maybe it doesn't matter if those have mismatches so using some sort of ad hoc rules like this that the community had observed evidence for we build this shortlist and in a perfect world we want that shortlist to be as big as possible and once we have that shortlist then we use machine learning so then we we build a model or we use a model where we for everything in the shortlist we're gonna apply that model and say how likely was it that there was a disruption here that there was off target activity remember this is just for one guide so now I have like you know potentially thousands of places that I've got a score from my machine learning model but the way people tend to use these tools not exclusively but probably the main use case is how I described before is they put in a gene and they say tell me the best guide and the best guide means it has high on target and low off target but I'm so low off target like someone wants to just scan the guides I want to see a single number for off target activity but I just told you that like there's stuff potentially happening everywhere and if I stopped at two which is actually what everyone most people have been doing who started to work on this problem then you just have this huge list of thousands of numbers and a user can't easily rank the guides by these thousands of numbers so you want to actually aggregate those into something that's sensible and so one sensible thing might be sort of like what's the probability overall that I'm in to disrupt the cell or if you're interested in you know particular parts of the genome then you might do more nuanced things and we can't create a tool for everyone so we just basically tried to create an aggregate score that says what's the overall chance of having a disruption if we use this died over everything of course to do this well we need to do step two well and so that sort of been the focus of this work and so right how do we just to give you a sense also of the explosion here like if I when I'm building this short list it's a back-of-the-envelope computation but if I allow one mismatch then genome-wide I might get roughly a hundred things that will be in my short list if I have to you know and so on and so forth so this explodes really really quickly like obviously we can't afford and this is just for one guide right and and we want to do this kind of for all kinds of things and so again we use this sort of value heuristic which I grayed out here because usually I talk to people who are just in the computation but so you allow kind of different kinds of Pam's and in fact there are many pieces of software out there that we don't have machine learning in them but are designed just to basically do a heuristic search like this and we do something that's very much along the lines of a lot of those those other systems so right so now we have this short list and now I'll tell you how do we build a model and how do we evaluate this model that does the scoring so right so there had been a few previous approaches to this that were not based on machine learning and we compared to these the best of them was actually CFD which was developed by John bench and there and in his group and that did actually very well but you now do better than that so now the schematic setup for the machine learning here is slightly different than the on target case so in the on target case you may recall I had only the RNA a guide here and only the activity here and the reason I did is that it was redundant to include the target because by design it was perfectly complimentary but now I need to tolerate mismatches so the input needs to basically be the the guide and the target as well because of the mismatch problem you can see already that's gonna be a bit harder to model so CFD is what John had devised and John has this incredible intuition and I don't actually know how he came up with it but when we started to look at it we realized that it was very similar to something a machine learning we would call naive Bayes and naive Bayes classifier and so we it was actually useful for us to back interpret it in that way because we could see what assumptions he had made and then sort of generalize away from that in in a principled way and that's basically how we developed this part of the model was to do that so I don't know if I need to go into details but the key thing about naivebayes is it basically it's a classifier and what it will the main assumption that it makes is that the features so these are different features for one example is that conditioned on the class they're independent of each other and that's you know usually when you do predictive modeling that's not the case it's like a very strong assumption which is helpful when you have combinatorial explosion and things like this if it's a suitable assumption and kind of in these kinds of data we almost don't have a choice but to make this kind of assumption and we do something very similar to that right so right so we generalize away from some of the modeling assumptions there which is one we as I mentioned we move away from classification to regression and we can show that helps both on the off target and we also have that in the on target paper and I don't know why people keep insisting on like truncating their real values into zero ones and then building classification on it but it's something people really like to do the other thing is that implicitly in CFD these features here were very very simple let me just get in one second so the simples were basically the features that see if the implicitly sort of had were the following it just said a list of where there was a mismatch the position number and whether it went from like a tea to a C or an a doji or something like that okay so it's okay some kind of logistic regression explain sort of the concepts of how wide we did better where where it started from and yeah and why I got to where we were and so you can imagine first of all decoupling the position here from the letters and and also using like the context of the sequence and all and the thermodynamics and all kinds of things and so that is actually what we do we don't use just those features the other thing we can do is use very nonlinear regression model for this and so I think that's what you're asking alex is like what is the model we actually use for these things and then we actually use boosted regression trees there again and then the other thing we did that I haven't seen anyone do so this is I should say this is all very non-standard so this is all stuff we just devised for this problem which is to generalize away from the naive Bayes or from CFD doing these kinds of things and the last thing we did is I said there's this key assumption in naive Bayes and likewise in CFD which is that the features and so that could be for example where there is a mismatch at what position that their independent condition on the class and that's a very strong assumption and we wanted to be able to back off that assumption a little bit but we couldn't because there's very little data so what we did is we ended up sort of tacking on a second layer model that mitigates that assumption a little bit and I'll explain what I mean by that so right so now what we do is we need to go from this pair to some feature vector and we use a lot of the things that I mentioned for the on target and then the additional thing we need to add in is this notion of mismatches which we encode in a few different ways so sort of the positions where they occur the letters they occur and also the conjunction of those so right now how do we actually build a sensible model so I mentioned that there's this fundamental problem which is there's a combinatorial explosion issue because number of mismatches and we have super little data here like in the thousands kind of thing and you saw how those mismatches just made things explode in the genome and we can't get anywhere near enough data to model that kind of thing so what we do is basically we make in a sense this naive Bayes assumption which is that we can break it up per mismatch so and what we so the starting piece we're gonna use in this modeling is we're gonna build a model which is very good at single mismatch prediction if there's only one mismatch how well like or how badly is there an off target activity here and amazingly what is in our model is data from just one gene CD 33 and it's actually not DNA cutting but it's actually protein knockout and we from just the CD 33 every possible mismatch that's our training data and we build this model here and it's and it turns out that it generalizes quite well we don't know how much better we would do if we had other data because we have other data now the question is imagine I have that model but I want to now be able to use two or three Mis or evaluate two or three mismatched examples how can I sort of bootstrap into that so one is that I can use the sort of naive Bayes assumption is if this is some sort of probability and things are class conditionally independent I can just multiply them together and that's effectively what CFD was doing with classification and it's limited feature space but what we're gonna do is say well you know what we have a very small amount of supervised labels for multiple mismatch I don't remember now the numbers but we have some data there it's not enough that we can just build an arbitrary number of mismatched whatl but we can use it as a sort of tack on to mitigate this assumption here and that's basically what we do is we use a very little amount of data to say how should I combine these so one option is to multiply them together according to independence and another is is something more general than that and that's basically what we do and so again this first model this is the single mismatch model we use boosted regression trees and then with the very limited amount of multiple mismatch data so in this example there's two mismatches and I want to learn a function M then we used l1 penalize linear regression and that's like a super super simple model for so a linear regression then someone penalized which like clamps it down even further and that's because we have so few data here and so this actually tells you what happens in that second step is or how are we mitigating that assumption of Independence or not so this is the second model where we're combining the single mismatch predictions and the way we combine them is we basically have a weight in the linear regression for was the mismatch if there was a mismatch at one two three were all the way up to 20 and then the N in the Pam which is the wild-card before the GG and then the other thing we include is the product of those things so the product of those things would actually just come be the conditional attendance assumption so if we didn't want to use machine learning we would use just this feature here and you can see that feature plays a huge role and basically the importance of everything else is a way to mitigate that assumption so if this if this was the only feature that was on and the others were zero it would show that this model wasn't doing anything with these data the fact that it's not that shows that it is able to do something extra I mean that bears out in the results when you evaluate things that makes any questions on this oh so this is the product of so this would be correspond to the oh sorry oh so I yeah this was this is it's confusing because these are really close together so what is this is the number of mismatches I'm very mismatched is actually highly predictive it's not a bad model and it's funny because if you well yeah I'm not gonna say that okay and then this is the sum also which doesn't necessarily make sense I mean the product makes sense because we know that if we didn't want to use the second layer machine learning we would just use products so we basically want to make sure it's in there so that we can recover the sort of naive Bayes version of things and these are things we just thought might help and turned out to help yeah it's noise or signal they're zero I'm trying to think of this this has been such a long project I trying to remember if we dug that deep into this I I'm not sure for certain but my intuition is based on the size of these and from what I remember is that it's it's not just noise I suspect it's real but it may be that you could do pretty well without it yeah so right I guess this is hanging actually longer than I expected so how do we evaluate this maybe I won't go into all the details of the data that we use to evaluate it interest of time but I'll just say that there's something very difficult about how to evaluate this problem that's not the case for the on target and the difficulty lies in the fact that there's an a inherent asymmetry and the kinds of errors that you make so imagine you're you want to use this model and my model tells you that over here there's no off-target activity but really there is so that's like a really bad kind of error because I'm going to go use that guide and is gonna be an off target activity if this is you know therapeutic you could kill someone if it's just a screening thing you'll add a lot of noise and the flipside is that if my model says over here there is off target activity but there actually isn't then I'm just not gonna use that guide and so it just constrains me a bit and how I do my experiments but the result is not nearly as grievous and it's a-you know the trade-off between these two things is domain-specific so if again if I'm in therapeutics I better make sure about the off targets and if I'm doing a drug screening I may care less and may be able to tolerate more and so it's not we know there's an asymmetry but we don't know exactly where we lie and so what we're gonna do is instead evaluate things with this newly devised metric which we thought a lot about and so so what it is it's based on a Spearman correlation again and on the very right hand side and we the wait just for visualization it turns out to be easier to look compared at CC top as another model and we basically evaluate how much better everything is then CC top it just things are a bit clearer to see and on the right-hand side here we're gonna use an unweighted Spearman so this correlates so this basically says our new model elevation does this much better than the baseline according to an unweighted Spearman correlation which is the same measure I was using for the on target and what happens is as I walk towards the left on the very left hand side here what I do is I use a weighted Spearman where each example is weighted by the true activity and so what that means is if an example like a particular guide and off target is not active it falls out of the evaluation so that critic captures an extreme version of the asymmetry that I mentioned right it says that if it's not active I don't even care about it but the thing is you can turn this knob and gracefully go back to the other one and so we get everywhere in the spectrum between those two things and what's nice so there's a theoretical reason to see this but you see that one model almost always dominates another model and so the beauty is that for practical reasons it doesn't matter where you are in that trade-off because this model the red ones always better than the yellow and the yellow is already always better than the blue and this kind of thing and so that ended up being a really nice way for us to evaluate these things but me sir I don't know when these became such bad resolution but and so we do this on a few different data that's with cross-validation and things like this and so this is just very quickly you can see here if you use classification how much more poorly you do then if you start to use different versions of regression so in this this purple line is exactly the same model and features as this as this brow or what is that whatever orange e line except the only difference is we've moved two classifications you see there's a huge hit there right so now let me just quickly go to the aggregation so now what I've told you is for a given guide I can I'm gonna build a shortlist I'm gonna get some probability or some score of off target activity at everything on that short list but now I want to aggregate those into a single number using machine learning so and people had actually there I think there was one or two aggregate things out there but people haven't really evaluated them in some case some of the most popular servers were out there people are using them and knowing had literally ever evaluated the aggregate score so we wanted to also make sure to do that quite carefully so right so what we're gonna do is there's gonna be some particular intended target sequence and we're gonna build a short list of potential off targets and this could be thousands of things for each of those I'm going to apply the two layer model that I just told you about so that's it tells you how to combine each of the single mismatched scores into an overall score so this point for three for example says if this was my intended target and I design a guide for that what's the sort of probability that all disrupt this region over here its 0.43 and then I get that for everything in my shortlist and what I need to do is I want to that's going to give me some distribution of scores and remember for each guide there's gonna be a different number of these so they don't sort of map to each other and so I need to look at this as a distribution and I have to do machine learning and I need to sort of feature eyes the distribution and that's how I'm gonna do machine learning there and so we can the sort of features of this distribution we use to get a final score is we use like the order statistics the minimum the maximum and one thing that turns out to be very important unsurprisingly is whether it lies in a genic region or not and so so that's how we compute this final score but one thing you might be asking is okay well what's the supervised data you use here like how can you even get a handle on that and so again luckily for us people much smarter than us in biology and John said hey you know what you should look at is this viability screen and so they do these viability screens where they target genes and then see how well with CRISPR and knock it out and see how well the cell survives and what's known is that a handful of genes are non-essential I don't know if it's a handful actually I forget the number but there's something this is fairly well annotated so you can go along the jima you can pick up the non-essential genes and by definition a non-essential gene is a gene where if you knock it out the cell survives and so why is that important so imagine I take the take non-essential gene and I target it with CRISPR and so that gene is an imagine I'm successful with the on target or I'm not either way the cell survives but imagine that it has off target activity the more off target activity it has the more likely it is that it will kill the cell or be non viable so this is a proxy to over all off target it may not be the exact aggregate we want I don't know what the exact aggregate we want is I don't know and if we if I had one it would probably be different for each person in the room who wanted to use it but I think we can all agree that it that's a pretty reasonable proxy and people have since then been leveraging this for this purpose as well and so that's the supervised label we use and you can see here that we compared to CFD which is John's and this is from the server I should actually say has been shut down the last revision we noticed that we couldn't access it anymore they basically said we know other people are doing you know a good job on this problem and we're gonna retire the service but we it's still in the paper and you can see that that you know it was not well tested and it's not performing super well and that our model is performing much much better than both of the other ones but you can see actually that the correlation is very low right it's only 10% so it means there's a lot of work still to be done here and I think probably the bulk of that is actually just getting more data because there's only so much you can do for this very difficult problem with these very limited and so now this will be live very soon since the paper has finally been accepted after years and so what we now have done is we've combined our on target and off target modeling into a front-end web page where we've actually pre computed for the whole human genome anywhere you might want where there's a gene where you might want to deploy CRISPR and do knockout it has pre computed the on and off target so the own target it's not a big deal because that's just one number but that off target that means for every single place we've scanned the whole genome we've built a shortlist we've scored every part and we've put it in so okay so first of all who can guess how many cores we had access to to do this how many would you have here I see okay so we use 17,000 cores and now who can guess how long we ran that 24 hours a day come on one guess oh my god six months no actually only three only three weeks they wouldn't have let us do it for six months but during the course of revision we did do it three times and so that was kind of I don't know how many places you can do that and that was kind of a lot of fun and when I was interviewing on the academic circuit and I would tell people this during my interview talk they would kind of freak out about how they were gonna get me those kind of resources so anyway so this will be live soon and you can so you can put in a transcript a gene name and then it will basically give you and you can sort by different things depending on what you care about one thing we don't do is that some people say is like well like just rank them by some overall quality but you can't I mean for every person has some different amount they care about the on target versus the off target and so you can just dump this on to excel and do whatever you want with it and so I think I'll mention this after if people care after people leave but just let me finish up since time is running out so this is very much joint work with my really really wonderful collaborators who Niccolo was here last week and john here at the broad as well as a whole bunch of other people who are at the broad or have been at the broad down here and also for the last couple of revisions you enlisted the help of of Ben and Keith and Alex who are actually the the developers of many of the state-of-the-art techniques to measure off target activities so the assay they developed that we originally actually have used in this paper is called guide C can we actually generated some new data to validate the model with their help and so right thanks very much [Applause]
Up Next

CRISPResso2 Tutorial: Quantifying & Visualizing Genome Editing
@LucaPinello
2.8K views•2020-05-20

Algae Biofuels: Harnessing Microalgae for Renewable Energy
@LosAlamosNationalLab
623 views•2020-12-03

Microbial Degradation of Plastics: Biodegradation Pathways & Sustainability
@majeedhammad
2.9K views•2021-04-11

CRISPR and Genetic Engineering: How Gene Editing Works and Why It Matters
@kurzgesagt
30.5M views•2016-08-10
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Biotechnology







































