This talk explores the limitations of current provable robustness approaches in adversarial machine learning, which focus on certifying robustness under LP norm threat models (small pixel-level perturbations), and presents ongoing research directions including alternative threat models based on optimal transport/Wasserstein distances, faster adversarial training methods using randomized FGSM, addressing overfitting challenges in robust training, and applying randomized smoothing to certify pre-trained black-box models. The speaker argues that while provable robustness provides valuable theoretical guarantees, the field must move beyond narrow LP norm assumptions to address more realistic security scenarios and broader notions of trustworthiness in machine learning systems.
Beyond Provable Robustness: Next Directions in Adversarial ML
Added:all right hi everyone it's great to virtually be here it's wonderful to be able to talk at this workshop I'm gonna talk today about some of our work that I'm calling beyond provable robustness so this is gonna be worth the kind of spans both talking about some of us that we've done in provable robustness which I'll talk about what that means in the context of adversarial robustness as well as sort of what I think the next steps are for the field and sort of from this perspective of sort of machine learning robustness what do I think is really going on here I should you know start out by saying that this work is is largely due to a lot of wonderful collaborators so here I have my students and this is also collaboration with a few other people as well but here I'm showing my students and a few other actually student collaborators so that the students have their pictures here all right so just to kind of dive right in this talk is going to be about kind of traditional notions of adversarial robustness in machine learning um so I'm gonna start off by first giving a little bit of background on that this is a very broad workshop that was speaking out here and so many the talks are gonna go quite far beyond or captured notions of robustness and trustworthiness that are well beyond what I'm talking about here what I'm gonna talk about in this talk is largely kind of a restricted notion of robustness that comes from adversarial robustness it's an area that I've worked in a lot and a lot of other people have worked in this field as well but I do want to preface everything by saying this this is of course a very narrow field and a very narrow definition of a robustness and so hopefully by the end of the talk I can actually also do some it's a little bit broader discussion and certainly the panel later do some broader discussion about where we might go from here to get beyond these sort of simple notions of robustness so I'm gonna start with some basic background on adversarial attacks what these things are in the context of deep learning systems and then I'm gonna talk about a lot of the work that we've done on provable robustness this is actually how we can build machine learning systems that are provably unable to be attacked in this way at least to a certain extent but then I actually want to sort of go beyond another that's the title of this talk I know kind of going going beyond provable robustness because I also want to talk about what we can do or what some of our current work is that thinks about ways to either get beyond that or thinking about issues that go beyond just provable they're certified for robustness because I think that this is this is one piece of the puzzle but it's only a small piece and I'll sort of mention some of the work that's ongoing in this area and then I'll finally finish with one slide on kind of my hopes for what's next in the field the field I've been working in for for some time now and I have some sort of strong opinions about what I think is maybe good or bad directions for where we go from here and I hope I can provide a little bit of motivation or or insight into where I think we might want to go from here all right so let's start off though with a background on adversarial attacks so I'm sure most people at this workshop have seen these kind of pictures before they're they're sort of classic in machine learning these days this is a figure from Alexander Maddie's group but the same thing actually has occurred in the same picture occurs in its basic form in many different settings the idea here is that machine learning systems despite how well they work are fundamentally I would argue broken you know in a very foundational way and the idea here is that you can take this image over here like you have a pig like you have this image on the left here and you add a little teeny bit of noise of that and to be clear this is a small multiple of this of this noise so it's the resulting image you get is the image on the right and it's undistinguishable from a normal Pig but yet a deep learning system will very highly come with very high confidence classify that as an airliner an airplane and the idea here the way this is working if that this is not a random noise that's being eyes the image what's actually happening is that we're solving an optimization problem to construct this noise so people were maximizing over some perturbation which I call which we call Delta here within some sets that we're calling uppercase Delta and both those are going to be very important actually throughout the rest of this presentation deltas little Delta we call the perturbation and that'll be called a threat model and so we're trying to find the perturbation within the threat model in this case this is that this threat model actually involves essentially very low magnitude perturbations you can't even see and and under this this class of potential perturbations we select the one that maximizes the loss or in this case actually max minimizes the loss on some different class like an airliner class as well as maximizing the loss in the original class of tick and this is sort of a very visceral example here because the image on the right looks the exactly the same to all of us in fact it is exactly the same image we can barely even see the difference here right and so if these machine light systems that we you know are putting all our trust in can't even figure this out then it seems like we're making some fundamental mistake here and there's been a lot of interest in this work and recently yours the one should also clarify that this work actually goes there's work that dates back much prior has sort of recent work in the area and actually this the work has been going on for a very long time and this in this sort of notion of attacking and adversely attacking machine learning systems so after I sort of present this I also want to answer another question is which is why should we care which is it sort of funny question to ask because it you know it seems so so obvious this is a problem but really I would argue you actually should ask this question about why we should care because it's not clear that we should care because you probably by all accounts don't have an adversary that's trying to change your inputs at the you know at the pixel level or if you do you have bigger problems right then you know someone accidentally changing your Pig to be an airplane you've already been completely breached and and nothing else is going to matter so the really matter if you can trust your machine learning or not I would argue there are actually two really important reasons why we want to think about these these adversarial examples despite this fact that probably these exact settings are not that realistic the first is that there are genuine security implications so these things are not just virtual they actually also can transfer to physical world attacks and I think actually I won't I won't bother talk about this too much because I believe that the next speaker actually is the person in this first photograph here so it's my colleague at Carnegie Mellon Leo Bauer and he's going to actually I think talk about this a little bit more so so we've also done some work for example an image and this one on the right here in obvious detection trying to you know put packages and images to make that all everything else in the scene disappear but I would say that even if you even ignoring the security implications or actually I'm gonna do for the most part and this is not the main motivation for me though it's a large motivation I would also argue that this says that these examples they say something very fundamental about the nature of deep learning about what it is we're learning and these classifiers and what really these decisions that these things make really look like because however they are working they're not working like humans are right because we know it's a pig nada not an airplane and so this says something very fundamental the nature of what deep classifies are doing I think and so because of this it's sort of a natural goal to try to design systems that are not going to be susceptible to these to these attacks in other words can we make a classifier that actually will classify this thing on the right as a pig despite the fact it's been adversely perturbed and the short answer is yes we can do this to a limited degree of success and this is actually what the first part of this talk is going to be about the simple idea here is that instead of minimizing the most machine learning systems or trains basically to minimize the expected loss of our actual data points right but what the this adversary robust setting is actually considering is something very different it's actually considering evaluating your system based upon the worst-case loss in this perturbation region so you're now perturbing your example as we did before and you're actually evaluating you know so with some perturbation in in some set and actually going to evaluate the power of this classifier under that worst-case performance and so if you're going to do this you better train your system not based upon the first objective not based upon this one here which is what most machine learning does but actually based upon the thing you really care about which is the robust objective now I actually want to pause for a moment and say you know I I don't I everyone the sort of objecting that we don't really care about this you know this objective of a very limited or very specific perturbation classes here I had completely agree and actually a lot of our current and ongoing work is precisely about how we how we can do things that come up with more realistic models for the threat I'll play that later though for now I'll talk about their relatively simple models and particular I'll talk largely about this threat model being what are called LP balls so you basically can can perturb your image within some norm limited upper tip with some normal aided perturbation now there are really two ways people go about doing this weird actually I should say to two good ways so the first way people do it is they just try something random hope it works and then a few weeks later some other researchers have broken that defense and doesn't really work at all so what there's two actual strategies which was sort of stood the test of time against these things and they basically boil down to sort of an empirical approach and a more certified or provable approach so the first approach that people use with adversarial training so this is this is actually goes back to support by Ian good fellow a Lascar akin and Alexander Madhuri and this basically the idea is very simple here it's that you sort of take gradient descent steps not at your original examples but at these worst-case perturbed examples and keep trying to play this game of you know adding the worst perturb example into your data set and playing this game and this actually works well empirically but but somehow it's unsatisfying because often times we can't really know this is going to work though we actually will return to this later on because it has actually stood up quite well under a lot of evaluation the second approach though which is one that I am going to spend more time focusing on are what are called certified or provable defenses so there's been a lot of work in this area including by a lot of my my group and others and the idea here is we actually want to get some guaranteed bound essentially on the nature of this objective so if we can somehow despite the fact these are complex functions complex nonlinear classifiers if we could actually somehow bound the performance of our loss that we might be able to get an actual bound here so this is actually what we're gonna do and I'm going to show you at least a few highlights of the ways we do that because there really are two approaches we've taken to achieving these certified and provable defenses the two approaches which I'll talk about in more detail essentially boiled down to convex relaxation and randomized smoothing and these form right now I think that the two most commonly used provable defenses against against the adversarial attacks okay so to get started now let me talk a little bit about this provable about provable robustness how do we actually attain these models that can give us guaranteed bounds and the idea here is actually is actually a fairly simple one so in deep the problem with deep networks the reason why deep networks have these problems is that you can take a relatively simple image here like the one on the right so you can take a relatively simple examples here like these X's and O's and consider some sort of region around these things that that's the that's the you know the nice in this case you know square boxes around them the problem is when you feed them through a deep network what comes out the other side are these very irregularly shaped you know hard to reason about combinatorially complex objects owing to the nonlinear nature of the of the deep network and this prose is a problem for sort of even sometimes understanding you know is there an adversarial example or not and so the the idea behind our first approach is convex relaxation approach is we're actually going to form convex tractable outer bounds over these regions which we call the adversarial poly tope and we're gonna perform robust optimization over this region as a particular was going through this first part is gonna apply to networks with value nonlinearities though it can be generalized to some more general activations as well so how do we do this I'm gonna do a quick overview that the maths here is a little bit well if math are both parts actually a little bit dumb a little bit involved but for a quick but for a quick version here's the following picture to have in mind so the problem here is that these regions these networks are fundamentally nonlinear and so could the and they're nonlinear because they have these activation functions it's not on their activations like your value or like a sigmoid of things like this and so the idea is if you want to get an outer bound on the on the quality of these predictions what we can actually do or the size of these pictures from cache should do is form an outer bound a relaxation on these nonlinearities so for example for the value if I were to know for you know maybe a lower and upper bounds on how how far we could we could span and the pre activation space for that values that's noted here by L and by you then I actually can form the convex relaxation of this rail u which looks like a little triangle here and we know that the worst-case perturbation and under under our original model here it's actually going to be a upper bounded by the worst case perturbation over this relaxed objective because this is a strict this thing here is a strict superset of the allowable activations here and now similarly this actually turns out to be hard to precisely optimize those things we actually go one more step and we actually form one more relaxation which takes the dual of this problem here it turns out this problem here with the gist this relaxation is now a linear program is actually tractable because all the constraints are now linear it's still too slow to solve though and so we end up forming a dual linear program and looking at specific solution to this dual and the details are actually are actually a little bit involved but basically what comes down to in the end is we get a fairly tractable way to form a closed form estimate of a provable upper bound on this worst case loss objective all right so that's that's the basic sort of technical idea of the approach here and now how does this work it works pretty well in practice so for instance if you take data set like Amnesty well I should qualify that it works pretty well some some of the time on smaller problems so maybe I shouldn't say a worse problem practice it works pretty well on some small problem so you can take a problem like it like Amnesty R if you train a normal Network on on this it gets pretty good error right it gets you know around the order of 1% error but the both the bound the national the actual robustness here is basically a hundred percent so even with a really small perturbation size this is the perturbation of epsilon being just 0.1 so you so the images are scale between 0 & 1 I can perturb them I compare each pixel by a value of 0.1 and even with this sort of very small amount of perturbation allowed you can basically completely fool the you can basically completely fool the the standard trains network here this is actually showing the bounds not the not the actual amount but basically the bound of the actual amount here are more or less the same thing if you train a robust linear classifiers it's actually been very well understood for a while those are things actually very laid the things like SVM's you can do a little bit better here so you can actually now get a robust an actual guaranteed robust bound of 17% but you pay a very big cost in accuracy you pee pee pee pee you know you've reduced your clean accuracy but a year robust accuracy done that and be your clean accuracy is reduced from you know one percent to 17 percent which is this not which is a pretty performance on one chemist and then finally with our method though you're actually able to maintain that same level of clean performance while being able to guarantee that no matter no matter how many papers are written after this after the fact right no matter people is some certainly teachers once I tried it to to write to break other defenses this will never be broken so so no classifier will ever or no attack under absolutely no attack under this threat model here so no attack with this epsilon here will ever attain more than more than three point seven percent error here and this bound has been improved since then by follow-on work this is that this work goes back to our original ICML paper 2018 now the problem here though is that you can ask you know well and this is pretty small right aren't we beyond that these days in in in deep learning and and of course yes but the problem here is that these methods as like a lot like a lot of a sheep robust training they scaled poorly so everything looks great on the M this size when I'm showing here and the different colors now are different size network so you can't really go small and this network or a larger one comes trainer ResNet we didn't bother on mms but you can train to rest nets to and there you know you can get pretty good performance here for robust performance but when I started going to domains like C far your accuracy falls off quite quickly so this is showing now results for these different models with a allowable perturbation on C far of 2 out of 2 and 15 so you're only allowed to perturb each pixel by two of its color values yet you know this go this takes C far from its accuracy we're used to of like you know six or five percent error all the way up now into the mid 40s when you're talking about certified robustness and I should mention you know the the I'm not showing it here but you have to pay a very hyper high cost I'll show it some other slides you pay very high cost in in clean performance to so these things also get you on the order of 20% or sometimes thirty percent clean act clean error as well so we thought you know we suffered horribly both in terms of getting a certified down than robust performance as well as giving a sort of clean label performance of a firm on our function on our underlying online data set so this is sort of a sad story for scaling maybe it's not a student to another set of techniques which is something about a different way of doing of getting a certified defense using a technique called randomized smoothing so the the intuition here is something I get this intuition before about kind of outer bounding the this this polytope of reachable points the intuition of randomized smoothing I think is a little different the idea is that in some sense you know if you have some query point here and which is the correct class you have a correct blue class maybe and there's an adversarial example what that kind of means is the aboriginal examples are sort of these like you know these regions that kind of jut into these like this likes you know spikes that kind of stick into our decision space really close to the to the target point in other words that sort of said something very fundamental about the smoothness or lack thereof of our classifiers and so one obvious way to fix this was actually has been considered by by others before us we were just dissident slightly different analysis on this but it's been also being accepted by others before us so the career at all and Liat all thought about this as well the idea is actually quite simple to avoid this you can just basically apply Gaussian smoothing to the function so you take your underlying classifier which is this you know signified by this sort of jagged region here and I sample our original point or points around that point of our Gaussian and I take them the majority sort of vote of all those points is my prediction so I don't pacify it according to just that there's the single classifier and I think of sort of a vote of many randomized classifiers and and take they're they're sort of joint average prediction and this has the effect of con smoothing these boundaries so the spikes that used to be there don't occur anymore and in fact it's something very sort of that very specific you can say about this so it turns out that you know this may seem surprising because in the smooth classifiers in the end you're just doing sort of random Gaussian sampling it isn't about that you know you're not drawing that many samples here but you can say that if the underlying smooth classifier correctly predicts the the label so that is when after you sample gaussians and take their average if that average predict is ago sort of is the correct class with with probability P then it turns out you can actually say that this this classifier at this point is also robust with some l2 radius so in a perturbation region on this case an l2 ball not an L infinity ball but it's sort of more or less they're very similar on the weights which basically scales like the the the variance of your randomized sampling times the inverse CDF of a Gaussian this is a very sort of simple formula here that actually can give you certified results but most important is this this method actually does scale to much larger domain so it scales that picture to things like image net and you can get image nets you know you can train image net sized classifiers that actually give you provable robustness and this is one of the first times or this technique really is the first one it's this class of techniques ours and other people's works was was the first sort of class of techniques that were able to get this certified robustness at the scale of image net and so particularly if you can you can take some of these this plot here is a little bit more of an involved version than the last ones I showed you but what this shows this shows actually certified accuracy on the on the y axis versus the the radius of certification so basically if I have if I consider a larger larger of a your perturbation radius here so for if the more you can preview examples the of course less guaranteed performance you can have but you know for example if we if we look at maybe to say this spot right here then this corresponds to what I'm saying here so this is saying that you know on image net I can perturb my my image with an l2 amount with a knelt with an l2 ball of radius 1 the because a very small amount of perturbation that that sort of would be you know covering one pixel the whole imbalance or occurring a bunch of them a little bit less but we can actually guarantee now a top one accuracy of 37 percent so you know definitely not going not blowing away any records on image net but it's the sort of you know some of the first times that sort of these non-people guarantees on infinitive impossible and also of course these results have been improved a lot upon since we since we published them um okay so without all being said I want to sort of give some sort of key takeaways or obstacles I mean they're more like obstacles in this approach here because I think people often sort of see these results and say oh well maybe for busting says solved you know we've we've we've done it we solved robustness and this is very far from the truth here so this problem is still really hard and the way I want emphasize that is we are certifying relatively tiny regions and our images compared to what humans can easily do right so if you look back here you know we're talking about one pixel total on these images and imagenet or maybe you know two out of twelve fifty five yea pixel values on see far for the L infinity ball these are tiny tiny perturbations crater what it's easy for humans and they they don't signify anything yet compared to the the robustness that we as humans have on these things and that we really want to have on on these sort of images the second sort of key takeaway is that there's a large when you train these robust models this actually holds true for both for the certified ones I mentioned here but also for the empirical ones when you train them you actually get much worse not only sort of maybe not that great rapport Minh so degrades or clean accuracy a lot so essentially making these models are busts and forces them to be smooth in a way that actually seems to drastically reduce the effectiveness of modern deep classifiers and this is this is a big problem here and finally I also now want to finally sort of get at this underlying point here which is this LP its LP norm threat model so the idea that somehow perturbations can or should lie within LP norms this is really much more of a simplifying assumption or a sort of a toy domain that we talked about at least initially because we can't even solve it even that's hard but of course any real threat model for any sort of real adversary of real security setting this is both far too strong in some sense because most adversaries wouldn't have complete that level control but it's also far far too weak because it doesn't capture nearly what a real adversary would do you know a real adversary would not be confined to tiny in percent of imperceptible perturbations of your image and so this problem I really do leave is actually still quite useful but it's important to also understand the limitations of the problem and I'll get back to this in a second later and talk to you in life and my final wrap-up slides because this these are a you know simple domains would tell us a lot about the nature of deep learning but they are not realistic security scenarios okay so that being said let me talk about some of our sort of current and ongoing work which is gonna be sort of a you know a quick pass to some of our current work which tries to go actually beyond some of these ideas approval robustness to extend either classical empirically robust training or to extend the notion of what we mean by robustness for example by considering other threat models besides the LP ball so I'm going to go through a few of these now it's gonna be quite quite quick actually but I want someone to highlight some of the ongoing work in this space so for example one of the works we've been doing is about exactly this topic I mentioned to you before on alternative threat models so a question that sort of comes up a lot is you know why on earth should we think about these you know LP perturbation balls as being the right metric for understanding how an attacker could function doesn't doesn't really seem like a reasonable way of doing things and it's definitely not and so in some of our prior work which was I see my last year we thought about a simple alternative model based upon a Vassar signed distance in image space so instead of thinking about sort of images as just these these things that you perturb you think about actually kind of moving the mass of an image or shifting the mass of an image in a kind of an optimal transport setting and so this actually can capture things like rotations or translations or distortions much better than than than traditional LP norms because for example if I take my original image here and I were to translate this thing over a little bit if I translate my image over a little bit then under you know L infinity norm this is a huge difference because you know a few pixels are perturbed a huge amount but of course this is the same image it's just translated over a little bit and foster-son distance can actually capture things like those pretty well or it gets at these things in a much better way than normal normal perturbation or a normal sort of additive perturbation models do and and again just being very sort of high-level here we that the actual challenge this is actually coming up with an efficient way of projecting on at the Vassar Stein ball this is the set of all examples that have this is less in some amount but we developed an efficient way of doing that and then we can actually see that when you apply these perturbations to things like c4 images they end up changing the image not just like with this random noise but they actually end up changing it at the boundaries of this image here so you actually have you sort of you know the biggest changes occur kind of at the boundaries suggesting it's kind of attacking more of the semantic properties of the image though it to be clear this is we're still very far away from sort of real semantic notions here that's one bit of work we've done in sort of building upon these notions approval of Brussels are trying to sort of understand them better another work we've done is on sort of actually going back to the empirical side of adversarial training we've thought about ways to make this process much faster actually want to have this physically because this this paper appears here at iclear in a few days so I should have gotten the actual time but anyway this is the name of the paper it's called fascist better than free ad for sale training revisited and and just sort of highlight highlight going after that you might wanna check it out in the iclear sessions if it sounds interesting but the basic idea here is that is that when it comes to adverse a training there are a few problems so so not only this and so ever so training is this other approach mentioned before which basically involves running an opposition method like participated dissent to find the worst case perturbation and then training against that worst case perturbation instead of the nominal examples of a class of your dataset now this is not the first method so before PGD was sort of became too popular you know became used popularly forever so training even simpler methods like the fast gradient sign method which is actually take a single gradient step to find the worst case example we're actually these other alternate method specially much more common and they're much faster but however methods like the fast gradient sign method which was used sort of for a while before these were these were actually sort of dismissed kind of early on because it turns out that models trained by them were actually very vulnerable to more fine-grained attacks and so the field kind of a band even thinking about using FG SM training for a personal robustness it turns out as we showed this I clear paper this is actually I'm just saying if this is actually very surprising so it's been quite thoroughly vetted by a lot of people since then but it turns out that actually you can kind of do some simple fixes to make single step addressed training actually work just just fine and so basically it involves you know instead of always starting at your example you actually first randomized a little bit and then you take a step and it turns out when you do this at least compared to the reported results these are RF GSM methods basically perform as well as as the reported results for PGD training so again on some nicely far see far you know we achieve very similar we achieve very similar levels of performance here in terms of the clean error and more importantly in terms of the robust error we get sort of very similar performance between our method and the orange here and kind of classical PG training from in the in the blue but of course the Bennett the real benefit here is that the fast gradient the fast ever training takes vastly less time to Train so we were can reduce a particular you know this this original model here this original adversely robust see far ten model took 84 hours of time to train on on a GPU whereas with this much faster method and some other tricks like half precision and stuff like that though it doesn't apply no me even with all this tricks you cannot even come close to bringing down episode training to this fast but with all these tricks we can train a CFR model in about six minutes which is which is a huge huge advance upon what's been quit what's what if it's currently done and even on the bigger days like like like imagenet we can now sort of feasibly train much much larger imagenet models in in you know in a reasonable time like 12 hours but actually all hope is not lost for PGD yet because that's what's been interesting about so the result I'm showing here we're actually from the the reported results on the pond Alexander my newspaper talking about pgps training but actually in another paper that we have another preprint we have we have available right now we also actually look more closely at the performance of adversarial training and notice something sort of very striking and have robustness which doesn't happen in deep learning so it turns out that in adversarial robustness the there's a much bigger problem that arises when it comes to overfitting so in Sandra deep learning you can't don't worry about overfitting right you train you train your your your algorithm for as many a box as you want you know you don't think about early stopping because we have enough you know regulation methods that have sort of fixed this thing and we sort of know how to we know how to make you keep getting training error lower and lower and your test error just keeps improving or these a plateaus and doesn't get any worse um it turns out the situation in robust training is very different and this is actually again largely interesting just from a standpoint of our understanding of deep learning so if you actually looked at that training of it did for PGD based training they trained for 200 at box but the final accuracy which is the one here I mean won't use a different color the final accuracy which is the one here is actually our final error is actually about eight points higher than the actual test error was right after they sort of dropped it did the first wait the first learning rate decay after training and the interesting thing here is is not so that means that PGD if you do it correctly and sort of do early stopping properly you can actually achieve state-of-the-art results and in fact this is maybe an a sad negative result but despite extensive tuning we found that really no more recent method so all the advances people have sort of been picking on tweaks of average sale training since the PGD sort of basic sort of PGD methods basically work worse than just PGD with good early stopping here and in fact there's that there's a recent leaderboard that some people maintain a value in these methods against against very sort of sophisticated attacks and and we are here which is the best performing method that does not use additional data so we just used to see if our data set these other ones use all augmented data not just with simple augmentations but actually new data period and without doing that we're able to actually get you know of the methods that don't do that we should have the best performance with simple PGD and early stopping and maybe even more interesting those papers that we also tried every other single regularization technique is used in deep learning people try these days so you know drop out and cut out and all sorts of things and nothing actually solved this problem so we fundamentally don't understand the nature of overfitting and adversarially robust deep learning and have a very different character than traditional deep learning and finally almost everything I've talked about here deals with sort of providing or defending classifiers that you can train yourself right so both the robust classifiers required that we actually train the classifier against us for a bust objective or require that we you know train in an adversarial fashion to get to get our part you know what does a trait adversarially empirical truth to make empirically robust models we need to use this adversarial training mechanism which is also requires that we you know be able to train our model from scratch so we also know that sort of standard off-the-shelf models for the most part art are not robust at all and this raises our finishing question is can we somehow make these models existing models pre train models robust it turns out we can so some reason that work which I'm also gonna highlight this is a work in collaboration with Hadi Salman and some colleagues at Microsoft Research it turns out that by taking our sort of normal image here and by applying randomized smoothing so you know adding this Gaussian noise to it but then most classifiers are really bad at Gaussian you know with the Gaussian always preaching customers cannot classify Gaussian perturbed examples so we actually also add to this before we close but I'm a custom trained decoder so because I'm trained denoiser and so we D noise yeah we add noise and then we D noise so that we can have this whole end-to-end smoothing randomized spinning process still apply here and by doing this we actually can certify sort of for the first time black box classifiers that no one has actually trained to be robust and we demonstrate those actually again that the details of these curves aren't too important than actually you know it's not meant to be visible but the actual parts which you hopefully can't see is the is the sort of headers here where we actually go ahead and create without of course having access to any of the of the code or data here we create robust versions of publicly available image classification API is like the azure API or the Google cloud vision API and others okay so with that all being said I want to take the little time I have left to sort of discuss what I hope can be next for the field I think we're making a lot of progress in this area both on the certified robustness side and we can say some sort of pretty impressive things that were you know two years ago we're not thought to be possible we now can say some pretty impressive things about the nature of of guarantees we can make about the nature of classifiers under under these attacks and we also know understand I know a lot more about sort of practical empirical robustness and making these things robust in practice and so the question that I sort of think about a lot now is is what's next for the field I mean the first point I want to make all of course you know which which I should highlight before anything for saying anything else is that it's not like we've solved this problem so we haven't actually sells underlying problem here and there is a lot more progress that can be made even in these very limited settings we're talking about so far so I think even under these very restrictive LP norm perturbation models there clearly is still massive room for improvement right we're nowhere close to the performance that we would expect you know that we're used to in standard deep learning systems and so the question is you know why is this gap still there why is this gaps seeming so persistent and you know why can't we make our models just better so that's sort of to start things out I only mention that me we haven't solved the problem but maybe it's sort of more fundamentally I think that there are a few things that there's food directions that I would sort of recommend for the field as a whole which maybe tie in betters the topics of this workshop I think are more about trustworthiness and robustness from a very very much broader perspective than a very limited perspective and I'm kind of addressing technically here so what my wishes for the field would be that would be is something like the following I think there needs to be a better understanding about the key motivations behind why we're looking at this robustness problem so I think a lot of people and this actually came out of discussions I had from talking about people at this workshop because I think oftentimes it's the the implication seems to be that this is the biggest problem with deep learning system so the oppressor does work everywhere they work great they're secure except for this problem of LP norm attacks and in my and this is hopefully as you can tell right now from this talk because it's Kleon true and I hope that we can get a better understanding of these key motivations which in my opinion are much less about actual security and more about the fundamental nature of deep learning of the classifiers we learn and they're sort of what they can do their smoothness and these kinds of properties about them but having said that I think we all seen a much better understanding so second third parts were are related here we all seen a much better understanding of where robustness of any type can sizably impact the the our arts of the practical security of machine learning and where this interface will be between robustness as I framed here sort of as this min max problem where you're you know it's very game theoretic formulation it's a very sort of mathematical formulation and narrow formulation where can this narrow formulation sort of play with our practical understanding of what we want in a robust classifier and finally if I will say one last thing I think it might be also a really good idea to move beyond the standard image sound video etc domains that deep learning has been so so successful and so far there are a lot of domains that are much estate lower dimensional where robustus is really important robustness and safety and verification are all very important things like control and other domains but by the way there's a huge history of robust control but there are a lot of other interesting domains I think we could consider that go beyond the kind of classical deep learning setups where these approaches may find in fact even more application because when it comes down to it you can never actually certify an image classifier right you can't say I can't give a spec for what all you know pedestrians in a road look like and so thinking about other domains could be very valuable here all right well with that being said thank you very much I'm looking forward to the panel in a few hours the papers and code are gonna be available on my website and I also want to thank my students as well as some of my external collaborators here and it's gonna great to great things everyone for attending this sort of virtual session and and hopefully answer some more questions at be at the panel
Up Next

Simulating Forces: Gravity and Wind in p5.js | The Nature of Code
@TheCodingTrain
106.3K views•2020-03-27

Building Real-Time ML Pipelines with Feature Stores and MLOps Frameworks
@ODSCAI
5.1K views•2022-02-20

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence








![[EEML'24] Martin Vechev - (Provable) Robustness of Neural Networks](https://i.ytimg.com/vi/fa2NY2Fs6X0/maxresdefault.jpg)


















![[DS Interface] A Closer Look at Accuracy vs. Robustness](https://i.ytimg.com/vi/pUsb5v99oqA/maxresdefault.jpg)











