Differential privacy is a mathematical framework that enables training machine learning models while preserving individual data privacy by strategically adding controlled noise to gradients during training, and Opacus is a PyTorch library that implements this technique with minimal code changes and performance overhead, allowing practitioners to track privacy budgets (epsilon) throughout the training process.
Differential Privacy with Opacus: Training PyTorch Models Privately
Added:[Music] hello everyone my name is davide and i'm an applied scientist at facebook ai today i'm excited to talk to you about differential privacy and how you can use it to train models with by torch let's go over the agenda real quick we're going to start off with an introduction to differential privacy we'll follow up with a more in-depth discussion on how to use it with machine learning and with my torch and there'll be a demo column and finally we'll conclude by talking about the awesome ecosystem that has formed around this project and i'll point you at some more resources if you're interested in following up all right so what is differential privacy rather than throwing a definition on you let's look at the problem that differential privacy can help us solve this is a collection of data with each data point being an image what we want to do with collections of data like this is learn something valuable from them something interesting that becomes visible only in aggregate but let's go back in for a second we have a ton of resolution here and we're able to see each individual data point but do we really need all this much information if all we care about is actually just the aggregate just the mona lisa could we do something like this here we have greatly reduced our resolution into the images so much so that we aren't able to say what they contain anymore we have preserved the privacy of those images but now the real question is of course what about doing our job what about the aggregate let me just show you the answer is remarkable almost nothing happens to do your job you don't need this much information from each and every of your training samples this is the core idea behind differential privacy by destroying a controlled amount of information you get to have your cake and eat it too you get to preserve both the privacy of your training data as well as still be able to do your job still be able to obtain the aggregate that you care about and this is what made the financial privacy the gold standard in privacy protection so much so that it has been used by the u.s census bureau for the 2020 census and it has been recognized by the european union under gdpr i mentioned aggregates previously and that was a vague term there is a reason for it it's vague because aggregates can really be anything they can be as simple as counting stuff or computing a mean or they can be as complex as training an advanced machine learning or deep learning model which is what we're going to talk about now one important thing about these models is that they memorize and we should really be careful about memorization in fact memorization violates privacy don't believe me let me show you this the image on the right was extracted from a classifier that was trained using the image on the left i want twitter at one point this was a classifier it was not again it was not a generative model this was a classifier that would read an image and classify um what person will personally belong to so the information about the training set can be extracted from a trade model and this is why it's important to think about privacy so how do we make machine learning more private well we're here so you can imagine that of course the answer is going to be with differential privacy so rather than asking what let's ask how and it turns out the differential privacy can actually be added at any point in your model's life cycle in the previous example we added it to the data and that's fine you can also add it to the training algorithm or to a trained model after after the fact or even to a model's predictions during deployment the gold standard is to do it during training and if it's by torch perfectly so that's what we're doing pythor just made it simple to train advanced models and now we're extending the simplicity to private models we built a library for this it's called the paccas and it's a library to limit memorization upakis is the python's domain library for differential privacy we built it so that it's fast it's safe and it's production ready now to show you what it is and what it does let's switch to colab it's much easier to show you stuff there and of course my video on collab can not be interactive so we we've also published the actual notebook for you to play with and this address let's see how easy it is to convert a project or yours to be trained using differential privacy to get us started on the notebook i have already written a bunch of import statements as well as define data set and data loader we're just going to turn a model once if our time i have also defined a model architecture already we are going to use a small model this way this demo trains fast but of course alpacas works with animal and just to get started let's train without privacy by writing a conventional training group this is all standard stuff organization batches move the batch to cuda compute the loss compute the gradients and have the optimizer update the parameters this works indeed and we're hitting about 28 30 iterations per sec and we are able to complete the network in less than 10 seconds great so far this has been the pythons we know and love now let's see how to make it more private all we need to do is import the pockets and define a privacy engine this abstraction takes care of applying differential privacy for you behind the scenes um it just needs a bunch of parameters such as your model um batch size total data size plus some privacy specific parameters alpha is the range of running differential privacy holders that the accountant can use to compute your epsilon noise multiplier is the parameter you want to tweak the most this is the standard deviation of the gaussian noise that you're going to add to your models per amps and finally max grad norm allows you to tweak the clipping threshold for each sample if the gradient for any given sample exceeds this value it will be set to this value instead and we get a warning don't worry this is entirely normal we expose an option to turn on secure rng for production runs but it comes with the performance cost so we keep it off by default this way we don't slow you down during experimentation once you construct your privacy engine all that's left is attaching it to the optimizer and once this is done the privacy engine will be notified by the optimizer whenever anything happens and you don't need to change your training code to make it private in fact we can literally just go ahead and copy paste the exact same training code that we used before put it down here and we can see that indeed this runs just fine and there's a small performance cost we were able to hit 28 before and now we are at 26 compared to 28 iterations per sec um the privacy engine will also remember everything that he did and it can return the preface expanded at any given point in time so in this case we have returned epsilon equals 1.48 for alpha equals 11.
okay let's switch back to the deck finally i wanted to talk about our community because the response has just been amazing we launched on august 31st and it was fantastic to see that we have already hit more than 600 stars on github we have more than a thousand monthly downloads last month and finally more than 200 of you have already crowned this repo and tinker with it which is just an estimate to how popular privacy preserving ai is becoming which is very comforting speaking of stars this is the growth of our stars over time we found it quite interesting there's there's two fun things about this graph one is that you can clearly see when we officially launched and two um it's just remarkable that about 200 of you um actually found us when we were still codepythogp under a different repo facebook research and already actively engaged with us in fact there are even a couple of papers that we found that um already use this the software in this form um so thank you for for alpha testing this with us we took your feedback to heart and a lot of it made its way into what became a box so thanks and finally shout out to one of our key collaborators openmind they are a community of thousands of developers who are building applications with privacy in mind as part of our collaboration oppakas is becoming a dependency for their libraries such as such as bisift and they recently wrote a couple of tutorials on how to use apocalypse there's one on how to use apaches in general um it's called the financially private deep learning under 20 lines of code and recently they they wrote another one on how to use it with biseft to achieve differentially private federated learning um thank you critic and theo for writing those tutorials and hopefully i have picked your interest in these few minutes that we went together if you want to get your hands dirty and start playing i recommend you just start with our website or because ai we put a bunch of stuff to get you started in there there are a bunch of tutorials examples faqs go check it out we're also working on a series of mediums that will act as uh from a gentle introduction to the core concepts of the psgd and they will take you all the way to more advanced concepts both in the software engineering behind it as well as in the theory the first entry is out so i recommend you check it out help us make a more private and thank [Music] you
Up Next

Secure Aggregation for Federated Learning: An ACM CCS 2017 Talk
@associationofcomputingmach2690
4.5K views•2018-02-08

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence



















![[AI를 위한 수학] 딥린이를 위한 필수 수학 패키지!](https://i.ytimg.com/vi/frkVgBvp850/maxresdefault.jpg)


















