In the Flower framework for federated learning, the ServerApp and ClientApp are fundamental components that work together to enable collaborative model training across distributed clients. The ServerApp uses a strategy (such as FedAvg) to sample clients, distribute the global model, aggregate local updates through weighted averaging, and coordinate training rounds. The ClientApp loads local data, applies global model parameters via set_weights, performs local training using standard ML loops, and returns updated parameters along with metrics. Data partitioning can be customized using different partitioners like IID or Dirichlet to simulate various real-world data distribution scenarios, with lower alpha values creating more heterogeneous (non-IID) distributions that better reflect practical federated learning challenges.
Understanding Flower Apps: ClientApp and ServerApp in Federated AI Simulations
Added:this is the second video in the video tutorial series for flower simulations in the first video you learn how to create your flower app and how to run it in simulation using flower new and flower run in this second video we're going to go back to the same flower app we created and we're going to dive into how The Client app and the server app looks like we will be changing the model slightly and completely replacing the data set we will also change how the data set gets partitioned so let's begin so let me remove once more this out of the way here I have vs code exactly the same as we left it here we have our app with a p project. Tomo that defines the dependencies that we use remember we were using a pytorch we we use the pytorch template and that's why we have the set of dependencies here what we're going to do now is take a closer look at how the server app and The Client app look like let's begin with the server app so the server app can be conveniently constructed by means of a server function which has this signature it receives a context that I will provide more details of what it is later in this tutorial series but for now we can see it as a place where configs leave there are two type of config that is the Run config that can hold hyper parameters that we might want to override at run time for example here the number of server rounds that the server app is going execute fraction fit which for example allow us to configure the strategy here we're going to be using the Federated average strategy is one of 20 plus strategies available in flower you can think of the strategy as the brain of a Federated learning application it performs actions such as sampling clients that should be involved in a particular round it is in charge of communicating them the global model along with some instructions on how to perform the training it will also receive the resulting locally trained models from the client apps and therefore it should Define a mechanism by which all these updated botels can be aggregated with fed a Fed average the aggregation will be rather simple it will be just performing a weighted average of all those models but there are more sophisticated strategies available in flower as well to the strategy we can pass a set of initial parameters this will become the starting point of our model in this case what we are doing is just converting the parameters of a simple model based on a class called net if we look into it it looks like a simple CNN probably you have seen this same model if you have taken other tutorials on how to use spor and if you are more familiar with other Frameworks you can see that with pyto is fairly easy to understand how this model is constructed and how it can be used so very briefly this model has two convolutional layers followed by three fully connected layers what we here is we we call a utility function we have written called get weights if we inspect it what it does and let me highlight here that the way this fun utility function is implemented is totally dependent on one the machine learning framework we're using here we're using python so this is the way of extracting the learnable parameters of a python model we do it through the state dict if you're using in another framework later on like tensor flow or maybe something completely different like XG boost or your custom uh way of representing models you would of course do this in a different way and what this utility function is returning is returning a list of numai arrays because numai is what um the flower strategy understand how to operate with we are getting those numai arrays and then we are converting to an internal representation we call parameters this is a small detail you don't have to worry too much about it but if you have qu if you are curious about it you can check the flower documentation or go directly into the source code the strategy we're using here is fed average as I said we decided to parameterize the fraction of for fit this means what's the percentage of clients we're going to sample on each round we do it by means of this hyper parameter and you might ask yourself okay where are where are these defined if we inspect the by project opal they are defined here this section of the of the by project opal allow us to define the Run config so by default it will do three rounds which probably you notice when we run the application in the first video by default fraction feed is going to be 0.5 this means that half of the clients that are connected will be involved in every fit round and this is another hyper parameter that here we are not making use of because this will be read by the client tabs of course you are welcome to add more hyper parameters here and use them where where you see it whether it is in a server app or in the client app once we have the strategy Define there is another type of object we need to Define which is called the server config which which primarily we want to just use it to specify the number of rounds remember the number of rounds we have read that information from the Run config but there is really nothing stopping you from not using the Run config and just setting it like this but using it from the Run config it brings so much more versatility and you will see that in action very soon and then finally we wrap everything into a super data structure that we call the server app components we pass the strategy and the config and that's all we need to Define our server app we will come back to the server app by the way in the next video where we will see how we can further customize the the strategy using callbox if we now take a look at the client app we'll see that the client app also can be parameterized constructed better say through a client function and if you are paying a lot of attention you will notice that this client function has a very similar in fact identical signature to the one we use for the server function and this is quite useful because we already know how we can extract parameters defined in the rank config like local AO which was is defined here we just need to do context do run underscore config and read that variable you know that here we are also initializing our model this is not as strictly needed to be do to be done here but in this template that's how we do it and then we are also reading some other configs that are unique for simulations these configs as you know they are not part of the Run config because this is something you can think of it as a of a property inherent to client ABS in simulation this is number of partitions this is equivalent to saying how many clients are there in total in in my simulation and partition ID this is another way of saying what is the client ID of this of this of this client app that is about to be constructed so to put some examples non partitions here is going to be equivalent to number of supernes so there will be 10 and then partition ID this could this will be assigned at runtime by the simulation engine and would be any number from 0 to 9 so these two uh variables are important when we call now the low data function because it will allow us to as we'll see one ensure the data set has been downloaded and two parp the data set H according into how many partitions as many as clients are in total and which partition should be used to constructed train loader and validation loader well the partition that is supposed to be assigned to this particular client that that is being constructed in this function and that's given by the partition ID before going into the low data function H we can see that the return object of this function is a is a flower client you are totally free on how you want to Define this class which is here it doesn't have to be called flower client it doesn't have to receive all these arguments as long as it inherits from a class a numai client class or an or a similar class called simply client that you can find in the documentation you are good to go in this example we decided to pass all these arguments so that's what we do again you don't have to pass all of them and you can add more naturally so here we are simply storing the variables such as the model the train loader the validation loader and some hyper parameters like the number of epo we want to discover at runtime what is the device available to our client app we do it it is important to do this either inside the client or in the body of this function instead of doing it here outside and the reason for that will become more apparent in the later videos of this tutorial Series where we talk about how the simulation engine works but the short and short answer is that the simulation you with the simulation engine you can control what devices are exposed to your client apps but if you define the device Outside The Client app you won't be able to control that that aspect so it is very much preferable to put this either in the client or you can put it somewhere here and then pass it to the client client apps allow for the definition of several methods for the implementation of several methods but the two most important ones are more relevant ones are the fit method which receives the parameters of a global model from the strategy that remember it lives in the server app it will call a utility function called set weights that applies those parameters to the local model of this of this client app and then you can execute a typical training H function to train that model locally us using the data that is owned by this client and this client alone let's take a look at how these set weights look like similarly to when we were discussing the get weight function set weights is is defined based on the machine learning framework you're using and also the model you are using here we are using a pytor so this is the way to do with pytor if you have a more exotic model you might want to change this so yeah by default this should work uh right away there is another way to change this slightly which could be for example Torch from numai this could bring more ver versatility slightly more versatility depending on what type of model you implement so but I'm not going to do that here just to let you know that you can do that and and of course if for example you are not federating the entire model just a section of it you will Implement that logic here to apply only a subset of model parameters to the model your client has locally but for now you don't have to worry about that we can go back to this in a more advanced tutorial series for flower if we go back to the client function here we can see that once our model has got the weights applied from those received by the global model we are calling a function called train we pass the model we pass the train loader and we pass some hyper parameters in this case how many local EPO local EPO should should be performed as well as what device to use if we zoom in into the train function you will see that this this looks very very typical as a standard as a training Loop in python as it gets we Define what is our loss to use we're going to use the cross entropy LW because that's fine for image classification we're going to use the atam optimizer but you are welcome to change it to something else and this is the training loop we're going to perform we're going to perform this many EPO using the train loader this is how we extract the images and the labels from the available batches from our data set and then we perform back propagation as usual if we go back to the client tab the return Arguments for the fit method usually come in a form of a replate which first return the parameters of the model that has been updated locally by this client this is done again with the same F function get weights at we use earlier in the server app to extract the parameters of our model once more this might depend this might require some adjustment if you use a very exotic model or different framework while passing communicating back the updated model probably is the most important uh return argument from the fit method there are other arguments that are can be particularly relevant depending on what you are doing at the strategy here we are returning what's the the size of the data set of this client this is useful when doing the aggregation for example if you want to give more weight to the clients to the update sent by clients that have more more training examples one way to think about it is if one client have many more examples and others maybe the update they are contributing to the global model should be should have some level of priority but you might argue that this is not a good way of doing it so you can still use this this this scaler here to change the way models are aggregated and finally the last argument we call it metrics you will learn about metrics in the next video a bit more and in the following one the idea here is that you can set up uh Python and dictionary communicating key value Pur where the keys are strings and the values are integers or floats typically integers and floats is enough to cover most used cases but in the one of the followup videos we'll learn how to send any arbitrary data structure so that's it for the fit method the evaluate method is fairly similar it also received the parameters from the global model we also need to apply them to the local model and then instead of train what we will call is an a function that does that evaluates the global model on the local data in this case we are calling this function a test but you might want to call it in a different way and if you look at this this looks like a very standard evaluation Loop written in Python for the task of image classification the output of this function is just the loss and accuracy and that is precisely what we be communicating back to the to the server app to the strategy in particular the loss again we can communicate how many tring examples in this case how many evaluation examples were in this in the data set of this client and we can communicate different metrics some applications in this application which is image classification it makes sense to have accuracy but in other tasks you might want to Federate maybe there is not such a thing as accuracy maybe you have F1 metrics or maybe you have something entirely custom for your use case you would communicate this in in this way sending it as a metric we'll see in the next video how we can for example aggregate all these metrics to have an average accuracy of our Global model when it is evaluated in our Federation so once more as uh was shown in the first video if you if we want to run this this client app for example if I do to see what's in my desktop just we see we have this directory for my app I could CD into it and execute it but that's really stly not needed for example I could I also do flower run and then just give a name of my app and this will execute it it will do it exactly as before do some preparations to load everything in the simulation engine and then k a start the process starting with round one doing fit evaluate then round two FIT evaluate and finally round three before we jump into how to modify slightly the model and how to change the data set let me mention again about the benefits of the of defining some of the hyper parameters here for example let's say I want to do five fated learning rounds in instead of the default three of course I could change it like this if I now do flower run just like before this will run five rounds but this is maybe not so convenient if you need to do these changes uh regularly for example what you can do you can use the Run config uh argument and override any of the settings here defined so for example can change n server rounds and I can change it to five and if I want the number of local EPO to be something different than one let's say two local EPO I will say to two this will do five round of fated learning and made in the same way that the flower engine makes this config available to server app and Client app when we override those parameters at runtime that updated config is also made available to the server app and Client app you can see that we let this run is running for four and five rounds so that's something useful to know and as I mentioned you can add here the hyper parameters that are more relevant to your app what we're going to see now is how we can change our data set [Music] so we know we haven't I think we haven't talked about it too much but if we if we go back to our client function we have this low data that we never zoomed into it if I do right click and we look into it here we have the cyar 10 data set that is downloaded and partitioned using flower data sets which partitioner are we using we are using the IID partitioner what this is going to do is create as many partitions as we set here but extract them in an IID fashion which means that this is going to be a purly unrealistic way of partitioning the data because all clients will have roughly the same number of training examples from all the classes available in this sideart 10 data set and this is fine this is fine to start prototyping your apps but at some point you maybe want to do something different maybe you want to change the data set or maybe you want to even change how the data set is partitioned so let's see how to how to do that and for that I'm going to load we're going to load the fashion em data set and I will show you how you can do that super easily with flower data sets so in order to do that what we're going to do is we're going to go to hugging phase and we can really search for any data set we want and it's it's going to work with flow data sets so let's go for fashion amist if I'm able to find it let's yes here passion so if you don't know about this data set this data set is just like amist and classes but instead of numbers we have different clothes like shoe t-shirt trowers and so on so if we want to make use of this data set in flower data sets we simply copy the name of this data set so I'm going to copy it here just like that replacing CER 10 with flower data with fashion mist and then we know that this data set is gray scale Cipher 10 is rgp so some of the things we need to do are replacing the way the the prep processing is perform so now we only have one chall instead of three if we go up to our model the inputs to our model are going to be RG are going to be great scale so let's modify this from 3 to one and the fashion amist images are 28 by 28 instead of the 32 by 32 from the cyar data sets so this means that we need to do some small adjustments to account for the difference in width and height in some Frameworks this might not be needed but in Python it is so we're going to change this from 5x5 to 4x4 and we are almost there one of the key things to know when you work with Hing face data set is that each data set is com is you can think of it as a big dictionary with different features and here we have a feature called image but if I look for another data set in particular for cyar 10 CER 10 doesn't call the images images they call it IMG G so we need to make this small change in our code you will see right now what I mean by it we don't want to look for the IMG key because it doesn't exist so we're going to replace everything with image which is the way this data set calls them we don't have to worry about the label because label is also the name used by this fashion data set so if I replace image everywhere then we are good to go so just as a short recap what we have done is we have replaced the cyar 10 data set with fashion Mist let me replace the we have adjusted the how the pr processing works because before we have RGB image and now we have a gray scale and because the because of that we have also changed the number of input channels and because the images are slightly smaller we needed to update just a tiny bit our model here and here finally because how Hing face it sets work we needed to change the key used to access our images from IMG to image that was not so difficult what if I want to run this code now uh will it work of course it will work let's run it let's run our app again without overwriting the the Run config and let's see what happens so the simulation engine will do the preparations behind the scenes once again if the data set is not available it will be downloaded and it will be partitioned again using the IID partitioner and the application will run just as you would expect we have run the simulation of 10 Federated learning clients where there are five clients sample for fit and then all the clients are sample for evaluation and this is these are the metrics we get by default at the end which is the aggregated losses the aggregated Federated evaluation losses sent by every client where do these numbers come from those numbers are the weighted aggregation of this loss communicated by every client at every round let's see now for example how we can change the partitioning of our data set this is an optional stage but I think it will be interesting for you H to know about it by default we use a ID partitioner which is fairly simple let's use something a bit more interesting and more challenging so flower data sets have quite a few partitioners that you can make use of if I open the flower website you can go to the flower documentation fairly easily from here documentation here flower framework you will see all the documentation about it but what going what we are going to see now is take a look at the flower data sets and in particular we want to see how let me bring this back to light we want to see other partitioners you can see here all the different partitioning mechanisms that there are available as of today here is our ID partitioner but what what we're going to do is we're going to use the dlet partitioner which is probably one of the partitioners more widely used in the research at least for now so yeah let's replace our partitioner with DL partitioner the [Music] partitioner we're going to replace it here interestingly the RL partitioner has a few other compulsory arguments one of them is Partition by and this what this is asking us to to do to set is what column in our data set should be used to partition and you can look into the documentation exactly what this means but what what we're going to be doing is we're going to Partition by label in this way we can control easily H what's the the ratio of images of it of each class that is assigned to each client in this way we can easily with a single hyper parameter control the degree of non idess in our partitioning so we're going to do this by label and then the other algorith the final um hyper parameter we want to set is the alpha alpha if you set it to a very high number like 1,000 you will the resulting partitions will be near IID if you set it to a very low number for example 0.1 this will result in a very non ID uh setting you can see all this in action if you go to the flower dat documentation and go to the visualized and you run through the visualized label distribution h tutorial which you can open in collab very easily but for now you have to trust me and I will set it to a relatively small value of 1.0 actually let me pause for a second and interrupt this video to show you how this data set partitioner that we are using now the D SL partition let's see how the resulting partitions look like we said that with the alpha hyperparameter we can control the degree of non the lower the more heterogeneous the partition become so let's let's see that in action and for that let me point you to the flower data set documentation and in particular we're going to run just a couple of cells from the visualization label distributions uh tutorial I've already opened it on CA but you cannot you can do it this easily by following the link here and the only change I've done is to replace the data set that was here by default which was Cypher 10 with the fashion M data set which is the one we are using for this tutorial series I have also increased the alpha to 1,000 which the theory tells us that this will create highly homogeneous partitions so let's see if that is the case with this we're creating our own partitioner and we can call what of the buil-in visualization function from flower data sets to visualize the partitions so let's take a look as we predicted here we can see that all 10 partitions from our population they look almost identical all partitions partition zero through nine have roughly the same train examples for all classes roughly the same examples of t-shirt roughly the same examples of ankle boot bag sneaker and so on this is fine if we want to test out some ideas fairly fast and we just want to do some sanity check that every component in our algorithm is working properly but soon after that is much better if we move to a partitioning a mechanism that brings some heterogeneity into the Federation and why the reason why we want to do that as soon as possible is because eventually all the algorithms will develop we want to run them out there in the real world and in the real world data is highly heterogeneous think about it for example if we have if we want to Federate the training of a image classifier that is going to run in people's smartphones you know that not everybody's going to have the same type of images some users might have a lot of images from dogs others might have a lot of food pictures others maybe pictures from their workplace and things like that similarly if you think not about users but maybe organizations and want to collaborate let's say for example hospitals not all hospitals have patients from the same demographics so some hospitals and because of that the overall distribution of data is heterogeneous how can we simulate that easily with flower data sets we can choose a different partitioner or we can customize existing partitioners and with dlet we can do that easily by lowering the alpha if we lower it we will introduce more non so let's see going what happens if we go from a th to a 10 this is not going to be very dramatic but we are going to start seeing some differences previously it was very homogeneous and now we going to start to see that well not all partitions have the same number of clients for all classes so for example if we take a look at partition number seven we can see that it has many less t-shirts than partition number five for example what if we keep decreasing the alpha let's say to one we will see that the difference are much more pronounced for example partition number four has all these many images of sandals and if we compare it to for example partition number seven it barely has any so this is starting to introduce some um more heterogeneity in our partitions and this is going to be a much more challenging setup for our training algorithm and model just before we conclude this Interruption of the tutorial let me show you that yes with flowerd you can very easily increase the number of partitions let's do 20 partitions and everything works as you would expect so this is everything I wanted to show you uh with regards to the partitioning of the data as mentioned earlier there are many other partitioners you can choose from but for this tutorial we will stick with the D partitioner and with alpha 1 which is a reasonable Alpha okay let's get back to the video when I run this the likely outcome would be that the loss doesn't go down by so much because it is again a much more challenging setup so let's run it and see what happens so now that we're running it for a second time the data set hasn't have to be downloaded it's already cach in the system and we can load it pretty efficiently what changes every time we run the data set is how the partitions are created now we got exact the same number of partitions 10 but they are very different in terms of the proportion of classes are assigned to each client and you can see that the loss didn't go down nearly as much as before now we got the the lowest loss we 0.64 whereas before well I I erased it but it was I think around 0.5 so this is everything we wanted to cover this in this video tutorial where we have done a very deep dive into the client app and the server app that is more we're going to update and customize in both the server app on the client up but that will be left for the following videos I hope you enjoy see you in the next video where we will customize the server app through callbox that will be passing to the strategy bye
Up Next

Differential Privacy with Opacus: Training PyTorch Models Privately
@PyTorch
3K views•2020-11-25

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

Secure Aggregation in Flower: Salvia+ Protocols for Federated Learning
@flowerlabs
814 views•2022-06-20

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence






































