Physics-informed neural networks (PINNs) embed partial differential equations (PDEs) into neural network loss functions using automatic differentiation, enabling mesh-free solutions to both forward and inverse PDE problems; DeepXDE is a Python library that implements this framework by allowing users to define geometry, PDEs, boundary/initial conditions, and neural networks through a modular API, with key features including residual-based adaptive refinement for efficient training and support for complex geometries via constructive solid geometry.
DeepXDE: Solving Differential Equations with Deep Learning | PINN Library
Added:yeah is it Lulu from Brian Mosely so in this talk I've the library impacts te for solving both Ford and also inverse differential equations so when we try to your deep learning for partial different equations PD years people have developed different approaches in the past almost 20 years some approaches can only be used on some to some specific PD problem for example if we consider the pink the domain of the PDA is like square it symmetry like domain or if the PD is parabolic then we can use the formula to convert a high dimension quality parabolic PDE into a local PD or if the PD has a corresponding variational energy form so this purchase are specific to the PD problem and they're also general approaches can be applied for any PDE for example we can use galerkin projection or PD and it was in this fog area which focused on the stone form of speedy basically that is the physics for neural networks in this symposium we have already had a few talks focus on this part in stay matters talk and the Paris talk and this morning especially the modules talk and also criest also mention this part just in case you missed all these talks like y'all give up with brief introduction of this methodology using a very simple setup called invisible clapping cloaking so in this project let's consider we have elect of the odd fine which propagates from our left hand side to the right hand side indicated by the arrows and the next let's consider we have a small circle in the blue color in this domain and this part in this small circle it has different permittivity from three which is different from epsilon for the surrounding environment here too the existence of this mosque fear then we can see on a red figure in the downstream the lake of electrical field is perturbed due to the existence of this object and in this invisible cloaking we try to add another layer with sorry which is in the pink color surrounding this blue region such that by special design this eep stone to such that the downstream allegro field as true in the red figure looks like unperturbed which means if we measure the electric field in the downstream we cannot know whether there is objects in upstream or not so this is the setup of this problem so mathematically even if someone and if zone 3 and also given this location and the size of these two circles which kind of fun if zone 2 which is a function of s coordinate and also y coordinate such that the far wall which is the field electric field for the our side domain which is unperturbed which means far wall should be equal to our far one target function so this is the measurements we have for farm and if we try to for these measurements to infer was the crowd responding via the east onto and as we know this this problem can be described by the Helmholtz equation as shown in the here the laplacian of phi i plus y so my K square for I is equal zero is equal to zero for three different regions outside region blue region and also pink region and also we have a the boundary condition for the outer circle and the inner circle so for each circle we have the continuity over the Phi 1 and Phi 2 and also the no more tivity or partial for impartial and the partial of Phi 2 partial M should be proportional to user 1 and user to so this is a PDE and also boundary conditions we have so next we travel combining the information of the measurements for far Wang and also the information over the PDE here together otto in for water is crowd fiat for herbs on to and as I mentioned we are using formula new networks for solving this problem the basic idea of freezing in formed in the analogs inverse simple be we are impaired at the PDE into the most function of neural networks by automatic deformation so let's look at a figure here where here is a sketch of this method let's consider first thing we have a standard Network could be a feel for the network could be ResNet or all types of network so for the input of an analog it has two parts the S coordinate and also y coordinate so there are two scalars and its output of the network include four different fears far along Phi 2 and Phi 3 which are three different electrofield on three different domains and more importantly the last output is equal to which is our target electro permittivity field or the king of ping region so the key part is to the time the loss function the loss function includes a few different components the first component is a standard supervised learning so because we know fire Wang should be equal to follow and target following target is our measurements so this is the standard measure standard a supervised learning loss function and the next video ad is a loss function for the PDE into the network for example because for the other standard region for this outside region firewall it should satisfy this hand post equation then we can select randomly a few of different locations for the other side domain and they evaluate was a causal value for an option for one class if someone K square for one so so ideally this term should equal to zero to satisfy our PDE but initially definitely does not so we add addition again in with a loss function and the Guru thing here when we use Noonan walk in that we can use automatic a different differentiation who compute this order the derivative of our Wong with respect to the x coordinate and also y coordinate so this laplacian term can be computed exactly and similarly for fire to field and a53 feared we have the themelis function and also we have this we can compare the solution for the boundary condition in a similar way so combining all these components together the total dysfunction is the way via submission or we each turns then we minimize is those function by traditional walk to learn what is Phi 1 Phi 2 Phi 3 and also was the epsilon 2 so this framework is it's very simple and easy to understand and also you can see essentially we are so we are solving the inverse problem of a PDE PDE most problem then was the advantage here so it's a first advantage is that this method is mesh free and also even particle free so we can bypass the process of drawing the mesh it's really time consuming no other machines and more importantly peace can seamlessly integrate the data information and the information oh please I lost half of it is like momentum conservation of mass conservation and also in these paper we show that by using some technology in deep learning this framework can be easily to be used to handle the PDS with black box boundary condition initial condition or purpose for interest or even the initial condition of our new condition has a very large noise we demonstrate that we can handle and up to 20 percent no ice on their turns so it's a pretty passed to handle all the noises so just a quick question what do you mean buying black box boundary conditions or initial conditions yes so usually when we saw when we solving a PD we know the exact formula over the boundary condition and in your condition then we can draw the mesh and we know all the mesh mesh palms on boundary and initial but if we only know the boundary condition or initial condition at some location just on some scheduler locations which is already there we cannot we cannot modify then if we use the standard for an element or find it find difference we have to get help at first to our interpolation over the boundary condition or in equation to have the whole field for the initial and the boundary then we can draw the mesh to have the value on any locations the black box means we own in law or we only the variables and some skeletons rather than the whole field yes and there the first point is that this framework can be easily be extended to all types the PDS for example in the corral given equations first our PDE or stock up all these different types of peas can be handled in the same framework of course we need extra techniques and a matter big a-frame before in esteem so there are fewer the one hideous of his method and the Methodist import yourself so after the training of the Nano walk the drug plot is the permittivity of this inner region learn the fontanelle walk with the crowd Chris pawning electrofield will showing the lead figure so this is the method it's easy to understand and next we try to have some understanding of this method so usually when we try to understand a deep learning framework there are three approaches we should investigate including approximation optimization and agitation so first for the ultimate issue so we oh so we know that no networks can be used to learn any continuous function he was an animal science is large enough so this is guaranteed by the universe' approximation theorem show them lots of years ago but here we added a PDE information into the host function then the first question is that whether or not if we increase the network size we can always find another one a network service that loss function go to 0 which means whether the loss function including the 30 root of the network can also go to 0 and this is the truth which is proved by pinkos in 1999 basically this theorem showed that the network cannot only be used to up to approximate a function and it can also be used to approximate the partial derivatives of functions simultaneously so this this theory is much more corn than the standard in most offices post-mission theorem so this is the first question so we know that the loss function over the network can go to zero if the lower size is large enough ideally and then when we have the loss function V is the PD we need then what was the augmentation behavior and here I've had to use two of simple examples to show that by using the PDE in the loss function then the training behavior or the learning mode of the network would be totally changed let's consider a two over simple cases so for the cases we used a standard then I walk to learn a function which is FX is equal to a summation of five times as you can see he quit this FX has five different frequencies so we do a standard of supervised learning we stand pause from doing some data points for this function and the Trini Network then let's look at a figure a to nab the figure so we have five crew frequencies of this target function so during the training of dialogue we do the free free analysis on the network output and if we investigate each components for example let's look at the first frequency once the red color sorry when the color turns red which means on this specific frequency the network matches the our target function so as we can see for the first frequency it requires almost two thousand iterations to make sure to Nana walk matches the target of frequency however when we go to high frequency for example the last one then I roughly it requires twenty thousand iterations so as we can see from this figure the network learns from low frequency to high frequency so this is for the standard approximation approximation of 1d function and this behavior has also been observed in image cast vehicle from image classification problem like image net as shown in these two figures for in the two papers last year so these are the true in general but if we consider the pings let's solve a Poisson equation with exactly the same function as the first one so the solution of the Poisson equation the same as case a then we can do the same thing we can look at the graph figure the drug figure is a training procedure of the network on this Poisson equation there as we can see from this figure clearly or all the frequencies are learned almost and somewhat changed somewhat easily rather than for low frequency to hyper with it so in these two figures you can see by adding the PDE into the solution the training directory or the learning mode of nano walk would be totally changed so this is our new market among we do not have the theoretical yes Syria yet but for numerical examples we can see how this noon training or learning of nano walk we changed from the pings so sorry I can just clarify it to the in a you just train directly the non network on your function whereas in B you use pin to to calculate the f of X yes yes for a in a I use a standard analog to twing twing effects some points so standard supervised learning and impede videos the pin of your videos pin to solve the Poisson equation and it's a solution of equation is the same as a solution over the first equation yes the solution are the same okay then then let's move next so as I mentioned before when we try to use pins for a peep for swapping a PDE we have to evaluate the loss function or some locations sampled on phone the field and now you will a naive way is that we can use the uniform haunts or the domain the worst in many cases however this uniform Holmes may not be efficient for some PDS for example let's consider poverty pigeon because the bulk of the equation the solution of the bogus equation is very sharp around the point X equal to 0 then if we use the uniform points then we need a loss of points to achieve already good accuracy so it's commedia or not it's not efficient and here we try to propose a new method search that we can today yours last training pause to achieve a rather good error the method of you purpose is called residual based adaptive refinement so idea a very simple it's very similar to the adaptive refinement mesh refinement in the fun element field so the idea is that as follows when we train to network we can adaptively add more and more points on the locations where the PD Raider is large for example between the network on 1,000 iterations then we check was the PT residual on the domain on low so cones and then we select those points with large errors as our new training points or in the deepening field in given any community his master can be called as active learning and here I draw two figures so for example as in Figure C they are for your crosses this cross points add new EDD points during the training so as we can see this method can automatically figure out where to add a new the new points because around the axle okay X equal to 0 in this Shabab region we should add more and more points so this this framework automatically figured we should put more points in this domain and as a comparison if we use the uniform jersey reports to achieve the same accuracy we need 10,000 training pawns on domain but if we use is adaptive way we can reduce normally pawns from Ken thousand to two thousand roughly so here I give a brief introduction and some understanding from different angles for the pings and in the second part of this talk just a quick question on the adaptive training there yes have you looked into using a general generative adversarial networks for improving this training because I mean that you I mean this is that would be the standard technique for figuring out where a neural network is most likely to be incorrect right and so if you have something most likely incorrect that might be a nice way for do it for building an adaptive algorithm yes yes hey this this idea I saw Billy already be used in a fun element mesh refinement we haven't using our frame over yet but a definite is that would be reinvestigated yes that's right yes so in the next second part of this talk I will give a brief introduction of the usage of depends te so in summary in depends the solving differential equations is no more than specifying the program using the PRD modulus so as we know a PD problem should be defined from the geometry the equations and also a boundary initial condition and then when we use networks we need to we need to specify what's network for the size or was the type of the network so in the following examples I will show you how easily we can use depress the for solving different types of PDS by only specifying these components then all the remaining part we Arpita automatically by the software the first example is a very simple I'm independent problem post-invasion let's consider a possible inclusion or 2d domain in the air shape these terrible conditions so so so force online we need to define the geometry of this region which is the polygon so we can use D de Kock geometry for polygon to define the polygon domain so Shyama tree is the modular indepence the year that handles all different geometries in 1d 2d and 3d then we need to define what's the PD we try to solve by on a magical different nation and as the only grammar here we need years from tensorflow is T a baton gradients so tariff gradients will compute that you the first order derivative of the netherworld output with the respected and net above input so for example we can use TF gradients Y X so let's go back one line so 91 we need to find was a PD so this X is network input and Y is network output so when we generally can use chapter Aquarians 1x to have the first audio cue then we repeat this again to have the second order cube and then we combine the second order to him or Y for X and what also addition to match the notation this Ruby yo y X and D Y D Y so this is the la pasión term and minus one so these already draw over the PD so we can define the PDE using he after audience that next we define was a boundary condition so the partition could be you should know you man robbing periodic or even a more general boundary condition defined by any type of equations so we need to define the boundary condition we need to specify two things the first was in the location of a new condition and the next was the values for the body condition bajo de deux chevaux new keynesian and a by default he pecks de would use the entire geometry boundary as a boundary condition of our PDE of course you can change which location to be the boundary condition for example I only want to my level part of my domain two maps we map our new condition oh I need to find some partial part of my boundary to my panel collision so budget positive as default if your the entire geometry boundary as my new condition and next we specify was a value for the dish American user and then we combine these two together into D de edad de chelly but gives you a PC so we have body collision so in this slide we define what's the geometry or the PDE was the boundary condition the next we combine everything together into another class called data which is the virtual data for PD so we define our key de tournay natal PD as the geometry was PD it was value kanesha and then we specify how many 20 pounds we try to use inside the domain and also on the boundary then next we select was Network we try to use for example if we can use stuff it feel fold and we can also use ResNet or other types of network here I use a feel for network as a example when a specific specify was the size is for hidden layers with 50 neurons player and the activation function is high poly Canyon and wins the collage color a uniform initialization method then we compile this model this data net using the atom in each other.you method means the linear 8 10 to the minus 3 then we say okay let's twist a network with 50,000 times so just took a question the compound is the static thing in ten to the floor or what is it study graph yes so in the compiler thing in a way or to it if your period is a tensor flow of static graph and everything for the training and so if you use the eager or if you spray torch no you can use paid torture contiune this device t only supposed tends flow whereas a future walk I still consider to use Python as warm backhander yeah so yeah so here these a few lines as you can see a series ripped apart just combine everything together and specifies some values for the training for the network size and so on and as you can see here tell those thoughts here for each component was a data was a net and was a training because in Paris the Consulate's supports also useful techniques for example for the compiler by default we need learning rate was the optimization method but if he can see more for tomorrow we can we need to say okay we try to use a language scheduling we try to decrease lineage during the training or we we are to the early stopping once the more eternal any rate decreased to our search hard we shall stop or if the language never changes based on magnitude leisure stopped so everything here there are lots of features can be selected from by the users but it's the slide show here is the main part then of the tweeny we can see for this airship bottle occasion the ping can achieve relatively good area compared to the true solution from spatula mode then we go from an independent problem to time dependent problem let's consider a division equation in 1d space 1d space and also one it Han there are only three components are different from the previous program so first this geometry is time space domain so besides to be defined it the physical the space domain as a using interval for minus 1 to 1 we need to define the time domain through from to be from 0 to 1 then our space time domain is just across the product between our space domain and all autonomy so in this way we can very easily to define the space time to me then we define with an initial condition so to define the initial condition is almost exactly as what we do for the boundary condition so by a few lines and then we define the data by combining geometry PDE boundary condition in the air condition and a specified was a number of trillion points inside the boundary or the dough or inside its domain on the palm tree and also on the initial part so as you can see when we call from time eating my dinner to accompaniment basically everything is the same except that we need to find extra geometry that in the still example you go from one purple one different equation to a system of different equations furthermore here I considered Louis a system which is OD system as shown here the only difference of this problem in that when we define this PDE Khan today we have a system so we need for this system we need a three components of this equation all over this problem each component is cribben is corresponding to Y equation on the learning system so we have three graduals returned by this dysfunction and similarly we need to find three initial initial boundary condition one for each component is Sh what is oh dear sister the next question next problem is we go from PDE PD system to integral equations so as I mentioned before the framework can be work can be very easily to expand in to the integral differential just and here I show you a y example how can we solve integral different equations using the bank's t let's consider the award Hera IDE is defining here the right hand side is the integration with the kernel exponential kernel then to define this problem first we need to define was a kernel or the integration is its exponential then by defining this kernel then the software would automatically construct a curve swanee matrix basically the integration matrix for this in equation kouno so this matrix will be cured automatically then we can define the right-hand side for the integration which is the unique which is the matrix times this Y vector then for the left hand side is exact same as previous in your TF taquerias to compute the left hand side d rutube then we combined the left and the right so our idea problem and here we also have the data we have the Kuna we have the IDE and here as I mentioned he sold where we automatically construct is Colonel there's matrix integration matrix and the insides of software we use the constant quality for the integration so you need to specify was a culture degree for it about 20 we use 20 degree and after training we can see we can successfully solve this water IDE and that the last example is inverse problem so we go from food problem to us problem and as you can see we only need a few lines for defining and solving a inverse problem furthermore let's consider a diffusion reaction system and that the diffusion equation T and it is a reaction configure KF are knowing we learned from our measurements so because D and K F are also variables to be trained up to be learned tell me that you find K F idea as T after the variables before we define this PDE so this is the first step and in the same step we need tell or we need input support measurements into this framework for example we have some measurements or the see a whistle a concentration of this a component we know the measurements of CA and some locations observe X and the corresponding measurements is observed CA so these our measurements then basically we use two extra lines by defying by defining these measurements as specific as a special type of each upon you Keynesian so basically we just a copy paste these two lines then we can define these measurements inside our depends the framework and every is exact same as would be in a previously then he trains on a walk and if we can monitor was the value of D and was a value of K F give inning and after a few iterations we can learn the D and K F successfully so one one feature of capacity is that it can automatically handle very complicated geometry using the Tanika called constructive solid geometry so in keeping see here are a few primitive geometry already defined including 1d interval and a 2d triangle rectangle polygon disk and also cuboid also come on and a 3d cobalt and a sphere so these are two nafeel geometries in 1d 2d and a3 or did you find and what we define yes three okay can I ask a question here yes so regarding the boundary conditions I mean in the in the pin this is this is typically so it goes in the loss function and so on so it sort of feels that straight up on and say you want to apply directly boundary conditions oh well the you could just specify like points and values for the solution of those points right so at some so so my question is do you know why why do you need to have this support you know to describe the geometry and so on rather than just having the users who specifies you know just coordinates of points so first if the geometry is is communication for example polish going here we try to solve the problem for suppose we solve the PD power on the geometry then if we manually define all the points then basically we need to define the points along this seven different segments is fun consuming it's doable but if what if especially to a 3d if we have some complicated a 3d geometry then we need try to specify or manually choose the locations on the geometry is it's boring and time-consuming so here I needed I try to give a very easy way to select what is the boundary condition on the geometry you so basically mist makes the user so easy to do easy to put behind the problem so the ways we will have a few standard geometry then we define three boolean operations Union difference and in ensuring this loud figure if we have a is a rectangle P is a circle then this union is a battle e is a plus B and a difference is a minus P and intersection is the region between a and B then for any complicated geometry we can easily control its trauma tree by the simple polygon circle and in return go using the three standard operations so just why as you can see in the previous example the user only need to define what's the boundary condition and yeah basically what what the locations are used in the boundary condition and a by default or the physical boundary condition are in the boundary condition so then if then the user can still turn by selecting order no points on hungry especially for this type or communicated a geometry so quick question here yes when you're representing it for example let's say you use a sphere are you still going to use a 3-dimensional neural network or you can wear a motorized it to be two dimensions on the sphere right are you using the manifold coordinates or you have to use the rahbaniya coordinates well it's the fear basically the importing is X Y and a Z oh yeah the coordinates currently I use a standard coordinate mm-hmm and then do you just have to make sure that you sample when doing the training process you're just sampling from the sphere or is there more that needs to be done in order to make it so that way it's only learning on this year right because it's it's it's a more general three dimensional object I guess it only needs to satisfy the PD on the sphere right that's the condition you're holding I know so fear means the we have a PDE defined in this industry as the boundary condition is on the surface of the sphere but we have also PB Insider that's okay okay yeah yeah it's not 30 not pretend the 2d manifold in the 3d space okay it's truth as as I demonstrated from the five examples to solve a problem using this Ebanks te the users is only defined was the geometry was the equations with boundary condition initial condition and select was the inner network and it is no redundant code so then the user do not require to select where the training points manually so it's the then the code is really it's a really short and comprehensive and the user also do not require know that has flow so the only kind of follow grandma the user required that is the TF got Williams and the cantle a depends the support or most as almost all yours for techniques in deep learning field especially when the planning for the for solving PDEs for example we can cook for to do the answer in qualification or we can do the earliest pumping or we can dynamically serve saving and loading and walk there are many years for techniques has already been implemented in the software roughly all techniques we require in practice and even if the currently the kind of software does not satisfy your new requirements for example you have some verse specific language like the yes the empowering Paris talk he has a new idea of new network which is not implemented in capacity in yet then it's very easy to do because the framework of de Paris de is very structured well as a traitor and a highly configurable each component can be so so the user can insert a very small piece of code in each component and it works and they are also attending a course called core back which means during the training without going into the details or without modifying the instead code of diversity the user can plug in some code to modify or change the behavior with training within a callback mechanism and also continuity backs the support supports a few other functions the modified as planning we proposed in these people and also learning nonlinear fairy-book we proposed in different in this paper so there are extra functions somewhere can be installed by either people or Conda and also here the web starts over and a welcome for reporting any box suggestions or pull requests yeah yeah okay thank you very much for this presentation so we're quite a bit behind schedule so maybe so maybe what maybe one or two questions I have a question when you sneeze stop the work yourself forward problem yeah because in traditional you know traditional methods you need to for example for time dependent equations you need to care about stability like CFL conditions since like that and in your framework you you to not and Aetna users choose the scattered points so do you have any mechanisms so that you choose the correct scale points q current in a stability or convergence yeah actually this is a good question so just a break he knows convergence of this methodology of pins is the operation so we have numerous examples to show that in fact we to know for example for the for Sami PD usually the standard among muscles we need a CFL condition and peace be to knowledge ratio for a CFL condition but it still works I'm not sure why so this is the operation I think I think that there is an answer to that you're not actually evolving things you assume that U is a function of X and T so you learn a function of X and T and then if you want a later time you evaluate it at a later time so this does not give you any trouble with stability but at the same time its restrictive because if you want to go into x which are out of the domain on which you train you cannot so my thing it's explainable y-you don't have an issue because you actually learn a solution you don't learn an update you don't learn another loop in evolution operator yes I mean look if you look at your setup you you learn U of X and T right you don't learn how to go from one time to the next so there is no problem with stability because you don't actually have to evolve something by iteratively applying it like you would do let's say with an Euler method or with any method to solve a PDE yes time step here you're positing for me information is job yes but you the representation of the solution in a domain a spatial temporal domain X and T you see here whatever picture but you're the treating the information is there propagates from the initial condition to a longer time right yeah I mean it's like if I give you if I give you the solution at at time zero point five you don't have a way to tell me you know how to go from 0.5 to something else you'll learn the representation of the solution of the PDE in the whole domain spatio-temporal domain I don't I don't see any evolution in your setup yeah I mean so this the spin method is mortis familiar with or more similar to things that are fully implicit PD solvers right so if you look at some of the old control theory literature you'll find stuff about fully implicit discretization and they have similar I mean they have really good stability properties for some my reasons and then I think that where you probably find most of the proofs here are probably with people who've used radial basis functions for solving partial differential equations in a very similar way there have been some stability proofs but it really does come from the fact that it is fully implicit but I agree there then if you learn a fully implicit discretization then you cannot necessarily evolve it beyond the bounds that you that you placed in the first place yes yes yes because think of it I mean you specify the domain of T to be 0 1 you specify the domain of X to be minus 1 and 1 so if I ask you for the solution that T is equal to 2 you have to actually retrain because your solution was I mean the the the representation of the solution was learned inside this spatial temporal domain so if you want to to predict the value of the solution outside of that let's say the time of 0 1 if you go to 2 it's like trying to do extrapolation and you know that neural networks are not good at extrapolation like that yes that's true yeah so so I think this is why you don't you don't have basically it's the stability constraint okay thanks any other questions or so actually I had a question about you say intensive for knowledge but you need to define your new network and so on right in other words or it predefined in them in their software but see well I was a user okay yeah you're the only little cord is functions that's it and then why don't you let the user define the new network intensive flow oh it's to about so as I mentioned this modular can be easily extended about it so for now magically that's basically it define apply to the software reveal a few standard networks
Up Next

Introduction to Bivectors: Geometric Algebra Explained
@sudgylacmoe
11.7K views•2025-01-29

Conformal Geometric Algebra Explained: Lines, Spheres & Inverse Kinematics
@bivector
1.5K views•2024-09-05

Fourier Series Introduction: The Big Idea Explained
@DrTrefor
387K views•2021-05-03

The Mathematical Impossibility of Accurate World Maps
@Vox
23.3M views•2016-12-02
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Mathematics
















![[Deepmath] 7.4. Différentiation automatique](https://i.ytimg.com/vi_webp/xQKw1xqNtvE/maxresdefault.webp)








![Deep Operator Networks (DeepONet) [Physics Informed Machine Learning]](https://i.ytimg.com/vi_webp/CDCyOHXDRcI/maxresdefault.webp)












