Visual localization is the problem of estimating the six degrees of freedom (three for rotation, three for translation) of a camera pose given an image captured relative to a reference scene representation. The field encompasses multiple strategies: image retrieval (finding similar images in a database), hierarchical localization (combining retrieval with local feature matching), absolute pose regression (directly predicting pose from images using CNNs), and scene coordinate regression (predicting 3D coordinates for each pixel). Key challenges include photometric changes (day/night), low-texture environments, viewpoint variations, and scale differences. Modern approaches leverage deep learning, with techniques like NetVLAD and ACE achieving state-of-the-art results by learning to match features and regress poses from images.
Visual Localization Methods: From Image Retrieval to Pose Regression
Added:Thank you very much for the kind introduction. I'm very sorry it's a torture for you to pronounce French name. I get that you know that's a that's fine and you did a great job actually. So thank you very much. So like good morning everybody. I I think this is the last day you might all be very exhausted. So don't worry today it will be very practical. I I won't like throw at you too much equations you know just you know fasten your seat belt sit enjoy that that's not going to be difficult. Okay. uh we'll talk actually about something rather applicative. This is um visual localization. I will try to give a sort of overview of like the different strategy we can use and at the end of this talk I hope that if one day you need to address a similar kind of problem you would be like okay I know what to use what tool to deploy and you know for further details then you just need to dig a bit in the literature.
We'll talk about old stuff more recent things. So it's it's quite a broad overview you would say. Um so the topic we'll talk about is visual localization and to illustrate that do you know what is geoger?
>> Do you guys have ever played Jo guesser or do you know what is geoger?
>> Yeah. So c can you explain what it is to others >> exactly? So basically you have a random picture. It says that's a very good one.
Actually you describe what is visualization. You did my job. So you have a picture and then the user has to pinpoint the geographic location where he thinks this picture has been located.
Okay. How do you think we can do that? I mean as a human we look at the image and we're like hey this picture is from Nepal for instance. What could make you think a picture is acquired in Nepal?
You know you might have some clue. You might see the language. You might you know like maybe those kind of grass is growing in Nepal. know it you know and this is embedded somewhere in your brain and sometime the clue you use to localize yourself to know that this picture is from Korea this picture is from there we cannot always really um really say like oh this is because this or that this is just a overall feeling now you know like so sometime it's a clear landmark sometime is not it can be a texture it can be anything you know and even those uh geoger players but it's quite funny I don't know if we talk a lot about that in this uh in this school actually um they they will take advantage of certain aberration of the image certain like um how to say uh glitch or you know they know that a certain car appears in Australia so they will guess so they use also some I will say they're cheating a little bit and your neur your network will do the same thing actually you know that's funny to see we take the same shortcut anyway I'm drifting too much away here but yeah basically visual localization is the idea of how we basically localize ourself self like using a picture. Okay, so it's a bit like a GPS but using images in some ways, right? Um, okay. What I wanted to add to that, um, so yeah, we I before we dig further into like the technical details of that, uh, I want to interrupt this presentation briefly for a short commercial break.
Okay, that my university tells me you have to do that. Okay, and people are curious actually. So because I'm a kind of weird guy, you know, I have three hats. One in Nami and you know that's one of the reason I'm here today. Uh I could collaborate with fantastic people here in Nepal but most of the time I'm actually located in Korea but I'm actually working for an American university which is located in New York State in Stony. This is Stonybrook University in Long Island. Okay. But we are basically just a branch of Stonybrook and we're called Sunni Korea.
So basically we offer exactly the same kind of program as the student will receive in the US but at a well located in Korea if you guys don't want to be in US basically um so we offer the same diploma everything is the same actually just the location is different and then the fitting is different the students are different as well and all our students are required just to spend one year in the states okay that's uh so this is you know different experience that they can get we are opening now so we also offer PhD and master program and we open a new master uh degree about data science. So if you guys are looking for a master degree you can type sunora you can contact me whatever okay so feel free to do that. Uh now let's go back to visual localization. You know before we start talking about that again I had like I was talking with with some research assistant in NAMI and I I have to confess something about this. You know I think I know exactly how to resolve this problem how to localize it in camera in space at a sentiment level using a computer. But after two weeks in Nepal I realized that as a human being I'm the worst visual localizer ever.
Okay. So trust me in the technical things but if you follow me in Kmand do like yeah don't trust me. So we have seen like a rough definition of what is visual localization with you know like using this uh you know image of geoger somehow like let's try now to see a more formal definition of what is visual localization. So visual localization is the problem of estimating six degrees of freedom of a camera pose which a given image was captured relative to a reference scene rep reference scene uh representation. I think I lost my pointer. Never mind. I don't know where I put it. It's okay. Um yeah. So basically what we want is more than just you know rough approximation of where the picture has been acquired like in joer. Oh yeah the pointer is here. Can you sorry thank you very much.
So we want more than just a rough you know approximation of where the camera is located. We really want to know exactly where the camera is in space and we want to know it's it's okay I it's it's on. Yeah. Uh we want to know exactly where it is and you know like always we start with like a representation of the environment. So basically a map it can be any sort of map. Here the example is a 3D point cloud but really it can be anything.
We'll see more example later. And then if I give you this query image the question is like where is this image being captured? And we want to know basically exactly where it is. So we talk about six degrees of freedom. What are those six degrees of freedom >> rotation and translation right? Because here we work in a 3D space. It's not always like this. It depends like in what environment you want to localize yourself. But in general this is what we want indeed. like you want to know like three rotation three degree of freedom in translation XYZ in the 3D space and you want to know where you're looking at with this camera. So it's basically three additional degrees of freedom for the rotation. You can think in terms of your angle like you have run the rotation around each axis you know row pitch somehow and then it will tell you precisely where this picture has been acquired in the 3D world. So can you guess actually where this thing can be useful?
3D >> definitely in 3D reconstruction uh in many other stuffs as well like in autonomous driving you know like what what you don't see in autonomous driving is like behind the hood of the system there are like uh a huge basically localization in some HD maps that have been created that basically uh contain all sort of information about the road environment it's and that's totally right used a lot in 3D reconstruction so we'll talk later about image retrieval and all those kind of thing and this is something that is at the center of like structure for motion and many other stuffs. Uh I think the the first thing that comes into people's mind when we talk about visual localization is like you know those Pokémon Go kind of game you know or like any actually augmented or virtual reality you know like if you want to place some virtual object in the scene very often you need to relocize yourself in that scene. Okay. Uh if you guys are doing a little bit of robotic navigation slam then typically we use those thing you know to avoid the drift over time of your motion estimation for instance. Okay. Uh it can be used for any mobile localization. Here is an example from some friends in neighbor lab in Korea uh where they basically tackle some very challenging problem in a very complex environment. And there are many more than this. You know I even have a friend he's trying to localize a camera in your digestive tract. So he's mapping the you're inside using some so that you have to swallow a sort of camera and then after for the next operation you can take another pill somehow and it tries to localize it inside. So there are many many things I don't work with those kind of uh medical you know challenges thankfully but there are many application of visual localization.
So you can also use like you can localize yourself in different representations you know we are talking about Joe guesser what is the representation in jog guessesser what is the map in which you localize yourself >> is in jogresser is not a 3D map in jogesser you need to pinpoint your location on what >> yeah it's a map right this is the planet actually that that's the entire earth right this is a projection of of of a sphere in that case right so how many degrees of freedom you have to solve is two only the longitude latitude right that that's an easy problem I told you that's in general we we we will try to go further than this we'll try to estimate six but again it depends in what kind of representation you want to look at yourself so one example I gave before is a 3D point cloud okay like this and in that case probably you will have to address six degrees of freedom um we will see another example that is much more common than just a 3D point cloud is basically what we call a structure from motion 3D map and we'll discover that those kind of representation they're the most widely used for place localization because they contain very rich informations okay and it can be some other stuff you know so something quite similar to what we have seen with joer it can be a road map okay we call those things actually SD map and there are very interesting paper those day like coming up with with that okay uh it can be a blueprint It can be a CAD model. It can be anything you know you can be creative whenever you have like your representation of of an environment whatever it is then you can think of a new localization problem basically.
Okay. And the point of visual localization of course is not to use any GPS or you know you can eventually add them as additional information but I will say the main goal is like how we can replace any other uh sensor just using images. So the most common input you will provide to a visual localization network is basically a basic perspective image from your cell phone. You know you have no distortion.
Easy peasy. Okay. And in general we start with those kind of images to make it simple. But there are actually over kind of modalities you can use that will be very beneficial for to localize yourself actually. And one of them that I particularly appreciate is omniirectional vision. So for instance instead of giving a simple you know like low narrow field of view image you can also use those kind of omnidirectional image and you see that it will be it's actually in general very interesting to use those kind of images. Why? Because you have much more landmarks. You have much more chance to see something that is of interest for your localization because if you use a very narrow field of view and you look at the ground that's very likely you won't be able to really localize yourself. Right? So omniirectional images are interesting.
You can get those kind of images using some uh catadoptic system or multiple cameras whatever you want. Okay. Uh if you want to localize yourself maybe during the night or you know you can use thermal camera you can have some prior information about the depth using stereo system or you can even use like chunks of videos you know and also it will be more rich in information. Okay, instead of just using one image and know you can take any modalities you want. You can take even camera something very exotic if you want. Okay, to some extent some people will assume as well that localizing yourself with LAR input will also be considered to some extent to visualization even though it's not really an image you know but anything that a perceptual sensor will give you will be considered as visualization input.
But you know it sounds like a simple problem right like I give you an image you tell me where this image has been located in a given referential. Okay but in reality this is much more complicated than we think because there are many challenges that occurs when you apply those strategies in a real setup. Okay.
So one of them that is very obvious is that imagine like you capture your scene originally using you know like uh during the daytime. Okay. And now you want to localize yourself during the night.
Well, now you have a little bit of a problem because those scene they don't look alike anymore, right? And then how can you really know that this image in the dark is the same as this one? Okay, even though you see the viewpoint is very close, nonetheless all the descriptor of that image, they will be widely different. They will be completely shifted actually. So you want something that is invariant to those phototric transformation.
you have other challenges like for instance how you localize yourself in like um an environment with very low texture and even for us human you know this is very difficult it's like you know when when you're in hospital like the corridor they're always the same how do you know at what floor you're you're located you know those kind of thing like they're also very very complicated to deal with you can al you can also have like some drastic viewpoint different like if I take a picture of this room of you guys from here or a picture from where Suman is located right behind you. That's what he's doing right now actually. Then it's very you know like the viewpoint will make the scene look completely different. So that's yet another challenge. Uh the seasons can be different if you want to do long-term localization or sometime you can have occlusion when you work in a dynamic environment. Some of your landmarks might be basically detected on some moving object and then well good luck to localize yourself after that.
And another challenge might be the scale. For instance, imagine you want to localize yourself in a huge city is not the same as if you want to localize yourself in a tiny room. Okay. So, and now I want to so here are the challenges. Now I want to talk a little bit about the metrics because later we'll see different techniques and then we might need actually those things. So the metric is actually rather simple.
You want to basically know what is the discrepancy between the translation you have estimated and the ground truth translation. So that's just the L2 distance between them. That's a very typical one when you work in 3D. And for the rotation, that's a bit different.
That looks a bit more exotic, but this is actually quite simple. We basically use this formula most of the time. And it will give you a symbol num, a single number that will be representative of how good your rotation is. Okay? Because if you look at this, you know, like rotation matrix, there are some weird animals, you know, there are orthogonal matrices determinant one. And another specificity of orthogonal matrices is that if you multiply itself with is transpose or inverse then this thing should become identity. So if you did a good job at estimating your pose then it means that this thing here will be an identity matrix. So if you take the trace you should get three. Three minus one is two. You take half of it it should be one. R co 1 is zero. And you see you have something that somehow will be representing your rotation error. So we use basically this metric when you want to know for one sample how good you have been you know in rotation and translation but when you work like on a large collection of image we tend to actually use the success rate as a more robust metric. So the success imagine like you you have like a data set that you want to localize 10,000 images you know in uh in an environment. Imagine you did a super good job. You're super accurate for almost all the images, but one image is located super far away. You have a huge outlier. So it means that the mean of that will be bad. Even though actually your algorithm, your network is good. Okay. So what we do is we just look at the percentage of image that has been well categorized for different level of accuracy in rotation translation. And if you read basically visual localization paper most of the time we rely on those kind of metric to determine like if we did a good job or not. Okay. So here there are few example you cannot see well but I guess you will have the slide. Um and then also the community and visualization is very big and depending on the challenge you want to tackle I can guarantee you that you will always find a data set that you can use to basically test and train your strategy. So some of them will have some challenge for day and night. Some of them will be some different weather condition. Some of them will be indoor.
Some of them will be like on very large scale etc. And you know there are many many more than this actually if you want to test your technique. So now like when you address the prem of visual localization you can sometime depending on the application you might have different granularity you might have different like basically uh accuracy that you want to achieve. Okay.
And like there are basically two level of accuracy that in general we consider in visual localization. The first one is uh image retrieval. Okay. So the idea here is you have you're given a huge data set of image. Let's say you have 10 millions images of different place across the world. Okay. And what you want to know is what image in this database is the closest from your query image basically. Okay. And then if you find an image that look very similar to your core image, well, you can start from the reasonable assumption that hey, they have been captured more or less at the same location. Okay, so if you're lucky and you have basically all your image in your database that are geocalized, you can say I'm in Paris here. I'm in front of the Lou. Okay, something like this. But of course, it's not very accurate. The real goal of visual localization in general, this is basically to localize yourself in like much more accurately than this. Okay, at a very exact location where the image has been captured and it's a topic that actually has attracted quite a lot of people for the past 20 years at least and people came up with different type of strategies. Um here is a taxonomy that I really like that has been uh proposed by some guys in neighbor lab Europs they're like expert in that in that field and here it looks total mess for you at this point. Okay, my goal through this lecture hopefully is that at the end of the lecture you come back to that thing and ah it makes total sense. Okay, I hope I'm not sure but we'll cover actually most of those strategies. So it's maybe a bit ambitious. I hope we'll have the time to finish. Um and well hopefully it will be more clear to you later. Okay. So before we directly start with visual localization we need to understand in what we want to localize ourself and most of the time we localize ourself in a 3D map that has been creating with structure from motion. So I will briefly talk about structure for motion that's not the core of the lecture so there will be no details just I want you conceptually to understand what is going on when you basically try to make a tree construction using images. So structure for motion is just how we is the the I would say the oldest probably strategy to reconstruct a send in 3D from a set of images. So you start with images of whatever you want to reconstruct maybe you can take many images of that room and then you can basically feed so like here for instance is the pipeline from co map. So call map is if if whenever if you want to start really uh like what I encourage you to do is you download call map and you try to push some image in there. Okay. And the way call map is working or any SFM is working is that first you will extract some local features on your images. All of them you can have like 10 images, 100 images, 1 million images doesn't matter. You will extract some local feature that looks interesting. Okay. And you will basically try to match those features across images. So you see they're local feature on the top image here. This is an example of image matching and you will do it for any possible pairs of images you can have. Okay. So here I will say this is the most difficult stage in 3D reconstruction. How to find those correspondence is very difficult.
Okay. What comes next is basically pure geometry at this point. Okay. Like given those pair of images how I can together find the poses of all my cameras and how I can triangulate all those point to reconstruct the scene. Okay. and we do some you know like additional refinement whatever okay and you should end up with actually the beautiful tri construction of the scene all the poses of your cameras of the of all the images you use etc etc so here is somehow how image like feature matching is working imagine you want to find the correspondence between those two image images you can first detect some distinctive points on image one I'm not sure if you see them very well but there are little red dots here that appears at some somehow interesting location, something that will be repeatable somehow. You do the same on image two, sorry.
Okay, you detect a set of key points and what you can do next is like you will try to extract some meaningful feature in the locality of all those key points.
Okay, you extract a little patch and you can use a neural network or whatever you want to basically have some representative vector for each of those local key point. And now what you can do is you can try to associate them one with another. Okay, this is a bipartite matching somehow like um and for instance you take you this first descriptor you can be like hey which one is the most similar to me. Okay and then you will be like maybe they're matching and you will take the next one and you will try to do the same and again and again. Okay so if you're able to do that then good job you are you are done with image matching somehow. Okay, except that you can have some outlier contamination. you can have many problem that can occur. Okay, so just to come back a little bit to key point descriptor because we'll talk a bit more about this. As I mentioned before, this is basically a vector that is encoding the local appearance of a key point.
Okay. And what we want is those uh descriptor to be uh phototrically invariant. So it means that if I want to localize, I want to match point between an image I take now and if you turn off turn on all the light, the light will be changing. I still want basically those key point to match and this should be ideally invariant to geometric transformation as well. Like if I want to do matching between widely different viewpoints, I still want those local key point to match. Okay. And for a very long time we were relying on endcrafted features like sift where you know you take this little patch you will extract the gradient locally and you will basically create an histogram of this gradient and you can consider that this histogram is representative of this local patch somehow uh we had some binary strategy that are still in use in uh in uh in slam or any realtime system and a bit more recently I would say in the past like seven years at least then we have like we have seen the emergence of tons of learn feature here is for instance super point where a neural network is directly regressing um the 2D key point location as a heat map and the descriptor for each and every pixel in the image somehow okay and those learn feature they tend to be better and better at the beginning they were barely uh competing with crafted feature now they are like very interesting thing coming out that's a very interesting part of the literature here. Uh anyway, so if you're able to match images and if you're lucky enough, maybe you will be able to have at least five key points that are matching. If you have this, then congratulation. You can estimate what we call the essential matrix. So you can basically together estimate the poses between your cameras and you can reconstruct the 3D structure of the scene by triangulating those points.
Okay? So you can do it for two images.
That's fine. So now you can basically you know have like a sort of seed map from two image you will create a tiny map. Of course it doesn't cover everything. So what you will do next is that you will try to aggregate new images to that because you have this tiny threepoint cloud. You can try to localize another camera in this global referential. Okay. And now you know the pose of this camera. So you can basically up match points between this new image in your map with the previous images and you will triangulate new set of points. So your map is going to grow and then you can aggregate more images to that and you will end up idally with basically your entire map. Okay. Uh if you if you end like adding all the images in your data set. So of course it's a bit more complicated. Okay. You have some refinement going on etc. But at the end of the day, what you get this is those kind of 3D representation. And what many students like they're missing when they start 3D reconstruction is that they only care about the point cloud. They're like, "Hey, good job. We have the point cloud. That's what we wanted." In reality, um like the output of the structure for motion is much richer than just a 3D point cloud.
Actually, three point cloud is not maybe the most interesting in what we we are estimating. So what you end up with is a very complex data structure that will contain indeed the 3D points but also for each of those 3D points you can have the list of descriptor that has been used to triangulate these points. Okay.
So you have like basically you know how this 3D point looks like somehow from different viewpoints. Okay. You will have a visibility list meaning like what image or seeing this particular 3D point. Okay. You will also have like of course all your original data, all the images. But you now also have the exact pose where those images have been captured. You also know the uh characteristic, the geometric characteristic of your cameras. You know the calibration. You know the model of your camera. Uh and you know many other stuffs. Okay. You have a co visibility graph between your images. Meaning like what camera is looking at the same place as another. Okay. So you see that you have all this information and in general when we do visual localization we take full advantage of that thing. Okay, we don't just consider a 3D point cloud. We consider that we have a three point cloud that is enriched with descriptors with many things. Okay, so knowing that we can now move to the simplest and most naive localization technique ever. Okay.
Uh you will see that's a very stupid one and we'll discuss about that. So you know now that you know that the threepoint cloud can contain some descriptors and that you can extract descriptor from an image. Hey you're like hey I can match my 3D point cloud with my image. Okay and then if I have some 2D 3D correspondence look at this like think think of some line in 3D space and you try to align it to 3D points you will find that up you can align your camera only in one position and all the degrees of freedom will be locked. Okay. So look at this image.
That's exactly what is happening here.
So it means that the premise is simple actually if you can have 2D 3D correspondence using those descriptor matching you're done. Okay. So you know like actually like to estimate the exact pose of this camera in general we use what we we call like a perspective endpoint algorithm. I'm not going to go through the details that's extremely simple. That's actually just an approach. If you give 2D 3D correspondence you can get the pose.
Okay. Um it's simply solving actually polomial system of equation that actually is even simpler because you can rewrite as a third order polomial form.
It will give you a couple of solution.
Easy peasy. It's known from the 17th century probably even long before that.
Okay. So you know well now we can use that thing.
We know to match descriptor. We know how to get the pose. So you know we can try this. Imagine you have an image. You have a 3D map. Each 3D point has a descriptor. You match those guys. Boom.
you get that thing right so you see that sometime you have some outliers that's okay we can use like some robust technique like ransack again I'm not going to go through the details but well you can end up with basically a nice pose of your camera okay so do you think it's going to work actually we do that Sorry.
>> Oh, not not Siri. I say sorry. Siri is is buzzing me now. How do I stop Siri?
Okay. So, uh yeah, actually it's working. Even I will tell you it's working extremely well if you do that.
Okay, but it will take you a huge computational cost and we'll see later that that's not the only problem here.
There is something that has to do with scale. Like if you work in a small environment, this thing is going to be a killer. is going to work super well.
Okay, actually if you have infinite compute, this thing might be actually one of the best idea. Okay, but the thing it takes a lot of computation and why it has to do with the way we're comparing those descriptors. So I will try to give you a very quick overview on descriptor matching. So imagine you have like a query image and a 3D map. You can extract descriptor from your image and you have those descriptor that are embedded in your 3D map here. So you see that you want to associate those descriptor with the descriptor in the map. So what you can do is like you can take the first descriptor here and try to compute its distance with all the other descriptor in your tree map. So for instance maybe the distance between D1 and D1 prime is 10 and you will do it for all other points in the map. Okay.
So now the next thing you will do so here you can consider the one with the shorter distance is the best is the best match. Right? In practice it doesn't work so well. So what we do is we take the top two matching like 10 and 20 in that case the smallest one and we compute the ratio between them. If this ratio is lower than certain threshold we're like hey come on that's a bit ambiguous they look too similar so I will throw away this much. Okay so this is what we call the ratio test and then you repeat that operation for everything else but here you can quickly realize that you end up with a n square problem.
Okay this thing is quadratic and you don't want to be quadratic. So another cool solution will be that you take your descriptor in your 3D map and you create a KD tree of that a binary tree. Okay.
And then if you do that then you can do some approximate nearest neighbor and you can apply exactly the same strategy but now you can use a data structure that is efficient. Okay. And this thing will speed up your matching by 100 time at least. Okay. When you scale to something big at least. Okay. Because now you end up with a n login problem.
Good. You can also use other kind of tricks like for instance using a dictionary of word a visual word we'll come back to that later I'm not going to talk about this too much and if you want to go further because you see so far what we have seen is very handcrafted right this is like hey we have some rules we set some random threshold okay and you know it's not working that well actually that that's uh that's always a bit tricky but like more recently uh some deep learning based strategy came out and they're very smart you know like Take those two images here and you see I don't know if you see very well guys but there are like some orange points here and some blue points here. Okay. Then if I ask you as a human, how can you find a corres? Can you find the correspondences between them?
It's easy, right? It's super easy, you know, and then but the thing is like because for us, how why it is easy is the is the world question here, you know, because like it's easy because we see the context, we see the relationship between those points. But the algorithm I talked about before, no, no, that's not what they see. What they see is that, you know, now try to do the same.
try to match those things with those things. It's going to be a pain. You will probably be wrong sometimes. Okay?
Because we lost we were losing like any location information where you know like we lose all the context. All right? And this is what the author of Superglue figured out. They were like hey dude you know like they write this in their paper. I think it's very interesting.
When asked to match a given ambiguous key point, human look back and forth at both images. They sift through tentative matching key points, examine each and look for contextual cues that help disambiguate the true match from other self similarities. And look the last one. So this at an iterative process that can focus its attention on specific location. And here the main key point is like the attention again attention is coming back here. You know attention is all you need is also true for keepoint matching you know because at the end of the day when you human are matching those point you always consider the context you were always consider relationship between all those points and this is what they mimicked and that was very smart using a graph neural network and I'm sorry to talk again about graph neural network you might be so bored of this architecture after so many days um and then they use a actually attention graph neural network where they will compute So you see the input is the visual descriptor for both images and also the location of those key points in the image you know and the enance basically the key point with the position and cutings and they compute a self attention for each image individually and a cross attention then basically through message passing you can enrich those feature okay and then now you can pass it to what we call a syncor syncorn algorithm do you guys know what is the angarian algorithm we have some ungarians around don't way.
>> Yeah. Yeah. Do Do you know what is the Angarian algorithm?
>> I don't >> That's super cool. You need to read about that, man. Like this is it should be your pride and that's a that's a very smart strategy to basically solve in optimal way by graph matching. Okay. And this is what is used here. But the traditional ungarin algorithm is not differentiable. So it means you cannot train a network through that. So they used instead a differentiable version of that that is called synhorn that will give you basically an optimal assignment between your keyoint matching and this thing is extremely good. So if you're interested in keepoint matching this one actually is one of the most interesting paper and now most of the keyoint matching strategies they're built upon that foundation. So what you see we have like fantastic way to basically match key points match descriptors between the 3D space and the 2D image. All right, they're all great. But now the prem is like and some of them will be faster, some will be slower, whatever. But now, you know, when you move to basically a very large scene, imagine like you have a huge 3D map like this of New York City that will contain like millions and millions of points. Okay, so imagine you have millions and millions of descriptor. If you want to locate localize this query image, imagine you want to know where this particular key point is located in this 3D map. So you will take this part you will extract some feature out of it and you will be like hey let's compare it against millions and millions of key points first it will be slow number two is that oh maybe this key point is here look here we're looking at a window that if you go in New York City you have those kind of window everywhere so maybe this point is here maybe this point is here maybe this point is here or here or here and you will end up with potentially like tens of thousands of possible location for that point they are too ambiguous is too slow it's not going to work. Again, if you have infinite compute, hey, you're going to make it work, right? You can try all possible combination, you will do it, right? But in general, that's not what we want. We want something effective. And to do that, we perform localization in multiple stages. First, we take our query image and we do what we call a global matching. You remember what I talk about image retrieval. That's what we do here. And we're like, hey, this image has been acquired near this place.
So instead of comparing against like all the three point of the scene which will never work, what we do is we just extract a local 3D point here and we will match actually again a very tiny location and now you can basically go from course to find you can estimate the six degrees of freedom of your query image in the map. So how this global matching is performed this is typically image retrieval. So basically the idea is you have a database of image they can be geollocized or not doesn't matter and you're like which one is the most similar to me. Okay so in another word is like you have this database of image you will have some feature extractor can be a CNN it can be whatever you want and you will basically extract representative global feature for each of them. So I take an image I put that in my descriptor space I take the next one I put this guy in my descriptor space blah blah you will do it for all the image in your database. Okay. And you will store that you know like um in something you can query fast basically.
And then you will do the same for your query image up it will be somewhere in your latent space and to know which one is the closest you can look in the latent space which one is the closest to that and then you should end up with the most similar image. Easy peasy. So the world is how we have those feature extractor here right? how we learn that and the traditional technique like long before deep learning we basically uh got inspired again about those NLP people okay I'm a bit sour at them because like for a long time computer vision was considered like the pinnacle of AI it's not the case anymore now we are very inspired by by them with the transformer but back in the day we were still inspired a bit but what they were do doing you know and especially with this idea of bugawward so do you know what bugawward is by the Yeah. Can you can you tell me what it is?
Okay. I cannot hear you. Don't worry. Is like s imagine you have a dictionary.
Okay. Like a normal dictionary like it's just the list of all the words in your language in English whatever. Okay. And you take a text and you will basically every time you see a word you will do plus one in this word. Okay. So you create a histogram of word and this histogram somehow will be representative of what is inside of your text. So imagine you have a a football magazine.
What word will come very often? You will have a high frequency of I don't know Messi maybe he's not in the dictionary but you will have high frequency of ball win uh I don't know like you know all those things related run whatever okay if you read at a more like a text about theology about you know you will see God coming up maybe Jesus whatever okay so you will have different frequencies of words depending on the content that you're reading. So you see that now you can retrieve text. So that's a very old idea in NLP in natural language processing but now how do you transpose that to images you know now we want to use what we call visual word okay but what is a visual word right doesn't really exist we don't have a dictionary of pieces of image right and then so what is it is basically just a quantiz uh like this is a database um of a quantiz is visual feature that we extract. Okay, so don't worry, we'll come back to that in a minute. Like how we create basically a dictionary of visual word is actually rather simple.
Okay, so what we do is we take lots of images, tons of them, like you take 10 millions of images for each of those images, you will extract some local feature. Okay, and you extract some descriptor. It can be whatever you want, super point, it can be uh sift, it can be whatever you extract all those guys.
you can then you have now the representation in the vector space okay and then you cluster those guys with came in you say I want 10,000 clusters each of those cluster is considered as one of your word that's all okay so you have your vocabulary tree congratulation in practice that's a bit more complex we do it hierarchically but I will skip on that it doesn't matter so now that you have a dictionary what you can do is that if you have an image you want to characterize you can extract key points and you look for each of those key point their descriptor to what cluster they belong and then you create that histogram and then now you can basically make some localization easy okay so that's the idea I will skip this one because we don't have much time and there were an extension uh of bag of war that came after that which is called vlad vector of locally aggregated descriptor so in the vlad paper what what the author basically realized is that the fact that you discretize is not always very good because imagine like you can have like many words in W1 for instance associated to that cluster but you know if all the points are on this side of this cluster or if all the points are on the other side actually they can look widely different because of the quantization error. So the idea of Vlad is instead of just counting things is that for each word we will associate a residual vector. Okay, you will basically accumulate those vectors here with respect to the center of the cluster. Okay. And then then you can have a better I would say representation something more expressive somehow. Okay.
So of course the the only backside of it I won't go through too much details is that you end up with something super massive. You end up with basically one vector per word. Okay that's extremely sparse and is going to be very huge because K is the number of words you have in your dictionary and potentially you can have like hundreds of thousands of them. Okay. But anyway that's working pretty well and those techniques have been used for quite long time and in 2014 when you know like people talk a lot about that image imageet challenge and then alexet that came up etc. Then people were like, "Hey, now we have those fantastic uh CNN that came out.
Why not instead of using SE and all those old school stuff, let's use basically some pre-trained uh CNN and let's try to do the job." And that's exactly what people have done. And it was not working that great actually.
Okay, at the beginning they they came up with a fairly good idea about normalization but it was not catching up so well against like traditional techniques because those things they were just trained on imageet it was not fine-tuned for the task here right and people had no idea how to do that until netvlad came out and that's a very cool paper that really changed all the game okay it got like inspired by Vlad but tried to replicate every single component of Vlad using basically deep learning So the first one is you know like in the normal pipeline is that you extract some descriptor right here they say okay we are not going to use those old school thing boom a huge CNN is going to do the job to extract the descriptors next is basically normally what we do is the aggregation okay we try to come the number of words for instance or we try to create those residual vector from blood here they replace it with like actually something that is rather complex I would say from nowadays um standard to basically have a sort of assignment and more than this is that they will learn actually during the training uh like the cluster of visual word as well. So very smart actually uh development. So here there are few equation but we don't really care it's just a recap on how bag word is working. So ble word is very simple. You basically for each word in your dictionary for you will look which descriptor belong to that word. So you see that this operator here is simply you do plus one if this is closest to the cluster you're looking at and zero if not. Okay, you create basically just an histogram. Uh netlad is almost the same except that now like if the descriptor belongs to the cluster you're looking at then you will compute the resial vector that you will add to this word. That's all. Okay. So the main problem here is that when people try to mimic that with deep learning is that this operator here is really not differentiable. Whenever you have something that do does one or zero, you cannot differentiate that thing. What does it imply?
>> You can do backward propagation. That's not going to work. Right? So what's uh a very stupid uh AI people? What are we always doing?
We take the soft of that. We take the soft max always you know easy peasy so you will tell me hey not much novelty right that's exactly what they do now they have a soft assignment you see like you have here on the top the exponential of the distance to all the cluster to one cluster and here is basically the normalization you have in softmax that's all right and if you here you have a hyperparameter if you put this guy very high it will look like very much like a hard assignment basically so here you we and replace this hard ass assignment by soft one. Congratulation, you're done. Now you can back propagate.
But NetVlad, they're much smarter than this. They were like, "Yeah, but you know, like how do we know the center of those clusters?" You know, what we want is to train together a good feature extractor, but we also want to train, we want to embed somehow like those center cluster directly in a neural network during the training. And for this, oh, they simply used another MLP to do the job during the training.
Again I'm not going to go through the details but that's what it's doing. So together in during the training you learn good descriptor to extract and you learn basically some cluster of visual word somehow. Okay. And well already that's quite complicated but they were facing even more problem than this is that at that time there were no like clear data set to train that thing. So they did something they took Google Street time machine. So you know like you have like Google Street View and you can go through time as well. Okay, so they took that thing and so you know they they they train with triplets of images. You have your query image and you will have a negative image is very simple. You know, you take an image very far away, you know, like in the in the Google Street View. Imagine like I take a picture in Nepal and my negative sample can be uh a picture from Mexico.
I know it's going to be negative. It's not the same place. Now, what is more difficult is that the positive sample you're not sure because Google Street View doesn't tell you anything about the orientation. So, they had to do some mining. That's actually a quite complex process. And then when you end up with basically one enchor image, your core image, one positive sample and one negative sample, you can do something super fun is that you can basically use what we call a uh contrastive learning here metric loss. Here is actually a triplet loss. So what you want to do is in your latent space, you know, you you don't have a classification prime. You just want to learn the proper feature such that your positive sample will get closer in the latent space but your negative sample will be basically set far apart from it. Okay. And this is a very central method for so many things and that's exactly what they use back in the day. And so using all those things together it was very not trivial and yeah they did a very good job at it. And okay without talking too much about the result that was a huge breakthrough.
Okay, so here again is just for image retrieval. I have database of image which one is the closest to the image I captured. And now that you know that you can combine what we have seen to before you can combine the like local descriptor matching with image retrieval. Okay, so you can basically extract some feature from your image.
You will do some image retrieval to restrict the area you want to search and then you can do feature matching and they work so well together. Okay, I will even tell you basically those kind of hybrid strategy. They're the most scalable and they are for most of the time the most accurate as well in most scenarios. Okay, so here is what we call hierarchical or hybrid but is better visual localization and if you don't really know what to use this thing is always going to work. Is it always the best? Hey, not always. And I would say, you know, yeah, I just presented NetVlad, but there are so many so many works that are coming like every year and they're fantastic and you know, now just if you use Dino V2 or Dino V3 directly for localization without any finetuning is already going to be pretty good honestly. Okay. Um but you know like there's still many papers coming up now we rely on huge transformers. We it's a bit different right? Um but anyway you can have something very accurate. So to come back to hierarchical localization I told you it's very good but this is not necessarily the fastest this is not necessarily the best for privacy preservation because look you have all those data you know your struct you will need your structure from motion map if you want to do that so it means that you have all your images stored somewhere and then for privacy that's not super good. So anyway, you know, like people were trying to find solution how we can improve that pipeline. Super difficult to be better than this. And in 2015 or so, someone came up with a bit of a crazy idea. Very simple idea. It's like, you know what? What if we simply use a basic CNN? This CNN outputs only six parameter. This is the pose of the camera. You give me an image as input, I give you the output of the pose, you know, for a given scene. That's all and the architecture there is nothing. No the wallpoint. So this thing is posnet actually and we can have a look on how this thing is working. That's a very old video we can move around but basically the only input is this image and it tells you where this camera is located directly and it takes a couple of milliseconds because the network is very tiny. Okay. So they propose some data set and you know back in the day it was a real revolution. Okay. Even though there is no architecture, there is nothing. It was very um very new. So you might wonder how to train that thing.
Okay. So the training require some image and post pairs that in general you're going to obtain by structure from motion. But you don't need the 3D structure of the scene. You can throw it away. Okay. So for your training, what you want to do is like you want your network to remember the relation between image and pose. Okay. So you will take basically a portion of this uh image you use for the scene reconstruction as the training set and some over to test how you're going okay and you're simply going to train your network in that way okay because like you know the ground trous pose for a given scene okay imagine we reconstruct this room in 3D and then I want to train posnet on that I will just take like all the image per pose I will train my network and if I want to localize myself now I take a new image H and I can reuse my network that I've trained before. Okay, so we have two regression head in general.
One for translation, one for rotation and you can basically regress that. You need just to be a bit careful on how you regress your rotation because there are many ways to represent rotations. Some are better than others. Okay, I won't go through the details. Quition is a pretty good choice I would say or rotation matrix is doing well. Never use uler angle. That's all. You know, if you do that, you should be fine. So back in the day they were using the Google net architecture that now is quite outdated.
You know back in back in the day uh ResNet was not even existing and then Google net is like some an architecture with some early exit to for a better flow of the gradient. It's almost like assemble in some ways anyway it doesn't matter. So they simply remove the soft max and they replace it with basically some regression. That's all. Okay. So something that was interesting for the time is that because is that they were showing the importance of pre-training and that was huge back in the day I would say now it's obvious to all of you because we went through all these self-supervised learning dyno etc. We know we know that we need basically a good pre-train network back in the day it was not that obvious. So it was also like a very interesting insight that they were proposing. So as I told you the output is simply you know like the translation and rotation as a quaternian. So this is you know just a tiny vector here and they are the better parameter to balance them because like this quaternian potentially you know the maximum norm of a quaternion is one but what is the maximum translation it depends on the size of your scene. Okay like for instance in that room is going to be like 30 m by 30 m or something.
Okay so you need to rebalance those losses somehow. Okay, so they use a better parameter that back in the day was tuned for every scene. So that was a bit of a pain and they ended up with actually pretty uh amazing result for the time. But if you were digging a bit into this, you were quickly realizing that this strategy is extremely inaccurate. Okay, it's like not working as well as advertised. And then people try to scale it up to bigger scene and boom, it's not working. You know, it can be because of cat catastrophic forgetting, can be anything. Okay, but it's not accurate whatever you do. It has some advantages as well. For instance, this is actually quite robust.
So even in a scene where you don't have much texture is doing the job. Okay, and so many follow-up work came out after that. But in general, those direct regression technique, they are not that accurate. Okay, so here are a few results I can pass. So well some advantages very fast inference robust in low texture you don't need to have this explicit 3D structure you know you don't even need the training images you can throw them away so you see that your privacy problem is resolved okay so uh the main problem is that this is not very good it doesn't scale it's is not accurate so it doesn't do the job so people were like hey but nonetheless that's interesting and then came the sin coordinate triggeration technique So it's actually quite similar to what we have seen right before the only difference is that now I so the idea of scene regression people were like yeah we want more geometric insight okay we want to those insight to be closer to what we did before in our previous you know like traditional pipeline okay keepoint matching pose estimation stuff like this and so they started from the same idea I give you an image I have a network but instead of predicting the directly the pose only six parameters or seven parameters. I will basically try to regress three like a tensor with three channels inside and for each pixel basically what we try to get is the 3D location of this pixel of the projection of this pixel in the 3D word in the scene coordinate. Okay.
And that's that's the reason why we call them scene coordinate regression. So you see this 3D point cloud has the specificity that it exists in the frame referential of the scene in which you want to localize. So you can understand that as a point map regression. It's another term that you might have heard of if you work in 3D. It's a onetoone mapping between the pixel in your image and their corresponding 3D points in the absolute reference of the scene. Okay.
And again here what does it mean? It means that we have some 2D coordinate and they're mapping in the 3D scene. And we have seen before that if you have this 2D 3D matching, we're done. You know, job is done. Look at this. You use your network. You take your image using your network. You get the 3D position of those point in your map. Look at this.
You use runs sack and PNP like a perspective endpoint that we have seen before. Boom. You localize yourself.
Done. Right? And then people realized that this thing is actually much better.
But how do you train that thing you know and the first people trying to do that it was in 2013. So you see at this time neural net was emerging somehow like people were using over stuffs and actually what they did is they took a data set where they had some RGBD image and they simply trained for each image how they can regress the 3D coordinates in the scene basically using a random forest. Okay. So so was it working? Yes. But this thing was but you so to train you need like RGB depth all the pose. So you need so many things in many situation you don't have the depth okay and especially not the dense depth right and it was extremely slow I mean this thing I had to use it at some point because one reviewer asked hey why do don't you compare against super old technique thank you that's a pain okay that taking forever to basically converge to something this is of course sin specific like poset that we have seen before and it doesn't scale at all so later on some people were working in nontik you know nontic is the company basically doing like Pokémon Go and those things. Okay.
Uh they were like can we do something better? You know we don't have this depth to train and they got another idea. They were like let's keep that network that can predict. So instead of random forest let's use neural net first. Okay. So a CNN and we are going to regress basically um like some 3D coordinates. Okay. But many of those point at the beginning of the training they will be useless. They will be bad.
Okay. So what they basically proposed to do is to sample sets of points randomly in that 3D space. Okay. And they try to basically estimate the pose and to compute the reprojection error for many uh candidates. Okay. And they use basically a neural another neural network to estimate the quality of basically each hypothesis that you get.
and then again a soft max and you sample through the soft max to have like the best hypothesis and then you do back propagation and you train and if you do that you will b you don't need any depth okay you just hope that at some point the the training is going to converge to something okay and this is surprisingly working well but the main problem of that so this technique is called DSAC it takes like more than a day for tiny sin so if you want basically to use this to spray in a localizer in that room. Yeah, like good luck with that. It's going to take you one entire day. So, super slow.
Much more recently in 2023, the same author proposed something called ACE.
And this thing is super cool. It's cool.
But so, it stands for accelerated coordinate encoding. And what they did is like they have trained now a backbone uh on like so many images. Okay. So, this thing is frozen. Okay. And then when they do their training they do multiple thing here. So they take like for each patch you will have one corresponding feature somehow. Okay. And what they do they shuffle those feature they will mix up all the images together. So all the patches okay and for all the patches they will try to predict their 3D position using simply a tiny MLP. Okay. And then they only basically guide their training through the reprojection error on the images.
Okay. And why do they shuffle here is because you want basically to train for each batch on as many possible variety of image as you can. Okay, because if you train image per image then the training will be extremely slow. Okay, because basically you will keep adjusting one with another. Okay, and this shuffling idea is so good actually.
Okay, and you see that there are many advantages in what they have proposed.
First you have this thing that basically doesn't need to be a retrain is fixed and the only thing you train is that given some features I have a tiny MLP that will regress the position for each point in my image for each pixel and it means that when you want to train for one scene it takes now five minutes okay and that's a very cool network to play with because it's super fast okay so very nice one and then yeah as I told you like this is much much much faster here is ace like it takes a couple of And here you have dac star that we talked about before 200 minutes u and then you have like yeah over dag it takes forever. Okay so it was a huge huge progress in those kind of sync coordinate regressoration techniques and these days I can see like plenty of new ideas that are emerging. So now like how you can train the similar kind of thing but without using any prior SFM reconstruction uh how we can localize using those technique in large scene many many things that are very cool that are emerging. I really like those kind of paper even though at this point they cannot really compete yet against the hierarchical localization technique we have seen before you know pose image retrieval and pose estimation. So now we will just conclude on that because like you might be confused we have seen so many things and then yeah what to use when so here is basically a little summary of all of this. So if you use absolute pose regression okay like basically poset again fast inference you don't need a tree map to train that's super simple to do you just take any CNN backbone boom you're done okay but this not going to scale it's going to be sin specific and not super accurate now coordinate regression they're quite close to absolute pose regression um but they're bit more accurate the rest is mostly the same okay uh and finally hierarchical localization like amazing accuracy very good scaling but then you need to go through this entire SFM reconstruction you have bad privacy that I forgot to write here it's memory intensive etc etc okay and there are few more techniques but here is I would say the main families that you will face okay so there are plenty of future direction that you can follow actually in um in visual localization I try to just put a few of them like where the field is going uh One of them is for instance like some multimodel kind of localization. Here is my work actually published last year is how we can localize a vehicle using plenty of cameras installed on the car and how we can localize it in a HD map of a city.
Okay. Uh and that's fun because now you you know you don't work with 3D map you don't work with images this is a graph and again here we use some form of graph neuron network we use we use some simore algorithm those kind of thing. uh now I can see that the field is moving towards huge uh pre-trained model like anywhere else right obviously okay as well this is like everything now in image field will be something like that uh there are many other fun things to do like cross model localization now what if what I want to localize instead of normal images thermal image how you going to do that okay um it can be cross view localization it can be you know so many things like this like plenty of challenges that are not fully resolved like here is a bit extreme is like how you localize this image into this satellite view. Okay. And they look very very different. Okay. So some people also I was talking about Jeser recently like maybe six months ago like I I wonder if it's Luke Vanul I I don't remember who published that thing but they made like basically an AI much better than any Joe guesser in the world right it's super accurate. Uh so yeah you have plenty of cool challenge you can tackle basically uh something I see a lot uh I'm editor in a journal so I see those paper coming out every day these days is basically how you can embed some you know like uh language model into that as well okay like instead of just giving an image to localize yourself what about if you localize with text okay say I see uh some people in a room with red chairs and where am I okay those kind of thing or like imagine you have like your robot and you can tell the robot go to the location where there is a fire extinguisher. Okay, how do you do that?
How can you localize this basically uh element in the scene? So you need some form of grounded uh reconstruction of the scene somehow. Okay, there also a lot of very cool work. Here is another work that we have done actually with NAMI maybe two years ago is like the premise how we localize ourself in implicit representation like nerve gian splat those kind of thing and many people have been working on that as well recently that's pretty fun uh you also have like some more technical related challenge how I can make those thing work on embedded systems like there are all those new very cool data set and new hardware basically for um uh ego localization somehow Okay, like with with those I don't know like Google glasses or what is it like the new glasses from meta those kind of thing.
Uh still the challenge of how you scale to a city it's still something difficult. There is a challenge that I'm working bit with these days is like how you preserve the privacy when you basically localize yourself. Um so yeah there are very fun techniques that have been deployed for that.
And then here just the last slide to conclude we'll actually be right on time. That's great. Visualization is a core problem with application in many different fields. And what will come next is basically yeah how we scale it, how we make it work with plenty of different kind of data uh and you know many other stuff that will come up and privacy is one of them that I like. I told you. So I hope it gave you a pretty good overview of this field and uh yeah we still have time for a few questions.
So, thank you so much Yeah, thank you for such a wonderful and interesting talk. I must say I am right for most like 50 years ago when you were doing a PhD and I was a master student uh and today's talk was a very good pressure for me and I can see that you guys advance a lot. My one question would be uh what would you say has been the major breakthrough is the speed accuracy or lucid.
So like recently the the very huge breakthrough we we have seen like actually in the past two years that's very recent you know we we just add our sort of attention is all unit moment recently is that um we we see new network emerging and that we call like point map regression basically and this is like some you can forget completely about geometry okay like geometry is somehow gone now they start to be very good with that you take an image or two images and it will directly give you basically the structure structure the pose everything the calibration you need nothing at this point image result you have the density reconstruction and those methods now they expand to series of images and they scale they start to scale very well so you know this is now you can do very cool things in robotics you don't need to know anything you know you put your image you get what you want and that's actually I think the the biggest breakthrough now and now people combine those thing with gian splat as well so directly regress gian splatting so not only you have the 3D structure not only you have Suppose not only you have the calibration now you have like a realistic phototric reconstruction as well. So you see we start really to get to something very tangible very useful and yeah for me that's the biggest breakthrough and that's a bit sad because all the theory we knew before is now replaced by a huge transformer and this is working right that's the the magic of deep learning and this is what we have to work with anyway. So that's the path I'm following a little bit these days.
Okay.
So I have like a few quick question if you're like combining like slam and nerve that propagating that rendering loss back to the slam in order to like update the port or not. Oh well, okay.
Uh, so you you can refine the pose using Nerf eventually. That can be a way. You can even use like some self-calibrated nerf if this is what you mean. I mean, combining Nerf and SL to me this is a bit antagonistic because Nerf is very slow and slam is very fast. That's you know that's something I'm always buzzing a bit with because uh maybe in that case and I think I told you last time I think ocean splat might be a bit better and that can be a good combination because you know you can start with some key point base you know localization and then if you have some phototric sort of reconstruction of the scene you know implicitly it can be nerfed it can be gushions but doesn't matter then you can have probably a much finer localization and much more robust to textureless environment so that can be a good combination but people are moving on that already quite fast. I would say there are like a lot of work in that in that direction. Old one will be niceam and there are many other things popping up actually. Yeah. So yeah, this is something I think people have explored already.
>> Can I have like another question like if you're doing like bundle injustment and like you loop and then you are updating the post then how are you like updating like first reconstructed map or like not reconstructed map? Uh okay when you do loop closure and then you assume that for each you have like some question you have some splats in your map right ah okay here is difficult because you know the splat you have they won't necessarily correspond to the 3D point you have triangulated right because you know you have this division that you make during the gussion splat and then those thing they're not grounded to the 3D map somehow so it's not it's not a use case I've seen but that's a fun thing I like how you move them accordingly that's not easy to be frank with you I don't know even if it's feasible or it will take lot of resources because one way it will be to refine the interrogation splat you know and you can have some rough approximation on how you will move them with the nearest 3D points I guess but it's going to be a total mess it's going to be very computationally expensive to do loop closure in a gian splat environment uh like this I think but that's a challenge that's a very interesting question >> well And so I think the way is like to like maybe like retrain the system or >> Yes. Yes. You will have to you will have to do that. There is no way you can skip that thing. Totally.
>> Thank you.
My question is uh regarding how do we localize dynamic system which are not stationary in uh elevation.
>> Uh what what do you mean by dynamic system? Exactly.
>> Just like they are moving right.
>> Yeah. Yeah. You will use the same kind of idea like you know you take okay if you want to be real time you don't need to localize all the time. Okay. You can say at time t I'm located here right. So maybe you will be a bit late but in the meantime you can do some visual autometry in the meantime. You will know how your drone has moved right. So you can basically take the toes you have located and you can add the visual autoometry to that and then you see you can have a low frequency localization.
Okay. And then you can update that using the visualmetry and you can be real time. The way we are going to do it is that you can use for instance for drone a satellite view is going to do the job and you know you you can have the six degrees of freedom. So you know exactly the elevation you know everything right?
So it does it doesn't matter if it's a drone or a car it's going to work for anything.
>> All right then if we don't have any further question thank you again so much and uh it will be time for the tea break. Thank you.
Up Next

NeRF: Neural Radiance Fields for Photorealistic View Synthesis | ECCV 2020 Talk
@benmildenhall3169
23.5K views•2020-08-04

Building Real-Time ML Pipelines with Feature Stores and MLOps Frameworks
@ODSCAI
5.1K views•2022-02-20

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence






![Кватернионы | Вращение в 3D [Самая суть]](https://i.ytimg.com/vi_webp/tBYXR_fAXSQ/maxresdefault.webp)











![Intro to Deep Learning -- L12 Intro to Convolutional Neural Networks (Part 1) [Stat453, SS20]](https://i.ytimg.com/vi_webp/7ftuaShIzhc/maxresdefault.webp)
![[AI for ME] L10-2, CNN](https://i.ytimg.com/vi/u03QN8lJsDg/maxresdefault.jpg)





![[논문리뷰]카메라 위치 파악을 위한 복셀 표현의 Covisibility 기반 참조 이미지 클러스터링](https://i.ytimg.com/vi_webp/V3O4fkmb6lo/sddefault.webp)



![[논문 리뷰] From Coarse to Fine : Robust Hierarchical Localization at Large Scale](https://i.ytimg.com/vi/rTmXeHCvt3U/maxresdefault.jpg)




![3D Gaussian Splatting [Paper Review]](https://i.ytimg.com/vi/xTp88ZOtm58/maxresdefault.jpg)




