Instant Neural Graphics Primitives (Instant NGP) is a technique that reduces the computational cost of neural graphics primitives by replacing deep neural networks with a multiresolution hash table of trainable feature vectors, achieving a speedup of several orders of magnitude (from 30 hours to 30 seconds) while maintaining high-quality 3D reconstruction from multi-view images.
Instant Neural Graphics Primitives with Multiresolution Hash Encoding
Added:foreign [Music] my name is Emil and today I'll be presenting instant neural Graphics Primitives with a multi-resolution hash encoding this paper was published by Nvidia employees actually Nvidia researchers it's actually quite interesting paper and the reason why I chose this paper because it has promising applications and so the potential for for its applications is actually quite more interesting kind of you know it's inner workings actually um so let's go over it and let's see what it is so the contents are basically this this is what I I will be presenting I will first introduce you what is going on here and second I'll talk about the applications uh you know potential applications and then I'll talk about previous work and finally I will try to explain like what is the core Innovation behind this work so basically why why did it make a difference basically let's proceed so imagine that you take you take let's say 16 to 30 images from different angles uh for this particular building and you're given a task for reconstructing a 3D image of it right let's say mesh right could be anything and it doesn't matter and the the requirement is let's say that you have write about enough let's say you know 100 images right doesn't really matter but all of them are from different angles and you need to like reconstruct a smooth like perfect 3D representation of it basically we need to be able to scroll it from side to side and we need to be able to kind of see a smooth transition as we slide and we have to be able to like three the size of the C the sides of it and behind it and so forth so basically as you can see like from these two images like here it is taken from the input you can see that it's taken from uh from this side and the output is also from this side right so uh this is the kind of close representation to the input and in the second one we just scroll to the other side and this was the input photo from our collection and this is how it's rendered as you can see in here I saw around here and it kind of messed up but um you know overall you kind of did a good job in kind of reconstructing the whole thing and the reason it messed up I think because like uh because we have taken multiple photos from different angles sometimes lightning got in the way and from on the one picture it got it was kind of you know more brighter than the other pictures so you couldn't learn this particular place so maybe that's why that's why it looks like this so let me introduce this uh neural radius field right so this was also a paper published in actually 2020 uh the paper that I'm presenting today uh multi-resolution hash encoding it's actually has been published in 2022. and um basically what Nerf does is it takes input images I'm sorry for the you know poor image quality here but what we have to what we have right here is like 100 images of this basically drums but taken from different angles and then uh optimizes it you know using neural networks and then it creates a 3D representation of it so basically this is what we've got and it made a big news for a while because it was very uh very effective basically the quality of this High you know dimensional reconstruction was very high so let's see how Nerf accomplished this first basically what what we have here is a deep neural network and it takes 5D input and it outputs a volume density and the color density so what are those so basically input is just x y z the the coordinates it's uh Theta and Phi and what Theta and Phi represent those are basically uh just angles view directions basically right so from which angle are you looking at it so basically uh you know it's basically New Direction and the output is uh volume density and volume density is is basically you know it answers if at some let's say location if something exists there essentially right so let's say that you have 3D image and at any particular location you want to know what we have there right essentially and it's it's only a function of location because what we have somewhere does not depend on the direction right either we have it or we don't so it's a functional location and the color density however depends on both location and the view Direction because if you look at at an object from different angles but because of the lightning because of the distance right the color may change so basically we both need a direction and we need location so the way that uh that the neural network works is that essentially it takes 100 images and it tries to overfit them essentially so this is the intentional overfitting basically and one way that it does it is suppose that you have uh your you have the view Direction Let's see like right here but as you can see so basically you emit a light range in One Direction and you basically look at each point strike here right so uh this is this is just a graph of what this is so it goes on and here it encounters something and then here it says okay there is nothing this is just an empty space right so we first find nothing and then we find something and then once again nothing right so uh this is basically what volume density means and the color is pretty much you know the color and the way that this this is optimized is let's say you start from random like let's say you start from scratch and then you basically calculate uh L2 loss from um you know from your actual images right uh to to your let's say recreated uh recreated 3D output right because you can because you have multiple images you can actually calculate this rendering loss um by taking uh your actual images as reference so let's see so let's see uh let's first actually discuss what the problems with this is so the thing is this is actually a high quality uh this actually has high quality performance just the problem is that it takes too long to train because it's deep neural network and it and the rendering time is actually I mean in in days not not just in hours but actually in days it takes one two two days to uh to to actually get the output and um and you even get this performance on a very powerful GPU so it's not just a typical GPU even on powerful GPU it takes like days to a couple of days to actually uh get the output so what this instant NGP NGP basically this multi-resolution hash encoding is that it first uh it decreased the processing time it greatly reduce the processing time and it did not have to sacrifice any uh quality from basically so let's see what the inner workings of this instant NGP basically is right so this is kind of the whole architecture essentially and as you can see we do not have a deep neural network here we only have two hidden layer neural networks so the reason why it's so fast and we'll get to to the speed part as well so the reason why it's so fast is because uh this is not the neural network right it only has two hidden Lanes but then how does it optimize the parameters right so how does he how does it check the parameters how how does it uh reach the uh let's say comparable performance so the way that that they do it essentially is the only difference between their job uh right and the previous Nerf is essentially about how you represent the spatial data so basically it's all about parametric encoding right so what the way that they do it is using cache tables so they use multiple hash tables for different resolutions and they they essentially update when when the when they apply back propagation they not only update the neural networks parameters but they also update the future parameters that are in the hash table as well so basically instead of representing the whole thing as voxels they uh they uh using linear interpolation they just uh you know embed them in inside of a hash table and then from hash table they provide the whole thing as an input to the neural network itself so because the whole thing is differentiable back propagation can be applied to the hash table to the entries in the higher Stables as well so basically as a neural network updates the entries in the hash table update as well and with this they kind of reduce the um processing time so let's go forward so this is actually the result so basically uh Nerf Nerf's processing time was in 30 hours right so basically uh for some it was 32 34 like 36 so 30 hours compared to 30 seconds so basically what what multi uh hashing multi-resolution hashing achieved is that it reduced from 30 hours to 30 seconds so this is a very huge deal and uh let's uh actually I forgot to talk about the applications as well so because it yeah it has very high performance in terms of quality and speed um you can write many different software applications for it right so let's say uh for we have today we have different hardware for 3D scanners but with this you don't even need 3D scanners you can just take multiple pictures and let's say you take 16 pictures and from 16 pictures generate kind of high quality 3D mock-up right 3D model that's one thing and for example for entertainment industry what they can do is usually working to create uh you know 3D models what they instead could do is to just draw 2D pictures and then create 3D models based on that so it can have a huge uh Improvement for gaming industry as well and you can also create multiple uh you know different type of augmented reality or virtual reality applications for it and the best part is it's actually open source you can see the code you can edit it you can run it on your own software uh I wanted to show the demo actually because I knew that the whole thing is open source it's just that unfortunately my my computer's GPU is actually not so good and uh so yeah I I checked the videos of people who run it and they they complained only about the GPU because their GPU was not good enough they will not be able to recreate kind of high quality image so uh yeah that's that's it uh I hope you enjoyed it
Up Next

3D Gaussian Splatting Explained: Real-Time Neural Rendering
@bilawalsidhu
153.9K views•2023-11-05

BitTorrent Protocol Explained: Piece Selection & Peer Choking
@StevenGordonAU
481 views•2013-02-22

HTTP Requests Explained: GET, POST, PUT, DELETE
@codecademy
103.1K views•2021-10-07

Enigma Machine Mechanics: WWII Encryption Explained
@JaredOwen
13.2M views•2021-12-11
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Computer Science





















![3D Gaussian Splatting [Paper Review]](https://i.ytimg.com/vi/xTp88ZOtm58/maxresdefault.jpg)

![3D Gaussian Splatting for Real-Time Radiance Field Rendering [Paper Review]](https://i.ytimg.com/vi/7WuKq1WfIL0/maxresdefault.jpg)

![Neuralangelo: High-Fidelity Neural Surface Reconstruction [Kim Yu-Ji]](https://i.ytimg.com/vi/5sEYfiv_AOw/maxresdefault.jpg)













