PyTorch's Dataset and DataLoader are essential abstractions for efficient data handling in deep learning training pipelines. The Dataset class defines how data is stored and accessed through two key methods: length() (total samples) and get_item() (retrieve individual samples). The DataLoader wraps the Dataset to automate batching (grouping samples into mini-batches), shuffling (randomizing order each epoch for better generalization), and parallel data loading using multiple CPU workers. This system ensures data is fed to the GPU efficiently while maintaining consistent and scalable training workflows.
PyTorch Dataset & DataLoader: Efficient Data Pipelines Guide
Added:So far we have learned about tensors, gradients, autograd and randomness in the last video. Now it's time to handle real data because without efficiently feeding your data into models, everything else falls apart. And that's exactly where PyTorch's data set and data loader come into play. Let's break this down first and then we'll jump into working code example.
Imagine you're training a neural network on thousands of cat and dog images. You could manually load them all using list directory, process them using pillow and feed them batch by batch. But that would be painfully slow, unscalable and inconsistent. PyTorch solves this problem with two elegant abstractions.
Data set and data loader. Data set is how the data is stored and fetched and data loader is how the data is patched, shuffled and fed to the GPU memory efficiently. Let's see how data set and data loader works internal. Data set defines two function length and get item. Length tells us how many samples are available. Get item provides utility to fetch one sample. On the other hand, data loader wraps the data set and helps you shuffle data every iteration or every epoch. It also helps you load batches of samples and it uses multiple CPU workers for parallel loading of data. Data loader also automatically pushes data to GPU efficiently. So in short, data set defines what your data is and data loader defines how your data flows through the training pipeline.
Let's have a look at this with example.
So first thing I'm going to do is create the notebook and I'm going to name this data set loader dot tutorial.
And as always, let's import torch.
Let's check the version of the torch.
Torch dot version. And then let's import the data set and data loader utility. So I'm going to say from torch dot utils dot data import data set and data loader.
Now the idea here is to create a custom data set that will return the square of the number. So I'm going to say class or let's actually write our step. So step one is create a custom data set and this is actually a comment and let's create our class. So class I'm going to name this square data set and let's inherit this with data set class.
Let's create a constructor for this. So in it and self and length equals to 10.
And if you remember from our tensor video, we can use a range function which is pretty similar to Python arrange or range function.
So a range and let's pass in our length.
Then we can override the length function self and let's return the length return length of our data.
So self dot data this should be dot data.
Now similarly we can override get item function as well.
This will take in self and let's take the index as well.
and then x = to self dot data using index and y will be the square of x.
So get item is the function where we are actually defining how to retrieve single sample from the batch and that's it. So let's see what's going on here. So square data set is our custom data set class and in it or the constructor defines or prepares your data and length function tells PyTorch how many sample exist in the data set and get item tells PyTorch or defines how to get one sample from the data set. Now step two will be to initialize this data set and create a data loader or initialize data loader for our data set. So let's write step two initialize data set and data loader data set equals to square data set and let's pass the length as 10 only.
And then let's define data loader.
Data loader equals to we have to pass our data set to data loader class and batch size argument and I'm going to pass this value as three. And let's pass this argument shuffle as true.
And that's it. So what this data loader is doing it automates batching which breaks data into mini groups. In this case that mini group is of size three.
Then it also does shuffling for us. So that's randomizing order each epoch which is good for training uh because our model can generalize better and converge better. And then it takes one more argument of num workers which if you set to more than zero it loads data in parallel and that's helpful for uh parallel loading of data. I'm not using num workers here but obviously when we actually train our model I'm going to use num workers as well to load data more efficiently. Now let's actually iterate through our data loader and see how does data look like. So I'm going to say step Three will be iterate through batches.
And let's see for batch index comma x comma y and let's use enumerate method. So we get index as well.
data loader and let's print all the things. So I'm going to say print batch index + one to get the accurate number.
And let's print x first.
and single quotation here as well. Now let's print y as well.
So y why y. So that's it. Let's look at our data and see how it looks like. Of course, this should be within quotation.
And let's rerun this. So this is what our data looks like. So data loader has automatically created batches. Data loader has grouped our data into batches of three. So three three and then last one as one. So this is all done by data loader. And when you feed these batches into your model, data loader will make sure to change or shuffle them as well.
And I can show you that if I rerun this, you see batches or the tensors or the values within our tensor has actually changed for every epoch. If I rerun this, you'll see values have changed again. So that's the beauty of data loader and data set. Now if you don't understand this or what exactly data loader and data set is I would suggest to rewatch the video and even then if you don't understand in the next video we'll extend this one step further and use real data set or mnest data set and load it using data loader and data set and do some transformation. So don't worry if you didn't understood or if it's too much as of now because trust me you'll understand this once we actually start working with the real data set. So with that thank you so much for watching and see you in the next part where we'll understand how to load real data set using data set and data loader and how to perform some transforms and augmentation on our data. So, thank you so much and see you in the next
Up Next

Federated Learning with Flower and PyTorch: A Step-by-Step Tutorial
@flowerlabs
14.6K views•2023-08-08

Triumph of Orthodoxy Icon: Byzantine Art & History Explained
@BenCallan
2.1K views•2024-08-06

FastAPI vs Flask vs Django: Choosing the Right Python Web Framework
@TechWithTim
302.5K views•2024-05-26

Game of Thrones Opening Credits: A Cinematic Analysis
@gameofthrones
46.3M views•2011-04-18
Related Study Plans & Knowledge Roadmaps
Structured learning paths in General & Interdisciplinary Studies













![[Mì Python] Bài 4. Python với Keras (Phần 1)](https://i.ytimg.com/vi/hPhnqTtidnA/maxresdefault.jpg)

















![[Technion ECE046211 Deep Learning W24] Maximizing CPU and GPU Utilization in PyTorch](https://i.ytimg.com/vi/tIoa8axf9MI/maxresdefault.jpg)







