This comprehensive lecture covers machine learning methods for time series forecasting, explaining why traditional statistical methods like ETS and ARIMA remain competitive despite the popularity of deep learning approaches. The talk explores the historical controversy between simple and complex forecasting methods through competitions like M3 and M4, introduces global modeling paradigms that train one model across multiple time series, and addresses critical challenges including non-stationarity, seasonality handling, and proper evaluation metrics. The instructors emphasize that while deep learning methods like Transformers show promise, they often require large datasets and careful benchmarking against simpler methods, making traditional approaches still highly relevant for many forecasting applications.
Machine Learning for Forecasting: Methods, Models, and Pitfalls
Added:um yeah exactly so I'm I'm I'm Kristoph bermier we are going to talk today about machine learning for forecasting um one thing I should say is we we I I got about two hours of material I think so um yeah bear with me the other thing is we had planned maybe some lab or something but then we thought well we have so much material uh let's just uh stick to this so so we won't have any like lab type of things so that means my PhD student abishek he will just present some of the theory as well initially we had planned for lab um good with that let's get into it uh with some introduction so H the The Talk today is is mainly about what some people would call plain old machine learning poml so um we will cover deep learning a little bit uh but not in in much detail uh I mean I think there are better resources out there probably and people more competent to talk about it we will also not talk a lot about like generative foundational models GPT like like things I I I will talk about that in a moment quite quickly but then we will talk about other things so if that's what you came here for then sorry to uh to to you know make you unhappy and um but I but I hope you find it still um enjoyable so I think the main intention of this training course is that you understand like the basic problems that General Solutions and forecasting for both modeling and evaluation so that you are unable to judge uh results of complicated ml methods right so my idea of today's talk is something like if if this shiny new deep learning method comes along in a new paper that you are enable to say hold on what do they actually do um does this make sense uh for me to to try to use that um yeah so so why should we exactly so why should we still look at plain well and when I say ploid machine learning I'm talking about any machine learning that is not deep learning things like gradient boosted trees and uh maybe even feed forward you know Network things like that um yeah so why should we still look into this is now if nowadays it's all about deep learning um and sure I mean Transformers craft near networks they may be are the F and a lot of papers coming out in this space however one thing that I have found is that many of these recent Transformer papers have flawed experimental setups uh and greatly misrepresenting their capabilities on the use data sets which means their evaluation is often quite weak and um what they promise they often cannot really hold um for example they use um only series with thousands of data points per series and why you know companies like Amazon Google also Sal they they may be getting good results with these Transformers but they have like hundreds of thousands of series um another thing to keep in mind is that these models they have been largely Irrelevant in them five competition even though many of these architectures were already available at that time I mean obviously there have been coming out a lot more since then but some things yeah were definitely already there so the question is why did people not use them so if you think about your typical forecasting problem that could be something like this I I have about 1,000 series their daily series should I try the Deep learning model or not um I to me I think that the the answer is not so clear right now so maybe it works maybe it doesn't and you can also have a look at our uh data repository at forecasting data.org where we also rent some of these methods and and the results are quite mixed right so sometimes these methods work quite well sometimes other things work better yeah and the other uh thing to address is obviously like why why should we even bother about this if in two years every everything will be done with some GPT variant anyway so what w't all of this be soon obsolete um I guess look the answer is these foundational time series models they are definitely coming I mean we've seen time GPT coming out um but you see now I have all these questions like will they work on any possible frequency because one thing that you have to understand about forecasting is that forecasting can mean a lot of things to a lot of people so you could have like one yearly time series with 15 data points in total you could have one time series with a a million data points and 10 co-variants you could have a 100,000 time series that are all you know 50 data points long and they could have all of these different frequencies you could have yearly series monthly daily half hourly minutely secondly so will they work with all of these possible frequencies I think well maybe they could uh I think well at least on maybe the most 20 most common ones another really big question for me is will they work with um with covariates because in many forecasting problems you you have you know things like the weather is available or um promotions that you know of different things how will they deal with this I I think it's not so clear right now so maybe with some form of embedding or dimensionality reduction I've just two days ago I learned about pfn which is like a transformer for tabular data that works up to 100 features uh so I think that's a very interesting thing maybe that can in some way be adapted to GPT you could try some stacking where basically you just take whatever this GPT method gives you and you feed that into some into your own algorithm as an additional input so I think there are definitely you know ways to address this but right now we we don't know any of these things and and then obviously the other big question is will this work with your constraints that you have in production with your Hardware your time frames your privacy data sharing concerns so I think many companies will not be happy to just send all of their time series to some external service to get them forecasted um will you need you know do you need a five minute lead prediction real time uh does it need to run on some Edge device on your wind farm all of these things uh yeah depend depends on whether it makes sense to run these models and then well one thing would be obviously then will there be open source pre-trained models out there at some point to get around some of these things so right now still we don't know um and I'm I'm sure there will be very interesting things happening the the years to come uh but well I mean right now we have Solutions at work uh with machine learning so I will just uh talk about where we at right now um and many of these things I'm sure can be also useful to to what's for what's to come good so let me very briefly talk about my background um why should you be yeah okay exactly very good question from from ban like how interpretable will these models be uh like time gbt uh results um exactly I think that's a very big research question as well right now and and not only in the time series domain right for these uh for these um large language models people also ask these questions I I was in a talk uh yeah two days ago where they asked some of these questions I think there will be some answers at some point but but it's still just you know quite open of research okay now um quickly about my background so you know why should you bother listening uh to my talk so I I have been in forecasting since the start of my PhD uh here at University of Granada about 2009 I then went on to Mones University working for example with Rob Hindman I recently did a sabatical in Industry working at meta uh now I'm back at un University of Granada uh doing research and forecasting and so some topics I have worked on would be forecasting for for detection in wind turbines for past outbreaks in green houses for predictive maintenance and renewable energy production forecasting for wind and solar energy demand forecasting energy price forecasting uh also retail demand forecasting supply chain type of applications um yeah so so so that gives you kind of an idea of what I forecast right so I I think forecasting yeah we can do sales forecasting in supply chain and Retail in energy there are many forecasting uh problems like demand price renewable energy production also building efficiency predictive maintenance for Road ra railroad mining other types of Machinery uh but strangely enough uh there are also some well interesting enough there are also quite some things that I don't forecast and that's what usually people ask me first like oh do you do this do you do that and two notable examples are the weather so because for the weather we have lots of domain knowledge and specialized models so we leave it to the experts in that domain which are the meteorologists but we often use weather forecasts as some external inputs to our models another thing I don't really for is the stock market um and why is that so I mean there are some obvious uh Pros to forecasting the stock market you have heaps of freely available high frequency data you have a very clear motivation and the use case which is you want to make money so I think that's why a lot of academics uh you know machine learning academics wanting to write papers about forecasting that's why they are drawn to a stock market forecast in because they just have the data anywhere available and and it's an obvious use case so but there are also a lot of drawbacks to stock market forecasting so the share price is usually not a function of its own past but of its anticipated future right so um yeah you have very uh extremely low signal to noise ratio and another thing that machine Learners that are not like in finance for example do often not understand is that there is uh some well markets tend to be close to efficient um and what that means so efficient market means that any information that is available is already included in the current price so that means in an efficient market you should not be able to do forecasting uh Beyond a knife forecast and then you know people got like Robert angler got a noble price in 2003 for methods of analyzing economic time series with time varying volatility uh which is about volatility forecasting right so in in finance it's usually all about predicting volatility predicting risk and not returns um yeah because there is kind of this consensus that predicting stock market returns from just using some lags which means just the part some you know just the past of the series that will not work no matter what method you're using um obviously what could work is if you if you get some you know very good cor variates I mean what people do is they they kind of check s uh sentiment on social media like on Twitter and then they try to use this sentiment to predict stock prices I mean I think it's all about like getting the right co-variates but if you just uh use a the past of of a stock price that's not going to work very well all this is also true in similar ways for exchange rates right I mean exchange rates are in this sense quite similar to stock prices uh this is very important still we will get to be get back to this later so a lot of like very recent literature in very good machine learning conferences tried to predict exchange rates um and yeah it doesn't work we we get to to it soon soon enough in about an hour so um yeah as I said that this talk is the main audience of this talk is machine Learners who are kind of new to forecasting in a certain way um yeah let me start off with a bit of forecasting terminology right so if forecast p is talk talk about model fitting parameter estimation that would be what a machine learner call as model training in Sample would be the training set out of sample would be the test set forecast Horizon that's kind of the target variable like do you want to forast one day ahead two days ahead forecast origin that's from where you do your forecasting from so that's usually the last known observation like the value from today for example rolling origin would be that your forecast origin changes so um for example over the test set you you you forecast from every point in the test set with information known up to that point and then fixed origin is just the forecast origin is fixed so it doesn't change um yeah and that's terminology I will use throughout this talk um some more terminology forecast combination that's what machine Learners call in sampling uh lags independent variables regressors covariates predictors that's what machine Learners would call features or inputs um a dummy variable and indicator variable that's what machine Learners would call one hot encoding seasonality seasonal period That's a cyclic change in the mean of the series and its length is known a priori and will not change in the future so I think this is very important so people in machine learning people often think like oh we have to find the seasonality uh and then it it could change and then I'm always like no that doesn't make a lot of sense I mean if I have a daily series I have a lot of reason to assume that I have a seasonal cycle of length seven right and let's say I have a series that has only trading days but or only only weekday my seasonal cycle would be five and if I go now five years into the future this cycle would still be five or would still be seven right so it's a quite strong assumption that makes sense but now if you say oh this seasonal cycle I just detected from the data and then I have it variable and it can change it can sometimes be seven sometimes eight it's usually not at all Justified um yeah so so that's the idea of seasonality doesn't change and it's known beforehand uh trend is usually a smooth change in the mean of the series um yeah good so now I I want to spend a little bit of time not too much to to just uh talk about the you know the the most basic forecasting methods that have come out of um the field of forecasting right so there is something called the knife forecast where we just use the last known observation as our forecast and if you think about it what you effectively do you're waiting your last observation with a one and all the other observations with a zero right so because you only take the last observation and you ignore the other observations that you have so mean forecast is kind of the opposite so you just calculate the mean of all your time series and you say this is my forecast and that means you you weigh all the observations equally but now if you think about it mean and na forecast obviously two extreme cases right in one only the last observation in the other one all the observations equally so because a central assumption in forecasting is often that the more recent past is more important than the more distance past so a quite obvious idea is what about we just wait the observations for example with an exponential decay and this is what we would call exponential smoothing so here is an example so this is a a time service this is our forecast this uh flat uh blue line um and the red line what this shows you is kind of the exponential weights I mean I've I've shifted them down so that you can see them better in the plot so they are not really at minus four but and you see kind of the last observation gets the most weight and then you weight it down and that's why uh the forecast is not exactly the last observation but it shifted a bit down uh yeah okay and this is what is called Simple exponential smoothing uh now exponential smoothing uh is basically this idea of simple exponential smoothing applied to components uh so you would have level seasonality and Trend component uh it dates back to the N to 19 57 this method and it was used for a long time as a quite ad hog method without any theoretical underpinning that was then one of well maybe the most important work of Rob Hinman to give it this solid statistical foundation and States space models and ETS actually stands for exponential smoothing or for error Trend and seasonality which are usually yeah the different parts that you fit yeah so another very important method is the autor regressive moving average or Amor method so so here the idea is that you basically just you see you have a linear model uh where you basically you see you get your current observation X at time t uh you go back to x t minus p+ one so this is like you input window you just have like a a linear model over this input window that you shift through the ser iies and then you also have a linear model over these over the errors that you made um over the last observations and now people always say well this is a linear model so in many papers you would read things like oh these traditional methods are all linear and now we do machine learning this is nonlinear well and for me it's always when I read this sentence I'm always like well but hold on I mean etss that even has EXP ential in its name and it's not it's not really linear but but if we now talk about linear what does linear even mean right so for example this model that we're looking at here this is linear in the lags and it's linear in the errors but when you now actually start fitting an armor model then you are like well hold on where do I get these errors from so these errors uh you need to actually step through the series to get these errors so you need to estimate some initial conditions and step through the whole series there's no CL closed form solution if you look at the r code of Arma it actually uses a bfgs nonlinear fitting procedure so and that's what makes armor models often actually quite slow so if you have a very long time series fitting an armor model is quite slow so yeah so for me it's always when people say well it's a linear model well but you fit it as a nonlinear model uh and and it's certainly not linear you know just in time in the time steps so yeah I never like this sentence um okay anyway so now we've we've seen what so basically this part here is the AR part this is called the AR part the linear model over the legs the linear model over the errors that's called the ma part so now what there's an ARA model so now what does the I stand for the I is integrated so the question is do you want to model the values directly or do you want to model the change of the last value so basically you're just differencing the series before you use the Rema uh so this addresses non-stationarity but it's actually just like a pre-processing step right so you difference the series before you do the armor um why would you difference because hopefully it makes the series stationary uh the problem is you do lose some information about the scale uh in the series in your pre-processing um and you know sometimes you would see papers where people say oh um like machine learning methods neuron Network they can only do with stationary data so let's run an ARA because that works on non-stationary data um and then whatever comes out of the Rema we fit that into the neuron Network and then I think well but the only only thing that the Rema does to make a series stationary is it does differencing of the series so why don't you just difference the series and then you run your neuron Network on this different Series right okay um good we we get back to that a bit later so now one other thing that I think is important mentioning um is this result that any stationary ar1 model has an equivalent ma Infinity model and any invertible ma1 model has an equivalent AR Infinity model so that means you can actually approximate this ma part in the armor model with a higher order AR part uh in practice this higher order is often not even that high so well you can use I don't know 20 or 50 or something like that like in in econometrics people often use ar2 AR3 and then ar5 is already quite high order so in that sense uh you don't really need this ma part I see it often more like a reparameterization to get fewer parameters smaller input windows at the cost of more complex model fitting um okay I think that's uh everything I wanted to to say about like these traditional methods just that we are kind of all on the same page when I say ARA ETS uh for for more details the standard resources obviously Rob hindman's and George asos's book that is very good that I I can recommend you um I put here fbp2 that still uses the old forecast package there fpp3 that uses the fa package um anyway yeah good so now I wanted to to go through a quite brief history of machine learning and forecasting before then we get to really the the message and and and how it's all done um so I mean well uh history of machine learning and forecasting in forecasting we obviously had these M competition so it started with M1 in 82 m stands for marakis uh who who who ran these compet who runs these competitions so and in the M3 com ah yeah and one thing I I should say is I think it is important to understand a little bit the history of machine learning and forecasting because if you are forecaster you will often see that machine learn if you are machine learner you will often see that forecasters are a bit um you know reserved to to to new machine learning results claiming that they have you know solved forcasting because this has a long tradition uh so I think understanding the history a bit gives you a bit of context of why as a machine learner you might get some backlash okay so I want to start talking about the M3 competition there were M1 M2 competitions before um and this M3 comp competition had as a central uh conclusion that complex methods do not necessarily outperform simple methods and in that context actually ETS and ARA were considered complex models and the simple models were things like a random walk with grift which means um So Random walk means like a knif forecast you just take the last observation and drift means you put a kind of a a constant change on it um so there was only one neuron Network participating in this competition it was quite bad uh the competition Was Won by the ca method which was later shown to be an average of linear regression and simple exponential smoothing with drift um and now this M3 data set that consists of yearly quarterly monthly data the yearly data have as low as 15 data points a monthly data has maybe 120 data points per series um now this data set was basically The Benchmark data set in forecasting for almost 20 years uh it has seen a lot of focus maybe a bit too much uh and I think on this data set to perform well as machine learning methods is quite difficult so yeah so basically since the M3 people thought forecasting methods need to be simple neuron networks don't work very well and and this idea continued up until easily 2010 uh 2012 maybe so there was this long-standing controversy in the forecasting field whether machine learning or statistical methods work better and and I think yeah the idea is that the forecasters knew that simple methods work best well because of them three but also afterwards there were competitions called nn3 nn5 particularly four machine learning methods and even those were won by simpler methods uh and machine Learners why knew that neuron networks work best because machine Learners always know that neuron networks work best because it's a universal approximator so there are all these papers out than the literature uh in the new network literature that are things like you know a novel method X for stock market forecasting uh yeah they often use cherry pick data sets no proper benchmarking you know if if you see a paper that talks about stock market forecasting and then they say oh and we've tested our method on these three stock market time series then for me the first question is always well if I go onto Yahoo finance I can can easily download 500 a thousand stock market time series why are you using these three right and if the paper doesn't present you with an with a really good answer chances are that it's just the three methods Series where they method work well yeah just as a side note also every time I hear somebody say informa is the state-ofthe-art in forecasting um I get reminded that this is still very relevant today because um well I don't want to open a can of worm worms here but the original and foral paper does not have a good uh bench marks and good um experimentation so I don't think it's sufficient to really show that the method works it's also not showing that the method doesn't work right I'm not saying this but um yeah okay so yeah there was this controversy and and that actually motivated Spurs Mar rakus to write a paper in 2018 uh about statistical and machine learning forecasting methods where he benchmarked ml methods again against statistical benchmarks in this paper the machine learning methods lost and the paper is published in plus one because it actually got rejected at two to three neural network Outlets beforehand uh reasons for rejection where that uh you know um there are many mls that have proven to overcome the results provided uh that's what the editors and the reviewers said they didn't name any in particular so it seems MC darkis was quite upset over this and he did what he seemingly does best when he is upair he organized another competition the M4 this time so the context here is that in the M1 already it was similar so basically he had written a paper about uh Rema not working very well he got a lot of backlash and then he organized the M competition okay so the idea of the M4 was obviously machine Learners can submit their forecasts they show their methods work well and we settled this kind of question right yeah exactly so I think this is a bit of context that should show you why the forecasting field kind of got quite hung up on this controversy um I still find it a bit surprising as forecasting actually emerged from statistics just the same way that machine learning later emerged from statistics namely by a focus on out of sample accuracy instead of model properties um and that actually also led to some things uh being discovered quite independently in both Fields I think the best example is forecast combination in forecasting that has been around since 1969 whereas assembling and machine learning was invented kind of in the 90s um forecasting for a long time at this idea we need simple methods whereas in machine learning you had this idea of why you need to regularize um yeah so forecasters have really been used to M promising much and hardly delivering anything uh however this has really changed in recent years since M4 um I would argue well that also the ignorant arrogance and ignorance that machine Learners often have you know kind of this disruptive mindset like oh I just go into this research field with my Transformer and I eclips all the 40 Years of research that have happened there before uh this hasn't helped either it's still not helping um and also keep in mind forecasting it's usually not like image classification or NLP where you do have just very complex input data structures but zero noise uh and no co-variates and so on so forcasting is a bit difficult uh still machine learning now is usually the way to go uh I still think forecasting and machine learning can still learn a lot from each other so also as a machine learner it makes a lot of sense to look at what's uh there in in forecasting okay so let's get back to the M4 competition that was about 100,000 series across many different domains Yeah [Music] question okay good um yeah so and for competition so 100,000 series across many different domains hourly daily weekly monthly quarterly years yearly series you also needed to provide prediction intervals uh the results um I would would argue that most forecasters back then expected that a combination approach of statistical methods would win uh however those actually got second and third and to everybody's big surprise the winner was uh well now it's called esnn a high that is hybrid between recurrent neuron Network and exponential smoothing uh by isamic SM uh yeah it was a big surprise to many people that machine learning is competitive and and and that it works so well uh I kind of saw it coming uh for the re simple reason that two years before there was a much less U uh kind of uh noticed competition the CIF 2016 competition where actually I got beaten already by by Slavic Smo um it's an interesting competition because it was just 72 monthly series of lengths about 22 to 108 data points so I you know when I saw this competition I was like no way that machine learning works well on this data set like 72 series with these short Series so I I participated just with plain ETS and with back DTS which was an algorithm um I had developed earlier together with some other people um and so actually back DTS won several subcategories uh and it also won on median Sate however the competition metric was mean asate and back DTS actually was quite bad it got nines and ETS got third in first and second place where lstms uh from Slavic SM globally trained across all the series so and you see I think this is still a very interesting use case because here this is a small data set and still machine learning is able to outperform uh everything else and this was quite new so yeah I think the most important thing here is globally trained across Series right so um yeah and this gets us to the next part which is um now global models right so I think this is a very important concept now in forecasting as well the term Global models was introduced by by Tim yanovski in 2020 uh it is arguably not a very good name for various reasons uh but it's still the name we we got so we stick to it for now um so the idea was that traditionally you take one time series and you see that as data set what that means is you build one model for each time series um and then in the M4 for example you have these 100,000 series you try to make them very diverse so that you can you know get a good picture of which method works well on which but the idea was you train one model per series and now what happens is that actually the low sampling frequency and the non-stationarity structural breaks make usually that we just don't have enough data to fit complex machine learning models um yeah and as I said M3 and four data sets are put together under this Paradigm very different series very different frequencies so yeah and think about it if you have like a monthly series well or let's say you have a yearly series like how many years of data can you realistically get five years 20 years 20 years on a yearly series means 20 data points and even if you can get 20 years of data like how relevant with the data from 15 years ago still be to your current forecasting right probably the whole business has changed like a lot of things have changed so that the past from 10 years ago is just not relevant so this is kind of the limiting factor normally in getting data in forecasting um yeah but now the idea in global models is that you really have this Paradigm Shift of you see a set of Time series as a data set for example a set of Series in retail smart meters and so on and now you build one model across all of these series as I said we call that Global modeling also cross learning multitask learning with linear models we often call it pool regression yeah and now suddenly you have enough data uh and the Machine learning models are actually quite competitive so and I I think the way to think about it is the following if you have a local model that trains per series uh let's say you fit etss etss will maybe have only five parameters or maybe you know less than 10 parameters but now if you have 10,000 series on each series you fit a model with five parameters you end up with 50,000 parameters so now if you fit a global model just one model across your 100 about your 10,000 series you can obviously use a model for example with 5,000 parameters which is a lot smaller than a model with 50,000 perameters but now you have one model with 5,000 perameters so it's a quite large and complex model yeah so and now the idea is obviously that the complexity uh in local models grows when the data set grows because for every new time series you fit a new model you get let's say five new parameters so it grows whereas if uh the complexity of a GL glob model can stay the same right no you can just add more series and train the same model uh yeah so Global models can afford to be more complex uh and now the question is how to add this complexity well you can add it as longer memory so you just make larger input Windows uh you can obviously use nonlinear or non-parametric models so you would now use your neuron Network radient booster tree your Transformer whatever you want uh and you can also do some data partitioning uh and what that means is now you're draining the global model not anymore globally across all the data but only on some subsets and that's a reason why Global model is not a good name because often the global models are not trained globally but on subsets of your data um one very important point I want to make is that a global model is not a multivariant model uh why is that well so so in a global model you learn across the series across all series together but you predict every series in isolation uh and that means they can work on data sets where different series have different lengths uh they are not aligned like M3 and4 they do not take into account interactions between the series and in fact the concepts are actually actually orthogonal so methods can be locally univariate that would be ETS ARA and so on globally univ VAR that's what we talk mostly about in the following local multivariate and Global multivariate um yeah exactly before we talked about local univarate now global Univar later also a bit local multivariate we won't cover Global multivariant um but yeah let me let's see if I can actually do this I wanted to like paint here a little bit okay so I think that's what always helps me qu quite to to understand so so let's say these are all your time series okay um now if you have a multivariant model what you would probably do is you do something like this you take you need to take an input window that captures all the series together in one input window because you want to look at kind of the interactions and so on and now this this window you shift that just that way through your Series so you see in the multivariate model all of them need to be aligned uh so that you can take the window at a certain point in time and now this is how you shift the window and now you already see one problem of multivariant modeling is often that this input window is incredibly big let's say you have 20,000 Series this window is really big so you need to come up with a way of of making this Spar taking only some of the series and so on in a global model this is actually a bit different so if we again have all our our series um in a global model the input window well I should say input window during training is still just on this time series and then sequentially you go to the next time series and you basically go through your data like this yeah uh and then if you get an and then if you do forecasting you I mean your input window is just from one Series so you just go through that series that you forecast right now right so that's the difference between Global and multivariant models okay so yeah now well some history of global models they go back to the early 2000s then uh one interesting thing is that actually pool regression is a quite standard technique and statistics there's a book from 2007 talking quite a lot about it uh it just people didn't have the idea of using that for forecasting until about 2015 and then obviously 2016 uh was when well Slavic SM developed some of his ideas in the CIF competition then 2017 was de uh when at Alis mqr Inn that was also 2017 we had some early works as well and since then the field has quite exploded and has so much work um another quite interesting um resources are the ker competitions because and that has a quite nice paper from 2020 talking about these competitions because there were competitions already in 2014 2015 2017 where people actually used Global models gradient booster trees neuron networks uh that work quite well then obviously in the M5 that was held in 2020 that was dominated by light gbms uh de NB it's also was successful as I said before it's interesting that the temporary Fusion Transformer TFT that was already around at that time um nobody used it uh with one of my students I tried to use it we couldn't get good results um so yeah uh a team of my students they actually won 17th Place out of over 5,000 participants and what they did was they used an emble of a light GBM and a pooled linear regression uh and the interesting thing is is that actually just this linear regression would have been 19th Place yeah uh okay does the term transfer learning apply when using Global models that's a very good question [Music] um I think it does right so I remember writing a paper recently uh about so so you see there are a lot of U I mean there's another term called concept drift and actually we wrote a paper recently about concept drift in forecasting because so there we go um so yeah concept drift transfer learning um I mean in Time series obviously the series would have non-stationarities so that is kind of the same as concept drift uh and the difference is more that these machine learning methods they usually don't deal with it I mean something like etss that automatically follows the series so you don't need to do like concept drift adaptation of ETS um and yeah transfer learning is similar I mean transfer learning uh can be used for concept drift I think Global models obviously it's similar I mean uh I guess it always depends for for me a question is often um like people do things like they they train on a bunch of Time series and then they would test on other time series and for me the question would be why why would you do that right so usually you you you you would have history of all of the series that you forecast so yeah well obviously there could be Series where you don't have history and then you uh but that's a whole different problem but yeah it's definitely related yeah um okay yeah so Global models have shown success to a quite surprising unreasonable degree uh the idea was for a very long time that these series have to be related or similar in some way so that you can learn something useful across them uh then people started asking like what what does related mean is that correlated or what does it mean so it's certainly yeah in terms of the DG in of the data generating process [Music] um yeah and and then actually uh Pablo Montero Mano and Rob Hinman they had this paper in 2020 about global models where they basically showed that Global models can produce the same forecast as a local model uh without any assumption about similarity so the series don't have to be related it's much more about uh fitting complex patterns across series uh and the global model will even fit if you have many simple patterns that are quite different across the series uh this result well is remarkable because similar results in in multitask learning for regression don't exist so it's it's something that works on forecasting and I think one very important thing that you have to understand about global modeling is that um the question is what is your evaluation about if if your evaluation is you want to perform well on an average error across all your series then Global modeling will enable you to get this average error across all the series down right so that's the that's the point like uh people before always thought well I need to be good on every series uh and that's what I do with the local models and now I and in the end I average to see what I get now if what you in the end are optimizing for is the average error across all your series that's what a global model can do for you if you're interested in only the error of Five series out of your 10,000 series the global model might not uh do what you hope it does right so yeah so it's I mean people often say that a global model it kind of it sees these related patterns and it can translate them into other series I guess that's happening but it's also really about the global model optimizes directly this average error over everything good okay so now finally yeah let's get into really like the the main topic the machine learning methods for forecasting uh yeah okay that is a very good question like how many series do we need for okay so well to just to reiterate I think think about it in terms of how many data points do I have across Series right like you can it can make perfect sense to build a global model just for two time series right it's much more important how many data points do you have across the series and then well it does depend on how similar are they like if if these two series are very similar H you and you have to effectively only learn one concept it's obviously going to be easier if the two uh are are very different yeah okay uh actually we we we also have a paper I I didn't cite that here um that's uh have a malag a simulation study for Global models or something like that it's called where we look at exactly this like how much data do you really need how do do the series need to be for the global modeling to work okay machine learning methods for forecasting uh yeah exactly it is very similar to multitask learning yes good machine learning methods for forecasting so yeah we've seen nowadays they they work so I think in the forecasting field there's no question anymore if they do work they do work if you do Global modeling uh you can also use them for single time series if they are long enough right if you have one time series uh that has half a million data points absolutely use machine learning models uh and yeah you often have longer series because of the finer granularities you can have additional metadata uh and then I think the basic setup would be always your nonlinear Auto regression nonlinear non parametric Auto regression so you uh yeah you so you just get like lags have an input window then you can use your favorite machine learning method out of the box we've seen that can approximate an armor model um you just need to choose more legs uh and then it should work quite well so oh yeah okay thanks check for sending that link so but now so what is the biggest problem when we when we use machine learning methods for forecasting the biggest problem is nonstationarity so non-stationarity means the data distribution changes over time and actually uh it's funny but if you think about it actually uh many or even most real world problems have a Time component and changing distributions so just think about you develop an algor ISM for detecting cars on the street and you are given a data set from the 1970s right so cars looked very different in the 1970s maybe even like people look different the streets look different so so there is like this kind of concept drift over time pretty much in everything but in time serat is usually much more explicit and has more impact uh and to just make this very obvious so this is a Time series I just generated some random walk uh series which means you know it could resemble like a stock price and now this is my training set up to about 800 uh and now uh I train my models I've trained a random Forest svm neural network just with a couple of lags and now I just tested on the testing set right and what do you see here well I mean like the testing said that has values of of close to 80 the algorithm has never seen those values in the training set so it's not predicting them right it's kind of just going flat here so I mean that's maybe a bit a trivial case but but you see I think that really uh okay we got a question about retail um let's get back to that later right I I hope we have time but um yeah it's a very good question and let's see if we can answer it at the end so yeah so what you see here is that um I mean it's it's kind of an obvious example but I think it shows very well the problem like if you try to predict values you have never seen it's not going to work out of the box right and now the problem with non-stationarity is that I mean there's this uh this joke well you know it's it's a mass joke so I'm not sure how funny it really is but so there's this joke of the dividing the word into linear and nonlinear is a bit like dividing the word into bananas and non- bananas uh so because you know there are so many different forms of nonlinear of nonlinearity um so linear very clearcut concept nonlinear just everything else and for stationerity it's like the same thing right so we it's quite obvious what stationerity is but then non so stationerity is that you know the distribution doesn't change over time um but now non-stationarity there can be so many different ways of non-stationarity you can have a change in the mean which is seasonality or Trend you could have a change in variance like heteroscedasticity you could have stochastic Trends um and just if we look at some series you know when when people say ah I do time series forecasting I have a method for time series forecasting I mean looking here at these series they are so different right the the first one that has this multiplicative Trend and it multiplicative seasonality and like an upwards Trend this second one that is like a stochastic Trend the electricity demand series down here that has a very prominent uh daily and weekly seasonality not so much of a trend so stationeri can be very different and the way to deal with them uh changes right so one thing that I found um yeah so one thing that I found well and and again sorry for people uh that I didn't clearly announce that we we'll do two hours but but yeah anyway you can obviously watch the recording so yeah so how to achieve stationarity um I mean or how to deal with non-stationarity right so in economics and econometrics and finance many series kind of look like random walks so that's why differencing is a common tool there to achieve stationerity I mean we've seen the Rema model earlier that does differencing and differencing is often kind of seen as well this is a way to deal with stationerity with nonstationarity but you know I mean it only soles certain forms of non-stationarity and not others so you are losing in and you are losing information about the scale uh I mean you can solve that by maybe having like another input with scale um and while differencing can help to make machine learning models more robust but there are situations where it won't work right so uh this is an example I mean if my original I mean obviously it's a bit a artificial example but let's say my series is this exponential Trend Now by definition if you difference an exponential Trend you get back an exponential Trend so you see like uh the black is my original series I do a difference I get the blue series I do another difference I get the red series and you see differencing doesn't help you the series is still non-stationary uh yeah another example I want to show is this wind power forecasting example here that I while it's kind of artificially made up data but you see one thing that in wind power well you have your wind turbine it you have a very clear minimum and maximum like you you cannot generate less power than zero power um and the maximum is that wind turbine will generate maybe you know one megabot and and that's where it cuts out right and that's what we see up here like once you hit the maximum the wind turbine just doesn't produce any anymore now if you you get a difference of your series I mean you see here this just gives you a zero here so you don't really uh you lose this information that here in this patch of zeros the series has no way of going up anymore because it's already at the maximum it can only go down right so with the differencing you kind of lose this information okay um yeah now how can you model a trend uh well the the problem is it's it's not very well specified what a trend even is so it's usually just like a smooth version of the series and the problem is well if you D Trend you still need to forecast the trend and that is often just as difficult as predicting the series itself so DET trending the series well it can be done but it's often not solving the problem uh another another option is obviously to do things like logarithms or box Cox transforms uh it makes exponential Trends linear it's stabilizes sarian uh choosing this Lambda is often not so easy but yeah that's the way how you can model Trends um one thing that we do a lot is this Window Wise normalization that actually Slavic SM it back then in 2016 in this CIF competition and that we' adopted since then so so the idea is you you take your input window and we've talked about it before having a large input window is often good so you take a large input window now you can just kind of take the mean of that large input window and you normalize by that and that actually works quite well so yeah you want to avoid saturation issues of sigo t h activations uh it's quite simp similar to bch normalization uh and now if you think about it instead of saturating on the absolute value of the training data as what we've seen before uh it's now saturating on the absolute value of the steepness of the trend of the training data which is often not such a bad idea like you should probably not predict trends that are steeper than anything you've seen before that's like a recipe of getting a very bad forecast um yeah and if you combine that with something like a log that you do before uh yeah you can predict very steep Trends so what you see here is uh H the the black that's all the training data and now you see well you can't really see it but the green line it really follows very closely that blue line so the the machine learning method is now really predicting this super steep Trend even if it hasn't seen ever anything like that um yeah so like Window Wise normalization that works really well uh you see what um this is like an example right so so we've seen the air passenger series before that was like uh well maybe if we go back quickly that was here this top Series so it kind of it has a quite steep Trend and now if you do Window Wise normalization the series actually looks like this so and like that it's obviously a lot better model yeah um yeah what else is there to be said about Trends so uh be careful with strong Trends uh because exponential Trends in your series will slow eventually like every exponential Trend will slow eventually just because you will run out of so for example you know um if you want to forecast Facebook growth has Facebook had like a exponential grow over the last 20 years um yes well let's assume yes but now at some point they will just run out of people in the world to give accounts to so that Trend will slow eventually right so that's like a knowledge we have that ex that Trends will slow eventually even linear Trends are usually quite bold assumptions so experts have come up with this idea of dam Trends so what's important about D Trends is that they are usually not justified from the data you just build them into your model to be conservative about the forecasting and the idea is forecasting should always be conservative okay uh okay we got a a very good question about uh so basically that in statistics you can often like make assumptions about uh the outcome uh yeah well I I guess in global models the idea is uh to model everything from the data I mean talking about assumptions I've talked about seasonalities I've talked about trend that could slow I mean that would be assumptions you might want to build into the model another thing is obviously the probabilistic forecasting uh to which we will hopefully get uh soon as well okay now how do you model seasonality so uh there was some discussions about whether neuron networks can model seasonality at all or not some early work suggested they can model seasonality later work suggested you need to deseasonalize your data first um we now came to the conclusion machine learning models can obviously model seasonality but if they have enough data um yeah so we did this simple experiment we generated a sine wave let the neon Network learn it so um you know maybe two full periods is not enough but 20 full periods can be enough uh yeah so we usually assume we know the seasonality beforehand um yeah and and then the question is you know should you deseasonalize the data before you run the model or not so it depends on how much data you have right so you can and then uh yeah how to model seasonality you can have these seasonal indicators where you just have a categorical variable telling you whether this is Monday Tuesday Wednesday yeah just like one hot encoded uh that's problematic if there are many seasons now what this gives you are these kind of abrupt changes right if you if you model let's say the day of months you count here from one to 30 and then it goes back to one which uh it's not great for some algorithms so what people would then do is they would get these more continuous like uh s cosine waves and they're actually called Furia terms so uh yeah so like you would often use fuer terms as additional inputs for your method um yeah for deseasonalization there are different methods like STL mstl St Str um yeah that D seasonaly and then well as I say um we can deseasonalize the series and only feed the trend and remainder component into the algorithm uh you can think about that as some form of boosting where at first you use one mod model that model seasonality and then the remainder you feed that into your machine learning model uh it it does put expert knowledge into the model right and and I think the idea here is that uh if you do not have enough data to model seasonality it can be a good idea well if you don't have enough data to learn the seasonality from the data it may be a good idea to actually do these decompositions beforehand right so what we've talked aled about before this thing of well we have a monthly data set 70 time series we want our machine learning model to perform well uh yeah you probably want to model the seasonality yourself and not leave it to the [Music] algorism okay yeah um maybe let's skip about this one yeah I I think this is very important as well the normalization right so because uh in some forecasting problems the forecasts are within a pre-specified domain right something like wind speed wind power maybe even electricity price you know beforehand that all the values you will ever get are between let's say zero and 25 so non normalization is very easy you just divide by your maximum that you know you will never be over and that's it right but in other time series this is very tricky so if you have things like share prices web page hits uh the price of Bitcoin you know businesses that just grow fast I mean one thing I often see is that a lot of research that we see is done in electricity and in electricity we hardly have any Trends or they are very kind of small and then you work in a company uh that is a startup company Everything grows a lot so all the time series have these super strong Trends uh and and then you need to use very different tools to model these time series um yeah so again if the domain is limited the scale has information so if you are at zero you know you cannot go further down uh then sometimes things happen like that if the level is already high Trends tend to be less steep you see for a company that has currently 100 customers it's obviously a lot easier to double their customer base and for a company like Facebook that has two billion users to double that is is a lot more expensive uh a lot more difficult so um steep Trends will slow if your values go High um yeah so you can include information about the scale as input um error measures are often scale and variance so I think the example I want to give here is this time series right so i' I've taken this air passenger series I've made it artificially a bit a bit even steeper so and now sometimes you see paper where papers where people say well we we do mean normalization because that's you know what you have to do now if we look at this time series uh in red I've shown you the mean of this time series so if we now divide this series by the mean what we get is this series so you see the values here they are like to the power of 11 here it's a value of 10 so we've brought it down to uh to be values from 0o to 10 but the shape has not changed at all right so normalization here by just dividing by the mean has not helped at all on the other hand and this uh what we've seen before this per window normalization that obviously um works better okay one quite important other point I want to make is this direct output versus iterative one step ahead so these traditional methods etss ARA they all optimize for one step ahead accuracy uh and now if you predict further out let's say you know you have a daily forecast you predict one 24 points out uh you feedback the forecast into the model that can lead to error accumulation and it's often better to predict all the horizons needed um and now methods like neuron networks they can have multiple outputs they can have output windows with gradient booster trees what you can do is well they produce only one output at a time but you can actually um well you can iterate out uh so you just feedback the forecast and iterate out uh the model um you can also build well obviously a direct model where you predict directly The Horizon that you want and then you can also do mixed forms of it so I think that's something that is quite interesting so let's say you need to predict 100 points out you can build 10 models that all predict the next 10 days and then you feedback the 10 points and and we said you get to to the um window that you need so this is relevant if you want to go for really long Horizon so let's say you have a daily time series you want to predict 400 Days out then you need to do things like that um yeah and also to mention the M5 winning method was actually an symble of direct and iterative uh models um okay just quickly checking the chat uh yeah so gradient boost trees uh they were quite successful in many competitions like GBM the obvious choice uh cut boost uh performs ordered gradient boosting and is supposed to be especially suitable for time series we've tried it it works well especially with its default parameters light GBM is still usually our go-to method uh yeah seasonality and Trend handling as discussed earlier like fer terms Window Wise normalization you build one model uh per horizon or or or for a subset of Horizons and then you roll that out if you need more uh it works really well if you have many external variables you do the feature engineering uh and you can also use differences of the L legs and the original lags things like that um yeah then well holiday effects obviously monthly series things like number of trading days in the months you can put in promotions out of stock events weather and so on uh yeah and and these methods well I mean that was kind of the type of modeling that won the M5 um I would still argue that even now in in many cases you know depending on your data set size and so on that would still be like my go-to modeling method embling um Works in forecasting just as well as in any other area of machine learning has been heavily used in keg competitions ensembles of gradient boosted trees neuron Network linear models like Po regression even ensembles of local and Global models um yeah and it's uh it's known under the term forecast combination and forecasting and also there we know that it works it's a long time okay good so I think well that finally gets us to uh yeah so this concludes like the part of how do you do kind of traditional uh forecasting um obviously for the traditional uh well traditional forecast I mean forecasting with traditional machine learning methods um for that you obviously need to to uh think a lot about feature engineering we've talked about that you need to think about how to deal with Trend seasonality um now we get to the Deep learning I mean many of these deep learning models that we will look at they actually internally also do things like fua terms uh and and many of them actually don't have any Trend handling and they are only tested on series that have hardly any trend right like um energy time series and so on so keep that always in mind like this machine learning the question is always how do you deal with the non-stationarities how do you deal with the trend how do you deal with the seasonalities uh and then some of these models would have that already build in uh the seasonality handling at least and then Trends uh it's a bit a a neglected topic to a certain amount in the forecasting field because people often don't have trends in their Series so chances are that if you use a deep learning method on series with a lot of trends that it's not going to work out of the box uh okay with that uh let's get into the the Deep learning uh and I think I'll hand over to abishek for that if thanks yeah okay shall SC share my screen or shall I use yours uh that's a very good question what about I I sto sharing uh then you can share uh because then I take a quick break and I'm right back yeah okay okay it's good okay so uh let's look at uh some of the deep learning uh techniques for forecasting so as we know like deep learning is pretty uh popular methods to use uh today and uh usually like uh the methods that people use for forecasting comes from NLP methods so uh you look at people using RNN it's because uh we started using it because NLP space it's really uh similar to time serieses where they have sequences we have sequences so it's very natural to adop uh techniques from NLP into for forecasting like lstms attentions and Transformers Etc right but there are also a lot of differences between NLP problems and uh forecasting problems uh because with NLP don't really have this uh idea of uh H the recent observations being more important than the others but in forecasting space that's very common uh mentality right and then there is also long-term dependencies in uh nlps but for us like long-term dependencies are usually very simple and uh it's just seasonalities which is very easy to uh credit usually well there are a couple of uh good tutorials and uh materials on um on forecasting with deep deep learning methods uh which we have listed here if you're interested you could uh check the uh references for them and let's start with uh looking at the rec neuron networks right so Rec neuron Network seems to be an obvious choice when it comes to deep learning because uh it works for sequence data in NLP so it has to work here as well right and then uh the uh very uh it's the idea of that working is because of Auto regession because it holds a internal state memory where it remembers what happened in the past right there is a memory that holds there and then so it's similar to how Auto Recreation works but then in practice we also include uh sliding window techniques like input and output Windows because uh that helps ease the uh process of Auto regression for neural networks so you make the state less important but uh focus more on the recent history by using those windows right and and then uh there are a couple of papers here we have listed uh which talks about uh the best practices of using fre neural network for forecasting uh and um majority of them uh use uh re majority of the recent papers use Global model method for uh rnns but you can also model uh univ serieses as a un model as well if you have enough uh length of uh the series so if you have enough history you could always uh go for that as well and uh there are a couple of problems that uh RNN SP as well right so if you look at RNN we have the problem of managing gradient which is a pretty common issue that people face when they train RNN models but but uh it's getting better so if you use lstm it kind of mitigates the problem of Vanishing gradients but there is also uh it's still the problem doesn't disappear it still exists so that's when like people started moving towards convolution neural networks and our convolution neural networks uh works a lot faster than RNN as well because uh you just have a ond convolution and uh you act on the entire window but the problem of uh ution is that it doesn't have a state memory which means that you should uh feed in entire seasonal window as your uh input windows or like basically you need a long Windows uh with dilations since it doesn't see everything uh all at once so there are a couple of other new techniques that are arising in the CNN as well uh which are like temporal convolutions dilations and wavenet uses causal convolutions so there are other approaches that are slightly different from the traditional convolution uh approaches that are coming up and uh it's being uh recorded and evidently proven that it's working better and lot faster than rnns as well so if like you have uh longer History of Time series and you don't need smaller Windows uh the approach CN approach would be more preferable than rnn's but for rnns uh especially with like time serieses with like short input windows or something or like uh series with holes uh or let's say uh irregular uh sizes or unequal sizes lengthes lengths of series RN would be a better preferred uh approach for deep learning methods there are a couple of uh specialized architect textes that have been published recently uh which includes uh the uh DPR from Amazon uh which uses a generative RNN model and does probabilistic forecasting and then there is a deeps state space model which uses rnns to parameterize the linear State space uh models and then there is uh deep factors for uh forecasting where they have two two different they model two different patterns one is the local probabilistic patterns and then there is a global time series patterns and they combine them together as uh to give a better so the local uh predictions also uh understands what's happening in a global level of Time series and then there is NBS which uses basis functions to decompose the series into uh seasonality and Trend and then uh they model the uh residual uh using stacking architecture as well so and N beats has uh really good uh evidence of working well uh because we have the M5 and M4 competitions where the winners have been using them can we get yeah let let me let me talk about these and then later you can uh well let's let's see if we can go a bit quicker and we still have time for the hierarchical stuff and then can talk about that um let let me share my screen again [Music] maybe good yeah you see that yes okay good so talking about the uh Transformers yeah so I mean there are these early works of Transformers for forecasting like this temporal Fusion Transformer uh the temporal Fusion Transformer IT addresses specifically common time series problems uh like incorporating static uh and dynamic cor variates both past covariates and future covariates into the Transformers it helps you to obtain prediction intervals and so on uh so I think well it's it's it's very good work in that sense and it's reported uh to work well in practice in many situations so um people are using it right it does work um keep in mind TFT was already in existence at the time of the M5 so I'm not sure why uh was not used by any winning team maybe because the data was too intermittent um yeah okay and now since uh ah yeah good question future coate and past coate yeah true I did didn't explain that so basically uh a future cor variant would be something like something that you already know is going to happen something like a holiday a promotion so for example H if you know so we know 25th of December this year will be Christmas so it's something about the future that we know whereas the past cover it would be something like um I know yesterday the price was this and because of that I feed that into my model to predict tomorrow's value yeah so some covariates are known and and others yeah yeah exactly but so you need both uh so the the comment from the chat is you need both past and future holidays yes but the way that you model it in your window right so in your input window uh you put in only kind of the you say uh in this this day in the future I know it's going to be a holiday right and sure in your training set you you have known that from the past so you modeled it accordingly but conceptually as the current input to your model it's something that you know about the future yeah good so yeah now at Transformers I mean well I've talked about it at the beginning there are these uh Foundation models they are coming so I'm by no means claiming Transformers are not good for forecasting uh but what I'm saying is that they are not the Silver Bullet at least not right now right so sometimes they work sometimes they don't let me tell you why I think that so I mean since about 2019 there has been this Myriad really of variant that people have come up with inmer auto former ETS former Rob former fed former and 10 others similar to these um I think well they you they all have their plausi ideas of why what they do Mak sense and why it should work so uh definitely um one thing that I find a bit strange is that they usually solve what they call Long sequence time series forecasting uh which to me seems a somewhat invented problem like I don't know like or well or let's call it academic problem right so people solve this problem because the papers before them have solved that but is that really a problem relevant in practice I don't know I mean at least there are many other it's only a small subset of what we would care about in forecasting and now what many of these papers do is they evaluate on the same small amount of data sets yeah as I said that only represents a small subset of what you could encounter as forecasting problems strangely enough they all or many of them use an exchange rate data set uh where they failed to Benchmark against the naive forecast uh and we actually have shown in a paper that uh many of them lose against naive uh and even if they do win against naive I mean predicting an exchange rate 720 days into the future which is basically two years into the future you're predicting the exchange rate between you know US dollar to euro I mean you you tell that to a finance person and they're like that's doesn't make any sense so be careful with this also this exchange rate data set it only has trading days in it um and then some people find like a seasonality of you know cycle seven cycle 30 things like that and then I'm like well but your month is only 20 days long so how can you find a mon see seasonality link 30 so uh there's another paper that kind of showed that many of them also lose against linear models uh and that one really important thing is that most of the benchmarks they use in these original papers are one step ahead methods that are then rolled out for hundreds of steps whereas these informal models to direct modeling directly for like very large output Windows um yeah and what this paper shows is that if you do the same thing with linear models all directly trained uh actually the the they are not as performant as they claim to be okay so having said all of this as a cautionary T so I mean I've seen here situations where you know you spend three hours running on a very powerful GPU you're training to get a result that is worse than naive that's obviously word of caution but to be fair I mean people like Sal they have reported they use Transformer architectures in production they work well uh but they obviously train on really large data sets like they have 350,000 daily series of s product sales and so on so if you have large amounts of data Transformers are probably uh a good uh message if you have small amounts of data um you definitely want to Benchmark against like the gradient boosted trees I've talked about before the uh and maybe even simpler methods right I mean things like uh if you do electricity demand forecasting a seasonal naive should be always your benchmark and because it will be quite difficult to beat okay and similarly if you do like share price exchange rate forecasting naive forecast needs to be a venture okay so now some resources GL on TS uh it's it's a very nice Library python library from Amazon that implements many of these deep learning architectures then our resource forecasting data.org where we have a repository as many data sets we have code end results where we run many of these methods and where you will see that the results are quite mixed sometimes on I mean there are very small data sets on which trans formal methods work well and then there are relatively large data sets where other things work better and now of course you could say well you didn't tune the hyperparameters properly um maybe you know this and that but the question is if you use these methods in production yourself then that's also a problem you will have right so like how to tune the parameters correctly and so on I think that goes goes into the mix so in that sense I think this is a quite realistic scenario where we run these methods and you know maybe with very intricate understanding of how the methods work maybe you could get better results but I guess the same is true for every method okay uh so let's have a look into multivariate forecasting so we've heard before that it's not the same as Global modeling uh because series can influence each other through cannibalization substitution effects and so on I I think it could be like the Holy Grail in retail demand forecasting uh it's a quite difficult problem active research area uh yeah so the main problems as I've said before are like how to scale the methods to really thousands of Time series uh how to have products come and go how to model changing relationships and usually we have some spareness constraints here uh I mean some methods I I want to mention here is U uh well paper from 2016 about Matrix factorization then 2018 LST net that's a quite famous paper where they do CNN cnns that feed into RNN with skip connections 2019 these are kind of the same people from 2016 again Matrix factorization with some causal convolution and attention then 2019 we have a paper from from the Amazon team with scan copas this method is actually implemented in gants under a name called de deep V uh then we have uh some graph neural networks there is a paper w at all 2020 where they kind of learn the the the graph from the data then abishek uh who has presented to you just now has a a quite nice paper from 13 where we look at like how to initialize the adjacency Matrix uh yeah and I mean this is still a quite active research area we also have some some work happening here but very often these methods have kind of scaling scaling issues yeah okay forecast evaluation that's a very important topic and one that I like to talk about way too much so I thought maybe in this talk let's keep it short uh we've already had some at least one question in in the chat so yeah we we'll have to skip this part today but I think uh by now we have put together some quite nice resources so I I recently wrote a paper uh together with hansika he malag about forecast evaluation for data scientists common pitfalls and best practices so the question that we had in in the chat about which error measure should you use so in this paper you know we gathered all the error measures we could find in the literature I think it's about 40 of them and then we we we talk about when is an error measure good when is it not good the interesting thing is that there is no single error measure where you cannot think of a use case where it does or a situation where it does something that you don't like so yeah it's surprisingly difficult which error measures to use we have this paper covering it I also wrote like a bit more an applied paper in the international Journal of Applied forecasting about the same topic uh I have a talk at ISF this year you can check that out on my web page but apart from that I think yeah we skip over this topic uh and we get straight to the probabilistic forecasting so yeah there I mean that was also a question in the chat already already like how how do you do like how do you get like a a distribution right because so far we've pretty much only talked about Point forecasting and many people would argue that a point forecasting in itself is not what you want like you want a distribution right uh I agree with that to a certain amount um so yeah so you have probabilistic forecasting very mature research topic as well these are all the topics I all the ways I've come across how you can do probabilistic forecasting it's not an exhaustive list I think there are others um but anyway so what you can do is you can do analytical prediction intervals you can do bootstrapping you can use like a Bayesian model with mcmc sampling for example uh you can forecast the parameters of a distribution you can do quanti regression with pinball loss uh uh you can do and you can determine the uncertainty empirically through back testing that is also called conform a prediction uh and you have another approach that is not as recent anymore that is this level set approach and and another and another couple of approaches that I didn't put here so let's start with the first one the analytical prediction intervals so this works well for some well understood models uh you would usually assume normally distributed errors that's what etss and Rema do for example and one problem that you have is that the intervals tend to be too narrow with that right so here I show you an example uh the the inner interval that is I think 80% U prediction interval and the out interval is 95% uh prediction interval um yeah the next thing you can do is simulation and bootstrapping that's usually quite slow because you need to run the model many times the intervals would be usually too narrow as well uh because you consider normally well with similar like if you just rerun your neural network with different seats on the same data the intervals would be too uh small because they only consider the parameter uncertainty of the model and they don't consider the data uncertainty the modeling uncertainty uh yeah you can bootstrap the residuals uh or again assume a distribution then you simulate forecasting paths by feeding back these bootstrap values uh and well if you want to know more about this that's also in in in the fpp book um yeah mcmc sampling that works well for for certain models we have a model called GT which is some like ban exponential smoothing version that does that it works quite well but it's slow um now this uh uncert determine uncertainty empirically through back testing now that's an interesting one so to me I mean my understanding is that this is really the same thing as conform a prediction right conform a prediction obviously put some theory behind it um it it has been used by companies in practice for a long time it leads to more realistic prediction intervals um you potentially need a lot of past data and rolling forecast so the idea is you have a validation set and now you check kind of off your method uh how how big do I need to make the interval on this validation set for the interval to have the right size um yeah it works well uh but now you have this additional validation set so the question is well how could we get rid of that so you want to do some cross validation or bootstrapping schemes as well uh but but this is a a way that that people nowadays use to just transform Point forecasts to probabilistic forecasts yeah forecasting the parameters or distribution so you could assume a normal distribution and then you just forecast mu and sigma of this that's actually what D does d has a normal distribution and a negative binomial distribution there's this NG boost method which is like a grad in boost tree that does that it was quite hyped in 2020 when it came out but I've not seen used a lot I I think it's quite slow apparently uh the drawback of this is you have to assume a certain distribution which is good if you have a limited amount of data uh and knowledge of your distribution but if you have a lot of data maybe you want to just learn this distribution from the data and this is kind of what you can do with quanti regression so uh yeah it's implemented for example in this when at I 2017 paper that's a quite nice kind of classic approach uh you don't need to do any distribution assumption it's fast to compute easy to implement uh the drawback is you only get certain quantis in practice maybe just five or seven quantis are enough anyway right so do you just need a prediction interval or do you need the full distribution um yeah with NE networks you can even fit different quanti at the same time uh and you can interpolate between these quanti to get a full distribution which was done by this paper in 2019 pinball loss function I think I'm going to skip over this um now the question is how do you evaluate uh probabilistic forecasts and I guess the message here is it doesn't get any easier than for the point forecast so you can have this simple hit Miss rule uh where you say well how many of my point like if I have an 80% forecast interval do really 80% of my ACT l in that interval the problem is you can come up with trivial Solutions where where you make the interval sometimes very small and sometimes very large um so you need to some other um ways to evaluate as well and then I think this one I don't want to go too deep into it so you have this mean scaled interval score that measures the size of the interval and then penalizes if you are lying outside of the interval by how much you lie outside of it uh what is this okay sorry what happened okay sorry I Som skip to the wrong slide uh yeah so msis so then you have other measures like scale pinball loss weighted scale pinball loss CRPS uh things like logarithmic score energy score variogram score so yeah as I'm saying I think my message here is just um forecast evaluation is difficult with Point forecasts with uh probabilistic forecasts it's also difficult but but there are obviously ways to do it okay now we want to take the last five minutes maybe or so to go into some special forecasting problems uh we actually well there are many of those right the ones I could think of are external variables intermittent data hierarchical forecasting forecasting in retail interpretability causal inference uh I have material for all all of those but I think we want to look well I wanted to look into intermittent data and hierarchical forecasting I think we might or well maybe very briefly go into just this intermittent forecasting so there's this classical method Cron's method where the idea is you have two series one that omits all the zeros and you uh only model the demand when it's non zero and the other One models the weight time between the events uh and you can do exactly that and and then it uses simple exponential smoothing for both Series so now you can obviously do this but with a a way more sophisticated model than SCS that's what would be called like a zero inflated model where you have one model that predicts the probability of zero and the second model that predicts the value under the assumption it's non zero you have some other like you could use specialized loss functions for example the light GBM has this TWD loss function implemented um and I think these are the two big approaches and now which one works better I I guess it depends a bit on the amounts of zeros like if you have really like more than 90% zeros I would probably try this zero inflated one if you have fewer then just try this specialized loss function uh and with that we get to the last topic which is the hierarchical forecasting and I think I'll hand back to abishek for that um yeah so we have about nine minutes left so yeah just uh go through that and uh yeah if you yeah and I try to answer some of your questions in the chat in the meantime so if you have questions just pop them in the chat uh yeah yeah so uh let's go to hierarchical forecasting so hierarchical forecasting is a very common problem you see in a lot of spaces uh where you have spatial data or different categories or different departments uh and hierarchy is not just uh spatial there is also temporal hierarchies as well where you predict something at a weekly uh frequency and then you want to sum them up to monthly frequency as well right so the with hierarchical forecasting is that they have to reconcile properly so let's say that you're predicting sales at a particular store and then if you uh aggregate them back to uh Regional level they have to be a perfect aggregate of it right so uh if you say let's say I have 10 stores and each uh prediction is 10 and then I I make another prediction for region level and it doesn't match up uh then uh it's probably not uh cons consistent enough right and there is a good overview by Rob hangman which we have cited here so if you guys want to read more you could uh look into that so there are a couple of approaches to do this reconciliation uh to make sure that your uh lower level forecast or the upper level forecast are all consistent with each other so the ways of doing that uh would be one uh way is a top down approach where you predict the top level series and then disaggregate it to the bottom levels so you build models only for the top level serieses and not for the bottom level you just for the bottom levels you just disaggregate them and the question is how do you disaggregate one way of doing it would be using ratios from the past or train a model to disaggregate uh but it's still challenging to do and then there is bottom up approach where you predict the bottom level serieses and then just aggregate them up by addition or something right but the problem is uh the bottom level serieses are usually intermittent or noisy they don't have proper seasonalities or trends that you have in the bottom level uh imagine like uh retail scenarios right like you have uh different products and uh most products don't get sold uh at uh every particular store so you have a lot of intermittent sales in uh individual store level but when you aggregate them up you get a better SE I Trend at the national level or Regional level then there is the middle out approach where you predict the middle level uh hierarchies and then you disaggregate and uh aggregate them both to get the upper levels and the lower levels as well then there is uh the other approach uh which is quite recent the optimal reconciliation approach right so uh optimal reconciliation approach is where like you predict all the different levels and then reconcile them to uh make it fit perfectly or make it more consistent across all the levels the way that you do that is uh by doing it in two steps one step is your forecasting step which forecast for all the level all the serieses and then the reconciliation step where we use uh some kind of optimization to make sure the consistency is uh uh consistency is is more uh forced upon it so the way that you do that is Le Square optimization or uh minty optimization you could see uh the trace minimization paper from vicam so those papers talk about how to do this and then uh there is uh the other paper from uh panag Tillis so this one talks about uh more about the uh probabilistic uh uh reconciliation right so when it comes to reconciliation Point reconciliation okay so you just have to make sure that it sums up to the higher levels so it's pretty simple pretty straightforward but when it comes to probabilistic reconciliation how do you say that it is uh well how do you make this reconciliation so because you don't have a single point forecast instead you have uh probab probability density so how do you make sure that they are coherent across all the uh different quantales right and then uh that's where this paper comes in where they use uh the uh geometric view where uh they uh they consider each of these problemistic uh forecast or the problemistic density uh uh density functions to integrate uh into one uh over all the different different variables together so which means that you don't just add up for each and every uh individual but the uh the the formula is more like the marginal probabilistic density of the higher level should be the sum of uh all the individual variables of the uh Pro probabilistic density functions of the lower levels so in that way you can model the uh probabilistic hierarchical forast and some of the other works which involve machine learning techniques are being listed here so you could look at those papers as well if you're interested then uh there are other new techniques that are coming up uh in the machine learning space uh for example the end to end uh coherent probabilistic forecast where uh they do a end to endend training so all the previous methods you see two different me uh two different steps one is the uh forecast step where you for for different levels or particular level and then the reconciliation step which Aggregates or reconciles their forecast right but then these approaches where uh they do endtoend optimization just focuses on one one uh one single Optimizer that optimizes for both the forecasting and the reconciliation step which is more uh efficient as well but also uh since you optimize for a single objective it's much more uh easier and much more better to use but then the problem with these type of approaches is some of these approaches follow the regularization way of uh reconciliation which means that these reconciliations cannot guarantee consistency which means some so it it optimizes to be consistent but in some scenarios you you won't have a consistent forecast where the bottom levels might not sum up to the top levels and uh here are a couple of uh the uh the repository where the end to end forecasting uh approach is uh existing so you could look at the uh code there if you want to use them and uh hand over to Kristoff for the conclusion yeah okay so we managed just on time amazing yeah so uh well as conclusion so well I I I think one thing we wanted to show you here is that forecasting as a field has come a long way it's it's a great time to do forecasting as a machine learner or data scientist uh machine learning methods they become more and more competitive uh but still be aware of some of the common pitfalls for forting such as evaluation uh benchmarks data leakage uh and non-stationarity well and actually the so most of them we actually haven't covered them too much today right so evaluation benchmarks that's all on this paper that I pointed to uh we also talk about data leakage non-stationarity we talked a lot about that like that you need to address that in certain ways uh but yeah with that I think we conclude and uh Happy forecasting everyone thank you so much [Music] um good I think that's it uh from us we are on time I mean I I personally could stay some more time and answer questions if you want but obviously it has been two hours so anybody who needs to leave please be free to do so obviously thank you very much you very much really interesting session so let's give our audience another five more minutes to ask any question if they have would that be fine I I see one question on the chat yeah about the assembling um yeah I think assembling is always a good idea right so uh embling is as close to a free lunch as you can get right uh the the only thing to be aware of with the embling is uh many many methods already do heavy make heavy use of embling internally uh and the other thing is that uh embling works best if your methods are very different from each other right so ideally you have if you have like two completely separate teams that develop a totally separate solution to the problem and then at the very end you assemble these two together that is what usually works best um yeah and and also like the the the methods you emble together uh they should be quite good already so often if you just emble together one good s uh solution and one that is a lot worse that often also doesn't work yeah yeah exactly so it was a it was a quite quick going through everything um uh yeah I guess we and and well obviously towards the end where we then really talked about you know the the Deep learning and the hierarchical stuff and so on we we only very briefly mentioned a couple of things um but yeah I hope we gave you enough pointers for the literature okay what's the rec what's the recommended modeling approach in the context of retail forecasting ah yeah okay true good find I didn't talk about that um yeah so so I think here it really depends right um so the okay so the question is how do you address uh forecasting in a retail context for example where there are lots of new products every year uh I think it depends right so if you have no history uh then you need to find similarity in some other ways right so if you say I introduce this new product what do I know about this product I would have some meta information and then it becomes more of a maybe a clustering problem where you say these new products are similar to these old ones and in that way you try to infer what what could be their possible uh demand another uh that's when you have no history let's say you have one data point of history or five data points of History so that's where maybe it could could make sense to actually train a neuron Network right because in a neuron Network you can have a very short input window maybe just five data points as an input window you you have the state that still models everything so you don't need like in a CNN you need this large input window to cover all of your seasonality so in that situation often your recurrent new networks work well because you have some series that have very very long history and other that that have just like a very few data points of History um but then apart from that obviously Global models are good because they would uh learn from the other series as well so you don't have to predict with only five data points for example um yeah and then that's it I mean some methods even claim that they can directly work with this out of the box right I mean if you think of let's say you use a a a gradient booster tree a gradient booster tree can have missing values uh in your data so and let's say you have a lot of meta data that you put in as well right so you have an input window but you also have a lot of other information about your products that you put in and now maybe your whole input window is just Naas but the the method will hopefully still produce a reasonable forecast yeah I think that's what I have to say about this question good okay it's been a very interesting session and very interesting few hours thank you very much Dr Kristoph and abishek for accepting our invitation and we are instructors for this session we will share all the materials and the recordings soon and really appreciate your uh your session and your knowledge sharing with us thank you very much thanks very much was good fun thanks bye-bye see you
Up Next

Implementing WaveNet from Scratch in PyTorch (Paper Explained)
@CanConTech
14.8K views•2021-10-25

Building Real-Time ML Pipelines with Feature Stores and MLOps Frameworks
@ODSCAI
5.1K views•2022-02-20

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence





































