The unreasonable difficulty of time series forecasting

(suzyahyah.github.io)

89 points | by suzyahyah 3 days ago

19 comments

  • c7b 1 hour ago
    Well. Lots of math that boils down to 'predicting the future is hard'. Especially when the future is one of social construction, that's what gets lost a bit here. Predicting the future is easier for planetary motions than for Bitcoin.
    • gofreddygo 29 minutes ago
      The unreasonable difficulty of predicting (actual) future effects based on past effects with no regards to their cause.
  • klodolph 3 hours ago
    I see that a lot of these are markets.

    Yes, it’s hard to predict markets. Because anybody who can successfully predict markets, does so, makes money, and changes the market so their predictions lose their edge.

    Time series forecasts are a lot easier if you are forecasting, say, disk use in your servers or whatnot. (By “easy” I mean you can do a simple prediction and get useful insights.)

    • jldugger 2 hours ago
      It's true that markets are more _adversarial_. But there's still a lot of trouble with distribution shifts even in server metrics. As an example, our SRE team got paged a few times in the past month for traffic drops due to the World Cup. This stresses the nowcasting alert in several dimensions:

      - there's no seasonal pattern to the matches, they happen sorta randomly.

      - they drive increased query traffic in the hour or so before the game

      - then during the game usage drops, sometimes to below "normal" depending on time of day and who's playing

      So... now the accuracy of your forecasting tool depends on correctly predicting when world cup matches happen, and also who wins them!

      edit: and this is just one recent example. others involve severe weather, national gameshows, earthquakes, and when you celebrate christmas.

      • kjellsbells 12 minutes ago
        The UK power grid operators famously plan for massive demand surges at eg the end of major football matches. I cant imagine what their forecasters thought of the England-Mexico nail biter. Half the nation heading to put the kettle on, half glued to their seats. Would love to see the charts of demand now that the World Cup is done...
      • pishpash 23 minutes ago
        Unless disconnected from the world, or very thinly coupled, most processes become hard to predict in a generalizable way for this reason.
    • gregw2 2 hours ago
      Between the difficulty and value propositions for forecasting timeseries "disk usage" versus "financial markets" there are some rather relevant time series such as "company sales" or "demand for company product" that come up again, and again, and again but are neither "easy" nor "predicting-the-stock-market-hard".
  • EdwardAF-IT 2 hours ago
    A very interesting article, suzyahyah. I especially appreciated your definition of stationarity, a concept with which I struggled in my own time series class. If I understand correctly, it sounds like the basic premise is that a fundamentally statistical methodology (LLMs) can't realistically predict a non-stationary data generation, which makes sense.

    Separately, I've wondered for some time if there might be some reliable way to predict non-stationary data. While I don't have the answer, it occurs to me that it will possibly be a non-statistical method due to the fundamental incompatibilities. However, it also occurs to me that, given enough information, every data-generating process actually could be predicted. For instance, in the stock example, if you could model every single input into the system of a single company's stock, including every variable affecting every human that might conduct a transaction of it (daunting and unrealistic as that might be, but this is a thought experiment), then I believe the problem of prediction stops being non-stationary and in fact becomes completely deterministic, if complex. In such a scenario, wouldn't you be able to accurately make your prediction? I believe that perhaps chaos theory could present us with some solutions here where pure statistics (or, rather, simple statistics) cannot.

    Just my 2 cents..

    • chas 1 hour ago
      Chaos theory does provide us with a tool here: The Lyapunov exponent[1] and the related Lyapunov time[2]. These allow you to characterize how far into the future you can expect to be able to predict the behavior of a dynamical system. To Prerok’s example of predicting the weather, this is why 3 day weather forecasts tend to be great, 10 day weather forecasts tend to be good, and weather forecasts two months out are just the general average trends for the climate in that time and place.

      There are chemical systems where Lyapunov time is small enough that you can only predict seconds or minutes into the future and astronomical systems that are nonlinear and chaotic but have a long enough Lyapunov time that you can make reasonable predictions for millions of years. For both of these scales, the Lyapunov time still bounds how far into the future you can expect your predictions to remain near to the actual behavior of the system.

      I do not know the Lyapunov times of the financial markets. That said, mathematically-sophisticated professional analysis frequently get their predictions wrong in major ways, so I expect there is a pretty hard bound on predictive quality caused by a short Lyanpunov time of the markets themselves.

      [1] https://en.wikipedia.org/wiki/Lyapunov_exponent

      [2] https://en.wikipedia.org/wiki/Lyapunov_time

    • evo 1 hour ago
      I think there's a close linkage between "predictable" and "stationary". Ultimately, if you strip everything away, either there's a core where the future looks like the past (which is equivalent to being stationary), or there isn't. That core could be "the laws of physics and base conditions are stationary", and everything else is deterministic functions applied on top, but the fundamental process is trying to find that stationary core.

      In the finance space, the stationary core is often some 'stylized fact' that you're hypothesizing will hold true. This could be e.g., the momentum factor, that if you strip away the noise, there's an underlying trend that will hold over an extended duration.

      • pishpash 33 minutes ago
        That's a bit of a vacuous statement. Stationarity makes sample statistics meaningful due to the LLN, sure, but almost nothing is stationary (and even if something is, there is no way to know anyway, you can only assume). If you condition on enough variables and do enough transformations you may get something seemingly stationary and therefore trivially predictable. But all the practical complexity is in that structure you need to specify. The true prediction "problem", when people work on such "problems", really is in the stuff besides the thing that you can just use averages to predict.
    • prerok 1 hour ago
      The weather is an example of a non-linear dynamic system. Try as we might, we still cannot predict it really even a few days in advance. The stock market is much worse, even, so, no, it cannot be predicted.
  • smokel 3 hours ago
    This article focuses mostly on point forecasting. There are many other things that are interesting to forecast.

    Consider for example the use case of forecasting the average speed on a road segment with a maximum speed of 70mph. Forecasting whether that will be 69.8 or 70.3 is not very relevant. What is relevant is forecasting when the speed drops below a traffic jam threshold. But the exact timing of that might be impossible to forecast due to the inherently chaotic behavior of traffic. Forecasting the probability of a traffic jam occurring may be more interesting to practitioners.

  • gregw2 2 hours ago
    I was a little surprised to find someone experienced in time series modeling not mentioning bayesian modeling nor the use of ensemble models... but I do think its more fundamental points were worth making.
  • a-dub 2 hours ago
    > Given the lack of forecasting signal, the obvious next step then is to seek out external (exogenous) features in the real world that can help prediction models.

    this seems right to me. maybe another interesting approach would be a fusion llm+ts model that does multiple-input-single-output with input metadata and causality narrative. so it "thinks" about what data it has and how predictive it may be of the target variable and when something "interesting" occurs it uses the big priors to synthesize a good guess at what it would look like.

    so you'd have something like the time series data plus textual narratives of the causality stories as the training data.

  • verteu 1 hour ago
    As touched on in the article: The data's not IID, and there hasn't been very much "time" to generalize across.
  • joe_the_user 15 minutes ago
    I remember in 2017, it was said that "regular" machine learning techniques were better for time series than neural networks. It seems like this blog post is indirectly references that belief and gives. It seems like the idea is simpler ML can be tuned to a given time series whereas neural network require lots of training.

    Also, one method works better than another is more meaningful than "predicting the future is hard".

  • giyanani 2 hours ago
    > Given the lack of forecasting signal, the obvious next step then is to seek out external (exogenous) features in the real world that can help prediction models. After all many real world time series are event-driven (FX, bitcoin) and hence they are exposed to shocks and drifts which can also be measured or accounted for in the data generating process.

    This is the approach I'm using at my job, which is incident detection with customer metrics. We're tagging our time series data with common features -- such as country, customer type, etc -- with the idea that we can do a graph-like search to find exogenous variables. We can also use this to identify time series that have a similar "data generating process" and are simply different "realizations" of each other.

    We don't need great time series forecasts, just something that detects large deviations quickly. We can then add in an existing dataset of _known_ incidents, indexed by the same common features, as a training/validation set.

    • mr_toad 2 hours ago
      > We can then add in an existing dataset of _known_ incidents, indexed by the same common features, as a training/validation set.

      Beware of the Anna Karenina principle. Well behaved data might be explicable by the same common features, but often the anomalies all have unique characteristics.

      https://en.wikipedia.org/wiki/Anna_Karenina_principle

  • vjk800 1 hour ago
    Again and again nerds coming from maths/CS/etc to finance are surprised to find that financial time series are actually impossible to predict. After spending almost a decade in finance now with a similar background, the arrogance of the "let's just throw in some neural network/whatever and be done with it" attitude now amuses me. After all, if it was easy - or even possible with any kind of effort - to predict (even within some error margins or probability) what the stock prices are tomorrow, anyone doing it would quickly become a billionaire or trillionaire. If the person could keep doing it, at some point they would own enough of the stock market so that their actions would affect the prices and whatever pattern they found would vanish.

    Financial markets are not a natural phenomenon that exist unchanged regardless of whoever is observing them. Their dynamics continuously change in response to collective actions of all of the humanity.

    Oil price changed quite a bit when the US attacked Iran. If you are trying to predict the price of oil, your model would have to be able to predict Trump ordering an attack on Iran. Does your model include a full simulation of the mind of the president of the United States (and every other person who have any kind of impact on the world events)? If not, then your time series forecasts are not going to be that great.

    • pishpash 17 minutes ago
      Which naturally begs the question whether the excess returns in finance (outside of services provided for liquidity matching, risk transformation, etc.) is not just exploiting inside information of one kind or another. If you have inside information, your time series forecasts are going to be excellent and no simulation is needed.
  • hintymad 3 hours ago
    The examples in the posts suggest that the past does not contain all the patterns, or information in general, about the future. If so, isn't it natural that point forecast will fail in some cases?
    • pishpash 14 minutes ago
      The past cannot contain all information about the future, or there is some superdeterminism going on, by definition.
  • BrokenCogs 3 hours ago
    Not at all unreasonable, why should predicting the future be easy for those of us living in the simulation?
  • xnx 4 hours ago
    "It is difficult to make predictions, especially about the future."
  • joe_the_user 3 hours ago
    The question "what probability distribution generated this" is very hard to answer if you're only getting short window before the distribution changes. A lot of distributions could generate relatively "short" set dispersed data points and deciding which distribution is "really" creating the points might not even be meaningful. A distribution is a mathematical device, not something with a physical existence.
    • mjburgess 3 hours ago
      Probability distributions don't generate data, they describe our uncertainty about its generation retrospectively. So there is no answer to that question. The lack of such an answer is at the heart of why such inference is hard: in most cases, there isnt a stable underlying reality of anything doing the distributing.
      • joe_the_user 1 hour ago
        I think we're saying roughly the same thing. It's natural starting with a statistical model to say it generates data (as the article sometimes does) but in reality, you often/generated have data generated by a possibly deterministic and completely different process but satisfying (if you're lucky) the statistical conditions. Even in testing random/stochastic models, you would usually use pseudo-random generators - which are deterministic.
  • brcmthrowaway 3 hours ago
    Well, Jane Street cracked it
    • nyeah 2 hours ago
      Not by just feeding in a lot of old time series.