←Home KnowML
TL;DR

Forecasting is ordinary regression with one constraint bolted on: the training rows must all be older than the test rows. Break that and your error metric measures interpolation and says nothing about forecasting.

Everything else follows. Stationarity and differencing are needed because a model fitted on the past must survive a future with a different level. Rolling-origin backtesting is needed because k-fold does not respect the arrow of time.

The thing that breaks is almost never the model. It is a feature that quietly knew the answer, and the tell is a validation score that looks too good.

01The constraint that makes this a different subject

Look at the diagram before reading anything else. The rest of this page is consequences of the gap between the two halves of it.

The same 40 observations, split two ways
RANDOM k-FOLD — WHAT ORDINARY CV DOES both neighbours of this test point are in train one of them is the future ROLLING ORIGIN — WHAT FORECASTING NEEDS fold 1 fold 2 fold 3 the forecast origin moves forward, never back t train gap (embargo) test Top: 8 of 40 points held out at random. Bottom: three honest origins.
Look at the pulsing cell in the top strip. It is a test point whose immediate neighbours on both sides are training points, so a model only has to interpolate between two nearly identical values. In the bottom half no test point has a training point to its right. That is the only difference between the two designs, and it is usually worth several times the error.

In ordinary supervised learning the rows are exchangeable. A house in row 900 tells you nothing special about the house in row 901, so shuffling before splitting is free.

Time series rows are not exchangeable. Row $t$ and row $t+1$ are the same system one step apart, and they are usually almost equal. Random k-fold cross-validation puts one in train and the other in test.

Three separate leaks come out of that, and they compound:

  • Temporal leakage. Training rows are dated after test rows. The model has seen the level, the regime and the shocks it is being asked to predict.
  • Leakage through smoothness. Even with no future rows in train, a neighbouring row carries nearly the same target. Any model that averages nearby points gets the answer handed to it.
  • Preprocessing leakage. A scaler, an imputer or a target encoder fitted on the whole series before splitting. The test statistics are baked into the transform.

None of this is hypothetical. Here is the gap on a series with a mild trend.

Try it Score the same model twice — shuffled, then honestly
import numpy as np
rng = np.random.default_rng(0)

n = 300                                 # a drifting, cyclical series:
t = np.arange(n)                        # the level in month 300 is not
y = .05 * t + np.sin(t / 6) + rng.normal(0, .1, n)   # the month-1 level

L = 3                                   # three lags, the standard setup
X = np.stack([y[L-1-k : n-1-k] for k in range(L)], 1)
target = y[L:]

def mae(tr, te, k=3):                   # 3-NN: any model that averages
    d = ((X[te][:, None] - X[tr][None]) ** 2).sum(-1)   # nearby points
    nn = np.argsort(d, 1)[:, :k]                        # behaves this way
    return np.abs(target[tr][nn].mean(1) - target[te]).mean()

order = np.arange(len(target))
cut = int(.8 * len(target))
shuffled = rng.permutation(order)       # the ordinary k-fold split
print("X", X.shape, "target", target.shape)
a = mae(shuffled[:cut], shuffled[cut:])
b = mae(order[:cut], order[cut:])       # every test point after train
print("shuffled CV MAE %.3f" % a)
print("rolling-origin  %.3f   (%.1fx worse)" % (b, b / a))
Same data, same model, same metric: 0.195 against 0.993. The shuffled split reports an error 5.1x smaller, and nothing about it is a bug you could find by reading the code. Every test row had a near-twin in the training set. Ship the model that scored 0.195 and production will hand you 0.993, which is the number you should have been optimising all along.
Why a careful-looking pipeline still leaks

People rarely shuffle on purpose. They call a generic splitter, or a hyperparameter search with a default cv=5, and the shuffle happens inside a library.

The defence is not vigilance. It is that the only splitter you ever use on a series is one that cuts by date, so no code path exists that could shuffle. Scikit-learn's TimeSeriesSplit is the minimum version of this.

02Stationarity, and why differencing is the first move

A model fitted on the past is only useful if the past and the future are the same kind of thing. Stationarity is the formal version of that hope.

A series is weakly stationary when three things hold: the mean does not depend on time, the variance does not depend on time, and the covariance between two points depends only on the lag between them.

Almost no interesting series is stationary as recorded. Sales grow and prices wander. The classical answer is to transform until what is left is stationary, fit there, then invert the transform.

Differencing is the main tool. Replace each value with the change since the previous value:

$$\nabla y_t = y_t - y_{t-1}, \qquad \nabla_m y_t = y_t - y_{t-m}$$

The first removes a trend. The second, seasonal differencing with period $m$, removes a repeating annual or weekly pattern.

The number of ordinary differences is the $d$ in ARIMA. The number of seasonal differences is the $D$ in SARIMA, and that is all those letters mean.

Try it Split a series in half and watch differencing make the halves agree
import numpy as np
rng = np.random.default_rng(0)

n = 400                                 # a random walk with drift, the
step = rng.normal(.3, 1.0, n)           # canonical non-stationary series
y = step.cumsum()

def halves(s, name):                    # same statistic, two windows
    a, b = s[:len(s)//2], s[len(s)//2:]
    print("%-13s mean %6.2f vs %6.2f    var %6.2f vs %6.2f"
          % (name, a.mean(), b.mean(), a.var(), b.var()))

print("y", y.shape, "   diff(y)", np.diff(y).shape, "\n")
halves(y, "raw")                        # nothing about it holds still
halves(np.diff(y), "differenced")       # both halves agree: stationary
halves(np.diff(y, 2), "twice-diffed")   # variance paid, nothing gained
The raw halves disagree about everything. Mean 34.99 against 82.81, variance 404.72 against 133.58. One difference and they line up: means 0.31 and 0.21, variances 0.93 and 1.05. Now look at the third row. Differencing a second time roughly doubles the variance, to 1.79 and 2.20, and buys nothing. That is over-differencing, and it is the standard way people make a forecast noisier while believing they are being careful.

Two tests are used to decide how many differences you need, and they are mirror images. The ADF test has non-stationarity as its null hypothesis. The KPSS test has stationarity as its null. Running both and agreeing is more informative than running either alone.

Differencing is safe, standardising is not

Differencing only looks backwards: computing $y_t - y_{t-1}$ uses nothing dated after $t$, so it can be applied to the whole series before splitting without leaking.

Subtracting the mean and dividing by the standard deviation is different. Both statistics are computed over the test period too. Fit the scaler on the training window only, exactly as you would with any other model.

03Decomposition, seasonality and the baseline you skipped

Most series are a slow trend, a repeating cycle and a remainder. Separating them is both a diagnostic and a forecasting method in its own right.

The additive decomposition writes the series as a sum of three parts:

$$y_t = T_t + S_t + R_t$$

A series splits into trend-cycle, seasonal and remainder components. If the seasonal swing grows with the level, use a multiplicative form instead, or take logs first and stay additive.

Three ways to do the split, in increasing order of usefulness:

  • Classical decomposition. A centred moving average for the trend, then average the detrended values by season. Simple, and it throws away the first and last few points.
  • STL (Seasonal-trend decomposition using Loess) lets the seasonal shape change slowly over the years and is robust to outliers. It is the default worth reaching for.
  • Prophet-style decomposable models. Fit trend, Fourier-series seasonality and holiday effects as an explicit curve-fitting problem, then forecast each piece forward.

Prophet deserves a fair hearing. It handles missing data, irregular sampling and holiday calendars without complaint, and an analyst can override the trend changepoints by hand. It is a curve fit, though, and not a stochastic process: it models the shape of the series and largely ignores the autocorrelation in the remainder.

Baselines are not a formality

Three forecasts cost nothing and beat a surprising fraction of real models:

  • Naive. Tomorrow equals today. Optimal for a random walk, which is more series than you would like.
  • Seasonal naive. This Tuesday equals last Tuesday. The baseline to beat on anything with a weekly or annual cycle.
  • Drift. Extend the straight line from the first observation to the last.

Report your model's error as a ratio against one of these. That ratio is what MASE formalises, and it is the only error number that travels between series on different scales.

Try it Put a fitted AR model up against a forecast with zero parameters
import numpy as np
rng = np.random.default_rng(0)

m, n = 7, 84                            # 12 weeks of noisy daily data
t = np.arange(n)
season = np.array([0., 1.2, 1.0, .8, 1.4, 3.5, 3.2])
y = 10 + .02 * t + season[t % m] + rng.normal(0, .8, n)
tr, te = y[:56], y[56:]                 # train 8 weeks, test 4 weeks

naive = np.full(te.shape, tr[-1])            # repeat the last value
snaive = np.array([y[56 + i - m * (i // m + 1)]
                   for i in range(len(te))]) # same weekday, last week
p = 14                                       # AR(14) by OLS on 42 rows
A = np.stack([tr[p-1-k : 55-k] for k in range(p)], 1)
w = np.linalg.lstsq(A, tr[p:], rcond=None)[0]
hist, ar = list(tr), []
for _ in te:                                 # recursive multi-step
    ar.append(float(w @ hist[-1:-p-1:-1])); hist.append(ar[-1])

print("train", tr.shape, "test", te.shape, "AR rows", A.shape)
base = np.abs(snaive - te).mean()
for nm, f in [("naive", naive), ("seasonal naive", snaive),
              ("AR(14) OLS", np.array(ar))]:
    e = np.abs(f - te).mean()
    print("%-15s MAE %.3f   vs snaive %.2fx" % (nm, e, e / base))
The zero-parameter forecast wins. Seasonal naive scores 0.881; the fourteen-parameter AR fit scores 1.022, or 1.16x worse. Fourteen coefficients estimated from (42, 14) rows is not enough data, and the errors compound through 28 recursive steps. Give it two years instead of eight weeks and the AR model wins comfortably. That is the point: 1.022 means nothing until you know that 0.881 was free.

04The classical toolkit, in one pass

ARIMA and exponential smoothing are the two families that dominated forecasting for forty years. Both are still the right answer for a single short series.

Start with the two building blocks. An autoregressive term regresses the series on its own past. A moving average term regresses it on its own past errors:

$$y_t = c + \sum_{i=1}^{p}\phi_i\,y_{t-i} + \sum_{j=1}^{q}\theta_j\,\varepsilon_{t-j} + \varepsilon_t$$

The $\phi$ terms are AR($p$), the $\theta$ terms are MA($q$). Apply this to a $d$-times differenced series and you have ARIMA($p,d,q$).

SARIMA adds a second copy of the same structure at the seasonal lag, written ARIMA($p,d,q$)($P,D,Q$)$_m$. Seven numbers, which is why automatic order selection exists. auto.arima and its Python equivalents search the grid by AICc.

Exponential smoothing comes at the same problem from the other end. Simple exponential smoothing forecasts a weighted average of the past where the weights decay geometrically:

$$\hat{y}_{t+1} = \alpha y_t + (1-\alpha)\hat{y}_t$$

It has one parameter. Holt adds a trend component, Holt-Winters adds a seasonal one, and the ETS taxonomy enumerates the combinations of error, trend and seasonality.

FamilyFitsReach for it whenWeak at
ETS / Holt-Winterslevel, trend, seasonone short series, clear seasonalitycovariates, long horizons
ARIMA / SARIMAautocorrelation structurestationary after differencingmany series, nonlinearity
State space + Kalmanlatent level and slopegaps, irregular sampling, online updatesspecification effort
Prophettrend + Fourier + holidaysbusiness series, analyst overridesautocorrelated remainders

The third row is worth dwelling on. A state space model splits the world into a hidden state that evolves and a noisy observation of it. The Kalman filter updates that state one observation at a time.

Two properties make it valuable. Missing observations are free, because you skip the update step. And the filter produces the exact likelihood, so the model can be fitted by maximum likelihood rather than by a bespoke recursion. Most modern ETS and ARIMA implementations are state space models underneath.

Where you are

You have the constraint, the transformations that make a series modellable, and the classical families that fit one series at a time.

The rest of the page is what happens when you have ten thousand series instead of one, which is where machine learning enters and where leakage gets much easier to commit.

05Turning a series into a table

Every tree and every neural network on this page is a regression on a feature matrix. Building that matrix is the whole job, and it is where leakage lives.

The move is mechanical: for each timestamp, build a row whose target is $y_{t+h}$ and whose features use nothing dated later than $t$. Five kinds of feature do almost all the work:

  • Lags of the target. $y_t, y_{t-1}, \dots$ plus the seasonal lags $y_{t-m}, y_{t-2m}$. The single most important family.
  • Rolling statistics. Mean, standard deviation, min and max over the last 7, 28 or 90 periods. These encode level and volatility cheaply.
  • Calendar features. Day of week, week of year, month, holiday flags, days to the next holiday. Cheap and known in advance.
  • Known-future covariates. Planned price, scheduled promotion, published holiday calendar. You may use their value at $t+h$, because you will know it.
  • Past-only covariates. Observed weather, competitor traffic, upstream demand. You may use their value at $t$ and no later.

The distinction in the last two bullets is the one people get wrong. It is also the heart of multivariate forecasting.

The multivariate trap, stated precisely

You want to forecast store sales and you have foot traffic, which correlates beautifully. Feeding today's foot traffic in as a feature produces a stunning backtest.

In production you are forecasting next week. You do not have next week's foot traffic. You would have to forecast it first, and its forecast error propagates into yours. Either lag the covariate so it is available, or accept that you now have two forecasting problems.

Rolling windows carry the same hazard in miniature. A 7-day rolling mean centred on $t$ includes three days after $t$. A 7-day mean ending at $t$ does not. One of those is a feature and the other is the answer.

There is also a choice about the horizon itself:

  • Recursive. Train one one-step model and feed its own predictions back in. Compact, and errors compound, as the AR baseline above showed.
  • Direct. Train a separate model for each horizon $h$. No compounding, more models, and no guarantee the forecasts are mutually consistent.

06Gradient-boosted trees, and why they keep winning

The honest summary of the last decade: for most tabular forecasting problems with many series, LightGBM on lag features is still the model to beat.

Once the series is a table, the usual tabular champion applies. Gradient boosting handles hundreds of mixed features, learns interactions without being told, tolerates missing values and trains in minutes on a laptop.

The evidence is the M-competitions, which are the closest thing forecasting has to a shared benchmark. In M4 (2018, 100,000 series) the pure machine-learning entries did badly. The winner was a hybrid of exponential smoothing and a recurrent network, and second place was an ensemble of classical statistical methods.

In M5 (2020, Walmart sales) the picture reversed. Gradient-boosted trees on engineered features dominated the leaderboard, and the winning accuracy entry was built on LightGBM. The difference was the data: many related series, rich covariates, real hierarchy.

A tree cannot extrapolate, ever

A decision tree predicts the mean of a training leaf. Outside the range of feature values it saw, it returns the nearest edge value and flatlines. On a growing series that means every long-horizon forecast is too low, systematically.

The fix is to remove the trend before the tree sees it. Difference the target, or model the ratio to a rolling baseline, then add the trend back afterwards. Check this first when a tree-based forecast looks oddly flat.

Trees also need a reason to know which series a row belongs to. Series identity, category and location go in as static features, so a single global model can serve thousands of series while still specialising.

07Deep forecasting: four architectures

The contribution of deep learning here is not a better function class. It is the global model: one network trained across every series at once.

Classical methods fit each series independently, so a series with 40 observations gets 40 observations of evidence. A global model trained on 40,000 related series learns the shape of seasonality once and shares it, and that is where the gains come from.

Four designs cover most of the field:

  • DeepAR (2017). An autoregressive recurrent network trained across many series that outputs the parameters of a distribution instead of a point. Forecasts are produced by sampling forward, so uncertainty comes for free.
  • N-BEATS (2019). A deep stack of fully connected blocks, no recurrence at all. Each block emits a backcast and a forecast, and the backcast is subtracted before the next block sees the input. Basis functions can be constrained to be polynomial or sinusoidal, which makes the decomposition readable.
  • Temporal Fusion Transformer (2019). Built for the messy case: static covariates, known-future covariates and past-only covariates, all at once. Variable selection networks score which inputs matter, gating layers let it skip unused capacity, and it emits quantiles directly.
  • PatchTST (2022). Chops the series into patches, embeds each patch as a token, and treats every channel independently. Patching shortens the sequence and makes attention see local shape rather than single points.

Dilated causal convolutions are the fourth family, inherited from WaveNet. They give a long receptive field with fixed-cost inference and no recurrence. Page 13 covers that stack in its original setting.

The paper that made the field check its work

In 2022, Zeng et al. asked whether transformers were helping at all on long-horizon benchmarks. They compared the published results against DLinear, a single linear layer applied to a decomposed series.

The linear model matched or beat several transformer architectures. The lesson is not that attention is useless. It is that the benchmarks were weak and the baselines were not being run, which is the same failure the previous section is about.

Take the ordering seriously: baseline first, trees second, and deep models when you have many related series, rich covariates and enough history to justify them.

08Foundation models for forecasting

Since 2023 several groups have pretrained a single model on very large collections of series, then forecast new ones with no fitting at all.

The recipe borrows directly from language modelling. Take a scaled window of values, turn it into tokens or patches, and train a sequence model to predict the continuation. Four public examples:

  • Chronos (Amazon, 2024). Scales and quantises values into a fixed vocabulary, then trains an off-the-shelf T5 on them. The series is literally treated as text.
  • TimesFM (Google, 2023). A decoder-only model over input patches, with an output patch longer than the input patch so it can forecast far in one step.
  • Moirai (Salesforce, 2024). Targets the awkward parts directly: multiple frequencies, variable numbers of covariates, and probabilistic output.
  • Lag-Llama (2023). A decoder-only transformer over lagged features, and one of the first open univariate probabilistic attempts.

What they buy you is a usable forecast on a new series in seconds, with no tuning. For a cold-start series, or for the long tail of series nobody has time to model, that is a real win.

There are two cautions. First, the zero-shot claims are only as good as the separation between the pretraining corpus and the evaluation set, and that separation is hard to audit from outside. Second, most of these models take the target series and little else, so any covariate you have worked hard to build has nowhere to go.

Treat a foundation model as a strong baseline that arrives for free. Then check whether LightGBM on your own features beats it, because on a well-understood problem it often does.

Where you are

That completes the model ladder: baselines, classical families, trees on lag features, global deep models, pretrained models.

Two things remain, and they matter more than the choice of model. Producing a forecast that admits what it does not know, and measuring any of it honestly.

09Probabilistic forecasts, quantiles and anomalies

A point forecast is a decision waiting to be made badly. Inventory, staffing and capacity all need a distribution, and a single number is not enough.

The cheap way to get one is quantile regression. Train the same model several times with a different loss for each target quantile. The pinball loss for quantile $\tau$ is asymmetric:

$$L_{\tau}(y,\hat{y}) = \max\big(\tau(y-\hat{y}),\ (\tau-1)(y-\hat{y})\big)$$

At $\tau=0.9$ under-predicting costs nine times as much as over-predicting, so the fitted value settles at the 90th percentile.

Whatever produces the interval, check its calibration. If your 90% interval contains the truth 62% of the time, it is not a 90% interval. Check coverage separately at each horizon, because intervals must widen as the horizon grows and often do not.

The metrics deserve the same care as the model:

  • MAE targets the median, RMSE targets the mean. On skewed demand data these are different forecasts, and optimising one while reporting the other is a common own goal.
  • MAPE divides by the actual value. It explodes at zero and it punishes over-forecasting more than under-forecasting. Intermittent demand breaks it completely.
  • MASE scales by the in-sample naive error, so it is comparable across series and has a clear reading: below 1 beats the naive forecast.
  • CRPS and weighted quantile loss score the whole predictive distribution. These are what probabilistic competitions use.

Anomaly detection is forecasting with the residual kept

The standard construction is to forecast one step ahead, then compare. A point is anomalous when it falls far outside its own prediction interval. Everything upstream on this page applies, because the quality of the anomaly flag is the quality of the forecast.

Two things cause most false alarms. Seasonality that the model does not represent, so every Monday looks like a spike. And historical anomalies left in the training data, which widen the intervals until nothing trips them.

10Backtesting that you can trust

Back to where the page started. A backtest is a simulation of deployment, and it is only worth anything if the simulation refuses to look forward.

Rolling-origin evaluation is the whole method: pick an origin, train on everything before it, forecast the next $h$ periods, score, move the origin forward, repeat. The bottom half of the diagram is the picture, and chapter 5.10 of Forecasting: Principles and Practice is the canonical write-up.

Four decisions inside that loop:

  • Expanding or sliding window. Expanding keeps all history and assumes the old regime still applies. Sliding uses a fixed-length window and adapts faster after a structural break.
  • Horizon. It must equal the horizon you forecast in production. A model tuned at one step ahead is not the model you want at 28.
  • Gap. Leave an embargo between the end of training and the start of the test window whenever features use rolling statistics or labels arrive late.
  • Refit policy. Refitting at every origin is honest and slow. Refitting monthly is what you will do in production, so backtest that instead.

Score per horizon and then aggregate, never the other way round. A model that is excellent at $h=1$ and useless at $h=28$ can show a perfectly respectable average.

The leaks, roughly in the order they bite:

  • A shuffled split, from a default cv=5 nobody read.
  • A scaler, imputer or target encoder fitted before the split.
  • A rolling window that includes the current point, or is centred.
  • A covariate used at $t+h$ that is not known at $t$.
  • Feature selection or hyperparameter search run once over all the data, then "validated" on part of it.
  • Restated data: a warehouse table that was corrected later, so the backtest sees numbers that did not exist at the time.
  • Deduplication or outlier removal performed globally before splitting.

The last two are the hardest to catch, because nothing in the modelling code is wrong. The data itself is the wrong vintage.

11Build this

The claim on this page is that leakage is worth more than modelling, and that is measurable in an afternoon.

Project Build one honest backtest, then break it five ways and price each break ~3 hours · numpy only

Generate a series you fully understand, build a rolling-origin harness, and then introduce one leak at a time. You are not trying to build a good forecaster. You are producing a table of how many points of MAE each mistake is worth, on data where you know the truth.

  1. Generate 2,000 points: a linear trend, a period-7 season, an AR(1) remainder, and two level shifts you place yourself. Keep the components, so you can check what any model recovers.
  2. Write the harness. Ten rolling origins, expanding window, horizon 14, scored per horizon and then averaged. Include seasonal naive as the baseline in every fold.
  3. Fit two models through the harness: ridge regression on lag features, and a depth-3 regression tree you write yourself. Record MASE per fold.
  4. Now break it, one change at a time, re-running the whole harness each time: shuffle the split; fit the scaler on all the data; use a centred rolling mean; add a covariate observed at $t+h$; select features on the full series.
  5. Tabulate the MASE each break reports against the honest number. Sort the table by the size of the gap.
You'll know it worked when the honest MASE sits somewhere near 1 and at least two of the broken variants report well below it. The broken runs will not look broken. They will look like progress, which is the entire problem.
What the tree adds. Run the tree on the raw target and then on the differenced target, across all ten origins. On the raw target its error grows steeply with the horizon, because it cannot predict a value above anything it saw in training. On the differenced target that growth mostly disappears. Plot both error-versus-horizon curves on one chart and the extrapolation limit becomes something you have seen rather than something you were told.

12What breaks

Forecasting failures are rarely a model that cannot fit. They are a number that was never measuring what you thought.

  • The backtest was excellent and production is not. The default outcome of a shuffled split. Nothing in the code looks wrong, and the gap only appears after deployment.
  • A covariate was not available at forecast time. Someone joined a table on timestamp without asking when the row was written. The feature importance chart looks wonderful.
  • The tree forecast flatlines. A growing series, no differencing, and every long-horizon prediction pinned to the top of the training range.
  • Nobody ran seasonal naive. Six weeks of modelling to land 8% worse than repeating last Tuesday, discovered by someone else.
  • MAPE on a series that touches zero. The metric goes to infinity, someone clips it, and the clipped version now rewards under-forecasting.
  • Intervals that never widen. The 90% band at horizon 28 is the same width as at horizon 1, so coverage collapses exactly where the decisions are risky.
  • A structural break mid-history. A pricing change or a pandemic sits inside the training window, and an expanding window keeps averaging over a regime that no longer exists.
  • Restated history. The warehouse corrected last quarter's numbers. Your backtest read the corrected version, and the live model never will.
  • Hierarchies that do not add up. Store forecasts summed do not equal the regional forecast, and the planning system needs them to.

13Where you meet this in the wild

The constraint is the same everywhere, and the consequences of getting it wrong differ a lot.

Retail demand planning

Tens of thousands of intermittent, hierarchical series with promotions and holidays attached. The forecast feeds an order that costs money whether it is too high or too low.

Global gradient-boosted trees on lag and calendar features, quantile outputs, reconciled across the hierarchy.

Cloud capacity planning

Strong daily and weekly cycles, a long-run growth trend, and sharp step changes when a large customer onboards. The cost of under-forecasting is an outage.

Decomposition plus an upper quantile, with explicit handling of known launches.

Energy load forecasting

The textbook case for known-future covariates: the weather forecast is available in advance, and temperature drives load nonlinearly.

Models that separate past-only from known-future inputs, such as TFT, earn their complexity here.

Financial returns

Prices are close to a random walk, so the naive forecast is near optimal and the signal is tiny. Leakage is fatal and not merely embarrassing, since anything profitable is arbitraged away.

Not a standard forecasting problem. Expect purged and embargoed cross-validation, and be sceptical of any strong backtest.

14Interview questions

BeginnerWhy can you not use ordinary k-fold cross-validation on a time series?

Because k-fold shuffles, which puts training rows after test rows in time and destroys the thing you are trying to measure. Forecasting means predicting values you have not observed, so an evaluation is only meaningful if every training point predates every test point. Two separate problems appear otherwise. Temporal leakage means the model has seen the future level, trend and shocks of the period it is being scored on. Leakage through smoothness means that even without explicit future rows, the immediate temporal neighbour of a test point sits in the training set and carries nearly the same value, so any model that averages nearby points is effectively interpolating rather than forecasting. The correct design is rolling-origin evaluation: pick an origin, train on everything before it, forecast the next h periods, score, then move the origin forward and repeat.

BeginnerWhat does it mean for a series to be stationary, and why does it matter?

Weak stationarity means three properties do not change over time: the mean, the variance, and the covariance between two points at a given lag. It matters because classical models estimate a fixed set of parameters from history and then apply them to the future, which is only valid if the statistical structure is the same in both periods. A series with a trend fails the mean condition and a series whose swings grow with its level fails the variance condition. The standard remedies are differencing, which replaces each value with the change since the previous value and removes a trend, seasonal differencing at the seasonal period, and a log or Box-Cox transform to stabilise variance. The ADF test takes non-stationarity as its null hypothesis and KPSS takes stationarity as its null, so running both gives a more confident answer than either alone. Over-differencing is a real cost: a second difference usually inflates the variance of the residual without removing anything.

IntermediateExplain ARIMA(p, d, q) and what SARIMA adds.

ARIMA has three parts. The AR component regresses the current value on p of its own past values. The I component is the number of times d the series has been differenced to make it stationary, so the model is fitted to the differenced series and forecasts are integrated back afterwards. The MA component regresses the current value on q of its own past forecast errors, which captures the effect of shocks that decay over a few periods. SARIMA adds a second copy of the same AR, differencing and MA structure operating at the seasonal lag m, written ARIMA(p,d,q)(P,D,Q)m, so it can capture a weekly or annual cycle without needing an enormous ordinary AR order. Orders were traditionally chosen by inspecting the ACF and PACF plots, which show where autocorrelation cuts off, but in practice automatic search by AICc over a grid of orders is the normal approach.

IntermediateYou are forecasting store sales and have foot-traffic data. How should you use it?

The deciding question is whether the covariate is known in advance at the time you make the forecast. Foot traffic is observed, not scheduled, so at an origin of today forecasting next week you will not have next week's value. Using it contemporaneously produces an excellent backtest and a model that cannot run in production. Three options are legitimate. Lag the covariate far enough that the value you need is always available at the origin, which is usually the right answer. Forecast the covariate first and accept that its error propagates, which means your evaluation must use the forecast rather than the actual. Or drop it in favour of covariates that genuinely are known ahead, such as holidays, planned promotions and published prices. The Temporal Fusion Transformer makes this distinction explicit by taking separate inputs for known-future and past-only covariates, which is a useful discipline even if you never use that model.

IntermediateWhy do gradient-boosted trees struggle with trending series, and what do you do about it?

A tree predicts the mean of the training examples in a leaf, so its output is bounded by the range of targets it saw during training. Outside the range of feature values it observed it simply returns the nearest edge value and stays flat. On a series with a persistent upward trend, every forecast beyond the training range is therefore biased low, and the bias grows with the horizon, which usually shows up as a forecast curve that goes suspiciously flat. The remedy is to remove the trend before the tree sees it: difference the target and have the model predict the change rather than the level, or divide by a rolling baseline and model the ratio, then invert the transform on the output. An alternative is to fit a simple linear or seasonal component first and let the tree model the residual. This is also the main structural reason the classical toolkit still matters, since ARIMA and ETS extrapolate trends by construction.

DeepDesign a backtest for a demand forecasting system that refits weekly and forecasts 28 days ahead.

Mirror production exactly. Choose origins spaced one week apart across at least a year so every season is represented, and at each origin train only on data available at that timestamp, forecast 28 days, then move forward. Refit at the weekly cadence production uses rather than at every origin, because a backtest that refits more often than the live system overstates its accuracy. Insert an embargo between the training cut and the forecast window, sized to the worst-case data latency, so late labels and boundary-spanning rolling features cannot leak. Score per horizon rather than averaging, since day 1 and day 28 are different quantities and the aggregate hides that. Use MASE so series of very different volumes can be pooled, include seasonal naive in every fold, and check quantile coverage per horizon. Finally, verify the data vintage: if the warehouse restates history, the backtest must read values as they stood at each origin, or it is measuring a system that could never have existed.

DeepWhen would you choose a deep forecasting model over gradient-boosted trees?

The decision is about the shape of the data, not the sophistication of the model. Deep models earn their cost when you have many related series, because one global network trained across all of them shares seasonal structure that a per-series model has too little data to estimate. That is the real contribution of DeepAR, N-BEATS and their successors. They also help when you need a full predictive distribution, when covariates are too heterogeneous for hand-built features, or when the horizon is long enough for patch-based models such as PatchTST to have something to work with. Against that, the M5 competition was won on gradient-boosted trees over engineered lag features, and the DLinear paper showed a single linear layer matching several published transformer results on long-horizon benchmarks. The defensible position is an ordering: seasonal naive first, then trees on lag features, then a global deep model only once a rolling-origin backtest shows it beats both.

DeepYour model's 90% prediction interval only contains the actual value 62% of the time. What do you investigate?

Coverage that far below nominal means the model is systematically overconfident, and four causes are worth separating. First, check whether coverage is uniformly bad or degrades with horizon, because intervals that fail to widen indicate the uncertainty is not propagating accumulated error, which happens when a recursive multi-step forecast sizes its band from one-step residuals. Second, check for heteroscedastic or heavy-tailed residuals, since a Gaussian interval fitted to demand with occasional spikes is far too narrow. Third, intervals derived from in-sample residuals assume the fitted model is correct, so a structural break in the test period produces errors the interval was never sized for. Fourth, check the evaluation itself, because coverage measured on a leaked split is measured against errors that were artificially small. The remedies are fitting quantiles directly with pinball loss, calibrating width on a held-out rolling-origin period, conformal methods adapted to sequential data, and always reporting coverage per horizon.

15Go deeper

📘
Book
Forecasting: Principles and Practice (3rd ed)
Hyndman & Athanasopoulos — free online, and the reference for everything in sections 02 to 04. Start here.
📄
Paper
DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks
Salinas et al. — the global probabilistic model that started the modern line — arXiv:1704.04110
📄
Paper
N-BEATS: Neural basis expansion analysis for interpretable time series forecasting
Oreshkin et al. — backcast/forecast residual blocks, no recurrence — arXiv:1905.10437
📄
Paper
Temporal Fusion Transformers for Interpretable Multi-horizon Forecasting
Lim et al. — separates static, known-future and past-only covariates, and outputs quantiles — arXiv:1912.09363
📄
Paper
A Time Series is Worth 64 Words: Long-term Forecasting with Transformers (PatchTST)
Nie et al. — patching and channel independence, the design most later work copies — arXiv:2211.14730
📄
Paper
Are Transformers Effective for Time Series Forecasting?
Zeng et al. — the DLinear paper from section 07. Read it before believing any long-horizon benchmark — arXiv:2205.13504
📄
Paper
Chronos: Learning the Language of Time Series
Ansari et al. — quantise the values into tokens and train a T5. The clearest foundation-model recipe — arXiv:2403.07815
📄
Review
Forecasting: theory and practice
Petropoulos et al. — an encyclopaedic survey by 80-odd authors, including the M-competition evidence in section 06 — arXiv:2012.03854
🏁
Competition
M5 Forecasting — Accuracy
Walmart hierarchical sales. The public solution write-ups are the best available record of what gradient boosting on lag features looks like in anger.
🔧
Tool
StatsForecast
Fast AutoARIMA, ETS and the naive baselines, with a rolling-origin cross-validation call built in. Run this before writing a model.

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

Other techniques for this problem

A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.

My Notes — 17 Time Series & Forecasting

Free notes

Highlights on this page