Imagine trying to predict the path of a single raindrop in a storm. You can’t know its exact position with certainty, but you know it’s part of a larger system-the storm itself-which has certain rules and patterns. In economics and finance, a time series dataset, like the daily price of a stock or quarterly GDP, is just like that single raindrop. We only see one path, but itโs drawn from a vast, invisible “storm” of possibilities. This underlying storm, the theoretical engine that generates our data, is what econometricians call a stochastic process.
Understanding this concept is the absolute bedrock of time series analysis. It’s the difference between simply describing data and truly understanding the process that created it. A stochastic process is formally defined as a sequence of random variables indexed by time. Our one set of data-say, 10 years of monthly inflation figures-is just one “realization” of the countless possible paths the process could have generated. Your dataset might show inflation rising, but another “realization” from the same underlying process might have shown it falling. Our goal is to use the single path we have to infer the rules of the storm itself.
Table of Contents
- What is this ‘data-generating process’?
- The building blocks: Key types of stochastic processes
- The foundation: White noise
- The “memory” model: Autoregressive (AR) process
- The “shock” model: Moving average (MA) process
- The hybrid model: ARMA process
- The all-important question: Is your data stationary?
- Why non-stationarity is a huge problem
- The great deception: Beware of spurious regression
- How spurious regression invalidates your results
- Putting it all together
What is this ‘data-generating process’?
Let’s stick with our storm analogy. The “process” is the entire weather system, governed by laws of physics (like temperature, pressure, and humidity). The “realization” is the specific path one raindrop takes as it zig-zags to the ground. In economics, the stochastic process (also called the Data Generating Process or DGP) represents the underlying economic structure. For example, a country’s GDP isn’t just a random number. It’s the result of a complex process involving millions of decisions by consumers, firms, and the government, all interacting over time. A time series of GDP data is just the one path that history happened to take.
Why does this distinction matter? Because in econometrics, we are detectives. We see the evidence-the data realization-and we have to deduce the nature of the culprit, the stochastic process. By understanding the process, we can do two crucial things: first, we can understand the past, and second, we can forecast the future. If we can model the rules of the storm, we can create a probabilistic forecast of where other raindrops are likely to fall.
To do this, econometricians have built a toolkit of fundamental “pure” processes that act as building blocks. Almost any complex economic time series can be described as a combination of these simple, foundational models.
The building blocks: Key types of stochastic processes
If we want to model a complex process, we start with the simplest possible components. In time series, these components describe how the value of our variable today (let’s call it Y_t) relates to its own past values or to past random shocks. The four main types are White Noise, Autoregressive (AR), Moving Average (MA), and their combination, ARMA.
The foundation: White noise
A White Noise process is the simplest and most boring process imaginable, but it’s also the most important. Think of it as pure, unpredictable randomness. A series is white noise if each value is a random draw from a distribution with a mean of zero, a constant variance, and zero correlation with all other values in the series.
In simple terms, a white noise series is one where the past tells you nothing about the future. Knowing the value today gives you no statistical edge in predicting the value tomorrow. Itโs the “static” you hear on an old radio. In modeling, this is our ideal endpoint. After we build a sophisticated model (say, for GDP growth), we want the leftover errors-the part our model can’t explain-to be nothing but white noise. If there are patterns left in the errors, it means our model missed something.
The “memory” model: Autoregressive (AR) process
An Autoregressive (AR) process is one where the current value of the series depends on its own past values, plus a white noise error term. The name says it all: it’s a regression of the series on itself (auto-regression).
A simple AR(1) model, the most common type, is written as: Y_t = c + ฯ * Y_t-1 + ฮต_t
Letโs translate that:
- Y_t is the value today.
- c is a constant (the baseline level).
- Y_t-1 is the value from the previous period (e.g., yesterday).
- ฯ (phi) is a coefficient that measures how much “memory” the series has. It shows how strongly yesterday’s value influences today’s value.
- ฮต_t (epsilon) is the white noise shock for today-the new, unpredictable information.
If ฯ is close to 1, the series has a strong “memory”; today will be very similar to yesterday. Think of a large container ship: its position today is almost entirely dependent on its position a minute ago. If ฯ is 0, the series is just white noise. Many economic series, like inflation or interest rates, show this kind of inertia and are often modeled using AR processes.
The “shock” model: Moving average (MA) process
A Moving Average (MA) process is a bit different. Instead of having memory of its past values, it has memory of its past random shocks (the error terms). The current value of the series is a function of the current shock and past shocks. Be careful: this is not the same as a “moving average” calculation used for smoothing data!
A simple MA(1) model is written as: Y_t = ฮผ + ฮต_t + ฮธ * ฮต_t-1
Letโs translate this one:
- Y_t is the value today.
- ฮผ (mu) is the mean or average of the series.
- ฮต_t is the random shock for today.
- ฮต_t-1 is the random shock from yesterday.
- ฮธ (theta) is a coefficient that measures how much of yesterday’s shock carries over to today.
What does this mean? Imagine an unexpected economic event, like a surprise announcement by the Reserve Bank of India (RBI). This is a shock (ฮต). An MA process describes how this single, unexpected event continues to ripple through the economy. The MA(1) model says the economy today (Y_t) is affected by today’s new shock (ฮต_t) and also by the lingering effects of yesterday’s shock (ฮต_t-1). Think of it as an aftershock. An MA process is good for modeling events that have a sudden, short-lived impact.
The hybrid model: ARMA process
Naturally, many real-world processes have both kinds of memory. They have inertia (like an AR process) and are also affected by lingering shocks (like an MA process). By combining them, we get the Autoregressive Moving Average (ARMA) model. This model explains today’s value using both its own past values and the past error terms, giving us a flexible and powerful tool for modeling complex dynamics.
The all-important question: Is your data stationary?
Before we can use any of these powerful models, we have to ask the most important question in time series analysis: is the process stationary? This single property changes everything.
In simple terms, stationarity means that a process’s statistical properties are constant over time. A stationary series will look roughly the same whenever or wherever you sample it. More technically, we usually look for weak stationarity (or covariance stationarity), which requires three conditions:
- Constant Mean: The average value of the series is the same for all time periods. It doesn’t drift up or down.
- Constant Variance: The “spread” or volatility of the series is the same over time. It doesn’t get wildly more erratic in some periods and calmer in others.
- Constant Covariance: The relationship between the value at one time (Y_t) and another time (Y_t-k) depends only on the lag k, not on the time t itself. For example, the correlation between January’s and February’s data should be the same as the correlation between August’s and September’s.
A white noise process is the perfect example of a stationary series. In contrast, a non-stationary series is one whose properties change. The most common type of non-stationarity is having a trend (a changing mean) or heteroskedasticity (a changing variance).
Why non-stationarity is a huge problem
Most of classical statistics and regression analysis is built on the assumption of stationarity. We assume we are drawing samples from a distribution with stable parameters. But if the process is non-stationary, the parameters themselves are changing. The mean in 1990 is different from the mean in 2020. This breaks all our standard statistical tools.
Consider a series like India’s nominal GDP or its population over the last 50 years. Both have a clear upward trend. This is a non-stationary process called a “random walk with drift.” The mean is clearly not constant. If we try to use this data in a regression, we fall into one of the most dangerous traps in econometrics.
The great deception: Beware of spurious regression
This brings us to the problem of spurious regression. This is a statistical illusion where you find a strong, statistically “significant” relationship between two variables that are, in reality, completely unrelated.
This happens when you run a regression using two or more non-stationary time series. If both series are trending (even for different reasons), the regression model will simply pick up on the fact that they are both moving in the same general direction. It will mistakenly conclude that one is “explaining” the other.
A classic, silly example: you might find a very high R-squared and a highly significant t-statistic if you regress the number of births in India against the total coffee production in Brazil over the last 30 years. Both series have been trending upwards. Does this mean Brazilian coffee causes babies in India? Of course not. The relationship is spurious, or “fake.” The regression is just capturing the shared underlying trend (general global economic and population growth) and nothing more.
How spurious regression invalidates your results
When you run a spurious regression, you get a set of results that look fantastic on paper but are complete nonsense.
- High R-squared: The model will report that your independent variable “explains” a large portion of the variation in your dependent variable.
- Significant t-statistics: The model will tell you (with a very low p-value) that the relationship is statistically significant and not due to chance.
- Low Durbin-Watson statistic: This is the classic red flag. The model’s errors will be highly autocorrelated, meaning this period’s error is strongly correlated with last period’s error. This is a sign that the model has failed to capture the true underlying dynamics.
Using a spurious regression for forecasting or policy advice is disastrous. You would be making decisions based on a relationship that doesn’t actually exist. The “fix” is to first make the data stationary before modeling it. This is often done by differencing the data-that is, instead of modeling the level of GDP, we model the change in GDP from one period to the next. This change is often stationary, and modeling it allows us to uncover the true, non-spurious relationships between variables.
Putting it all together
So, what’s the takeaway? When you look at a time series chart, don’t just see a line. See it as one piece of evidence-a single realization-from a deeper, unobservable stochastic process. Your first job as an analyst is to characterize that process. Is it an AR process with memory of its past? An MA process that remembers past shocks? Most importantly, is it stationary? Answering these questions is the only way to move from simply describing the past to building models that can reliably forecast the future and avoid the trap of spurious relationships.
What do you think?
When you look at a time series like the stock market or your monthly sales data, do you find it more helpful to think of it as having “memory” (an AR process) or being driven by “shocks” (an MA process)? Can you think of another real-world example of what might be a spurious regression?
Leave a Reply