What you did yesterday almost certainly affects what you do today. This simple idea, known as “persistence” or “state dependence,” is everywhere in economics. A company’s investment last year influences its investment this year. A country’s GDP last quarter is a strong predictor of its GDP this quarter. A person’s spending habits last month shape their spending today. This makes perfect sense, but for economists trying to measure these relationships, this “dynamic” element creates a massive statistical headache: endogeneity.

When we try to study data that follows the same people, firms, or countries over time (known as panel data), and we include the past value of a variable to predict its current value, standard methods like Ordinary Least Squares (OLS) or Fixed Effects (FE) break down. They produce biased results. This is where a clever and powerful solution comes in: the Arellano-Bond estimator. Developed by econometricians Manuel Arellano and Stephen Bond in 1991, this estimator provides a way to get reliable answers from dynamic panel data. It’s a cornerstone of modern econometrics, and it works by combining two smart ideas: first-differencing and the Generalized Method of Moments (GMM).

Table of Contents

Understanding the core problem: Dynamics and endogeneity

Let’s imagine we want to understand what drives corporate investment. We collect data for 1,000 companies (our $N$, or cross-sectional units) over 10 years (our $T$, or time periods). We might model a firm’s current investment ($Investment_{it}$) based on its current profits ($Profits_{it}$) and, crucially, its investment from last year ($Investment_{it-1}$).

Our model looks something like this:

$Investment_{it} = \delta Investment_{it-1} + \beta Profits_{it} + \alpha_i + \epsilon_{it}$

Here, $\alpha_i$ is the “fixed effect.” This is a critical term. It captures all the unique, unobservable, and time-invariant characteristics of each firm. Think of it as Firm A’s “aggressive investment culture” or Firm B’s “naturally cautious management style.” This “style” doesn’t change year to year, and it affects *all* of that firm’s investment decisions.

Why standard methods fail: The ‘Nickell Bias’

Here’s the problem: a firm’s “investment culture” ($\alpha_i$) obviously affected its investment last year ($Investment_{it-1}$). This means the fixed effect ($\alpha_i$) is correlated with one of our explanatory variables ($Investment_{it-1}$). This correlation is the definition of endogeneity, and it poisons our estimates.

You might think, “I know how to solve that! I’ll use a Fixed Effects (FE) estimator!” An FE estimator works by subtracting the time-mean from every variable (a process called “de-meaning”), which perfectly removes the $\alpha_i$. Problem solved, right?

Unfortunately, no. When you “de-mean” the $Investment_{it-1}$ variable, you create a *new* correlation between it and the de-meaned error term. This specific problem is famously known as “Nickell bias,” (named after Stephen Nickell). This bias is severe when $T$ (the number of time periods) is small, which is exactly the kind of data we often have. So, both OLS and FE estimators are biased and inconsistent. We need a different approach.

The Arellano-Bond solution: A two-part strategy

Arellano and Bond proposed a brilliant escape route. Their strategy is to first transform the equation to remove the fixed effect, and then use a clever set of instrumental variables to solve the endogeneity problem that remains.

Step 1: First-differencing to remove fixed effects

Instead of “de-meaning,” the Arellano-Bond method starts by first-differencing the equation. This means subtracting last year’s equation from this year’s equation for each firm.

If $Investment_{it} = \delta Investment_{it-1} + \beta Profits_{it} + \alpha_i + \epsilon_{it}$

And $Investment_{it-1} = \delta Investment_{it-2} + \beta Profits_{it-1} + \alpha_i + \epsilon_{it-1}$

Subtracting the second from the first gives us:

$\Delta Investment_{it} = \delta \Delta Investment_{it-1} + \beta \Delta Profits_{it} + \Delta \epsilon_{it}$

Notice what happened: the fixed effect $\alpha_i$ is gone (since $\alpha_i – \alpha_i = 0$)! This is a huge victory. But, we’ve traded one problem for another.

Look at the new lagged variable, $\Delta Investment_{it-1}$ (which is $Investment_{it-1} – Investment_{it-2}$). Now look at the new error term, $\Delta \epsilon_{it}$ (which is $\epsilon_{it} – \epsilon_{it-1}$). Both of these new terms contain $\epsilon_{it-1}$. They are, by construction, correlated. We *still* have an endogeneity problem.

Step 2: Using the past as an instrument

This is the genius of the method. We need to find an instrumental variable (IV) for our “problem” variable, $\Delta Investment_{it-1}$. An instrument needs to satisfy two conditions:

  1. Relevance: It must be correlated with the problem variable ($\Delta Investment_{it-1}$).
  2. Exogeneity: It must be *uncorrelated* with the error term ($\Delta \epsilon_{it}$).

Where can we find such a variable? Arellano and Bond’s insight was to use lagged levels of the dependent variable.

Let’s think about this. For the equation at time $t=3$, our problem variable is $\Delta Investment_{i2}$ (which is $Investment_{i2} – Investment_{i1}$). Our error is $\Delta \epsilon_{i3}$ (which is $\epsilon_{i3} – \epsilon_{i2}$).

What about using $Investment_{i1}$ (investment in period 1) as an instrument?

  1. Is it relevant? Yes, $Investment_{i1}$ is clearly correlated with $\Delta Investment_{i2}$ (since it’s literally part of the term).
  2. Is it exogenous? Is $Investment_{i1}$ correlated with the error $\Delta \epsilon_{i3}$? No! $Investment_{i1}$ was determined by things that happened in period 1 (like $\epsilon_{i1}$), but it has no reason to be correlated with new, unexpected shocks in periods 2 or 3 (i.e., $\epsilon_{i2}$ and $\epsilon_{i3}$).

It’s a valid instrument! And it gets better. For the equation at time $t=4$, the problem variable is $\Delta Investment_{i3}$. The valid instruments are now $Investment_{i2}$ *and* $Investment_{i1}$. As we move forward in time, the list of available, valid instruments grows. We end up with a large set of instruments for our differenced equation.

Enter the Generalized Method of Moments (GMM)

We now have an equation and a set of instruments. In fact, we have *more* instruments than we strictly need. This is called “overidentification.” We can’t use standard IV estimation. Instead, we use the Generalized Method of Moments (GMM).

What is GMM and what are ‘moment conditions’?

GMM is a powerful estimation framework. It’s based on “moment conditions.” A moment condition is simply the statistical statement of our exogeneity assumption: that our instruments are uncorrelated with the error term. In math, we say $E[Z_{it} \cdot \Delta \epsilon_{it}] = 0$, where $Z$ is our set of instruments.

In our sample, this average product will never be *exactly* zero due to random chance. GMM works by finding the parameter estimates (our $\delta$ and $\beta$) that make the sample moment conditions *as close to zero as possible*. It does this by minimizing a weighted average of all the squared moment conditions. This “stacking” of many instruments across many individuals is what gives the estimator its power, and it’s why it works so well for “small T, large N” panels.

The one-step and two-step GMM estimators

The prompt mentions a one-step and two-step process, which relates to the “weighting” GMM uses.

  • One-Step GMM: This is the first pass. It uses a simple, assumed weighting matrix to combine all the moment conditions. This gives us *consistent* estimates of our parameters.
  • Two-Step GMM: This step aims for more *efficiency*. It takes the residuals from the one-step estimation and uses them to build an *optimal* weighting matrix. This matrix gives more weight to the more informative instruments (those with less “noise”). It then re-runs the estimation using these new weights.

In theory, the two-step estimator is superior. However, in small samples, its standard errors can be biased downwards, making us overconfident in our results. To fix this, researchers often use a finite-sample correction developed by Windmeijer, which is now standard practice in most statistical software.

How do we know if the estimator worked?

The Arellano-Bond estimator isn’t a magic wand. Its validity rests on crucial assumptions, which we must test.

  1. Sargan/Hansen Test of Overidentifying Restrictions: This tests the overall validity of our instruments. The null hypothesis is “all our instruments are valid (i.e., uncorrelated with the errors).” In this case, we *want* to fail to reject the null hypothesis. A high p-value is good news, suggesting our instruments are clean. A low p-value is bad news, indicating that some of our lagged variables are not valid instruments.
  2. Arellano-Bond Test for Serial Correlation: This test is more subtle. The entire method assumes that the *original* errors ($\epsilon_{it}$) are not serially correlated. This implies that the *differenced* errors ($\Delta \epsilon_{it}$) *must* be serially correlated at order one (AR(1)), because $\Delta \epsilon_{it}$ and $\Delta \epsilon_{it-1}$ both share the $\epsilon_{it-1}$ term. However, there should be *no* serial correlation at order two (AR(2)), because $\Delta \epsilon_{it}$ and $\Delta \epsilon_{it-2}$ share no common terms.

Therefore, when we run the test, we *want* to see a significant p-value for AR(1) (confirming our setup) and an insignificant p-value for AR(2) (confirming our underlying assumption). If the AR(2) test is significant, our assumptions are violated, and the estimates are not reliable.

In conclusion, the Arellano-Bond estimator is a vital tool. It allows researchers to navigate the tricky problem of endogeneity in dynamic panels by combining first-differencing to remove fixed effects with a GMM procedure that cleverly uses the model’s own past as instruments. By understanding how it works, we can more accurately measure how the past shapes the present.

What do you think? Can you think of another real-world economic question (besides corporate investment) where this kind of dynamic panel model would be essential? Given the complexity and the number of assumptions, what do you see as the biggest risk of misinterpreting results from an Arellano-Bond estimation?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.worldbank.org/en/research/brief/panel-data-analysis-nickell-bias
  2. https://www.stata.com/features/dynamic-panel-data-estimators/
  3. https://www.cemfi.es/~arellano/windmeijer.pdf

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Advanced Econometric Methods

1 Discrete Dependent Variable Models

  1. Introduction
  2. Qualitative Choice Analysis
  3. The Regression Approach
  4. The Latent Regression Approach
  5. The Probit Model
  6. The Logit Model
  7. Estimation and Inference

2 Censored and Truncated Regression Models

  1. Characteristics of Qualitative Response Models
  2. Tobit Model
  3. Truncated Regression Model
  4. Sample Selection Model
  5. Models with Multiple Choices

3 Autoregressive (AR) Models

  1. Structure of AR Models
  2. Reasons for Inclusion of Lags in AR Models
  3. Use of Lag Operator in AR Models
  4. Inter-temporal Effect of Shocks in AR Models
  5. Relevance of AR Models to Economic Theory
  6. Yule-Walker Equations in AR Models
  7. Estimation of Parameters of AR Model
  8. Use of AR Models in Financial Economics

4 Distributed Lag Models

  1. Distributed Lag Models
  2. Koyck Model
  3. Autoregressive Models
  4. A More General Dynamic Model
  5. Jorgensonโ€™s Rational Lag Model
  6. Partial Adjustment Model
  7. Adaptive Expectations Model
  8. Interpretation of Coefficients
  9. Estimation and Inference

5 Estimation of System of Equations

  1. Seemingly Unrelated Regression Equations (SURE)
  2. Generalized Least Squares (GLS)
  3. Feasible Generalized Least Squares (FGLS)
  4. Maximum Likelihood Estimates
  5. Hypothesis Testing
  6. Treating Autocorrelation
  7. Interrelated Factor Demand

6 Introduction to Simultaneous Equations Models

  1. Simultaneous Equations Model (SEM)
  2. Structural Form and Reduced Form
  3. Identification Problem
  4. Order Condition
  5. Rank Condition
  6. General Structure of SEM
  7. Simultaneity Bias

7 Estimation of Simultaneous Equations Models

  1. Limited Information Systems
  2. Full Information Systems

8 Specification Issues of Time Series Data Models

  1. Stochastic Process
  2. Detection of Unit Root โ€“ Graphical Examination
  3. Detection of Unit Root โ€“ Statistical Tests
  4. The KPSS Test
  5. Test for Unit Root in the Presence of Structural Break
  6. Relations among Non-Stationary Series
  7. Limitations of Engle-Granger Test

9 Modelling Univariate Time Series

  1. Autoregressive Models
  2. Moving Average Models
  3. ARMA Models
  4. Integrated Processes and the ARIMA Models
  5. Box-Jenkins Methodology
  6. ARIMA Modelling in Software R

10 Vector Auto-Regression (VAR) Models

  1. Specification and Estimation of VAR
  2. Uses of VAR
  3. Innovation Accounting
  4. Vector Autoregression of Non-Stationary Data

11 Modelling Volatility

  1. The Autoregressive Conditional Heteroscedasticity (ARCH) Model
  2. Properties of the ARCH Model
  3. Test for ARCH Effects
  4. Generalized-ARCH (GARCH) Model
  5. Extensions of the GARCH Model

12 Introduction to Panel Data Models

  1. Introduction
  2. Panel Data Models
  3. Fixed Effects Model
  4. Random Effects Model
  5. Choice between Fixed Effects and Random Effects Models
  6. Hausman Test

13 Dynamic Panel Data Analysis

  1. Static Panel Data Model
  2. Specification of Dynamic Panel Data Model
  3. Estimation Methods of Dynamic Panel data Models
  4. Arellano-Bond Estimator
  5. System-GMM Method of Estimation
  6. Problems with the Arellano-Bond Approach
  7. Maximum Likelihood Estimator

14 Introduction to Generalised Method of Moments Estimation

  1. Need for Generalized Method of Moments
  2. Additional Moments Restrictions and Generalized Method of Moments
  3. Leading Example of GMM: IV Regression in Overidentified Models
  4. Variance Estimation and Optimal GMM
  5. Estimating Optimal GMM โ€“ Two-Step GMM Estimator
  6. Test of Overidentifying Restrictions