In a perfect world, our economic models would be simple. Weโ€™d run a regression, say, of education on wages, and the resulting coefficient would tell us exactly how many more rupees youโ€™d earn for one extra year of schooling. But our world isn’t perfect. Itโ€™s messy, complex, and full of hidden connections. What if people who get more education are also naturally more ambitious? Suddenly, our simple regression is contaminated. Itโ€™s attributing the effect of ‘ambition’ to ‘education’, giving us a biased answer. This problem, known as endogeneity, is one of the biggest challenges in econometrics. To solve it, economists developed a clever tool: Instrumental Variables (IV) regression. But even that tool has its limits, leading us to a more powerful and flexible framework: the Generalized Method of Moments (GMM). It turns out, GMM isn’t just a new, complex technique; it’s the master framework that our familiar IV methods belong to.

Table of Contents

The core problem: When our variables have hidden baggage

Let’s stick with our example: estimating the effect of education ($X$) on wages ($Y$). The standard Ordinary Least Squares (OLS) method works only if a critical assumption holds: the education variable ($X$) must be uncorrelated with the ‘error term’ ($u$). The error term is a catch-all bucket for *everything else* that determines wages but isn’t in our model-like ability, family connections, or ambition. If ambitious people get more education, then $X$ and $u$ are correlated, and our OLS estimate is biased.

This is where an instrument comes in. An instrumental variable ($Z$) is a third variable that has a very special set of skills:

  • Relevance: The instrument ($Z$) must be correlated with our problematic variable, education ($X$). For example, maybe we find that the ‘distance to the nearest college’ affects how much education people get.
  • Exclusion: The instrument ($Z$) must be completely uncorrelated with the error term ($u$). This means distance to college should *only* affect wages *through* its effect on education, and not in any other hidden way. (We’d have to argue, for instance, that it’s not just capturing a rural vs. urban wage gap).

The instrument acts like a ‘clean’ channel. It isolates *only* the part of education that was ‘pushed’ by the instrument (e.g., the extra schooling people got *because* a college was nearby) and uses that clean variation to estimate the effect on wages. This is the essence of IV estimation. The core assumption is that our instrument ($Z$) is ‘exogenous’, meaning it’s not correlated with the error $u$. In population terms, this is written as $E[Z’u] = 0$.

The ‘exactly identified’ case: One tool for one job

The simplest IV scenario is what we call exactly identified. This happens when the number of instruments ($L$) is exactly equal to the number of endogenous variables ($K$). In our example, we have one problematic variable (education, $K=1$) and one instrument (distance to college, $L=1$).

Because we assume $E[Z’u] = 0$ in the population, we try to mimic this in our sample data. We replace the error $u$ with its formula from the model, $u = y – X\beta$. Our condition becomes $E[Z'(y – X\beta)] = 0$.

To estimate $\beta$ from our sample, we use the method of moments. We just take the sample version of that population condition and set it to zero:

$$ \frac{1}{n} \sum_{i=1}^{n} Z_i(y_i – X_i\beta) = 0 $$

Or, in matrix form, $Z'(y – X\beta) = 0$.

Since we have $L$ instruments and $K$ parameters to estimate, and $L=K$, we have exactly one equation for each unknown. We can solve this system of equations algebraically to find our one-and-only IV estimate for $\beta$. Itโ€™s neat, tidy, and gives us a single answer.

The ‘overidentified’ case: When you have too many tools

But what if we find *more* good instruments? Let’s say we have our ‘distance to college’ ($Z_1$). We also find that a ‘local tuition subsidy policy’ ($Z_2$) affected education but not wages directly. And maybe even the ‘quarter of birth’ ($Z_3$), famously used by economists, which affects school start dates. Now we have one endogenous variable (education, $K=1$) but three instruments ($L=3$).

This is an overidentified model. And it’s a fantastic problem to have! It means we have *more* information than we strictly need. But it also breaks our simple IV estimator.

Why? We now have three moment conditions we want to set to zero:

  1. $Z_1′(y – X\beta) = 0$
  2. $Z_2′(y – X\beta) = 0$
  3. $Z_3′(y – X\beta) = 0$

We have three equations, but still only one unknown parameter $\beta$. It’s almost certain that no single value of $\beta$ can make all three equations equal zero at the same time. We’re “overdetermined.” We can’t just pick one instrument and throw the others away; that would be incredibly wasteful. We need a way to combine the information from all three instruments in an optimal way.

GMM: The master framework for combining information

This is where the Generalized Method of Moments (GMM) comes in. GMM provides a comprehensive solution for exactly this problem. The philosophy of GMM is simple: If we can’t make all the sample moment conditions *exactly* zero, let’s find the parameter $\beta$ that makes them as close to zero as possible, all at the same time.

GMM defines a “criterion function” that measures the total “farness from zero” of all our moment conditions combined. We can write our $L$ sample moments as a vector $g(\beta) = Z'(y – X\beta)$. GMM then tries to minimize a quadratic form of this vector:

$$ J(\beta) = [g(\beta)]’ W [g(\beta)] $$

That is, $J(\beta) = (y – X\beta)’Z \cdot W \cdot Z'(y – X\beta)$

The GMM estimator is the $\beta$ that makes this $J(\beta)$ function as small as possible. But what is that new $W$ matrix in the middle? That is the weighting matrix. Itโ€™s the heart of GMM. The weighting matrix tells us how to *combine* the different moment conditions. It answers questions like: “How much should we ‘care’ about the ‘distance to college’ moment vs. the ‘tuition subsidy’ moment?”

A smart choice of $W$ will give more weight to instruments that provide more precise information and less weight to instruments that are noisy. This makes GMM an incredibly powerful and efficient estimator. It allows us to use *all* our instruments in a statistically optimal way. This flexibility is why researchers, such as those at the Reserve Bank of India, use GMM to model complex economic systems where simple assumptions often fail.

The big reveal: Two-Stage Least Squares (2SLS) is just GMM

So, what does this all have to do with our standard IV models? This is the most elegant part. It turns out that the most common method used to estimate overidentified models, Two-Stage Least Squares (2SLS), is actually just a special case of GMM.

Here’s how it works. The “optimal” GMM estimator-the one that is most efficient-uses a very specific weighting matrix $W$ that is the inverse of the covariance matrix of the moments. Calculating this can be complicated. However, if we make a big simplifying assumption-that our original model’s errors ($u$) are homoskedastic (meaning they have a constant variance)-then the optimal weighting matrix $W$ simplifies to be proportional to $(Z’Z)^{-1}$.

And what happens when you plug $W = (Z’Z)^{-1}$ into that big GMM criterion function and find the $\beta$ that minimizes it? The math simplifies perfectly, and the resulting estimator is *identical* to the 2SLS estimator.

You may know 2SLS by its two-step procedure:

  1. First Stage: Regress the “bad” endogenous variable $X$ on *all* the instruments $Z$ (e.g., $Education = \alpha_0 + \alpha_1 Z_1 + \alpha_2 Z_2 + \alpha_3 Z_3 + error$). Save the predicted values, $\hat{X}$. This $\hat{X}$ is the ‘clean’ part of $X$, stripped of its correlation with the error term.
  2. Second Stage: Regress the original outcome variable $Y$ on those predicted values $\hat{X}$ (e.g., $Wages = \beta_0 + \beta_1 \hat{X} + error$). The resulting $\beta_1$ is the 2SLS estimate.

This direct link shows that 2SLS is not a separate technique but rather a specific, intuitive application of the GMM principle. GMM is the “parent” concept. It provides the general theory for how to combine moment conditions, and 2SLS is what you get when you apply that theory to a linear IV model with the assumption of homoskedasticity. When that assumption fails (i.e., we have heteroskedasticity), we can use a “robust” GMM estimator that is more efficient than 2SLS.

So, while you may have learned IV and GMM as separate topics, they are deeply connected. GMM is the unifying theory that shows us how to handle everything from a simple, exactly-identified model to a complex, overidentified one, all by following a single principle: get as close to zero as you can.

What do you think? Given that “better” instruments (more relevant, more exogenous) are hard to find, what are the potential risks of using “weaker” instruments just to have an overidentified model? And does knowing that 2SLS is a form of GMM change how you think about its assumptions?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://online.stat.psu.edu/stat501/lesson/13/13.1
  2. https://ocw.mit.edu/courses/14-382-econometrics-spring-2017/resources/mit14_382s17_lec10/
  3. https://www.rbi.org.in/Scripts/PublicationsView.aspx?id=19916
  4. https://stats.stackexchange.com/questions/56608/what-is-the-relationship-between-gmm-and-2sls

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Advanced Econometric Methods

1 Discrete Dependent Variable Models

  1. Introduction
  2. Qualitative Choice Analysis
  3. The Regression Approach
  4. The Latent Regression Approach
  5. The Probit Model
  6. The Logit Model
  7. Estimation and Inference

2 Censored and Truncated Regression Models

  1. Characteristics of Qualitative Response Models
  2. Tobit Model
  3. Truncated Regression Model
  4. Sample Selection Model
  5. Models with Multiple Choices

3 Autoregressive (AR) Models

  1. Structure of AR Models
  2. Reasons for Inclusion of Lags in AR Models
  3. Use of Lag Operator in AR Models
  4. Inter-temporal Effect of Shocks in AR Models
  5. Relevance of AR Models to Economic Theory
  6. Yule-Walker Equations in AR Models
  7. Estimation of Parameters of AR Model
  8. Use of AR Models in Financial Economics

4 Distributed Lag Models

  1. Distributed Lag Models
  2. Koyck Model
  3. Autoregressive Models
  4. A More General Dynamic Model
  5. Jorgensonโ€™s Rational Lag Model
  6. Partial Adjustment Model
  7. Adaptive Expectations Model
  8. Interpretation of Coefficients
  9. Estimation and Inference

5 Estimation of System of Equations

  1. Seemingly Unrelated Regression Equations (SURE)
  2. Generalized Least Squares (GLS)
  3. Feasible Generalized Least Squares (FGLS)
  4. Maximum Likelihood Estimates
  5. Hypothesis Testing
  6. Treating Autocorrelation
  7. Interrelated Factor Demand

6 Introduction to Simultaneous Equations Models

  1. Simultaneous Equations Model (SEM)
  2. Structural Form and Reduced Form
  3. Identification Problem
  4. Order Condition
  5. Rank Condition
  6. General Structure of SEM
  7. Simultaneity Bias

7 Estimation of Simultaneous Equations Models

  1. Limited Information Systems
  2. Full Information Systems

8 Specification Issues of Time Series Data Models

  1. Stochastic Process
  2. Detection of Unit Root โ€“ Graphical Examination
  3. Detection of Unit Root โ€“ Statistical Tests
  4. The KPSS Test
  5. Test for Unit Root in the Presence of Structural Break
  6. Relations among Non-Stationary Series
  7. Limitations of Engle-Granger Test

9 Modelling Univariate Time Series

  1. Autoregressive Models
  2. Moving Average Models
  3. ARMA Models
  4. Integrated Processes and the ARIMA Models
  5. Box-Jenkins Methodology
  6. ARIMA Modelling in Software R

10 Vector Auto-Regression (VAR) Models

  1. Specification and Estimation of VAR
  2. Uses of VAR
  3. Innovation Accounting
  4. Vector Autoregression of Non-Stationary Data

11 Modelling Volatility

  1. The Autoregressive Conditional Heteroscedasticity (ARCH) Model
  2. Properties of the ARCH Model
  3. Test for ARCH Effects
  4. Generalized-ARCH (GARCH) Model
  5. Extensions of the GARCH Model

12 Introduction to Panel Data Models

  1. Introduction
  2. Panel Data Models
  3. Fixed Effects Model
  4. Random Effects Model
  5. Choice between Fixed Effects and Random Effects Models
  6. Hausman Test

13 Dynamic Panel Data Analysis

  1. Static Panel Data Model
  2. Specification of Dynamic Panel Data Model
  3. Estimation Methods of Dynamic Panel data Models
  4. Arellano-Bond Estimator
  5. System-GMM Method of Estimation
  6. Problems with the Arellano-Bond Approach
  7. Maximum Likelihood Estimator

14 Introduction to Generalised Method of Moments Estimation

  1. Need for Generalized Method of Moments
  2. Additional Moments Restrictions and Generalized Method of Moments
  3. Leading Example of GMM: IV Regression in Overidentified Models
  4. Variance Estimation and Optimal GMM
  5. Estimating Optimal GMM โ€“ Two-Step GMM Estimator
  6. Test of Overidentifying Restrictions