Imagine you’re trying to figure out how much money people earn after getting a specific degree, like a Master’s in Economics. You send out a survey, but only those who are currently employed and feel good about their salaries bother to reply. What happens to your research? Your estimate of the average salary will be much higher than the reality, right? This isn’t just a simple mistake; itโ€™s a deep, hidden problem in data analysis called sample selection bias, and itโ€™s a huge headache for economists and researchers everywhere. If you only look at the ‘survivors’ or the ‘responders,’ youโ€™re ignoring a vital part of the story-and that’s where the brilliant, Nobel Prize-winning work of economist James Heckman comes in.


Table of Contents

The lurking problem of endogenous censoring

The core issue we’re tackling is called endogenous censoring. In plain English, โ€˜censoringโ€™ means we donโ€™t observe the data we want for everyone in our study. If you’re studying wages, you only see the wages of people who are working; the wages of those who aren’t are โ€˜censored.โ€™ When this censoring is ‘endogenous,’ it means the decision to be included in the sample (e.g., choosing to work, choosing to answer the survey) is related to the very outcome you are trying to measure (e.g., the wage itself).

This bias doesn’t just happen randomly. It systematically creeps into data through various channels, all of which mean your observed sample is not a truly representative slice of the overall population. Understanding these sources is the first step toward correcting the data.

Key sources of selection bias

Selection bias arises from the choices individuals make, often driven by underlying economic or social factors. Here are some of the most common ways this hidden bias sneaks into economic data:

  • Non-response bias: This is the simplest and perhaps most common source. Think about a survey asking about household income. People with very low incomes might feel ashamed and not respond, while those with very high incomes might be too busy or too private. Similarly, people who are unemployed or in the informal sector may not be captured in official labour force surveys, leading to an overly optimistic view of formal employment rates.
  • Attrition or survivorship bias: This occurs when observing a process over time, and only the โ€˜survivorsโ€™ remain in the sample. If you study the performance of a group of startup businesses over five years, those that failed and dropped out are no longer in your data. If you only analyze the surviving firms, your conclusions about the โ€˜averageโ€™ startup’s success will be heavily skewed. This bias can be particularly relevant when studying longitudinal data in developing economies, like tracking beneficiaries of a government scheme, where people might move or records are lost.
  • Volunteer bias: Also known as self-selection. People who volunteer for a study, a medical trial, or even a training program are usually different-they might be more motivated, healthier, or have more time. If an economist studies the impact of a financial literacy course only on those who willingly signed up, the effect of the course itself might be overstated because the participants were already more financially aware or motivated than the general population.
  • Hawthorne effect: While less about sample inclusion, itโ€™s a related issue where the act of being observed changes behaviour. If workers know their productivity is being measured for a study, they might temporarily work harder, meaning your results don’t reflect their typical, everyday performance.

Introducing the Nobel solution: the Heckman model

For decades, economists struggled to correct this selection problem robustly. They knew the problem existed, but mathematically untangling the true economic effect from the bias was notoriously difficult. That changed with the work of James Heckman, who won the Nobel Memorial Prize in Economic Sciences in 2000 for his contributions to the analysis of selective samples.

The Heckman two-step procedure (or Heckman Correction) provides a rigorous econometric method for dealing with sample selection bias, particularly in cases of non-randomly missing data, such as the wage example where we only observe the wages of people who have chosen to work.

The two equations: selection and outcome

The model is built on the simple yet powerful idea that we can model the selection process itself. The Heckman model essentially uses two separate, but related, equations:

  1. The Selection Equation: This equation models the decision to be included in the sample. For our wage example, this would be a Probit model estimating the probability that a person is employed (i.e., their wage is observed). The factors here might include variables like education, age, marital status, and the presence of small children.
  2. The Outcome Equation: This is the equation we actually care about. It models the main variable of interest-the wage in our example-as a function of things like years of experience, education, and skills. This is the equation that will provide the unbiased estimate once the correction is applied.

The goal is to use the selection equation to generate a factor that quantifies the bias, and then use that factor as a control variable in the outcome equation. This mathematically separates the bias from the true effect you are trying to measure.


The crucial exclusion restriction

Hereโ€™s the part that is key to the model’s success, but often the trickiest part for researchers to implement: the exclusion restriction. This is the secret ingredient that allows the model to work its magic and cleanly separate the selection effect from the outcome effect.

The exclusion restriction requires that you include at least one variable in the selection equation (the probability of being in the sample) that is *not* included in the outcome equation (the variable you are studying). In technical terms, this variable must influence the likelihood of being observed but must have no direct causal effect on the outcome variable itself.

Finding the ‘identification instrument’

Think of it as finding a unique identifier for the bias. This special variable is often called the identification instrument or the Heckman instrument. Without it, the selection and outcome equations are mathematically too similar, making it impossible to disentangle the selection effect.

For the classic labour supply example (studying observed wages), a common instrument is non-labour income or the number of children under a certain age (like six years old). The reasoning goes like this: the presence of young children might strongly influence a parent’s decision to work (the selection decision) but should not, theoretically, have a direct impact on the wage they command *once they are working* (the outcome). Similarly, a spouseโ€™s income, or non-labour income like rental property earnings, makes it easier for someone to choose *not* to work, but once they *do* decide to work, that outside income doesn’t directly raise their hourly wage rate.


The two-step estimation method in practice

The beauty of the Heckman model, as opposed to complex single-step maximum likelihood estimation, is its relatively straightforward, two-step approach (Heckman, 1979). This method is what made the correction accessible to a generation of econometricians.

Step 1: estimating the selection equation and the Inverse Mills Ratio

In the first step, you run the Probit model (the selection equation) to predict the probability of inclusion in the sample (e.g., the probability of employment). Using the results of this model, you calculate a crucial term for every observation in your sample: the Inverse Mills Ratio (IMR), often denoted by the Greek letter $\lambda$ (lambda).

The IMR is a measure of the selection bias. Mathematically, it’s derived from the ratio of the probability density function (PDF) to the cumulative distribution function (CDF) of the error term in the selection equation. In plain English, the IMR quantifies the likelihood that an observation is *selected* into the sample, relative to the population. A high IMR for an observation means that the forces that caused that data point to be observed were unusual, suggesting a strong selection effect.

Step 2: correcting the outcome equation

The second step takes the original outcome equation (the OLS regression of the variable you care about) and adds the calculated Inverse Mills Ratio as a new, additional explanatory variable. The equation now looks something like this:

$$Y_i = \beta_0 + \beta_1 X_{i1} + \dots + \beta_k X_{ik} + \gamma \lambda_i + \epsilon_i$$

Where $Y_i$ is the outcome variable (wage), $X_{ik}$ are the original explanatory variables (experience, education), and $\lambda_i$ is the IMR. The $\gamma$ coefficient measures the extent of the selection bias. If $\gamma$ is statistically significant, it confirms that sample selection bias was present and the original estimates were indeed biased.

The coefficients $\beta_1$ through $\beta_k$ in this new, augmented regression are now consistent and unbiased estimates of the true effects, corrected for the sample selection problem. This two-step process is conceptually clear and robust enough to handle the complex non-randomness that plagues economic data.

While the two-step method is simpler to implement, it’s worth noting that it is not as statistically efficient as the single-step Maximum Likelihood Estimation (MLE) method, which estimates both equations simultaneously. However, due to its conceptual simplicity and robustness, the two-step method remains a cornerstone of applied econometrics.


Real-world impact: correcting for bias in policy

The Heckman correction isn’t just an academic exercise; it has profound implications for policymaking. If a government program aimed at increasing female labour force participation is evaluated, researchers must account for the fact that the women who *choose* to participate are likely different from those who don’t. Without the Heckman model, the reported success of the program could simply be an artifact of self-selection bias.

For instance, in analyzing the effectiveness of Indiaโ€™s various skill development programs, researchers must contend with volunteer bias-the most motivated people sign up. The Heckman model allows researchers to generate more accurate estimates of the true effect of the training itself, isolating it from the participant’s pre-existing motivation, thus helping policymakers design more effective, targeted interventions.

The modelโ€™s flexibility means itโ€™s used in diverse fields, from studying the impact of smoking on health outcomes (where the decision to smoke is non-random) to the effect of firm size on profitability (where only successful firms survive long enough to be studied).

What do you think? Can you think of a local context, perhaps in market research or a public health survey, where non-response bias might completely skew the key findings? What variables, relating only to the decision to respond, would you use as the exclusion restriction in that scenario?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://niti.gov.in/data-quality-challenges

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Advanced Econometric Methods

1 Discrete Dependent Variable Models

  1. Introduction
  2. Qualitative Choice Analysis
  3. The Regression Approach
  4. The Latent Regression Approach
  5. The Probit Model
  6. The Logit Model
  7. Estimation and Inference

2 Censored and Truncated Regression Models

  1. Characteristics of Qualitative Response Models
  2. Tobit Model
  3. Truncated Regression Model
  4. Sample Selection Model
  5. Models with Multiple Choices

3 Autoregressive (AR) Models

  1. Structure of AR Models
  2. Reasons for Inclusion of Lags in AR Models
  3. Use of Lag Operator in AR Models
  4. Inter-temporal Effect of Shocks in AR Models
  5. Relevance of AR Models to Economic Theory
  6. Yule-Walker Equations in AR Models
  7. Estimation of Parameters of AR Model
  8. Use of AR Models in Financial Economics

4 Distributed Lag Models

  1. Distributed Lag Models
  2. Koyck Model
  3. Autoregressive Models
  4. A More General Dynamic Model
  5. Jorgensonโ€™s Rational Lag Model
  6. Partial Adjustment Model
  7. Adaptive Expectations Model
  8. Interpretation of Coefficients
  9. Estimation and Inference

5 Estimation of System of Equations

  1. Seemingly Unrelated Regression Equations (SURE)
  2. Generalized Least Squares (GLS)
  3. Feasible Generalized Least Squares (FGLS)
  4. Maximum Likelihood Estimates
  5. Hypothesis Testing
  6. Treating Autocorrelation
  7. Interrelated Factor Demand

6 Introduction to Simultaneous Equations Models

  1. Simultaneous Equations Model (SEM)
  2. Structural Form and Reduced Form
  3. Identification Problem
  4. Order Condition
  5. Rank Condition
  6. General Structure of SEM
  7. Simultaneity Bias

7 Estimation of Simultaneous Equations Models

  1. Limited Information Systems
  2. Full Information Systems

8 Specification Issues of Time Series Data Models

  1. Stochastic Process
  2. Detection of Unit Root โ€“ Graphical Examination
  3. Detection of Unit Root โ€“ Statistical Tests
  4. The KPSS Test
  5. Test for Unit Root in the Presence of Structural Break
  6. Relations among Non-Stationary Series
  7. Limitations of Engle-Granger Test

9 Modelling Univariate Time Series

  1. Autoregressive Models
  2. Moving Average Models
  3. ARMA Models
  4. Integrated Processes and the ARIMA Models
  5. Box-Jenkins Methodology
  6. ARIMA Modelling in Software R

10 Vector Auto-Regression (VAR) Models

  1. Specification and Estimation of VAR
  2. Uses of VAR
  3. Innovation Accounting
  4. Vector Autoregression of Non-Stationary Data

11 Modelling Volatility

  1. The Autoregressive Conditional Heteroscedasticity (ARCH) Model
  2. Properties of the ARCH Model
  3. Test for ARCH Effects
  4. Generalized-ARCH (GARCH) Model
  5. Extensions of the GARCH Model

12 Introduction to Panel Data Models

  1. Introduction
  2. Panel Data Models
  3. Fixed Effects Model
  4. Random Effects Model
  5. Choice between Fixed Effects and Random Effects Models
  6. Hausman Test

13 Dynamic Panel Data Analysis

  1. Static Panel Data Model
  2. Specification of Dynamic Panel Data Model
  3. Estimation Methods of Dynamic Panel data Models
  4. Arellano-Bond Estimator
  5. System-GMM Method of Estimation
  6. Problems with the Arellano-Bond Approach
  7. Maximum Likelihood Estimator

14 Introduction to Generalised Method of Moments Estimation

  1. Need for Generalized Method of Moments
  2. Additional Moments Restrictions and Generalized Method of Moments
  3. Leading Example of GMM: IV Regression in Overidentified Models
  4. Variance Estimation and Optimal GMM
  5. Estimating Optimal GMM โ€“ Two-Step GMM Estimator
  6. Test of Overidentifying Restrictions