Ever wondered how economists predict whether a consumer will buy a product, a firm will default on a loan, or a voter will choose a particular candidate? These are all binary choices-outcomes that fall into one of two categories, like ‘Yes’ or ‘No’, ‘1’ or ‘0’. Standard linear regression, or Ordinary Least Squares (OLS), struggles here because its predicted values can easily fall outside the logical bounds of a probability (0 to 1). Enter the Probit and Logit models-the superheroes of binary choice modeling. They use a brilliant technique called Maximum Likelihood Estimation (MLE) to provide not just estimates, but predictions that respect the 0-1 probability range, giving us a far more accurate look into the mechanics of economic decision-making.

Table of Contents

The maximum likelihood estimation (mle): the heart of probit and logit

Unlike OLS, which minimizes the sum of squared errors, Probit and Logit models are estimated using Maximum Likelihood Estimation. Think of MLE as a detective trying to find the best possible explanation for the evidence at hand. In our case, the “evidence” is the set of observed binary outcomes (e.g., who bought the product), and the “explanation” is the set of model parameters ($\beta$ coefficients) that makes the observed data most likely.

Building the likelihood function from bernoulli draws

The foundation of this process lies in the Bernoulli distribution, which governs a single trial with two outcomes (Success, $Y_i=1$, or Failure, $Y_i=0$). For each observation $i$, the probability of a ‘Success’ is $P_i$, and the probability of a ‘Failure’ is $1-P_i$. The key difference between Logit and Probit models is how they define this probability $P_i$ based on the explanatory variables, $\mathbf{X}_i\mathbf{\beta}$:

  • Logit Model: Uses the Logistic Cumulative Distribution Function (CDF). $P_i = \Lambda(\mathbf{X}_i\mathbf{\beta}) = \frac{e^{\mathbf{X}_i\mathbf{\beta}}}{1+e^{\mathbf{X}_i\mathbf{\beta}}}$.
  • Probit Model: Uses the Standard Normal CDF, denoted $\Phi(\mathbf{X}_i\mathbf{\beta})$. $P_i = \Phi(\mathbf{X}_i\mathbf{\beta})$.

For a sample of $N$ independent observations, the joint probability of observing the entire dataset is the product of the individual probabilities. This product forms the Likelihood Function, $L(\mathbf{\beta} | \mathbf{Y}, \mathbf{X})$.

Since working with products is mathematically tricky, we usually work with the Log-Likelihood Function, $\ln L(\mathbf{\beta})$. This log-likelihood is a sum, which is much easier to manage. The log-likelihood function for Probit/Logit models is often written as: $$ \ln L(\mathbf{\beta}) = \sum_{i=1}^{N} \left[ y_i \ln(P_i) + (1-y_i) \ln(1-P_i) \right] $$ Where $y_i$ is the observed outcome (0 or 1) and $P_i$ is the model’s predicted probability, dependent on $\mathbf{\beta}$ through the Logit or Probit function. The job of the MLE is to find the vector of coefficients, $\mathbf{\hat{\beta}}$, that maximizes this log-likelihood function. [Image: Log-Likelihood Function Optimization]

The iterative optimization process

Unlike OLS, where the optimal $\mathbf{\beta}$ can be found with a simple closed-form equation, the maximization of the log-likelihood function for Probit and Logit models requires numerical optimization. Since the log-likelihood function is concave, it ensures a unique global maximum.

The estimation process is iterative-itโ€™s like a mountain climber finding the highest peak in the dark:

  1. Start: Choose an initial guess for the coefficient vector, $\mathbf{\beta}^{(0)}$ (often the OLS estimates).
  2. Climb: Compute the first and second derivatives of the log-likelihood function at the current guess (the score vector and the Hessian matrix). These tell the algorithm the slope and curvature of the ‘mountain’.
  3. Update: Use an algorithm (like the Newton-Raphson method) to calculate a new, better guess for $\mathbf{\beta}^{(1)}$, moving in the direction of the steepest ascent towards the maximum.
  4. Repeat: The process continues ($\mathbf{\beta}^{(2)}, \mathbf{\beta}^{(3)}, \dots$) until the change in the log-likelihood function between iterations is negligibly small. This point is the maximum, yielding the final MLE estimates, $\mathbf{\hat{\beta}}$.

Interpreting coefficients and marginal effects

This is where things diverge significantly from OLS. In a linear model, a coefficient $\beta_j$ tells you directly: “A one-unit increase in $X_j$ leads to a $\beta_j$-unit change in $Y$.” In Probit and Logit models, the coefficients themselves are less intuitive-they represent the change in the latent variable or the link function (log-odds in Logit, or the Probit index), not the change in the probability of interest, $P(Y=1)$.

Why coefficients aren’t marginal effects

Remember that the probability $P_i$ is connected to the linear index $\mathbf{X}_i\mathbf{\beta}$ through a non-linear S-shaped curve (the CDF, $\Phi$ or $\Lambda$). Because of this non-linearity, the effect of a one-unit change in an independent variable $X_j$ on the probability $P_i$ is not constant; it depends on the current value of $X_j$ and all other variables in the model. This is called heteroskedasticity inherent in the non-linear functional form.

For example, in a study on automobile purchase (a binary choice, $Y=1$), increasing a very poor person’s income by โ‚น10,000 may have almost no effect on their probability of buying a car. Increasing a middle-income person’s income by the same amount, however, could drastically push them over the purchase threshold, leading to a large change in probability. The effect of income is clearly different across income levels.

Calculating and interpreting the marginal effect

To correctly gauge the impact of $X_j$ on the probability $P(Y=1)$, we must calculate the Marginal Effect (ME). The marginal effect for a continuous variable $X_j$ is the partial derivative of the probability function with respect to $X_j$:

$$ ME_j = \frac{\partial P(Y=1|\mathbf{X})}{\partial X_j} = f(\mathbf{X}\mathbf{\beta}) \cdot \beta_j $$

Here, $f(\cdot)$ is the probability density function (PDF) corresponding to the CDF ($\phi$ for Probit, $\lambda$ for Logit). Since $f(\cdot)$ is always positive, the sign of the marginal effect is determined by the sign of the coefficient, $\beta_j$.

The standard way to report the ME is as the Average Marginal Effect (AME), calculated by averaging the individual marginal effects over all observations in the sample:

$$ AME_j = \frac{1}{N} \sum_{i=1}^{N} \left[ f(\mathbf{X}_i\mathbf{\hat{\beta}}) \cdot \hat{\beta}_j \right] $$

The AME is interpreted as the average change in the probability of $Y=1$ for a one-unit change in $X_j$. For instance, an AME of 0.05 for ‘Age’ means that a one-year increase in age, on average, increases the probability of the outcome by 5 percentage points.


Handling dummy variables and the odds ratio

The marginal effect of a dummy variable

When an explanatory variable, say $D$, is a dummy variable (e.g., $D=1$ for ‘Female’, $D=0$ for ‘Male’), the marginal effect is a discrete change, not a derivative. We compute the change in predicted probability when the dummy variable flips from 0 to 1, holding all other variables constant:

$$ ME_D = P(Y=1 | D=1, \mathbf{X}_{\text{other}}) – P(Y=1 | D=0, \mathbf{X}_{\text{other}}) $$

Like the continuous case, this discrete change is typically calculated at a specific point (e.g., the mean of the other regressors) or, preferably, as the Average Discrete Change (ADC), which averages this difference over all observations. An ADC of -0.08 for the ‘Female’ dummy means being female is associated with an 8 percentage point lower probability of the outcome, everything else being equal.

The power of the logit odds ratio

The Logit model has a unique advantage in interpretation through the Odds Ratio. Recall that the logit model is essentially estimating the logarithm of the odds ratio, which is the ratio of the probability of success to the probability of failure: $$ \ln\left(\frac{P(Y=1)}{P(Y=0)}\right) = \mathbf{X}\mathbf{\beta} $$ The coefficient $\beta_j$ is the change in the log-odds for a one-unit change in $X_j$. While the log-odds is itself hard to interpret, by exponentiating the coefficient, we get the Odds Ratio (OR), $e^{\beta_j}$.

The odds ratio $e^{\beta_j}$ is interpreted as the factor by which the odds of the outcome ($Y=1$) change for a one-unit increase in $X_j$.

  • If $e^{\beta_j} > 1$, the odds of $Y=1$ increase.
  • If $e^{\beta_j} < 1$, the odds of $Y=1$ decrease.
  • If $e^{\beta_j} = 1$, $X_j$ has no effect on the odds.

In medical and social sciences, the odds ratio provides an extremely relatable interpretation. For instance, if the OR for smoking on heart disease is 2.5, it means the odds of a smoker developing heart disease are 2.5 times the odds of a non-smoker, controlling for other factors. This feature often makes the Logit model preferred in disciplines where communicating odds is standard practice.


Hypothesis testing and goodness-of-fit

Hypothesis testing: wald, lr, and lm tests

To assess the statistical significance of coefficients in MLE, we rely on three powerful, asymptotically equivalent test statistics. These are crucial for comparing a restricted model (the null hypothesis, $H_0$) against an unrestricted model (the alternative hypothesis, $H_A$):

  1. Wald Test: This is the most common test, similar to the t-test in OLS, and it is available in almost all statistical software outputs. It tests whether a single coefficient or a set of coefficients is equal to zero. It uses the difference between the unrestricted estimate ($\mathbf{\hat{\beta}}$) and the restricted estimate (usually 0) and scales it by the inverse of the estimated covariance matrix.
  2. Likelihood Ratio (LR) Test: This test is based on the ratio of the maximized likelihood functions of the unrestricted model, $L_U$, and the restricted model, $L_R$. The test statistic is calculated as $\text{LR} = -2[\ln L_R – \ln L_U]$. A large LR statistic indicates that the unrestricted model significantly improves the fit over the restricted one, leading to the rejection of $H_0$.
  3. Lagrange Multiplier (LM) Test (or Score Test): This test only requires the estimation of the restricted model. It checks how far the slope (score) of the log-likelihood function is from zero when evaluated at the restricted coefficient estimates. If the slope is steep, the model’s likelihood could be significantly improved by moving away from the restricted estimates, leading to the rejection of $H_0$.

In practice, all three tests usually lead to the same conclusions in large samples, but the LR test is often considered more robust for model comparison.

Goodness-of-fit: Pseudo $r^2$ and mcfadden’s LRI

A major challenge in discrete choice modeling is that the standard OLS $R^2$ is inappropriate because the observed $Y$ is a binary 0/1 variable, not a continuous outcome. Instead, we use Pseudo $R^2$ measures to evaluate the model’s overall fit (or “goodness-of-fit”). These measures compare the fit of the estimated model to a baseline model-the “null” model-that only includes an intercept.

The most widely reported pseudo $R^2$ is McFadden’s Likelihood Ratio Index (LRI). It is calculated as:

$$ LRI = R^2_{\text{McFadden}} = 1 – \frac{\ln L_{\text{full}}}{\ln L_{\text{null}}} $$

Where $\ln L_{\text{full}}$ is the log-likelihood of the full model with all regressors, and $\ln L_{\text{null}}$ is the log-likelihood of the null model (intercept only). Since $\ln L_{\text{full}}$ is always greater than or equal to $\ln L_{\text{null}}$, and both are negative, the index ranges from 0 to a value less than 1 (sometimes called the “perfect fit” case, which is rarely achieved).

  • A value near 0 indicates that the model is no better than the null model.
  • A value closer to 1 suggests a very good fit. However, a McFadden’s LRI between 0.2 and 0.4 is often considered an excellent fit in social science and economics, where the complexity of human behavior makes perfect prediction difficult.

Understanding Probit and Logit models, along with the correct application of MLE and the interpretation of marginal effects, is the foundation for rigorous analysis in areas like labor economics, finance, and consumer behavior, guiding policy and business strategy, particularly in a complex and rapidly evolving market like India’s, where the organised retail sector alone is projected to nearly double in value by 2030, according to the India Brand Equity Foundation (IBEF).

What do you think? Given the complexity of calculating marginal effects, how might a policy-maker inadvertently misinterpret the results of a Probit or Logit model if they only look at the raw coefficients? Can you think of an economic decision, other than buying a product, that is clearly binary and where the marginal effect would be expected to change significantly depending on the initial conditions?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.statlect.com/fundamentals-of-statistics/probit-model-maximum-likelihood
  2. https://www.econometrics-with-r.org/11.3-estimation-and-inference-in-the-logit-and-probit-models.html
  3. https://www.researchgate.net/publication/323367939_Random_effects_probit_and_logit_understanding_predictions_and_marginal_effects
  4. https://clas.ucdenver.edu/marcelo-perraillon/content/marginal-effects-lisbon
  5. https://online.stat.psu.edu/stat504/lesson/2/2.4
  6. https://cran.r-project.org/package=margins/vignettes/TechnicalDetails.pdf
  7. https://ibef.org/news/india-s-retail-sector-projected-to-nearly-double-to-rs-1-68-00-650-crore-us-1-93-trillion-by-2030-deloitte-ficci

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Advanced Econometric Methods

1 Discrete Dependent Variable Models

  1. Introduction
  2. Qualitative Choice Analysis
  3. The Regression Approach
  4. The Latent Regression Approach
  5. The Probit Model
  6. The Logit Model
  7. Estimation and Inference

2 Censored and Truncated Regression Models

  1. Characteristics of Qualitative Response Models
  2. Tobit Model
  3. Truncated Regression Model
  4. Sample Selection Model
  5. Models with Multiple Choices

3 Autoregressive (AR) Models

  1. Structure of AR Models
  2. Reasons for Inclusion of Lags in AR Models
  3. Use of Lag Operator in AR Models
  4. Inter-temporal Effect of Shocks in AR Models
  5. Relevance of AR Models to Economic Theory
  6. Yule-Walker Equations in AR Models
  7. Estimation of Parameters of AR Model
  8. Use of AR Models in Financial Economics

4 Distributed Lag Models

  1. Distributed Lag Models
  2. Koyck Model
  3. Autoregressive Models
  4. A More General Dynamic Model
  5. Jorgensonโ€™s Rational Lag Model
  6. Partial Adjustment Model
  7. Adaptive Expectations Model
  8. Interpretation of Coefficients
  9. Estimation and Inference

5 Estimation of System of Equations

  1. Seemingly Unrelated Regression Equations (SURE)
  2. Generalized Least Squares (GLS)
  3. Feasible Generalized Least Squares (FGLS)
  4. Maximum Likelihood Estimates
  5. Hypothesis Testing
  6. Treating Autocorrelation
  7. Interrelated Factor Demand

6 Introduction to Simultaneous Equations Models

  1. Simultaneous Equations Model (SEM)
  2. Structural Form and Reduced Form
  3. Identification Problem
  4. Order Condition
  5. Rank Condition
  6. General Structure of SEM
  7. Simultaneity Bias

7 Estimation of Simultaneous Equations Models

  1. Limited Information Systems
  2. Full Information Systems

8 Specification Issues of Time Series Data Models

  1. Stochastic Process
  2. Detection of Unit Root โ€“ Graphical Examination
  3. Detection of Unit Root โ€“ Statistical Tests
  4. The KPSS Test
  5. Test for Unit Root in the Presence of Structural Break
  6. Relations among Non-Stationary Series
  7. Limitations of Engle-Granger Test

9 Modelling Univariate Time Series

  1. Autoregressive Models
  2. Moving Average Models
  3. ARMA Models
  4. Integrated Processes and the ARIMA Models
  5. Box-Jenkins Methodology
  6. ARIMA Modelling in Software R

10 Vector Auto-Regression (VAR) Models

  1. Specification and Estimation of VAR
  2. Uses of VAR
  3. Innovation Accounting
  4. Vector Autoregression of Non-Stationary Data

11 Modelling Volatility

  1. The Autoregressive Conditional Heteroscedasticity (ARCH) Model
  2. Properties of the ARCH Model
  3. Test for ARCH Effects
  4. Generalized-ARCH (GARCH) Model
  5. Extensions of the GARCH Model

12 Introduction to Panel Data Models

  1. Introduction
  2. Panel Data Models
  3. Fixed Effects Model
  4. Random Effects Model
  5. Choice between Fixed Effects and Random Effects Models
  6. Hausman Test

13 Dynamic Panel Data Analysis

  1. Static Panel Data Model
  2. Specification of Dynamic Panel Data Model
  3. Estimation Methods of Dynamic Panel data Models
  4. Arellano-Bond Estimator
  5. System-GMM Method of Estimation
  6. Problems with the Arellano-Bond Approach
  7. Maximum Likelihood Estimator

14 Introduction to Generalised Method of Moments Estimation

  1. Need for Generalized Method of Moments
  2. Additional Moments Restrictions and Generalized Method of Moments
  3. Leading Example of GMM: IV Regression in Overidentified Models
  4. Variance Estimation and Optimal GMM
  5. Estimating Optimal GMM โ€“ Two-Step GMM Estimator
  6. Test of Overidentifying Restrictions