Ever wondered how economists predict whether a consumer will buy a product, a firm will default on a loan, or a voter will choose a particular candidate? These are all binary choices-outcomes that fall into one of two categories, like ‘Yes’ or ‘No’, ‘1’ or ‘0’. Standard linear regression, or Ordinary Least Squares (OLS), struggles here because its predicted values can easily fall outside the logical bounds of a probability (0 to 1). Enter the Probit and Logit models-the superheroes of binary choice modeling. They use a brilliant technique called Maximum Likelihood Estimation (MLE) to provide not just estimates, but predictions that respect the 0-1 probability range, giving us a far more accurate look into the mechanics of economic decision-making.
Table of Contents
- The maximum likelihood estimation (mle): the heart of probit and logit
- Building the likelihood function from bernoulli draws
- The iterative optimization process
- Interpreting coefficients and marginal effects
- Why coefficients aren’t marginal effects
- Calculating and interpreting the marginal effect
- Handling dummy variables and the odds ratio
- The marginal effect of a dummy variable
- The power of the logit odds ratio
- Hypothesis testing and goodness-of-fit
- Hypothesis testing: wald, lr, and lm tests
- Goodness-of-fit: Pseudo $r^2$ and mcfadden’s LRI
The maximum likelihood estimation (mle): the heart of probit and logit
Unlike OLS, which minimizes the sum of squared errors, Probit and Logit models are estimated using Maximum Likelihood Estimation. Think of MLE as a detective trying to find the best possible explanation for the evidence at hand. In our case, the “evidence” is the set of observed binary outcomes (e.g., who bought the product), and the “explanation” is the set of model parameters ($\beta$ coefficients) that makes the observed data most likely.
Building the likelihood function from bernoulli draws
The foundation of this process lies in the Bernoulli distribution, which governs a single trial with two outcomes (Success, $Y_i=1$, or Failure, $Y_i=0$). For each observation $i$, the probability of a ‘Success’ is $P_i$, and the probability of a ‘Failure’ is $1-P_i$. The key difference between Logit and Probit models is how they define this probability $P_i$ based on the explanatory variables, $\mathbf{X}_i\mathbf{\beta}$:
- Logit Model: Uses the Logistic Cumulative Distribution Function (CDF). $P_i = \Lambda(\mathbf{X}_i\mathbf{\beta}) = \frac{e^{\mathbf{X}_i\mathbf{\beta}}}{1+e^{\mathbf{X}_i\mathbf{\beta}}}$.
- Probit Model: Uses the Standard Normal CDF, denoted $\Phi(\mathbf{X}_i\mathbf{\beta})$. $P_i = \Phi(\mathbf{X}_i\mathbf{\beta})$.
For a sample of $N$ independent observations, the joint probability of observing the entire dataset is the product of the individual probabilities. This product forms the Likelihood Function, $L(\mathbf{\beta} | \mathbf{Y}, \mathbf{X})$.
Since working with products is mathematically tricky, we usually work with the Log-Likelihood Function, $\ln L(\mathbf{\beta})$. This log-likelihood is a sum, which is much easier to manage. The log-likelihood function for Probit/Logit models is often written as: $$ \ln L(\mathbf{\beta}) = \sum_{i=1}^{N} \left[ y_i \ln(P_i) + (1-y_i) \ln(1-P_i) \right] $$ Where $y_i$ is the observed outcome (0 or 1) and $P_i$ is the model’s predicted probability, dependent on $\mathbf{\beta}$ through the Logit or Probit function. The job of the MLE is to find the vector of coefficients, $\mathbf{\hat{\beta}}$, that maximizes this log-likelihood function. [Image: Log-Likelihood Function Optimization]
The iterative optimization process
Unlike OLS, where the optimal $\mathbf{\beta}$ can be found with a simple closed-form equation, the maximization of the log-likelihood function for Probit and Logit models requires numerical optimization. Since the log-likelihood function is concave, it ensures a unique global maximum.
The estimation process is iterative-itโs like a mountain climber finding the highest peak in the dark:
- Start: Choose an initial guess for the coefficient vector, $\mathbf{\beta}^{(0)}$ (often the OLS estimates).
- Climb: Compute the first and second derivatives of the log-likelihood function at the current guess (the score vector and the Hessian matrix). These tell the algorithm the slope and curvature of the ‘mountain’.
- Update: Use an algorithm (like the Newton-Raphson method) to calculate a new, better guess for $\mathbf{\beta}^{(1)}$, moving in the direction of the steepest ascent towards the maximum.
- Repeat: The process continues ($\mathbf{\beta}^{(2)}, \mathbf{\beta}^{(3)}, \dots$) until the change in the log-likelihood function between iterations is negligibly small. This point is the maximum, yielding the final MLE estimates, $\mathbf{\hat{\beta}}$.
Interpreting coefficients and marginal effects
This is where things diverge significantly from OLS. In a linear model, a coefficient $\beta_j$ tells you directly: “A one-unit increase in $X_j$ leads to a $\beta_j$-unit change in $Y$.” In Probit and Logit models, the coefficients themselves are less intuitive-they represent the change in the latent variable or the link function (log-odds in Logit, or the Probit index), not the change in the probability of interest, $P(Y=1)$.
Why coefficients aren’t marginal effects
Remember that the probability $P_i$ is connected to the linear index $\mathbf{X}_i\mathbf{\beta}$ through a non-linear S-shaped curve (the CDF, $\Phi$ or $\Lambda$). Because of this non-linearity, the effect of a one-unit change in an independent variable $X_j$ on the probability $P_i$ is not constant; it depends on the current value of $X_j$ and all other variables in the model. This is called heteroskedasticity inherent in the non-linear functional form.
For example, in a study on automobile purchase (a binary choice, $Y=1$), increasing a very poor person’s income by โน10,000 may have almost no effect on their probability of buying a car. Increasing a middle-income person’s income by the same amount, however, could drastically push them over the purchase threshold, leading to a large change in probability. The effect of income is clearly different across income levels.
Calculating and interpreting the marginal effect
To correctly gauge the impact of $X_j$ on the probability $P(Y=1)$, we must calculate the Marginal Effect (ME). The marginal effect for a continuous variable $X_j$ is the partial derivative of the probability function with respect to $X_j$:
$$ ME_j = \frac{\partial P(Y=1|\mathbf{X})}{\partial X_j} = f(\mathbf{X}\mathbf{\beta}) \cdot \beta_j $$
Here, $f(\cdot)$ is the probability density function (PDF) corresponding to the CDF ($\phi$ for Probit, $\lambda$ for Logit). Since $f(\cdot)$ is always positive, the sign of the marginal effect is determined by the sign of the coefficient, $\beta_j$.
The standard way to report the ME is as the Average Marginal Effect (AME), calculated by averaging the individual marginal effects over all observations in the sample:
$$ AME_j = \frac{1}{N} \sum_{i=1}^{N} \left[ f(\mathbf{X}_i\mathbf{\hat{\beta}}) \cdot \hat{\beta}_j \right] $$
The AME is interpreted as the average change in the probability of $Y=1$ for a one-unit change in $X_j$. For instance, an AME of 0.05 for ‘Age’ means that a one-year increase in age, on average, increases the probability of the outcome by 5 percentage points.
Handling dummy variables and the odds ratio
The marginal effect of a dummy variable
When an explanatory variable, say $D$, is a dummy variable (e.g., $D=1$ for ‘Female’, $D=0$ for ‘Male’), the marginal effect is a discrete change, not a derivative. We compute the change in predicted probability when the dummy variable flips from 0 to 1, holding all other variables constant:
$$ ME_D = P(Y=1 | D=1, \mathbf{X}_{\text{other}}) – P(Y=1 | D=0, \mathbf{X}_{\text{other}}) $$
Like the continuous case, this discrete change is typically calculated at a specific point (e.g., the mean of the other regressors) or, preferably, as the Average Discrete Change (ADC), which averages this difference over all observations. An ADC of -0.08 for the ‘Female’ dummy means being female is associated with an 8 percentage point lower probability of the outcome, everything else being equal.
The power of the logit odds ratio
The Logit model has a unique advantage in interpretation through the Odds Ratio. Recall that the logit model is essentially estimating the logarithm of the odds ratio, which is the ratio of the probability of success to the probability of failure: $$ \ln\left(\frac{P(Y=1)}{P(Y=0)}\right) = \mathbf{X}\mathbf{\beta} $$ The coefficient $\beta_j$ is the change in the log-odds for a one-unit change in $X_j$. While the log-odds is itself hard to interpret, by exponentiating the coefficient, we get the Odds Ratio (OR), $e^{\beta_j}$.
The odds ratio $e^{\beta_j}$ is interpreted as the factor by which the odds of the outcome ($Y=1$) change for a one-unit increase in $X_j$.
- If $e^{\beta_j} > 1$, the odds of $Y=1$ increase.
- If $e^{\beta_j} < 1$, the odds of $Y=1$ decrease.
- If $e^{\beta_j} = 1$, $X_j$ has no effect on the odds.
In medical and social sciences, the odds ratio provides an extremely relatable interpretation. For instance, if the OR for smoking on heart disease is 2.5, it means the odds of a smoker developing heart disease are 2.5 times the odds of a non-smoker, controlling for other factors. This feature often makes the Logit model preferred in disciplines where communicating odds is standard practice.
Hypothesis testing and goodness-of-fit
Hypothesis testing: wald, lr, and lm tests
To assess the statistical significance of coefficients in MLE, we rely on three powerful, asymptotically equivalent test statistics. These are crucial for comparing a restricted model (the null hypothesis, $H_0$) against an unrestricted model (the alternative hypothesis, $H_A$):
- Wald Test: This is the most common test, similar to the t-test in OLS, and it is available in almost all statistical software outputs. It tests whether a single coefficient or a set of coefficients is equal to zero. It uses the difference between the unrestricted estimate ($\mathbf{\hat{\beta}}$) and the restricted estimate (usually 0) and scales it by the inverse of the estimated covariance matrix.
- Likelihood Ratio (LR) Test: This test is based on the ratio of the maximized likelihood functions of the unrestricted model, $L_U$, and the restricted model, $L_R$. The test statistic is calculated as $\text{LR} = -2[\ln L_R – \ln L_U]$. A large LR statistic indicates that the unrestricted model significantly improves the fit over the restricted one, leading to the rejection of $H_0$.
- Lagrange Multiplier (LM) Test (or Score Test): This test only requires the estimation of the restricted model. It checks how far the slope (score) of the log-likelihood function is from zero when evaluated at the restricted coefficient estimates. If the slope is steep, the model’s likelihood could be significantly improved by moving away from the restricted estimates, leading to the rejection of $H_0$.
In practice, all three tests usually lead to the same conclusions in large samples, but the LR test is often considered more robust for model comparison.
Goodness-of-fit: Pseudo $r^2$ and mcfadden’s LRI
A major challenge in discrete choice modeling is that the standard OLS $R^2$ is inappropriate because the observed $Y$ is a binary 0/1 variable, not a continuous outcome. Instead, we use Pseudo $R^2$ measures to evaluate the model’s overall fit (or “goodness-of-fit”). These measures compare the fit of the estimated model to a baseline model-the “null” model-that only includes an intercept.
The most widely reported pseudo $R^2$ is McFadden’s Likelihood Ratio Index (LRI). It is calculated as:
$$ LRI = R^2_{\text{McFadden}} = 1 – \frac{\ln L_{\text{full}}}{\ln L_{\text{null}}} $$
Where $\ln L_{\text{full}}$ is the log-likelihood of the full model with all regressors, and $\ln L_{\text{null}}$ is the log-likelihood of the null model (intercept only). Since $\ln L_{\text{full}}$ is always greater than or equal to $\ln L_{\text{null}}$, and both are negative, the index ranges from 0 to a value less than 1 (sometimes called the “perfect fit” case, which is rarely achieved).
- A value near 0 indicates that the model is no better than the null model.
- A value closer to 1 suggests a very good fit. However, a McFadden’s LRI between 0.2 and 0.4 is often considered an excellent fit in social science and economics, where the complexity of human behavior makes perfect prediction difficult.
Understanding Probit and Logit models, along with the correct application of MLE and the interpretation of marginal effects, is the foundation for rigorous analysis in areas like labor economics, finance, and consumer behavior, guiding policy and business strategy, particularly in a complex and rapidly evolving market like India’s, where the organised retail sector alone is projected to nearly double in value by 2030, according to the India Brand Equity Foundation (IBEF).
What do you think? Given the complexity of calculating marginal effects, how might a policy-maker inadvertently misinterpret the results of a Probit or Logit model if they only look at the raw coefficients? Can you think of an economic decision, other than buying a product, that is clearly binary and where the marginal effect would be expected to change significantly depending on the initial conditions?
References
- https://www.statlect.com/fundamentals-of-statistics/probit-model-maximum-likelihood
- https://www.econometrics-with-r.org/11.3-estimation-and-inference-in-the-logit-and-probit-models.html
- https://www.researchgate.net/publication/323367939_Random_effects_probit_and_logit_understanding_predictions_and_marginal_effects
- https://clas.ucdenver.edu/marcelo-perraillon/content/marginal-effects-lisbon
- https://online.stat.psu.edu/stat504/lesson/2/2.4
- https://cran.r-project.org/package=margins/vignettes/TechnicalDetails.pdf
- https://ibef.org/news/india-s-retail-sector-projected-to-nearly-double-to-rs-1-68-00-650-crore-us-1-93-trillion-by-2030-deloitte-ficci
Leave a Reply