Ever wondered why, when faced with a simple choice-like buying an Indian-made car (Option 1) or an imported one (Option 0)-we can’t just use a regular linear regression model? You’re trying to predict a ‘yes’ or ‘no’ (a discrete choice), not a continuous value like income or sales. This is where the brilliant concept of the Latent Regression Approach comes into play, creating an elegant bridge between the rigid rules of econometric modeling and the messy reality of human decision-making, which is fundamentally rooted in microeconomic theory.
Table of Contents
The theoretical foundations in utility maximization
At its heart, the latent regression approach is simply a mathematical formalization of the most fundamental idea in microeconomics: people make decisions to maximize their utility (satisfaction or benefit). While we, as economists, only observe the final choice (Y=1 or Y=0), we know that decision is driven by a hidden, unobserved process.
Imagine an individual, say an entrepreneur in Bengaluru, deciding whether to invest in a startup (Y=1) or keep the money in a high-yield fixed deposit (Y=0). The decision rests on the unobservable utilities:
- $U_1$: Utility from investing in the startup (potential for massive growth, personal satisfaction, but high risk).
- $U_0$: Utility from keeping the money in a fixed deposit (guaranteed return, zero management stress, but low growth).
The individual will choose to invest (Y=1) if and only if $U_1 > U_0$. We don’t see $U_1$ or $U_0$ directly, but we can model their difference, which we’ll call Net Utility ($U^*$).
$$U^* = U_1 – U_0$$
The observed choice ($Y$) is then just a simple trigger:
- If $U^* > 0$, then $Y = 1$ (Invest).
- If $U^* \le 0$, then $Y = 0$ (Do not Invest).
This simple framework connects the discrete outcome $Y$ to a continuous, but unobserved, utility difference $U^*$. This Net Utility, $U^*$, is what we call the latent variable.
Breaking down the net utility equation
To make the latent utility $U^*$ measurable, we decompose it into two parts, much like in a standard regression:
$$U^* = \mathbf{X}\mathbf{\beta} + \epsilon$$
Where:
- $\mathbf{X}\mathbf{\beta}$ (The Observed Part): This is the systematic or deterministic component. It captures the influence of all the factors we *can* observe about the entrepreneur (their age, education, current income, market conditions in India, etc.) on their Net Utility. $\mathbf{\beta}$ is the vector of coefficients we want to estimate.
- $\epsilon$ (The Unobserved Part): This is the random or stochastic error term. It represents everything we *cannot* observe: the entrepreneur’s risk appetite, their gut feeling, the advice they got from a friend, or any other inherent traits that influence their decision but aren’t in our data.
The choice rule then becomes:
$$Y = 1 \text{ if } \mathbf{X}\mathbf{\beta} + \epsilon > 0$$
$$Y = 0 \text{ if } \mathbf{X}\mathbf{\beta} + \epsilon \le 0$$
Rearranging this allows us to define the probability of the event occurring:
$$P(Y=1) = P(\mathbf{X}\mathbf{\beta} + \epsilon > 0) = P(\epsilon > -\mathbf{X}\mathbf{\beta})$$
If we assume the error term $\epsilon$ is symmetrically distributed around zero (which is common), this simplifies to the central pillar of discrete choice modeling:
$$P(Y=1) = P(\epsilon \le \mathbf{X}\mathbf{\beta}) = F(\mathbf{X}\mathbf{\beta})$$
Here, $F$ is the cumulative distribution function (CDF) of the error term $\epsilon$. The choice of this CDF $F$ is what distinguishes the two most popular latent regression models: Logit and Probit.
The unobserved propensity variable $Y^*$
The variable $U^*$ from the previous section is often represented simply as $Y^*$, the Unobserved Propensity Variable. Think of $Y^*$ as the individual’s “readiness” or “inclination” to choose $Y=1$. It’s the continuous score of latent potential, even though the final observed outcome is discrete.
A classic example is a person’s decision to join the labor force. The observed outcome $Y$ is a binary variable (1 = In the Labor Force, 0 = Not in the Labor Force). The unobserved propensity $Y^*$ is the individual’s underlying inclination to work, which is influenced by factors like education, spousal income, and local employment rates.
- If the propensity score ($Y^*$) exceeds a certain threshold (often normalised to zero), the person chooses to work ($Y=1$).
- If it falls below the threshold, they do not work ($Y=0$).
This latent variable interpretation is incredibly powerful because it gives the model’s coefficients ($\mathbf{\beta}$) a much clearer, utility-based meaning. Unlike the confusing interpretation of coefficients in a simple Linear Probability Model (LPM), the $\mathbf{\beta}$ coefficients in a latent regression model tell us how a one-unit change in an observable factor $\mathbf{X}$ changes the unobserved net utility ($Y^*$) or the propensity to choose $Y=1$.
Identification and scale in latent models
The beauty of the latent regression model comes with an important caveat known as the identification problem, specifically concerning the scale of the latent variable. This is a critical step for translating the theoretical concept into a practically estimable model.
The need to fix the error variance
Recall the latent model: $Y^* = \mathbf{X}\mathbf{\beta} + \epsilon$.
The observed probability is $P(Y=1) = P(\mathbf{X}\mathbf{\beta} + \epsilon > 0)$, which depends on the mean and variance of the error term, $\epsilon$. Since $Y^*$ itself is unobserved, we can arbitrarily scale it without changing the final choice $Y$.
Consider multiplying the entire equation by a positive constant, $\lambda$:
$$\lambda Y^* = \lambda \mathbf{X}\mathbf{\beta} + \lambda \epsilon \Rightarrow Y^{} = \mathbf{X}\mathbf{\beta}^{} + \epsilon^{}$$
Since $\lambda$ is positive, the choice condition remains the same: $Y^{} > 0$ still implies $Y=1$. However, the new parameters ($\mathbf{\beta}^{**}$) and the new error variance (which is now $\lambda^2$ times the original variance) would be different, but they would yield the exact same observed choices. This means the model parameters ($\mathbf{\beta}$ and $\sigma^2_{\epsilon}$) are not uniquely determined-the model is not identified.
To overcome this, we must impose a normalization constraint. In latent regression models, this constraint is typically achieved by arbitrarily fixing the variance of the error term, $\sigma^2_{\epsilon}$. Since $\mathbf{\beta}$ and $\sigma^2_{\epsilon}$ are indistinguishably intertwined in determining the probability scale, fixing $\sigma^2_{\epsilon}$ essentially sets the scale for the unobserved propensity variable, allowing the model to uniquely identify the remaining parameters, $\mathbf{\beta}$.
- For the Probit Model: We assume the error $\epsilon$ follows a standard normal distribution, meaning we fix the variance at $\sigma^2_{\epsilon} = 1$. This defines the CDF $F(\cdot)$ as the standard normal CDF, $\Phi(\cdot)$.
- For the Logit Model: We assume the error $\epsilon$ follows a logistic distribution, which has a variance of $\sigma^2_{\epsilon} = \pi^2/3 \approx 3.29$. The fact that this variance is fixed (though not equal to one) is what identifies the model. This defines the CDF $F(\cdot)$ as the logistic CDF.
This crucial step of fixing the error variance is the reason Probit and Logit models, while mathematically different, tend to produce very similar predictions for the probability of the outcome $Y=1$.
Case study: Nakosteen and Zimmer’s migration model
To see the latent regression approach in action, let’s look at the influential work on internal migration by Nakosteen and Zimmer (1982), and later refinements. They viewed the decision to migrate as a classic utility maximization problem.
The core question: Why do some people migrate (Y=1) while others stay (Y=0)?
They hypothesized that an individual migrates if the expected lifetime net benefits of moving are positive. This net benefit is the latent variable, $Y^*$.
$$Y^* = \text{Net Benefit from Migration} = (\text{Benefits}) – (\text{Costs})$$
### The utility decomposition
1. Benefits: Primarily the expected wage differential (higher expected wages in the destination region versus the origin region). 2. Costs: Direct costs (moving expenses, selling property) and psychic costs (leaving family, friends, and familiar surroundings-a major factor in the Indian context of strong social ties). 3. The Latent Component: This is where the model shines. The unobserved error term, $\epsilon$, captures things like an individual’s innate entrepreneurial spirit, risk-taking propensity, or “adventurousness,” which aren’t measured in the census but strongly influence the migration decision.
The econometric challenge they tackled was self-selection. People who choose to migrate (Y=1) are not a random sample of the population; they are likely inherently more dynamic and entrepreneurial than non-migrants. If you only compare the post-migration wages of migrants to the wages of non-migrants, you might mistakenly attribute higher wages to the act of migration itself, when the migrant’s innate, unobserved traits ($\epsilon$) were the real cause.
### The impact
By using a latent regression framework (specifically, a type of selection model often estimated using Probit for the migration decision), Nakosteen and Zimmer could model the migration choice and the subsequent earnings *jointly*. This allowed them to control for the impact of the unobserved latent characteristics, leading to more accurate and reliable estimates of the true economic return (wage increase) from migration. This method has become the standard for analyzing any discrete economic choice where the characteristics of those who choose one option versus another are fundamentally different.
What do you think? Can you recall a time you made a major “yes/no” decision (like changing jobs or taking a loan)? What were the unobserved, ‘gut-feeling’ factors that played the role of the error term ($\epsilon$) in your personal latent regression?
References
- https://egyankosh.ac.in/bitstream/123456789/111172/1/Unit-1.pdf
- https://prepp.in/news/e-492-microeconomics-indian-economy-notes
- https://pmc.ncbi.nlm.nih.gov/articles/PMC2897159/
- https://www.econometricstutor.co.uk/logit-and-probit-models-definition-of-logit-and-probit-models
- https://egyankosh.ac.in/bitstream/123456789/23444/1/Unit-12.pdf
- https://www.researchgate.net/publication/4914909_Migration_and_self-selection_Measured_earnings_and_latent_characteristics
- https://www.econstor.eu/bitstream/10419/245236/1/10.1080-23322039.2019.1609155.pdf
Leave a Reply