We make โyes or noโ decisions every single day. Do you buy that cup of coffee? Do you take the bus or drive? Do you accept a new job offer? These are all binary choices-a simple 1 or 0, a “yes” or a “no.” In economics, we are obsessed with understanding these choices. If we can understand what drives them, we can build models to predict everything from consumer behavior to how people will respond to a new government policy. But modeling these choices is surprisingly tricky. Our first instinct, as data-driven people, might be to reach for the simplest tool in our toolkit: linear regression. But as we’ll soon discover, trying to fit a straight line to a “yes/no” world is fraught with problems. This is the story of the Linear Probability Model (LPM)-what it is, why itโs our first logical guess, and why it ultimately fails in spectacular fashion.
Table of Contents
- What makes us choose? Linking utility to probability
- Can we just use regular regression? Introducing the linear probability model (LPM)
- The five critical drawbacks of the LPM
- Problem 1: The nonsensical probability problem
- Problem 2: The case of the changing error (heteroscedasticity)
- Problem 3: The errors aren’t normal
- Problem 4: A straight line for a curvy relationship
- Problem 5: The mythical constant marginal effect
- What’s next? Moving beyond the straight line
What makes us choose? Linking utility to probability
Before we can model a choice, we need a theory for *why* people choose. In economics, that theory is built on one simple, elegant concept: utility. You can think of utility as a measure of satisfaction, happiness, or value an individual gets from a particular outcome. When faced with a set of options, the theory of rational choice posits that an individual will choose the option that maximizes their utility.
Let’s make this real. Imagine you’re deciding whether to buy a new smartphone. You have two choices: “Buy” (Y=1) or “Don’t Buy” (Y=0). You have a certain level of utility for buying the phone, let’s call it U_buy. This utility is based on factors like the phone’s features, its price, and how much you enjoy new gadgets. You also have a utility for *not* buying the phone, U_nobuy, which is based on the satisfaction you get from saving that money or using your current phone.
You will buy the phone if, and only if:
U_buy > U_nobuy
This is the decision rule. But here’s the catch: as researchers, we can’t perfectly observe your utility. It’s a latent, unobservable concept. We can, however, model it. We can say that your utility is based on observable factors (like your income, the phone’s price) and unobservable factors (like your personal mood that day, a sudden recommendation from a friend). We call these unobservable parts the “error” or “disturbance” term.
This is where probability enters the picture. The probability that you will buy the phone is simply the probability that the utility of buying is greater than the utility of not buying. Our goal in discrete choice modeling is to link the observable factors (income, price) to this *probability* of a “yes” outcome. This utility maximization framework is the solid, theoretical foundation for all discrete choice models, including the more advanced Logit and Probit models.
Can we just use regular regression? Introducing the linear probability model (LPM)
Given that we want to connect a set of explanatory variables (let’s call them X’s, like income, age, price) to an outcome (Y), the first tool we learn in statistics is Ordinary Least Squares (OLS) linear regression. The equation is beautifully simple:
Y_i = β0 + β1X1i + β2X2i + ... + εi
The Linear Probability Model (LPM) is born from the idea of just… trying this. We have a binary outcome, so let’s just make Y a dummy variable. We set Y=1 if the event happens (e.g., the person buys the phone) and Y=0 if it doesn’t. Our X variables are our explanatory factors: X_1 could be income, X_2 could be the person’s age, and so on.
When we do this, something magical seems to happen. The expected value of a binary (0/1) variable is simply the probability that the variable equals 1. In mathematical terms, E(Y_i | X_i) = P(Y_i=1 | X_i). This means our standard regression equation can be re-interpreted:
P(Y_i=1 | X) = β0 + β1X1i + β2X2i + ...
This is incredibly appealing! The model is linear, so it’s easy to estimate and, even better, it’s easy to interpret. The coefficient β1 is simply the change in the *probability* of Y happening for a one-unit change in X_1. For example, if β1 (for income, in thousands) is 0.05, it means that for every additional $1,000 of income, the probability of the person buying the phone increases by 0.05, or 5 percentage points. Itโs simple, itโs elegant, and itโs unfortunately deeply flawed.
The five critical drawbacks of the LPM
The simplicity of the LPM is a siren’s call, luring us onto the rocks of bad statistical inference. While it’s built on a simple regression, it violates several of the fundamental assumptions of that same model, leading to a host of problems. Let’s break down the five most critical failures.
Problem 1: The nonsensical probability problem
This is the most glaring and intuitive flaw. A probability, by its very definition, must lie between 0 and 1. It is a number that represents a chance, from 0% (impossible) to 100% (certain). The LPM, however, has no idea what a probability is. It is a simple linear equation, a straight line. And straight lines, by their nature, go on forever in both directions.
[Image: Graph of LPM predictions showing line below 0 and above 1]
Let’s go back to our homeownership-on-income example. Imagine our estimated model is: P(Owns_Home) = -0.1 + 0.01 * Income (where income is in thousands of dollars). What happens if we plug in a low income, like $5,000? The model predicts: P(Owns_Home) = -0.1 + 0.01 * 5 = -0.05 A negative 5% probability of owning a home. This is meaningless. Now, what about a high income, like $120,000? P(Owns_Home) = -0.1 + 0.01 * 120 = 1.1 A 110% probability of owning a home. This is equally absurd. Because the model is a straight line, for any dataset with a wide enough range of X values, it will *always* end up predicting probabilities that are nonsensical. This isn’t just a small quirk; it’s a fundamental failure to represent the very thing we are trying to model.
Problem 2: The case of the changing error (heteroscedasticity)
This is the most serious *statistical* flaw, and it’s a bit more technical. One of the core assumptions of OLS linear regression is homoscedasticity. This is a fancy word that means the variance (or spread) of the error terms (εi) is constant across all levels of the X variables. The LPM violates this assumption by its very design.
Think about it. For any set of X values, the predicted Y is P_i. The actual Y can only be 0 or 1. Therefore, the error term, εi = Y_i - P_i, can only take on two possible values:
- If
Y_i = 1(the person buys the phone), the error is1 - P_i - If
Y_i = 0(the person doesn’t buy), the error is0 - P_i, or just-P_i
The variance of this two-point error term can be shown to be: Var(εi) = P_i * (1 - P_i) Notice what this says. The variance of the error is not constant. It *depends on P_i*. And since P_i is a function of our X variables, the variance of the error changes with X. This is the definition of heteroscedasticity. The variance is smallest when P_i is near 0 or 1, and it’s largest in the middle when P_i is 0.5.
Why does this matter? When we have heteroscedasticity, our OLS coefficient estimates (the β‘s) are still unbiased, but they are no longer “efficient” (they aren’t the best-guess estimates we can make). More importantly, the formulas our software uses to calculate the standard errors of those coefficients are wrong. This means our t-statistics, p-values, and confidence intervals are all incorrect. We might conclude a variable is statistically significant when it isn’t, or vice-versa. The entire foundation of our hypothesis testing crumbles.
Problem 3: The errors aren’t normal
Another key assumption of OLS, especially for our hypothesis tests to be valid in small samples, is that the error terms are normally distributed (they follow a bell curve). As we just established, the error term in an LPM can only take on two possible values. A variable that can only be one of two things is the very opposite of a normally distributed variable, which can take on an infinite number of values. This violation means that our t-tests and F-tests are not strictly valid. Now, thanks to the Central Limit Theorem, this problem becomes less of a concern in very large samples, but it’s still a theoretical strike against the model.
Problem 4: A straight line for a curvy relationship
This flaw is related to the first one, but it’s more about the model’s fundamental shape. Does it make sense that the relationship between an X variable (like income) and a probability is a straight line? Probably not. Think about the probability of owning a high-end luxury car.
- For people with very low incomes, the probability is near 0, and a small $5,000 raise (from $20,000 to $25,000) will likely do almost nothing to that probability.
- For people with very high incomes, the probability is already near 1. A $5,000 raise (from $500,000 to $505,000) will also do almost nothing.
- But for someone in the “middle,” a $5,000 raise (from $150,000 to $155,000) might significantly increase the probability, as it pushes them over a threshold of affordability.
The true relationship isn’t a straight line; it’s an S-shaped curve. It starts flat near 0, gets steep in the middle, and then flattens out again near 1. The LPM, by imposing a linear, straight-line form, is a misspecification of the true functional form. It just doesn’t capture the reality of how probabilities change.
Problem 5: The mythical constant marginal effect
This problem is a direct consequence of the previous one. Because the LPM is a straight line, the coefficient β1 is constant. This is what we call the “marginal effect.” In our example, the model assumes that a $1,000 increase in income has the exact same impact on probability, regardless of whether you’re a student earning $10,000 a year or a CEO earning $2,000,000. As our luxury car example showed, this is simply not realistic. The marginal effect of income on the probability of buying a car should be small at the extremes and large in the middle. The LPM is incapable of capturing this dynamic and instead forces a single, constant effect, which is almost certainly wrong.
What’s next? Moving beyond the straight line
Given this pile-up of critical, theory-busting flaws, it’s clear the Linear Probability Model is not a good choice for serious analysis. Its simplicity is a trap. The good news is that these very failures point us directly toward a solution.
We need a model that:
- Forces its predictions to stay between 0 and 1.
- Allows for a non-linear, S-shaped relationship.
- Has marginal effects that can change depending on the values of X.
This is exactly what the Logit (or Logistic Regression) and Probit models are designed to do. Instead of assuming the probability itself is a linear function of X, these models assume that a “transformation” of the probability is linear. They use a non-linear function (called a link function) to map the straight line β0 + β1X from the real number line (from -∞ to +∞) into the (0, 1) probability space. This “S-curve” function is the key to fixing all of the LPM’s problems, creating models that are both theoretically sound and far more accurate in practice.
What do you think? Can you think of another “yes/no” decision from your own life that you’re sure follows an S-shaped curve relative to some factor, like price or income? Before reading this, would you have thought twice about using a simple line to model a binary choice?
References
- https://www.publichealth.columbia.edu/research/population-health-methods/discrete-choice-model-and-analysis
- https://en.wikipedia.org/wiki/Linear_probability_model
- https://statisticalhorizons.com/linear-vs-logistic-probability-models-which-is-better-and-when/
- https://statisticsbyjim.com/regression/heteroscedasticity-regression/
- https://www.sfu.ca/~dsignori/buec333/lecture%2018.pdf
Leave a Reply