Imagine trying to understand how different pieces of a puzzle fit together when some of those pieces are invisible. That’s essentially what researchers faced when they wanted to study complex relationships between variables in the early twentieth century. The story of Structural Equation Modeling is one of brilliant minds across different disciplines-from biology to psychology to statistics-each contributing a piece to solve this puzzle. Today, this powerful statistical technique helps economists, social scientists, and researchers worldwide understand intricate relationships in their data.

Table of Contents

The foundation: Pearson’s correlation coefficient

The journey begins in 1896, when British mathematician Karl Pearson published groundbreaking work on the correlation coefficient. Building upon earlier contributions by Francis Galton and Auguste Bravais, Pearson developed what we now call the product-moment correlation coefficient. This mathematical tool gave researchers their first reliable way to measure the strength of relationships between two variables.

Think of correlation as a way to answer questions like: “Do students who study more hours tend to score higher on exams?” Pearson’s coefficient, represented by the letter “r,” could tell you not just whether a relationship existed, but how strong it was. Values range from negative one to positive one, with zero indicating no relationship at all. This seemingly simple metric became the cornerstone for all future developments in analyzing relationships between variables.

What made Pearson’s work revolutionary was that it provided a standardized, mathematical approach to understanding associations. Before this, researchers relied heavily on subjective observations. Now they had a quantifiable index that could be calculated using the least squares criterion, enabling the creation of linear regression models to predict outcomes.

Spearman introduces factor analysis

Enter Charles Spearman, an English psychologist with an unconventional background. After serving fifteen years in the British Army, Spearman pursued his passion for psychology and made discoveries that would reshape the field. In 1904, he published his seminal work on what he called the “general factor” of intelligence, or the g factor.

Spearman noticed something intriguing: when children performed well in one subject at school, they typically performed well in others too. This pattern of correlations suggested that something underlying all these different abilities might exist. Using Pearson’s correlation coefficient as his primary tool, Spearman developed a statistical technique to identify items that correlated with each other, which he termed factor analysis.

The birth of the g factor theory

Spearman’s two-factor theory of intelligence proposed that cognitive performance could be explained by two types of factors: a general ability common to most tasks, and specific abilities unique to particular tests. His work demonstrated that when you measured various mental abilities, about forty to fifty percent of the variation in performance could be attributed to this single general factor.

This was more than just an academic exercise. Spearman’s factor analysis provided a method to look beyond surface-level measurements and identify latent variables-characteristics that couldn’t be directly observed but could be inferred from patterns in the data. Later researchers, including Louis Thurstone and Raymond Cattell, would expand on these ideas, but Spearman had laid the groundwork for understanding how multiple observed indicators could point to underlying constructs.

Wright’s path analysis brings causality into focus

While psychologists were exploring intelligence, a biologist named Sewall Wright was wrestling with a different challenge. Working at the United States Department of Agriculture around 1918, Wright needed to understand the complex causal relationships that determined traits like guinea pig coloration and plant transpiration rates.

Wright realized that correlation coefficients alone couldn’t answer his questions about cause and effect. He needed a way to model how variables influenced each other through multiple pathways. This led him to develop path analysis, a method that used correlation coefficients and regression analysis together to map out causal relationships among observed variables.

From guinea pigs to social sciences

Wright’s innovation was representing causal hypotheses through path diagrams-visual maps showing how one variable might influence another through direct and indirect routes. Each arrow in the diagram represented a causal pathway, with standardized coefficients indicating the strength of each connection. This visual and mathematical approach made it possible to decompose complex relationships into their component parts.

Interestingly, path analysis was slow to gain traction in biology, perhaps because it challenged the purely correlational thinking that dominated at the time. Some believe the famous phrase “correlation does not imply causation” may have originated with Wright himself. However, by the 1950s, econometricians rediscovered his work, and by the 1960s, sociologists embraced it enthusiastically. Path analysis became recognized as a form of simultaneous equation modeling, allowing researchers to estimate multiple interconnected relationships at once.

The JKW model: bringing it all together

By the late 1960s and early 1970s, researchers had these powerful but separate tools: correlation and regression from Pearson, factor analysis from Spearman, and path analysis from Wright. What if these techniques could be combined into a unified framework? This question led to the creation of modern Structural Equation Modeling.

Three researchers working independently-Karl Jöreskog, Ward Keesling, and David Wiley-developed remarkably similar models around the same time. Their combined contribution became known as the JKW model, which elegantly integrated path models with confirmatory factor models. This unified approach could handle both latent variables (like intelligence or satisfaction) and observed variables (like test scores or survey responses) within a single analytical framework.

The LISREL revolution

The real breakthrough came in 1973 when Karl Jöreskog released LISREL (Linear Structural Relations), the first software program specifically designed for structural equation modeling. Working at the Educational Testing Service in Princeton, New Jersey, Jöreskog partnered with Dag Sörbom to create software that made these complex calculations accessible to researchers.

Before LISREL, performing structural equation modeling required extensive manual calculations and deep statistical expertise. The software automated the computational heavy lifting, allowing researchers to focus on theory development and interpretation. It could estimate complex models using maximum likelihood methods, test model fit, and provide detailed output about relationships among variables.

The impact was transformative. Researchers in psychology, sociology, economics, education, and marketing quickly adopted SEM as their go-to method for testing theoretical models with multiple constructs and complex relationships. Today, while many other SEM software packages exist, LISREL remains significant as the pioneer that made this sophisticated analysis practical for applied researchers.

Why this history matters today

Understanding the evolution of Structural Equation Modeling isn’t just an exercise in historical trivia. Each stage in its development addressed specific limitations and opened new possibilities for research. Pearson gave us the ability to measure associations. Spearman showed us how to uncover hidden factors underlying our measurements. Wright demonstrated how to think about and model causal pathways. And Jöreskog, Keesling, and Wiley integrated these insights into a comprehensive framework.

Today’s researchers using SEM stand on the shoulders of these giants. When you specify a measurement model to capture latent constructs, you’re applying Spearman’s insights. When you draw paths between variables in your structural model, you’re following Wright’s approach. And when you test whether your theoretical model fits the observed data, you’re using methods refined through decades of statistical development that began with Pearson’s correlation coefficient.

The beauty of SEM lies in its ability to test entire theoretical frameworks at once, rather than examining relationships piecemeal. It allows researchers to acknowledge and account for measurement error, test competing models, and estimate both direct and indirect effects. From understanding consumer behavior to testing economic theories to evaluating social programs, SEM has become an indispensable tool in the researcher’s toolkit.

What do you think? How might the integration of these different statistical traditions continue to evolve with modern computing power and new data sources? What kinds of questions in economics and social sciences might benefit most from understanding these complex, multivariate relationships?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.britannica.com/biography/Karl-Pearson
  2. https://en.wikipedia.org/wiki/Charles_Spearman
  3. https://explorable.com/spearman
  4. https://www.publichealth.columbia.edu/research/population-health-methods/path-analysis
  5. https://en.wikipedia.org/wiki/LISREL

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Research Methods in Economics

1 Research Methodology- Conceptual Foundation

  1. Research Methodology and its Constituents
  2. Theoretical Perspectives
  3. Approaches to Social Enquiry
  4. Research Strategies
  5. Research Process
  6. Hypothesis: Its Types and Sources
  7. The Nature, Sources and Types of Data
  8. Measurement Scales of Variables

2 Approaches to Scientific Knowledge- Positivism and Post Positivism

  1. Positivist Philosophy of Science
  2. Attack on Positivist Philosophy of Science
  3. Karl Popper’s Philosophy of Science
  4. Criticism against Karl Popper’s Philosophy of Science
  5. Thomas Kuhn’s Philosophy of Science
  6. Popper Versus Kuhn

3 Models of Scientific Explanation

  1. Unified View of Rules of Positivism
  2. Search for the Criterion of Cognitive Significance
  3. Rules of Logic or Rules of Correct Reasoning
  4. Hypothetico-Deductive Model
  5. Covering-Law Models
  6. Critical Appraisal of Covering-Law Models
  7. Explanation in Non-Physical Sciences

4 Debates on Models of Explanation in Economics

  1. Classical Political Economy and Ricardo’s Method
  2. Robbins, Positivism and Apriorism in Economics
  3. Hutchison and Logical Empiricism in Economics
  4. Milton Friedman and Instrumentalism in Economics
  5. Paul Samuelson and Operationalism
  6. Theory – Assumptions Debate in Economics: A Long View
  7. Amartya Sen on Heterogeneity of Explanation in Economics

5 Foundations of Qualitative Research- Interpretativism and Critical Theory Paradigm

  1. Interpretive Paradigm
  2. Critical Theory Paradigm
  3. Applications in Research: Illustrative Cases

6 Research Design and Mixed Methods Research

  1. Types of Research
  2. Research Design
  3. Research Design vs. Research Methods
  4. Research Methods
  5. The Rationale for Mixed Methods Research
  6. Forms of Mixed Methods Research Designs
  7. Case Studies of Mixed Methods Research Design

7 Data Collection and Sampling Design

  1. Method of Data Collection
  2. Tools of Data Collection
  3. Sampling Design
  4. Non-Random Sampling
  5. Random or Probability Sampling
  6. Methods of Random Sampling
  7. The Choice of an Appropriate Sampling Method

8 Measurement and Scaling Techniques

  1. Concept of Measurement
  2. Measurement Issues in Research
  3. Scales of Measurement
  4. Criteria for Good Measurement
  5. Errors in Measurements
  6. Scaling Techniques
  7. Comparative Scaling Techniques
  8. Non-Comparative Scaling Techniques

9 Two Variable Regression Models

  1. The Issue of Linearity
  2. The Non-deterministic Nature of Regression Model
  3. Population Regression Function
  4. Sample Regression Function
  5. Estimation of Sample Regression Function
  6. Goodness of Fit
  7. Functional Forms of Regression Model
  8. Classical Normal Regression Model
  9. Hypothesis Testing

10 Multivariable Regression Models

  1. Regression Model with Two Explanatory Variables
  2. Interpretation of Regression Coefficients
  3. Inclusion and Exclusion of Variables
  4. Generalisation to n-explainatory Variables
  5. Problem of Multi-co-linearity
  6. Problem of Hetero-scedasticity
  7. Problem of Autocorrelation
  8. Maximum Likelihood Estimations

11 Measures of Inequality

  1. Positive Measures
  2. Gini Index
  3. Lorenz Curve
  4. Normative Measures

12 Construction of Composite Index in Social Sciences

  1. Composite Index: The Concept
  2. Steps in Constructing Composite Index
  3. Dealing with Missing Values and Outliers
  4. Methods to Construct Composite Index
  5. Principal Component Analysis (PCA)
  6. Merits and Limitations of Composite Index

13 Multivariate Analysis- Factor Analysis

  1. Factor Analysis: Concept and Meaning
  2. Historical Background of Factor Analysis
  3. The Orthogonal Factor Model
  4. Communalities
  5. Methods of Estimation
  6. Factor Rotation
  7. Oblique Rotation
  8. Factor Scores
  9. Methods for Estimation of Factor Scores

14 Canonical Correlation Analysis

  1. Canonical Correlation Analysis (CCA): Concept and Meaning
  2. Assumptions of Canonical Correlation
  3. Canonical Correlation Analysis as Generalization of the Multiple Regression Analysis
  4. Steps and Procedure Involved in Computation of CCA Results
  5. Illustration of CCA
  6. Interpretation of CCA Results
  7. Limitations of Canonical Correlation

15 Cluster Analysis

  1. Cluster Analysis: Concept and Meaning
  2. Steps and Algorithm Involved in Cluster Analysis
  3. Methods of Cluster Analysis
  4. Partitioning Cluster Methods
  5. Hierarchical Cluster Methods
  6. Other Approaches: Two-step Cluster Analysis
  7. Interpretation of the Results

16 Correspondence Analysis

  1. Correspondence Analysis: Concept and Its Features
  2. Steps and Algorithm Involved in Correspondence Analysis Technique
  3. Basic Concepts and Definitions
  4. Reduction of Dimensionality
  5. Biplots
  6. Interpretation of the Results of Correspondence Analysis
  7. Multiple Correspondence Analysis

17 Structural Equation Modeling

  1. History of Structural Equation Modelling (SEM)
  2. Why do we Conduct Structural Equation Modelling?
  3. Assumptions of SEM
  4. Concepts and Terminology used in SEM
  5. SEM Models Specification
  6. Steps in SEM
  7. Software Programs for SEM
  8. Advantages and Disadvantages of SEM

18 Participatory Method

  1. What is Participatory Research?
  2. Methods of Participatory Research: Observation Method
  3. Focused Interview
  4. Oral Histories
  5. Life History
  6. Case Study Method
  7. Narratives
  8. Focus Group Discussion
  9. Grounded Theory
  10. Analysis of Qualitative Data
  11. Criticism of Participatory Methods
  12. Advantages of Participatory Research

19 Content Analysis

  1. Historical Background of Content Analysis
  2. Content Analysis: Concept and Meaning
  3. Terms Used in Content Analysis
  4. Approaches of Content Analysis
  5. Procedure Involved in Content Analysis
  6. Uses of Content Analysis
  7. Advantages and Disadvantages of Content Analysis

20 Action Research

  1. Historical Background of Action Research
  2. Definition of Action Research
  3. Principles of Action Research
  4. Characteristics of Action Research
  5. Models of Action Research
  6. Steps Involved in Action Research
  7. Advantages and Disadvantages of Action Research

21 Macro-Variable Data- National Income, Saving and Investment

  1. The Indian Statistical System
  2. National Income and Related Macro Economic Aggregates – System of National Accounts (SNA)
  3. National Income and Related Macro Economic Aggregates – Estimates of National Income and Related Macroeconomic Aggregates
  4. National Income and Related Macro Economic Aggregates – The Input-Output Table
  5. National Income and Related Macro Economic Aggregates – Regional Accounts – Estimates of State Income and Related Aggregates
  6. National Income and Related Macro Economic Aggregates – Regional Accounts – Estimates of Districts Income
  7. National Income and Related Macro Economic Aggregates – National Income and Levels of Living
  8. Saving
  9. Investment

22 Agricultural and Industrial Data

  1. Agricultural Data
  2. Industrial Data

23 Trade and Finance

  1. Trade
  2. Merchandise Trade
  3. Services Trade
  4. Finance
  5. Public Finances
  6. Currency, Coinage, Money and Banking
  7. Financial Markets

24 Social Sector

  1. Employment, Unemployment and Labour Force
  2. Education
  3. Health
  4. Shelter and Amenities
  5. Social Consequences of Development
  6. Environment
  7. Quality of Life