Mastering Correlation Analysis & Causation
Explore how variables co-vary, distinguish true cause-and-effect from spurious relationships, and master statistical techniques including Scatter Plots, Karl Pearson's Coefficient (r), and Spearman's Rank Correlation (R) based on the official NIOS Economics curriculum.
1. Meaning of Correlation & Causation Fallacy
Definition of Correlation
Correlation refers to the statistical measure of association or mutual relationship between two or more variables. When an association exists, a change in the value of one variable is accompanied by a systematic change in the average value of another variable.
Correlation vs. Causation (Co-variation vs Cause)
A high degree of statistical correlation does NOT imply a cause-and-effect relationship. Correlation merely measures co-variation, not causation. When two variables appear strongly correlated due to a third underlying factor or sheer coincidence, it is called a Spurious Correlation.
Textbook Examples: Spurious Correlation & Third Variable Influence
| Observed Spurious Correlation | Third Variable / True Explanation |
|---|---|
| Ice Cream Sales & Beach Drowning Deaths | Summer Season / Temperature: Warm months increase both ice cream consumption and swimming activity. |
| Shoe Size & Reading Performance in Children | Age: Older children naturally have larger shoe sizes and higher reading skills. |
| Number of Doctors & Death Rates in Towns | Population Density: Dense regions require more doctors and experience more total deaths. |
| Number of Police Officers & Crime Rates | Population Size: Larger cities employ more police and have higher overall crime counts. |
| Imported Oranges & Road Accidents | Sheer Coincidence: Pure statistical coincidence without any logical connection. |
2. Types & Degrees of Correlation
Positive Correlation
Both variables move in the same direction (when X increases, Y increases; when X decreases, Y decreases).
- Height and Weight of children
- Advertising expenditure and Sales
- Price and Supply of goods
- Rainfall and Crop yield
Negative Correlation
Variables move in opposite directions (when X increases, Y decreases, and vice versa).
- Price and Demand for normal goods
- Volume and Pressure of gas (Boyle's Law)
- TV registrations and Cinema attendance
- Car age and Market resale value
Linear vs. Non-Linear (Curvilinear) Correlation
Degrees of Correlation Matrix (-1 to +1)
| Degree of Correlation | Positive Coefficient (r) | Negative Coefficient (r) |
|---|---|---|
| Perfect Correlation | +1.0 | -1.0 |
| High Degree | +0.75 to +1.0 | -0.75 to -1.0 |
| Moderate Degree | +0.25 to +0.75 | -0.25 to -0.75 |
| Low Degree | 0 to +0.25 | 0 to -0.25 |
| Absence of Correlation | Zero (0) | |
3. Scatter Plots & Karl Pearson's Coefficient (r)
A Scatter Plot (Dot Diagram Method)
A visual graphical representation where data pairs (X, Y) are plotted as points on a graph. The pattern of dots reveals the direction and strength of correlation without calculating numerical values.
B Karl Pearson's Coefficient of Correlation (r)
Gives a precise numerical measure of linear relationship between two quantitative variables.
Where x = X - X̄ and y = Y - Ȳ.
Where Cov(X, Y) = Σ(X - X̄)(Y - Ȳ) / N.
Essential Mathematical Properties of 'r':
- Limits: The value of r always lies between -1 and +1 (-1 ≤ r ≤ +1).
- Pure Number: r is a dimensionless coefficient, completely independent of units of measurement.
- Change of Origin & Scale: r is independent of both change of origin and change of scale. Adding, subtracting, multiplying, or dividing data by constants does NOT change r.
4. Spearman's Rank Correlation Coefficient (R)
Developed by Charles Edward Spearman, this method measures correlation based on the ranks assigned to observations rather than exact numerical values. Ideal for qualitative variables (e.g., beauty, honesty, intelligence, leadership).
Where D = R₁ - R₂ (difference in ranks) and ΣD = 0 always.
Where m is the number of times an item repeats. Tied items receive the average rank.