Vector notation
When reading papers, you will frequently see the following notation:
y = \mathbf{x}\beta + u
What’s going on here? The variables have been collected as vectors:
y = \underbrace{\begin{bmatrix} 1 & x_{1} & x_{2} & \dots &x_{k}\end{bmatrix}}_{\mathbf{x}}\underbrace{\begin{bmatrix} \beta_0 \\ \beta_1 \\ \beta_2 \\ \vdots \\ \beta_{k} \end{bmatrix}}_\beta + u
When written in vector notation \mathbf{x} always precedes \beta, as \beta\mathbf{x}\neq \mathbf{x}\beta.
Approach #1: DGP
The model represents the true data generating process at the population level \Rightarrow
we can generate data in a manner akin to sampling from the population distribution.
e.g.
\begin{aligned}
&y_i = \beta_0 + \beta_1 x_{i} + u_i \\
\text{with}\;&\beta_0=0.5,\; \beta_1=0.6,\; x_{i}\sim U(0,1),\; u_i \sim N(0, 0.6^2)
\end{aligned}
When econometricians uses the word ‘population’, they are not referring to an enumerable collection of units, as in a census. They mean a probability distribution — the joint distribution over the random variables of interest, from which the observed sample is drawn. As we shall shortly see, the population parameters are defined in terms of the joint distribution of the random variables.
Approach #2: Prediction
Suppose you want to predict y using information from a set of variables \mathbf{x}=[x_1,x_2,\dots,x_k]. The ‘best’ predictor of y is the conditional expectation (function; CEF)
E[y | \mathbf{x}] = m(\mathbf{x})
Where the CEF is a function of \mathbf{x}.1
This is the starting point for Data Science.
- ‘Fancier’ methods like Lasso, Random Forests, and Neural Networks are simply newer methods to estimate m(\mathbf{x}). They have core advantages over linear regression models that relate specifically to ‘big’ data.
It turns out that you can write (see CEF Material),
y = m(\mathbf{x}) + u
where u has the nice property E[u| \mathbf{x}] = 0.
If you assume that m(\mathbf{x}) is linear (in parameters), then the best predictor of y is
E[y | \mathbf{x}] = \beta_0 + \beta_1x_{1} + \beta_2x_{2}+\dots + \beta_kx_{k}
and,
y = \beta_0 + \beta_1x_{1} + \beta_2x_{2}+\dots + \beta_kx_{k} + u
Population Regression Function
The function m(\mathbf{x}) is called the population regression function and \mathbf{x}\beta is referred to as the linear population regression function (or linear projection). It may be that m(\mathbf{x})=\mathbf{x}\beta.
Approach #3: Economics
The traditional approach to Econometrics is rooted in the formal economic modelling. A famous example is the Solow Growth model:
Step 1: Specify production function
Popular choice is Cobb-Douglas
Y_t = A_tK_t^\alpha L_t^{1-\alpha}
Step 2: Express in terms of per-worker variables:
y_t = A_tk_t^\alpha
where y_t=Y_t/L_t and k_t=K_t/L_t.
Step 3: Log linearize
\ln y_t = \ln A_t + \alpha \ln k_t
Now we have a linear model with a slope parameter \alpha - the parameter of interest.
BUT we have a problem:
A_t - which represents TFP - is not observed in the data
There is no error term! The model is deterministic.
Step 4: Make an assumption
Assume that TFP follows a particular time-series process (i.e. DGP):1
\ln A_t = \ln A_0 +gt+u_t
where gt is a deterministic trend and u_t a stochastic shock.
NOW we have a model of the form
\ln y_t = \ln A_0 + gt + \alpha \ln k_t+u_t
which we can rewrite as a linear regression model with 3 unknown parameters:
\ln y_t = \beta_0 + \beta_1t + \beta_2 \ln k_t + u_t
Approach #4: Causal Inference
A more modern approach, popularized by applied microeconomists, is not rooted in formal economic modelling. Instead, it focuses on a set of research designs (methodologies) that target the causal relationship between two variables. These are often taught as Microeconometrics.1
For example,
\text{Marginal Tax Rate} \rightarrow \text{Labour Supply}
A researcher might choose to estimate this relationship within a linear regression model framework, by specifying the estimating equation:
hrswrk_i = \beta_0 + \beta_1 \text{MTR}_i + u_i
The estimate of \beta_1 tells us about how a change in MTR affects hrswrk in the data.
BUT if we were to estimate this relationship in the data, do we really think it will tell us something causal? Consider:
Can we account for these differences in the model?
hrswrk_i = \beta_0 + \beta_1 MTR_i + \beta_2 income_i + \beta_3 yrsedu_i + u_i
This is where we get into the idea of adding additional regressors as “control” variables.
Classical Linear Regression Model
MLR #1
These assumptions have the abbreviation MLR for Multiple Linear Regression.
MLR-1: Linear in parameters
The population model is given by,
y = \beta_0 + \beta_1x_{1} + \beta_2x_{2}+\dots + \beta_kx_{k} + u
It may seem obvious, but it is nonetheless an assumption that the model is correctly specified as linear in parameters.
MLR #2
MLR-2: Random sampling
Random sample of n observations, \{\mathbf{x}_i,y_i\}:\; i=1,\dots,n, drawn from the population model
Random sampling is sometimes referred to as independently and identically distributed (iid).
- The ‘identically distributed’ part can be relaxed (i.e. error variance)
Together with MLR-1, it gives us
y_i = \beta_0 + \beta_1x_{i1} + \beta_2x_{i2}+\dots + \beta_kx_{ik} + u_i\qquad i=1,\dots,n
- Needed for Law of Large Numbers and Central Limit Theorem
- Sampling is the primary source of uncertainty the model.
MLR #3
MLR-3: No perfect collinearity
No exact linear relationship among regressors (including the constant).
This assumption is key for identification. It ensures that each parameter is uniquely identified.
Example
Consider a regression of hours-worked against temperature, measured in both F^\circ and C^\circ,
hrswrk_i = \beta_0 + \beta_1 tempC_i + \beta_2 tempF_i + u_i
The model has 3 parameters.
BUT there is a linear formula that links temperature in F^\circ\Rightarrow C^\circ: F^\circ = 9/5\cdot C^\circ+32.
- Plug this into the equation:
\begin{aligned}
hrswrk_i=& \beta_0 + \beta_1 tempC_i + \beta_2 (9/5 \cdot tempF_i + 32) + u_i \\
=& (\beta_0 + \beta_2 \cdot 32) + (\beta_1 + \beta_2 \cdot 9/5) tempC_i + u_i \\
=& \gamma_0 + \gamma_1 tempC_i + u_i \\
\end{aligned}
The model has 2 parameters now.
- The third was not identified because the variable tempF_i is perfectly collinear with tempC_i and the constant in the model.
MLR #4
MLR-4: Zero conditional mean
The error term has expected value of 0 given the any value of the explanatory variables
E[u | \mathbf{x}] = 0
With random sampling, then MLR-4 \Rightarrow E[u_i | \mathbf{x}_i] for i=1,\dots,n.
weaker than full independence: u\perp \mathbf{x};
but stronger than uncorrelatedness: Cov(\mathbf{x}, u) = 0
We can also show the following:
- MLR-4 implies that the unconditional mean of the error term is 0 (Law of Iterated Expectations; see CEF notes).
E[u | \mathbf{x}] = 0 \Rightarrow E[u] = 0
- MLR-4 implies the error term is uncorrelated with all regressors
E[u | \mathbf{x}] = 0 \Rightarrow E[ux_{j}]=Cov(u, x_{j}) = 0\quad j=1,2,\dots,k
MLR-4 implies that we can treat \mathbf{x} as fixed across repeated samples (even if they are not; i.e. also random variables).
- If x is fixed, there is only source of uncertainty: u.
- Allows us to treat x as non-random in derivations; e.g. variance of estimator.
MLR #5
MLR-5: Homoskedasticity
The variance of the error term is the same given any value of \mathbf{x}.
Var(u|\mathbf{x}) = \sigma^2
With random sampling, then MLR-5 \Rightarrow Var(u_i|\mathbf{x}_i) = \sigma^2 for i=1,\dots,n.
Needed to define the variance of the estimator
Needed for the famous BLUE property of OLS (Gauss Markov Theorem)
It can be relaxed; such errors are referred to as heteroskedastic
MLR #6
MLR-6: Normality
The conditional distribution of the error term is normal.
u|\mathbf{x}\sim N(0,\sigma^2)
Needed to know the finite distribution of the estimator
Can be relaxed, but then only asymptotic distribution is known
SLR model
Let’s consider the simplest case: a bivariate regression model
y_i = \beta_0 + \beta_1 x_i + u_i
Under MLR. 1-4 (or Wooldridge’s SLR. 1-4):
\begin{aligned}
E[y_i | x_i] =& E[\beta_0 + \beta_1 x_i + u_i | x_i] \\
=& E[\beta_0 | x_i] + E[\beta_1 x_i | x_i] + E[u_i | x_i] \\
=& \beta_0 + \beta_1 x_i + E[u_i | x_i] \\
=& \beta_0 + \beta_1 x_i
\end{aligned}
Identification
Recall, under MLR-4
- E[u_i]=0
- E[u_i x_i]=0
We can use these two moments to demonstrate the identification of \beta_0 and \beta_1.
Step 1: Substitute u_i = y_i - \beta_0 -\beta_1x_i into (1)
0 = E[u_i] = E[y_i - \beta_0 - \beta_1x_i]\Rightarrow \beta_0 = E[y_i]-\beta_1 E[x_i]
Step 2: Substitute u_i = y_i - \beta_0 -\beta_1x_i into (2)
0 = E[u_i x_i]=E[(y_i - \beta_0 -\beta_1x_i) x_i] = E[y_ix_i]-\beta_0 E[x_i]-\beta_1 E[x_i^2]
Step 3: Substitute the solution for \beta_0
0 = E[y_ix_i]-\beta_0E[x_i]-\beta_1E[x_i^2] = E[y_ix_i] - (E[y_i]-\beta_1 E[x_i])E[x_i] - \beta_1 E[x_i^2]
Manipulate and solve
0 = \underbrace{E[y_ix_i]-E[y_i]E[x_i]}_{Cov(y_i,x_i)}-\beta_1(\underbrace{E[x_i^2]-E[x_i]^2}_{Var(x_i)})
Arrive at the neat solution:
\beta_1 = \frac{Cov(y_i,x_i)}{Var(x_i)}
and,
\beta_0 = E[y_i] - \beta_1 E[x_i]
Both population parameters can be written as expressions of “observable” moments in the data. In this sense, they are identified.
Let’s plot E[y | x]=0.5 + 0.6 x, the (linear) population regression function.
Around the E[y_i|x_i], the error term is normaly distributed: u_i\sim N(0,0.6^2)
This is what it would look like if you sampled n=50 errors for x from that distribution.
![]()
For a given x, more of the observations will be concentrated around the conditional mean.
Interpretation: Continuous regressor
This then gives us a clear interpretation of \beta_0 and \beta_1
\begin{aligned}
\beta_0 =& E[y_i | x_i=0] \\
\beta_1 =& \frac{dE[y_i | x_i]}{dx_i}
\end{aligned}
Note, \beta_0 may not have a “realistic” interpretation.
- e.g., if y is level of exports (to US) and x is USD exchange rate of country i.
The interpretation of \beta_1 as a derivative assumes that x_i is a continuous variable.
Why not?
\beta_1 = \frac{dy_i}{dx_i}
- Without MLR. 1-4 (zero condition mean), we can’t rule out a relationship bewteen u and x
- e.g., u could include powers of x: u_i= \beta_2 x_i^2 + v_i
- With MLR. 1-4, E[y_i|x_i] = \beta_0 + \beta_1 x_i (i.e. the population regression function)
Interpretation: Binary regressor
What if x\in\{0,1\}?
- We cannot differentiate E[y_i|x_i].
Solution: difference
\begin{aligned}
E[y_i|x_i=\textcolor{blue}{0}] =& \beta_0 + \beta_1\times \textcolor{blue}{0} = \beta_0 \\
E[y_i|x_i=\textcolor{red}{1}] =& \beta_0 + \beta_1\times \textcolor{red}{1} = \beta_0 + \beta_1 \\
&\\
\Rightarrow \beta_1 =& E[y_i|x_i=\textcolor{red}{1}]-E[y_i|x_i=\textcolor{blue}{0}]
\end{aligned}
MLR model
In the multivariate case we have models of the form (k=3),
y_i = \beta_0 + \beta_1 x_{i1} + \beta_2x_{i2}+\beta_3 x_{i3} + u_i
As with the SLR case, under MLR 1-4:
E[y_i|\mathbf{x}_i] = \beta_0 + \beta_1 x_{i1} + \beta_2 x_{i2} + \beta_3 x_{i3}
Interpretation: Continuous regressor (linear)
Suppose, the linear population regression function is also linear in x.1
We can now think of each slope coefficient as
\beta_j = \frac{\partial E[y_i | \mathbf{x}_i]}{\partial x_{ij}} \quad j=1,2,3
This is a partial derivative.
Change in the conditional mean of y for a 1 unit change in regressor x_j, holding x_2 and x_3 fixed
Why additional regressors in a model are referred to as “control variables”; their presence changes the interpretation of the slope coefficient.
Differential
The mean of the outcome will change in the following manner with the regressors:
\Delta E[y_i|\mathbf{x}_i] = \frac{\partial E[y_i | \mathbf{x}_i]}{\partial x_{i1}} \Delta x_{i1} + \frac{\partial E[y_i | \mathbf{x}_i]}{\partial x_{i2}} \Delta x_{i2} + \frac{\partial E[y_i | \mathbf{x}_i]}{\partial x_{i3}} \Delta x_{i3}
So, if you ask what is the change in (expected value of y) from a 10-unit change in x_{1}:
\Delta E[y_i|\mathbf{x}_i] = \beta_1 \times 10
Interpretation: Continuous regressor (polynomial)
Suppose, the linear population regression function includes higher-order polynomials of some x’s.1
E[y_i|\mathbf{x}_i] = \beta_0 + \beta_1 \textcolor{red}{x_{i1}} + \beta_2 \textcolor{red}{x_{i1}^2} + \beta_3 x_{i3}
The derivative (w.r.t. x_1) now depends on the value of x_1:
\frac{\partial E[y_i | \mathbf{x}_i]}{\partial \textcolor{red}{x_{i1}}}= \beta_1 + 2\beta_2 x_{i1}
Interpretation: Continuous regressor (interactions)
Suppose, the linear population regression function includes higher-order polynomials of some x, but had some interactions.
E[y_i|\mathbf{x}_i] = \beta_0 + \beta_1 \textcolor{red}{x_{i1}} + \beta_2 \textcolor{red}{x_{i1}}\cdot x_{i3} + \beta_3 x_{i3}
The derivative (w.r.t. x_1) now depends on the value of x_3:
\frac{\partial E[y_i | \mathbf{x}_i]}{\partial \textcolor{red}{x_{i1}}}= \beta_1 + \beta_2 x_{i3}
Interpretation: binary variable
Suppose, that x_{1}\in\{0,1\}.
E[y_i|\textcolor{blue}{x_{i1}=0},x_{i2},x_{i3}] = \beta_0 + \textcolor{blue}{0} + \beta_2 x_{i2} + \beta_3 x_{i3}
and,
E[y_i|\textcolor{red}{x_{i1}=1},x_{i2},x_{i3}] = \beta_0 + \textcolor{red}{\beta_1} + \beta_2 x_{i2} + \beta_3 x_{i3}
Thus,
\beta_1 = E[y_i|\textcolor{red}{x_{i1}=1},x_{i2},x_{i3}]-E[y_i|\textcolor{blue}{x_{i1}=0},x_{i2},x_{i3}]
Interpretation: log-level
Suppose,
E[\textcolor{red}{\ln(y_i)}|\mathbf{x}_i] = \beta_0 + \beta_1 x_{i1} + \beta_2 x_{i2} + \beta_3 x_{i3}
\beta_j = \frac{\partial E[\textcolor{red}{\ln(y_i)} | \mathbf{x}_i]}{\partial x_{ij}}
The coefficient is therefore measured in log-units of y.
The relation to a change in the (expected) level of y is given by,
\%\Delta\; E[y_i|\mathbf{x}_i] = (exp(\beta_j)-1)\times 100
For reasonably small values of \beta_1 (i.e. within the range [-0.1,0.1]) this can be approximated by,
\%\Delta\; E[y_i|\mathbf{x}_i] \approx \beta_j\times 100
A 1-unit change in x_{ij} is associated with a \approx \beta_j\times 100 percentage change in the expected value of y.
This referred to as a semi-elasticity.
Interpretation: level-log
Suppose,
E[y_i|\mathbf{x}_i] = \beta_0 + \beta_1 \textcolor{blue}{\ln(x_{i1})} + \beta_2 x_{i2} + \beta_3 x_{i3}
Then,
\beta_1 = \frac{\partial E[y_i | \mathbf{x}_i]}{\partial \textcolor{blue}{\ln(x_{i1})}}
The coefficient is measured in y.
A 1 percent increase in x is exactly x\times1.01. This is equivalent to a change in \ln(x) of,
\ln(x_i\times1.01) - \ln(x_i) = \ln(1.01) \approx 0.01
Thus, a 1 percent increase in (the level of) x is associated with a \beta_1\times 0.01 = \beta_1/100 increase in the expected value of y. Or, more accurately
\Delta E[y_i|x_i] = \beta_1\times \ln(1.01)
This is also a semi-elasticity.
Interpretation: log-log
Suppose,
E[\textcolor{red}{\ln(y_i)}|\mathbf{x}_i] = \beta_0 + \beta_1 \textcolor{blue}{\ln(x_{i1})} + \beta_2 x_{i2} + \beta_3 x_{i3}
Then,
\beta_1 = \frac{\partial E[\textcolor{red}{\ln(y_i)}|\mathbf{x}_i]}{\partial \textcolor{blue}{\ln(x_{i1})}}
This is an elasticity measure: \beta_1 is the % change in the expected value of y from a 1 % change in x