The Conditional Expectation Function
Expectation operator
For a continuous variable, we define the expectation the integral over all possible values of \(y\) and the probability density function (\(f_y\)).
\[ E[y_i] = \int t \cdot f_{y}(t)dt = \int t dF_{y}(t) \]
The letter \(t\) is used to denote the argument of the integral; i.e. all the values \(y\) can take on. The final equality also shows that we can define the \(E[]\) using the cumulative density function (\(F_y\)). This definition is some used in Microeconomic theory.
It is important to know that the \(E[y_i]\) is a non-random scalar. It is just a number; e.g. 5.623.
The use of integrals can seem a little confusing/complicated/daunting. If \(y\) takes on \(m\) discrete values - e.g., \(y\in\{\text{y}_1,\text{y}_2,...,\text{y}_m\}\) - then we can define the expectation in terms of the probability mass function (\(p_y\)).
\[ E[y_i] = \sum_{j=1}^{m} \text{y}_j \cdot p_{y}(\text{y}_j) \]
where \(p_{y}(\text{y}_j) = Pr(y_i = \text{y}_j)\).
Conditional Expectations
The expected value of \(y\) for a given value of variable \(x\) is defined in terms of the conditional expectation: \(E[y_i|x_i=\text{x}]\). Here, the value of \(x\) has been fixed at \(x_i=\text{x}\). The value of the expected value will therefore depend on the conditional distribution of \(y|x\).
\[ E[y_i|x_i = \text{x}] = \int t \cdot f_{y|x}(t|x_i = \text{x})dt = \int t dF_{y}(t|x_i = \text{x}) \]
Just like the unconditional expectation of \(y\), \(E[y_i]\), this value is a non-random scalar.
Conditional Expectation Function
The Conditional Expecation Function (CEF) - denoted \(E[y_i|x_i]\) - is a random function. It is a function that returns the conditional expectation of \(y\) for each value of \(x\). And since \(x\) is a random variable, the CEF is a random function.
We define the conditional expectation function as
\[ E[y_i|x_i] = \int t \cdot f_{y|x}(t|x_i)dt = \int t dF_{y|x}(t|x_i) \]
This should look very similarly; with the exception that \(x\) is not fixed. If we fix \(x_i = \text{x}\) (as above), then the value at which we are evaluating the function is no longer random. The result is the constant conditional expectation: \(E[y_i|x_i = \text{x}]\).
Case of Binary \(x\)
In the case where \(x\) is a binary random variable (i.e., dummy variable), the CEF is linear. We can write the CEF of \(y\) given \(x\) as,
\[ E[y_i|x_i] = E[y_i|x_i=0] + x_i\cdot\big(E[y_i|x_i=1]-E[y_i|x_i=0]\big) \]
The above function returns \(E[y_i|x_i=0]\) when \(x_i=0\) and \(E[y_i|x_i=1]\) when \(x_i=1\). This expression for the CEF is used Lecture 6 and related to the discussion of potential outcomes on page 50 of Wooldridge.
Law of Iterated Expectations
The Law of Iterated Expectations says that given two random variables1 \([y_i,x_i]\), we can express the unconditional expected value of \(y_i\) as the expected value of the conditional expectation of \(y_i\) on \(x_i\).
\[ E[y_i] = E\big[E[y_i|x_i]\big] \]
Where the outside expectation is with respect to \(x_i\),2 since the CEF is a random function of \(x_i\). We can expand this as follows,
\[ E[y_i] = \int t \cdot f_{y}(t)dt = \int\int t \cdot f_{y|x}(t|v)dtf_x(v)dv = E\big[E[y_i|x_i]\big] \]
Example 1 Suppose \(y\) and \(x\) are both discrete random variables: \(y_i\in\{1,2\}\) and \(x_i\in\{3,4\}\). With the joint distribution:
| \(x_i=3\) | \(x_i=2\) | |
| \(y_i=1\) | 1/10 | 3/10 |
| \(y_i=2\) | 2/10 | 4/10 |
We can then define the two marginal distributions,
| \(y_i=1\) | \(y_i=2\) |
| 4/10 | 6/10 |
and,
| \(x_i=3\) | \(x_i=4\) |
| 3/10 | 7/10 |
Likewise, we know the conditional distribution \(f_{y|x}\); which we get by dividing the joint distribution by the marginal distribution of \(x\). Each column of the conditional distribution should add up to 1.
| \(x_i=3\) | \(x_i=4\) | |
| \(y_i=1\) | 1/3 | 3/7 |
| \(y_i=2\) | 2/3 | 4/7 |
Now we can calculate the following objects:
- \(E[y_i]\)
\[ \begin{aligned} E[y_i] =& 1\cdot Pr(y_i=1)+2\cdot Pr(y_i=2) \\ =&1\cdot 4/10+2\cdot 6/10 \\ =&16/10 \end{aligned} \]
- \(E[y_i|x_i=3]\)
\[ \begin{aligned} E[y_i|x_i=3] =& 1\cdot Pr(y_i=1|x_i=3)+2\cdot Pr(y_i=2|x_i=3) \\ =&1\cdot 1/3+2\cdot 2/3 \\ =&5/3 \end{aligned} \]
- \(E[y_i|x_i=4]\)
\[ \begin{aligned} E[y_i|x_i=4] =& 1\cdot Pr(y_i=1|x_i=4)+2\cdot Pr(y_i=2|x_i=4) \\ =&1\cdot 3/7+2\cdot 4/7 \\ =&11/7 \end{aligned} \]
- \(E\big[E[y_i|x_i]\big]\)
\[ \begin{aligned} E\big[E[y_i|x_i]\big] =& E[y_i|x_i=3]\cdot Pr(x_i=3)+ E[y_i|x_i=4]\cdot Pr(x_i=4) \\ =&5/3\cdot3/10+11/7\cdot 7/10 \\ =&16/10 \end{aligned} \]
We have therefore demonstrated the law of iterated expectations.
We can extend this principle to conditional expectations. Suppose you have three random variables/vectors \(\{y_i,x_i,z_i\}\), we can express the conditional expected value of \(y_i\) on \(x_i\) as the (conditional) expected value of the conditional expectation of \(y_i\) on \(x_i\) and \(z_i\).
\[ E[y_i|x_i] = E\big[E[y_i|x_i,z_i]|x_i\big] \]
Here the outside expectation is with respect \(z_i\) conditional on \(x_i\). It utilizes the conditional distribution \(f_{z|x}\) to form the outside expectation,
\[ E[y_i|x_i] = \int t \cdot f_{y|x}(t|x_i)dt = \int\int y \cdot f_{y|x,z}(t|x_i,u)dtf_{z|x}(u|x_i)du = E\big[E[y_i|x_i,z_i]|x_i\big] \]
Properties of the CEF
The following three theorems can be found in a range of Econometrics textbooks and Microeconometrics texts.
Theorem 1 We can express the observed outcome \(y_i\) as a sum of \(E[y_i|x_i]+\varepsilon_i\) where \(E[\varepsilon_i|x_i]=0\) (i.e., mean independent).
Proof.
\(E[\varepsilon_i | x_i] = E[y_i - E[y_i | x_i] | x_i] = E[y_i | x_i] - E[y_i | x_i] = 0\)
\(E[h(x_i)\varepsilon_i] = E[h(x_i)E[\varepsilon_i | x_i]] = E[h(x_i) \times 0] = 0\)
Theorem 2 \(E[y_i|x_i]\) is the best predictor of \(y_i\).
Proof. \[ \begin{aligned} (y_i - m(x_i))^2 =& \left((y_i - E[y_i | x_i]) + (E[y_i | x_i] - m(x_i))\right)^2 \\ =& (y_i - E[y_i \| x_i])^2 + (E[y_i | x_i] - m(x_i))^2 \\&+ 2(y_i - E[y_i | x_i]) \times (E[y_i | x_i] - m(x_i)) \end{aligned} \]
The last term (cross product) is mean zero. Thus, the function is minimized by setting \(m(x_i) = E[y_i | x_i]\).
Theorem 3 [ANOVA Theorem] The variance of \(y_i\) can be decomposed as \(V(E[y_i|x_i])+E(V(y_i|x_i))\)
Proof. \[ \begin{aligned} V(y_i)=&V(E[y_i|x_i] + \varepsilon_i) \\ =&V(E[y_i|x_i])+V(\varepsilon_i) \\ =&V(E[y_i|x_i])+E[\varepsilon_i^2] \end{aligned} \] The second line follows from Theorem 1.1 (independence) and
\[ E[\varepsilon_i^2]=E\left[E[\varepsilon_i^2|x_i]\right]=E\left[V(y_i|x_i)\right] \]

