The Conditional Expectation Function

Expectation operator

For a continuous variable, we define the expectation the integral over all possible values of \(y\) and the probability density function (\(f_y\)).

\[ E[y_i] = \int t \cdot f_{y}(t)dt = \int t dF_{y}(t) \]

The letter \(t\) is used to denote the argument of the integral; i.e. all the values \(y\) can take on. The final equality also shows that we can define the \(E[]\) using the cumulative density function (\(F_y\)). This definition is some used in Microeconomic theory.

It is important to know that the \(E[y_i]\) is a non-random scalar. It is just a number; e.g. 5.623.

The use of integrals can seem a little confusing/complicated/daunting. If \(y\) takes on \(m\) discrete values - e.g., \(y\in\{\text{y}_1,\text{y}_2,...,\text{y}_m\}\) - then we can define the expectation in terms of the probability mass function (\(p_y\)).

\[ E[y_i] = \sum_{j=1}^{m} \text{y}_j \cdot p_{y}(\text{y}_j) \]

where \(p_{y}(\text{y}_j) = Pr(y_i = \text{y}_j)\).

Conditional Expectations

The expected value of \(y\) for a given value of variable \(x\) is defined in terms of the conditional expectation: \(E[y_i|x_i=\text{x}]\). Here, the value of \(x\) has been fixed at \(x_i=\text{x}\). The value of the expected value will therefore depend on the conditional distribution of \(y|x\).

\[ E[y_i|x_i = \text{x}] = \int t \cdot f_{y|x}(t|x_i = \text{x})dt = \int t dF_{y}(t|x_i = \text{x}) \]

Just like the unconditional expectation of \(y\), \(E[y_i]\), this value is a non-random scalar.

Conditional Expectation Function

The Conditional Expecation Function (CEF) - denoted \(E[y_i|x_i]\) - is a random function. It is a function that returns the conditional expectation of \(y\) for each value of \(x\). And since \(x\) is a random variable, the CEF is a random function.

We define the conditional expectation function as

\[ E[y_i|x_i] = \int t \cdot f_{y|x}(t|x_i)dt = \int t dF_{y|x}(t|x_i) \]

This should look very similarly; with the exception that \(x\) is not fixed. If we fix \(x_i = \text{x}\) (as above), then the value at which we are evaluating the function is no longer random. The result is the constant conditional expectation: \(E[y_i|x_i = \text{x}]\).

Case of Binary \(x\)

In the case where \(x\) is a binary random variable (i.e., dummy variable), the CEF is linear. We can write the CEF of \(y\) given \(x\) as,

\[ E[y_i|x_i] = E[y_i|x_i=0] + x_i\cdot\big(E[y_i|x_i=1]-E[y_i|x_i=0]\big) \]

The above function returns \(E[y_i|x_i=0]\) when \(x_i=0\) and \(E[y_i|x_i=1]\) when \(x_i=1\). This expression for the CEF is used Lecture 6 and related to the discussion of potential outcomes on page 50 of Wooldridge.

Law of Iterated Expectations

The Law of Iterated Expectations says that given two random variables1 \([y_i,x_i]\), we can express the unconditional expected value of \(y_i\) as the expected value of the conditional expectation of \(y_i\) on \(x_i\).

\[ E[y_i] = E\big[E[y_i|x_i]\big] \]

Where the outside expectation is with respect to \(x_i\),2 since the CEF is a random function of \(x_i\). We can expand this as follows,

\[ E[y_i] = \int t \cdot f_{y}(t)dt = \int\int t \cdot f_{y|x}(t|v)dtf_x(v)dv = E\big[E[y_i|x_i]\big] \]

Example 1 Suppose \(y\) and \(x\) are both discrete random variables: \(y_i\in\{1,2\}\) and \(x_i\in\{3,4\}\). With the joint distribution:

\(f_{y,x}\)
\(x_i=3\) \(x_i=2\)
\(y_i=1\) 1/10 3/10
\(y_i=2\) 2/10 4/10

We can then define the two marginal distributions,

\(f_y\)
\(y_i=1\) \(y_i=2\)
4/10 6/10

and,

\(f_x\)
\(x_i=3\) \(x_i=4\)
3/10 7/10

Likewise, we know the conditional distribution \(f_{y|x}\); which we get by dividing the joint distribution by the marginal distribution of \(x\). Each column of the conditional distribution should add up to 1.

\(f_{y|x}\)
\(x_i=3\) \(x_i=4\)
\(y_i=1\) 1/3 3/7
\(y_i=2\) 2/3 4/7

Now we can calculate the following objects:

  1. \(E[y_i]\)

\[ \begin{aligned} E[y_i] =& 1\cdot Pr(y_i=1)+2\cdot Pr(y_i=2) \\ =&1\cdot 4/10+2\cdot 6/10 \\ =&16/10 \end{aligned} \]

  1. \(E[y_i|x_i=3]\)

\[ \begin{aligned} E[y_i|x_i=3] =& 1\cdot Pr(y_i=1|x_i=3)+2\cdot Pr(y_i=2|x_i=3) \\ =&1\cdot 1/3+2\cdot 2/3 \\ =&5/3 \end{aligned} \]

  1. \(E[y_i|x_i=4]\)

\[ \begin{aligned} E[y_i|x_i=4] =& 1\cdot Pr(y_i=1|x_i=4)+2\cdot Pr(y_i=2|x_i=4) \\ =&1\cdot 3/7+2\cdot 4/7 \\ =&11/7 \end{aligned} \]

  1. \(E\big[E[y_i|x_i]\big]\)

\[ \begin{aligned} E\big[E[y_i|x_i]\big] =& E[y_i|x_i=3]\cdot Pr(x_i=3)+ E[y_i|x_i=4]\cdot Pr(x_i=4) \\ =&5/3\cdot3/10+11/7\cdot 7/10 \\ =&16/10 \end{aligned} \]

We have therefore demonstrated the law of iterated expectations.

We can extend this principle to conditional expectations. Suppose you have three random variables/vectors \(\{y_i,x_i,z_i\}\), we can express the conditional expected value of \(y_i\) on \(x_i\) as the (conditional) expected value of the conditional expectation of \(y_i\) on \(x_i\) and \(z_i\).

\[ E[y_i|x_i] = E\big[E[y_i|x_i,z_i]|x_i\big] \]

Here the outside expectation is with respect \(z_i\) conditional on \(x_i\). It utilizes the conditional distribution \(f_{z|x}\) to form the outside expectation,

\[ E[y_i|x_i] = \int t \cdot f_{y|x}(t|x_i)dt = \int\int y \cdot f_{y|x,z}(t|x_i,u)dtf_{z|x}(u|x_i)du = E\big[E[y_i|x_i,z_i]|x_i\big] \]

Properties of the CEF

The following three theorems can be found in a range of Econometrics textbooks and Microeconometrics texts.

Theorem 1 We can express the observed outcome \(y_i\) as a sum of \(E[y_i|x_i]+\varepsilon_i\) where \(E[\varepsilon_i|x_i]=0\) (i.e., mean independent).

Proof.

  1. \(E[\varepsilon_i | x_i] = E[y_i - E[y_i | x_i] | x_i] = E[y_i | x_i] - E[y_i | x_i] = 0\)

  2. \(E[h(x_i)\varepsilon_i] = E[h(x_i)E[\varepsilon_i | x_i]] = E[h(x_i) \times 0] = 0\)

Theorem 2 \(E[y_i|x_i]\) is the best predictor of \(y_i\).

Proof. \[ \begin{aligned} (y_i - m(x_i))^2 =& \left((y_i - E[y_i | x_i]) + (E[y_i | x_i] - m(x_i))\right)^2 \\ =& (y_i - E[y_i \| x_i])^2 + (E[y_i | x_i] - m(x_i))^2 \\&+ 2(y_i - E[y_i | x_i]) \times (E[y_i | x_i] - m(x_i)) \end{aligned} \]

The last term (cross product) is mean zero. Thus, the function is minimized by setting \(m(x_i) = E[y_i | x_i]\).

Theorem 3 [ANOVA Theorem] The variance of \(y_i\) can be decomposed as \(V(E[y_i|x_i])+E(V(y_i|x_i))\)

Proof. \[ \begin{aligned} V(y_i)=&V(E[y_i|x_i] + \varepsilon_i) \\ =&V(E[y_i|x_i])+V(\varepsilon_i) \\ =&V(E[y_i|x_i])+E[\varepsilon_i^2] \end{aligned} \] The second line follows from Theorem 1.1 (independence) and

\[ E[\varepsilon_i^2]=E\left[E[\varepsilon_i^2|x_i]\right]=E\left[V(y_i|x_i)\right] \]

Footnotes

  1. This can be extended to random vectors.↩︎

  2. Some texts use the notation \(E_X\big[E[y_i|x_i]\big]\) to demonstrate that the outside expectation is with respect to \(x_i\).↩︎