dis normal(2).97724987


EC3301 - Martinmas Semester
Consider a random sample (iid: independently and identitically distributed) drawn from the distribution
\[ y_i\sim N(2,25) \]
\[ \begin{aligned} \mu = 2 \\ \sigma = 5 \end{aligned} \]
\[ \begin{aligned} E[\bar{y}] =& E\Bigg[\frac{1}{n}\sum_{i=1}^n y_i\Bigg] \\ =& \frac{1}{n}\sum_{i=1}^n E[y_i] \\ =& \frac{1}{n}\sum_{i=1}^n 2 \\ =& \frac{2n}{n} \\ =& 2 \end{aligned} \]
\[ \begin{aligned} Var(\bar{y}) =& E[(\bar{y}-\underbrace{E[\bar{y}]}_{=2})^2] \\ =& E\Bigg[\bigg(\frac{1}{n}\sum_{i=1}^n (y_i-2)\bigg)^2\Bigg] \\ =& E\Bigg[\frac{1}{n^2}\sum_{i=1}^n (y_i-2)^2+2\times\frac{1}{n^2}\sum_{i=1}^{n-1} \sum_{j=i+1}^{n} (y_i-2)(y_j-2)\Bigg] \\ =& \frac{1}{n^2}\sum_{i=1}^n \underbrace{E\big[(y_i-2)^2\big]}_{Var(y_i)} + 2\times \sum_{i=1}^{n-1} \sum_{j=i+1}^{n} \underbrace{E\big[(y_i-2)(y_j-2)\big]}_{Cov(y_i,y_j)=0} \\ =& \frac{nVar(y_i)}{n^2} \\ =& \frac{25}{n} \end{aligned} \]
\[ \bar{y} \sim N(2,25/n) \]
Since \(\bar{y} \sim N(2,1/4)\),
\[ \frac{\bar{y}-2}{0.5} \sim N(0,1) \]
\[ \begin{aligned} Pr(\bar{y}<3) =& Pr(\bar{y}-2<3-2) \\ =& Pr\bigg(\frac{\bar{y}-2}{0.5} < \frac{1}{0.5}\bigg) \\ =& Pr(z<2) \end{aligned} \]
For the next few questions we will assume no knowledge of the true mean and variance of \(y\); only that the sample was drawn from a normal distribution \(y_i\sim N(\mu,\sigma^2)\). You draw a random sample of 100 observations from this distribution and the sample has the following characteristics:
Variable | Obs Mean Std. dev. Min Max
-------------+---------------------------------------------------------
y | 100 2.218156 2.082121 -2.61975 6.654652
Mean estimation Number of obs = 100
--------------------------------------------------------------
| Mean Std. err. [95% conf. interval]
-------------+------------------------------------------------
y | 2.218156 .2082121 1.805018 2.631294
--------------------------------------------------------------
The reported standard deviation is the square-root of the sample variance. The sample variance (\(\hat{\sigma}^2\)) is an estimate of the unobserved distribution variance (\(\sigma^2\)). In this sample, \(\hat{\sigma} = 2.082121\). Note, the estimator is
\[ \hat{\sigma}^2 = \frac{1}{n-1}\sum_{i=1}^n (y_i-\bar{y})^2 \]
where division by \(n-1\) results from the fact that the demeaned \(y\) has \(n-1\) degrees of freedom (since the variable has mean zero). It is for this same reason that the \(t\)-distribution in Q.6 has \(n-1\) degrees of freedom.
The reported standard error is the square root of the \(Var(\bar{y})\). As we proved above, \(Var(\bar{y}) = \sigma^2/n\). That means that \(sd(\bar{y}) = \sigma/\sqrt{n}\).
However, we do not know the true value of \(\sigma\). In our estimation of the \(sd(\bar{y})\), we replace \(\sigma\) with \(\hat{\sigma}\): \(se(\bar{y}) = \hat{\sigma}/\sqrt{n}\).
As \(n=100\), we find that \(se = 0.2082121\); exactly equal to the \(\hat{\sigma}/10\).
\[ \frac{\bar{y}-2}{\hat{\sigma}/\sqrt{n}} \sim t_{n-1} \]
\[ \frac{\bar{y}-1.85}{\hat{\sigma}/\sqrt{n}} \underset{H_0}{\sim} t_{n-1} \]
Compute the test statistic:
The mean command stores estimates in a similar way to regress. They are both e-class commands.
_b[] function can be used to call the stored coefficient estimate for a particular variable. The full set of coefficients is stored as in the vector e(b)._se[] function can be used to call the SE of a particular coefficient. This is computed using the covariance matrix stored as e(V).To see the full list of stored estimates, type ereturn list. Note, this does not work for summarize, which is a r-class command. After sum you can run return list to see the format of its stored estimates.
It is good practice to used stored estimates in this way as it prevents typing errors and improves replicability.
R users tend define a model as a list before summarizing it: e.g. model <- lm(yvar ~ xvar, data) followed by summary(model). You can then reference the slope coefficient of the stored model using coef(model)["xvar"] or model$coefficients["xvar"] (replacing xvar with the name of the variable). The get the standard error, you need to do a bit more work: summary(model)$coefficients["xvar", "Std. Error"] references a value from the summary or sqrt(diag(vcov(model)))["xvar"] (compute the square-root of the diagonal of the covariance matrix).
Rejection rule:
The test static remains the same.
Critical values of \(t\)-distribution (which is symmetric):
Since \(|t\text{-stat}|<\text{t}_{99,0.975}\), we fail to reject.
Rejection rule:
Critical values of \(t\)-distribution (which is symmetric):
We reject \(H_0\) and conclude that \(\mu > 1.85\).
The only difference is the critical value. The one-side test has a lower critical value than the two-sided test.
In the two-sided test:
\[ \begin{aligned} Pr(\text{Reject}\;H_0|\mu = 2) =& Pr\Bigg(\bigg|\frac{\bar{y}-1.85}{se}\bigg|>\text{t}_{99,0.975}\Bigg|\mu = 2\Bigg) \\ =& Pr\Bigg(\frac{\bar{y}-1.85}{se}<\text{t}_{99,0.025}\Bigg|\mu = 2\Bigg) + Pr\Bigg(\frac{\bar{y}-1.85}{se}>\text{t}_{99,0.975}\Bigg|\mu = 2\Bigg) \\ =& Pr\Bigg(\frac{\bar{y}-2}{se}<\text{t}_{99,0.025}-\frac{0.15}{se}\Bigg|\mu = 2\Bigg) \\ &+ Pr\Bigg(\frac{\bar{y}-2}{se}>\text{t}_{99,0.975}-\frac{0.15}{se}\Bigg|\mu = 2\Bigg) \\ =& t_{n-1}\big(\text{t}_{99,0.025}-0.15/se\big) + 1-t_{n-1}\big(\text{t}_{99,0.975}-0.15/se\big) \\ \end{aligned} \]
where \(t_{n-1}(\cdot)\) is the CDF of the \(t\)-distribution.
In the one-sided test:
\[ \begin{aligned} Pr(\text{Reject}\;H_0|\mu = 2) =& Pr\Bigg(\frac{\bar{y}-1.85}{se}>\text{t}_{99,0.95}\Bigg|\mu = 2\Bigg) \\ =& Pr\Bigg(\frac{\bar{y}-2}{se}>\text{t}_{99,0.95}-\frac{0.15}{se}\Bigg|\mu = 2\Bigg) \\ =& 1-t_{n-1}\big(\text{t}_{99,0.95}-0.15/se\big) \\ \end{aligned} \]
This shows that the one-sided test is a more powerful test.
In each case, try to prove the following properties of the OLS estimators (\(\hat{\beta_0},\hat{\beta_1}\)) for a simple linear regression model
\[ y_i = \beta_0 + \beta_1 x_i + u_i \]
Substitute in the definitions of \(\hat{\beta}_0\):
\[ \begin{aligned} \sum_{i=1}^n \hat{u}_i =& \sum_{i=1}^n \big(y_i - (\bar{y}-\hat{\beta}_1\bar{x})-\hat{\beta}_1 x_i\big) \\ =& \sum_{i=1}^n\big( (y_i - \bar{y})-\hat{\beta}_1 (x_i-\bar{x})\big) \\ =& \underbrace{\sum_{i=1}^n (y_i - \bar{y})}_{=0} + \hat{\beta}_1 \underbrace{\sum_{i=1}^n(x_i-\bar{x})}_{=0} \\ =& 0 \end{aligned} \]
This demonstrates that the intercept plays an important role in ensuring this result.
This follows from the definition of \(\hat{\beta_0}\), but can also be shown using the decomposition of \(y_i = \hat{y}_i + \hat{u}_i\)
\[ \begin{aligned} \frac{1}{n}\sum_{i=1}^n y_i =& \frac{1}{n}\sum_{i=1}^n \hat{y}_i + \underbrace{\frac{1}{n}\sum_{i=1}^n\hat{u}_i}_{=0} \\ =& \frac{1}{n}\sum_{i=1}^n (\hat{\beta_0} + \hat{\beta_1}x_i) \\ =& \hat{\beta_0} + \hat{\beta_1}\frac{1}{n}\underbrace{\sum_{i=1}^n x_i}_{\bar{x}} \end{aligned} \]
Substitute in the definitions of \(\hat{u}_i\) from 1.
\[ \sum_{i=1}^n x_i \hat{u}_i =\sum_{i=1}^n x_i (y_i - \hat{\beta}_0-\hat{\beta}_1 x_i) \]
Now substitue in \(\hat{\beta}_0\) and re-arrange
\[ = \sum_{i=1}^n x_i (y_i - \bar{y}) -\hat{\beta}_1 \sum_{i=1}^nx_i(x_i-\bar{x}) \]
Substitute in the definition of \(\hat{\beta}_1\)
\[ \begin{aligned} =& \sum_{i=1}^n x_i (y_i - \bar{y}) -\frac{\sum_{i=1}^n (x_i-\bar{x})(y_i-\bar{y})}{\sum_{i=1}^n(x_i-\bar{x})^2} \sum_{i=1}^nx_i(x_i-\bar{x}) \\ =& 0 \end{aligned} \]
The result follows from the fact that,
\[ \sum_{i=1}^n x_i (y_i - \bar{y}) = \sum_{i=1}^n (x_i-\bar{x})(y_i-\bar{y}) \] and,
\[ \sum_{i=1}^nx_i(x_i-\bar{x}) = \sum_{i=1}^n(x_i-\bar{x})(x_i-\bar{x})=\sum_{i=1}^n(x_i-\bar{x})^2 \]
See Lecture 2 and Wooldridge pp. 27 and pp. 757.
The result follows directly from 1. and 3.
\[ \begin{aligned} \widehat{Cov}(x_i,\hat{u}_i) =& \frac{1}{n-1}\sum_{i=1}^n(x_i-\bar{x})(\hat{u}_i-\underbrace{\bar{\hat{u}}_i}_{=0}) \\ =& \frac{1}{n-1}\underbrace{\sum_{i=1}^nx_i\hat{u}_i}_{=0}-\frac{1}{n-1}\sum_{i=1}^n\bar{x}\hat{u}_i \\ =& -\frac{\bar{x}}{n-1}\underbrace{\sum_{i=1}^n\hat{u}_i}_{=0} \\ =& 0 \end{aligned} \]
Since covariance \(=0\) implies correlation \(=0\), we often discuss the result in 3. as the uncorrelatedness of the residual and regressor(s). The result holds for multivariate models too.
One of the most widely estimated regression model in economics is the Mincer equation (Mincer 1974; Lemieux 2006). It is a simple linear regression model that regresses log of wages (or earnings) on years of education and a quadratic term in years of potential experience. Where potential experience is measured using years since school completion.
\[ \ln(wage_i) = \beta_0 + \beta_1 edu_i + \beta_2 exp_i + \beta_3 exp_i^2 + u_i \]
This simple model can be rationalized by a model in which workers choose the length of time the choose to remain in school and delay earning in income in the labour market (Card 2001).
\[ \frac{\partial E[\ln(wage_i)|edu_i,exp_i]}{\partial exp_i} = \beta_2 + 2\times\beta_3 exp_i \]
The function will be convex (concave) if the second derivate is positive (negative). This then depends on the sign of \(\beta_3\), which is not known.
While one can count the number of years someone spends in education, the learning – or human capital – accumulated through education may not not increase in a linear way with years of education. To account for these non-linearities, you could:
For the remaing questions we are going to use a random sample of the 2024 CPS (Source: NBER MORG). The sample includes only male, full-time, workers and excludes self-employed workers. These estimates are based on a random sub-sample of 300 individuals from this dataset.
(Note: Below code run with echo to enable preserve/restore functionality.)

When working with micro-level data (i.e., unit is a firm, individual, household, etc.) you often need to aggregate the data to a lower level (e.g., region, group category). In Stata, this can be done using the collapse command. In R, you can use group_by() function within dplyr. This leads to an irreversible change in the structure of the dataset: you can aggregate data, but not disaggregate it. For this reason, it is a good idea to retain a version of the disaggregated data. In R, this is relatively straight forward because you can assign the aggregated file to a new dataframe. In Stata, you cannot have two datasets open at once. The solution therefore is to preserve a version of the disaggregated data and restore it once you are done with the aggregated file.
Note, preserve and restore must be run together. The command uses temporary files, so if you just run preserve and then later try to restore, the temporary file will no longer exist and you will need to re-open the disaggregated file and reproduce all other changes you made to the dataset.
For parts of the domain (i.e. years of education) the relationship does appear to be linear. This is for years of education from 10-19. Average wages decline beyond 19 and are flat (although noisy) below 10. This may be a sign that level of education matters above a minimum threshold; e.g. completing high school.
(Note: Below code run with echo to enable preserve/restore functionality.)

A similar pattern exists: linear within the domain 10-19 years of education. The log transformation is monotonic, but has a disproportionate impact on larger numbers (e.g. outliers). This may suggest that there are few outliers in the data.
Below is an estimate of the Mincer equation using the same sub-sample:
Source | SS df MS Number of obs = 300
-------------+---------------------------------- F(3, 296) = 33.82
Model | 26.8176018 3 8.93920059 Prob > F = 0.0000
Residual | 78.2323313 296 .264298417 R-squared = 0.2553
-------------+---------------------------------- Adj R-squared = 0.2477
Total | 105.049933 299 .351337569 Root MSE = .5141
------------------------------------------------------------------------------
lnwage | Coefficient Std. err. t P>|t| [95% conf. interval]
-------------+----------------------------------------------------------------
edu | .0998468 .0107859 9.26 0.000 .07862 .1210736
exp | .0208725 .0081481 2.56 0.011 .004837 .0369081
expsq | -.0002387 .000162 -1.47 0.142 -.0005575 .0000801
_cons | 1.732044 .1753344 9.88 0.000 1.386984 2.077104
------------------------------------------------------------------------------
edu.A 1 year increase in years of education increases predicted wages by approximately \(\sim 10\%\), holding other factors constant. This coefficient is right on the edge of values for which it is appropriate to use the \(\hat{\beta}\times 100\) approximation. The exact value is given by,
which is slightly higher at 10.5%.
A query, using Claude, cited a few key studies and meta-analyses. The conclusion was a range of 0.08-0.10; although it did specifiy a lower range of 0.06-0.07 for rich countries. This estimate is therefore slightly higher.
This test is designed to check whether or model estimates a return to education that is significantly higher than the using 0.08 value. It is a one-sided test, so we will reject if
\[ t\text{-stat} = \frac{\hat{\beta_1}-0.08}{se(\hat{\beta_1})}>t^{-1}_{300-4}(0.99) \]
Where \(t^{-1}_{n-k-1}(\cdot)\) is the inverse of the CDF, so \(t^{-1}_{300-4}(0.99)\) is the \(99^{th}\) percentile of a \(t\)-distribution with 296 degrees of freedom.
1.8400626
2.3390116
1.6500177
We fail to reject this null hypothesis at 1% significance level. We would reject at a 5% level.
This is a two-sided test. We will reject if,
\[ |t\text{-stat}| = \Bigg|\frac{\hat{\beta_2}}{se(\hat{\beta_2})}\Bigg|>t^{-1}_{300-4}(0.975) \]
We reject this null hypothesis.
This is a one-sided test. We will reject if,
\[ t\text{-stat} = \frac{\hat{\beta_3}}{se(\hat{\beta_3})}<t^{-1}_{300-4}(0.05) \]
We fail to reject this null hypothesis.
Since our estimate of \(\hat{\beta}_3<0\), the data suggests that the relationship is concave. Wages are increasing in experience, but at a decreasing rate.
For a quadratic function – \(y = ax^2 + bx + c\) – the turning point is given by \(-b/2a\)
Wages reach their peak around 44 years of experience. This is close to retirement for most workers, so it suggests that wages continue to rise throughout a worker’s career.
The following model was inspired by Frank and Goyal (2009) Capital Structure Decisions: Which Factors Are Reliably Important?. The authors use a database of publicly-traded (US-listed) firms, from 1950-2003. They have a rich dataset of firm-level and macro-economic variables. In this example, we do not have that. Instead, we have data from the most recent SEC filings for 2026-Q2. Using this smaller cross-sectional dataset, we can estimate the model:
\[ leverage_i = \beta_0 + \beta_1 size_i + \beta_2 profitability_i + \beta_3 tangibility_i + u_i \]
where,
leverage: total liabilities divided by total assetssize: log of total assetsprofitability: operating income divided by total assetstangibility: physical assets (e.g., property and equipment) divided by total assetsThe sample has been restricted to firms with assets of at least $1 million and a tangibility score within \([0,1]\).
Variable | Obs Mean Std. dev. Min Max
-------------+---------------------------------------------------------
leverage | 271 1.311067 2.914094 0 22.16866
size | 274 18.3593 2.601397 14.00696 26.29069
profitabil~y | 271 -.673556 2.286625 -29.20916 1.338037
tangibility | 274 .1079863 .1624912 0 .9984815
Source | SS df MS Number of obs = 268
-------------+---------------------------------- F(3, 264) = 29.28
Model | 572.0114 3 190.670467 Prob > F = 0.0000
Residual | 1719.38448 264 6.51282001 R-squared = 0.2496
-------------+---------------------------------- Adj R-squared = 0.2411
Total | 2291.39588 267 8.58200705 Root MSE = 2.552
------------------------------------------------------------------------------
leverage | Coefficient Std. err. t P>|t| [95% conf. interval]
-------------+----------------------------------------------------------------
size | -.2137509 .0642591 -3.33 0.001 -.3402764 -.0872253
profitabil~y | -.5102825 .0724029 -7.05 0.000 -.6528431 -.3677218
tangibility | .868687 .9541517 0.91 0.363 -1.010029 2.747403
_cons | 4.792315 1.208169 3.97 0.000 2.413441 7.171188
------------------------------------------------------------------------------
The \(\mathbf{R}^2=0.2496\). This tells us that 25% of the variation in the outcome is explained by the model (i.e. the regressors). It is difficult to say whether this is a high or low value, since different variables can be explained more/less by observable regressors. We can say that the value is no smaller than those found in the cited paper, which is somewhat surprising given that the authors include more regressors in their model (and \(\mathbf{R}^2\) is weakly increasing in the number of regressors).
_cons coefficient. Does this interpretation relate to a real firm?Assuming that the error term has conditional mean zero, we can interpret the intercept as \(E[y_i|\mathbf{x}_i=0]\). In this setting, this corresponds to the leverage of a firm with
size\(=0\Rightarrow\) total assets of 1.profitability\(=0\Rightarrow\) zero profitstangibility \(=0\Rightarrow\) zero tengible assetsThere is unlikely to be a firm of this nature in the SEC filings. Some firms may have zero profits, but SEC firms are unlikely to have total assets of $1.
size, consider a 10% increase in assets. For profitability and tangibility, the regressor is a ratio. Consider the impact of an increase in the ratio of \(\Delta x = 0.01\) (or 1%).size: holding other factors fixed, a 10% increase in assets increases predicted leverage byA reduction of 0.02 in leverage, or in 2% in liabilities as a share of assets.
profitability: holding other factors fixed, a 1% increase in profitability reduced predicted leverage (liabilities/assets) by 0.51%.tangibility: holding other factors fixed, a 1% increase in tangibility increases predicted leverafe by 0.87%. However, this coefficient is not significantly different from 0.Here are a few factors to consider:
Time: this is a more recent dataset covering a very different macroeconomic environment. For example, interests have been historically low since the 2008 financial crisis. The authors also pool many years of data.
Sample selection: the datasets have different sample selection rules, which may lead to a sample representing a very different subset of US firms.
Model: the authors use a model with many more regressors: firm characteristics and macroeconomic variables not observed in this dataset. The estimates reported here may be effected by the absence of these other regressors. [Note: this is something called omitted variable bias, which we will discuss in Week 8.]