In this section we examine the an example of multivariate regression model: the hedonic regression model. These models are widely used in real estate to estimate the value of various characteristics of a house: its size, number of rooms, school access, neighbourhood quality, etc. Having decomposed the price, one can then predict the value of another house using its characteristics. In this case, we will make use a small dataset made available with the Wooldridge (2025, 131) textbook which includes the price and housing characteristic information for 506 nieghbourhoods (unit of observation) in the Boston area.
How would this interpretation of \(\beta_3\) from the above equation change if \(proptax\) was measured in $ and not $ per $1000? And \(\beta_4\) if \(crime\) was measured in crimes committed per 100 people. Would these changes affect the other coefficients in the model?
The following questions relate to the regression output below:
Provide an accurate interpretation of the rooms coefficient (i.e. semi-elasticity).
Provide an interpretation of the estimated ldist and crime coefficients.
What is the predicted (median) house price of a 4 bedroom home, subject to a \(30,000\) property tax, 1 (weighted) mile from 5 economic centers, in a neighbourhood that reports 20 crimes per capita and a student-teacher ratio 16. [You can apply coefficients to the nearest 3 decimal points]
CautionOut of date
The data comes a rather old study by Harrison Jr and Rubinfeld (1978) titled Hedonic housing prices and the demand for clean air. The answer you get for this question will not seem realistic given the housing market today. The mean of the price variable (which captures median house price in a neighbourhood) is $22,512. That would not be enough to buy a new car in the US today.
Test the null hypothesis \(H_0:\beta_5\geq-0.025\) at a signficance level of 0.05. You will need the following information on \(t_{n-k-1}^{-1}(0.05)\)
dis invt(e(df_r),0.05)
-1.6479069
Replicate the 95% CI interval for ldist shown above. You need the following information on \(t_{n-k-1}^{-1}(0.975)\),
dis invt(e(df_r),0.975)
1.9647198
Replicate the reported \(\mathbf{R}^2\) using the information reported above under ‘Source’. [Hint: ‘SS’ stands for Sum of Squares.]
Replicate the reported \(F\)-statistic (‘F(5, 500)’) using the reported \(\mathbf{R}^2\). Is the model statistically significant?
Note
Before the next two questions, I will store some relevant Stata output from the first model:
This will be wiped from memory after estimating a different model.
Test the \(H_0: \beta_{\ln(dist)}=\beta_{\ln(proptax)}=0\) using the reported information below. [The code the \(\mathbf{R}^2\) from a model that excludes ldist and lproptax.]
quireg lprice rooms crime stratiodis e(r2)
.60187703
Test the \(H_0: \beta_{\ln(dist)}=0\) using the reported information below. [The code the \(SSR\) from a model that excludes ldist.] Compute the relevant p-value for this test. What is the relationship between the \(F\)-statistic computed here and the \(t\)-statistic shown in the original output.
Explain each component of this equation and how you would compute it.
What happens to the variance of the estimator if \(x_1\) and \(x_2\) are highly correlated?
ImportantMulticollinearity
In the limit, you can consider what happens as \(R_{j}^2\rightarrow 1\). This is the problem of multicollinearity discussed on p. 91 of Wooldridge (2025).
How does the equation above relate to the material from Lecture 2 on partialling out (i.e. Frisch-Waugh Theorem)?
Would the problem of correlated regressors be a problem in the evaluation of a Randomized Control Trial? As we saw in the regression from Karlan, Rigol, and Roth (2026, 24) (in Lecture 1), authors often estimate the impact of treatment in a randomized control trial by estimating a linear regression model with additional controls.
What would the impact of an increase in sample size be?
Randomized control trials are famously low powered. Given the discussion above, what are some of the ways that researchers might increase the power of their test to reject the null hypothesis of no effect (\(H_0: \beta=0\)).
If the error term in the underlying model is heteroskedastic, will the estimated variance be smaller or larger?
When the error variance is not homoskedastic, why can you no longer use the \(t\) or \(F\) distribution? [This is an advanced question, not covered by material in the lectures. See what you can find out.]
References
Harrison Jr, David, and Daniel L Rubinfeld. 1978. “Hedonic Housing Prices and the Demand for Clean Air.”Journal of Environmental Economics and Management 5 (1): 81–102.
Karlan, Dean, Natalia Rigol, and Benjamin N. Roth. 2026. “When Microenterprises Grow, Are Consumers Better Off? Evidence from Large Loans to Microenterprises in Chile.” Working Paper 35729. Working Paper Series. National Bureau of Economic Research. https://doi.org/10.3386/w35729.
Wooldridge, Jeffrey M. 2025. “Introductory Econometrics: A Modern Approach.”