Tutorial 2

EC3301 - Martinmas Semester

Hedonic regression models

In this section we examine the an example of multivariate regression model: the hedonic regression model. These models are widely used in real estate to estimate the value of various characteristics of a house: its size, number of rooms, school access, neighbourhood quality, etc. Having decomposed the price, one can then predict the value of another house using its characteristics. In this case, we will make use a small dataset made available with the Wooldridge (2025, 131) textbook which includes the price and housing characteristic information for 506 nieghbourhoods (unit of observation) in the Boston area.

  1. Consider the model below:

    \[ price_i = \beta_0 + \beta_1 rooms_i + \beta_2 proptax_i + \beta_3 dist_i + \beta_4 crime_i + \beta_5 statio_i + u_i \]

    where,

    • \(price\): median housing price (USD)
    • \(rooms\): average number of rooms
    • \(proptax\): property tax per $1000
    • \(dist\): weighted distance (in miles) to 5 employment centers
    • \(crime\): crimes committed in last year (per capita)
    • \(statio\): average student-teacher ratio of schools in the area

    Provide an interpretation for each coefficient and predict its sign. For \(proptax\), think about a $100 increase in tax.

  2. How does the interpretation of \(\beta_1\), \(\beta_2\), and \(\beta_3\) change in the model below:

    \[ \ln(price_i) = \beta_0 + \beta_1 rooms_i + \beta_2 \ln(proptax_i) + \beta_3 \ln(dist_i) + \beta_4 crime_i + \beta_5 statio_i + u_i \]

  3. How would this interpretation of \(\beta_3\) from the above equation change if \(proptax\) was measured in $ and not $ per $1000? And \(\beta_4\) if \(crime\) was measured in crimes committed per 100 people. Would these changes affect the other coefficients in the model?

The following questions relate to the regression output below:

reg lprice rooms lproptax ldist crime stratio

      Source |       SS           df       MS      Number of obs   =       506
-------------+----------------------------------   F(5, 500)       =    167.09
       Model |  52.9136182         5  10.5827236   Prob > F        =    0.0000
    Residual |  31.6686068       500  .063337214   R-squared       =    0.6256
-------------+----------------------------------   Adj R-squared   =    0.6218
       Total |   84.582225       505  .167489554   Root MSE        =    .25167

------------------------------------------------------------------------------
      lprice | Coefficient  Std. err.      t    P>|t|     [95% conf. interval]
-------------+----------------------------------------------------------------
       rooms |   .2633338   .0174317    15.11   0.000     .2290854    .2975821
    lproptax |  -.2005414   .0403609    -4.97   0.000    -.2798393   -.1212435
       ldist |   .0043865   .0266419     0.16   0.869    -.0479574    .0567304
       crime |  -.0127542    .001601    -7.97   0.000    -.0158996   -.0096087
     stratio |  -.0334352   .0059361    -5.63   0.000    -.0450979   -.0217725
       _cons |   10.13379   .2869904    35.31   0.000     9.569931    10.69764
------------------------------------------------------------------------------
  1. Provide an accurate interpretation of the rooms coefficient (i.e. semi-elasticity).

  2. Provide an interpretation of the estimated ldist and crime coefficients.

  3. What is the predicted (median) house price of a 4 bedroom home, subject to a \(30,000\) property tax, 1 (weighted) mile from 5 economic centers, in a neighbourhood that reports 20 crimes per capita and a student-teacher ratio 16. [You can apply coefficients to the nearest 3 decimal points]

CautionOut of date

The data comes a rather old study by Harrison Jr and Rubinfeld (1978) titled Hedonic housing prices and the demand for clean air. The answer you get for this question will not seem realistic given the housing market today. The mean of the price variable (which captures median house price in a neighbourhood) is $22,512. That would not be enough to buy a new car in the US today.

  1. Test the null hypothesis \(H_0:\beta_5\geq-0.025\) at a signficance level of 0.05. You will need the following information on \(t_{n-k-1}^{-1}(0.05)\)
dis invt(e(df_r),0.05)
-1.6479069
  1. Replicate the 95% CI interval for ldist shown above. You need the following information on \(t_{n-k-1}^{-1}(0.975)\),
dis invt(e(df_r),0.975)
1.9647198
  1. Replicate the reported \(\mathbf{R}^2\) using the information reported above under ‘Source’. [Hint: ‘SS’ stands for Sum of Squares.]

  2. Replicate the reported \(F\)-statistic (‘F(5, 500)’) using the reported \(\mathbf{R}^2\). Is the model statistically significant?

Note

Before the next two questions, I will store some relevant Stata output from the first model:

scalar r2_ur = e(r2)
scalar ssr_ur = e(rss)
scalar df_ur = e(df_r)

This will be wiped from memory after estimating a different model.

  1. Test the \(H_0: \beta_{\ln(dist)}=\beta_{\ln(proptax)}=0\) using the reported information below. [The code the \(\mathbf{R}^2\) from a model that excludes ldist and lproptax.]
qui reg lprice rooms crime stratio
dis e(r2)
.60187703
  1. Test the \(H_0: \beta_{\ln(dist)}=0\) using the reported information below. [The code the \(SSR\) from a model that excludes ldist.] Compute the relevant p-value for this test. What is the relationship between the \(F\)-statistic computed here and the \(t\)-statistic shown in the original output.
qui reg lprice rooms lproptax crime stratio
dis e(rss)
31.670324

Properties of OLS

Consider the following decomposition of \(y_i\) based upon the OLS estimator,

\[ y_i = \hat{\beta}_0 + \hat{\beta}_1 x_{1i}+ \hat{\beta}_2 x_{2i} + \hat{u}_i \]

corresponding to the multivariate model

\[ y_i = \beta_0 + \beta_1x_{1i} + \beta_2 x_{2i} + u_i \]

  1. Under MLR. 1-4, you can show that \(E[\hat{\beta}_1] = \beta_1\). What is this property called and what does it mean?

  2. In Lecture 3, we defined

    \[ se(\hat{\beta}_1) = \frac{\hat{\sigma}}{\sqrt{n} sd(x_1) \sqrt{1-R^2_1}} \]

    Explain each component of this equation and how you would compute it.

  3. What happens to the variance of the estimator if \(x_1\) and \(x_2\) are highly correlated?

ImportantMulticollinearity

In the limit, you can consider what happens as \(R_{j}^2\rightarrow 1\). This is the problem of multicollinearity discussed on p. 91 of Wooldridge (2025).

  1. How does the equation above relate to the material from Lecture 2 on partialling out (i.e. Frisch-Waugh Theorem)?

  2. Would the problem of correlated regressors be a problem in the evaluation of a Randomized Control Trial? As we saw in the regression from Karlan, Rigol, and Roth (2026, 24) (in Lecture 1), authors often estimate the impact of treatment in a randomized control trial by estimating a linear regression model with additional controls.

    \[ y_i = \alpha + \beta T_i + \mathbf{x}_i\delta + \varepsilon_i \]

  3. What would the impact of an increase in sample size be?

  4. Randomized control trials are famously low powered. Given the discussion above, what are some of the ways that researchers might increase the power of their test to reject the null hypothesis of no effect (\(H_0: \beta=0\)).

  5. If the error term in the underlying model is heteroskedastic, will the estimated variance be smaller or larger?

  6. When the error variance is not homoskedastic, why can you no longer use the \(t\) or \(F\) distribution? [This is an advanced question, not covered by material in the lectures. See what you can find out.]

References

Harrison Jr, David, and Daniel L Rubinfeld. 1978. “Hedonic Housing Prices and the Demand for Clean Air.” Journal of Environmental Economics and Management 5 (1): 81–102.
Karlan, Dean, Natalia Rigol, and Benjamin N. Roth. 2026. “When Microenterprises Grow, Are Consumers Better Off? Evidence from Large Loans to Microenterprises in Chile.” Working Paper 35729. Working Paper Series. National Bureau of Economic Research. https://doi.org/10.3386/w35729.
Wooldridge, Jeffrey M. 2025. “Introductory Econometrics: A Modern Approach.”