flowchart LR A[Data] --> B[Econometric model ] B --> C[Estimator] C --> E[Inference]
Introduction to Econometrics
Pre-sessional Lecture
Reading
This lecture covers material from Wooldridge (2025):
- Chapter 1 - “The Nature of Econometrics and Economic Data”
Econometrics
Econometrics? What is it?
“Econometrics is based upon the development of statistical methods for estimating economic relationships, testing economic theories, and evaluating and implementing government and business policy.” (Wooldridge 2025, 2)
There are a few key elements here:
- statistical methods
- economic relationships
- testing theory
- policy evaluation
Purpose of Econometrics
Econometrics encompasses a set of tools used by Economists to empirically verify claims made about “the real world” (as observed in the data):
Purpose 1: Test economic theory.
e.g., Efficiency wage theory: paying workers in excess of their marginal product will increase effort (improve retention or reduce shirking).
How would we test this theory?
- Need data on worker wages and a measure of effort (?)
- Ideally, we want to compare similar workers, doing the same job, but being paid different wages.
The impact of economic policy is often theoretically ambiguous,
e.g., according to standard economic models, minimum wages should
- decrease in employment under perfect competition
- increase in employment under monopsony
Purpose 2: Evaluate the impact of economic policy.
Statistical significance: Can we reject a null (zero) effect?
Sign: Do we find a positive or negative impact?
Magnitude: How big was the effect (on average)? 1
Evidence on the impact of a policy can be a used as a (in)direct test of economic theory. For example, the literature on Minimum Wages - which broadly finds no disemployment effect - is cited as evidence of monopsony power in the labour market.
Interjection: green boxes
You will notice in the notes that there is a green box!
Econometrics also encompasses the estimation of structural economic models. For example, estimating the macroeconomic Real Business Cycle model. In this setting, the goal is to estimate the structural parameters of the model, which then allows you to do policy evaluation and/or counterfactual analysis.
This module will largely deal with the estimation of reduced-form models. These are models which focus on the relationship between economic variables, but do not necessarily identify structural parameters from an economic model.
- Non-examinable material
- Extra information for the curious
- All other boxes ARE examinable
- Will not appear in future lecture slides
Data
A lot is going to depend on the nature of the data.
Cross-sectional data
- Each observation represents a different unit (individual, firm, state, municipality,…), observed at the same period of time
- e.g. British Social Attitudes Survey
- Notation: \(y_{i}\)
Time series data
- Each observation represents the same unit, observed over multiple periods of time
- e.g. UK CPI
- Notation: \(y_{t}\)
More complex complex data can include both a cross-sectional and time dimension
Repeated/pooled cross-section data
- A collection of multiple cross-sections
- e.g. British Social Attitudes for 1983-2025
- Notation: \(y_{it}\)
Longitudinal/panel data
- Multiple units (individual, firm, state) are observed over multiple periods of time
- e.g. UK Household Longitudinal Study (“Understanding Society”)
- Notation: \(y_{it}\)
DGP
Datasets vary in another important dimension: the data generating process (DGP)
Experimental data
- Key variables are the outcome of a randomized experiment (i.e. probabilistic)
- Controlled experiments: researcher determines the assignment mechanism
- Field (i.e. RCTs) and lab experiments are increasingly common in Economics
Observational (non-experimental) data
- The true data generating process (or assignment mechanism) is unknown
- Also referred to as retrospective data; e.g. surveys, administrative records, etc.
ImportantNatural ExperimentsNatural experiments occur when the assignment mechanism is outside the control of the researcher, but still creates the conditions of an experiment (i.e. treated and control). This assignment need not be random, but must be exogenous to factors relevant to the outcome. These are considered observational studies.
Econometric analysis
Inference will depend on the properties of the estimator
The properties of the estimator will, in turn, depend on the assumptions of the model
The assumptions of the model have to fit (within reason) the data generating process.
This is a not a description of the research process, which typically begins with a research question and testable hypothesis. In reality, applied economic research questions are often constrained by the nature of the data and econometric model. This creates a feedback loop between the econometric analysis and research question. This is particularly the case in observational studies where the researcher does not know the data generating process (or control assignment in the experiment).
Causal Inference
When testing economic theory or evaluating policy, it is important to be able to separate out causation from correlation.
For example, to test efficiency wage theory
\[ E[e_i|w_i = high]-E[e_i|w_i = low] \]
Experiments solve this problem through random assignment
- Higher wages are paid to a random subset of workers.
- In Lecture 6, we will discuss the Potential Outcomes Framework that underpins this theory.
Causal Inference
In observational studies, we can’t rely on random assignment. Instead,
- Control variables: adjust for confounding factors
- Instruments: exploit exogenous sources of variation
- Transformations: transform the model to remove sources of bias
Ceteris Paribus
For example, to test efficiency wage theory, we might compare
\[ E[e_i|w_i = high,c]-E[e_i|w_i = low,c] \]
where \(c\) is a set of worker characteristics:
- e.g. \(c=\{education_i,age_i,gender_i,ethnicity_i\}\)
Now the comparison is between two “similar” groups of workers.
A Latin phrase meaning “all other things being equal”. A phrase that is used to infer causation from a comparison of similar, but not equal, groups.
Example
Brown (2011)
Paper: brown2011quitters “Quitters Never Win: The (Adverse) Incentive Effects of Competing with Superstars”
Theory: “Superstar Effect”:
in tournaments,
with unequal distribution of talent,
may be optimal for less talented to “give up” (reduce effort) when competing against a more talented individual
Data
“Professional golf tournaments, where effort relates relatively directly to performance, present an opportunity to examine empirically the influence of a superstar.” (pp. 983)

Unconditional differences

But these are different tournaments with potentially different players!
Econometric model
The author estimates a model that controls for certain player/event characteristics:
\[ \begin{aligned} strokes_{ij} =& \beta_1 star_j\times HRanked_i + \beta_2star_j\times LRranked_i \\ &+ \beta_3star_j\times URanked_i + \alpha_1 HRanked_i + \alpha_2 LRranked_i \\ &+ \gamma_0 + \gamma_1 X_i + \gamma_2 Y_j + \varepsilon_{ij} \end{aligned} \]
where
- HRanked = top 20; LRranked = 21-200; URanked = unranked
- \(X_i\): player characteristics
- \(Y_j\): event controls (major dummy, weather, purse, field quality, viewership)
Estimation & Inference

EC3301 Housekeeping
Syllabus
| Week | Topic |
|---|---|
| 1 | Linear regression model |
| 2 | Ordinary least squares |
| 3 | Properties of OLS and inference |
| 4 | Hypothesis testing and heteroskedasticity (recorded) |
| 5 | Large sample properties of OLS |
| 6 | 🎉 Independent Learning Week 🎉 |
| 7 | Dummy variables and potential outcomes framework |
| 8 | Model misspecfication and omitted variables |
| 9 | Proxy variables and instrumental variables |
| 10 | Estimation with instrumental variables (recorded) |
| 11 | Simple panel data models and difference in differences |
Key dates
- Week 4: Class Test 1
- 50 minutes, 25% weight
- Covering Lectures 1-3 (and Tutorial 1)
- Week 10: Class Test 2
- 50 minutes, 25% weight
- Covering Lectures 1-9 (with emphasis on 4-8)
- Week : Project
- Individual, 50% weight
- Due
Support
Econometric Labs
- 5x 2 hours
- Weeks 2,4,8,9,10
Tutorials
- 4x 1 hour
- Weeks 3,5,7,11
Textbook
- Wooldridge (2025) Introductory Econometrics: A Modern Approach, Eighth Edition, Cengage International Edition
- You can also use an earlier edition.
Office Hours
- Tues & Wed 11:00-12:00, G3 Castlecliffe
- No reservations required
Overview
| Week | Lecture | Classes | Assessment |
|---|---|---|---|
| 1 | Lecture 1 | ||
| 2 | Lecture 2 | Lab 1 | |
| 3 | Lecture 3 | Tutorial 1 | |
| 4 | Lecture 4 (recorded) | Lab 2 | Class Test 1 |
| 5 | Lecture 5 | Tutorial 2 | |
| 6 | 🎉 | 🎉 | 🎉 |
| 7 | Lecture 6 | Tutorial 3 | |
| 8 | Lecture 7 | Lab 3 | |
| 9 | Lecture 8 | Lab 4 | |
| 10 | Lecture 9 (recorded) | Lab 5 | Class Test 2 |
| 11 | Lecture 10 | Tutorial 4 |
Contact
Moodle Forum: Please use the Moodle forums to ask questions that may be of benefit to other students.
- Culture of supportive online participation.
Email: neil.lloyd@st-andrews.ac.uk
Office: G3 Castlecliffe
Please do not contact me directly via MS Teams. Use the Moodle forums or email.

Code
Lecture material will contain parallel examples in Stata and R for all relevant code
display "Hello EC3301!"Hello EC3301!
print("Hello EC3301!")[1] "Hello EC3301!"
Stata and R?
This is a bit of an experiment: and you are the ‘guinea pigs’
Labs will be ‘run’ in Stata:
- majority of students learnt Stata in EC2203;
- tutors are primarily trained in Stata (as are most economists);
- BUT labs are largely self directed;
- you can choose to complete the R lab (during the lab)
Tutorial material may contain some Stata output, only
Why add R?
- Continuity of learning: some students have worked with R in during sub-honours
- Assessment: you will be allowed to use R for the project
- Learning Outcomes: econometric analysis \(\neq\) programming
Stata or R?

- easy to learn and use
- intuitive GUI
.dtadata format- simple command-based ‘language’
- programming light

- flexible language
- advanced programming
- industry standard (for data science)
- open source \(=\) FREE
- packages, packages, and more packages
(See Lab 1 for more information.)
Survival guide
Thriving in 3301
Attend class: lectures, labs, tutorials
Engage: ask questions, provoke discussion, question everything
Listen: keep track of what the instructor emphasizes
Keep up: follow up on things you don’t understand
Collaborate: \(\text{teaching}=\max\{\text{learning}\}\)
Use LLMs smartly: beware the false sense of ‘learning’
References
Bibliography
Footnotes
Linked to the idea of economic significance and used in cost-benefit analysis.↩︎

