Financial Econometrics
R Tutorial Guidance
By
Y IZHI WANG & P ROF. S AMUEL A. V IGNE
Trinity Business School
U NIVERSITY OF D UBLIN
Computer (Windows or MacOS) Required
This tutorial uses R-4.0.3.pkg and RStudio-1.4.1103.dmg
Electronic copy available at: https://ssrn.com/abstract=3863563
Electronic copy available at: https://ssrn.com/abstract=3863563
TABLE OF C ONTENTS
Page
1
2
A BRIEF INTRODUCTION OF R
1
1.1
Not an introduction’s an introduction . . . . . . . . . . . . . . . . . . . . . . . . . . .
1
1.2
What is R? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3
1.3
Why R? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
1.4
What is RStudio? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
1.5
The Useful R Package . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
1.5.1
Data import . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
1.5.2
Data sort . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5
1.5.3
Data visualization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6
1.6
R Tips . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
1.7
Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
7
HOW TO DOWNLOAD AND INSTALL R
9
2.1
How to Download and Install R for Mac . . . . . . . . . . . . . . . . . . . . . . . . . .
9
2.1.1
Step 1. Click the Link to Download R for Mac . . . . . . . . . . . . . . . . . .
9
2.1.2
Step 2. Click the Link to Download R for (Mac) OS X . . . . . . . . . . . . .
10
2.1.3
Step 3. Click the Link to R-4.0.3.pkg (notarized and signed) . . . . . . . . .
10
2.1.4
Step 4. Installation Package of R for Mac is Downloading . . . . . . . . . .
11
2.1.5
Step 5. Install R for Mac . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
11
2.1.6
Step 6.Click the Link to Download RStudio for Mac . . . . . . . . . . . . . .
12
2.1.7
Step 7. Click the Link to RStudio Desktop . . . . . . . . . . . . . . . . . . . .
12
2.1.8
Step 8. Click the Link to Open Source Edition Download . . . . . . . . . . .
13
2.1.9
Step 9. Click the Link to RStudio Desktop Free Download . . . . . . . . . .
13
2.1.10 Step 10. Click the Link to macOS 10.13+ . . . . . . . . . . . . . . . . . . . .
14
2.1.11 Step 11. Installation Package of RStudio for Mac is Downloading . . . . . .
14
2.1.12 Step 12. Install RStudio for Mac . . . . . . . . . . . . . . . . . . . . . . . . . .
15
How to Download and Install R for Windows . . . . . . . . . . . . . . . . . . . . . . .
16
2.2.1
Step 1. Click the Link to Download R for Windows . . . . . . . . . . . . . . .
16
2.2.2
Step 2. Click the Link to Download R for Windows . . . . . . . . . . . . . . .
16
2.2
i
Electronic copy available at: https://ssrn.com/abstract=3863563
TABLE OF CONTENTS
3
Step 3. Click the Link to install R for the first time . . . . . . . . . . . . . .
17
2.2.4
Step 4. Click the Link to Download R 4.0.3 for Windows . . . . . . . . . . .
17
2.2.5
Step 5. The Installation Package of R for Windows is Downloading . . . . .
18
2.2.6
Step 6. Install R for Windows . . . . . . . . . . . . . . . . . . . . . . . . . . .
18
2.2.7
Step 7. Click the Link to Download RStudio for Windows . . . . . . . . . . .
19
2.2.8
Step 8. Click the Link to RStudio Desktop . . . . . . . . . . . . . . . . . . . .
19
2.2.9
Step 9. Click the Link to Open Source Edition Download . . . . . . . . . . .
20
2.2.10 Step 10. Click the Link to RStudio Desktop Free Download . . . . . . . . .
20
2.2.11 Step 11. Click the Link to Download RStudio for Windows . . . . . . . . . .
21
2.2.12 Step 12. Installation Package of RStudio for Windows is Downloading . . .
21
2.2.13 Step 13. Install RStudio for Windows . . . . . . . . . . . . . . . . . . . . . . .
22
BASIC OF R
23
3.1
The Home Interface of RStudio . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
23
3.2
Assignment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
24
3.2.1
Assignment code . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
24
3.2.2
Data type . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
29
3.2.3
Numeric Data Character Data Transformation . . . . . . . . . . . . . . . .
36
Data Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
38
3.3.1
One Dimensional Data Structure - Vector . . . . . . . . . . . . . . . . . . . .
38
3.3.2
Two Dimensional Data Structure - Matrix . . . . . . . . . . . . . . . . . . . .
41
3.3.3
Vector & Matrix Transformation . . . . . . . . . . . . . . . . . . . . . . . . .
44
3.3.4
Multiple Dimensional Data Structure - Data Frame . . . . . . . . . . . . . .
46
3.3.5
Vector, Matrix & Data Frame Transformation . . . . . . . . . . . . . . . . .
49
3.3.6
Multiple Dimensional Data Structure - List . . . . . . . . . . . . . . . . . . .
53
3.3.7
Matrix List Transformation . . . . . . . . . . . . . . . . . . . . . . . . . . . .
57
Data Import and Export . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
58
3.4.1
Import data from an external data source . . . . . . . . . . . . . . . . . . . .
58
Data Export . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
67
3.3
3.4
3.5
4
2.2.3
BASIC OPERATION OF R
69
4.1
Mathematical Operation of R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
69
4.1.1
Addition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
69
4.1.2
Subtraction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
70
4.1.3
Multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
71
4.1.4
Division . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
72
4.1.5
Find quotient division . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
73
4.1.6
Find remainder division . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
74
4.1.7
Find power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
75
ii
Electronic copy available at: https://ssrn.com/abstract=3863563
TABLE OF CONTENTS
4.2
4.3
5
4.1.8
Find root . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
76
4.1.9
Find sum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
77
4.1.10 Find product . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
78
4.1.11 Find cumulation sum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
79
4.1.12 Find cummulation product . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
80
4.1.13 Find results to a specific decimal . . . . . . . . . . . . . . . . . . . . . . . . .
81
4.1.14 Matrix multiplication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
85
4.1.15 Find the inverse matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
88
4.1.16 Find the system of linear equations . . . . . . . . . . . . . . . . . . . . . . . .
90
Logical Operation of R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
93
4.2.1
If equal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
93
4.2.2
If unequal . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
95
4.2.3
If greater than . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
96
4.2.4
If greater than or equal to . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
97
4.2.5
If less than . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
98
4.2.6
If less than or equal to . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
99
4.2.7
If include or not . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100
4.2.8
Logical operation symbol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 101
if-else Statement Iteration Statement . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
4.3.1
if-else statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102
4.3.2
Iteration statement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108
FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.1
5.2
111
Data generation function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
5.1.1
Random number . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111
5.1.2
Replicate function and sequence function . . . . . . . . . . . . . . . . . . . . 116
5.1.3
Sampling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 122
Descriptive data analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
5.2.1
Arithmetic mean value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
5.2.2
Weighted mean value . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 127
5.2.3
Median . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128
5.2.4
Mode . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
5.2.5
Minimum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
5.2.6
Maximum . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
5.2.7
Range . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129
5.2.8
Skewness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
5.2.9
Kurtosis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130
5.2.10 Descriptive data summary function . . . . . . . . . . . . . . . . . . . . . . . . 130
5.2.11 Covariance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
iii
Electronic copy available at: https://ssrn.com/abstract=3863563
TABLE OF CONTENTS
5.2.12 Coefficient . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 134
6
DATA VISUALIZATION OF R - ggplot2
6.1
6.2
7
135
Data Visualization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 135
6.1.1
Continuous univariate distribution graph . . . . . . . . . . . . . . . . . . . . 135
6.1.2
Categorical variate distribution graph . . . . . . . . . . . . . . . . . . . . . . 139
6.1.3
The relationship between two continous variates . . . . . . . . . . . . . . . . 140
6.1.4
The relationship between two categorical variables . . . . . . . . . . . . . . 142
6.1.5
Relationship between categorical variable & continuous variate . . . . . . 144
Details adjusting of graph . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145
6.2.1
Title . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 145
6.2.2
Axis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147
6.2.3
Legend . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149
6.2.4
Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151
FINANCIAL ECONOMETRICS ANALYSIS OF R
7.1
7.2
7.3
7.4
7.5
Analysis of Variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 153
7.1.1
Install and load all the packages . . . . . . . . . . . . . . . . . . . . . . . . . 153
7.1.2
Data set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 155
7.1.3
Normal distribution testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156
7.1.4
Homogeneity of variance testing . . . . . . . . . . . . . . . . . . . . . . . . . . 157
7.1.5
One way analysis of variance . . . . . . . . . . . . . . . . . . . . . . . . . . . . 158
Linear Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159
7.2.1
Data set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159
7.2.2
Model.1 and Model.1 results summary . . . . . . . . . . . . . . . . . . . . . . 160
7.2.3
Regression diagnostic . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 161
7.2.4
Model optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163
Logistic Regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
7.3.1
Data set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
7.3.2
Model construction and model analysis summary . . . . . . . . . . . . . . . 168
7.3.3
Model optimization and optimized model analysis summary . . . . . . . . . 169
7.3.4
Confusion matrix . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 170
Heteroscedasticity Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171
7.4.1
Data set . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 171
7.4.2
Detecting heteroskedasticity - The Breusch-Pagan Test . . . . . . . . . . . . 172
Correlation Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 173
7.5.1
7.6
153
Autocorrelation test_Durbin - Watson Test . . . . . . . . . . . . . . . . . . . 175
Cointegration Test . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 177
7.6.1
Install and library the packages . . . . . . . . . . . . . . . . . . . . . . . . . . 177
iv
Electronic copy available at: https://ssrn.com/abstract=3863563
TABLE OF CONTENTS
7.7
7.6.2
Set up a regression equation with two variables (y1 , y2 ) . . . . . . . . . . . 178
7.6.3
Engle-Granger Two Steps cointegration test . . . . . . . . . . . . . . . . . . 179
ARCH Model and GARCH Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 187
7.7.1
Step one: set a working directory . . . . . . . . . . . . . . . . . . . . . . . . . 187
7.7.2
Step two: install and library the packages . . . . . . . . . . . . . . . . . . . . 189
7.7.3
Step three: Load the data into RStudio . . . . . . . . . . . . . . . . . . . . . . 190
7.7.4
Step four: calculate log returns . . . . . . . . . . . . . . . . . . . . . . . . . . 191
7.7.5
Step five: descriptive statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . 192
7.7.6
Step six: Ljung box test on returns . . . . . . . . . . . . . . . . . . . . . . . . 194
7.7.7
Step seven: Ljung box test on squared returns . . . . . . . . . . . . . . . . . 195
7.7.8
Step eight: ARCH test on returns . . . . . . . . . . . . . . . . . . . . . . . . . 196
7.7.9
Step nine: defining the GARCH model . . . . . . . . . . . . . . . . . . . . . . 197
7.7.10 Step ten: fitting the GARCH model . . . . . . . . . . . . . . . . . . . . . . . . 199
7.7.11 Step eleven: manually compute the Ljung box and ARCH test . . . . . . . . 204
7.7.12 Step twelve: DCC estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 212
7.7.13 Step thirteen: plot method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 220
8
FINANCIAL ECONOMETRICS R CODES DICTIONARY
223
8.1
A brief introduction of R . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 223
8.2
CSV file import . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 228
8.3
Prepare scatter plot . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 230
8.4
GLM examples_initial setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 232
8.5
Illustration of using GLM data set in R . . . . . . . . . . . . . . . . . . . . . . . . . . 238
8.6
GLM logistic regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 250
8.7
Poisson regression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 256
8.8
Diagnostics and model building . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 260
8.9
Introduction to multivariate analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 269
8.10 Multivariate & special distributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . 276
8.11 Nonlinear . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 285
8.12 Multivariate & special distributions testing . . . . . . . . . . . . . . . . . . . . . . . 299
8.13 Principal components analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 308
8.14 Simulating distributions for hypothesis testing . . . . . . . . . . . . . . . . . . . . . 316
8.15 Introduction time series analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 325
8.16 Autocorrelation & Evaluating Performance . . . . . . . . . . . . . . . . . . . . . . . . 332
8.17 Smoothing methods & Regression - Based TS models . . . . . . . . . . . . . . . . . . 339
8.18 Autoregressive (AR) & ARIMA models, forecasting binary outcomes . . . . . . . . . 351
8.19 Panel data regression analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 358
8.20 Advanced error component models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 369
8.21 Robust inference & Endogeneity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 387
v
Electronic copy available at: https://ssrn.com/abstract=3863563
TABLE OF CONTENTS
8.22 Dynamic models & Panel time series . . . . . . . . . . . . . . . . . . . . . . . . . . . 406
8.23 Confirmatory factor analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 422
8.24 Analysis of confirmation factor analysis models kabc . . . . . . . . . . . . . . . . . . 424
8.25 Structural equation models illustration . . . . . . . . . . . . . . . . . . . . . . . . . . 426
8.26 Analysis of structural equation models - roth . . . . . . . . . . . . . . . . . . . . . . 430
8.27 Analysis of structural regression models . . . . . . . . . . . . . . . . . . . . . . . . . 431
8.28 Vector autoregressive models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 435
8.29 VAR, SVAR, VECM, SVECM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 441
8.30 Vector autoregression example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 450
8.31 Multilevel linear models (MLM) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 463
8.32 Three-level multilevel linear models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 467
8.33 Longitudinal data analysis using multilevel models . . . . . . . . . . . . . . . . . . 470
8.34 Graphing data in multilevel contexts . . . . . . . . . . . . . . . . . . . . . . . . . . . 471
8.35 Multilevel generalized linear models (MGLM) . . . . . . . . . . . . . . . . . . . . . . 483
8.36 Generalized linear models recap . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 487
8.37 Introduction to Bayesian probability . . . . . . . . . . . . . . . . . . . . . . . . . . . . 491
8.38 Markov Chain Monte Carlo . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 497
8.39 LDA Topic Modelling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 508
vi
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
1
A BRIEF INTRODUCTION OF R
1.1
Not an introduction’s an introduction
Please run these codes, then you can get a red heart plot. This is the the romantic of Econometrics.
i n s t a l l . packages ( " t i d y v e r s e " )
library ( tidyverse )
t = seq ( 0 , 2 * pi , by = 0 . 1 )
x = 16 * s i n ( t )^3
y = 13 * cos ( t ) − 5 * cos ( 2 * t ) − 2 * cos ( 3 * t ) − cos ( 4 * t )
a = ( x − min ( x ) ) / ( max( x ) − min ( x ) )
b = ( y − min ( y ) ) / ( max( y ) − min ( y ) )
g g p l o t ( data=NULL, aes ( x=x , y=y ) ) +
geom_line ( aes ( c o l o r =I ( ’ white ’ ) ) ) +
geom_polygon ( aes ( f i l l = ’ red ’ ) , show . legend = F ) +
s c a l e _ x _ c o n t i n u o u s ( l a b e l s = NULL) +
s c a l e _ y _ c o n t i n u o u s ( l a b e l s = NULL) +
1
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 1. A BRIEF INTRODUCTION OF R
theme_bw ( ) +
theme ( panel . g r i d . major = element_blank ( ) ,
panel . g r i d . minor = element_blank ( ) ,
panel . border = element_blank ( ) ,
a x i s . t i c k s = element_blank ( ) ,
a x i s . t i t l e = element_blank ( ) ) +
annotate ( ’ text ’ , x=median ( a ) , y=median ( b ) ,
l a b e l = ’ Welcome t o the world o f Econometrics and R’ , s i z e =5)
ggsave ( ’ heart . png ’ , p l o t = l a s t _ p l o t ( ) , dpi = 300)
Figure 1.1: Red heart plot
Source:RStudio
2
Electronic copy available at: https://ssrn.com/abstract=3863563
1.2. WHAT IS R?
Figure 1.2: Red heart plot screenshot
Source:RStudio
1.2
What is R?
R is a language and free software environment for statistical computing and graphics. The R
language is widely used among statisticians and data miners for developing statistical software
and data analysis.
R is a GNU project. What is the GNU project? GNU is an operating system that is free
software. It respects users’ freedom. The GNU operating system consists of GNU packages as
well as free software released by third parties. The development of GNU made it possible to use a
computer without software. R can be considered as a different implementation of S. There are
some essential differences. Still, much code written for S runs unaltered under R.
R was created by Ross Ihaka and Robert Gentleman at the University of Auckland, New
Zealand, and is developed by the R Development Core Team. R is named partly after the first
names of the first two R authors and partly as a play on the name of S.
R provides a wide variety of statistical (linear and nonlinear modeling, classical statistical
tests, time-series analysis, classification, clustering, and et al.) and graphical techniques, and is
highly extensible.
3
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 1. A BRIEF INTRODUCTION OF R
1.3
Why R?
1) R is easy and straightforward to learn, read & write.
2) R is an example of a FLOSS (Free/Libre and Open Source Software).
3) Varieties of packages.
4) R scripts can be shared freely.
1.4
What is RStudio?
RStudio is an integrated development environment (IDE) that allows users to interact with R
more readily. RStudio is similar to the standard RGui but is considerably more user friendly. It
has more drop-down menus, windows with multiple tabs, and many customization options. The
first time users open RStudio, users will see three windows. A forth window is hidden by default
but can be opened by clicking the File drop-down menu, then New File, and then R Script.
1.5
The Useful R Package
1.5.1
Data import
1) feather: feather is file format designed for efficient on-disk serialisation of data frames that
can be shared across programming languages (e.g. Python and R).
2) readr: The goal of readr is to provide a fast and friendly way to read rectangular data (like
csv, tsv, and fwf).
3) readxl: Import excel files into R.
4) openxlsx: This R package simplifies the creation of .xlsx files by providing a high level
interface to writing, styling and editing worksheets.
5) googlesheets: Interact with Google Sheets from R.
6) haven: Import and Export ’SPSS’, ’Stata’ and ’SAS’ Files.
7) httr: The aim of httr is to provide a wrapper for the curl package, customised to the demands
of modern web APIs.
8) rvest: rvest helps users scrape information from web pages.
9) xml2 : The xml2 package is a binding to libxml2, making it easy to work with HTML and
XML from R.
4
Electronic copy available at: https://ssrn.com/abstract=3863563
1.5. THE USEFUL R PACKAGE
10) Rselenium: This is a set of R Bindings for Selenium 2.0 Remote WebDriver.
11) webreadr: webreadr provides utilities for reading access log data in R.
12) DBI: The DBI package defines a common interface between the R and database management systems (DBMS).
13) RMySQL: RMySQL can connect the database of MySQL.
14) RPostgres: RPostgres is an DBI-compliant interface to the postgres database.
15) ROracle: Oracle Database interface (DBI) driver for R. This is a DBI-compliant Oracle
driver based on the OCI.
16) bigrquery: Easily talk to Google’s ’BigQuery’ database from R.
17) PivotalR: PivotalR is a package that enables users of R, the most popular open source
statistical programming language and environment, to interact with Greenplum Database
and the PostgreSQL for big data analytics.
18) dplyr: dplyr is a grammar of data manipulation, providing a consistent set of verbs that
help you solve the most common data manipulation challenges.
19) data.table: The fread() function in data.table can quickly read big dataset.
20) git2r: Provides access to ’Git’ repositories to extract data and running some basic ’Git’
commands.
21) rlist: rlist is a set of tools for working with list objects.
1.5.2
Data sort
1) tidyr: The goal of tidyr is to help you create tidy data.
2) dplyr: dplyr is a grammar of data manipulation, providing a consistent set of verbs that
help you solve the most common data manipulation challenges.
3) purrr: purrr enhances R’s functional programming (FP) toolkit by providing a complete and
consistent set of tools for working with functions and vectors.
4) broom: summarizes key information about statistical objects in tidy tibbles. This makes it
easy to report results, create plots and consistently work with large numbers of models at
once.
5) zoo: An S3 class with methods for totally ordered indexed observations. It is particularly
aimed at irregular time series of numeric vectors/matrices and factors.
5
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 1. A BRIEF INTRODUCTION OF R
1.5.3
Data visualization
1) ggplot: ggplot2 is a system for declaratively creating graphics, based on The Grammar of
Graphics
2) ggthemes: Some extra themes, geoms, and scales for ’ggplot2’.
3) ggmap: A collection of functions to visualize spatial data and models on top of static maps
from various online sources (e.g Google Maps and Stamen Maps).
4) ggiraph: Create interactive ’ggplot2’ graphics using ’htmlwidgets’.
5) ggstance: A ’ggplot2’ extension that provides flipped components: horizontal versions of
’Stats’ and ’Geoms’, and vertical versions of ’Positions’.
6) GGally: The R package ’ggplot2’ is a plotting system based on the grammar of graphics.
7) ggalt: A compendium of ’geoms’, ’coords’ and ’stats’ for ’ggplot2’.
8) ggpubr: ’ggpubr’ provides some easy-to-use functions for creating and customizing ’ggplot2’based publication ready plots.
9) ggforce: ggforce is a package aimed at providing missing functionality to ggplot2 through
the extension system introduced with ggplot2 v2.0.0.
10) ggrepel: Provides text and label geoms for ’ggplot2’ that help to avoid overlapping text
labels.
11) ggraph: ggraph is an extension of ggplot2 aimed at supporting relational data structures
such as networks, graphs, and trees.
12) ggpmisc: Package ’ggpmisc’ (Miscellaneous Extensions to ’ggplot2’) is a set of extensions to
R package ’ggplot2’ (>= 3.0.0) with emphasis on annotations and highlighting related to
fitted models and data summaries.
13) geomnet: Package ’ggpmisc’ (Miscellaneous Extensions to ’ggplot2’) is a set of extensions to
R package ’ggplot2’ (>= 3.0.0) with emphasis on annotations and highlighting related to
fitted models and data summaries.
14) ggExtra: ggExtra is a collection of functions and layers to enhance ggplot2. The flagship
function is ggMarginal, which can be used to add marginal histograms/boxplots/density
plots to ggplot2 scatterplots.
15) gganimate: The grammar of graphics as implemented in the ’ggplot2’ package has been
successful in providing a powerful API for creating static visualisation.
6
Electronic copy available at: https://ssrn.com/abstract=3863563
1.6. R TIPS
16) plotROC: Functions are provided to generate an interactive ROC curve plot for web use,
and print versions.
17) ggspectra: The goal of ’ggspectra’ is to make it easy to plot radiation spectra and similar data,
such and transmittance, absorbance and reflectance spectra, producing fully annotated
publication- and presentation-ready plots.
18) ggnetwork: Geometries to plot network objects with ’ggplot2’.
19) ggradar: This package exports the functionality of ggplot2 to create radar charts.
20) ggTimeSeries: This R package offers novel time series visualisations.
1.6
R Tips
1) Self-practice is the best R teacher.
2) Clear R Console space: Ctrl + L.
3) Create a new R Script: Ctrl+Shift+N.
4) Add R Notes: Ctrl+Shift+C.
5) Pay attention to initials.
6) Objective names can contain letters, numbers, underlines, and dots.
7) Objective names cannot start with underlines and numbers, and numbers cannot follow
the objective names which start with dots.
1.7
Data
All the data files which are used in this tutorial can be downloaded from:
Wang, Yizhi (2021), "DATA_Financial Econometrics - R Tutorial Guidance", Mendeley Data, V1,
doi: 10.17632/wykh9hfmrr.1 http://dx.doi.org/10.17632/wykh9hfmrr.1
7
Electronic copy available at: https://ssrn.com/abstract=3863563
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
2
HOW TO DOWNLOAD AND INSTALL R
2.1
How to Download and Install R for Mac
2.1.1
Step 1. Click the Link to Download R for Mac
Please click:
https://ftp.acc.umu.se/mirror/CRAN/
Figure 2.1: The Home Page of Download R for Mac
Source: R. https://ftp.acc.umu.se/mirror/CRAN/
9
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.1.2
Step 2. Click the Link to Download R for (Mac) OS X
Please click:
Download R for (Mac) OS X
Figure 2.2: The Interface of Download R for (Mac) OS X
Source: R. https://ftp.acc.umu.se/mirror/CRAN/
2.1.3
Step 3. Click the Link to R-4.0.3.pkg (notarized and signed)
Please click:
R-4.0.3.pkg (notarized and signed)
Figure 2.3: The Interface of R-4.0.3.pkg (notarized and signed)
Source: R. https://ftp.acc.umu.se/mirror/CRAN/
10
Electronic copy available at: https://ssrn.com/abstract=3863563
2.1. HOW TO DOWNLOAD AND INSTALL R FOR MAC
2.1.4
Step 4. Installation Package of R for Mac is Downloading
Now, the Installation Package of R for Mac can be downloaded automatically.
Figure 2.4: The Interface of Installation Package of R for Mac is Downloading
Source: R. https://ftp.acc.umu.se/mirror/CRAN/
2.1.5
Step 5. Install R for Mac
Please click the Installation Package of R to install R for Mac. The home page of R are as follows.
Figure 2.5: The Home Page of R for Mac
Source: R
11
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.1.6
Step 6.Click the Link to Download RStudio for Mac
Please click:
https://rstudio.com/products/rstudio/
Figure 2.6: The Home Page of Download RStudio for Mac
Source: RStudio. https://rstudio.com/products/rstudio/
2.1.7
Step 7. Click the Link to RStudio Desktop
Please click RStudio Desktop.
Figure 2.7: The Interface of RStudio Desktop
Source: RStudio. https://rstudio.com/products/rstudio/
12
Electronic copy available at: https://ssrn.com/abstract=3863563
2.1. HOW TO DOWNLOAD AND INSTALL R FOR MAC
2.1.8
Step 8. Click the Link to Open Source Edition Download
Please find Open Source Edition, and click DOWNLOAD RSTUDIO DESKTOP.
Figure 2.8: The Interface of RStudio Open Source Edition Download
Source: RStudio. https://rstudio.com/products/rstudio/
2.1.9
Step 9. Click the Link to RStudio Desktop Free Download
Please find RStudio Desktop Free, and click DOWNLOAD.
Figure 2.9: The Interface of RStudio Desktop Free Download
Source: RStudio. https://rstudio.com/products/rstudio/download/
13
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.1.10
Step 10. Click the Link to macOS 10.13+
Please find macOS 10.13+ and click it.
Figure 2.10: The Interface of the Link of macOS 10.13+
Source: RStudio. https://rstudio.com/products/rstudio/download/download
2.1.11
Step 11. Installation Package of RStudio for Mac is Downloading
Now, the Installation Package of RStudio for Mac can be downloaded automatically.
Figure 2.11: The Interface of Installation Package of RStudio is Downloading
Source: RStudio. https://rstudio.com/products/rstudio/download/download
14
Electronic copy available at: https://ssrn.com/abstract=3863563
2.1. HOW TO DOWNLOAD AND INSTALL R FOR MAC
2.1.12
Step 12. Install RStudio for Mac
Please click the Installation Package of RStudio to install RStudio. The installation results are as
follows.
Figure 2.12: The Home Page of RStudio for Mac
Source:RStudio.
15
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.2
How to Download and Install R for Windows
2.2.1
Step 1. Click the Link to Download R for Windows
Please click:
https://ftp.acc.umu.se/mirror/CRAN/
Figure 2.13: The Home Page of Download R for Windows
2.2.2
Step 2. Click the Link to Download R for Windows
Please click:
Download R for Windows
Figure 2.14: The Interface of Download R for Windows
16
Electronic copy available at: https://ssrn.com/abstract=3863563
2.2. HOW TO DOWNLOAD AND INSTALL R FOR WINDOWS
2.2.3
Step 3. Click the Link to install R for the first time
Please click:
install R for the first time
Figure 2.15: The Interface of install R for the first time
2.2.4
Step 4. Click the Link to Download R 4.0.3 for Windows
Please click:
Download R 4.0.3 for Windows (85 megabytes, 32/64 bit)
Figure 2.16: The Interface of Download R 4.0.3 for Windows
17
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.2.5
Step 5. The Installation Package of R for Windows is Downloading
Now, the Installation Package of R for Windows can be downloaded automatically.
Figure 2.17: The Interface of Installation Package of R for Windows is Downloading
2.2.6
Step 6. Install R for Windows
Please click the Installation Package of R to install R for Windows. The installation results are as
follows.
Figure 2.18: The Home Page of R for Windows
18
Electronic copy available at: https://ssrn.com/abstract=3863563
2.2. HOW TO DOWNLOAD AND INSTALL R FOR WINDOWS
2.2.7
Step 7. Click the Link to Download RStudio for Windows
Please click:
https://rstudio.com/products/rstudio/
Figure 2.19: The Home Page of Download RStudio for Windows
2.2.8
Step 8. Click the Link to RStudio Desktop
Please click RStudio Desktop
Figure 2.20: The Interface of RStudio Desktop
19
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.2.9
Step 9. Click the Link to Open Source Edition Download
Please find Open Source Edition, and click DOWNLOAD RSTUDIO DESKTOP
Figure 2.21: The Interface of RStudio Open Source Edition Download
2.2.10
Step 10. Click the Link to RStudio Desktop Free Download
Please find RStudio Desktop Free, and click DOWNLOAD.
Figure 2.22: The Interface of RStudio Desktop Free Download
20
Electronic copy available at: https://ssrn.com/abstract=3863563
2.2. HOW TO DOWNLOAD AND INSTALL R FOR WINDOWS
2.2.11
Step 11. Click the Link to Download RStudio for Windows
Please find DOWNLOAD RSTUDIO FOR WINDOWS 1.4.1103|156.96MB and click it.
Figure 2.23: The Interface of the Link-DOWNLOAD RSTUDIO FOR WINDOWS
2.2.12
Step 12. Installation Package of RStudio for Windows is Downloading
Now, the Installation Package of RStudio for Windows can be downloaded automatically.
Figure 2.24: The Interface of Installation Package of RStudio is Downloading
21
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 2. HOW TO DOWNLOAD AND INSTALL R
2.2.13
Step 13. Install RStudio for Windows
Please click the Installation Package of RStudio to install RStudio. The installation results are as
follows.
Figure 2.25: The Home Page of RStudio for Windows
22
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
3
BASIC OF R
3.1
The Home Interface of RStudio
Figure 3.1: The home interface of R
Source:RStudio
23
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.2
Assignment
3.2.1
Assignment code
3.2.1.1
Symbolical assignment
Object Naming <- Object Number
For example:
variable.name1<-(1:36)
1) Please type variable.name1<-(1:36) into R Script space.
Figure 3.2: C3p2
Source:RStudio
24
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
2) Please select c(1:36) and press Ctrl+Enter or click the Run button.
Figure 3.3: Select c(1:36) and click run
Source:RStudio
3) Results are as follows:
Figure 3.4: The results of select c(1:36) and click run
Source:RStudio
25
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
4) Please move the cursor behind variable.name1<-(1:36) and press Ctrl+Enter or click the
Run button.
Figure 3.5: The results of move cursor behind variable.name1<-(1:36) and run
Source:RStudio
5) Please select variable.name1, and press Ctrl+Enter or Click the Run button.
Figure 3.6: The results of the select variable. name1 and click run
Source:RStudio
6) Now, it can be found that the numbers of c(1:36) have been assigned to variable.name1.
26
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
3.2.1.2
Function assignment
assign ("Object Naming" <- Object Number)
For example:
assign("variable.name1", c(1:66))
1) Please type assign("variable.name1", c(1:66)) into R Script space.
Figure 3.7: The results of type assign("variable.name1", c(1:66)) into R script
Source:RStudio
27
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
2) Please move the cursor behind assign("variable.name1", c(1:66)), and press Ctrl+Enter
or Click the Run button.
Figure 3.8: The result of type codes and run it
Source:RStudio
3) Now, it can be found that the numbers of c (1:66) have been assigned to variable.name1.
28
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
3.2.2
Data type
3.2.2.1
Numeric type of data
Assignment
num.1<- 666
Notes: It means to assign the number 666 to an object named num.1
1) Please type num.1<- 666 into R Script space.
Figure 3.9: Type num.1<- 666 into R script space
Source:RStudio
29
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
2) Please move the cursor behind num.1<- 666, and press Ctrl+Enter or click the Run
button.
Figure 3.10: The result of move the cursor behind num.1<- 666 and run it
Source:RStudio
3) Please select num.1, and then press Ctrl+Enter or click the Run button.
Notes: Now, the number 666 has been assigned to an object named num.1.
Figure 3.11: The result of move the cursor behind num.1<- 666 and run It
Source:RStudio
30
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
4) Please type mode(num.1), and then press Ctrl+Enter or click the Run button.
Notes: It gives you the type of function of object num.1 which is "numeric".
Figure 3.12: The results of type mode(num.1) and run it
Source:RStudio
5) Please type is.numeric(num.1), and then press Ctrl+Enter or click the Run button.
It tells you whether the object num.1 is a numeric type or not.
The result shows: [1] TRUE
So, object num.1 is numeric.
Figure 3.13: The result of type is.numeric(num.1) and run It
Source:RStudio
31
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.2.2.2
Character type of data
Assignment
char.1<-"Hello,Trinity College Dublin"
Notes: It means to assign the character Hello, Trinity College Dublin to an object
char.1.
1) Please type: char.1<-"Hello,Trinity College Dublin" into R Script space.
Figure 3.14: Type char.1<-"Hello,Trinity College Dublin" into R script space
Source:RStudio
32
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
2) Please move the cursor behind char.1<-"Hello,Trinity College Dublin", and press
Ctrl+Enter or click the Run button.
Figure 3.15: The results of move the cursor and run It
Source:RStudio
3) Please select char.1, and then press Ctrl+Enter or click the Run button.
Now, the character Hello, Trinity College Dublin has been assigned to char.1.
Figure 3.16: The result of select char.1 and run it
Source:RStudio
33
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
4) Please type mode(char.1), and then press Ctrl+Enter or click the Run button.
It can know the type of function char.1 is a character type.
Figure 3.17: The result of type mode(char.1) and run it
Source:RStudio
5) Please type is.character(char.1), and press Ctrl+Enter or click the Run button.
It can take an exam on whether the char.1 is a character type.
The result shows: [1] TRUE
It means the char.1 is a character type.
Figure 3.18: The result of type character(char.1) and run it
Source:RStudio
34
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
6) Please type is.numeric(char.1), and press Ctrl+Enter or click the Run button.
It can take an exam on whether the char.1 is a numeric type.
The result shows: [1] FALSE
It means the char.1 is a numeric type.
Figure 3.19: The results of type is.numeric(char.1) and run it
Source:RStudio
35
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.2.3
Numeric Data Character Data Transformation
1) Please type: num.1<-666 into R Script space, and press Ctrl+Enter or click the Run
button.
Figure 3.20: The result of type num.1<-666 and run it
Source:RStudio
2) Please type: char.1<-"888" into R Script space, and press Ctrl+Enter or click the Run
button.
Figure 3.21: The result of type char.1<-"888" and run it
Source:RStudio
36
Electronic copy available at: https://ssrn.com/abstract=3863563
3.2. ASSIGNMENT
3) Please type: as.numeric(char.1) into R Script space, and press Ctrl+Enter or click the
Run button.
It shows char.1, a character data has been transferred to numeric data.
Figure 3.22: The result of char.1 has been transferred to numeric data
Source:RStudio
4) Please type: as.character(num.1) into R Script space, and press Ctrl+Enter or click the
Run button.
It shows num.1, which is a numeric data has been transferred to character data.
Figure 3.23: The result of num.1 has been transferred to character data
Source:RStudio
37
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.3
Data Structure
3.3.1
One Dimensional Data Structure - Vector
For example:
assign("variable.name1", c(1:66))
Notes: c stand for a conjunction function.
1) Please type: vec.1<-c(1:66) into R Script space.
Figure 3.24: Type vec.1<-c(1:66) into R script space
Source:RStudio
38
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
2) Please move the cursor behind vec.1<-c(1:66), and press Ctrl+Enter or click Run button.
Now, the numbers from 1 to 66 have been assigned to vec.1.
Figure 3.25: Move the cursor behind vec.1<-c(1:66) and run it
Source:RStudio
3) Please select the vec.1, and then press Ctrl+Enter or click the Run button.
It shows the vec.1 is a vector, which has vectors from vector 1 to vector 66.
Figure 3.26: Select vec.1 and run it
Source:RStudio
39
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
4) Please type class(vec.1), and then press Ctrl+Enter or click the Run button.
It can know the type of Function vec.1 is an integer vector type.
Figure 3.27: Type class(vec.1) and run it
Source:RStudio
5) Please type is.vector(vec.1), and then press Ctrl+Enter or click the Run button.
It can take an exam about if vec.1 is a vector.
The result shows: [1] TRUE
It means the vec.1 is a vector.
Figure 3.28: The result of type is.vector(vec.1) and run it
Source:RStudio
40
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
3.3.2
Two Dimensional Data Structure - Matrix
For example:
mat.1<-matrix(c(1:64), nrow=8, ncol=8, byrow=T)
Notes:
1) c stands for a conjunction function.
2) c(1:64) stands for the matrix contains 64 factors.
3) nrow=8 stands for the matrix has eight rows.
4) Ncol=8 stands for the matrix has eight columns.
5) Byrow=T stands for the matrix is arranged by row.
1) Please type: mat.1<-matrix(c(1:64), nrow=8, ncol=8, byrow=T) into R Script space.
Figure 3.29: Type mat.1<-matrix(c(1:64), nrow=8, ncol=8, byrow=T) into R
Source:RStudio
41
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
2) Please move the cursor behind mat.1<-matrix(c(1:64), nrow=8, ncol=8, byrow=T),
and press Ctrl+Enter or click Run button.
Now, the factors from 1 to 66 have been assigned to mat.1
Figure 3.30: Move the cursor and run it
Source:RStudio
3) Please select the mat.1, and press Ctrl+Enter or click Run button.
It shows the mat.1 is a matrix, which consists of the numbers from 1 to 64.
Figure 3.31: Select mat.1 and run it
Source:RStudio
42
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
4) Please type class(mat.1), and press Ctrl+Enter or click Run button.
It can know that the type of function mat.1 is a matrix array type.
Figure 3.32: Type class(mat.1) and run it
Source:RStudio
5) Please type is.matrix(mat.1), and press Ctrl+Enter or click Run button.
It can take an exam if mat.1 is a matrix.
The result shows: [1] TRUE
It means the vec.1 is a matrix.
Figure 3.33: The result of type is.matrix(mat.1) and run it
Source:RStudio
43
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.3.3
Vector & Matrix Transformation
1) Please type: vec.1<-c(1:66) into R Script space, and press Ctrl+Enter or click the Run
button.
Figure 3.34: The result of type vec.1<-c(1:66) and run it
Source:RStudio
2) Please type: mat.1<-matrix(c(1:64), nrow=8, ncol=8, byrow=T) into R Script space, and
press Ctrl+Enter or click the Run button.
Figure 3.35: The result of type codes and run it
Source:RStudio
44
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
3) Please type: as.matrix(vec.1) into R Script space, and press Ctrl+Enter or click the Run
button.
It shows vec.1, which is a vector that has been transferred to a matrix.
Figure 3.36: The results of vec.1 has been transferred to a matrix
Source:RStudio
4) Please type as.vector(mat.1) into R Script space, and press Ctrl+Enter or click the Run
button
It shows mat.1, which is a matrix that has been transferred to a vector.
Figure 3.37: The results of mat.1 has been transferred to a vector
Source:RStudio
45
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.3.4
Multiple Dimensional Data Structure - Data Frame
For example:
df.1<-data.frame(col1=c(1:4),col2=c("a","b","c","d"),col3=c(T,F,T,F))
Notes:
1) df stands for data frame.
2) col1, col2, and col3 stand for the names of each column.
3) c(1:4), c("a","b","c","d") and c(T,F,T,F)) stand for the factors in each column.
1) Please type:
df.1<-data.frame(col1=c(1:4),col2=c("a","b","c","d"),col3=c(T,F,T,F))
Figure 3.38: The results of type code into R
Source:RStudio
46
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
2) Please move the cursor behind
df.1<-data.frame(col1=c(1:4),col2=c("a","b","c","d"),col3=c(T,F,T,F))
Then, press Ctrl+Enter or click Run button. Now, the factors for each column have been
assigned to df.1.
Figure 3.39: Move the cursor and run it
Source:RStudio
3) Please select the df.1, and then press Ctrl+Enter or click the Run button. It can show all
the factors from the data frame.
Figure 3.40: The results of select df.1 and run it
Source:RStudio
47
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
4) Please type class(df.1), and then press Ctrl+Enter or click the Run button.
It can know the df.1 is a data frame.
Figure 3.41: The result of type class(df.1) and run it
Source:RStudio
5) Please type is.data.frame(df.1), and press Ctrl+Enter or click the Run button.
It can know the df.1 is a data frame.
The result shows: [1] TRUE
It means the df.1 is a data frame.
Figure 3.42: The result of type class(df.1) and run it
Source:RStudio
48
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
3.3.5
Vector, Matrix & Data Frame Transformation
1) Please type: vec.1<-c(1:66) into R Script space, and press Ctrl+Enter or click the Run
button.
Figure 3.43: The result of type vec.1<-c(1:66) and run it
Source:RStudio
2) Please type:
mat.1<-matrix(c(1:64), nrow=8, ncol=8, byrow=T)
into R Script space, and press Ctrl+Enter or click the Run button.
Figure 3.44: The result of type vec.1<-c(1:66) and run it
Source:RStudio
49
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3) Please type:
df.1<-data.frame(col1=c(1:4),col2=c("a","b","c","d"),col3=c(T,F,T,F))
into R Script space, and press Ctrl+Enter or click the Run button.
Figure 3.45: Type code and run it
Source:RStudio
4) Please type: as.data.frame(vec.1) into R Script space, and press Ctrl+Enter or click the
Run button.
It shows vec.1, which is a vector that has been transferred to a data frame.
Figure 3.46: The results of vec.1 has been transferred to a data frame
Source:RStudio
50
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
5) Please type: as.data.frame(mat.1) into R Script space, and press Ctrl+Enter or click the
Run button.
It shows mat.1, which is a matrix that has been transferred to a data frame.
Figure 3.47: The results of mat.1 has been transferred to a data frame
Source:RStudio
6) Please type: as.vector(df.1) into R Script space, and press Ctrl+Enter or click the Run
button.
The output results of as.vector(df.1) are the same as the output results of df.1. It means
df.1, which is a data frame can not be transferred to a vector.
Figure 3.48: The results of df.1 can not be transferred to a vector
Source:RStudio
51
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
7) Please type: as.matrix(df.1) into R Script space, and press Ctrl+Enter or click the Run
button.
It shows df.1, which is a data frame that has been transferred to a Matrix.
Figure 3.49: The results of df.1 has been transferred to a matrix
Source:RStudio
52
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
3.3.6
Multiple Dimensional Data Structure - List
For example:
ls.1<list(component1=c(1:10),component2=matrix(c(1:16),nrow=4),component3=data.frame(col1=c(1:4)))
Notes:
1) ls stands for list.
2) The list supports a wide range of data, and it can contain vector, matrix, data frame, and
others.
3) The component1 is a vector. It contains ten vectors from vector 1 to vector 10.
4) The component2 is a matrix, which has four rows and four columns. It contains 16 numbers
from 1 to 16.
5) The component3 is a data frame, which has one column. It contains four numbers from 1 to
4.
1) Please type:
ls.1<-list(component1=c(1:10),
component2=matrix(c(1:16),nrow=4),component3=data.frame(col1=c(1:4)))
into Script space.
Figure 3.50: The results of type codes into R
Source:RStudio
53
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
2) Please move the cursor behind
ls.1<-list(component1=c(1:10),
component2=matrix(c(1:16),nrow=4),component3=data.frame(col1=c(1:4)))
And press Ctrl+Enter or click the Run button.
Now, the factors for each component have been assigned to ls.1.
Figure 3.51: Move the cursor and run it
Source:RStudio
3) Please select the ls.1, and then press Ctrl+Enter or click the Run button. It shows all the
factors in the list.
Figure 3.52: The results of select ls.1 and run it
Source:RStudio
54
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
4) Please type class(ls.1), and then press Ctrl+Enter or click the Run button. It can know
the ls.1 is a list.
Figure 3.53: The result of type class(ls.1) and run it
Source:RStudio
5) Please type is.list(ls.1), and press Ctrl+Enter or click the Run button.
It can take an exam about if ls.1 is a list.
The result shows: [1] TRUE
It means the ls.1 is a list.
Figure 3.54: The result of type is.list(ls.1) and run it
Source:RStudio
55
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
6) Please type unlist(ls.1), and press Ctrl+Enter or click the Run button.
It can release the list and show all the components.
Figure 3.55: The results of type unlist(ls.1) and run it
Source:RStudio
56
Electronic copy available at: https://ssrn.com/abstract=3863563
3.3. DATA STRUCTURE
3.3.7
Matrix List Transformation
1) Please type: mat.1<-matrix(c(1:8), nrow=2, ncol=2, byrow=T) into R Script space, and
press Ctrl+Enter or click the Run button.
Figure 3.56: The results of type code and run it
Source:RStudio
2) Please type: as.list(mat.1) into R Script space, and press Ctrl+Enter or click the Run
button.
It shows mat.1, which is a matrix that has been transferred to a list.
Figure 3.57: The results of type as.list(mat.1) and run it
Source:RStudio
57
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
3.4
Data Import and Export
3.4.1
Import data from an external data source
•txt file
1) Please type txt.t<-read.table(file.choose(), header=T, sep="\t") into R Script, and
press Ctrl+Enter or click the Run button.
Notes:
a) "header" is a logicality parameter, and it can take the first row of a data set as a new
data set’s column name.
b) "sep" stands for the method of how to separate the new data set. sep="\t" stands for the
new data set separate by space.
Figure 3.58: Type txt.t<-read.table(file.choose(), header=T, sep="\t") and Run It
Source:RStudio
58
Electronic copy available at: https://ssrn.com/abstract=3863563
3.4. DATA IMPORT AND EXPORT
2) Please find the txt file needs to be imported, and click the Open button.
Figure 3.59: Find txt file and open it
Source:RStudio
3) Now, the txt file has been imported successfully.
Figure 3.60: Successfully imported the txt file
Source:RStudio
59
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
4) Please select txt.t in the R Script, and press Ctrl+Enter or click the Run button.
Now, the data set from the txt file can be shown in RStudio.
Figure 3.61: Select txt.t and run it
Source:RStudio
60
Electronic copy available at: https://ssrn.com/abstract=3863563
3.4. DATA IMPORT AND EXPORT
•csv file
1) Please type csv<-read.csv(file.choose()) into R Script, and press Ctrl+Enter or click
the Run button.
Notes: read.csv function automatically set header=T and sep="\t" in RStudio.
Figure 3.62: Type csv<-read.csv(file.choose()) and run it
Source:RStudio
61
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
2) Please find the csv file needs to be imported, and click the Open button.
Figure 3.63: Find csv file and open it
Source:RStudio
3) Now, the csv file has been imported successfully.
Figure 3.64: Find csv file and open it
Source:RStudio
62
Electronic copy available at: https://ssrn.com/abstract=3863563
3.4. DATA IMPORT AND EXPORT
4) Please select csv in the R Script, and press Ctrl+Enter or click the Run button.
Now, the data set from the csv file can be shown in RStudio.
Figure 3.65: Select csv and run it
Source:RStudio
63
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
•Excel-xls\xlsx file
1) Please type install.packages("readxl") into R Script, and press Ctrl+Enter or click the
Run button.
Figure 3.66: Successfully download and install readxl packages
Source:RStudio
2) Please type library(readxl) into R Script, and press Ctrl+Enter or click the Run button.
Figure 3.67: Load the readxl package successful
Source:RStudio
64
Electronic copy available at: https://ssrn.com/abstract=3863563
3.4. DATA IMPORT AND EXPORT
3) Please type excel<-read_excel(file.choose()) into R Script, and press Ctrl+Enter or
click the Run button.
Figure 3.68: Type excel<-read_excel(file.choose()) and run it
Source:RStudio
4) Please find the Excel file needs to be imported, and click the Open button.
Figure 3.69: Find excel file and open it
Source:RStudio
65
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
5) Now, the Excel file has been imported successfully
Figure 3.70: Successfully imported the excel file
Source:RStudio
6) Please select excel in the R Script, and press Ctrl+Enter or click the Run button.
Now, the data set from the Excel file can be shown in RStudio.
Figure 3.71: Select excel and run it
Source:RStudio
66
Electronic copy available at: https://ssrn.com/abstract=3863563
3.5. DATA EXPORT
3.5
Data Export
txt file
1) Please type write.table(txt.t,"txt_t for R Tutorial2", sep="\t",row.names=F) into R
Script, and press Ctrl+Enter or click the Run button.
Notes:
a) "txt.t" is the txt file that has been imported into RStudio. For this example, the "txt.t" is
the "txt_t for R Tutorial" file that has been imported into RStudio in the Chapter 5.2.1.
b) "txt_t for R Tutorial2" is the new file name when it is exported.
c) ’sep="\t"’ means when the data set is exported, the data set also separate by space.
d) "row.names=F" means the row names of the data set in RStudio do not be kept when it
be exported.
Figure 3.72: The results of type codes and run it
Source:RStudio
67
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 3. BASIC OF R
2) Please type txt_t for R Tutorial2 file in the Documents folder.
Notes: all the exporting data are stored at the Documents folder automatically.
Figure 3.73: Find txt_t for R Tutorial2 file in the documents folder
Source:RStudio
3) Now, The results of the txt_t for R Tutorial2 file are exported successfully can be shown
on the computer.
Figure 3.74: The results of successfully export txt_t for R Tutorial2 file
Source:RStudio
68
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
4
BASIC OPERATION OF R
4.1
Mathematical Operation of R
4.1.1
Addition
For example:
c(1, 3, 5, 7) + c(10,20)
Notes: It has a Cyclic Extension Algorithm. Because on the left of the plus sign has four numbers
(1, 3, 5, 7), and on the right of the plus sign only has two numbers (10, 20). The Cyclic Extension
Algorithm for this calculation is the numbers on the right of the plus sign need to be circulated to
match the numbers on the left of the plus sign, such as 1+10, 3+20, 5+10 and 7+20.
Please type: c(1, 3, 5, 7) + c(10,20) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.1: The results of type c(1, 3, 5, 7) + c(10,20) and run it
Source:RStudio
69
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.2
Subtraction
For example:
c(1, 3, 5, 7) - c(10,20)
Notes: It has a Cyclic Extension Algorithm. Because on the left of the minus sign has four
numbers (1, 3, 5, 7), and on the right of the minus sign only has two numbers (10, 20). The Cyclic
Extension Algorithm for this calculation is the numbers on the right of the minus sign need to
be circulated to match the numbers on the left of the minus sign, such as 1-10, 3-20, 5-10 and 7-20.
Please type: c(1, 3, 5, 7) - c(10,20) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.2: The results of type c(1, 3, 5, 7) - c(10,20) and run it
Source:RStudio
70
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.3
Multiplication
For example:
c(4, 8, 12, 16) × c(2,4)
Notes: It has a Cyclic Extension Algorithm. Because on the left of the multiple sign has four
numbers (4, 8, 12, 16), and on the right of the multiple sign only has two numbers (2, 4). The
Cyclic Extension Algorithm for this calculation is the numbers on the right of the multiple sign
need to be circulated to match the numbers on the left of the multiple sign, such as 4×2, 8×4,
12×2, 16×4.
Please type: c(4, 8, 12, 16) × c(2,4) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.3: The results of type c(4, 8, 12, 16) × c(2,4) and run it
Source:RStudio
71
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.4
Division
For example:
c(4, 8, 12, 16) ÷ c(2,4)
Notes: It has a Cyclic Extension Algorithm. Because on the left of the division sign has four
numbers (4, 8, 12, 16), and on the right of the division sign only has two numbers (2, 4). The
Cyclic Extension Algorithm for this calculation is the numbers on the right of the division sign
need to be circulated to match the numbers on the left of the division sign, such as 4÷2, 8÷4,
12÷2, 16÷4.
Please type: c(4, 8, 12, 16) ÷ c(2,4) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.4: The results of type c(4, 8, 12, 16) ÷ c(2,4) and run it
Source:RStudio
72
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.5
Find quotient division
For example:
7÷2
Please type: 7%/%2 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.5: The result of type 7%/%2 and run it
Source:RStudio
73
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.6
Find remainder division
For example:
7÷2
Please type: 7%/%2 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.6: The result of type 7%/%2 and run it
Source:RStudio
74
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.7
Find power
For example:
53
Please type: 53 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.7: The result of type 53 and run it
Source:RStudio
75
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.8
Find root
For example:
p
64
ˆ
Please type: 64(1/2)
into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.8: The result of type 64ˆ(1/2) and run it
Source:RStudio
76
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.9
Find sum
For example:
1 + 2 + 3 + ... + 98 + 99 + 100
Please type: sum(1:100) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.9: The result of type sum(1:100) and run it
Source:RStudio
77
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.10
Find product
For example:
1 × 2 × 3 × ... × 98 × 99 × 100
Please type: prod(1:100) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.10: The result of type prod(1:100) and run It
Source:RStudio
78
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.11
Find cumulation sum
For example:
cumsum(1:100)
Please type: cumsum(1:100) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.11: The result of type cumsum(1:100) and run it
Source:RStudio
79
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.12
Find cummulation product
For example:
cumprod(1:50)
Please type: cumprod(1:50) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.12: The result of type cumprod(1:50) and run It
Source:RStudio
80
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.13
Find results to a specific decimal
4.1.13.1
Round up and round down
For example:
Round π to the nearest hundredth
1) Please type: options(digits=3) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.13: The result of type options(digits=3) and run it
Source:RStudio
81
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
2) Please type: round(pi,3) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.14: The result of type round(pi,3) and run it
Source:RStudio
82
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.13.2
Ceiling
For example:
Find the ceiling of π
Please type: ceiling(pi) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.15: The result of type ceiling(pi) and run it
Source:RStudio
83
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.13.3
Floor
For example:
Find the floor of π
Notes: the floor of œÄ is the maximal integral smaller than π.
Please type: floor(pi) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.16: The result of type floor(pi) and run it
Source:RStudio
84
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4.1.14
Matrix multiplication
For example:
"
(4.1)
1 2 3
4 5 6
#
1 4
×
2 5
3 6
1) Please type: mat.1<-matrix(c(1,2,3,4,5,6), nrow=2, byrow=T) into R Script, and press
Ctrl+Enter or click the Run button.
Figure 4.17: The results of type code and run it
Source:RStudio
85
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
2) Please type: mat.1 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.18: The results of select mat.1 and run it
Source:RStudio
3) Please type: mat.2<-matrix(c(1,2,3,4,5,6), nrow=3, byrow=F) into R Script, and press
Ctrl+Enter or click the Run button.
Figure 4.19: The results of type code and run it
Source:RStudio
86
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
4) Please type: mat.2 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.20: The results of select mat.2 and run it
Source:RStudio
5) Please type: mat.1 % * % mat.2 into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.21: The results of type mat.1 % * % mat.2 and run it
Source:RStudio
87
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.15
Find the inverse matrix
For example:
"
(4.2)
1 2
#
3 4
1) Please type: mat.1<-matrix(c(1,2,3,4), nrow=2, byrow=T) into R Script, and press
Ctrl+Enter or click the Run button.
Figure 4.22: The results of type code and run it
Source:RStudio
88
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
2) Please type: mat.1 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.23: The results of select mat.1 and run it
Source:RStudio
3) Please type: solve(mat.1) into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.24: The results of type solve(mat.1) and run it
Source:RStudio
89
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.1.16
Find the system of linear equations
For example:
x + 2y = 7,
(4.3)
3x + 4y = 8,
1) Please type: mat.1<-matrix(c(1,2,3,4), nrow=2, byrow=T) into R Script, and press
Ctrl+Enter or click the Run button.
Figure 4.25: The results of type code and run it
Source:RStudio
90
Electronic copy available at: https://ssrn.com/abstract=3863563
4.1. MATHEMATICAL OPERATION OF R
2) Please type: mat.2<-matrix(c(7,8),nrow=2) into R Script, and press Ctrl+Enter or click
the Run button.
Figure 4.26: The results of type code and run it
Source:RStudio
3) Please type: solve(mat.1, mat.2) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.27: The results of type solve(mat.1, mat.2) and run it
Source:RStudio
91
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4) As a result:
(4.4)
x + 2y = 7,
3x + 4y = 8,
=
x = −6
y = 6.5
92
Electronic copy available at: https://ssrn.com/abstract=3863563
4.2. LOGICAL OPERATION OF R
4.2
Logical Operation of R
4.2.1
If equal
For example:
6 and 6
6 and 8
1) Please type: 6 == 6 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] TRUE, which means 6 is equal to 6.
Figure 4.28: The result of type 6 == 6 and run it
Source:RStudio
93
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
2) Please type: 6 == 8 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] FALSE, which means 6 is unequal to 8.
Figure 4.29: The result of type 6 == 8 and run it
Source:RStudio
94
Electronic copy available at: https://ssrn.com/abstract=3863563
4.2. LOGICAL OPERATION OF R
4.2.2
If unequal
For example:
6 and 8
Please type: 6!=8 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] TRUE, which means 6 is unequal to 8.
Figure 4.30: The result of type 6! = 8 and run it
Source:RStudio
95
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.2.3
If greater than
For example:
6 and 8
Please type: 6 > 8 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] FALSE, which means 6 is not greater than 8.
Figure 4.31: The result of type 6 > 8 and run it
Source:RStudio
96
Electronic copy available at: https://ssrn.com/abstract=3863563
4.2. LOGICAL OPERATION OF R
4.2.4
If greater than or equal to
For example:
6 and 8
Please type: 6 >= 8 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] FALSE, which means 6 is not greater than or equal to 8.
Figure 4.32: The result of type 6 >= 8 and run it
Source:RStudio
97
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.2.5
If less than
For example:
6 and 8
Please type: 6 < 8 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] TRUE, which means 6 is less than 8.
Figure 4.33: The result of type 6 < 8 and run it
Source:RStudio
98
Electronic copy available at: https://ssrn.com/abstract=3863563
4.2. LOGICAL OPERATION OF R
4.2.6
If less than or equal to
For example:
6 and 8
Please type: 6 <= 8 into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] TRUE, which means 6 is less than or equal to 8.
Figure 4.34: The result of type 6 < 8 and run it
Source:RStudio
99
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.2.7
If include or not
For example:
c(6,8)
Please type: 6 %in% c(6,8) into R Script, and press Ctrl+Enter or click the Run button.
The result shows: [1] TRUE, which means 6 is included in c(6, 8).
Figure 4.35: The result of type 6%in%c(6, 8) and run it
Source:RStudio
100
Electronic copy available at: https://ssrn.com/abstract=3863563
4.2. LOGICAL OPERATION OF R
4.2.8
Logical operation symbol
And: &
Or: |
Not: !
Exclusive: xor()
101
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.3
if-else Statement Iteration Statement
4.3.1
if-else statement
4.3.1.1
if-else function
For example:
x <- 8
ifelse (x>=6, x, x*2)
x <- 5
It means if x greater than 6 or equal to 6, then output x. If x less than 6, then multiply x with 2.
1) Please type: x <- 8 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.36: The result of type x < −8 and run it
Source:RStudio
102
Electronic copy available at: https://ssrn.com/abstract=3863563
4.3. IF-ELSE STATEMENT ITERATION STATEMENT
2) Please type: ifelse (x>=6, x, x*2) into R Script, and press Ctrl+Enter or click the Run
button.
x < −8 greater than 6, the result is : [1] 8
Figure 4.37: The result of type i f else(x >= 6, x, x ∗ 2) and run it
Source:RStudio
3) Please type: x <- 5 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.38: The result of type x < −5 and run it
Source:RStudio
103
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4) Please move cursor behind ifelse (x>=6, x, x*2) again, and press Ctrl+Enter or click the
Run button.
x < −5 less than 6, so 5 multiply with 2. The result is: [1] 10.
Figure 4.39: Move the cursor behind i f else(x >= 6, x, x ∗ 2) and run it
Source:RStudio
104
Electronic copy available at: https://ssrn.com/abstract=3863563
4.3. IF-ELSE STATEMENT ITERATION STATEMENT
4.3.1.2
if-else structure
For example:
x <- 8
if(x>=6)print(x)elseprint(x*10)
x <- 5
It means if x greater than 6 or equal to 6, then output x. If x less than 6, then multiply x with 10.
1) Please type: x <- 8 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.40: Move the cursor behind x < −8 and run it
Source:RStudio
105
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
2) Please type: if(x>=6)print(x)elseprint(x*10) into R Script, and press Ctrl+Enter or
click the Run button.
x <- 8 greater than 6. The result is: [1] 8
Figure 4.41: The results of type code and run it
Source:RStudio
3) Please type: x <- 5 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.42: The result of type x < −5 and run it
Source:RStudio
106
Electronic copy available at: https://ssrn.com/abstract=3863563
4.3. IF-ELSE STATEMENT ITERATION STATEMENT
4) Please move cursor behind if(x>=6)print(x)elseprint(x*10) again, and press Ctrl+Enter
or click the Run button.
x <- 5 less than 6, so 5 multiply with 10. The result is: [1] 50
Figure 4.43: Move the cursor behind ifelse code again and run it
Source:RStudio
107
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
4.3.2
Iteration statement
4.3.2.1
for statement
For example:
for (a in (1:10))print(a+10)
Notes: for statement has an iteration range, and all the iteration factors in the iteration range
should carry out the order.
Please type: for (a in (1:10))print(a+10) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 4.44: The results of type code and run it
Source:RStudio
108
Electronic copy available at: https://ssrn.com/abstract=3863563
4.3. IF-ELSE STATEMENT ITERATION STATEMENT
4.3.2.2
while statement
For example:
a < −10
while(a<=20) { print(a)(a < −a + 1)}
Notes: while statement sets conditions into the iteration range. Only the iteration factors which
meet the conditions should carry out the order.
1) Please type: a <- 10 into R Script, and press Ctrl+Enter or click the Run button.
Figure 4.45: The result of type a < −10 and run it
Source:RStudio
109
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 4. BASIC OPERATION OF R
2) Please type: whil e(a <= 20) { print(a)(a < −a + 1)} into R Script, and press Ctrl+Enter or
click the Run button.
Figure 4.46: The results of type codes and run it
Source:RStudio
110
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
5
FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.1
Data generation function
5.1.1
Random number
5.1.1.1
Density function of the normal distribution: dnorm
For example:
dnorm(8, mean=10, sd=2)
Notes: It stands for a probability density function in which the mean is 10, and the standard
deviation is 2. When the random variable is 8, what is the result of the probability density
function?
Please type dnorm(8, mean=10, sd=2) into R Script, and press Ctrl+Enter or click the Run
button.
The result is: [1] 0.1209854, which means when the random variable is 8, the result of the
probability density function is 0.1209854.
111
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
Figure 5.1: The result of type dnorm(8, mean=10, sd=2) and run it
Source:RStudio
112
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
5.1.1.2
Cumulative density function of the normal distribution: pnorm
For example:
pnorm(8, mean=10, sd=2)
Notes: It stands for a cumulative density function in which the mean is 10, and the standard
deviation is 2. what is the cumulative probability when the random variable less than 8?
Please type pnorm(8, mean=10, sd=2) into R Script, and press Ctrl+Enter or click the Run
button.
The result is: [1] 0.1586553, which means the cumulative probability when the random
variable less than 8 is 0.1586553.
Figure 5.2: The result of type pnorm(8, mean=10, sd=2) and run it
Source:RStudio
113
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.1.1.3
Quantile function of the normal distribution: qnorm
The quantile function is simply the inverse of the cumulative density function. Therefore, the
quantile function maps from probabilities to values.
For example:
qnorm(0.1586553, mean=10, sd=2)
It stands for a quantile function in which the mean is 10, and the standard deviation is 2. When
the cumulative probability is 0.1586553, please find the random variable?
The result is: [1] 8. It means when the cumulative probability is 0.1586553, the random variable
is 8.
Figure 5.3: The result of type qnorm(0.1586553, mean=10, sd=2) and run it
Source:RStudio
114
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
5.1.1.4
Random sampling from the normal distribution: rnorm
For example:
rnorm(100, mean=10, sd=2)
Notes: It stands for 100 random numbers are generated. These random numbers follow the
normal distribution. The mean of these random numbers is 10, and the standard deviation of
these random numbers in 2.
Please type rnorm(100, mean=10, sd=2) into R Script, and press Ctrl+Enter or click the
Run button.
Figure 5.4: The result of type rnorm(100, mean=10, sd=2) and run it
Source:RStudio
115
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.1.2
Replicate function and sequence function
5.1.2.1
rep function
Example 1:
rep(1:6, times=3)
Notes: It means to replicate the numbers from 1 to 6 for three times
Please type rep(1:6, times=3) into R Script, and press Ctrl+Enter or click the Run button.
Figure 5.5: The result of type rep(1:6, times=3) and run it
Source:RStudio
116
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
Example 2:
rep(1:6, each=3)
Notes: It means to replicate each number among 1 and 6 for three times.
Please type rep(1:6, each=3) into R Script, and press Ctrl+Enter or click the Run button.
Figure 5.6: The result of type rep(1:6, each=3) and run it
Source:RStudio
117
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
Example 3:
rep(1:6, each=3, times=2)
Notes: It means firstly to replicate each number among 1 and 6 for three times. And then replicate
the results two times.
Please type rep(1:6, each=3, times=2) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 5.7: The result of type rep(1:6, each=3, times=2) and run it
Source:RStudio
118
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
Example 4:
rep(1:6, each=3, times=2, length=20)
Notes: It means firstly to replicate each number among 1 and 6 for three times. And then replicate
the results two times. In the end, only 20 numbers are kept.
Please type rep(1:6, each=3, times=2, length=20) into R Script, and press Ctrl+Enter or
click the Run button.
Figure 5.8: The results of type codes and run it
Source:RStudio
119
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.1.2.2
seq function
Example 1:
seq(from=0, to=88, by=11)
Notes: It means finding the results in the 0 to 88 sequence by add 11 each time.
Please type seq(from=0, to=88, by=11) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 5.9: The results of type seq(from=0, to=88, by=11) and run it
Source:RStudio
120
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
Example 2:
seq(from=0, to=88, by=11, length=5)
Notes: It means finding the numbers in the 0 to 88 sequence by add 11 each time. In the end,
only keep 5 numbers
Please type seq(from=0, to=88, by=11, length=5) into R Script, and press Ctrl+Enter or
click the Run button.
The result is: Error in seq.default(from = 0, to = 88, by = 11, length = 5) : too many
arguments. Because replicate function and sequence function only can cover no more than three
arguments.
Figure 5.10: The results of type codes and run it
Source:RStudio
121
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.1.3
Sampling
5.1.3.1
Random sampling
For example:
sample(1:10, size=20, replace=T, prob=c(rep(0.05,9),0.55))
Notes: It means in this random sampling, the population is 10, the sample size is 20. It is a
simple random sampling with replacement(srswr). The sampled probability of the numbers from
1 to 9 is 5%, and the sampled probability of the number 10 is 55%.
Please type sample(1:10, size=20, replace=T, prob=c(rep(0.05,9),0.55)) into R Script, and
press Ctrl+Enter or click the Run button.
Figure 5.11: The results of random sampling
Source:RStudio
122
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
5.1.3.2
Stratified sampling
For example:
strata(iris, stratanames = "Species", size = (3:5), method=c("srswor"))
Notes: It means the target of the stratified sampling is the iris’s species. Three samples will be
sampled in the first strata, four samples will be sampled in the second strata, and five samples
will be sampled in the third strata. It is a simple random sampling without replacement(srswor).
1) Please type install.packages("sampling") into R Script, and press Ctrl+Enter or click
the Run button.
Figure 5.12: The results of type install.packages("sampling") and run it
Source:RStudio
123
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
2) Please type library(sampling) into R Script, and press Ctrl+Enter or click the Run
button.
Figure 5.13: The results of type library(sampling) and run It
Source:RStudio
3) Please type head(iris) into R Script, and press Ctrl+Enter or click the Run button.
Notes: iris is a data set that comes from RStudio, and head function can get the first several
lines of a data set.
Figure 5.14: The results of type head(iris) and run it
Source:RStudio
124
Electronic copy available at: https://ssrn.com/abstract=3863563
5.1. DATA GENERATION FUNCTION
4) Please type table(iris$Species) into R Script, and press Ctrl+Enter or click the Run
button.
Notes: It can get the species data of iris.
The results are:
Table 5.1: The results of type table(iris$Species)
setosa
50
versicolor
50
virginica
50
Figure 5.15: The results of type table(iris$Species) and run it
Source:RStudio
125
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5) Please type strata(iris, stratanames = "Species", size = (3:5), method=c("srswor")),
and press Ctrl+Enter or click the Run button.
Figure 5.16: The results of stratified sampling
Source:RStudio
126
Electronic copy available at: https://ssrn.com/abstract=3863563
5.2. DESCRIPTIVE DATA ANALYSIS
5.2
Descriptive data analysis
For example
set.seed(66)
test<-sample(1:10, 30, replace=T)
Notes: "set.seed function" can make the sampling results keep the same for each time. And users
can choose the set.seed function’s number whatever they like.
5.2.1
Arithmetic mean value
set.seed(66)
test<-sample(1:10, 30, replace=T)
mean(test)
sum(test)/length(test)
> s e t . seed ( 6 6 )
> t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
> mean ( t e s t )
[ 1 ] 5.333333
> sum( t e s t ) / length ( t e s t )
[ 1 ] 5.333333
5.2.2
Weighted mean value
5.2.2.1
Method one
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
weight<-sample(1:5,30,replace=T)
weighted.mean(test,weight)
> s e t . seed ( 6 6 )
> t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
> weight<−sample ( 1 : 5 , 3 0 , r e p l a c e =T )
> weighted . mean ( t e s t , weight )
[ 1 ] 5.352381
127
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.2.2.2
Method two
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
weight<-sample(1:5,30,replace=T)
t(test) %*% weight/sum(weight)
> s e t . seed ( 6 6 )
> t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
> weight<−sample ( 1 : 5 , 3 0 , r e p l a c e =T )
> t ( t e s t ) %*% weight / sum( weight )
[ ,1]
[ 1 , ] 5.352381
5.2.2.3
Method three
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
weight<-sample(1:5,30,replace=T)
crossprod(test,weight)/sum(weight)
> s e t . seed ( 6 6 )
> t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
> weight<−sample ( 1 : 5 , 3 0 , r e p l a c e =T )
> crossp ro d ( t e s t , weight ) / sum( weight )
[ ,1]
[ 1 , ] 5.352381
5.2.3
Median
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
median(test)
128
Electronic copy available at: https://ssrn.com/abstract=3863563
5.2. DESCRIPTIVE DATA ANALYSIS
> s e t . seed ( 6 6 )
> t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
> median ( t e s t )
[1] 5
5.2.4
Mode
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
sort(table(test),decreasing=T)
> s e t . seed ( 6 6 )
> t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
> s o r t ( t a b l e ( t e s t ) , decreasing=T )
test
5
4
1
2
8 10
6
7
9
3
6
5
3
3
3
2
2
2
1
5.2.5
3
Minimum
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
min(test)
5.2.6
Maximum
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
max(test)
5.2.7
Range
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
range(test)
129
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.2.8
Skewness
Codes:
s e t . seed ( 6 6 )
t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
sum( t e s t −mean ( t e s t ) ) ^ 3 / sd ( t e s t ) ^ 3 / length ( t e s t )
5.2.9
Kurtosis
Codes:
s e t . seed ( 6 6 )
t e s t <−sample ( 1 : 1 0 , 30 , r e p l a c e =T )
sum( t e s t −mean ( t e s t ) ) ^ 4 / sd ( t e s t ) ^ 4 / length ( t e s t )
5.2.10
Descriptive data summary function
5.2.10.1
Method one
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
summary(test)
Figure 5.17: Method one descriptive data summary results
Source:RStudio
130
Electronic copy available at: https://ssrn.com/abstract=3863563
5.2. DESCRIPTIVE DATA ANALYSIS
5.2.10.2
Method two
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
summary(test)
install.packages("timeDate")
install.packages("timeSeries")
install.packages("fBasics")
library(timeDate)
library(timeSeries)
library(fBasics)
basicStats(test)
Figure 5.18: Method two descriptive data summary results
Source:RStudio
131
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.2.10.3
Method three
Codes:
set.seed(66)
test<-sample(1:10, 30, replace=T)
install.packages("psych")
library(psych)
describe(test)
Figure 5.19: Method three descriptive data summary results
Source:RStudio
132
Electronic copy available at: https://ssrn.com/abstract=3863563
5.2. DESCRIPTIVE DATA ANALYSIS
5.2.10.4
Descriptive data summary for data sets
Codes:
set.seed(66)
test.df <- data.frame(x=sample(1:20,5,replace=T),
y=rnorm(5,mean=3,sd=1))
apply(test.df,1,summary)
Notes: In "apply function", the number of "1" stands for calculating by row, and the number of "2"
stands for calculating by column.
Figure 5.20: The results of descriptive data summary for data sets
Source:RStudio
133
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 5. FINANCIAL ECONOMETRICS DATA PROCESSING OF R
5.2.11
Covariance
Codes:
set.seed(66)
test.df <- data.frame(x=sample(1:20,5,replace=T),
y=rnorm(5,mean=3,sd=1))
cov(test.df)
5.2.12
Coefficient
Codes:
set.seed(66) test.df <- data.frame(x=sample(1:20,5,replace=T),
y=rnorm(5,mean=3,sd=1))
cor(test.df)
134
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
6
DATA VISUALIZATION OF R - GGPLOT 2
6.1
Data Visualization
6.1.1
Continuous univariate distribution graph
For example:
library ( tidyverse )
s e t . seed ( 6 6 )
graph . data.1< − data . frame ( x . c=rnorm ( 2 0 0 ,mean=5 , sd = 2 ) )
135
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
6.1.1.1
Histogram
Codes:
s e t . seed ( 6 6 )
graph . data.1< − data . frame ( x . c=rnorm ( 2 0 0 ,mean=5 , sd = 2 ) )
library ( tidyverse )
g g p l o t ( graph . data . 1 , mapping=aes ( x=x . c ) ) +
geom_histogram ( bins =20 , f i l l = " white " , c o l o r =" black " )
Notes:
a) There are three essential parts to use R to make a graph, data sets, mapping, and geom.
b) "aes" is the abbreviation of aesthetics. ggplot2 uses aes ( ) function to make a relationship
between graph properties and data set.
c) "bins=20" stands for there are 20 different groups in the graph.
d) "fill" stands for the color of each bar chart in the histogram, and “color” stands for the
frame color of these bar charts.
Figure 6.1: Histogram
Source:RStudio
136
Electronic copy available at: https://ssrn.com/abstract=3863563
6.1. DATA VISUALIZATION
6.1.1.2
Geom density
Codes:
g g p l o t ( graph . data . 1 , aes ( x=x . c ) ) +
geom_density ( )
Figure 6.2: Geom density
Source:RStudio
137
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
6.1.1.3
Composite graph
Codes:
g g p l o t ( graph . data . 1 , aes ( x= x . c , y = . . d e n s i t y . . ) ) +
geom_histogram ( bins =20 , f i l l =" white " , c o l o r =" black " ) +
geom_line ( s t a t =" d e n s i t y " , c o l o r =" blue " , s i z e =2)+
geom_line ( s t a t =" d e n s i t y " , a d j u s t = 1 . 5 , c o l o r =" red " , s i z e =2)
Figure 6.3: Composite graph
Source:RStudio
138
Electronic copy available at: https://ssrn.com/abstract=3863563
6.1. DATA VISUALIZATION
6.1.2
Categorical variate distribution graph
Codes:
graph . data.1< − data . frame ( x . d=sample (LETTERS[ 1 : 8 ] , s i z e =200 , r e p l a c e =T ) )
g g p l o t ( graph . data . 1 , aes ( x=x . d ) ) +
geom_bar ( )
Figure 6.4: Categorical variate distribution graph
Source:RStudio
139
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
6.1.3
The relationship between two continous variates
For example:
s e t . seed ( 6 6 )
graph . data.2< − data . frame ( v1=sample ( 1 : 3 0 , 2 0 0 , r e p l a c e =T)+ rnorm ( 2 0 0 , sd = 2 ) )
graph . data.2< − graph . data . 2 %>% mutate ( v4=2+3* v1+rnorm ( 2 0 0 , sd = 5 ) )
6.1.3.1
Scatter graph
Codes:
g g p l o t ( graph . data . 2 , aes ( x=v1 , y=v4 ) ) +
geom_point ( )
Figure 6.5: Scatter graph
Source:RStudio
140
Electronic copy available at: https://ssrn.com/abstract=3863563
6.1. DATA VISUALIZATION
6.1.3.2
Insert a fitting line into scatter graph
Codes:
g g p l o t ( graph . data . 2 , aes ( x=v1 , y=v4 ) ) +
geom_point ( ) +
geom_smooth ( method="lm " )
Figure 6.6: The results of insert a fitting line into scatter graph
Source:RStudio
141
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
6.1.4
The relationship between two categorical variables
For example:
s e t . seed ( 6 6 )
graph . data.2< − data . frame ( v2=sample (LETTERS[ 1 : 5 ] , 2 0 0 , r e p l a c e =T ) ,
v3=sample ( month . name, 2 0 0 , r e p l a c e =T ) )
6.1.4.1
Shown as different size of black dot
Codes:
s e t . seed ( 6 6 )
graph . data.2< − data . frame ( v2=sample (LETTERS[ 1 : 5 ] , 2 0 0 , r e p l a c e =T ) ,
v3=sample ( month . name, 2 0 0 , r e p l a c e =T ) )
Figure 6.7: Shown as different size of black dot
Source:RStudio
142
Electronic copy available at: https://ssrn.com/abstract=3863563
6.1. DATA VISUALIZATION
6.1.4.2
Shown as different colour shade
Codes:
graph . data . 2 %>%
count ( v2 , v3 ) %>%
g g p l o t ( mapping=aes ( x=v2 , y=v3 ) ) +
geom_tile ( mapping=aes ( f i l l =n ) )
Figure 6.8: Shown as different colour shade
Source:RStudio
143
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
6.1.5
Relationship between categorical variable & continuous variate
Codes:
s e t . seed ( 6 6 )
graph . data.2< − data . frame ( v1=sample ( 1 : 3 0 , 200 , r e p l a c e =T)+ rnorm ( 2 0 0 , sd = 2 ) ,
v2=sample (LETTERS[ 1 : 5 ] , 2 0 0 , r e p l a c e =T ) ,
v3=sample ( month . name, 2 0 0 , r e p l a c e =T ) )
graph . data.2< − graph . data . 2 %>% mutate ( v4=2+3* v1+rnorm ( 2 0 0 , sd = 5 ) )
6.1.5.1
Box-plot
Codes:
g g p l o t ( graph . data . 2 , aes ( x=v2 , y=v4 ) ) +
geom_boxplot ( )
Figure 6.9: Box-plot
Source:RStudio
144
Electronic copy available at: https://ssrn.com/abstract=3863563
6.2. DETAILS ADJUSTING OF GRAPH
6.2
Details adjusting of graph
6.2.1
Title
For example:
graph . data.3< − graph . data . 2 %>% group_by ( v3 ) %>% summarise (m. v4=mean ( v4 ) )
g g p l o t ( graph . data . 3 , aes ( x=v3 , y=m. v4 ) ) +
geom_col ( )
Codes:
graph . data.3< − graph . data . 2 %>% group_by ( v3 ) %>% summarise (m. v4=mean ( v4 ) )
g g p l o t ( graph . data . 3 , aes ( x=v3 , y=m. v4 ) ) +
geom_col ( ) +
l a b s ( t i t l e ="TCD" , s u b t i t l e ="TBS" ) +
theme ( p l o t . t i t l e =element_text ( h j u s t = 0 . 5 ) ,
p l o t . s u b t i t l e =element_text ( h j u s t = 1 ) )
Notes:
a) "hjust=0.5" means the title will in the middle of the graph;
b) "hjust=1" means the subtitle will on the right side of the graph;
c) "hjust=0" can set the title or subtitle to the left side of the graph.
145
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
Figure 6.10: Title adjusting
Source:RStudio
146
Electronic copy available at: https://ssrn.com/abstract=3863563
6.2. DETAILS ADJUSTING OF GRAPH
6.2.2
Axis
For example:
graph . data.3< − graph . data . 2 %>% group_by ( v3 ) %>% summarise (m. v4=mean ( v4 ) )
g g p l o t ( graph . data . 3 , aes ( x=v3 , y=m. v4 ) ) +
geom_col ( )
Codes:
g g p l o t ( graph . data . 3 , aes ( x=v3 , y=m. v4 ) ) +
geom_col ( ) +
l a b s ( t i t l e ="TCD" , s u b t i t l e ="TBS" , x="Month " , y =" P u b l i c a t i o n s " ) +
s c a l e _ y _ c o n t i n u o u s ( l i m i t s =c ( 0 , 8 0 ) ,
breaks=seq ( from =0 , t o =75 , by =15))+
s c a l e _ x _ d i s c r e t e ( breaks=month . name ,
l a b e l s =paste ( " nnnnew_ " , month . name , sep = " " ) ) +
theme ( p l o t . t i t l e =element_text ( h j u s t = 0 . 5 ) ,
p l o t . s u b t i t l e =element_text ( h j u s t = 1 ) ,
a x i s . t e x t . x=element_text ( angle =15 , h j u s t = 1 ) )
Notes:
a) "scale_y_continuous(limits=c(0,80)" means set 80 as the totally value of the y axis.
b) "breaks=seq(from=0,to=75,by=15))" means set the scale of the y axis from 0 to 75, and
divide by 15.
c) "labels=paste("nnnnew_",month.name,sep=""))" means add "nnnnew_"to each month name
on the x axis.
d) "axis.text.x=element_text(angle=15, hjust=1))" means arrange the new names on the x axis
as 15 o angle, and also keep the unit 1’s distance with the x axis.
147
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
Figure 6.11: Axis adjusting
Source:RStudio
148
Electronic copy available at: https://ssrn.com/abstract=3863563
6.2. DETAILS ADJUSTING OF GRAPH
6.2.3
Legend
For example:
g g p l o t ( graph . data . 2 , aes ( x=v1 , y=v4 , c o l o r =v2 ) ) +
geom_point ( )
Figure 6.12: Scatter graph example
Source:RStudio
149
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
Codes:
g g p l o t ( graph . data . 2 , aes ( x=v1 , y=v4 , c o l o r =v2 ) ) +
geom_point ( ) +
l a b s ( c o l o r =" group " ) +
scale_color_discrete ( labels=letters [1:5])+
theme ( legend . p o s i t i o n =c ( 1 , 0 ) ,
legend . j u s t i f i c a t i o n =c ( 1 , 0 ) ,
legend . background=element_blank ( ) ,
legend . key=element_blank ( ) )
Notes:
a) "labs(color="group")" can transfer the name of v2 from color to group.
b) "scale_color_discrete(labels=letters[1:5])" can transfer A, B, C,D, E to a, b, c, d, e.
c) "theme(legend.position=c(1,0)" and "legend.justification=c(1,0) help to move the legend to
the lower right."
d) "legend.background=element_blank()" and "legend.key=element_blank())" help the legend
to keep the same background color with the entire graph.
Figure 6.13: Legend adjusting
Source:RStudio
150
Electronic copy available at: https://ssrn.com/abstract=3863563
6.2. DETAILS ADJUSTING OF GRAPH
6.2.4
Background
For example:
g g p l o t ( graph . data . 2 , aes ( x=v1 , y=v4 , c o l o r =v2 ) ) +
geom_point ( )
Figure 6.14: Scatter graph example
Source:RStudio
151
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 6. DATA VISUALIZATION OF R - GGPLOT2
Codes:
g g p l o t ( graph . data . 2 , aes ( x=v1 , y=v4 , c o l o r =v2 ) ) +
geom_point ( ) +
theme ( panel . g r i d . major=element_blank ( ) ,
panel . g r i d . minor=element_blank ( ) ,
panel . background=element_rect ( f i l l =" white " ) ,
a x i s . l i n e =element_line ( c o l o r =" black " ) )
Notes:
a) "theme(panel.grid.major=element_blank()"and"panel.grid.minor=element_blank()" can delete
the major panel grid and minor panel grid.
b) "panel.background=element_rect(fill="white")" help to change the background color to
white.
c) "axis.line=element_line(color="black"))" gives the black color to axis.
Figure 6.15: Background adjusting
Source:RStudio
152
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
7
FINANCIAL ECONOMETRICS ANALYSIS OF R
7.1
Analysis of Variance
For example:
cholesterol
7.1.1
Install and load all the packages
i n s t a l l . packages ( " mvtnorm " )
i n s t a l l . packages ( "TH. data " )
i n s t a l l . packages ( "MASS" )
i n s t a l l . packages ( " s y r v i v a l " )
i n s t a l l . packages ( " multcomp " )
i n s t a l l . packages ( " t i d y v e r s e " )
i n s t a l l . packages ( " car " )
i n s t a l l . packages ( " carData " )
l i b r a r y ( mvtnorm )
library ( survival )
l i b r a r y (MASS)
l i b r a r y (TH. data )
l i b r a r y ( multcomp )
library ( tidyverse )
l i b r a r y ( carData )
153
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
l i b r a r y ( car )
154
Electronic copy available at: https://ssrn.com/abstract=3863563
7.1. ANALYSIS OF VARIANCE
7.1.2
Data set
Codes:
c h o l e s t e r o l %>% group_by ( t r t ) %>% summarise (m. resp=mean ( response ) ,
s . resp=sd ( response ) )
Figure 7.1: The results of relevant descriptive data
Source:RStudio
155
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.1.3
Normal distribution testing
Codes:
qqPlot ( lm ( response~ t r t , data= c h o l e s t e r o l ) ,
simulate=TRUE, main="Q−Q P l o t " , l a b e l s =FALSE)
Figure 7.2: The results of type codes and run it
Source:RStudio
Figure 7.3: The results of the normal distribution testing
Source:RStudio
Notes: It means samples comply with the normal distribution, because of all the black circles in
the area between the two blue dotted lines.
156
Electronic copy available at: https://ssrn.com/abstract=3863563
7.1. ANALYSIS OF VARIANCE
7.1.4
Homogeneity of variance testing
Codes:
b a r t l e t t . t e s t ( response~ t r t , data= c h o l e s t e r o l )
Figure 7.4: The results of the homogeneity of variance testing
Source:RStudio
Notes: the p-value = 0.9653 > 5%, which means the null hypothesis has a 96.53% possibility will
happen, and it can not reject the null hypothesis. Therefore, the homogeneity of variance testing
is valid.
157
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.1.5
One way analysis of variance
Codes:
summary ( aov ( response~ t r t , data= c h o l e s t e r o l ) )
Figure 7.5: The results of one way analysis of variance
Source:RStudio
Notes: the significance level = P r(< F) = 9.82 × 10−13 < 5% < 1%, it means when take different
treatments, the responses will dramatically different. Different treatments have a great impact
on the responses.
158
Electronic copy available at: https://ssrn.com/abstract=3863563
7.2. LINEAR REGRESSION
7.2
Linear Regression
For example:
data . set <−data . frame ( x1=sample ( 1 : 5 0 , 2 0 , r e p l a c e =F ) ,
x2=sample ( 3 : 3 0 , 2 0 , r e p l a c e =F ) )
data . set <−data . s e t %>% mutate ( y=x1 + 0 . 5 * ( x1 )^2+0.002 * x2+rnorm ( 2 0 ) )
7.2.1
Data set
Codes:
library ( tidyverse )
s e t . seed ( 6 6 )
data . set <−data . frame ( x1=sample ( 1 : 5 0 , 2 0 , r e p l a c e =F ) ,
x2=sample ( 3 : 3 0 , 2 0 , r e p l a c e =F ) )
data . set <−data . s e t %>% mutate ( y=x1 + 0 . 5 * ( x1 )^2+0.002 * x2+rnorm ( 2 0 ) )
Figure 7.6: The results of type codes and run it
Source:RStudio
159
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.2.2
Model.1 and Model.1 results summary
Codes:
model.1< −lm ( y~x1+x2 , data=data . s e t )
summary ( model . 1 )
Figure 7.7: The results summary of model 1
Source:RStudio
Notes:
a) P r (> | t|) of X 1 = 9.7 × 10−11 , the result is significant.
P r (> | t|) of X 2 = 0.73020, the result is not significant.
b) Multiple R-squared = 0.92, it means the general linear regression model can explain 92%
part of y.
c) "F-statistic" can explain the whole general linear regression model. The null hypothesis is
X 1 = X 2 = 0, and the coefficient of X 1 = X 2 = 0. The p-value of F-statistic= 4.723 × 1010 <
5% < 1%, there was not significant between X 1 and X 2 . Therefore, the results reject the
null hypothesis and the null hypothesis is not valid. In other words, it means the model is
significant.
160
Electronic copy available at: https://ssrn.com/abstract=3863563
7.2. LINEAR REGRESSION
7.2.3
Regression diagnostic
Codes:
par ( mfrow=c ( 2 , 2 ) )
p l o t ( model . 1 )
Notes: par(mfrow=c(2,2)) can arrange the plots in one interface as 2 × 2.
Figure 7.8: The results of regression diagnostic of model.1
Source:RStudio
Notes:
a) Residuals vs Fitted shows the linear diagnostic results. The standard situation is the
red line is a flat line, and the blank circle distributes along with the red line. As for this
example, the red line shows an up tendency. It means this model breaks some assumptions.
It should add some high-order coefficient to adjust this model.
b) Normal Q-Q shows the normal distribution diagnostic results. Blank circle distributes
along with the dotted line means there are no mistakes on the normal distribution diagnostic results.
c) Scale-Location shows the Heteroscedasticity diagnostic results. The standard situation
is the blank circle distribute along with the red line without the outlier. For this model,
161
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
on the left of the graph, the blank circle distribute in a smaller range. But on the right of
the graph, the black circle distribute in a larger range. It means the fluctuation on the left
part is smaller than the right part. As a result, the samples have heteroscedasticity. The
variance of different samples in different values has a significant difference.
d) Residuals vs Leverage shows the outliers about some samples on some values.
162
Electronic copy available at: https://ssrn.com/abstract=3863563
7.2. LINEAR REGRESSION
7.2.4
Model optimization
7.2.4.1
Step one
codes:
s e t . seed ( 6 6 )
model.2< −lm ( y~x1+I ( x1^2)+x2 , data=data . s e t )
summary ( model . 2 )
p l o t ( model . 2 )
Figure 7.9: The results summary of model.2
Source:RStudio
Notes:
a) Add high-order coefficient 2 to X 1 .
b) P r(> | t|) of X 1 = 6.89 × 10−11 , the result is significant.
P r(> | t|) of X 2 = 0.913, the result is not significant.
c) Multiple R-squared = 1, it means the general linear regression model can explain 100%
part of y.
d) "F-statistic" can explain the whole general linear regression model. The null hypothesis
is X 1 = X 2 = 0, the coefficient of X 1 = X 2 = 0 The p-value of F-statistic=2.2 × 10−16 < 5% <
163
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
1%, there was not significant between X 1 and X 2 . Therefore, the results reject the null
hypothesis and the null hypothesis is not valid. In other words, it means the model is
significant.
Figure 7.10: The results of regression diagnostic of Model.2
Source:RStudio
Notes: In Residuals vs Fitted graph, the linear diagnostic results show the blank circle distributes
along with a nearly flat line. Therefore, the high-order coefficient of I(X 1 )2 successfully optimize
the model.
164
Electronic copy available at: https://ssrn.com/abstract=3863563
7.2. LINEAR REGRESSION
7.2.4.2
Step two
Now, the only problem with this model is that X 2 is not significant. Therefore, stepwise regression
analysis is applied.
codes:
model.3< − step ( model . 2 )
summary ( model . 3 )
p l o t ( model . 3 )
Figure 7.11: The results of stepwise regression
Source:RStudio
Notes: The smaller the AIC, the better the model. When the X 2 is deleted, the model has the
smallest AIC. Therefore, y = x1 + I(x12 ) is the final optimal model.
165
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
Figure 7.12: The results summary of the final optimal model
Source:RStudio
Figure 7.13: The results of regression diagnostic of final optimal model
Source:RStudio
166
Electronic copy available at: https://ssrn.com/abstract=3863563
7.3. LOGISTIC REGRESSION
7.3
Logistic Regression
For example:
Breast Cancer Data From UCI
7.3.1
Data set
Codes:
breast . cancer . data<−read . csv ( f i l e . choose ( ) )
head ( breast . cancer . data )
breast . cancer . d a t a $ C l a s s i f i c a t i o n <− f a c t o r ( breast . cancer . d a t a $ C l a s s i f i c a t i o n , l e v e l s =c ( 1 ,
Notes: 1 stands for healthy and 2 stands for sicken.
Figure 7.14: The results of breast cancer data import
Source:RStudio
167
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.3.2
Model construction and model analysis summary
Codes:
model . l o g i t <−glm ( C l a s s i f i c a t i o n ~ . , breast . cancer . data , family =" binomial " )
summary ( model . l o g i t )
a) The symbol "." stands for the independent variable.
b) "glm" stands for generalized linear Model.
Figure 7.15: The results of model construction and model analysis summary
Source:RStudio
168
Electronic copy available at: https://ssrn.com/abstract=3863563
7.3. LOGISTIC REGRESSION
7.3.3
Model optimization and optimized model analysis summary
Codes:
step ( model . l o g i t )
model . l o g i t 2 <−glm ( C l a s s i f i c a t i o n ~BMI+Glucose+ I n s u l i n + R e s i s t i n ,
breast . cancer . data , family =" binomial " )
summary ( model . l o g i t 2 )
Figure 7.16: Model optimization and optimized model analysis summary
Source:RStudio
169
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.3.4
Confusion matrix
Codes:
prob<− p r e d i c t ( model . l o g i t 2 , breast . cancer . data , type =" response " )
pred<−as . f a c t o r ( i f e l s e ( prob > = 0 . 5 , " sicken " , " healthy " ) )
confuse . table <− t a b l e ( breast . cancer . d a t a $ C l a s s i f i c a t i o n , pred ,
dnn=c ( " Actual " , " Predicated " ) )
Notes: "dnn" stands for the row name and column name of the newly generated data table.
Figure 7.17: The results of confusion matrix
Source:RStudio
The results of the confusion matrix shows:
a) The model predicts there are 38 healthy cases, but there are 52 healthy cases actually.
b) The model also predicts there are 49 sicken cases, but there are 64 sicken cases actually.
170
Electronic copy available at: https://ssrn.com/abstract=3863563
7.4. HETEROSCEDASTICITY TEST
7.4
Heteroscedasticity Test
For example:
data . set <−data . frame ( x1=sample ( 1 : 5 0 , 2 0 , r e p l a c e =F ) ,
x2=sample ( 3 : 3 0 , 2 0 , r e p l a c e =F ) )
data . set <−data . s e t %>% mutate ( y=x1 + 0 . 5 * ( x1 )^2+0.002 * x2+rnorm ( 2 0 ) )
7.4.1
Data set
Codes:
library ( tidyverse )
s e t . seed ( 6 6 )
data . set <−data . frame ( x1=sample ( 1 : 5 0 , 2 0 , r e p l a c e =F ) ,
x2=sample ( 3 : 3 0 , 2 0 , r e p l a c e =F ) )
data . set <−data . s e t %>% mutate ( y=x1 + 0 . 5 * ( x1 )^2+0.002 * x2+rnorm ( 2 0 ) )
model.1< −lm ( y~x1+x2 , data=data . s e t )
summary ( model . 1 )
Figure 7.18: The results of type codes and run it
Source:RStudio
171
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.4.2
Detecting heteroskedasticity - The Breusch-Pagan Test
Codes:
i n s t a l l . packages ( " zoo " )
i n s t a l l . packages ( " lmtest " )
i n s t a l l . packages ( " b p t e s t " )
l i b r a r y ( zoo )
l i b r a r y ( lmtest )
b p t e s t ( model . 1 )
Notes: A formal, mathematical way of detecting heteroskedasticity is what is known as the
Breusch-Pagan test.
Figure 7.19: The results of the Breusch-Pagan test
Source:RStudio
Notes: While it doesn’t give the critical value to compare the test statistic, all testers need to look
at is the p-value to determine whether or not testers should reject the null. If the p-value is less
than the level of significance (in this case if the p-value is less than 5%), then testers reject the
null hypothesis. Since 0.4705 < 0.05, the null hypothesis is rejected.
172
Electronic copy available at: https://ssrn.com/abstract=3863563
7.5. CORRELATION TEST
7.5
Correlation Test
For example:
i n s t a l l . packages ( " pwr " )
l i b r a r y ( pwr )
pwr . r . t e s t ( n= , r = , s i g . l e v e l = , power = , a l t e r n a t i v e =)
Notes:
a) The function of pwr.r.test() can do a power analysis to correlation analysis.
b) n stands for the number of observations;
r stands for the effect size (Measure by linear correlation coefficient);
sig.level stands for the significance level;
power stands for the power level;
alternative points the significance test is a single side hypothesis test or a two-tail hypothesis test.
For example:
The relationship between depression and solitary.
The null hypothesis and research hypothesis are:
H0 = P ≤ 0.25
H1 > 0.25
Notes:
a) P is the overall correlation between these two psychology variables.
b) The significance level is set as 0.05, and if H0 is invalid, the confident to refuse H0 is 90%.
How many observation cases are this research need? The answers are as follows.
173
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
Figure 7.20: How many observation cases are this research need
Source:RStudio
As a result, for meeting the above requirements, 133.2803 ≈ 134 observations cases are needed to
evaluate the relationship between depression and solitary, so that can refuse the null hypothesis
in a 90% confidence when it is false.
174
Electronic copy available at: https://ssrn.com/abstract=3863563
7.5. CORRELATION TEST
7.5.1
Autocorrelation test_Durbin - Watson Test
For example:
X: Family Income/per year
Y: Family Food Consumption/per year
7.5.1.1
Preliminary regression
Codes:
l i b r a r y ( readxl )
df<− r e a d _ e x c e l ( f i l e . choose ( ) )
lm . df <− lm (Y ~ X, data = df )
summary ( lm . df )
Figure 7.21: The results of preliminary regression
Source:RStudio
Notes:
a) According to the preliminary regression results, P r(> | t|) = 0.00967 < 5% < 1%. It means
different family income has a different influence on family food consumption. This model is
significance.
b) But the Std. Error of X is too small, just equal to 0.01588. Therefore, this model could have
an autocorrelation character and should be further tested.
175
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.5.1.2
Autocorrelation test_Durbin-Watson Test
Codes:
i n s t a l l . packages ( " zoo " )
l i b r a r y ( ’ lmtest ’ )
dwtest ( lm . df )
Figure 7.22: The results of Durbin Watson Test
Source:RStudio
Notes: R can not calculate the function of Durbin-Watson Test’s critical value. But based on the
experience, when 0<DW<1.8, the model is a positive autocorrelation; when 2.2<DW<4, the model
is a negative autocorrelation.
According to the results of the Durbin-Watson Test above, DW = 1.4912, and the null hypothesis
should be rejected. Therefore, family income has a positive autocorrelation relationship with
family food consumption.
176
Electronic copy available at: https://ssrn.com/abstract=3863563
7.6. COINTEGRATION TEST
7.6
Cointegration Test
7.6.1
Install and library the packages
Codes:
i n s t a l l . packages ( " t s e r i e s " )
i n s t a l l . packages ( " lmtest " )
l i b r a r y ( lmtest )
library ( tseries )
Figure 7.23: The results of install and library the packages
Source:RStudio
177
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.6.2
Set up a regression equation with two variables ( y1 , y2 )
Notes: Applying the unit root test to the residual (resid) of this regression equation. The null
hypothesis is that these two variables do not have the relationship of cointegration, and the
alternative hypothesis is that these two variables have the relationship of cointegration. Applying
the Least Squares Method to estimate this regression equation and extract the residuals from
this regression equation to test.
Codes:
s e t . seed ( 1 2 3 )
a1 <− rnorm ( 5 0 0 )
a2 <− rnorm ( 5 0 0 )
y1 <− cumsum( a1 )
y2 <− 0 . 8 * y1 + a2
Figure 7.24: The results of setting up a regression equation with two variables
Source:RStudio
178
Electronic copy available at: https://ssrn.com/abstract=3863563
7.6. COINTEGRATION TEST
7.6.3
Engle-Granger Two Steps cointegration test
7.6.3.1
Set up the spurious regression
Codes:
lmod = lm ( y2 ~ y1 )
summary ( lmod )
Figure 7.25: The results of setting up the spurious regression
Source:RStudio
Results:
> lmod = lm ( y2 ~ y1 )
> lmod
Call :
lm ( formula = y2 ~ y1 )
Coefficients :
( Intercept )
y1
0.01522
0.79617
179
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
> summary ( lmod )
Call :
lm ( formula = y2 ~ y1 )
Residuals :
Min
1Q
Median
3Q
Max
− 2.8033 − 0.6848 − 0.0024
0.6331
2.7162
Coefficients :
Estimate Std . Error t value Pr( >| t |)
( I n t e r c e p t ) 0.015216
0.062105
0.245
0.807
y1
0.009294
85.662
<2e −16 ***
0.796166
−−−
S i g n i f . codes :
0 ’ * * * ’ 0.001 ’ * * ’ 0.01 ’ * ’ 0.05
’ . ’ 0.1 ’ ’ 1
Residual standard e r r o r : 1.012 on 498 degrees o f freedom
Multiple R−squared :
F− s t a t i s t i c :
0.9364 ,
Adjusted R−squared :
7338 on 1 and 498 DF,
0.9363
p−value : < 2 . 2 e −16
Multiple R-squared is equal to 0.9364, the Multiple R-squared is significance.
180
Electronic copy available at: https://ssrn.com/abstract=3863563
7.6. COINTEGRATION TEST
7.6.3.2
Extract residuals
Codes:
e r r o r = r e s i d u a l s ( lmod )
Figure 7.26: The results of extracting residuals
Source:RStudio
The cointegration regression equation can be denoted as Equation 7.1:
(7.1)
ln y2 = 0.015216 + 0.796166ln y1 + ε t .
181
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.6.3.3
Unit root test
Codes:
adf . t e s t ( e r r o r )
e r r o r . lagged = e r r o r [− c ( 9 9 , 1 0 0 ) ]
Figure 7.27: The results of unit root test
Source:RStudio
7.6.3.4
Build the Error Correlation Model (ECM.REG)
Codes:
dy1 = d i f f ( y1 )
dy2 = d i f f ( y2 )
d i f f . dat = data . frame ( embed ( cbind ( dy1 , dy2 ) , 2 ) )
colnames ( d i f f . dat ) = c ( " dy1 " , " dy2 " , " dy1 . 1 " , " dy2 . 1 " )
ecm . reg = lm ( dy2 ~ e r r o r . lagged + dy1 . 1 + dy2 . 1 , data = d i f f . dat )
summary ( ecm . reg )
182
Electronic copy available at: https://ssrn.com/abstract=3863563
7.6. COINTEGRATION TEST
Figure 7.28: The results of ECM.REG
Source:RStudio
Results:
> ecm . reg = lm ( dy2 ~ e r r o r . lagged + dy1 . 1 + dy2 . 1 , data = d i f f . dat )
> ecm . reg
Call :
lm ( formula = dy2 ~ e r r o r . lagged + dy1 . 1 + dy2 . 1 , data = d i f f . dat )
Coefficients :
( Intercept )
e r r o r . lagged
dy1 . 1
dy2 . 1
0.03482
0.73294
0.36319
− 0.44189
> summary ( ecm . reg )
Call :
lm ( formula = dy2 ~ e r r o r . lagged + dy1 . 1 + dy2 . 1 , data = d i f f . dat )
Residuals :
Min
1Q
Median
3Q
Max
183
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
− 3.9671 − 0.7942
0.0011
0.7468
4.1695
Coefficients :
Estimate Std . Error t value Pr( >| t |)
( Intercept )
0.03482
0.05585
0.623
0.533
e r r o r . lagged
0.73294
0.05605
13.076
< 2e −16 ***
dy1 . 1
0.36319
0.06458
5.624 3.12 e −08 ***
dy2 . 1
− 0.44189
0.03929 − 11.246
< 2e −16 ***
−−−
S i g n i f . codes :
0 ’ * * * ’ 0.001 ’ * * ’ 0.01 ’ * ’ 0.05
’ . ’ 0.1 ’ ’ 1
Residual standard e r r o r : 1.246 on 494 degrees o f freedom
Multiple R−squared :
0.4078 ,
Adjusted R−squared :
F− s t a t i s t i c : 113.4 on 3 and 494 DF,
0.4042
p−value : < 2 . 2 e −16
According to the output results of the summary (ecm.reg) function, the Error Correlation Model
can be expressed as Equation 7.2:
(7.2)
∆ ln y2 t = 0.03482 + 0.73294ecm t−1 + 0.36319∆ ln y1 t − 0.44189∆ ln y2 t + ε t .
184
Electronic copy available at: https://ssrn.com/abstract=3863563
7.6. COINTEGRATION TEST
7.6.3.5
Visualization
Codes:
par ( mfrow = c ( 2 , 2 ) )
p l o t ( ecm . reg )
Figure 7.29: The results of visualization
Source:RStudio
Figure 7.30: Zoom in the visualization results
Source:RStudio
185
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
Notes: Applying Diagnostic Plot to test the Error Correlation Model.
a) Residuals vs Fitted: residuals uniformly distributed.
b) Q-Q plot: errors are normal distribution.
c) Scale – Location: the range of the error distribution can be observed.
d) Residuals vs Leverage: outlier and leverage point can be tested.
Based on these four plots, the reliability of the Error Correlation Model can be tested. After
testing, this Error Correlation Model can be judged as reliable.
186
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7
ARCH Model and GARCH Model
7.7.1
Step one: set a working directory
Codes:
wd <− "C:\\ Users\\Lenovo\\Desktop \\"
setwd (wd)
Figure 7.31: The results of set the working directory
Source:RStudio
187
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
Figure 7.32: The location of the working directory
Source:RStudio
Notes:Please type your own location.
188
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7.2
Step two: install and library the packages
Codes:
i n s t a l l . packages ( " FinTS " )
i n s t a l l . packages ( "MTS" )
i n s t a l l . packages ( " rmgarch " )
i n s t a l l . packages ( " D i s t r i b u t i o n U t i l s " )
i n s t a l l . packages ( " rugarch " )
i n s t a l l . packages ( " moments " )
i n s t a l l . packages ( " p o r t e s " )
i n s t a l l . packages ( " normtest " )
i n s t a l l . packages ( " x l s x " )
i n s t a l l . packages ( " t s e r i e s " )
i n s t a l l . packages ( " urca " )
i n s t a l l . packages ( " zoo " )
i n s t a l l . packages ( " f o r e c a s t " )
i n s t a l l . packages ( " vars " )
i n s t a l l . packages ( " p a r a l l e l " )
i n s t a l l . packages ( " ggplot2 " )
l i b r a r y ( FinTS )
l i b r a r y ( rugarch )
l i b r a r y ( moments )
library ( portes )
l i b r a r y ( normtest )
library ( xlsx )
library ( tseries )
l i b r a r y ( urca )
l i b r a r y ( zoo )
library ( forecast )
l i b r a r y ( vars )
l i b r a r y (MTS)
library ( parallel )
l i b r a r y ( rmgarch )
l i b r a r y ( ggplot2 )
l i b r a r y ( readxl )
189
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.3
Step three: Load the data into RStudio
Codes:
Data <− read . csv ( f i l e . choose ( ) , header = T , sep = " , " , dec = " . " )
time <− Data [ , 1 ]
Figure 7.33: The results of load the data into RStudio
Source:RStudio
190
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7.4
Step four: calculate log returns
Codes:
data <− apply ( l o g ( as . data . frame ( cbind ( Data [ , 2 ] , Data [ , 3 ] ) ) ) , 2 , d i f f )
rownames ( data ) <− time [ − 1]
colnames ( data ) <− c ( " Crude Oil " , " SP500 " )
Figure 7.34: The results of calculate log returns
Source:RStudio
191
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.5
Step five: descriptive statistics
Codes:
p l o t ( as . Date ( time ) , Data$Crude . Oil , type =" l " , c o l =" blue " , lwd =2 ,
ylab ="Crude Oil " , main=" Daily c l o s i n g p r i c e s " )
p l o t ( as . Date ( time ) , Data$SP500 , type =" l " , c o l =" blue " , lwd =2 ,
ylab ="SP500 " , main=" Daily c l o s i n g p r i c e s " )
summary ( data )
sd ( data [ , 1 ] )
sd ( data [ , 2 ] )
skewness ( data [ , 1 ] )
skewness ( data [ , 2 ] )
k u r t o s i s ( data [ , 1 ] )
k u r t o s i s ( data [ , 2 ] )
Figure 7.35: The results of descriptive statistics
Source:RStudio
192
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
Results:
> p l o t ( as . Date ( time ) , Data$Crude . Oil , type =" l " , c o l =" blue " , lwd =2 ,
+ ylab ="Crude Oil " , main=" Daily c l o s i n g p r i c e s " )
> p l o t ( as . Date ( time ) , Data$SP500 , type =" l " , c o l =" blue " , lwd =2 ,
+ ylab ="SP500 " , main=" Daily c l o s i n g p r i c e s " )
> summary ( data )
Crude Oil SP500
Min . : − 0.2822061 Min . : − 0.1276522
1 s t Qu.: − 0.0131699 1 s t Qu.: − 0.0029184
Median : 0.0008961 Median : 0.0005561
Mean : − 0.0007245 Mean : 0.0001819
3rd Qu . : 0.0121698 3rd Qu . : 0.0046751
Max . : 0.2182347 Max . : 0.0896832
> sd ( data [ , 1 ] )
[ 1 ] 0.02669878
> sd ( data [ , 2 ] )
[ 1 ] 0.0114489
> skewness ( data [ , 1 ] )
[ 1 ] − 1.226625
> skewness ( data [ , 2 ] )
[ 1 ] − 1.374074
> k u r t o s i s ( data [ , 1 ] )
[ 1 ] 25.74864
> k u r t o s i s ( data [ , 2 ] )
[ 1 ] 29.79904
193
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.6
Step six: Ljung box test on returns
Codes:
Box . t e s t ( data [ , 1 ] , lag = 10 , type = c ( " Box−P i e r c e " , " Ljung−Box " ) )
Box . t e s t ( data [ , 2 ] , lag = 10 , type = c ( " Box−P i e r c e " , " Ljung−Box " ) )
Figure 7.36: The results of Ljung box test on returns
Source:RStudio
194
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7.7
Step seven: Ljung box test on squared returns
Codes:
Box . t e s t ( data [ , 1 ] ^ 2 , lag = 10 , type = c ( " Box−P i e r c e " , " Ljung−Box " ) )
Box . t e s t ( data [ , 2 ] ^ 2 , lag = 10 , type = c ( " Box−P i e r c e " , " Ljung−Box " ) )
Figure 7.37: The results of Ljung box test on squared returns
Source:RStudio
195
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.8
Step eight: ARCH test on returns
Codes:
ArchTest ( data [ , 1 ] , l a g s = 10 , demean = TRUE)
ArchTest ( data [ , 2 ] , l a g s = 10 , demean = TRUE)
Figure 7.38: The results of ARCH test on returns
Source:RStudio
196
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7.9
Step nine: defining the GARCH model
Codes:
f i t . spec <− ugarchspec ( variance . model = l i s t ( model = "sGARCH" , garchOrder =
c (1 , 1)) ,
mean . model = l i s t ( armaOrder = c ( 0 , 0 ) ,
i n c l u d e . mean = TRUE) ,
d i s t r i b u t i o n . model = "norm " )
Figure 7.39: The results of defining the GARCH model
Source:RStudio
Results:
> f i t . spec <− ugarchspec ( variance . model = l i s t ( model = "sGARCH" , gar
chOrder = c ( 1 , 1 ) ) ,
+ mean . model = l i s t ( armaOrder = c ( 0 , 0 ) ,
+ i n c l u d e . mean = TRUE) ,
+ d i s t r i b u t i o n . model = "norm " )
> f i t . spec *−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−* * GARCH Model Spec *
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Conditional Variance Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
GARCH Model : sGARCH( 1 , 1 )
197
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
Variance Targeting : FALSE
Conditio nal Mean Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Mean Model : ARFIMA( 0 , 0 , 0 )
Include Mean : TRUE
R T u t o r i a l Guidance f o r BU7510 F i n a n c i a l Econometrics
27
GARCH−in−Mean : FALSE
Conditio nal D i s t r i b u t i o n
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
D i s t r i b u t i o n : norm
Includes Skew : FALSE
Includes Shape : FALSE
Includes Lambda : FALSE
198
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7.10
Step ten: fitting the GARCH model
Codes:
u g a r c h f i t ( data = data [ , 1 ] , spec = f i t . spec )
u g a r c h f i t ( data = data [ , 2 ] , spec = f i t . spec )
Figure 7.40: The results of fitting the GARCH model
Source:RStudio
Results:
> u g a r c h f i t ( data = data [ , 1 ] , spec = f i t . spec )
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−* * GARCH Model F i t *
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Conditional Variance Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
GARCH Model : sGARCH( 1 , 1 )
Mean Model : ARFIMA( 0 , 0 , 0 )
D i s t r i b u t i o n : norm
Optimal Parameters
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Estimate Std . Error t value Pr( >| t |)
mu 0.000508 0.000566 0.89828 0.369035
199
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
omega 0.000021 0.000008 2.63534 0.008405
alpha1 0.128990 0.022793 5.65907 0.000000
beta1 0.844323 0.031233 27.03340 0.000000
Robust Standard Errors :
Estimate Std . Error t value Pr( >| t |)
mu 0.000508 0.000656 0.77468 0.438530
omega 0.000021 0.000014 1.47372 0.140558
alpha1 0.128990 0.062830 2.05300 0.040073
beta1 0.844323 0.072794 11.59874 0.000000
LogLikelihood : 2978.062
Information C r i t e r i a
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Akaike − 4.7661
Bayes − 4.7497
Shibata − 4.7661
Hannan−Quinn − 4.7599
Weighted Ljung−Box Test on Standardized Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ] 0.01233 0.9116
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 2 ] 0.27219 0.8114
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 5 ] 2.11549 0.5914
d . o . f =0
H0 : No s e r i a l c o r r e l a t i o n
Weighted Ljung−Box Test on Standardized Squared Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ] 3.809 0.05097
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 5 ] 5.103 0.14509
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 9 ] 7.398 0.16809
d . o . f =2
Weighted ARCH LM Tests
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
S t a t i s t i c Shape Scale P−Value
ARCH Lag [ 3 ] 0.4499 0.500 2.000 0.5024
ARCH Lag [ 5 ] 1.5315 1.440 1.667 0.5841
ARCH Lag [ 7 ] 2.2526 2.315 1.543 0.6639
Nyblom s t a b i l i t y t e s t
200
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
J o i n t S t a t i s t i c : 0.5733
Individual S t a t i s t i c s :
mu 0.05411
omega 0.20552
alpha1 0.31157
beta1 0.20089
Asymptotic C r i t i c a l Values (10% 5% 1%)
J o i n t S t a t i s t i c : 1.07 1.24 1 . 6
I n d i v i d u a l S t a t i s t i c : 0.35 0.47 0.75
Sign Bias Test
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
t−value prob s i g
Sign Bias 0.4219 0.67319
Negative Sign Bias 2.2015 0.02788 **
P o s i t i v e Sign Bias 0.7501 0.45335
J o i n t E f f e c t 11.1610 0.01089 **
Adjusted Pearson Goodness−of −F i t Test :
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
group s t a t i s t i c p−value ( g −1)
1 20 62.54 1.521 e −06
2 30 69.12 4.003 e −05
3 40 81.10 8.799 e −05
4 50 84.45 1.236 e −03
Elapsed time : 0.6327131
> u g a r c h f i t ( data = data [ , 2 ] , spec = f i t . spec )
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−* * GARCH Model F i t *
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Conditional Variance Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
GARCH Model : sGARCH( 1 , 1 )
Mean Model : ARFIMA( 0 , 0 , 0 )
D i s t r i b u t i o n : norm
Optimal Parameters
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Estimate Std . Error t value Pr( >| t |)
mu 0.000856 0.000159 5.3821 0.000000
201
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
omega 0.000004 0.000001 3.7991 0.000145
alpha1 0.245249 0.023547 10.4154 0.000000
beta1 0.719919 0.018490 38.9365 0.000000
Robust Standard Errors :
Estimate Std . Error t value Pr( >| t |)
mu 0.000856 0.000251 3.4120 0.000645
omega 0.000004 0.000004 1.1672 0.243133
alpha1 0.245249 0.051328 4.7780 0.000002
beta1 0.719919 0.054898 13.1138 0.000000
LogLikelihood : 4315.897
Information C r i t e r i a
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Akaike − 6.9101
Bayes − 6.8937
Shibata − 6.9101
Hannan−Quinn − 6.9039
Weighted Ljung−Box Test on Standardized Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ] 0.1547 0.6941
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 2 ] 0.2249 0.8391
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 5 ] 0.8582 0.8909
d . o . f =0
H0 : No s e r i a l c o r r e l a t i o n
Weighted Ljung−Box Test on Standardized Squared Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ] 0.8872 0.3462
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 5 ] 1.7651 0.6749
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 9 ] 2.7806 0.7946
d . o . f =2
Weighted ARCH LM Tests
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
S t a t i s t i c Shape Scale P−Value
ARCH Lag [ 3 ] 0.00297 0.500 2.000 0.9565
ARCH Lag [ 5 ] 1.03100 1.440 1.667 0.7237
ARCH Lag [ 7 ] 1.72138 2.315 1.543 0.7760
Nyblom s t a b i l i t y t e s t
202
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
J o i n t S t a t i s t i c : 3.4283
Individual S t a t i s t i c s :
mu 0.2518
omega 0.8010
alpha1 0.2251
beta1 0.1562
Asymptotic C r i t i c a l Values (10% 5% 1%)
J o i n t S t a t i s t i c : 1.07 1.24 1 . 6
I n d i v i d u a l S t a t i s t i c : 0.35 0.47 0.75
Sign Bias Test
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
t−value prob s i g
Sign Bias 2.81152 0.005008 ***
Negative Sign Bias 0.02397 0.980884
P o s i t i v e Sign Bias 0.04684 0.962645
J o i n t E f f e c t 12.05221 0.007206 ***
Adjusted Pearson Goodness−of −F i t Test :
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
group s t a t i s t i c p−value ( g −1)
1 20 83.99 3.785 e −10
2 30 113.06 7.203 e −12
3 40 116.62 1.152 e −09
4 50 137.98 2.117 e −10
Elapsed time : 0.3072
203
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.11
Step eleven: manually compute the Ljung box and ARCH test
Codes:
c r u d e _ o i l <− u g a r c h f i t ( data = data [ , 1 ] , spec = f i t . spec )
sp500 <− u g a r c h f i t ( data = data [ , 2 ] , spec = f i t . spec )
Box . t e s t ( as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize=TRUE) ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
Box . t e s t ( as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize=TRUE) ^ 2 ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
ArchTest ( as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize=TRUE) ) , l a g s = 20)
Box . t e s t ( as . numeric ( r e s i d u a l s ( sp500 , standardize=TRUE) ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
Box . t e s t ( as . numeric ( r e s i d u a l s ( sp500 , standardize=TRUE) ^ 2 ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
ArchTest ( as . numeric ( r e s i d u a l s ( sp500 , standardize=TRUE) ) , l a g s = 20)
Figure 7.41: The results of manually compute the Ljung Box and ARCH tests
Source:RStudio
204
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
Results:
> crude_oil
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
GARCH Model F i t
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Conditional Variance Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
GARCH Model
: sGARCH( 1 , 1 )
Mean Model
: ARFIMA( 0 , 0 , 0 )
Distribution
: norm
Optimal Parameters
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Estimate
Std . Error
t value Pr( >| t |)
mu
0.000508
0.000566
0.89828 0.369035
omega
0.000021
0.000008
2.63534 0.008405
alpha1
0.128990
0.022793
5.65907 0.000000
beta1
0.844323
0.031233 27.03340 0.000000
Robust Standard Errors :
Estimate
Std . Error
t value Pr( >| t |)
mu
0.000508
0.000656
0.77468 0.438530
omega
0.000021
0.000014
1.47372 0.140558
alpha1
0.128990
0.062830
2.05300 0.040073
beta1
0.844323
0.072794 11.59874 0.000000
LogLikelihood : 2978.062
Information C r i t e r i a
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Akaike
− 4.7661
Bayes
− 4.7497
Shibata
− 4.7661
Hannan−Quinn − 4.7599
205
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
Weighted Ljung−Box Test on Standardized Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ]
0.01233
0.9116
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 2 ]
0.27219
0.8114
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 5 ]
2.11549
0.5914
d . o . f =0
H0 : No s e r i a l c o r r e l a t i o n
Weighted Ljung−Box Test on Standardized Squared Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ]
3.809 0.05097
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 5 ]
5.103 0.14509
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 9 ]
7.398 0.16809
d . o . f =2
Weighted ARCH LM Tests
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
S t a t i s t i c Shape Scale P−Value
ARCH Lag [ 3 ]
0.4499 0.500 2.000
0.5024
ARCH Lag [ 5 ]
1.5315 1.440 1.667
0.5841
ARCH Lag [ 7 ]
2.2526 2.315 1.543
0.6639
Nyblom s t a b i l i t y t e s t
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Joint S t a t i s t i c :
0.5733
Individual S t a t i s t i c s :
mu
0.05411
omega
0.20552
alpha1 0.31157
beta1
0.20089
Asymptotic C r i t i c a l Values (10% 5% 1%)
Joint S t a t i s t i c :
1.07 1.24 1 . 6
Individual S t a t i s t i c :
0.35 0.47 0.75
Sign Bias Test
206
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
t−value
prob s i g
Sign Bias
0.4219 0.67319
Negative Sign Bias
2.2015 0.02788
P o s i t i v e Sign Bias
0.7501 0.45335
Joint Effect
11.1610 0.01089
**
**
Adjusted Pearson Goodness−of −F i t Test :
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
group s t a t i s t i c p−value ( g −1)
1
20
62.54
1.521 e −06
2
30
69.12
4.003 e −05
3
40
81.10
8.799 e −05
4
50
84.45
1.236 e −03
Elapsed time : 0.156615
> sp500 <− u g a r c h f i t ( data = data [ , 2 ] , spec = f i t . spec )
> sp500
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
GARCH Model F i t
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Conditional Variance Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
GARCH Model
: sGARCH( 1 , 1 )
Mean Model
: ARFIMA( 0 , 0 , 0 )
Distribution
: norm
Optimal Parameters
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Estimate
Std . Error
t value Pr( >| t |)
mu
0.000856
0.000159
5.3821 0.000000
omega
0.000004
0.000001
3.7991 0.000145
alpha1
0.245249
0.023547
10.4154 0.000000
207
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
beta1
0.719919
0.018490
38.9365 0.000000
Robust Standard Errors :
Estimate
Std . Error
t value Pr( >| t |)
mu
0.000856
0.000251
3.4120 0.000645
omega
0.000004
0.000004
1.1672 0.243133
alpha1
0.245249
0.051328
4.7780 0.000002
beta1
0.719919
0.054898
13.1138 0.000000
LogLikelihood : 4315.897
Information C r i t e r i a
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Akaike
− 6.9101
Bayes
− 6.8937
Shibata
− 6.9101
Hannan−Quinn − 6.9039
Weighted Ljung−Box Test on Standardized Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ]
0.1547
0.6941
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 2 ]
0.2249
0.8391
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 5 ]
0.8582
0.8909
d . o . f =0
H0 : No s e r i a l c o r r e l a t i o n
Weighted Ljung−Box Test on Standardized Squared Residuals
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
s t a t i s t i c p−value
Lag [ 1 ]
0.8872
0.3462
Lag [ 2 * ( p+q ) + ( p+q ) − 1 ] [ 5 ]
1.7651
0.6749
Lag [ 4 * ( p+q ) + ( p+q ) − 1 ] [ 9 ]
2.7806
0.7946
d . o . f =2
Weighted ARCH LM Tests
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
208
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
S t a t i s t i c Shape Scale P−Value
ARCH Lag [ 3 ]
0.00297 0.500 2.000
0.9565
ARCH Lag [ 5 ]
1.03100 1.440 1.667
0.7237
ARCH Lag [ 7 ]
1.72138 2.315 1.543
0.7760
Nyblom s t a b i l i t y t e s t
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Joint S t a t i s t i c :
3.4283
Individual S t a t i s t i c s :
mu
0.2518
omega
0.8010
alpha1 0.2251
beta1
0.1562
Asymptotic C r i t i c a l Values (10% 5% 1%)
Joint S t a t i s t i c :
1.07 1.24 1 . 6
Individual S t a t i s t i c :
0.35 0.47 0.75
Sign Bias Test
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
t−value
Sign Bias
prob s i g
Negative Sign Bias
2.81152 0.005008 ***
0.02397 0.980884
P o s i t i v e Sign Bias
0.04684 0.962645
Joint Effect
12.05221 0.007206 ***
Adjusted Pearson Goodness−of −F i t Test :
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
group s t a t i s t i c p−value ( g −1)
1
20
83.99
3.785 e −10
2
30
113.06
7.203 e −12
3
40
116.62
1.152 e −09
4
50
137.98
2.117 e −10
Elapsed time : 0.1954758
209
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
> Box . t e s t ( as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize=TRUE) ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
Box−Ljung t e s t
data :
as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize = TRUE) )
X−squared = 15.743 , df = 20 , p−value = 0.7325
> Box . t e s t ( as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize=TRUE) ^ 2 ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
Box−Ljung t e s t
data :
as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize = TRUE) ^ 2 )
X−squared = 16.273 , df = 20 , p−value = 0.6996
> ArchTest ( as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize=TRUE) ) , l a g s = 20)
ARCH LM−t e s t ; Null hypothesis : no ARCH e f f e c t s
data :
as . numeric ( r e s i d u a l s ( c r u d e _ o i l , standardize = TRUE) )
Chi−squared = 16.762 , df = 20 , p−value = 0.6684
> Box . t e s t ( as . numeric ( r e s i d u a l s ( sp500 , standardize=TRUE) ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
Box−Ljung t e s t
data :
as . numeric ( r e s i d u a l s ( sp500 , standardize = TRUE) )
X−squared = 24.689 , df = 20 , p−value = 0.2136
> Box . t e s t ( as . numeric ( r e s i d u a l s ( sp500 , standardize=TRUE) ^ 2 ) , lag =20 ,
type =" Ljung−Box " , f i t d f =0)
Box−Ljung t e s t
data :
as . numeric ( r e s i d u a l s ( sp500 , standardize = TRUE) ^ 2 )
X−squared = 16.219 , df = 20 , p−value = 0.7029
210
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
> ArchTest ( as . numeric ( r e s i d u a l s ( sp500 , standardize=TRUE) ) , l a g s = 20)
ARCH LM−t e s t ; Null hypothesis : no ARCH e f f e c t s
data :
as . numeric ( r e s i d u a l s ( sp500 , standardize = TRUE) )
Chi−squared = 15.514 , df = 20 , p−value = 0.7463
211
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.12
Step twelve: DCC estimation
7.7.12.1
Univariate normal GARCH (1,1) for each series
Codes:
garch11 . spec = ugarchspec ( mean . model = l i s t ( armaOrder = c ( 0 , 0 ) ) ,
variance . model = l i s t ( garchOrder = c ( 1 , 1 ) ,
model = "sGARCH" ) ,
d i s t r i b u t i o n . model = "norm " )
Figure 7.42: The results of univariate normal GARCH (1,1) for each series
Source:RStudio
Results:
> garch11 . spec = ugarchspec ( mean . model = l i s t ( armaOrder = c ( 0 , 0 ) ) ,
+
variance . model = l i s t ( garchOrder = c ( 1 , 1 ) ,
+
model = "sGARCH" ) ,
+
d i s t r i b u t i o n . model = "norm " )
> garch11 . spec
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
GARCH Model Spec
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
212
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
Conditional Variance Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
GARCH Model
: sGARCH( 1 , 1 )
Variance Targeting
: FALSE
Conditional Mean Dynamics
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Mean Model
: ARFIMA( 0 , 0 , 0 )
Include Mean
: TRUE
GARCH−in−Mean
: FALSE
Conditional D i s t r i b u t i o n
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Distribution
:
norm
Includes Skew
:
FALSE
Includes Shape
:
FALSE
Includes Lambda :
FALSE
213
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.12.2
DCC specification - GARCH (1,1) for conditional correlations
Codes:
dcc . garch11 . spec = dccspec ( uspec = multispec ( r e p l i c a t e ( 2 , garch11 . spec ) ) ,
dccOrder = c ( 1 , 1 ) ,
d i s t r i b u t i o n = "mvnorm " )
dcc . garch11 . spec
dcc . f i t = d c c f i t ( dcc . garch11 . spec , data = data [ , 1 : 2 ] )
c l a s s ( dcc . f i t )
slotNames ( dcc . f i t )
names ( dcc . f i t @ m f i t )
names ( dcc . fit@model )
Figure 7.43: DCC specification - GARCH(1,1) for conditional correlations
Source:RStudio
Results:
> dcc . garch11 . spec = dccspec ( uspec = multispec ( r e p l i c a t e ( 2 , garch11 . spec ) ) ,
+
dccOrder = c ( 1 , 1 ) ,
+
d i s t r i b u t i o n = "mvnorm " )
214
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
> dcc . garch11 . spec
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
DCC GARCH Spec
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Model
: DCC( 1 , 1 )
Estimation
:
2− step
Distribution
:
mvnorm
No . Parameters :
11
No . S e r i e s
2
:
> dcc . garch11 . spec
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
DCC GARCH Spec
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Model
: DCC( 1 , 1 )
Estimation
:
2− step
Distribution
:
mvnorm
No . Parameters :
11
No . S e r i e s
2
:
> dcc . f i t = d c c f i t ( dcc . garch11 . spec , data = data [ , 1 : 2 ] )
> dcc . f i t
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
DCC GARCH F i t
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Distribution
:
mvnorm
Model
:
DCC( 1 , 1 )
No . Parameters
:
11
[VAR GARCH DCC UncQ] : [0+8+2+1]
No . S e r i e s
:
2
215
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
No . Obs .
:
1248
Log−L i k e l i h o o d
:
7340.861
Av . Log−L i k e l i h o o d
:
5.88
Optimal Parameters
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
Estimate
Std . Error
t value Pr( >| t |)
[ Crude Oil ] .mu
0.000508
0.000618
0.822535 0.410772
[ Crude Oil ] . omega
0.000021
0.000016
1.340065 0.180224
[ Crude Oil ] . alpha1
0.128990
0.061989
2.080833 0.037449
[ Crude Oil ] . beta1
0.844323
0.076625 11.018899 0.000000
[ SP500 ] .mu
0.000856
0.000272
3.149145 0.001637
[ SP500 ] . omega
0.000004
0.000004
1.111306 0.266437
[ SP500 ] . alpha1
0.245249
0.049901
4.914712 0.000001
[ SP500 ] . beta1
0.719919
0.052742 13.649777 0.000000
[ J o i n t ] dcca1
0.000000
0.016584
0.000002 0.999999
[ J o i n t ] dccb1
0.937521
0.689475
1.359761 0.173905
Information C r i t e r i a
−−−−−−−−−−−−−−−−−−−−−
Akaike
− 11.747
Bayes
− 11.701
Shibata
− 11.747
Hannan−Quinn − 11.730
Elapsed time : 3.049927
> c l a s s ( dcc . f i t )
[ 1 ] " DCCfit "
a t t r ( , " package " )
[ 1 ] " rmgarch "
> slotNames ( dcc . f i t )
[ 1 ] " mfit "
" model "
216
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
> class
function ( x )
. Primitive ( " class " )
> slotNames
function ( x )
i f ( i s ( x , " c l a s s R e p r e s e n t a t i o n " ) ) names ( x@ sl ot s ) e l s e . slotNames ( x )
<bytecode : 0x000002b0deb3f0f0 >
<environment : namespace : methods>
217
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.12.3
Show DCC fit
Codes:
dcc.fit
Figure 7.44: The results of dcc.fit
Source:RStudio
Results:
> dcc . f i t
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
DCC GARCH F i t
*
*
*−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−*
Distribution
:
mvnorm
Model
:
DCC( 1 , 1 )
No . Parameters
:
11
[VAR GARCH DCC UncQ] : [0+8+2+1]
No . S e r i e s
:
2
No . Obs .
:
1248
Log−L i k e l i h o o d
:
7340.861
Av . Log−L i k e l i h o o d
:
5.88
Optimal Parameters
−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−−
218
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
Estimate
Std . Error
t value Pr( >| t |)
[ Crude Oil ] .mu
0.000508
0.000618
0.822535 0.410772
[ Crude Oil ] . omega
0.000021
0.000016
1.340065 0.180224
[ Crude Oil ] . alpha1
0.128990
0.061989
2.080833 0.037449
[ Crude Oil ] . beta1
0.844323
0.076625 11.018899 0.000000
[ SP500 ] .mu
0.000856
0.000272
3.149145 0.001637
[ SP500 ] . omega
0.000004
0.000004
1.111306 0.266437
[ SP500 ] . alpha1
0.245249
0.049901
4.914712 0.000001
[ SP500 ] . beta1
0.719919
0.052742 13.649777 0.000000
[ J o i n t ] dcca1
0.000000
0.016584
0.000002 0.999999
[ J o i n t ] dccb1
0.937521
0.689475
1.359761 0.173905
Information C r i t e r i a
−−−−−−−−−−−−−−−−−−−−−
Akaike
− 11.747
Bayes
− 11.701
Shibata
− 11.747
Hannan−Quinn − 11.730
Elapsed time : 3.049927
219
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 7. FINANCIAL ECONOMETRICS ANALYSIS OF R
7.7.13
Step thirteen: plot method
Codes:
plot(dcc.fit)
Notes: When plot(dcc.fit) is run, the following options will be shown on the Console:
Make a p l o t s e l e c t i o n ( or 0 t o e x i t ) :
1:
C ond iti ona l Mean ( vs Realized Returns )
2:
C ond iti ona l Sigma ( vs Realized Absolute Returns )
3:
C ond iti ona l Covariance
4:
C ond iti ona l C o r r e l a t i o n
5:
EW P o r t f o l i o P l o t with c o n d i t i o n a l d e n s i t y VaR l i m i t s
Selection :
Notes: That means there are five different plot can be output, and users just need to type the
Corresponding Number. In this example, Number 5 is typed.
Figure 7.45: EW portfolio plot with conditional density VaR limits
Source:RStudio
220
Electronic copy available at: https://ssrn.com/abstract=3863563
7.7. ARCH MODEL AND GARCH MODEL
7.7.13.1
Conditional standard deviation of each series’ plots
Codes:
plot(dcc.fit, which=2)
7.7.13.2
Conditional correlation
Codes:
plot(dcc.fit, which=4)
7.7.13.3
Extracting correlation series
Codes:
ts.plot(rcor(dcc.fit)[1,2,])
7.7.13.4
Forecasting conditional volatility and correlations
Codes:
dcc.fcst = dccforecast(dcc.fit, n.ahead=100)
class(dcc.fcst)
slotNames(dcc.fcst)
class(dcc.fcst@mforecast)
names(dcc.fcst@mforecast)
7.7.13.5
Show forecasts
Codes:
dcc.fcst
7.7.13.6
Plot forecasts
Codes:
4
221
Electronic copy available at: https://ssrn.com/abstract=3863563
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER
8
FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.1
A brief introduction of R
# A BRIEF INTRODUCTION OF R_EXERCISE
# NOTES: 1 . Symbol r e p r e s e n t s a COMMENT l i n e
# 2 . So everything typed a f t e r a # symbol i s ignored by the program
# 3 . I t mereLy outputs as t e x t in the Console Window below when executed
# 4 . Use <CTRL> + <ENTER> keys t o execute a p i e c e o f code
# 5. <CTRL> + L w i l l c l e a r the Console window i f i t g e t s t o o busy
# 1 .R works j u s t l i k e a c a l c u l a t o r
2+2
sqrt (25)
# 2 . Puts 1 t o 5 i n t o v a r i a b l e x
x <− 1 : 5
print ( x )
#NOTES: 1 . To assign a number t o an o b j e c t use assignment o pera tor <−
# 2 . R i s a l s o case s e n s i t i v e
# 3 . R s t o r e s numbers as arrays ( think column in Excel )
# 3 . Multiples 1 m i l l i o n by 1 percent
223
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
X <− 1E6 * 0.01
p r i n t (X)
# 4 . Puts the numbers 6 t o 10 i n t o y , note ’ c ’ f o r combine or concatenate
y <− c ( 6 , 7 , 8 , 9 , 10)
z <− X + y
print ( z )
# 5 . Assign a value t o multipl e o b j e c t s , a l l now have the value 3
a <− b <− c <− 3
print ( b )
c
a
# 6 .R commands are highly i n t u i t i v e e . g . remove or rm
#6.1 Removes the x v a r i a b l e
rm( x )
#6.2 Removes more than one o b j e c t at a time
rm( a , b )
#6.3 To remove a l l d efined o b j e c t s , use the l i s t = l s ( )
rm( l i s t = l s ( ) )
# 7 . To view the R s t y l e guide o n l i n e
browseURL ( " https : / / s t y l e . t i d y v e r s e . org " )
# 8 . Importing data
c o l a <− read . csv ( f i l e . choose ( ) , header = TRUE)
#Notes :
# 1 . Importing data f i l e s i s very quick and easy
# 2 . The most common f i l e format i s CSV (Comma Separated Values )
# 3 . Plus there i s no need t o s p e c i f y d e l i m i t e r s f o r missing data ,
l i k e with . t x t f i l e s
#8.1 Provides quick summary d e s c r i p t i v e
summary ( c o l a )
224
Electronic copy available at: https://ssrn.com/abstract=3863563
8.1. A BRIEF INTRODUCTION OF R
#8.2 Shows the f i r s 6 l i n e s
head ( c o l a )
#8.3 Shows the l a s t 6 l i n e s o f data
t a i l ( cola )
# 9 . Dataset
#9.1 R has b u i l t −in d a t a s e t s that can be used
require ( datasets )
data ( )
#9.2 To see more o n l i n e
browseURL ( " http : / /
s t a t . ethz . ch / R−manual / R−devel / l i b r a r y / d a t a s e t s / html / 0 0 Index . html " )
#9.3 Use ? f o r help . For example l e t s look at a i r m i l e s
? airmiles
#9.4 Shows the documentation in the Help window
airmiles
#9.5 Shows the data
str ( airmiles )
#Notes : 1 . s t r i s short f o r s t r u c t u r e
# 2 . We see i t i s time−s e r i e s , annual , with 24 years from 1937−60
# 9 . 6 Edit , then save a dataset
f i x ( airmiles )
#10. Packages
#10.1 Lots o f packages e x i s t f o r R, check out CRAN − Comprehensive R Archive Network
browseURL ( " http : / / cran . r−p r o j e c t . org / web / views / " )
225
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#Notes : 1 . Packages e x i s t f o r almost every research domain , so check them out
# 2 . Packages are s n i p p e t s o f code in R, can be s u b s t a n t i a l
#10.2 browse packages by name
browseURL ( " http : / / cran . r−p r o j e c t . org / web / packages / available_packages_by_name . html " )
#10.3 Check out other third −party sources t o o
browseURL ( " https : / / rdrr . i o / cran / c r a n t a s t i c . org / " )
#11. View current packages
library ( )
#11.1 To see c u r r e n t l y loaded packages
search ( )
#11.2 Get help i n s t a l l i n g packages
? i n s t a l l . packages
#Notes : S c r i p t s can be used t o i n s t a l l packages instead o f from the Tools menu
#11.3 Example o f graphics package f o r p l o t s − grammar o f graphics v e r s i o n 2
i n s t a l l . packages ( " ggplot2 " )
#11.4 load i n t o memory when the package need t o be used
l i b r a r y ( " ggplot2 " )
#11.5 A l t e r n a t i v e l y you can use r e q u i r e ( ) command
r e q u i r e ( " ggplot2 " )
#Notes : This i s p r e f e r a b l e f o r loading f u n c t i o n s
#11.6 How t o bring up documentation f o r an i n s t a l l e d package
l i b r a r y ( help =" ggplot2 " )
#11.7 I n s t a l l g r i d
i n s t a l l . packages ( " g r i d " )
226
Electronic copy available at: https://ssrn.com/abstract=3863563
8.1. A BRIEF INTRODUCTION OF R
#11.8 Another f u n c t i o n v i g n e t t e ( ) brings up l i s t s o f examples
v i g n e t t e ( package =" g r i d " )
#11.9 Detach a package from memory ( make i t i n a c t i v e )
detach ( " package : ggplot2 " , unload = TRUE)
#11.10 Example t o i n s t a l l then remove a package
i n s t a l l . packages ( " psytabs " )
#Notes : Formatting package f o r APA (USA) f o r P s y c h o l o g i c a l Research
remove . packages ( " psytabs " )
227
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.2
CSV file import
#CSV FILE IMPORT_COMMA SEPARATED VALUES (CSV)
#1.1 Read data
uscovid . csv <− read . csv ( f i l e . choose ( ) , header=TRUE)
#1.2 Check the imported f i l e s t r u c t u r e
s t r ( uscovid . csv )
#1.3 View the imported data − note R i s case s e n s i t i v e
View ( uscovid . csv )
#1.4 Look at a quick summary d e s c r i p t i o n o f the data
summary ( uscovid . csv )
#1.5 P l o t the data so we can v i s u a l i z e i t onscreen
p l o t ( Cases ~ Population , data=uscovid . csv ,
xlab =" Population Size " ,
ylab =" Covid −19 cases " ,
main=" Covid Cases by State Population " ,
xlim=c (500000 , 40000000) ,
ylim=c (10000 , 3500000) , l a s =1)
#1.6 C alcul ate c o r r e l a t i o n between Covid −19 cases and population s i z e
c o r ( uscovid . csv$Cases , uscovid . csv$Population )
#1.7 Try a simple l i n e a r r e g r e s s i o n model
model <− lm ( Cases ~ Population , data=uscovid . csv )
# 1 . 7 . 1 Shows us the b a s i c model equation
p r i n t ( model )
# 1 . 7 . 2 Outputs the r e g r e s s i o n model d i a g n o s t i c s
summary ( model )
#1.8 Second model uses Rate o f Covid −19 instead o f Case Numbers
model2 <− lm ( Rate ~ Population , data=uscovid . csv )
summary ( model2 )
#1.9 P l o t the Covid −19 r a t e data by State population s i z e
228
Electronic copy available at: https://ssrn.com/abstract=3863563
8.2. CSV FILE IMPORT
p l o t ( Rate ~ Population ,
data=uscovid . csv ,
xlab =" Population Size " ,
ylab =" Covid −19 Rate " ,
main=" Covid Rate by State Population " ,
xlim=c (500000 , 40000000) ,
ylim=c ( 0 , 0 . 1 5 ) ,
l a s =1)
#NOTES: the huge variance amongst smaller s t a t e s with widely d i f f e r i n g Covid −19
# r a t e s (2% t o 13%)
browseURL ( " http : / /
s t a t . ethz . ch / R−manual / R−devel / l i b r a r y / d a t a s e t s / html / 0 0 Index . html " )
229
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.3
Prepare scatter plot
#PREPARE SCATTER PLOT
#1.1 Read data from . csv f i l e
data <− read . csv ( f i l e . choose ( ) , header = TRUE) # c o l a _ c s v
head ( data )
#1.2 S c a t t e r P l o t
p l o t ( data , main = " S c a t t e r P l o t " )
#1.3 Add best − f i t l i n e t o the s c a t t e r p l o t
# 1 . 3 . 1 I n s t a l l Package and load the l i b r a r y
i n s t a l l . packages ( " hydroGOF " )
i n s t a l l . packages ( " zoo " )
l i b r a r y ( zoo )
l i b r a r y ( " hydroGOF " )
#1.4 F i t l i n e a r model using OLS
model = lm ( Cola ~ Temperature , data )
#1.5 Overlay best − f i t l i n e on s c a t t e r p l o t
a b l i n e ( model )
#1.6 C alcul ate RMSE
PredCola = p r e d i c t ( model , data )
RMSE = rmse ( PredCola , data$Cola )
#1.7 Summary model d i a g n o s t i c s
summary ( model )
#1.8 F i t t i n g Log−l i n e a r model
# 1 . 8 . 1 Transform the dependent v a r i a b l e
data$LCola = l o g ( data$Cola , base = exp ( 1 ) )
#1.8.2 Scatter Plot
p l o t ( LCola ~ Temperature , data
= data , main = " S c a t t e r P l o t " )
230
Electronic copy available at: https://ssrn.com/abstract=3863563
8.3. PREPARE SCATTER PLOT
# 1 . 8 . 3 F i t the best l i n e in log −l i n e a r model
model1 = lm ( LCola ~ Temperature , data )
a b l i n e ( model1 )
#1.9 Calcul ate RMSE
PredCola1 = p r e d i c t ( model1 , data )
RMSE = rmse ( PredCola1 , data$LCola )
#1.10 Summary model d i a g n o s t i c s f o r l o g ( c o l a ) model
summary ( model1 )
231
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.4
GLM examples_initial setup
#INITIAL SETUP_EXERCISE
#1.1 I n s t a l l needed packages
i n s t a l l . packages ( " faraway " )
i n s t a l l . packages ( "MASS" )
i n s t a l l . packages ( " car " )
i n s t a l l . packages ( " lme4 " )
i n s t a l l . packages ( " reshape2 " )
#1.2 Load the r e c e n t l y i n s t a l l e d packages as l i b r a r i e s
l i b r a r y ( faraway )
# glm support
l i b r a r y (MASS)
# negative binomial support
l i b r a r y ( car )
# regression functions
l i b r a r y ( carData )
# carData
l i b r a r y ( Matrix )
# Matrix
l i b r a r y ( lme4 )
# random e f f e c t s
l i b r a r y ( ggplot2 )
# p l o t t i n g commands
l i b r a r y ( reshape2 )
# wide t o t a l l reshaping
library ( xtable )
# n i c e t a b l e formatting
library ( knitr )
# kable t a b l e formatting
library ( grid )
# units function for ggplot
#1.3 LOAD THE DATA FRAME
data ( hsb )
s t r ( hsb )
#NOTES: can check out the documentation a s s o c i a t e d with t h i s data frame
? hsb
#1.4 Enter the GLM model command
g1 <− glm ( I ( prog ==" academic " ) ~ gender+race+ses+schtyp+
read+write+ s c i e n c e + s o c s t ,
family=binomial ( ) , data=hsb )
#1.5 Now examine the summary output
232
Electronic copy available at: https://ssrn.com/abstract=3863563
8.4. GLM EXAMPLES_INITIAL SETUP
summary ( g1 )
#1.6 Next we w i l l use the step command with the l i k e l i h o o d r a t i o t e s t (LRT)
t o reduce unnecessary v a r i a b l e s
step ( g1 , t e s t ="LRT" )
step ( g1 , t e s t ="LRT" )
#1.7 Can now f i t the model suggested by step
g2 <− glm ( I ( prog == " academic " ) ~ ses + schtyp + read +
write + s c i e n c e + s o c s t ,
family = " binomial " , data = hsb )
summary ( g2 )
#NOTES:
# 1 . There are d i f f e r e n c e s between the p−values reported in summary and what
was reported the the LRT t e s t in the f i n a l step o f the step ( ) f u n c t i o n above .
# 2 . The quasi−binomial family i s u s e f u l f o r modeling response v a r i a b l e s with
a bounded range .
# 3 . Residual p l o t s provide l i t t l e a s s i s t a n c e in evaluating binary models .
The d i a g n o s t i c s f o r the s e n s i t i v i t y o f the model t o the data are checked
using the same methods as was done f o r OLS models .
# 2 . COUNT RESPONSE VARIABLE
#NOTES:
# 1 . Count data i s t y p i c a l l y modeled using the poisson family .
This uses a l o g l i n k f u n c t i o n and a variance f u n c t i o n o f ? ? .
The modeled response i s the p r e d i c t e d l o g count .
# 2 . Will use the d i s c o v e r i e s dataset from the d a t a s e t s package f o r
our binary response model . This dataset i s the number o f great i n v e n t i o n s
f o r the years 1860 t o 1959. We w i l l s t a r t by loading the data . frame
and adding a v a r i a b l e t o represent the number o f years s i n c e 1860.
We w i l l assume that there i s no c o r r e l a t i o n between the years
233
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
t o f o c u s on the GLM model f i t .
#2.1 Load the d i s c o v e r i e s data frame
data ( d i s c o v e r i e s )
d i s c <− data . frame ( count=as . numeric ( d i s c o v e r i e s ) ,
year=seq ( 0 , ( length ( d i s c o v e r i e s ) − 1 ) , 1 ) )
#2.2 Can examine the documentation a s s o c i a t e d with t h i s data frame
? discoveries
#2.3 w i l l f i t the count o f i n v e n t i o n s with year and year squared .
# 2 . 3 . 1 Create new o b j e c t f o r year−squared
yearSqr=d i s c $ y e a r ^2
# 2 . 3 . 2 Run the GLM model s p e c i f i c a t i o n , as signi ng t o new o b j e c t p1
p1 <− glm ( count~year+yearSqr , family =" poisson " , data= d i s c )
# 2 . 3 . 3 Examine the summary output t a b l e
summary ( p1 )
#2.4 Check the goodness o f f i t o f t h i s model . We w i l l use the deviance o f
the r e s i d u a l s f o r t h i s t e s t
1 − pchisq ( deviance ( p1 ) , df . r e s i d u a l ( p1 ) )
#2.5 The p−value i s approximately . 0 0 1 . This deviance i s not l i k e l y t o
have occurred by chance , under the n u l l hypothesis o f
the deviances being xx . Therefore we have evidence o f
over−d i s p e r s i o n . The presence o f over−d i s p e r s i o n suggested the use
o f the F−t e s t f o r nested models . Will t e s t i f the squared term
can be dropped from the model .
drop1 ( p1 , t e s t ="F " )
234
Electronic copy available at: https://ssrn.com/abstract=3863563
8.4. GLM EXAMPLES_INITIAL SETUP
#NOTES: The p−value f o r yearSqr term i s very small , so we w i l l
r e t a i n yearSqr in the model
# 3 . QUASI−POISSON MODEL
#NOTES: The i n v e n t i o n count model from above needs t o be f i t
using the quasipoisson family . The quasipoisson family
w i l l account f o r the g r e a t e r variance in the data .
p2 <− glm ( count~year+yearSqr , family =" quasipoisson " , data= d i s c )
summary ( p2 )
#NOTES: There i s no change in the estimated c o e f f i c i e n t between
the quasipoisson f i t and the poisson f i t .
The s i g n i f i c a n c e o f the terms does change and a d i s p e r s i o n parameter
i s estimated .
#3.1 Before determining that the quasipoisson family i s appropriate ,
we w i l l check t o see i f the variance o f the r e s i d u a l s i s p r o p o r t i o n a l
t o the mean . We begin t h i s check by c r e a t i n g a new data . frame
which i n c l u d e s the r e s i d u a l s and f i t t e d values .
p1Diag <− data . frame ( d i s c ,
l i n k = p r e d i c t ( p1 , type =" l i n k " ) ,
f i t = p r e d i c t ( p1 , type =" response " ) ,
pearson= r e s i d u a l s ( p1 , type =" pearson " ) ,
r e s i d = r e s i d u a l s ( p1 , type =" response " ) ,
residSqr= r e s i d u a l s ( p1 , type =" response " ) ^ 2 )
#NOTES: Beware that there i s no Console output f o r p1Diag o b j e c t
#3.2 Will p l o t the square o f the r e s i d u a l t o the p r e d i c t e d mean .
Will add t o t h i s s c a t t e r p l o t a black l i n e f o r the Poisson assumed variance ,
a green l i n e f o r the quasi−Poisson assumed variance , and a blue curve f o r
the smoothed mean o f the square o f the r e s i d u a l .
g g p l o t ( data=p1Diag , aes ( x= f i t , y=residSqr ) ) +
235
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
geom_point ( ) +
geom_abline ( i n t e r c e p t = 0 , s l o p e = 1 ) +
geom_abline ( i n t e r c e p t = 0 , s l o p e = summary ( p2 ) $dispersion ,
c o l o r =" green " ) +
stat_smooth ( method=" l o e s s " , se = FALSE) +
theme_bw ( )
#NOTES: I d e a l l y the blue curve would be s t r a i g h t and i t would be
c o l l i n e a r with the green l i n e f o r the quasi−Poisson variance .
The g r e a t e r the d e v i a t i o n from the green l i n e the g r e a t e r the concern
i s about the p r o p o r t i o n a l i t y o f the variance t o the mean .
Here we have some i n d i c a t i o n that the variance may not be p r o p o r t i o n a l
t o the mean .
# 4 . NEGATIVE BINOMIAL MODEL
#4.1 now look t o see i f a negative binomial model might be a b e t t e r f i t .
The negative binomial r e q u i r e s the use o f the glm . nb ( ) f u n c t i o n .
The c a l l t o glm . nb i s s i m i l a r t o glm , except no family i s given .
nb1 <− glm . nb ( count~year+yearSqr , data= d i s c )
summary ( nb1 )
#NOTES: the c o e f f i c i e n t s have only a small change from the quasi−Poisson model .
#4.2 The drop1 f u n c t i o n i s used t o t e s t the s i g n i f i c a n c e o f the squared term f o r y
We use the l i k e l i h o o d r a t i o t e s t f o r negative binomial models .
drop1 ( nb1 , t e s t ="LRT" )
#NOTES: The squared term i s s i g n i f i c a n t and i s r e t a i n e d in the model .
#4.3 Will repeat the check o f the variance o f the r e s i d u a l s which was done
f o r the quasi−Poisson model . P l o t t i n g the square o f the r e s i d u a l
t o the f i t t e d values , with a black l i n e f o r Poisson , green l i n e f o r quasi−Poisson ,
a blue curve f o r smoothed mean o f the square o f the r e s i d u a l ,
and a red curve f o r p r e d i c t e d variance from the negative binomial f i t .
236
Electronic copy available at: https://ssrn.com/abstract=3863563
8.4. GLM EXAMPLES_INITIAL SETUP
nb1Diag <− data . frame ( d i s c ,
l i n k = p r e d i c t ( nb1 , type =" l i n k " ) ,
f i t = p r e d i c t ( nb1 , type =" response " ) ,
pearson= r e s i d u a l s ( nb1 , type =" pearson " ) ,
r e s i d = r e s i d u a l s ( nb1 , type =" response " ) ,
residSqr= r e s i d u a l s ( nb1 , type =" response " ) ^ 2 )
g g p l o t ( data=nb1Diag , aes ( x= f i t , y=residSqr ) ) +
geom_point ( ) +
geom_abline ( i n t e r c e p t = 0 , s l o p e = 1 ) +
geom_abline ( i n t e r c e p t = 0 , s l o p e = summary ( p2 ) $dispersion ,
c o l o r =" green " ) +
s t a t _ f u n c t i o n ( fun= f u n c t i o n ( f i t ) { f i t + f i t ^ 2 / 1 1 . 5 3 } , c o l o r =" red " ) +
stat_smooth ( method=" l o e s s " , se = FALSE) +
theme_bw ( )
#NOTES: Please note no c o n s o l e output , but check out the p l o t s .
The negative binomial variance curve ( red ) i s c l o s e
t o the quasi−Poisson l i n e ( green . )
#4.4 Although the means and variance p r e d i c t i o n s f o r the negative binomial
and quasi−Poisson models are s i m i l a r , the p r o b a b i l i t y f o r
any given i n t e g e r i s d i f f e r e n t f o r the two models .
The f o l l o w i n g code shows the p r e d i c t e d p r o b a b i l i t i e s o f
0 through 7 when the mean i s p r e d i c t e d t o be 4 .
data . frame ( number = 0 : 8 ,
prob_Poisson=round ( dpois ( 0 : 8 , ( 4 * summary ( p2 ) $ d i s p e r s i o n ) ) , 3 ) ,
prob_NBinom=round ( dnbinom ( 0 : 8 ,mu=4 , s i z e =summary ( nb1 ) $theta ) , 3 ) )
summary ( nb1 ) $theta
#NOTES: The i n t e r p r e t a t i o n o f the two models i s d i f f e r e n t as well as
the p r o b a b i l i t i e s o f the event counts . Examining the d i a g n o s t i c s
would be u s e f u l step in choosing between these two models .
237
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.5
Illustration of using GLM data set in R
#ILLUSTRATION OF USING GLM DATA SET IN R_EXERCISE
#1.1 I n s t a l l and Load the GLMsData Package
i n s t a l l . packages ( " GLMsData " )
l i b r a r y ( " GLMsData " )
#1.2 Make the data s e t lungcap a v a i l a b l e f o r use
data ( lungcap )
#1.3 Look at the v a r i a b l e names
names ( lungcap )
#1.4 Show the f i r s t 6 l i n e s o f data
head ( lungcap )
#1.5 Examine f i r s t 6 rows o f a p a r t i c u l a r v a r i a b l e
− using the $ symbol b e f o r e v a r i a b l e name
head ( lungcap$Age )
#1.6 Show the l a s t 6 l i n e s o f data
t a i l ( lungcap )
#1.7 As above , you can examine the l a s t 6 rows o f data f o r another one v a r i a b l e
t a i l ( lungcap$Gender )
#1.8 The s t r u c t u r e command shows you information about the data frame lungcap
s t r ( lungcap )
#1.9 The dimension command t e l l s you the numbers o f cases and v a r i a b l e s
in your data frame
dim ( lungcap )
#1.10 The summary command g i v e s you d e s c r i p t i v e data summary ( mean , median ,
max / min & i n t e r q u a r t i l e range )
summary ( lungcap )
238
Electronic copy available at: https://ssrn.com/abstract=3863563
8.5. ILLUSTRATION OF USING GLM DATA SET IN R
#1.11 Again , i t i s p o s s i b l e t o summarize a s p e c i f i c v a r i a b l e using the $
symbol then v a r i a b l e name
summary ( lungcap$Age )
summary ( lungcap$Smoke )
#1.12 Need t o e x p l i c i t l y t e l l R that Smoke i s a q u a l i t a t i v e ( nominal data )
lungcap$Smoke <− f a c t o r ( lungcap$Smoke ,
l e v e l s =c ( 0 , 1 ) ,
l a b e l s =c ( " Non−smoker " , " Smoker " ) )
#1.13 Now, c o r r e c t l y summarize t h i s nominal v a r i a b l e
summary ( lungcap$Smoke )
#1.14 I f need help on t h i s data frame , then simply use the ’ ? ’ b e f o r e the name
? lungcap
# 2 . PLOTTING DATA
#2.1 P l o t FEV versus AGE using the data−frame " lungcap " using good format ranges
p l o t (FEV~Age , data=lungcap ,
xlab ="Age ( in years ) " , # The x−a x i s l a b e l
ylab ="FEV ( in L ) " , # The y−a x i s l a b e l
main="FEV vs Age " , # The Graph main t i t l e
xlim=c ( 0 , 2 0 ) , # E x p l i c i t l y s e t s x−a x i s l i m i t s
ylim=c ( 0 , 6 ) , # E x p l i c i t l y s e t s y−a x i s l i m i t s
l a s =1) # Makes a x i s l a b e l s h o r i z o n t a l
#2.2 Now repeat t h i s p r o c e s s f o r FEV versus Height
p l o t ( FEV ~ Ht , data=lungcap , main="FEV vs height " ,
xlab =" Height ( in inches ) " , ylab ="FEV ( in L ) " ,
l a s =1 , ylim=c ( 0 , 6 ) )
#2.3 Also p l o t FEV by b i o l o g i c a l Gender
p l o t ( FEV ~ Gender , data=lungcap ,
main="FEV vs gender " , ylab ="FEV ( in L ) " ,
l a s =1 , ylim=c ( 0 , 6 ) )
239
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#2.4 P l o t FEV by Smoker or Non−Smoker
p l o t ( FEV ~ Smoke , data=lungcap , main="FEV vs Smoking s t a t u s " ,
ylab ="FEV ( in L ) " , xlab ="Smoking s t a t u s " ,
l a s =1 , ylim=c ( 0 , 6 ) )
#NOTES:
# 1 . Central dark l i n e i s Median ( not Mean ) , boxes represent 25−75 percent upper
& lower q u a r t i l e s that R h i g h l i g h t s i n d i v i d u a l o u t l i e r s ( c i r c l e s )
as obs > 1 . 5 x i n t e r q u a r t i l e range .
# 2 . Sur p r i s i n g l y , smokers appear t o have a l a r g e r FEV than non−smokers
# 3 . EXPLORING RELATIONSHIPS INVOLVING TWO VARIABLES AT A TIME
#NOTES: p l o t the data separate f o r smokers and non−smokers using s i m i l a r s c a l e s
on the axes
#3.1 F i r s t we do Age f o r Smokers versus Non−smokers
p l o t ( FEV ~ Age ,
data=subset ( lungcap , Smoke=="Smoker " ) , # Only s e l e c t smokers
main="FEV vs age\ n f o r smokers " , # \n means ‘ new l i n e ’
ylab ="FEV ( in L ) " , xlab ="Age ( in years ) " ,
ylim=c ( 0 , 6 ) , xlim=c ( 0 , 2 0 ) , l a s =1)
p l o t ( FEV ~ Age ,
data=subset ( lungcap , Smoke=="Non−smoker " ) , # Only s e l e c t non−smokers
main="FEV vs age\ n f o r non−smokers " ,
ylab ="FEV ( in L ) " , xlab ="Age ( in years ) " ,
ylim=c ( 0 , 6 ) , xlim=c ( 0 , 2 0 ) , l a s =1)
#3.2 repeat the p l o t s f o r Height ; smokers versus non−smokers
p l o t ( FEV ~ Ht , data=subset ( lungcap , Smoke=="Smoker " ) ,
main="FEV vs height\ n f o r smokers " ,
ylab ="FEV ( in L ) " , xlab =" Height ( in inches ) " ,
xlim=c ( 4 5 , 7 5 ) , ylim=c ( 0 , 6 ) , l a s =1)
p l o t ( FEV ~ Ht , data=subset ( lungcap , Smoke=="Non−smoker " ) ,
main="FEV vs height\ n f o r non−smokers " ,
240
Electronic copy available at: https://ssrn.com/abstract=3863563
8.5. ILLUSTRATION OF USING GLM DATA SET IN R
ylab ="FEV ( in L ) " , xlab =" Height ( in inches ) " ,
xlim=c ( 4 5 , 7 5 ) , ylim=c ( 0 , 6 ) , l a s =1)
#NOTES: Please note that we use == t o make l o g i c a l comparisons
# 4 . ADJUSTING VARIABLES TO DISTINGUISH BETTER
#4.1 F i r s t , a d j u s t Age so that the values f o r smokers and non−smokers are
s l i g h t l y separated
AgeAdjust <− lungcap$Age + i f e l s e ( lungcap$Smoke=="Smoker " , 0 , 0 . 5 )
#NOTES: The code i f e l s e ( lungcap$Smoke=="Smoker " , 0 , 0 . 5 ) adds zero t o
the value o f Age f o r youth l a b e l e d with Smoker ,
and adds 0 . 5 t o youth l a b e l e d otherwise ( that i s , non−smokers ) .
#4.2 Then we p l o t FEV against t h i s v a r i a b l e AgeAdjust
p l o t ( FEV ~ AgeAdjust , data=lungcap ,
pch = i f e l s e ( Smoke=="Smoker " , 3 , 2 0 ) ,
xlab ="Age ( in years ) " , ylab ="FEV ( in L ) " , main="FEV vs age " , l a s =1)
#NOTES:
# 1 . The input pch i n d i c a t e s the p l o t t i n g c h a r a c t e r t o use when p l o t t i n g ;
then , i f e l s e ( Smoke=="Smoker " , 3 , 20) means t o p l o t
with p l o t t i n g c h a r a c t e r 3 ( a ’ plus ’ sign )
i f Smoke takes the value " Smoker " , and otherwise
t o p l o t with p l o t t i n g c h a r a c t e r 20 ( a f i l l e d c i r c l e ) .
# 2 . More information may be found here using help p o i n t s
? points
#4.3 Before d i s t i n g u i s h e d the nominal v a r i a b l e ’ Smoker ’ vs ’ Non−smoker ’ ,
now use legend command
legend ( " t o p l e f t " , pch=c ( 2 0 , 3 ) , legend=c ( " Non−smokers " , " Smokers " ) )
#NOTES: The f i r s t input s p e c i f i e s the l o c a t i o n ( such as " c e n t e r " or " bottomright " ) .
The second input g i v e s the p l o t t i n g n o t a t i o n t o be explained
( such as the points , using pch , or the l i n e types , using l t y ) .
The legend input prov ides the explanatory t e x t .
241
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#4.4 For more i n f o see help legends
? ? legends
#4.5 Now, can use a b o x p l o t t o show the r e l a t i o n s h i p
b o x p l o t ( lungcap$FEV ~ lungcap$Smoke + lungcap$Gender ,
ylab ="FEV ( in L ) " , main="FEV, by gender\n and smoking s t a t u s " ,
l a s =2 , # Keeps l a b e l s perpendicular t o the axes
names=c ( " F: \nNon " , "F: \ nSmoker " , "M: \nNon " , "M: \ nSmoker " ) )
# 5 . INTERACTION PLOTS
#NOTES: Another way t o show the r e l a t i o n s h i p between three v a r i a b l e s i s t o use
an i n t e r a c t i o n p l o t , which shows the r e l a t i o n s h i p between the l e v e l s o f
two f a c t o r s and ( by d e f a u l t ) the mean response o f a q u a n t i t a t i v e v a r i a b l e .
#5.1 F i r s t , examine mean FEV by Non−smoker vs Smoker , and by Male vs Female
i n t e r a c t i o n . p l o t ( lungcap$Smoke , lungcap$Gender , lungcap$FEV ,
xlab ="Smoking s t a t u s " , ylab ="FEV ( in L ) " ,
main="Mean FEV, by gender\n and smoking s t a t u s " ,
t r a c e . l a b e l ="Gender " , l a s =1)
#5.2 Repeat t h i s step r e p l a c i n g mean FEV with mean Age
i n t e r a c t i o n . p l o t ( lungcap$Smoke , lungcap$Gender , lungcap$Age ,
xlab ="Smoking s t a t u s " , ylab ="Age ( in years ) " ,
main="Mean age , by gender\n and smoking s t a t u s " ,
t r a c e . l a b e l ="Gender " , l a s =1)
#5.3 Can see that Height and Age are s t r o n g l y c o r r e l a t e d
p l o t ( Ht ~ Age , data=subset ( lungcap , Gender=="F " ) , l a s =1 ,
ylim=c ( 4 5 , 7 5 ) , xlim=c ( 0 , 2 0 ) , # Use s i m i l a r s c a l e s f o r comparisons
main=" Females " , xlab ="Age ( in years ) " , ylab =" Height ( in inches ) " )
p l o t ( Ht ~ Age , data = subset ( lungcap , Gender=="M" ) , l a s =1 ,
ylim=c ( 4 5 , 7 5 ) , xlim=c ( 0 , 2 0 ) , # Use s i m i l a r s c a l e s f o r comparisons
242
Electronic copy available at: https://ssrn.com/abstract=3863563
8.5. ILLUSTRATION OF USING GLM DATA SET IN R
main=" Males " , xlab ="Age ( in years ) " , ylab =" Height ( in inches ) " )
# 6 . MODEL FITTING CONSIDERATIONS
#6.1 The r e l a t i o n s h i p between FEV and height i s not l i n e a r , so a l i n e a r model
i s not appropriate . However , p l o t t i n g the logarithm o f FEV
against height does show an approximate l i n e a r r e l a t i o n s h i p .
#6.2 Use the s c a t t e r . smooth f u n c t i o n t o add a smooth curve t o the s c a t t e r p l o t
s c a t t e r . smooth ( lungcap$Ht , lungcap$FEV , l a s =1 , c o l =" grey " ,
ylim=c ( 0 , 6 ) , xlim=c ( 4 5 , 7 5 ) , # Use s i m i l a r s c a l e s f o r comparisons
main="FEV" , xlab =" Height ( in inches ) " , ylab ="FEV ( in L ) " )
s c a t t e r . smooth ( lungcap$Ht , l o g ( lungcap$FEV ) , l a s =1 , c o l =" grey " ,
ylim=c ( − 0.5 , 2 ) , xlim=c ( 4 5 , 7 5 ) , # Use s i m i l a r s c a l e s f o r comparisons
main=" l o g o f FEV" , xlab =" Height ( in inches ) " ,
ylab =" l o g o f FEV ( in L ) " )
#NOTES: Perhaps running a l i n e a r model using Log (FEV) instead o f
j u s t FEV would s u f f i c e
# 7 . CONSIDER FITTING A LINEAR MODEL
data ( lungcap )
lungcap$Smoke <− f a c t o r ( lungcap$Smoke , l e v e l s =c ( 0 , 1 ) ,
l a b e l s =c ( " Non−smoker " , " Smoker " ) )
Xmat <− model . matrix ( ~ Age + Ht + f a c t o r ( Gender ) + f a c t o r ( Smoke ) ,
data=lungcap )
#NOTES: Here model . matrix ( ) i s used t o combine the v a r i a b l e s
as columns o f a matrix , a f t e r d e c l a r i n g Smoke as a f a c t o r
[ Note that % denotes matrix , so %*% denotes matrix m u l t i p l i c a t i o n ]
head ( Xmat )
243
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
XtX <− t ( Xmat ) %*% Xmat # t ( ) i s transpose ; %*% i s matrix multiply
y <− l o g ( lungcap$FEV ) # Assigning y as new v a r i a b l e f o r l o g (FEV)
inv . XtX <− s o l v e ( XtX ) # s o l v e returns the matrix i n v e r s e
XtY <− t ( Xmat ) %*% y
beta <− inv . XtX %*% XtY ; drop ( beta )
#NOTES:
# 1 . Note drop ( ) drops any unnecessary dimensions
( turns a s i n g l e column matrix t o a v e c t o r )
# 2 . The f i t t e d model has the systematic component
^?? = ? ? ? 1 . 9 4 4 + 0.02339Age + 0.04280Ht + 0.02932 Gender ? ? ? 0.04607Smoke ,
where Gender i s 0 f o r females and 1 f o r males ,
and Smoke i s 0 f o r non−smokers and 1 f o r smokers .
#7.1 Note that s l i g h t l y more e f f i c i e n t code would have been t o
compute ^?? by s o l v i n g a l i n e a r system o f equations :
beta <− s o l v e ( XtX , XtY ) ; beta
#NOTES: The r e s u l t i s the same
#7.2 An even more e f f i c i e n t approach would have been t o use the QR−decomposition :
QR <− qr ( Xmat )
beta <− qr . c o e f (QR, y ) ; beta
#NOTES: Once again , the r e s u l t i s the same
( more than one way t o achieve the same r e s u l t in R)
# 8 . ESTIMATING THE VARIANCE
#8.1 A f t e r estimating the beta c o e f f i c i e n t s , the f i t t e d values are
obtained as ??=X ? ? . The variance sigma−squared i s estimated
244
Electronic copy available at: https://ssrn.com/abstract=3863563
8.5. ILLUSTRATION OF USING GLM DATA SET IN R
from the RSS as usual
mu <− Xmat %*% beta
RSS <− sum( ( y − mu)^2 ) ; RSS
s2 <− RSS / ( length ( lungcap$FEV ) − length ( beta ) )
c ( s= s q r t ( s2 ) , s2=s2 )
#NOTES: Please note that using the l i n e a r model lm ( ) command
a l l o f these c a l c u l a t i o n s are performed a u t o m a t i c a l l y .
However , sometimes i t i s good t o see the i n d i v i d u a l computations being performed .
# 9 . ESTIMATING THE VARIANCE OF BETA COEFFICIENTS
#9.1 For the model r e l a t i n g FEV t o Age , Height , Gender and Smoking s t a t u s
var . matrix <− s2 * inv . XtX
var . b e t a j <− diag ( var . matrix ) # diag ( ) grabs the diagonal elements
# Diagonal elements o f var [ beta ] are the values o f var [ b e t a j ] from which
the estimated standard e r r o r s o f the i n d i v i d u a l parameters are computed ,
so se ( b e t a j ) = s q r t [ var ( b e t a j ) ]
s q r t ( var . b e t a j )
#NOTES: these are c a l c u l a t e d a u t o m a t i c a l l y by lm ( ) in R
#10. ESTIMATING THE VARIANCE OF FITTED VALUES
#NOTES: Previous p l o t s suggest a l i n e a r r e l a t i o n s h i p between l o g (FEV) and height .
#Suppose we wish t o estimate the mean o f l o g (FEV) f o r females ( that i s x3 = 0 )
# that smoke ( that i s , x4 = 1 ) ,
#aged 18 who are 66 inches t a l l using the e a r l i e r model
xg . vec <− matrix ( c ( 1 , 18 , 66 , 0 , 1 ) , nrow=1) # The f i r s t 1 i s the constant term
mu. g <− xg . vec %*% beta
var .mu. g <− s q r t ( xg . vec
%*% ( s o l v e ( t ( Xmat)%*%Xmat ) )
%*% t ( xg . vec ) * s2 )
c ( mu. g , var .mu. g )
# Shows the estimate o f l o g (FEV) f o r these parameters and second value i s variance
245
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#11. FITTING A LINEAR MODEL
#11.1 R e c a l l that Smoke has been d e c l a r e a f a c t o r p r e v i o u s l y
lm ( l o g (FEV) ~ Age + Ht + Gender + Smoke , data=lungcap )
#11.2 The constant term in the model i s included i m p l i c i t l y by R, s i n c e i t i s
almost always necessary . To e x p l i c i t l y exclude the constant in the model
( which i s unusual ) , use one o f these forms
lm ( l o g (FEV) ~ 0 + Age + Ht + Gender + Smoke , data=lungcap )
# No constant term ( or i n t e r c e p t )
lm ( l o g (FEV) ~ Age + Ht + Gender + Smoke − 1 , data=lungcap )
# No constant term
#11.3 R returns more information about the f i t t e d model by
d i r e c t i n g the output o f lm ( ) t o an output o b j e c t
LC.m1 <− lm ( l o g (FEV) ~ Age + Ht + Gender + Smoke , data=lungcap )
#11.4 The names o f the components o f LC.m1
names ( LC.m1 )
#11.5 For f u l l summary output
summary (LC.m1)
#11.6 To examine parameter c o e f f i c i e n t s d i r e c t l y
c o e f ( LC.m1 )
#11.7 The estimate o f standard d e v i a t i o n ( sigma )
summary ( LC.m1 ) $sigma
#NOTES: A d d i t i o n a l explanation on p−value s i g n i f i c a n c e codes
# The *** i n d i c a t e s a two−t a i l e d P−value between 0 and 0 . 0 0 1 ;
** i n d i c a t e s a two−t a i l e d P−value between 0.001 and 0 . 0 1 ;
* i n d i c a t e s a two−t a i l e d P−value between 0.01 and 0 . 0 5 ;
. i n d i c a t e s a two−t a i l e d P−value between 0.05 and 0 . 1 0 .
#11.8 This information can be accessed d i r e c t l y using c o e f ( summary ( ) ) :
246
Electronic copy available at: https://ssrn.com/abstract=3863563
8.5. ILLUSTRATION OF USING GLM DATA SET IN R
round ( c o e f ( summary ( LC.m1 ) ) , 5 )
#NOTES: Thus , some evidence e x i s t s that smoking s t a t u s i s s t a t i s t i c a l l y
s i g n i f i c a n t when age , height and gender are in the model .
I f gender was omitted from the model and the r e l e v a n t n u l l hypothesis r e t e s t e d ,
the t e s t has a d i f f e r e n t meaning : t h i s second t e s t determines
i f age i s s i g n i f i c a n t in the model adjusted only f o r height ( but not gender ) .
Consequently , we should expect the t e s t s t a t i s t i c and P−values t o be d i f f e r e n t ,
and so the c o n c l u s i o n may d i f f e r a l s o .
#12 FINDING A 95% CONFIDENCE INTERVAL
#12.1 Find a 95 percent CI f o r a l l 5 r e g r e s s i o n c o e f f i c i e n t s
using the c o n f i n t ( ) command
c o n f i n t ( LC.m1)
#12.2 Estimating a s p e c i f i c s c e n a r i o CI
#12.3 Suppose we wish t o estimate mu = E[ l o g (FEV ) ] f o r female smokers aged 18
who are 66 inches t a l l
#12.4 Using R, we f i r s t c r e a t e a new data frame c o n t a i n i n g the values o f
the explanatory v a r i a b l e s f o r which we need t o make the p r e d i c t
new . df <− data . frame ( Age=18 , Ht=66 , Gender="F" , Smoke="Smoker " )
#12.5 Then we use the p r e d i c t ( ) command t o compute the estimates o f mu
out <− p r e d i c t ( LC. m1, newdata=new . df , se . f i t =TRUE)
names ( out )
#12.6 Now choose the one we want
out$se . f i t
t s t a r <− qt ( df=LC. m1$df , p=0.975 ) # For a 95% CI
c i . l o <− o u t $ f i t − t s t a r * out$se . f i t # Estimate l e s s std e r r o r
c i . hi <− o u t $ f i t + t s t a r * out$se . f i t # Estimate plus std e r r o r
C I i n f o <− cbind ( Lower= c i . lo , Estimate= o u t $ f i t , Upper= c i . hi )
# Assign t o a new o b j e c t
247
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
CIinfo
#12.7 An approximate CI f o r E[FEV] − now i s in l i t r e s ( a unit we understand )
exp ( C I i n f o )
#12.8 This idea can be extended t o compute the c o n f i d e n c e i n t e r v a l s f o r
18 year o l d female smokers f o r varying heights
newHt <− seq ( min ( lungcap$Ht ) , max( lungcap$Ht ) , by =2)
newlogFEV <− p r e d i c t ( LC. m1,
se . f i t =TRUE,
newdata=data . frame ( Age=18 ,
Ht=newHt ,
Gender="F " ,
Smoke="Smoker " ) )
c i . l o <− exp ( newlogFEV$fit − t s t a r * newlogFEV$se . f i t )
c i . hi <− exp ( newlogFEV$fit + t s t a r * newlogFEV$se . f i t )
#12.9 Notice that the i n t e r v a l s do not have the same width over the whole range o
cbind ( Ht=newHt ,
FEVhat=exp ( newlogFEV$fit ) ,
SE=newlogFEV$se . f i t ,
Lower= c i . lo ,
Upper= c i . hi ,
CI . Width= c i . hi − c i . l o )
#13.ANALYSIS OF VARIANCE FOR REGRESSION MODEL
#NOTES:
# 1 . R e c a l l that DATA = FIT + RESIDUAL, so SST = SSReg + RSS,
and R−squared = SSReg / SST = 1 − RSS / SST
# 2 . So now compute RSS and SST , remembering that y = LOG(FEV)
mu <− f i t t e d ( LC.m1 ) ; RSS <− sum( ( y − mu)^2 )
SST <− sum( ( y − mean ( y ) )^2 )
c (RSS=RSS, SST=SST , SSReg = SST−RSS) # c i s the combine or concatenate f u n c t i o n
248
Electronic copy available at: https://ssrn.com/abstract=3863563
8.5. ILLUSTRATION OF USING GLM DATA SET IN R
#13.1 Can manually compute R−squared
R2 <− 1 − RSS / SST
#13.2 Next wish t o output the r e s u l t s onscreen
c ( " Output R2" = summary (LC.m1) $r . squared , " Computed R2" = R2 ,
" adj R2" = summary (LC.m1) $adj . r . squared )
#NOTES: Can compare these r e s u l t s t o summary (LC.m1) e a r l i e r
#13.3 Can a l s o compute the F− s t a t i s t i c ,
which i n c l u d e s the numerator and denominator df
summary (LC.m1) $ f s t a t i s t i c
249
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.6
GLM logistic regression
#GLM LOGISTIC REGRESSION EXAMPLE_EXERCISE
# 1 . Fair ’ s A f f a i r s
#1.1 I n s t a l l and load the packages
i n s t a l l . packages ( "AER" )
i n s t a l l . packages ( " lmtest " )
i n s t a l l . packages ( " zoo " )
i n s t a l l . packages ( " sandwich " )
i n s t a l l . packages ( " s u r v i v a l " )
l i b r a r y ( lmtest )
l i b r a r y ( zoo )
l i b r a r y ( sandwich )
library ( survival )
l i b r a r y (AER)
#1.2 Load the data
data ( A f f a i r s , package ="AER" )
#1.3 Let ’ s look at some d e s c r i p t i v e data
summary ( A f f a i r s )
table ( Affairs$affairs )
#NOTES: From these s t a t i s t i c s , you can see that that 52% o f respondents were female
# that 72% had children , and that the median age f o r the sample was 32 years .
#With regard t o the response v a r i a b l e ,
#75% o f respondents reported not engaging in an i n f i d e l i t y
# in the past year ( 4 5 1 / 6 0 1 ) .
#The l a r g e s t number o f encounters reported was 12 ( 6 % ) .
#1.4 Although the number o f i n d i s c r e t i o n s was recorded ,
#our i n t e r e s t here i s in the binary outcome ( had an a f f a i r / did not have an a f f a i r ) .
#We can transform a f f a i r s i n t o a dichotomous f a c t o r
# c a l l e d y n a f f a i r with the f o l l o w i n g code .
A f f a i r s $ y n a f f a i r [ A f f a i r s $ a f f a i r s > 0 ] <− 1
250
Electronic copy available at: https://ssrn.com/abstract=3863563
8.6. GLM LOGISTIC REGRESSION
A f f a i r s $ y n a f f a i r [ A f f a i r s $ a f f a i r s == 0 ] <− 0
A f f a i r s $ y n a f f a i r <− f a c t o r ( A f f a i r s $ y n a f f a i r ,
l e v e l s =c ( 0 , 1 ) ,
l a b e l s =c ( " No " , " Yes " ) )
#1.5 Can view the r e s u l t s with the t a b l e ( ) f u n c t i o n
table ( Affairs$ynaffair )
#1.6 This dichotomous f a c t o r can now be used as the outcome v a r i a b l e
# in a l o g i s t i c r e g r e s s i o n model
f i t . f u l l <− glm ( y n a f f a i r ~ gender + age + yearsmarried + c h i l d r e n +
r e l i g i o u s n e s s + education + occupation +rating ,
data= A f f a i r s , family=binomial ( ) )
#1.7 Call the summary output
summary ( f i t . f u l l )
#NOTES: From the p−values f o r the r e g r e s s i o n c o e f f i c i e n t s ( l a s t column ) ,
#we can see that gender , presence o f children , education ,
#and occupation may not make a s i g n i f i c a n t c o n t r i b u t i o n t o the equation
# (we cannot r e j e c t the hypothesis that the parameters are 0 ) .
#1.8 Let us f i t a second equation without them
#and t e s t whether t h i s reduced model f i t s the data as well
f i t . reduced <− glm ( y n a f f a i r ~ age + yearsmarried + r e l i g i o u s n e s s +
rating , data= A f f a i r s , family=binomial ( ) )
summary ( f i t . reduced )
#NOTES:
# 1 . Each r e g r e s s i o n c o e f f i c i e n t in the reduced model
#is statistically significant (p < .05)
# 2 . Because the two models are nested ( f i t . reduced i s a subset o f f i t . f u l l ) ,
#you can use the anova ( ) f u n c t i o n t o compare them .
251
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#1.9 For g e n e r a l i z e d l i n e a r models ,
#we w i l l want a chi −square v e r s i o n o f t h i s t e s t :
anova ( f i t . reduced , f i t . f u l l , t e s t =" Chisq " )
#NOTES: The non−s i g n i f i c a n t chi −square value ( p = 0 . 2 1 )
# suggests that the reduced model with f o u r p r e d i c t o r s f i t s
#as well as the f u l l model with nine p r e d i c t o r s , r e i n f o r c i n g our b e l i e f that gender
# children , education , and occupation do not add s i g n i f i c a n t l y t o
#the p r e d i c t i o n above and beyond the other v a r i a b l e s in the equation .
# Therefore , we can base our i n t e r p r e t a t i o n s on the simpler model .
# 2 . INTERPRETING THE MODEL PARAMETERS
#2.1 Let us examine the r e g r e s s i o n c o e f f i c i e n t s with the c o e f ( ) f u n c t i o n
c o e f ( f i t . reduced )
#NOTES:
# 1 . In a l o g i s t i c r e g r e s s i o n ,
#the response being modeled i s the l o g ( odds ) that Y = 1 .
#The r e g r e s s i o n c o e f f i c i e n t s g i v e the change in l o g ( odds )
# in the response f o r a unit change in the p r e d i c t o r v a r i a b l e ,
# holding a l l other p r e d i c t o r v a r i a b l e s constant .
# 2 . Because l o g ( odds ) are d i f f i c u l t t o i n t e r p r e t ,
#you can exponentiate them t o put the r e s u l t s on an odds s c a l e
#− thus e a s i e r f o r us t o understand and i n t e r p r e t ( using r a t i o s )
#2.2 Simply use the exp ( ) or ex po ne nt ia tio n f u n c t i o n
exp ( c o e f ( f i t . reduced ) )
#NOTES:
# 1 . Now you can see that the odds o f an e x t r a m a r i t a l encounter are in cre ased
#by a f a c t o r o f 1.106 f o r a one−year i n c r e a s e in years married
# ( holding age , r e l i g i o u s n e s s , and marital r a t i n g constant ) .
#Conversely , the odds o f an e x t r a m a r i t a l a f f a i r are m u l t i p l i e d by a f a c t o r o f
#0.965 f o r every year i n c r e a s e in age .
#The odds o f an e x t r a m a r i t a l a f f a i r i n c r e a s e with years married and decrease
#with age , r e l i g i o u s n e s s , and marital r a t i n g .
252
Electronic copy available at: https://ssrn.com/abstract=3863563
8.6. GLM LOGISTIC REGRESSION
#Because the p r e d i c t o r v a r i a b l e s cannot equal 0 ,
#the i n t e r c e p t i s not meaningful in t h i s case .
#2.2 I f desired , you can use the c o n f i n t ( ) f u n c t i o n
# t o obtain c o n f i d e n c e i n t e r v a l s f o r the c o e f f i c i e n t s .
#For example , exp ( c o n f i n t ( f i t . reduced ) ) would p r i n t 95% c o n f i d e n c e i n t e r v a l s
# f o r each o f the c o e f f i c i e n t s on an odds s c a l e .
exp ( c o n f i n t ( f i t . reduced ) )
#NOTES: Finally ,
#a one−unit change in a p r e d i c t o r v a r i a b l e may not be i n h e r e n t l y i n t e r e s t i n g .
#For binary l o g i s t i c r e g r e s s i o n , the change in the odds o f the higher value
#on the response v a r i a b l e f o r an n unit change in a p r e d i c t o r v a r i a b l e i s exp (
# 3 . ASSESSING THE IMPACT OF PREDICTORS ON THE PROBABILITY OF AN OUTCOME
#NOTES:
# 1 . I t i s e a s i e r t o think in terms o f p r o b a b i l i t i e s than odds .
#Can use the p r e d i c t ( ) f u n c t i o n t o observe the impact o f varying the l e v e l s o f
#a p r e d i c t o r v a r i a b l e on the p r o b a b i l i t y o f the outcome .
# 2 . The f i r s t step i s t o c r e a t e an a r t i f i c i a l dataset c o n t a i n i n g the values o f
#the p r e d i c t o r v a r i a b l e s are i n t e r e s t e d in .
#Then can use t h i s a r t i f i c i a l dataset with
#the p r e d i c t ( ) f u n c t i o n t o p r e d i c t the p r o b a b i l i t i e s
# o f the outcome event o c c u r r i n g f o r these values .
#3.1 Apply t h i s s t r a t e g y t o assess the impact o f marital r a t i n g s
#on the p r o b a b i l i t y o f having an e x t r a m a r i t a l a f f a i r .
t e s t d a t a <− data . frame ( r a t i n g =c ( 1 , 2 , 3 , 4 , 5 ) ,
age=mean ( A f f a i r s $ a g e ) ,
yearsmarried=mean ( Affairs$yearsmarried ) ,
r e l i g i o u s n e s s =mean ( A f f a i r s $ r e l i g i o u s n e s s ) )
#3.2 No c o n s o l e output , so type the name o f the newly created o b j e c t
testdata
#3.3 Next , use the t e s t dataset and p r e d i c t i o n equation t o obtain p r o b a b i l i t i e s :
253
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
testdata$prob <− p r e d i c t ( f i t . reduced , newdata=testdata , type =" response " )
testdata
#NOTES: From these r e s u l t s , can see that the p r o b a b i l i t y
# o f an e x t r a m a r i t a l a f f a i r decreases from 0.53
#when the marriage i s rated 1=very unhappy t o 0.15
#when the marriage i s rated 5=very happy
# ( holding age , years married , and r e l i g i o u s n e s s constant ) .
#3.4 Now look at the impact o f age ( from age 17 t o 57 , in decadal increments ) :
t e s t d a t a <− data . frame ( r a t i n g =mean ( A f f a i r s $ r a t i n g ) ,
age=seq ( 1 7 , 57 , 1 0 ) ,
yearsmarried=mean ( Affairs$yearsmarried ) ,
r e l i g i o u s n e s s =mean ( A f f a i r s $ r e l i g i o u s n e s s ) )
testdata
#3.5 Now we choose p r o b a b i l i t i e s
testdata$prob <− p r e d i c t ( f i t . reduced , newdata=testdata , type =" response " )
testdata
#NOTES: Here , can see that as age i n c r e a s e s from 17 t o 57 ,
#the p r o b a b i l i t y o f an e x t r a m a r i t a l encounter decreases from 0.34 t o 0 . 1 1 ,
# holding the other v a r i a b l e s constant ( C e t e r i s Paribus ) .
#Using t h i s approach ,
#you can e x p l o r e the impact o f each p r e d i c t o r v a r i a b l e on the outcome .
# 4 . OVERDISPERSION
#4.1 Applying t h i s ( phi ) r a t i o t o the A f f a i r s example , can have
deviance ( f i t . reduced ) / df . r e s i d u a l ( f i t . reduced )
#NOTES:
# 1 . This i s c l o s e t o 1 , suggesting no o v e r d i s p e r s i o n .
# 2 . Can a l s o t e s t f o r o v e r d i s p e r s i o n .
#To do t h i s , f i t the model twice ,
#but in the f i r s t i n s t a n c e
use family =" binomial "
#and in the second i n s t a n c e use family =" quasibinomial " .
# I f the glm ( ) o b j e c t returned in the f i r s t case i s c a l l e d f i t
254
Electronic copy available at: https://ssrn.com/abstract=3863563
8.6. GLM LOGISTIC REGRESSION
#and the o b j e c t returned in the second case i s c a l l e d f i t . od ,
#then pchisq ( summary ( f i t . od ) $ d i s p e r s i o n * f i t $ d f . r e s i d u a l , f i t $ d f . r e s i d u a l , lower = F )
# provides the p−value f o r t e s t i n g the n u l l hypothesis H0: ? ? = 1
# versus the a l t e r n a t i v e hypothesis H1 : . . . 1 . I f p i s small ( say , l e s s than 0 . 0 5 ) ,
#can r e j e c t the n u l l hypothesis .
#4.2 Applying t h i s t o the A f f a i r s dataset , can have
f i t <− glm ( y n a f f a i r ~ age + yearsmarried + r e l i g i o u s n e s s +
rating , family = binomial ( ) , data = A f f a i r s )
#4.3 Now second model
f i t . od <− glm ( y n a f f a i r ~ age + yearsmarried + r e l i g i o u s n e s s +
rating , family = quasibinomial ( ) , data = A f f a i r s )
#4.4 Now apply the t e s t
pchisq ( summary ( f i t . od ) $ d i s p e r s i o n * f i t $ d f . r e s i d u a l ,
f i t $ d f . r e s i d u a l , lower = F )
#NOTES: The r e s u l t i n g p−value ( 0 . 3 4 ) i s c l e a r l y not s i g n i f i c a n t ( p > 0 . 0 5 ) ,
# strengthening our b e l i e f that o v e r d i s p e r s i o n i s not a problem .
#We s h a l l return t o the i s s u e o f o v e r d i s p e r s i o n when we d i s c u s s Poisson r e g r e s s i o n
255
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.7
Poisson regression
#POISSON REGRESSION ILLUSTRATION
#1.1 I n s t a l l the robust package and load the l i b r a r y
i n s t a l l . packages ( " robust " )
i n s t a l l . packages ( " f i t . models " )
l i b r a r y ( robust )
l i b r a r y ( f i t . models )
#1.2 F i r s t take a look at the summary s t a t i s t i c s
data ( breslow . dat , package =" robust " )
names ( breslow . dat )
summary ( breslow . dat [ c ( 6 , 7 , 8 , 1 0 ) ] )
#1.3 Note that although there are 12 v a r i a b l e s in the dataset ,
#are l i m i t i n g the a t t e n t i o n t o the 4 d e s c r i b e d e a r l i e r .
#Both the b a s e l i n e and post −randomization number o f s e i z u r e s are highly skewed .
#1.4 look at the response v a r i a b l e in more d e t a i l .
#The f o l l o w i n g code produces the graphs
opar <− par ( no . readonly=TRUE)
par ( mfrow=c ( 1 , 2 ) )
attach ( breslow . dat )
h i s t (sumY, breaks =20 , xlab =" Seizure Count " ,
main=" D i s t r i b u t i o n o f Seizures " )
b o x p l o t (sumY ~ Trt , xlab =" Treatment " , main="Group Comparisons " )
par ( opar )
#NOTES: Can c l e a r l y see the skewed nature o f the dependent v a r i a b l e
#and the p o s s i b l e presence o f o u t l i e r s .
#At f i r s t glance ,
#the number o f s e i z u r e s in the drug c o n d i t i o n appears t o be smaller
#and has a smaller variance .
# (You would expect a smaller variance t o accompany a smaller mean
#with Poisson d i s t r i b u t e d data . ) Unlike standard OLS r e g r e s s i o n ,
# t h i s h e t e r o g e n e i t y o f variance i s not a problem in Poisson r e g r e s s i o n .
256
Electronic copy available at: https://ssrn.com/abstract=3863563
8.7. POISSON REGRESSION
#1.5 The next step i s t o f i t the Poisson r e g r e s s i o n :
f i t <− glm (sumY ~ Base + Age + Trt , data=breslow . dat , family=poisson ( ) )
summary ( f i t )
#NOTES: The output p rovid es the deviances , r e g r e s s i o n parameters ,
#and standard e r r o r s ,
#and t e s t s that these parameters are 0 .
#Note that each o f the p r e d i c t o r v a r i a b l e s i s s i g n i f i c a n t at the p < 0.05 l e v e l .
# 2 . INTERPRETING THE MODEL PARAMETERS
#2.1 The model c o e f f i c i e n t s are obtained
#using the c o e f ( ) f u n c t i o n or by examining the C o e f f i c i e n t s t a b l e
# in the summary ( ) f u n c t i o n output :
coef ( f i t )
#NOTES:
# 1 . In a Poisson r e g r e s s i o n ,
#the dependent v a r i a b l e being modeled i s the l o g o f the c o n d i t i o n a l mean l o g e ( y ) .
#The r e g r e s s i o n parameter 0.0227 f o r Age i n d i c a t e s that a one year i n c r e a s e in age
# i s a s s o c i a t e d with a 0.03 i n c r e a s e in the l o g mean number o f s e i z u r e s ,
# holding b a s e l i n e s e i z u r e s and treatment c o n d i t i o n constant .
#The i n t e r c e p t i s the l o g mean number o f s e i z u r e s
#when each o f the p r e d i c t o r s equals 0 .
#Because you cannot have a zero age and none o f the p a r t i c i p a n t s
#had a zero number o f b a s e l i n e s e i z u r e s , the i n t e r c e p t i s not meaningful
# in t h i s case .
# 2 . I t i s usually much e a s i e r t o i n t e r p r e t the r e g r e s s i o n c o e f f i c i e n t s
# in the o r i g i n a l s c a l e o f the dependent v a r i a b l e
# ( number o f s e i z u r e s , rather than l o g number o f s e i z u r e s ) .
#2.2 To accomplish t h i s , exponentiate the c o e f f i c i e n t s :
exp ( c o e f ( f i t ) )
#NOTES:
# 1 . Now see that a one−year i n c r e a s e in age m u l t i p l i e s
#the expected number o f s e i z u r e s by 1 . 0 2 3 , holding the other v a r i a b l e s constant .
257
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#This means that i ncr ease d age i s a s s o c i a t e d with higher numbers o f s e i z u r e s .
# 2 . More important , a one−unit change in Trt
# ( that i s , moving from placebo t o progabide )
# m u l t i p l i e s the expected number o f s e i z u r e s by 0 . 8 6 .
#Would expect a 20% decrease in the number o f s e i z u r e s f o r
#the drug group compared with the placebo group ,
# holding b a s e l i n e number o f s e i z u r e s and age constant .
# 3 . I t i s important t o remember that ,
# l i k e the exponentiated parameters in l o g i s t i c r e g r e s s i o n ,
#the exponentiated parameters in the Poisson model have a m u l t i p l i c a t i v e rather than
#an a d d i t i v e e f f e c t on the response v a r i a b l e .
#Also , as with l o g i s t i c r e g r e s s i o n , must evaluate your model f o r o v e r d i s p e r s i o n .
# 3 . OVERDISPERSION
#2.1 Perform the r a t i o t e s t
deviance ( f i t ) / df . r e s i d u a l ( f i t )
#NOTES: which i s c l e a r l y much l a r g e r than 1 .
#2.2 The qcc package provi des a t e s t f o r o v e r d i s p e r s i o n in the Poisson case .
# (Be sure t o download and i n s t a l l t h i s package b e f o r e f i r s t use . )
i n s t a l l . packages ( " qcc " )
#2.3 Test f o r o v e r d i s p e r s i o n in the s e i z u r e data using the f o l l o w i n g code :
l i b r a r y ( qcc )
qcc . o v e r d i s p e r s i o n . t e s t ( breslow . dat$sumY , type =" poisson " )
#NOTES: Not s u r p r i s i n g l y , the s i g n i f i c a n c e t e s t has a p−value l e s s than 0 . 0 5 ,
# s t r o n g l y suggesting the presence o f o v e r d i s p e r s i o n .
#2.4 Can s t i l l f i t a model t o the data using the glm ( ) function ,
#by r e p l a c i n g family =" poisson " with family =" quasipoisson " .
#Doing so i s analogous t o the approach t o l o g i s t i c r e g r e s s i o n
#when o v e r d i s p e r s i o n i s present
f i t . od <− glm (sumY ~ Base + Age + Trt , data=breslow . dat ,
family=quasipoisson ( ) )
258
Electronic copy available at: https://ssrn.com/abstract=3863563
8.7. POISSON REGRESSION
summary ( f i t . od )
#NOTES: Notice that the parameter estimates in the quasi−Poisson approach
#are i d e n t i c a l t o those produced by the Poisson approach .
#The standard e r r o r s are much l a r g e r , though .
#In t h i s case , the l a r g e r standard e r r o r s have l e d t o p−values f o r Trt ( and Age )
# that are g r e a t e r than 0 . 0 5 .
#When you take o v e r d i s p e r s i o n i n t o account ,
# there i s i n s u f f i c i e n t evidence t o d e c l a r e that the drug regimen reduces s e i z u r e
# counts more than r e c e i v i n g a placebo , a f t e r c o n t r o l l i n g f o r b a s e l i n e s e i z u r e
# r a t e and age .
259
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.8
Diagnostics and model building
#DIAGNOSTICS AND MODEL−BUILDING_EXERCISE
#NOTES: Use a systematic component i n v o l v i n g Ht , Gender and Smoke .
# ( p r e f e r r i n g Ht over Age as Ht i s a p h y s i c a l t r a i t ) .
#1.1 I n s t a l l and load the l i b r a r y , and load the data frame once more
i n s t a l l . packages ( " GLMsData " )
l i b r a r y (GLMsData)
data ( lungcap )
#1.2 Need t o r e c r e a t e the q u a l i t a t i v e o b j e c t ( binary v a r i a b l e )
# d i s t i n c t i o n between Smoker vs Non−smoker
lungcap$Smoke <− f a c t o r ( lungcap$Smoke ,
l e v e l s =c ( 0 , 1 ) ,
l a b e l s =c ( " Non−smoker " , " Smoker " ) )
#1.3 d e l i b e r a t e l y run a bad LM model s p e c i f i c a t i o n
LC. lm <− lm ( FEV ~ Ht + Gender + Smoke , data=lungcap )
#1.4 Next w i l l compute the r e s i d u a l s f o r t h i s model
# 1 . 4 . 1 F i r s t the raw r e s i d u a l s
r e s i d . raw <− r e s i d ( LC. lm )
# 1 . 4 . 2 Then the standardized r e s i d u a l s
r e s i d . std <− rstandard ( LC. lm )
# 1 . 4 . 3 Note no output t o c o n s o l e windows , so use
c ( Raw=var ( r e s i d . raw ) , Standardized=var ( r e s i d . std ) )
#NOTES: can c l e a r l y see that the standardized r e s i d u a l s have a variance c l o s e t o 1 ,
#as expected
# 2 . LEVERAGES FOR LINEAR REGRESSION MODELS
#2.1 For the Poor model d efined above ,
#the l e v e r a g e s are found using hatvalues ( ) f u n c t i o n
h <− hatvalues ( LC. lm )
260
Electronic copy available at: https://ssrn.com/abstract=3863563
8.8. DIAGNOSTICS AND MODEL BUILDING
#2.2 Wish t o examine the l a r g e s t two l e v e r a g e s
s o r t ( h , decreasing=TRUE) [ 1 : 2 ]
#2.3 The two l a r g e s t l e v e r a g e s are f o r Observations 629 and 631.
#Compare these l e v e r a g e s t o the mean value o f the l e v e r a g e s
#2.4 Mean l e v e r a g e
mean ( h ) ; length ( c o e f (LC. lm ) ) / length ( lungcap$FEV )
#2.5 Divide the two l a r g e s t l e v e r a g e s by the mean
s o r t ( h , decreasing=TRUE) [ 1 : 2 ] / mean ( h )
#2.6 Observations 629 and 631 are many times g r e a t e r than
#the mean value o f the l e v e r a g e s .
#Note that both o f these l a r g e l e v e r a g e s correspond t o male smokers :
s o r t . h <− s o r t ( h , decreasing=TRUE, index . return=TRUE)
#2.7 Provide the index where these occur
l a r g e . h <− s o r t . h$ix [ 1 : 2 ]
lungcap [ l a r g e . h , ]
#2.8 can check with a p l o t o f FEV against Ht f o r j u s t male smokers then :
p l o t ( FEV ~ Ht , main="Male smokers " ,
data=subset ( lungcap , Gender=="M" & Smoke=="Smoker " ) ,
# Only male smokers l a s =1 , xlim=c ( 5 5 , 7 5 ) , ylim=c ( 0 , 5 ) ,
xlab =" Height ( inches ) " , ylab ="FEV ( L ) " )
#2.9 Large values − we can f i l l in the c i r c u l a r p l o t symbols
p o i n t s ( FEV[ l a r g e . h ] ~ Ht [ l a r g e . h ] , data=lungcap , pch =19)
#2.10 Now add a legend e x p l a i n i n g what these f i l l e d in c i r c l e s
#are t o teh bottom r i g h t corner
legend ( " bottomright " , pch =19 , legend=c ( " Large l e v e r a g e p o i n t s " ) )
#NOTES: The two l a r g e s t l e v e r a g e s correspond t o the two unusual o b s e r v a t i o n s
# in the bottom l e f t corner o f the p l o t
# 3 . RESIDUAL PLOTS AGAINST x i : LINEARITY
261
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#3.1 B r i e f l y return t o PowerPoint s l i d e s ,
# continue with the LC. lm model example from b e f o r e .
#Assume data were c o l l e c t e d so that the responses are independent .
#Then , p l o t s o f r e s i d u a l s against the c o v a r i a t e can be cr eated :
#3.2 P l o t std r e s i d u a l s against Ht
s c a t t e r . smooth ( rstandard ( LC. lm ) ~ lungcap$Ht , c o l =" grey " ,
l a s =1 , ylab =" Standardized r e s i d u a l s " , xlab =" Height ( inches ) " )
#NOTES: The p l o t s o f r e s i d u a l s against height are s l i g h t l y non−l i n e a r ,
#and have i n c r e a s i n g variance . This suggests that the model i s poor .
#Of course , l i n e a r i t y i s not r e l e v a n t f o r gender or smoking status ,
#as these v a r i a b l e s take only two l e v e l s .
# 4 . PARTIAL RESIDUAL PLOTS
#4.1 Compute the p a r t i a l r e s i d u a l s using r e s i d ( ) f u n c t i o n
p a r t i a l . r e s i d <− r e s i d ( LC. lm , type =" p a r t i a l " )
#4.2 No response t o the c o n s o l e window ,
# so use the head f u n c t i o n t o examine f i r s t 6 l i n e s
head ( p a r t i a l . r e s i d )
#4.3 The e a s i e s t way t o produce the p a r t i a l r e s i d u a l p l o t s i s t o use termplot ( ) .
#Do so here t o produce the p a r t i a l r e s i d u a l s p l o t f o r Ht only
termplot ( LC. lm , p a r t i a l . r e s i d =TRUE, terms ="Ht " , l a s =1)
#NOTES:
# 1 . Termplot ( ) a l s o shows the i d e a l l i n e a r r e l a t i o n s h i p in the p l o t s .
#The p a r t i a l r e s i d u a l p l o t f o r Ht shows non−l i n e a r i t y ,
#again suggesting the use o f y =
E[ l o g ( f e v ) ] as the response v a r i a b l e
# 2 . The r e l a t i o n s h i p between FEV and Ht appears q u i t e strong a f t e r
# a d j u s t i n g f o r the other explanatory v a r i a b l e s .
#Note that the s l o p e o f the simple r e g r e s s i o n l i n e
# i s equal t o the c o e f f i c i e n t in the f u l l model .
#4.4 For example , compare the r e g r e s s i o n c o e f f i c i e n t s f o r Ht :
262
Electronic copy available at: https://ssrn.com/abstract=3863563
8.8. DIAGNOSTICS AND MODEL BUILDING
c o e f ( summary (LC. lm ) )
#4.5 Can run a quick check f o r the Height ( Ht ) v a r i a b l e
lm . Ht <− lm ( p a r t i a l . r e s i d [ , 1]~ lungcap$Ht )
c o e f ( summary ( lm . Ht ) )
#NOTES: The c o e f f i c i e n t s f o r Ht are e x a c t l y the same .
#The f u l l r e g r e s s i o n g i v e s l a r g e r standard e r r o r s than the simple l i n e a r r e g r e s s i o n
#however , because the l a t t e r over−estimates the r e s i d u a l degrees o f freedom
# 5 . PLOT RESIDUALS AGAINST MU−hat : CONSTANT VARIANCE
#5.1 P l o t std r e s i d u a l s against the f i t t e d values
s c a t t e r . smooth ( rstandard ( LC. lm ) ~ f i t t e d ( LC. lm ) , c o l =" grey " ,
l a s =1 , ylab =" Standardized r e s i d u a l s " , xlab =" F i t t e d values " )
#NOTES:
# 1 . This p l o t shows that the p l o t o f r e s i d u a l s against f i t t e d values
#has a variance that i s not constant
# 2 . In other words , there appears t o be an i n c r e a s i n g mean−variance r e l a t i o n s h i p .
# 3 . The p l o t a l s o shows non−l i n e a r i t y ,
#again suggesting that the model can be improved .
# 6 . Q−Q PLOTS AND NORMALITY
#6.1 Q−Q p r o b a b i l i t y p l o t
qqnorm ( rstandard ( LC. lm ) , l a s =1 , pch =19)
#6.2 Now add a r e f e r e n c e l i n e t o o
q q l i n e ( rstandard ( LC. lm ) )
#NOTES:
# 1 . The Q−Q p l o t suggests the normality assumption i s suspect
# 2 . The d i s t r i b u t i o n o f r e s i d u a l s appears t o have heavier t a i l s than the normal
263
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# d i s t r i b u t i o n in both d i r e c t i o n s , because the r e s i d u a l s curve above the l i n e on
#the r i g h t and below the l i n e on the l e f t .
#The p l o t a l s o shows a number o f l a r g e r e s i d u a l s ,
#both p o s i t i v e and negative , suggesting the model can be improved .
# 7 . OUTLIERS AND INFLUENTIAL OBSERVATIONS
#NOTES: Using r , Studentized r e s i d u a l s are e a s i l y found using rstudent ( )
#7.1 For the lungcap data , the r e s i d u a l p l o t shows no o u t l i e r s
# ( but does shows some l a r g e r e s i d u a l s , both p o s i t i v e and negative ) ,
# so r ’ and r ’ ’ are expected t o be s i m i l a r
summary ( cbind ( Standardized = rstandard (LC. lm ) ,
Studentized = rstudent (LC. lm ) ) )
#7.2 For the model LC. lm f i t t e d t o the lungcap data the example e a r l i e r ,
#the Studentized r e s i d u a l s can be computed by manually d e l e t i n g each o b s e r v a t i o n .
#7.3 For example , d e l e t i n g Observation 1 and
# r e f i t t i n g produces the Studentized r e s i d u a l f o r Observation 1 :
# 7 . 3 . 1 F i t the model * without * Observation 1 :
LC. no1 <− lm ( FEV ~ Ht + Gender + Smoke ,
data=lungcap [ − 1 , ] )
#NOTES: The negative index * removes * row 1
#7.4 The f i t t e d value f o r Observation 1 , from the o r i g i n a l model :
mu <− f i t t e d ( LC. lm ) [ 1 ]
#7.5 The estimate o f s from the new model , without Obs . 1 :
s <− summary (LC. no1 ) $sigma
#7.6 Hat value , f o r Observation 1
h <− hatvalues ( LC. lm ) [ 1 ]
r e s i d . stud <− ( lungcap$FEV [ 1 ] − mu ) / ( s * s q r t (1 −h ) )
264
Electronic copy available at: https://ssrn.com/abstract=3863563
8.8. DIAGNOSTICS AND MODEL BUILDING
r e s i d . stud
#7.7 There i s an easy way !
rstudent (LC. lm ) [ 1 ]
# 8 . INFLUENTIAL OBSERVATIONS
#8.1 Return t o model LC. lm
# 8 . 1 . 1 The o b s e r v a t i o n s with the s m a l l e s t and l a r g e s t values o f Cook ’ s d i s t a n c e are :
# Largest D
cd . max <− which . max( cooks . d i s t a n c e (LC. lm ) )
# Smallest D
cd . min <− which . min ( cooks . d i s t a n c e (LC. lm ) )
#8.2 Now examine output
c ( Min . Cook = cd . min , Max . Cook = cd . max)
#8.3 The values o f d f f i t s ,
#cv and Cook ’ s d i s t a n c e f o r these o b s e r v a t i o n s can be found as f o l l o w s :
out <− cbind ( DFFITS= d f f i t s (LC. lm ) ,
Cooks . d i s t a n c e =cooks . d i s t a n c e (LC. lm ) ,
Cov . r a t i o = c o v r a t i o (LC. lm ) )
#8.4 These s t a t i s t i c s f o r the o b s e r v a t i o n s cd . max and cd . min are :
# 8 . 4 . 1 SHow the values f o r these obs only
round ( out [ c ( cd . min , cd . max ) , ] , 5 )
#NOTES:
# 1 . From these three measures ,
# Observation 613 i s more i n f l u e n t i a l
#than Observation 69 according t o d f f i t s and Cook ’ s d i s t a n c e ( but not CR ) .
# 2 . Now examine i n f l u e n c e o f Observation 613 and 69 on each o f the r e g r e s s i o n parameters
#8.5 Show DBETAS f o r cd . min
d f b e t a s (LC. lm ) [ cd . min , ]
265
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#8.6 Show DBETAS f o r cd . max
d f b e t a s (LC. lm ) [ cd . max , ]
#Notea Omitting Observation 69 ( cd . min ) makes almost
#no d i f f e r e n c e t o the r e g r e s s i o n c o e f f i c i e n t s .
# Observation 613 i s c l e a r l y more i n f l u e n t i a l than Observation 69 , as expected .
#8.7 The R f u n c t i o n i n f l u e n c e . measures ( ) i s used t o
# i d e n t i f y p o t e n t i a l l y i n f l u e n t i a l o b s e r v a t i o n s according t o Rxfixxs c r i t e r i a :
LC. im <− i n f l u e n c e . measures ( LC. lm ) ; names (LC. im )
#8.8 The o b j e c t LC. im c o n t a i n s the i n f l u e n c e s t a t i s t i c s ( as LC. im$infmat ) ,
#and whether or not they are i n f l u e n t i a l according t o rs c r i t e r i a (LC. im$is . i n f ) :
# 8 . 8 . 1 Show f o r the f i r s t few obs only
head ( round (LC. im$infmat , 3 ) )
# 8 . 8 . 2 Now the other
head ( LC. im$is . i n f )
# 8 . 8 . 3 To determine how many e n t r i e s in the columns o f
#LC. im$is . i n f are TRUE, sum over the columns
# ( t h i s works because R t r e a t s FALSE as 0 and TRUE as 1 ) :
colSums ( LC. im$is . i n f )
#NOTES:
# 1 . Seven o b s e r v a t i o n s have high leverage ,
#as i d e n t i f i e d by the column l a b e l e d hat ,
#56 o b s e r v a t i o n s are i d e n t i f i e d by the cov ari an ce r a t i o as i n f l u e n t i a l ,
#but Cook ’ s d i s t a n c e does not i d e n t i f y any o b s e r v a t i o n as i n f l u e n t i a l .
# 2 . Also determine how many c r i t e r i a d e c l a r e o b s e r v a t i o n s
#as i n f l u e n t i a l by summing the r e l e v a n t columns o f LC. im$is . i n f over the rows :
#8.9 Omitting l e v e r a g e s ( c o l 8 )
t a b l e ( rowSums ( LC. im$is . i n f [ , −8] ) )
266
Electronic copy available at: https://ssrn.com/abstract=3863563
8.8. DIAGNOSTICS AND MODEL BUILDING
#NOTES: This shows that most o b s e r v a t i o n s are not d ecla red
# i n f l u e n t i a l on any o f the c r i t e r i a ,
#and 54 o b s e r v a t i o n s d eclar ed as i n f l u e n t i a l on j u s t one c r i t e r i o n .
#8.10 For Observations 69 and 613 e x p l i c i t l y :
LC. im$is . i n f [ c ( cd . min , cd . max ) , ]
#NOTES: Observation 613 i s s i g n i f i c a n t l y i n f l u e n t i a l based on d f f i t s and cv .
#8.11 p l o t o f these i n f l u e n c e d i a g n o s t i c s i s o f t e n useful ,
#using type ="h " t o draw histogram−l i k e ( or high−d e n s i t y ) p l o t s :
# 8 . 1 1 . 1 Cook ’ s Distance
p l o t ( cooks . d i s t a n c e ( LC. lm ) , type ="h " , main="Cook ’ s d i s t a n c e " ,
ylab ="D" , xlab =" Observation number " , l a s =1 )=
# 8 . 1 1 . 2 DFFITS
p l o t ( d f f i t s ( LC. lm ) , type ="h " , main="DFFITS " ,
ylab ="DFFITS " , xlab =" Observation number " , l a s =1 )
# 8 . 1 1 . 3 DFBETAS f o r beta_2 only ( that i s , column 3 )
d f b i <− 2
p l o t ( d f b e t a s ( LC. lm ) [ , d f b i + 1 ] , type ="h " , main="DFBETAS f o r beta2 " ,
ylab ="DFBETAS" , xlab =" Observation number " , l a s =1 )
# 9 . TRANSFORMING THE RESPONSE
#9.1 Start with the square r o o t transformation
LC. s q r t <− update ( LC. lm , s q r t (FEV) ~ . )
#9.2 P l o t f o r square−r o o t transformation
s c a t t e r . smooth ( rstandard (LC. s q r t )~ f i t t e d (LC. s q r t ) , l a s =1 , c o l =" grey " ,
ylab =" Standardized r e s i d u a l s " , xlab =" F i t t e d values " ,
main=" Square−r o o t transformation " )
#NOTES: This transformation produces s l i g h t l y i n c r e a s i n g variance .
#9.3 So t r y the next transformation on the ladder ,
#the commonly−used l o g a r i t h m i c transformation :
267
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
LC. l o g <− update ( LC. lm , l o g (FEV) ~ . )
#9.4 Now p l o t
s c a t t e r . smooth ( rstandard (LC. l o g )~ f i t t e d (LC. l o g ) , l a s =1 , c o l =" grey " ,
ylab =" Standardized r e s i d u a l s " , xlab =" F i t t e d values " ,
main="Log transformation " )
#NOTES: This p l o t show approximately constant variance and no trend .
#The l o g a r i t h m i c transformation appears s u i t a b l e ,
#and a l s o allows e a s i e r i n t e r p r e t a t i o n s than using the square r o o t transformation .
#A l o g a r i t h m i c transformation o f the response i s required
# t o produce almost constant variance .
#10.BOX−COX TRANSFORMATION
#NOTES: The maximum o f the Box−Cox p l o t i s achieved when Œª i s j u s t above zero ,
# confirming that a l o g a r i t h m i c transformation
# i s c l o s e t o optimal f o r achieving l i n e a r i t y , normality and constant variance
#10.1 Use MASS package as t h i s i s where boxcox ( ) l i v e s
l i b r a r y (MASS)
#10.2 Use the f o l l o w i n g command t o run the transformation and c r e a t e the p l o t
boxcox ( FEV ~ Ht + Gender + Smoke ,
lambda=seq ( − 0.25 , 0 . 2 5 , length =11) , data=lungcap )
#11. SPECIAL CASE
#11.1 Transforming Covariates and Response v a r i a b l e s
LC. lm . l o g <− lm ( l o g (FEV)~ l o g ( Ht ) , data=lungcap )
#11.2 Now p r i n t the output t o view i t
printCoefmat ( c o e f ( summary (LC. lm . l o g ) ) )
268
Electronic copy available at: https://ssrn.com/abstract=3863563
8.9. INTRODUCTION TO MULTIVARIATE ANALYSIS
8.9
Introduction to multivariate analysis
# 1 . INTRODUCTION TO MULTIVARIATE ANALYSIS_EXERCISE
#1.1 Type the f o l l o w i n g command t o get a s s i s t a n c e on MVA
? ?MVA
#1.2 Library the package o f MVA
l i b r a r y ( "MVA" )
#1.3 Type demo ( " Ch−MVA" ) t o see more
demo ( " Ch−MVA" )
#1.4 Library the package o f l a t t i c e
library (" lattice ")
#1.5 Import data
hypo
#1.6 Hypothetical example o f some m u l t i v a r i a t e data ( n = 10 obs x q = 7 var )
dim ( hypo )
#1.7 data . frame contain the matrix o f data
#NOTES:
# Subsets o f data may be e x t r a c t e d using the subset [ o pera tor ]
hypo [ 1 : 2 , c ( " health " , " weight " ) ]
#1.8 Examples o f m u l t i v a r i a t e data s e t s
#1.9 The f i r s t data s e t c o n s i s t s o f chest , waist , and hip measurements
#on a sample o f men and women and the measurements f o r 20 i n d i v i d u a l s
measure
#NOTES:
# Two questio ns might be addressed by such data :
# 1 . Could body s i z e and body shape be summarised in some way
#by combining the three measurements i n t o a s i n g l e number?
# 2 . Are there sub−types o f body shapes amongst the men
269
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#and amongst the women within which i n d i v i d u a l s are o f s i m i l a r shapes
#and between which body shapes d i f f e r ?
# The f i r s t question could be answered by P r i n c i p a l Components Analysis (PCA)
#− w i l l examine l a t e r
# The Second question could be i n v e s t i g a t e d using Cluster Analysis
#− w i l l look at l a t e r t o o
#1.10 The second s e t o f m u l t i v a r i a t e data c o n s i s t s o f the r e s u l t s o f
# chemical a n a l y s i s on Romano−B r i t i s h p o t t e r y made in three d i f f e r e n t r e g i o n s
# ( region 1 c o n t a i n s k i l n 1 , re gion 2 c o n t a i n s k i l n s 2 and 3 ,
#and region 3 c o n t a i n s k i l n s 4 and 5 ) .
pottery
dim ( p o t t e r y )
#NOTES:
# One question that might be posed about these data
# i s whether the chemical p r o f i l e s o f each pot suggest d i f f e r e n t types o f pots
#and i f any such types are r e l a t e d t o k i l n or region ?
#1.11 The t h i r d s e t o f m u l t i v a r i a t e data i n v o l v e s the examination s c o r e s
# o f a l a r g e number o f c o l l e g e students in s i x s u b j e c t s ;
exam
#NOTES:
# Here the main question o f i n t e r e s t might be whether the exam s c o r e s r e f l e c t
#some underlying t r a i t in a student that cannot be measured d i r e c t l y ,
#perhaps " general i n t e l l i g e n c e " ?
# The question could be i n v e s t i g a t e d by using e x p l o r a t o r y f a c t o r a n a l y s i s (EFA)
#− w i l l review l a t e r on
#1.12 The f i n a l s e t o f data we s h a l l c o n s i d e r in t h i s s e c t i o n
#was c o l l e c t e d in a study o f a i r p o l l u t i o n in c i t i e s in the USA.
#NOTES:
# The f o l l o w i n g v a r i a b l e s were obtained f o r 41 US c i t i e s :
270
Electronic copy available at: https://ssrn.com/abstract=3863563
8.9. INTRODUCTION TO MULTIVARIATE ANALYSIS
# 1 . SO2 : SO2 content o f a i r in micrograms per c u b i c metre ;
# 2 . temp : average annual temperature in degrees Fahrenheit ;
# 3 . manu : number o f manufacturing e n t e r p r i s e s employing 20 or more workers ;
# 4 . popul : population s i z e (1970 census ) in thousands ;
# 5 . wind : average annual wind speed in miles per hour ;
# 6 . p r e c i p : average annual p r e c i p i t a t i o n in inches ;
# 7 . predays : average number o f days with p r e c i p i t a t i o n per year .
U Sa i rp o ll u ti o n
dim ( U Sa i r po l l u t i o n )
#NOTES:
# 1 . What might be the question o f most i n t e r e s t about these data ?
# 2 . Very probably i t i s "how i s p o l l u t i o n l e v e l as measured
#by sulphur d i o x i d e c o n c e n t r a t i o n r e l a t e d t o the s i x other v a r i a b l e s ? "
# 3 . In the f i r s t i n s t a n c e at l e a s t , t h i s question suggests the a p p l i c a t i o n
# o f multiple l i n e a r r e g r e s s i o n ,
#with sulphur d i o x i d e c o n c e n t r a t i o n as the response v a r i a b l e
#and the remaining s i x v a r i a b l e s being the independent
# or explanatory v a r i a b l e s
# ( the l a t t e r i s a more a c c e p t a b l e l a b e l
#because the " independent " v a r i a b l e s are r a r e l y independent o f one another ) .
# 4 . But in the model underlying multipl e r e g r e s s i o n ,
# only the response i s considered t o be a random v a r i a b l e ;
#the explanatory v a r i a b l e s are s t r i c t l y assumed t o be f i x e d ,
#not random , v a r i a b l e s .
# 5 . In p r a c t i c e , o f course , t h i s i s r a r e l y the case ,
#and so the r e s u l t s from a mul tiple r e g r e s s i o n a n a l y s i s need t o be i n t e r p r e t e d
#as being c o n d i t i o n a l on the observed values o f the explanatory v a r i a b l e s .
# 6 . So , when answering the question o f most i n t e r e s t about these data ,
#they should not r e a l l y be considered m u l t i v a r i a t e
#− there i s only a s i n g l e random v a r i a b l e i n v o l v e d
#− a more s u i t a b l e l a b e l i s m u l t i v a r i a b l e
271
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 2 . COVARIANCE
#2.1 The cov ari an ce f o r the chest , waist & hip measurements , use the cov ( ) f u n c t i o n
#NOTES:
# however , have t o " remove " the c a t e g o r i c a l v a r i a b l e gender
#from the measure data frame by sub−s e t t i n g on the numerical v a r i a b l e s f i r s t :
cov ( measure [ , c ( " chest " , " waist " , " hips " ) ] )
#2.1 I f r e q u i r e the separate cov ar ian ce matrices o f men and women, can use
cov ( subset ( measure , gender == " female " ) [ , c ( " chest " , " waist " , " hips " ) ] )
cov ( subset ( measure , gender == " male " ) [ , c ( " chest " , " waist " , " hips " ) ] )
#NOTES:
# where the subset ( ) returns a l l o b s e r v a t i o n s corresponding t o
# females f i r s t statement ) or males ( second statement ) .
# 3 . CORRELATIONS
#3.1 The sample c o r r e l a t i o n matrix f o r the three v a r i a b l e s in above
# i s obtained by using the f u n c t i o n c o r ( ) in R :
c o r ( measure [ , c ( " chest " , " waist " , " hips " ) ] )
##4. DISTANCES
#4.1 The necessary R code f o r d i s t ( ) and s c a l e ( ) are :
d i s t ( s c a l e ( measure [ , c ( " chest " , " waist " , " hips " ) ] , c e n t e r = FALSE ) )
##5. THE MULTIVARIATE NORMAL DENSITY FUNCTION
#NOTES:
# Will now assess the body measurements data f o r normality ,
272
Electronic copy available at: https://ssrn.com/abstract=3863563
8.9. INTRODUCTION TO MULTIVARIATE ANALYSIS
#although because there are only 20 o b s e r v a t i o n s
# in the sample there i s r e a l l y t o o l i t t l e information
# t o come t o any convincing c o n c l u s i o n .
#5.1 The p l o t i s s e t up as f o l l o w s . F i r s t e x t r a c t the r e l e v a n t data
x <− measure [ , c ( " chest " , " waist " , " hips " ) ]
#5.2 Estimate the means o f a l l three v a r i a b l e s
# ( i . e . , f o r each column o f the data ) and the co var ian ce matrix
cm <− colMeans ( x )
S <− cov ( x )
#5.3 The d i f f e r e n c e s d i have t o be computed f o r a l l u n i t s in our data ,
# so i t e r a t e over the rows o f x using the apply ( ) f u n c t i o n
#with argument MARGIN = 1 and , f o r each row , compute the d i s t a n c e d i :
d <− apply ( x , MARGIN = 1 , f u n c t i o n ( x )
+ t ( x − cm) %*% s o l v e ( S ) %*% ( x − cm ) )
#5.4 The s o r t e d d i s t a n c e s can now be p l o t t e d
# against the appropriate q u a n t i l e s o f the chi −square3 d i s t r i b u t i o n obtained from qchisq (
qqnorm ( measure [ , " chest " ] , main = " chest " ) ; q q l i n e ( measure [ , " chest " ] )
qqnorm ( measure [ , " waist " ] , main = " waist " ) ; q q l i n e ( measure [ , " waist " ] )
qqnorm ( measure [ , " hips " ] , main = " hips " ) ; q q l i n e ( measure [ , " hips " ] )
p l o t ( qchisq ( ( 1 : nrow ( x ) − 1 / 2 ) / nrow ( x ) , df = 3 ) , s o r t ( d ) ,
xlab = e xpr ess ion ( paste ( c h i [ 3 ] ^ 2 , " Quantile " ) ) ,
ylab = " Ordered d i s t a n c e s " )
abline ( a = 0 , b = 1)
#NOTES:
# Now look at using the chi −square p l o t on a s e t o f data introduced e a r l i e r ,
#namely the a i r p o l l u t i o n in US c i t i e s .
#5.5 I t e r a t e over a l l v a r i a b l e s , t h i s time using a s p e c i a l function ,
# sapply ( ) , that l o o p s over the v a r i a b l e names :
layout ( matrix ( 1 : 8 , nc = 2 ) )
sapply ( colnames ( U S a i r p o l l u t i o n ) , f u n c t i o n ( x ) {
273
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
+ qqnorm ( U S a i r p o l l u t i o n [ [ x ] ] , main = x )
+ qqline ( USairpollution [ [ x ] ] ) } )
#NOTES:
#1 The p l o t s f o r SO2 c o n c e n t r a t i o n and p r e c i p i t a t i o n both d e v i a t e c o n s i d e r a b l y
#from l i n e a r i t y , and the p l o t s f o r manufacturing
#and population show evidence o f a number o f o u t l i e r s .
#2 But o f more importance i s the chi −square p l o t f o r the data .
#3 The R code i s i d e n t i c a l t o the code used t o
#produce the chi −square p l o t f o r the body measurement data .
#4 In addition , the two most extreme p o i n t s in the p l o t have been
# l a b e l l e d with the c i t y names t o which they correspond using t e x t ( ) .
x <− U S a i r p o l l u t i o n
cm <− colMeans ( x )
S <− cov ( x )
d <− apply ( x , 1 , f u n c t i o n ( x ) t ( x − cm) %*% s o l v e ( S ) %*% ( x − cm ) )
p l o t ( qc <− qchisq ( ( 1 : nrow ( x ) − 1 / 2 ) / nrow ( x ) , df = 6 ) , sd <− s o r t ( d ) ,
xlab = e xpr ess ion ( paste ( c h i [ 6 ] ^ 2 , " Quantile " ) ) ,
ylab = " Ordered d i s t a n c e s " , xlim = range ( qc ) * c ( 1 , 1 . 1 ) )
oups <− which ( rank ( abs ( qc − sd ) , t i e s = " random " ) > nrow ( x ) − 3 )
t e x t ( qc [ oups ] , sd [ oups ] − 1 . 5 , names ( oups ) )
abline ( a = 0 , b = 1)
#NOTES:
# 1 . This example i l l u s t r a t e s that the chi −square p l o t might a l s o be u s e f u l f o r
# d e t e c t i n g p o s s i b l e o u t l i e r s in m u l t i v a r i a t e data ,
#where i n f o r m a l l y o u t l i e r s are " abnormal "
# in the sense o f d e v i a t i n g from the natural data v a r i a b i l i t y .
# 2 . O u t l i e r i d e n t i f i c a t i o n i s important in many a p p l i c a t i o n s o f
#multivariate analysis either
#because there i s some s p e c i f i c i n t e r e s t in f i n d i n g anomalous o b s e r v a t i o n s
274
Electronic copy available at: https://ssrn.com/abstract=3863563
8.9. INTRODUCTION TO MULTIVARIATE ANALYSIS
# or as a pre−p r o c e s s i n g task b e f o r e the a p p l i c a t i o n o f some m u l t i v a r i a t e method
# in order t o preserve the r e s u l t s from p o s s i b l e misleading e f f e c t s produced
#by these o b s e r v a t i o n s .
275
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.10
Multivariate & special distributions
## BU8555 MULTIVARIATE & SPECIAL DISTRIBUTIONS
## NORMALITY
# You can ask R f o r help on t h i s t o p i c
help . search ( " normality " )
# Shapiro−Wilk Normality Test i s the one that we s h a l l look at
names
## WILK−SHAPIRO TEST
# I n s t a l l the package c a l l e d mvnormtest
i n s t a l l . packages ( " mvnormtest " )
l i b r a r y ( mvnormtest ) # Load package i n t o a p p l i c a t i o n
data ( a t t i t u d e )
# Load the a t t i t u d e data s e t
l i b r a r y ( ) # Check l i s t o f packages a v a i l a b l e t o you on your d e v i c e
data ( )
# Check the current l i s t o f data s e t s in R
attach ( a t t i t u d e )
# Attach data f i l e so we can use the v a r i a b l e
head ( a t t i t u d e , n = 10)
# Examine the f i r s t 10 l i n e s o f the data
summary ( a t t i t u d e )
# We can perform the usual d e s c r i p t i v e summary output
dim ( a t t i t u d e )
# We can q u i c k l y see how many obs and var t h i s data s e t has
# We need t o transpose the data s e t b e f o r e using the Shapiro−Wilk t e s t
# The t ( ) f u n c t i o n i s used t o transpose the data
mydata = t ( a t t i t u d e )
# Transpose the data
mshapiro . t e s t ( mydata ) # Run the W−S t e s t on the transposed data
# The Shapiro−Wilk p value o f .0002 i n d i c a t e s
# that the m u l t i v a r i a t e normality assumption does not hold .
# I t i n d i c a t e s that one or more i n d i v i d u a l v a r i a b l e s
#are not normally d i s t r i b u t e d .
## JARQUE−BERA TEST
# We s h a l l use another R package c a l l e d normtest
# J−B t e s t f o r obs and r e g r e s s i o n r e s i d u a l s i s l o c a t e d in normtest
276
Electronic copy available at: https://ssrn.com/abstract=3863563
8.10. MULTIVARIATE & SPECIAL DISTRIBUTIONS
i n s t a l l . packages ( " normtest " )
l i b r a r y ( normtest )
# Now we run the j b . norm . t e s t on the a t t i t u d e data s e t
j b . norm . t e s t ( a t t i t u d e )
# The Jarque−Bera t e s t f o r normality agreed with the Shapiro−Wilk t e s t
# that one or more v a r i a b l e s are not normally d i s t r i b u t e d .
## CHECK INDIVIDUAL VARIABLE SKEWNESS, KURTOSIS & NORMALITY
# To perform these checks in R we can use the normwhn . t e s t
# F i r s t i n s t a l l the normwhn . t e s t package
i n s t a l l . packages ( " normwhn . t e s t " )
l i b r a r y ( normwhn . t e s t )
# Next we run the f o l l o w i n g t e s t
normality . t e s t 1 ( a t t i t u d e )
# This normality t e s t i n c l u d e s i n d i v i d u a l v a r i a b l e skewness
#and k u r t o s i s values and an omnibus t e s t o f normality
# The r e s u l t s i n d i c a t e d that the v a r i a b l e s o v e r a l l
# did not contain skewness or k u r t o s i s . ( see PowerPoint f o r t a b l e )
# The omnibus normality t e s t i n d i c a t e d that
#the data were normally d i s t r i b u t e d , Z = 1 2 . 8 4 , df = 14 , p = . 5 4 .
## EXAMINING EACH INDIVIDUAL VARIABLE NORMALITY
# The R package n o r t e s t c o n t a i n s f i v e normality t e s t s ( see PowerPoint s l i d e s )
# We s h a l l f o c u s on the Anderson−Darling (A−D) t e s t
#as recommended by Stephens , M. A . ( 1 9 8 6 )
i n s t a l l . packages ( " n o r t e s t " )
library ( nortest )
attach ( a t t i t u d e )
# Let us perform a l l f i v e t e s t s on ’ rating ’
277
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
ad . t e s t ( r a t i n g )
# Anderson−Darling
cvm . t e s t ( r a t i n g )
# Cramer−Von Mises
l i l l i e . test ( rating )
# Kolmogorov−Smirnov
pearson . t e s t ( r a t i n g )
# Pearson Chi−square
sf . test ( rating )
# Shapiro−Francia
# We repeat f o r the v a r i a b l e ’ complaints ’
ad . t e s t ( complaints )
cvm . t e s t ( complaints )
l i l l i e . t e s t ( complaints )
pearson . t e s t ( complaints )
s f . t e s t ( complaints )
# Likewise we repeat the 5 t e s t s f o r v a r i a b l e ’ p r i v i l e g e s ’
ad . t e s t ( p r i v i l e g e s )
cvm . t e s t ( p r i v i l e g e s )
l i l l i e . test ( privileges )
pearson . net ( p r i v i l e g e s )
sf . test ( privileges )
# And a l s o f o r the v a r i a b l e ’ learning ’
ad . t e s t ( l e a r n i n g )
cvm . t e s t ( l e a r n i n g )
l i l l i e . t e st ( learning )
pearson . t e s t ( l e a r n i n g )
s f . t e s t ( learning )
# Next we do the same t e s t s f o r v a r i a b l e c a l l e d ’ r a i s e s ’
ad . t e s t ( r a i s e s )
cvm . t e s t ( r a i s e s )
l i l l i e . test ( raises )
pearson . t e s t ( r a i s e s )
sf . test ( raises )
278
Electronic copy available at: https://ssrn.com/abstract=3863563
8.10. MULTIVARIATE & SPECIAL DISTRIBUTIONS
# Now we perform the same t e s t s f o r the v a r i a b l e ’ c r i t i c a l ’
ad . t e s t ( c r i t i c a l )
cvm . t e s t ( c r i t i c a l )
l i l l i e . test ( c r i t i c a l )
pearson . t e s t ( c r i t i c a l )
sf . test ( c r i t i c a l )
# Finally , we repeat the t e s t s f o r the l a s t v a r i a b l e c a l l e d ’ advance ’
ad . t e s t ( advance )
cvm . t e s t ( advance )
l i l l i e . t e s t ( advance )
pearson . t e s t ( advance )
s f . t e s t ( advance )
# See PowerPoint f o r a summary t a b l e o f a l l 7 var x 5 t e s t s
# A l l f i v e normality t e s t s showed that
#the v a r i a b l e c r i t i c a l v i o l a t e d the normality assumption .
# This i s why the Shapiro−Wilk and Jarque−Bera t e s t s i n d i c a t e d that
#the m u l t i v a r i a t e normality assumption was not met .
# However , t h i s s i n g l e v a r i a b l e , c r i t i c a l , although not normally d i s t r i b u t e d ,
#would g e n e r a l l y not a f f e c t the power t o
# d e t e c t a d i f f e r e n c e in a m u l t i v a r i a t e s t a t i s t i c a l a n a l y s i s ( Stevens , 2009)
## DETERMINANT OF A MATRIX
## Return t o PowerPoint f o r c o n t e x t u a l background
# The R commands f o r c o r r e l a t i o n matrix are as f o l l o w s :
# Create square c o r r e l a t i o n matrix
mycor = c o r ( a t t i t u d e )
# Next , the determinant i s computed using the R f u n c t i o n det ( )
det ( mycor )
279
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# Please note :
#The d e f a u l t c o r r e l a t i o n type in c o r ( ) i s the Pearson c o r r e l a t i o n c o e f f i c i e n t .
# The R commands t o c r e a t e a variance −co va ria nce matrix
#and t o obtain the determinant o f the matrix are as f o l l o w s :
# Create a square variance −co var ia nce matrix
mycov = cov ( a t t i t u d e )
# Determinant o f the variance −co var ia nce matrix
det ( mycov )
# The determinant ( g e n e r a l i z e d variance ) o f the matrix i s p o s i t i v e ,
# t h e r e f o r e m u l t i v a r i a t e s t a t i s t i c a l a n a l y s i s can proceed .
# Note : I f you want decimal places , rather than s c i e n t i f i c notation ,
# i s s u e the f o l l o w i n g command :
o p t i o n s ( scipen = 999)
# An a d d i t i o n a l R f u n c t i o n can be u s e f u l
#when wanting a c o r r e l a t i o n matrix from a variance −co var ia nce matrix .
# The f u n c t i o n cov2cor ( ) s c a l e s a cov ari anc e matrix
#by i t s diagonal t o compute the c o r r e l a t i o n matrix .
# The command i s as f o l l o w s :
mymatrix = cov2cor ( mycov )
# Convert c ova ri anc e matrix t o c o r r e l a t i o n matrix
# Now p r i n t out c o r r e l a t i o n matrix
mymatrix
# You w i l l get t h i s same c o r r e l a t i o n matrix i f you l i s t the matrix mycor ,
#which was p r e v i o u s l y cre ated from the a t t i t u d e data s e t .
mycor
## EQUITY OF VARIANCE−COVARIANCE MATRIX
## Back t o S l i d e s f o r 1 min
# The R commands t o add a group membership v a r i a b l e are as f o l l o w s :
280
Electronic copy available at: https://ssrn.com/abstract=3863563
8.10. MULTIVARIATE & SPECIAL DISTRIBUTIONS
# Create group v a r i a b l e 15 boys and 15 g i r l s
group = rep ( c ( " boy " , " g i r l " ) , c ( 1 5 , 1 5 ) )
# Now merge a t t i t u d e f i l e with the new group v a r i a b l e
newdata = data . frame ( a t t i t u d e , group )
# We check the f i r s t 10 rows o f newdata
head ( newdata , n = 10)
# The within variance −co var ia nce matrices and a s s o c i a t e d determinants
# f o r each group can now be c a l c u l a t e d with the f o l l o w i n g R commands
o p t i o n s ( scipen = 999)
# Stops s c i e n t i f i c n o t a t i o n output
boys = newdata [ 1 : 1 5 , ]
# S e l e c t s only boys
boycov = cov ( boys [ , − 8 ] ) # Create c ova ri anc e matrix
det ( boycov )
# Determinant computed o f matrix
# and now we do the same f o r g i r l s
g i r l s = newdata [ 1 6 : 3 0 , ] # S e l e c t only g i r l s from the data s e t
g i r l c o v = cov ( g i r l s [ , − 8 ] ) # Create c ova ria nc e matrix
det ( g i r l c o v )
# Compute determinant
# The determinants o f the boys and g i r l s variance −co var ia nce matrices were p o s i t i v e ,
#thus m u l t i v a r i a t e s t a t i s t i c a l a n a l y s i s can proceed .
# We can obtain the d e s c r i p t i v e s t a t i s t i c s by using the R package psych
#and the describeBy ( ) f u n c t i o n .
# We i n s t a l l and load the psych package ,
#and then we i s s u e the command f o r the f u n c t i o n .
i n s t a l l . packages ( " psych " )
l i b r a r y ( psych )
# Run the describeBy ( ) f u n c t i o n
describeBy ( newdata , group = group )
# The standard d e v i a t i o n s , hence variance values ,
281
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#do look a l i t t l e d i f f e r e n t between the boys and the g i r l s a c r o s s the v a r i a b l e s .
# We can c r e a t e and l i s t the separate cov ari anc e matrices f o r each group
# in the newdata data s e t by using the l a p p l y ( ) function ,
#where the group v a r i a b l e i s d e l e t e d ( ? ? ? 8 ) as f o l l o w s
cov . l i s t = l a p p l y ( unique ( newdata$group ) , f u n c t i o n ( x )
cov ( newdata [ newdata$group==x , − 8 ] ) )
# We can l i s t the boys ’ c ov ari anc e matrix as f o l l o w s :
cov . l i s t [ 1 ]
# We can do l i k e w i s e f o r the g i r l s
cov . l i s t [ 2 ]
# The f i r s t approach obtained the separate group variance −co var ia nce matrices much
# e a s i l y ( l e s s s o p h i s t i c a t e d programming ) .
# We can e a s i l y convert these separate variance −co var ia nce matrices i n t o
# c o r r e l a t i o n matrices using the cov ( ) f u n c t i o n and the cov2cor ( ) f u n c t i o n
#as shown b e f o r e .
# The c ova ri anc e matrices exclude the grouping v a r i a b l e
#by i n d i c a t i n g a ? ? ? 8 value ( column l o c a t i o n in the matrix )
# in the s e l e c t i o n o f v a r i a b l e s t o i n c l u d e .
# The two s e t s o f R commands are as f o l l o w s
boys = newdata [ 1 : 1 5 , ]
boycov = cov ( boys [ , − 8 ] )
boycor = cov2cor ( boycov )
boycor
# and
g i r l s = newdata [ 1 6 : 3 0 , ]
g i r l c o v = cov ( g i r l s [ , − 8 ] )
g i r l c o r = cov2cor ( g i r l c o v )
girlcor
282
Electronic copy available at: https://ssrn.com/abstract=3863563
8.10. MULTIVARIATE & SPECIAL DISTRIBUTIONS
## BOX−M TEST
# Look at s i n g l e s l i d e on PowerPoint
# The b a s i c example shown here however d e c l a r e s the grouping v a r i a b l e as a f a c t o r ,
# t h e r e f o r e the complete code would be as f o l l o w s :
i n s t a l l . packages ( " b i o t o o l s " ) # I n s t a l l the R b i o t o o l s package
l i b r a r y ( b i o t o o l s ) # Load the package
data ( i r i s ) # Load t h i s data s e t
f a c t o r ( i r i s [ , 5 ] ) # S e l e c t v a r i a b l e s f o r matrix , minus grouping v a r i a b l e
boxM ( i r i s [ , − 5] , i r i s [ , 5 ] ) # Run f u n c t i o n with v a r i a b l e s and group v a r i a b l e s p e c i f i e d
# Note :
#The use o f i r i s [ , − 5] s e l e c t s v a r i a b l e s and i r i s [ , 5 ] i n d i c a t e s
#grouping v a r i a b l e in data f i l e
# In the i r i s data set , the n u l l hypothesis o f equal variance −co va ria nce matrices
#was r e j e c t e d .
# The groups did have d i f f e r e n t variance −co va ria nce matrices .
# A m u l t i v a r i a t e a n a l y s i s would t h e r e f o r e be suspect ,
# e s p e c i a l l y i f non−normality e x i s t e d among the v a r i a b l e s .
# Now we return t o our i l l u s t r a t i o n as b e f o r e
# The R code s t e p s are l i s t e d below .
# F i r s t , i n s t a l l the package from the main menu .
#Next , load the b i o t o o l s package and the data s e t .
#You must d e c l a r e the v a r i a b l e group as a f a c t o r
# b e f o r e running the Box M t e s t f u n c t i o n . The s e t o f R commands are as f o l l o w s :
i n s t a l l . packages ( " b i o t o o l s " )
# I n s t a l l the R b i o t o o l s package
library ( biotools )
# Load the package
data ( newdata )
# Load t h i s data s e t
f a c t o r ( group )
# Declare grouping v a r i a b l e a f a c t o r
boxM ( newdata [ , − 8] , newdata [ , 8 ] )
# Run Box−M f u n c t i o n
283
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# The r e s u l t s i n d i c a t e d that
#the boys and the g i r l s variance −co var ia nce matrices are equal ;
# that i s , the chi −square t e s t was non−s i g n i f i c a n t .
284
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
8.11
Nonlinear
# 1 . NONLINEAR_EXERCISE
#1.1 Nonlinear Least Squares
# 1 . 1 . 1 I n s t a l l and l i b r a r y packages
i n s t a l l . packages ( " Formula " )
i n s t a l l . packages ( " plotmo " )
i n s t a l l . packages ( " TeachingDemos " )
l i b r a r y ( Formula )
l i b r a r y ( TeachingDemos )
l i b r a r y ( plotmo )
library ( plotrix )
#Notes :
# The nonlinear nature o f d i s t r i b u t i o n i s c l e a r l y i d e n t i f i a b l e .
#1.2 Now, t o understand how a l i n e a r r e g r e s s i o n
# i s not s u i t a b l e f o r modeling such a d i s t r i b u t i o n ,
#a l i n e a r r e g r e s s i o n model and then a nonlinear one w i l l be b u i l t f o r comparing .
x <− seq ( 0 , 200 , 1 )
y< − (( r u n i f ( 1 , 8 , 2 5 ) * x ) / ( r u n i f ( 1 , 0 , 7 ) + x ) ) + r u n i f ( 2 0 1 , 0 , 1 )
plot (x , y )
# 1 . 1 . 1 l i n e a r model
LModel<−lm ( y~x )
# 1 . 1 . 2 Display some data s t a t i s t i c s using the summary f u n c t i o n
LMSummary<−summary ( LModel )
# 1 . 1 . 3 Results are :
LMSummary
# 1 . 1 . 4 the nonlinear model
NLModel<− n l s ( y~a * x / ( b+x ) , s t a r t = l i s t ( a=1 ,b = 0 . 1 ) )
NLMSummary <− summary ( NLModel )
NLMSummary
# 1 . 1 . 5 From the s t a t i s t i c s returned ,
285
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#the RSS or r e s i d u a l standard e r r o r f o r each model can be e x t r a c t e d t o compare them
LMSummary$sigma
NLMSummary$sigma
# NOTES:
# 1 . The l i n e a r model e r r o r i s about s i x times g r e a t e r than that o f
#the nonlinear one .
# 2 . This shows that the nonlinear model f i t s b e t t e r .
# 1 . 1 . 6 The same r e s u l t can be looked through a g r a p h i c a l a n a l y s i s
plot (x , y )
a b l i n e ( LModel )
l i n e s ( x , p r e d i c t ( NLModel ) , c o l =" red " )
#NOTES:
# 1 . In the code j u s t proposed ,
# f i r s t reported data on the graph so added the r e g r e s s i o n l i n e
#using the a b l i n e ( ) f u n c t i o n .
#This f u n c t i o n adds one or more s t r a i g h t l i n e s through the current p l o t .
# 2 . Then added the l i n e s using the l i n e s ( ) f u n c t i o n
# that r e p r e s e n t s the nonlinear model ( where s e l e c t e d c o l o r =red )
# 3 . By analyzing the f i g u r e ,
# i t i s p o s s i b l e t o confirm
# that the r e g r e s s i o n l i n e i s unable t o p r e d i c t the s t a r t i n g data .
# 4 . Conversely , the nonlinear model p e r f e c t l y f i t s the model .
# 2 . MULTIVARIATE ADAPTIVE REGRESSION SPLINES (MARS)
#2.1 load the data t r e e s
data ( t r e e s )
show ( t r e e s )
#2.2 To d i s p l a y a compact summary use the s t r u c t u r e command
str ( trees )
#2.3 To get more information , use the summary ( ) f u n c t i o n
summary ( t r e e s )
286
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
#NOTES:
# 1 . Some s t a t i s t i c s are returned with min , max, mean , median ,
#and q u a r t i l e s f o r each v a r i a b l e .
# 2 . To understand something more than t h i s d i s t r i b u t i o n ,
#we can perform a preliminary v i s u a l a n a l y s i s .
#Let ’ s f i r s t t r y t o f i n d out whether the v a r i a b l e s are r e l a t e d t o each other .
#2.4 Can do t h i s using the p a i r s f u n c t i o n
# t o c r e a t e a matrix o f sub axes c o n t a i n i n g s c a t t e r p l o t s o f the columns o f a matrix :
p a i r s ( tr ee s , panel = panel . smooth )
#NOTES:
# 1 . As seen in the f i g u r e ,
# s c a t t e r p l o t matrices are a great way t o determine roughly whether
#has a c o r r e l a t i o n between mul tiple v a r i a b l e s .
#This i s p a r t i c u l a r l y h e l p f u l in l o c a t i n g s p e c i f i c v a r i a b l e s
# that might have mutual c o r r e l a t i o n s , i n d i c a t i n g a p o s s i b l e redundancy o f data .
# 2 . In the diagonal are the name o f the measured v a r i a b l e s .
#The r e s t o f the p l o t s are s c a t t e r p l o t s o f the matrix columns .
# S p e c i f i c a l l y , each p l o t i s present twice ;
#the p l o t s on the i t h row are the same as those in the i t h column ( mirror image ) .
# 3 . From a f i r s t a n a l y s i s o f the previous f i g u r e ,
#can see s e v e r a l p l o t s showing a nonlinear r e l a t i o n s h i p between the v a r i a b l e s .
#This i s the case o f the p l o t showing the r e l a t i o n s h i p between Volume
#and Girth as well as between Volume and Height .
#This trend was confirmed by the smooth curves that are added at each p l o t .
# 4 . This f e a t u r e was s e t by adding the panel argument t o the p a i r s ( ) f u n c t i o n :
# 5 . The panel . smooth value draw a Lowess
# ( l o c a l l y weighted s c a t t e r p l o t smoothing ) curve .
# 6 . A Lowess curve i s a smooth curve through a s e t o f data p o i n t s .
# 7 . In t h i s curve each smoothed value
287
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# i s given by a weighted quadratic l e a s t squares r e g r e s s i o n
# over the span o f values o f the y a x i s s c a t t e r p l o t c r i t e r i o n v a r i a b l e .
#When each smoothed value i s given by a weighted l i n e a r l e a s t squares r e g r e s s i o n
# over the span , t h i s i s known as a Lowess curve . Returns a measure o f l o c a l r e g r e s s i
# 8 . From the a n a l y s i s o f the Lowess curves ,
# i t i s p o s s i b l e t o confirm that the d i s t r i b u t i o n
# i s l i a b l e t o be t r e a t e d with the MARS technique .
# Furthermore , by analyzing the s c a t t e r p l o t matrix produced ,
#can n o t i c e a g r e a t e r c o r r e l a t i o n between the Volume v a r i a b l e
#and the Girth v a r i a b l e .
#At the same time the Lowess curves made i t c l e a r that
#such a r e l a t i o n s h i p i s not l i n e a r ( maybe at times l i n e a r ) .
#2.5 Let ’ s see what happens i f put the Volume against Girth on a chart by
# adopting the l o g a r i t h m i c s c a l e f o r both axes :
p l o t ( Volume ~ Girth , data = t r e e s , l o g = " xy " ,
xlab =" LogGirth " , ylab ="LogVolume " )
#NOTES:
# 1 . From the a n a l y s i s o f the graph i t i s easy t o i d e n t i f y two zones that
#show trends that could be f i t t e d by two s t r a i g h t l i n e s , and then at times l i n e a r .
#Where i s the v a r i a b l e Height .
#Let ’ s r e i n t r o d u c e i t in the v i s u a l a n a l y s i s are working on .
#One way i s o f f e r e d by the c o p l o t ( ) f u n c t i o n .
#This f u n c t i o n produces two v a r i a n t s o f the c o n d i t i o n i n g p l o t s .
# 2 . Conditioning s c a t t e r p l o t s i n v o l v e s c r e a t i n g a multipanel display ,
#where each panel c o n t a i n s a subset o f the data .
#This subset can be e i t h e r o f the f o l l o w i n g :
# ( 1 ) Those o b s e r v a t i o n s that f a l l in a p a r t i c u l a r group ,
# or ( 2 ) They may represent the values
# that f a l l within a p a r t i c u l a r range o f the values o f a v a r i a b l e .
# 3 . Each i n d i v i d u a l panel i l l u s t r a t e s the r e l a t i o n s h i p
#between a p a i r o f v a r i a b l e s ,
288
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
# over part o f the range o f the two marginal c o n d i t i o n i n g v a r i a b l e s .
#The r e l a t i o n s h i p i s c o n d i t i o n a l on one marginal v a r i a b l e
# l y i n g in one p a r t i c u l a r i n t e r v a l , and the other l y i n g in a d i f f e r e n t i n t e r v a l .
# 4 . In our case , c o p l o t ( ) w i l l contain s c a t t e r diagrams f o r the l o g o f Volume
#as a f u n c t i o n o f the l o g ofGirth ,
# c o n d i t i o n e d by Height ( one response v a r i a b l e and one c o n d i t i o n i n g v a r i a b l e . )
#2.6 Lets f i r s t look at the command t o do t h i s :
c o p l o t ( l o g ( Volume ) ~ l o g ( Girth ) | Height , data = t r e e s ,
panel = panel . smooth )
#NOTES:
# 1 . F i r s t , a formula d e s c r i b i n g the form o f c o n d i t i o n i n g p l o t i s passed
# 2 . A formula o f the form y~x |a i n d i c a t e s
# that p l o t s o f y versus x should be produced c o n d i t i o n a l on the v a r i a b l e a .
# 3 . Then compare the name o f the dataset and the panel . smooth argument .
#The f o l l o w i n g f i g u r e shows a c o n d i t i o n i n g p l o t f o r l o g o f Volume
#as a f u n c t i o n o f the l o g o f Girth , c o n d i t i o n e d by Height :
# 4 . In the top part o f the graph
#are the ranges where the range o f e x i s t e n c e o f the Height v a r i a b l e was subdivided .
#At the bottom o f the graph are the s i x graphs that contain the l o g o f Volume
#as a f u n c t i o n o f the l o g o f Girth f o r each range o f Height .
#From the a n a l y s i s o f the Lowess curves ,
# i t i s c l e a r the d i f f e r e n t trend that LogVolume assumes
# in each range o f the Height v a r i a b l e . The time has come t o apply the MARS.
#2.7 I n s t a l l the earth package and load the l i b r a r y .
i n s t a l l . packages ( " earth " )
l i b r a r y ( earth )
#2.8 Use the earth ( ) f u n c t i o n that b u i l d s a r e g r e s s i o n model using
#the techniques
# in Friedman ’ s papers M u l t i v a r i a t e Adaptive Regression Splines and Fast MARS
289
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
MARSModel <− earth ( Volume ~ . , data = t r e e s )
#2.9 Summary r e s u l t s
summary (MARSModel)
#2.10 Estimate v a r i a b l e importance in the model j u s t cr eated :
EvimpMD<−evimp (MARSModel)
EvimpMD
#2.11 Now show the model r e s u l t s in the p l o t provided by the earth package :
p l o t (MARSModel)
#2.12 To conclude the analysis , w i l l c a l c u l a t e the accuracy o f the model
#and compare i t t o that obtained with a l i n e a r model .
#To c a l c u l a t e the accuracy ,
#must f i r s t make p r e d i c t i o n s with the newly b u i l t model .
#To do t h i s , w i l l use the p r e d i c t ( ) function , as f o l l o w s :
PredMARSModel <− p r e d i c t (MARSModel, t r e e s )
#2.13 Now w i l l c a l c u l a t e the mean squared e r r o r (MSE) that
#measures the average o f the squares o f the e r r o r s or d e v i a t i o n s :
mse1 <− mean ( ( trees$Volume − PredMARSModel) ^ 2 )
#2.14 The r e s u l t i s shown as f o l l o w s :
p r i n t ( mse1 )
#2.15 To evaluate the performance o f the model ,
#compare i t with a l i n e a r r e g r e s s i o n model .
#To do t h i s , w i l l do the same thing , that i s ,
# w i l l f i r s t b u i l d the model , so w i l l make p r e d i c t i o n s , and f i n a l l y c a l c u l a t e MSE:
LMModel <− lm ( Volume ~ . , data = t r e e s )
#2.16 So carry out the f o r e c a s t
PredLMModel <− p r e d i c t ( LMModel , t r e e s )
#2.17 Finally , c a l c u l a t e the MSE again
mse2 <− mean ( ( trees$Volume − PredLMModel ) ^ 2 )
p r i n t ( mse2 )
290
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
#NOTES: Comparing the two models , the MARS model shows a lower e r r o r .
# 3 . GENERALIZED ADDITIVE MODEL
#NOTES: Air . Flow r e p r e s e n t s the r a t e o f o p e r a t i o n o f the plant .
#Water . Temp i s the temperature o f c o o l i n g water c i r c u l a t e d through c o i l s
# in the absor pti on tower . Acid . Conc . i s the c o n c e n t r a t i o n o f the a c i d c i r c u l a t i n g ,
# −50, times 1 0 : that i s , 89 corresponds t o 58.9 percent a c i d . stack . l o s s
# ( the dependent v a r i a b l e ) i s ten times the percentage
# o f the ingoing ammonia t o the plant that escapes
#from the a bsor pti on column unabsorbed ;
# that i s , an ( i n v e r s e ) measure o f the o v e r a l l e f f i c i e n c y o f the plant .
#3.1 Compare the l i n e a r model with an a d d i t i v e model .
#The a n a l y s i s begins by uploading the dataset :
data ( s t a c k l o s s )
#3.2 Display a compact summary
str ( stackloss )
#3.3 For more information
summary ( s t a c k l o s s )
#3.4 Can perform a preliminary v i s u a l analysis ,
# t o f i r s t t r y and see i f the v a r i a b l e s are r e l a t e d t o each other .
#And can do t h i s with the p a i r s ( ) f u n c t i o n t o c r e a t e a matrix o f
#sub−axes c o n t a i n i n g s c a t t e r p l o t s o f the columns
p a i r s ( s t a c k l o s s , panel = panel . smooth )
#NOTES: The a n a l y s i s o f the graph confirms the n o n l i n e a r i t y
# o f the r e l a t i o n s h i p s between responses and p r e d i c t o r s .
#3.5 Now w i l l see what returns a l i n e a r r e g r e s s i o n model .
#This a n a l y s i s w i l l help us l a t e r t o measure GAM performance :
LModel <− lm ( stack . l o s s ~ Air . Flow + Water . Temp + Acid . Conc . ,
data= s t a c k l o s s )
291
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#3.6 Can q u i c k l y summarize the l i n e a r model
summary ( LModel )
#3.7 At f i r s t glance the r e s u l t s are not bad at a l l :
#the p r e d i c t o r s are s i g n i f i c a n t ( l e s s than Acid . Conc . )
#and R−squared i s high enough .
#Can b e t t e r analyze the r e s u l t s by doing a r e s i d u a l s a n a l y s i s :
p l o t ( LModel )
#NOTES:
# 1 . Hit <CTRL>+<ENTER> r e p e a t e d l y f o r each p l o t !
# 2 . The r e s i d u a l s a n a l y s i s o f the model
#does not r e v e a l p a r t i c u l a r l y obvious problems .
# I t t h e r e f o r e concludes that the model i s well s u i t e d t o the r e a l s i t u a t i o n ,
#with a very high R−squared value
#3.8 What happens now by b u i l d i n g GAM? F i r s t , must i n s t a l l the mgcv package
i n s t a l l . packages ( " mgcv " )
l i b r a r y ( mgcv )
#3.9 Can use the gam ( ) f u n c t i o n that b u i l d s a GAM
GAMModel <− gam( stack . l o s s ~ s ( Air . Flow , k=7) +
s ( Water . Temp, k=7) + s ( Acid . Conc . , k = 7 ) ,
data= s t a c k l o s s )
#NOTES:
# 1 . In the previous command , the formula i s e x a c t l y l i k e
#the formula f o r a GLM except that smooth terms s ,
#can be added t o the r i g h t −hand s i d e t o s p e c i f y
# that the l i n e a r p r e d i c t o r depends
#on smooth f u n c t i o n s o f p r e d i c t o r s ( or l i n e a r f u n c t i o n s o f these ) .
# 2 . In each smooth term , compare the k value that i s the dimension o f the b a s i s
#used t o represent the smooth term .
# I f k i s not s p e c i f i e d then b a s i s s p e c i f i c d e f a u l t s are used .
#Note that these d e f a u l t s are e s s e n t i a l l y a r b i t r a r y ,
#and i t i s important t o check that they are not so small
292
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
# that they cause over−smoothing ( t o o l a r g e j u s t slows down computation ) .
#The value k−1 r e p r e s e n t s the maximum number o f a c t u a l degrees o f
#freedom a t t r i b u t a b l e t o each s p l i n e ;
#then i t must be checked in output that the s p e c i f i e d k value i s g r e a t e r than
#the number o f degrees o f freedom a t t r i b u t e d t o the corresponding s p l i n e ,
# otherwise the estimates can be q u i t e d i s t o r t e d .
#In the example f o r a l l three s p l i n e s , a number o f nodes equal t o seven i s s e t .
#3.10 Again , use the summary ( ) f u n c t i o n t o examine the model output
summary (GAMModel)
#3.11 Graph o f the smoothness f u n c t i o n
p l o t (GAMModel, s e l e c t =2)
#NOTES:A comparison o f the l i n e a r model with the GAM can be done by using
#the anova ( ) f u n c t i o n .
#This f u n c t i o n computes a n a l y s i s o f variance ( or deviance ) t a b l e s f o r one or
#more f i t t e d model o b j e c t s and returns an o b j e c t o f c l a s s anova .
#3.12 These o b j e c t s represent a n a l y s i s o f variance
#and a n a l y s i s o f deviance t a b l e s . When given a s i n g l e argument ,
# i t produces a t a b l e which t e s t s whether the model terms are s i g n i f i c a n t .
#When given a sequence o f o b j e c t s ,
#anova t e s t s the models against one another in the order s p e c i f i e d :
anova ( LModel , GAMModel)
#NOTES: From the a n a l y s i s o f the r e s u l t s i t i s concluded that
#the d i f f e r e n c e between the two models i s s i g n i f i c a n t ( P= 0 . 0 3 4 ) :
#the a d d i t i v e model best s u i t s the l i n e a r model dataset .
# 4 . REGRESSION TREES
#NOTES: The f u e l consumption o f v e h i c l e s has always been studied by
#the major manufacturers o f the e n t i r e planet .
#In an era c h a r a c t e r i z e d by o i l r e f u e l i n g problems and
293
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#even g r e a t e r a i r p o l l u t i o n problems ,
# f u e l consumption by v e h i c l e s has become a key f a c t o r .
#In t h i s example ,
#we w i l l b u i l d a r e g r e s s i o n t r e e with the purpose o f
# p r e d i c t i n g the f u e l consumption o f
#the v e h i c l e s according t o c e r t a i n c h a r a c t e r i s t i c s .
#4.1 Load the data
data ( mtcars )
#4.2 To d i s p l a y a compact summary o f the dataset simply type :
s t r ( mtcars )
#4.3 Thus , confirm 11 v a r i a b l e s with 32 obs , now can perform d e s c r i p t i v e summary
summary ( mtcars )
#NOTES:
# 1 . Before s t a r t i n g with data analysis ,
#conduct an e x p l o r a t o r y a n a l y s i s t o understand
#how the data i s d i s t r i b u t e d and e x t r a c t preliminary knowledge .
# F i r s t t r y t o f i n d out whether the v a r i a b l e s are r e l a t e d t o each other .
# 2 . Can do t h i s using the p a i r s ( ) f u n c t i o n t o c r e a t e a matrix o f
#sub−axes c o n t a i n i n g s c a t t e r p l o t s o f the columns o f a matrix .
#To reduce the number o f p l o t s in the matrix
#we l i m i t our a n a l y s i s t o j u s t f o u r p r e d i c t o r s :
# c y l i n d e r s , displacement , horsepower , and weight .
#4.4 The t a r g e t i s the mpg v a r i a b l e that
# c o n t a i n s measurements o f the miles per g a l l o n o f 32 sample cars .
p a i r s (mpg~ c y l +disp+hp+wt , data=mtcars )
#NOTES:
# 1 . To s p e c i f y the response and p r e d i c t o r s ,
#used the formula argument .
#Each term g i v e s a separate v a r i a b l e in the p a i r s p l o t ,
# so terms must be numeric v e c t o r s .
#Response was i n t e r p r e t e d as another v a r i a b l e , but not t r e a t e d s p e c i a l l y .
294
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
# 2 . Observing the p l o t s in the f i r s t l i n e ,
# i t can be noted that f u e l consumption i n c r e a s e s as the number o f c y l i n d e r s ,
#the engine displacement , the horsepower , and the weight o f the v e h i c l e i n c r e a s e .
#4.5 At t h i s p o i n t can use the t r e e ( ) f u n c t i o n t o b u i l d the r e g r e s s i o n t r e e .
# F i r s t , must i n s t a l l the t r e e package abd load the l i b r a r y .
i n s t a l l . packages ( " t r e e " )
library ( tree )
#4.6 Can use the t r e e ( ) f u n c t i o n t o b u i l d the r e g r e s s i o n t r e e
RTModel <− t r e e (mpg~ . , data = mtcars )
#NOTES: Only two arguments are passed : a formula and the dataset name .
#The l e f t −hand s i d e o f the formula ( response ) should be e i t h e r a numerical v e c t o r
#when a r e g r e s s i o n t r e e w i l l be f i t t e d .
#The ri gh t −hand s i d e should be a s e r i e s o f numeric v a r i a b l e s separated by + ;
# there should be no i n t e r a c t i o n terms .
#Both . and − are allowed : r e g r e s s i o n t r e e s can have o f f s e t terms .
#4.7 The r e s u l t s summary
RTModel
#NOTES: These r e s u l t s d e s c r i b e e x a c t l y each node in the t r e e .
# Information on each node i s presented in indented format .
# I t i s used t o i n d i c a t e the t r e e t o p o l o g y : that i s , i t i n d i c a t e s the parent
#and c h i l d r e l a t i o n s h i p s ( a l s o r e f e r r e d t o as primary and secondary s p l i t s ) .
#Also , t o denote a terminal node an a s t e r i s k ( * ) i s used .
#4.8 From the a n a l y s i s o f the r e s u l t s we can see a s e l e c t i o n o f v a r i a b l e s ,
# in f a c t between the 10 a v a i l a b l e v a r i a b l e s only three wt , c y l ,
#and hp were s e l e c t e d . More information can be obtained from the summary ( ) f u n c t i o n :
summary ( RTModel )
#4.9 As already mentioned p r e v i o u s l y , the output o f summary ( ) i n d i c a t e s that
# only three o f the v a r i a b l e s have been used in c o n s t r u c t i n g the t r e e .
#In the c o n t e x t o f a r e g r e s s i o n tree ,
#the deviance i s simply the sum o f squared e r r o r s f o r the t r e e .
295
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#Now, we can p l o t the r e g r e s s i o n t r e e :
p l o t ( RTModel )
t e x t ( RTModel )
#NOTES: The f i r s t one p l o t s the r e g r e s s i o n tree ,
#the second one adds the t e x t on the branches t o explain the work flow .
# Then look at what the r e g r e s s i o n t r e e has returned t o us .
#The f i r s t thing that seems obvious i s a s o r t o f i n d i c a t i o n o f
#the importance o f v a r i a b l e s .
#Already the c h o i c e o f three p r e d i c t o r s f o r the ten a v a i l a b l e makes us r e a l i z e #
# that these three are the ones that
#most a f f e c t the f u e l consumption o f cars i n s e r t e d in the dataset .
#Now can add that the most important p r e d i c t o r i s the weight o f the v e h i c l e ;
# in f a c t a weight l e s s than 2.26 l b s leads us t o a terminal knot that
# g i v e s a consumption estimate ( 3 0 . 0 7 miles / ( US) g a l l o n ) .
#Can then see that immediately a f t e r
#we f i n d the number o f c y l i n d e r s o f the engine , and , f i n a l l y , the horsepower .
# 5 . SUPPORT VECTOR REGRESSION (SVR)
#5.1 I n s t a l l the required package e1071 and load the l i b r a r y
i n s t a l l . packages ( " e1071 " )
l i b r a r y ( e1071 )
#5.2 To perform a r e g r e s s i o n t r e e example ,
#as always do , begin from the data .
#This time we w i l l c r e a t e a s e r i e s o f sample data t o understand the
# f u n c t i o n i n g o f the algorithm b e t t e r :
x <− seq ( 0 , 7 , by = 0 . 0 1 )
y <− s i n ( x ) + rnorm ( x , sd = 0 . 1 )
#NOTES: In the f i r s t l i n e o f code , cr eated the p r e d i c t o r x
#by using the seq ( ) f u n c t i o n .
#This f u n c t i o n generates a reg ular sequence from zero t o seven
#with an increment o f 0 . 0 1 .
#In the second l i n e , cre ated the response y using the s i n ( ) f u n c t i o n .
296
Electronic copy available at: https://ssrn.com/abstract=3863563
8.11. NONLINEAR
#To make the d i s t r i b u t i o n more real ,
#we added n o i s e with the rnorm ( ) f u n c t i o n that
# generates random d e v i a t e s with the same length o f x and
#with a standard d e v i a t i o n o f 0 . 1
#5.3 Start by b u i l d i n g a simple l i n e a r r e g r e s s i o n model using the lm ( ) f u n c t i o n :
LModel<−lm ( y~x )
summary ( LModel )
#5.4 To understand how the l i n e a r r e g r e s s i o n model works b e t t e r ,
#can p l o t the d i s t r i b u t i o n and the r e g r e s s i o n l i n e :
plot (x , y )
a b l i n e ( LModel )
#NOTES: I t i s easy t o p e r c e i v e
# that a l i n e a r r e g r e s s i o n model cannot p r e d i c t t h i s type o f d i s t r i b u t i o n .
#5.5 In order t o be able t o compare the l i n e a r r e g r e s s i o n with the SVR,
# thatare going t o f i t , can c a l c u l a t e the MSE.
#The MSE measures the average o f the squares o f the e r r o r s or d e v i a t i o n s .
#To c a l c u l a t e MSE, we f i r s t w i l l c a l c u l a t e the p r e d i c t e d value o f the model :
PredLModel<− p r e d i c t ( LModel )
#5.6 Now, can c a l c u l a t e the MSE:
mse1 <− mean ( ( y − PredLModel ) ^ 2 )
#5.7 The r e s u l t i s shown as f o l l o w s :
mse1
#5.8 This i s now the moment t o b u i l d the SVR model :
SVModel <− svm ( x , y )
#5.9 The svm ( ) f u n c t i o n can be used t o
# carry out general r e g r e s s i o n and c l a s s i f i c a t i o n ( o f nu and e p s i l o n type ) ,
#as well as d e n s i t y estimation .
#A formula i n t e r f a c e i s provided t o i n s e r t more than one p r e d i c t o r .
#To obtain some s t a t i s t i c s o f the model , now can use the summary ( ) f u n c t i o n
summary ( SVModel )
297
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#5.10 To compare the SVR model with the l i n e a r r e g r e s s i o n model ,
#now c a l c u l a t e the MSE.
#But as have done before , must f i r s t c a l c u l a t e the p r e d i c t i o n s o f the new model :
PredSVModel <− p r e d i c t ( SVModel )
#5.11 Now, c a l c u l a t e the MSE again , f o r new model type
mse2 <− mean ( ( y − PredSVModel ) ^ 2 )
mse2
#5.12 The MSE o f the l i n e a r model i s roughly 34 times g r e a t e r than that
# o f the SVR model .
#To get a g r a p h i c a l overview o f the goodness o f f i t o f the SVR model , p l o t a graph :
plot (x , y )
p o i n t s ( x , PredSVModel )
a b l i n e ( LModel )
#NOTES:
# 1 . In the f i r s t l i n e o f the code j u s t proposed traced the data d i s t r i b u t i o n .
#In the second , have traced the p o i n t s
# that represent the p r e d i c t i o n s o f the SVR model .
# Finally , have traced the l i n e a r r e g r e s s i o n l i n e ,
#as shown in the f o l l o w i n g f i g u r e : ( see a l s o PowerPoint l a s t s l i d e )
# 2 . The f i g u r e shows the p r e d i c t i o n c a p a c i t y o f the SVR model .
#In addition , the i n a b i l i t y o f the l i n e a r model t o
# represent such d i s t r i b u t i o n i s e q u a l l y c l e a r .
298
Electronic copy available at: https://ssrn.com/abstract=3863563
8.12. MULTIVARIATE & SPECIAL DISTRIBUTIONS TESTING
8.12
Multivariate & special distributions testing
#MULTIVARIATE & SPECIAL DISTRIBUTIONS TESTING_EXERCISE
# 0 . Can ask R f o r help on t h i s t o p i c
help . search ( " normality " )
# 0 . Shapiro−Wilk Normality Test i s the one that we s h a l l look at
names
# 1 . WILK−SHAPIRO TEST
#1.1 I n s t a l l the package c a l l e d mvnormtest , and load the l i b r a r y
i n s t a l l . packages ( " mvnormtest " )
l i b r a r y ( mvnormtest )
#1.2 Load the data
data ( a t t i t u d e )
#1.3 Need t o transpose the data s e t b e f o r e using the Shapiro−Wilk t e s t .
#The t ( ) f u n c t i o n i s used t o transpose the data
# 1 . 3 . 1 Transpose the data
mydata = t ( a t t i t u d e )
# 1 . 3 . 2 Run the W−S t e s t on the transposed data
mshapiro . t e s t ( mydata )
#NOTES:
# 1 . The Shapiro−Wilk p value o f .0002 i n d i c a t e s that
#the m u l t i v a r i a t e normality assumption does not hold .
# 2 . I t i n d i c a t e s that one or
#more i n d i v i d u a l v a r i a b l e s are not normally d i s t r i b u t e d .
# 2 . JARQUE−BERA TEST
#NOTES:
# 1 . Use another R package c a l l e d normtest
# 2 . J−B t e s t f o r obs and r e g r e s s i o n r e s i d u a l s i s l o c a t e d in normtest
299
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#2.1 I n s t a l l and load the l i b r a r y
i n s t a l l . packages ( " normtest " )
l i b r a r y ( normtest )
#2.2 Now run the j b . norm . t e s t on the a t t i t u d e data s e t
j b . norm . t e s t ( a t t i t u d e )
#NOTES: The Jarque−Bera t e s t f o r normality agreed with
#the Shapiro−Wilk t e s t that one or more v a r i a b l e s are not normally d i s t r i b u t e d .
# 3 . CHECK INDIVIDUAL VARIABLE SKEWNESS, KURTOSIS & NORMALITY
# To perform these checks in R can use the normwhn . t e s t
#3.1 F i r s t i n s t a l l the normwhn . t e s t package and load the package
i n s t a l l . packages ( " normwhn . t e s t " )
l i b r a r y ( normwhn . t e s t )
#3.2 run the f o l l o w i n g t e s t
normality . t e s t 1 ( a t t i t u d e )
#NOTES:
# 1 . This normality t e s t i n c l u d e s i n d i v i d u a l v a r i a b l e skewness
#and k u r t o s i s values and an omnibus t e s t o f normality .
# 2 . The r e s u l t s i n d i c a t e d that the v a r i a b l e s o v e r a l l
# did not contain skewness or k u r t o s i s . ( see PowerPoint f o r t a b l e ) .
# 3 . The omnibus normality t e s t i n d i c a t e d
# that the data were normally d i s t r i b u t e d , Z = 1 2 . 8 4 , df = 14 , p = . 5 4 .
# 4 . EXAMINING EACH INDIVIDUAL VARIABLE NORMALITY
#NOTES: f o c u s on the Anderson−Darling (A−D) t e s t as recommended
#by Stephens , M. A . ( 1 9 8 6 )
#4.1 I n s t a l l and load the l i b r a r y
i n s t a l l . packages ( " n o r t e s t " )
300
Electronic copy available at: https://ssrn.com/abstract=3863563
8.12. MULTIVARIATE & SPECIAL DISTRIBUTIONS TESTING
library ( nortest )
attach ( a t t i t u d e )
#4.2 A l l f i v e t e s t s on ’ rating ’
ad . t e s t ( r a t i n g )
# Anderson−Darling
cvm . t e s t ( r a t i n g )
# Cramer−Von Mises
l i l l i e . test ( rating )
# Kolmogorov−Smirnov
pearson . t e s t ( r a t i n g )
# Pearson Chi−square
sf . test ( rating )
# Shapiro−Francia
#4.3 Repeat f o r the v a r i a b l e ’ complaints ’
ad . t e s t ( complaints )
cvm . t e s t ( complaints )
l i l l i e . t e s t ( complaints )
pearson . t e s t ( complaints )
s f . t e s t ( complaints )
#4.4 Likewise we repeat the 5 t e s t s f o r v a r i a b l e ’ p r i v i l e g e s ’
ad . t e s t ( p r i v i l e g e s )
cvm . t e s t ( p r i v i l e g e s )
l i l l i e . test ( privileges )
pearson . t e s t ( p r i v i l e g e s )
sf . test ( privileges )
#4.5 And a l s o f o r the v a r i a b l e ’ learning ’
ad . t e s t ( l e a r n i n g )
cvm . t e s t ( l e a r n i n g )
l i l l i e . t e st ( learning )
pearson . t e s t ( l e a r n i n g )
s f . t e s t ( learning )
#4.6 Do the same t e s t s f o r v a r i a b l e c a l l e d ’ r a i s e s ’
ad . t e s t ( r a i s e s )
cvm . t e s t ( r a i s e s )
l i l l i e . test ( raises )
pearson . t e s t ( r a i s e s )
sf . test ( raises )
301
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#4.7 Now perform the same t e s t s f o r the v a r i a b l e ’ c r i t i c a l ’
ad . t e s t ( c r i t i c a l )
cvm . t e s t ( c r i t i c a l )
l i l l i e . test ( c r i t i c a l )
pearson . t e s t ( c r i t i c a l )
sf . test ( c r i t i c a l )
#4.8 Finally , repeat the t e s t s f o r the l a s t v a r i a b l e c a l l e d ’ advance ’
ad . t e s t ( advance )
cvm . t e s t ( advance )
l i l l i e . t e s t ( advance )
pearson . t e s t ( advance )
s f . t e s t ( advance )
#NOTES:
# 1 . A l l f i v e normality t e s t s showed that
#the v a r i a b l e c r i t i c a l v i o l a t e d the normality assumption .
# 2 . This i s why the Shapiro−Wilk and Jarque−Bera t e s t s i n d i c a t e d that
#the m u l t i v a r i a t e normality assumption was not met .
# 3 . However , t h i s s i n g l e v a r i a b l e , c r i t i c a l ,
#although not normally d i s t r i b u t e d ,
#would g e n e r a l l y not a f f e c t the power t o
# d e t e c t a d i f f e r e n c e in a m u l t i v a r i a t e s t a t i s t i c a l a n a l y s i s ( Stevens , 2009)
# 5 . DETERMINANT OF A MATRIX
#5.1 Create square c o r r e l a t i o n matrix
mycor = c o r ( a t t i t u d e )
#5.2 Next , the determinant i s computed using the R f u n c t i o n det ( )
det ( mycor )
#NOTES:
#The d e f a u l t c o r r e l a t i o n type in c o r ( ) i s the Pearson c o r r e l a t i o n c o e f f i c i e n t .
#5.3 The R commands t o c r e a t e a variance −co var ia nce matrix and
# t o obtain the determinant o f the matrix are as f o l l o w s :
302
Electronic copy available at: https://ssrn.com/abstract=3863563
8.12. MULTIVARIATE & SPECIAL DISTRIBUTIONS TESTING
#Create a square variance −co var ia nce matrix
mycov = cov ( a t t i t u d e )
#5.4 Determinant o f the variance −co var ia nce matrix
det ( mycov )
#NOTES: The determinant ( g e n e r a l i z e d variance ) o f the matrix i s p o s i t i v e ,
# t h e r e f o r e m u l t i v a r i a t e s t a t i s t i c a l a n a l y s i s can proceed .
#5.5 I f you want decimal places , rather than s c i e n t i f i c notation ,
# i s s u e the f o l l o w i n g command :
o p t i o n s ( scipen = 999)
#NOTES: An a d d i t i o n a l R f u n c t i o n can be u s e f u l
#when wanting a c o r r e l a t i o n matrix from a variance −co var ia nce matrix .
#5.6 The f u n c t i o n cov2cor ( ) s c a l e s a cov ar ian ce matrix
#by i t s diagonal t o compute the c o r r e l a t i o n matrix . The command i s as f o l l o w s :
# Convert cov ari anc e matrix t o c o r r e l a t i o n matrix .
mymatrix = cov2cor ( mycov )
#5.7 Now p r i n t out c o r r e l a t i o n matrix
mymatrix
#5.8 Get t h i s same c o r r e l a t i o n matrix i f l i s t the matrix mycor ,
#which was p r e v i o u s l y cre ated from the a t t i t u d e data s e t .
mycor
# 6 . EQUITY OF VARIANCE−COVARIANCE MATRIX
#6.1 The R commands t o add a group membership v a r i a b l e are as f o l l o w s .
#Create group v a r i a b l e 15 boys and 15 g i r l s
group = rep ( c ( " boy " , " g i r l " ) , c ( 1 5 , 1 5 ) )
#6.2 Now merge a t t i t u d e f i l e with the new group v a r i a b l e
newdata = data . frame ( a t t i t u d e , group )
303
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#6.3 Check the f i r s t 10 rows o f newdata
head ( newdata , n = 10)
#6.4 The within variance −co var ia nce matrices and a s s o c i a t e d determinants
# f o r each group can now be c a l c u l a t e d with the f o l l o w i n g R commands
# 6 . 4 . 1 Stops s c i e n t i f i c n o t a t i o n output
o p t i o n s ( scipen = 999)
# 6 . 4 . 2 S e l e c t s only boys
boys = newdata [ 1 : 1 5 , ]
# 6 . 4 . 3 Create c ova ria nc e matrix
boycov = cov ( boys [ , − 8 ] )
# 6 . 4 . 4 Determinant computed o f matrix
det ( boycov )
#6.5 Now we do the same f o r g i r l s
# 6 . 5 . 1 S e l e c t only g i r l s from the data s e t
g i r l s = newdata [ 1 6 : 3 0 , ]
# 6 . 5 . 2 Create c ova ria nc e matrix
g i r l c o v = cov ( g i r l s [ , − 8 ] )
# 6 . 5 . 3 Compute determinant
det ( g i r l c o v )
#NOTES: The determinants o f the boys
#and g i r l s variance −co var ia nce matrices were p o s i t i v e ,
#thus m u l t i v a r i a t e s t a t i s t i c a l a n a l y s i s can proceed .
#6.6 Can obtain the d e s c r i p t i v e s t a t i s t i c s
#by using the R package psych and the describeBy ( ) f u n c t i o n .
# I n s t a l l and load the psych package ,
#and then we i s s u e the command f o r the f u n c t i o n .
i n s t a l l . packages ( " psych " )
l i b r a r y ( psych )
#6.7 Run the describeBy ( ) f u n c t i o n
describeBy ( newdata , group = group )
#NOTES: The standard d e v i a t i o n s , hence variance values ,
304
Electronic copy available at: https://ssrn.com/abstract=3863563
8.12. MULTIVARIATE & SPECIAL DISTRIBUTIONS TESTING
#do look a l i t t l e d i f f e r e n t between the boys and the g i r l s a c r o s s the v a r i a b l e s .
#6.8 Can c r e a t e and l i s t the separate cov ar ian ce matrices
# f o r each group in the newdata data s e t by using the l a p p l y ( ) function ,
#where the group v a r i a b l e i s d e l e t e d as f o l l o w s
cov . l i s t = l a p p l y ( unique ( newdata$group ) ,
f u n c t i o n ( x ) cov ( newdata [ newdata$group==x , − 8 ] ) )
#6.9 L i s t the boys ’ cov ari an ce matrix as f o l l o w s :
cov . l i s t [ 1 ]
#6.10 Do l i k e w i s e f o r the g i r l s
cov . l i s t [ 2 ]
#NOTES:
# 1 . The f i r s t approach obtained the separate group variance −co var ia nce matrices
#much e a s i l y ( l e s s s o p h i s t i c a t e d programming ) .
# 2 . Can e a s i l y convert these separate variance −co var ia nce matrices i n t o
# c o r r e l a t i o n matrices using the cov ( ) f u n c t i o n and the cov2cor ( ) f u n c t i o n
#as shown b e f o r e .
# 3 . The covar ian ce matrices exclude the grouping v a r i a b l e
#by i n d i c a t i n g a value ( column l o c a t i o n in the matrix )
# in the s e l e c t i o n o f v a r i a b l e s t o i n c l u d e .
#6.11 The two s e t s o f R commands are as f o l l o w s
boys = newdata [ 1 : 1 5 , ]
boycov = cov ( boys [ , − 8 ] )
boycor = cov2cor ( boycov )
boycor
# and
g i r l s = newdata [ 1 6 : 3 0 , ]
g i r l c o v = cov ( g i r l s [ , − 8 ] )
g i r l c o r = cov2cor ( g i r l c o v )
girlcor
# 7 . BOX−M TEST
#NOTES: The b a s i c example shown
305
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#here however d e c l a r e s the grouping v a r i a b l e as a f a c t o r ,
# t h e r e f o r e the complete code would be as f o l l o w s :
#7.1 I n s t a l l the R b i o t o o l s package and load the package and Load the package
i n s t a l l . packages ( " b i o t o o l s " )
library ( biotools )
#7.2 Load t h i s data s e t
data ( i r i s )
#7.3 S e l e c t v a r i a b l e s f o r matrix , minus grouping v a r i a b l e
factor ( i ri s [ ,5])
#7.4 Run f u n c t i o n with v a r i a b l e s and group v a r i a b l e s p e c i f i e d
boxM ( i r i s [ , − 5] , i r i s [ , 5 ] )
#NOTES:
# 1 . The use o f i r i s [ , − 5] s e l e c t s v a r i a b l e s and i r i s [ , 5 ]
# i n d i c a t e s grouping v a r i a b l e in data f i l e
# 2 . In the i r i s data set ,
#the n u l l hypothesis o f equal variance −co var ia nce matrices was r e j e c t e d .
# 3 . The groups did have d i f f e r e n t variance −co var ia nce matrices .
# 4 . A m u l t i v a r i a t e a n a l y s i s would t h e r e f o r e be suspect ,
# e s p e c i a l l y i f non−normality e x i s t e d among the v a r i a b l e s .
#7.5 The R code s t e p s are l i s t e d below .
# F i r s t , i n s t a l l the package from the main menu .
#Next , load the b i o t o o l s package and the data s e t .
#You must d e c l a r e the v a r i a b l e group as a f a c t o r
# b e f o r e running the Box M t e s t f u n c t i o n . The s e t o f R commands are as f o l l o w s :
# 7 . 5 . 1 I n s t a l l the R b i o t o o l s package
i n s t a l l . packages ( " b i o t o o l s " )
# 7 . 5 . 2 Load the package
library ( biotools )
306
Electronic copy available at: https://ssrn.com/abstract=3863563
8.12. MULTIVARIATE & SPECIAL DISTRIBUTIONS TESTING
# 7 . 5 . 3 Load t h i s data s e t
data ( newdata )
# 7 . 5 . 4 Declare grouping v a r i a b l e a f a c t o r
f a c t o r ( group )
# 7 . 5 . 5 Run Box−M f u n c t i o n
boxM ( newdata [ , − 8] , newdata [ , 8 ] )
# The r e s u l t s i n d i c a t e d that the boys
#and the g i r l s variance −co var ia nce matrices are equal ;
# that i s , the chi −square t e s t was non−s i g n i f i c a n t .
307
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.13
Principal components analysis
#PRINCIPAL COMPONENTS ANALYSIS (PCA)_EXERCISE
#1.1 I n s t a l l and Library Packages
i n s t a l l . packages ( "HSAUR2" )
i n s t a l l . packages ( " t o o l s " )
library ( tools )
l i b r a r y (HSAUR2)
i n s t a l l . packages ( "MVA" )
l i b r a r y (MVA)
demo ( " Ch−PCA" )
i n s t a l l . packages ( " l a t t i c e " )
library ( lattice )
#1.2 Check the data
blood_corr
blood_sd
#1.3 There are c o n s i d e r a b l e d i f f e r e n c e s between these standard d e v i a t i o n s ,
#and can apply PCA t o both the c ov ari anc e
#and c o r r e l a t i o n matrix o f the data using the f o l l o w i n g code
cov <− princomp ( covmat = blood_cov )
summary ( blood_pcacov , l o a d i n g s = TRUE)
blood_pc ac or <− princomp ( covmat = b l o o d _ c o r r )
summary ( blood_pcacor , l o a d i n g s = TRUE)
#NOTES:
# 1 . The blanks in t h i s output represent very small values .
# 2 . Examining the r e s u l t s ,
#can see that the p r i n c i p a l components o f the cov ari anc e matrix i s
# l a r g e l y dominated by a s i n g l e v a r i a b l e ,
308
Electronic copy available at: https://ssrn.com/abstract=3863563
8.13. PRINCIPAL COMPONENTS ANALYSIS
#whereas those f o r the c o r r e l a t i o n matrix have moderate−s i z e d c o e f f i c i e n t s
#on s e v e r a l o f the v a r i a b l e s .
# 3 . And the f i r s t component from the co va ria nce matrix accounts f o r almost 99%
# o f the t o t a l variance o f the observed v a r i a b l e s .
#The components o f the c ov ari anc e matrix are completely dominated by
#the f a c t that the variance o f the p l a t e v a r i a b l e i s roughly 400 times
# l a r g e r than the variance o f any o f the seven other v a r i a b l e s .
# 4 . Consequently , the p r i n c i p a l components from the co va ria nce matrix
#simply r e f l e c t the order o f the s i z e s o f the variances o f the observed v a r i a b l e s .
#The r e s u l t s from the c o r r e l a t i o n matrix t e l l us ,
# in p a r t i c u l a r , that a weighted c o n t r a s t o f the f i r s t f o u r
#and l a s t f o u r v a r i a b l e s i s the l i n e a r f u n c t i o n with the l a r g e s t variance .
# 5 . This example i l l u s t r a t e s
# that when v a r i a b l e s are on very d i f f e r e n t s c a l e s or
#have very d i f f e r e n t variances ,
#a p r i n c i p a l components a n a l y s i s o f the data should be performed
#on the c o r r e l a t i o n matrix , not on the cov ari anc e matrix .
#1.4 P l o t the Scree diagram ( please note type = l not 1 , as i s
p l o t ( blood_pcacor$sdev ^2 ,
xlab = " Component number " ,
ylab = " Component variance " ,
type = " l " ,
main = " Scree diagram " )
#1.5 P l o t the Log ( Eigenvalue ) diagram t o o
p l o t ( l o g ( blood_pcacor$sdev ^ 2 ) ,
xlab = " Component number " ,
ylab = " l o g ( Component variance ) " ,
type = ’ l ’ ,
main = " Log ( eigenvalue ) diagram " )
# 2 . EXAMPLES
309
Electronic copy available at: https://ssrn.com/abstract=3863563
’ l ’ for log )
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#2.1 Head data o f the f i r s t two sons
head_dat <− headsize [ , c ( " head1 " , " head2 " ) ]
colMeans ( head_dat )
cov ( head_dat )
#2.2 The p r i n c i p a l components o f these data ,
# e x t r a c t e d from t h e i r c ov ari anc e matrix , can be found using
head_pca <− princomp ( x = head_dat )
head_pca
#2.3 Now p r i n t formal t a b l e
p r i n t ( summary ( head_pca ) , l o a d i n g s = TRUE)
# Back t o PowerPoint s l i d e s f o r b r i e f i n t e r p r e t a t i o n
#2.4 Now c r e a t e the two p l o t s shown on the s l i d e s
a1 < − 183.84 − 0.721 * 185.72/0.693
b1 < − 0.721/0.693
a2 < − 183.84 − ( − 0.693 * 185.72/0.721)
b2< −− 0.693/0.721
p l o t ( head_dat , xlorab = " F i r s t son ’ s head length (mm) " ,
ylab = " Second son ’ s head length " )
a b l i n e ( a1 , b1 )
a b l i n e ( a2 , b2 , l t y = 2 )
#Note : l i n e t y p e ( l t y ) = 2 = dotted l i n e
#2.5 For the second s c a t t e r p l o t
xlim <− range ( head_pca$scores [ , 1 ] )
p l o t ( head_pca$scores , xlim = xlim , ylim = xlim )
# 3 . OLYMPIC HEPTATHLON RESULTS
#3.1 I n s t a l l and l i b r a r y package
i n s t a l l . packages ( "HSAUR2" )
l i b r a r y (HSAUR2)
310
Electronic copy available at: https://ssrn.com/abstract=3863563
8.13. PRINCIPAL COMPONENTS ANALYSIS
#3.2 Can see the l i s t o f 25 female co mpe ti to rs
data ( heptathlon )
show ( heptathlon )
#NOTES:
# 1 . But b e f o r e undertaking the p r i n c i p a l components analysis ,
# i t i s good data a n a l y s i s p r a c t i c e t o carry out
#an i n i t i a l assessment o f the data using one
# or another o f the graphics d e s c r i b e d b e f o r e .
# 2 . Some numerical summaries may a l s o be h e l p f u l
# b e f o r e we begin the main a n a l y s i s .
#And b e f o r e any o f these ,
# i t w i l l help t o s c o r e a l l seven events in the same d i r e c t i o n so that
#" l a r g e " values are i n d i c a t i v e o f a " b e t t e r " performance .
# 3 . The R code f o r r e v e r s i n g the values f o r some events ,
#then c a l c u l a t i n g the c o r r e l a t i o n c o e f f i c i e n t s between
#the ten events and f i n a l l y c o n s t r u c t i n g the s c a t t e r p l o t matrix o f the data i s
heptathlon$hurdles <− with ( heptathlon ,
max( hurdles)− hurdles )
show ( heptathloon$hurdles ) heptathlon$run200m <− with ( heptathlon ,
max( run200m)−run200m )
heptathlon$run800m <− with ( heptathlon ,
max( run800m)−run800m )
s c o r e <− which ( colnames ( heptathlon ) == " s c o r e " )
round ( c o r ( heptathlon [ , − s c o r e ] ) , 2 )
p l o t ( heptathlon [ , − s c o r e ] )
#NOTES: Gives a s c a t t e r p l o t matrix o f the seven heptathlon events a f t e r
# transforming some v a r i a b l e s so that
311
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# f o r a l l events l a r g e values are i n d i c a t i v e o f a b e t t e r performance e . g .
# think f a s t e r 800m compared t o longer j a v e l i n throw !
#3.3 Return t o s l i d e s f o r a d i s c u s s i o n o f t h i s matrix
#3.4 Removing the competitor from Papua New Guinea (PNG) , the r e l e v a n t R code i s :
heptathlon <− heptathlon[− grep ( "PNG" , rownames ( heptathlon ) ) , ]
s c o r e <− which ( colnames ( heptathlon ) == " s c o r e " )
round ( c o r ( heptathlon [ , − s c o r e ] ) , 2 )
#NOTES:
# 1 . The new s c a t t e r p l o t matrix i s shown .
# 2 . Several o f the c o r r e l a t i o n s are changed t o some degree from those shown
# b e f o r e removal o f the PNG competitor ,
# p a r t i c u l a r l y the c o r r e l a t i o n s i n v o l v i n g the j a v e l i n event ,
#where the very small c o r r e l a t i o n s between performances in t h i s event
#and the others have i ncr ease d c o n s i d e r a b l y .
# 3 . Given the r e l a t i v e l y l a r g e o v e r a l l change
# in the c o r r e l a t i o n matrix produced by omitting the PNG competitor ,
#we s h a l l e x t r a c t the p r i n c i p a l components o f
#the data from the c o r r e l a t i o n matrix a f t e r t h i s omission .
#3.5 The p r i n c i p a l components can now be found using
heptathlon_pca <− prcomp ( heptathlon [ , − s c o r e ] ,
s c a l e = TRUE)
p r i n t ( heptathlon_pca )
p l o t ( heptathlon [ , − s c o r e ] ,
pch = " . " ,
cex = 1 . 5 )
#Notes : p l o t c h a r a c t e r i s pch and s c a l e or t e x t cex =1.5 => 150%
#3.6 The summary method can be used f o r f u r t h e r i n s p e c t i o n o f the d e t a i l s
summary ( heptathlon_pca )
#3.7 The l i n e a r combination f o r the f i r s t p r i n c i p a l component i s 2
312
Electronic copy available at: https://ssrn.com/abstract=3863563
8.13. PRINCIPAL COMPONENTS ANALYSIS
a1 <− he pt a t h l o n _ p c a $ r o t a t i o n [ , 1 ]
a1
#NOTES: see that the hurdles and long jump events r e c e i v e the highest weight
#but the j a v e l i n r e s u l t i s l e s s important
#3.8 For computing the f i r s t p r i n c i p a l component ,
#the data need t o be re−s c a l e d a p p r o p r i a t e l y .
#The c e n t e r and the s c a l i n g used by prcomp i n t e r n a l l y
#can be e x t r a c t e d from the heptathlon_pca via
c e n t e r <− heptathlon_pca$center
s c a l e <− heptathlon_pca$scale
# Apply the s c a l e f u n c t i o n t o the data
#and multiply i t with the l o a d i n g s matrix
# in order t o compute the f i r s t p r i n c i p a l component s c o r e f o r each competitor
hm <− as . matrix ( heptathlon [ , − s c o r e ] )
drop ( s c a l e (hm, c e n t e r = center , s c a l e = s c a l e ) %*% h e p t a t h l o n _ p c a $ r o t a t i o n [ , 1 ] )
#3.9 or more conveniently ,
#by e x t r a c t i n g the f i r s t from a l l pre−computed p r i n c i p a l components
p r e d i c t ( heptathlon_pca ) [ , 1 ]
#NOTES: The f i r s t two components account f o r 75% o f the variance .
#3.10 A b a r p l o t o f each component ’ s variance shows the f i r s t two components dominate
p l o t ( heptathlon_pca )
#3.11 The c o r r e l a t i o n between the s c o r e given t o each a t h l e t e
#by the standard s c o r i n g system used f o r the heptathlon
#and the f i r s t p r i n c i p a l component s c o r e can be found from
c o r ( heptathlon$score , heptathlon_pca$x [ , 1 ] )
#NOTES: This i m p l i e s that the f i r s t p r i n c i p a l component i s in good agreement
#with the s c o r e assigned t o the a t h l e t e s by o f f i c i a l Olympic r u l e s ;
#a s c a t t e r p l o t o f the o f f i c i a l s c o r e and the f i r s t p r i n c i p a l component i s given
313
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#3.12 The f a c t that the c o r r e l a t i o n i s negative i s unimportant here because o f
#the a r b i t r a r i n e s s o f the s i g n s o f the c o e f f i c i e n t s
# d e f i n i n g the f i r s t p r i n c i p a l component ;
# i t i s the magnitude o f the c o r r e l a t i o n that i s important .
p l o t ( heptathlon$score , heptathlon_pca$x [ , 1 ] )
# 4 . AIR POLLUTION IN US CITIES
data ( " U S a i r p o l l u t i o n " )
show ( U S a i r p o l l u t i o n )
#4.1 C o r r e l a t i o n matrix and PC components
c o r ( US a i r p o l l u t i o n [ , − 1 ] )
usair_pca <− princomp ( U S a i r p o l l u t i o n [ , − 1] , c o r = TRUE)
data ( " U S a i r p o l l u t i o n " , package = "HSAUR2" )
panel . h i s t <− f u n c t i o n ( x ,
...) {
usr <− par ( " usr " ) ; on . e x i t ( par ( usr ) )
par ( usr = c ( usr [ 1 : 2 ] , 0 , 1 . 5 ) )
h <− h i s t ( x , p l o t = FALSE)
breaks <− h$breaks ; nB <− length ( breaks )
y <− h$counts ; y <− y / max( y )
r e c t ( breaks[−nB ] , 0 , breaks [ − 1] , y , c o l =" grey " , . . . )
}
USairpollution$negtemp <− USairpollution$temp * ( − 1)
USairpollution$temp <− NULL
p a i r s ( U S a i r p o l l u t i o n [ , − 1] , diag . panel = panel . h i s t ,
pch = " . " , cex = 1 . 5 )
#4.2 Now examine the summary output
summary ( usair_pca , l o a d i n g s = TRUE)
#4.3 B i v a r i a t e b o x p l o t s
l i b r a r y (MVA)
314
Electronic copy available at: https://ssrn.com/abstract=3863563
8.13. PRINCIPAL COMPONENTS ANALYSIS
#NOTES: Contains the bvbox f u n c t i o n used in the next p l o t
p a i r s ( u s a i r _ p c a $ s c o r e s [ , 1 : 3 ] , ylim = c ( − 6 , 4 ) , xlim = c ( − 6 , 4 ) ,
panel = f u n c t i o n ( x , y ,
...) {
t e x t ( x , y , ab brevi ate ( row . names ( U S a i r p o l l u t i o n ) ) ,
cex = 0 . 6 )
bvbox ( cbind ( x , y ) , add = TRUE)
})
#4.4 Sulphur d i o x i d e depending on c o n c e n t r a t i o n p l o t s
out <− sapply ( 1 : 6 , f u n c t i o n ( i ) {
p l o t ( USairpollution$SO2 , u s a i r _ p c a $ s c o r e s [ , i ] ,
xlab = paste ( "PC" , i , sep = " " ) ,
ylab = " Sulphur d i o x i d e c o n c e n t r a t i o n " )
})
#4.5 PCA
u s a i r _ r e g <− lm (SO2 ~ usair_pca$scores ,
data = U S a i r p o l l u t i o n )
summary ( u s a i r _ r e g )
#NOTES: Clearly , the f i r s t p r i n c i p a l component s c o r e i s the most p r e d i c t i v e o f
#sulphur d i o x i d e conce ntration ,
#but i t i s a l s o c l e a r that components
#with small variance do not n e c e s s a r i l y have small c o r r e l a t i o n s with the response .
315
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.14
Simulating distributions for hypothesis testing
# SIMULATING DISTRIBUTIONS FOR HYPOTHESIS TESTING_EXERCISE
#NOTES:
# 1 . To r e j e c t the n u l l hypothesis
# i f the average number o f a r r e s t s in a sample n=30 i s 0 . 3 .
# 2 . Do not have the a n a l y t i c a l t o o l s t o compute t h i s p r o b a b i l i t y ,
#because we do not know how t o compute the d i s t r i b u t i o n o f the sample mean .
#We can , however , write some code here in R t o do the j o b .
# 3 . Simulate data from the n u l l d i s t r i b u t i o n .
#1.1 In the f o l l o w i n g p i e c e o f code ,
#we generate one m i l l i o n random samples o f
# s i z e 30 from the h i s t o r i c a l ( n u l l hypothesis ) d i s t r i b u t i o n ,
#and f o r each sample we c a l c u l a t e the mean .
nloop =1000000
sampmean=1: nloop
f o r ( i l o o p in 1 : nloop ) {
x=sample ( 0 : 3 , 3 0 , r e p l a c e =TRUE,
prob=c ( 0 . 6 5 , 0 . 1 8 , 0 . 1 0 , 0 . 0 7 ) )
sampmean [ i l o o p ]=mean ( x )
}
h i s t ( sampmean , br =150)
sum( sampmean< 0 . 3 ) / nloop
#NOTES:
# 1 . Fingding 2.6% o f the sample means are below 0 . 3 ,
# so the t e s t s i z e f o r t h i s d e c i s i o n r u l e i s alpha = .026
# 2 . The d i s t r i b u t i o n in the histogram i s d i s c r e t e ,
#because there are a f i n i t e number o f p o s s i b l e values
# that the sample mean can take .
#The simulated sample mean approximates
#the d i s t r i b u t i o n o f a mean o f a sample o f s i z e 30 , chosen " at random " .
# 3 . Because simulated such a l a r g e number o f samples ,
316
Electronic copy available at: https://ssrn.com/abstract=3863563
8.14. SIMULATING DISTRIBUTIONS FOR HYPOTHESIS TESTING
#the d i s t r i b u t i o n i s very c l o s e t o the true d i s t r i b u t i o n .
# 4 . Now, suppose want t o t e s t s i z e t o be the t r a d i t i o n a l alpha = . 0 5 .
#What should the d e c i s i o n r u l e be ?
#1.2 Can use our m i l l i o n sample means t o
# f i n d the value c f o r which ( approximately ) 5% are l e s s than c .
#The f o l l o w i n g R code accomplishes t h i s ,
# given the p r e v i o u s l y simulated d i s t r i b u t i o n :
ss= s o r t ( sampmean )
ss [ 0 . 0 5 * nloop ]
#NOTES:
# 1 . Finding the f i f t h p e r c e n t i l e o f the d i s t r i b u t i o n i s c = 0.333
# 2 . In other words ,
# i f r e j e c t H0 i f the sample average number o f a r r e s t s i s 0.333 or smaller ,
#then t e s t s i z e i s c l o s e t o alpha = . 0 5 .
# 3 . What i s the power o f t h i s t e s t
# i f the true d i s t r i b u t i o n i s given as in the second t a b l e ?
#1.3 Can again use simulations t o f i n d t h i s ,
#but t h i s time we sample from the d i s t r i b u t i o n a s s o c i a t e d with
#the a l t e r n a t i v e hypothesis :
nloop =1000000
sampmean=1: nloop
f o r ( i l o o p in 1 : nloop ) {
x=sample ( 0 : 3 , 3 0 , r e p l a c e =TRUE, prob=c ( 0 . 8 2 , 0 . 1 0 , 0 . 0 5 , 0 . 0 3 ) )
sampmean [ i l o o p ]=mean ( x )
}
h i s t ( sampmean , br =100)
sum( sampmean< = 0 . 3 3 3 ) / nloop
#NOTES: Finding that the power i s about 0.611
# 2 . SIX−SIDED DIE EXAMPLE
317
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES: 1 . The f o l l o w i n g code simulates the d i s t r i b u t i o n o f our t e s t s t a t i s t i c
#when the n u l l hypothesis i s true ( d i e i s f a i r ) .
# 2 . Within the loop , the d i e i s r o l l e d 60 times ,
#and the t e s t s t a t i s t i c i s computed .
# 3 . The values o f the T s t a t i s t i c are s t o r e d in t s t a t
n=60; nloop =1000000
t s t a t =1: nloop
f o r ( i l o o p in 1 : nloop ) {
r o l l s =sample ( 1 : 3 , 6 0 , r e p l a c e =TRUE, prob=c ( 1 / 2 , 1 / 3 , 1 / 6 ) )
nr=sum( r o l l s ==1)
nb=sum( r o l l s ==2)
ng=sum( r o l l s ==3)
t s t a t [ i l o o p ]= abs ( nr −30)+abs ( nb −20)+abs ( ng −10)
}
h i s t ( t s t a t , br =100 ,
main=" n u l l d i s t r i b u t i o n " ,
f r e q =FALSE,
ylab =" P r o b a b i l i t y " )
#NOTES: One m i l l i o n simulated t e s t s
#when the n u l l hypothesis i s true
# g i v e s a very p r e c i s e d i s t r i b u t i o n o f the t e s t s t a t i s t i c .
#2.1 I f the d e c i s i o n r u l e i s t o r e j e c t the n u l l hypothesis
#when T i s at l e a s t 17 , can f i n d the t e s t s i z e using :
sum( t s t a t >=17)/ nloop
#NOTES: t o f i n d a t e s t s i z e o f about 0.042
#2.2 I f r e q u i r e alpha <= . 0 1 ,
#we f i n d that the d e c i s i o n r u l e must be " r e j e c t H0 i f >=22" ,
#which pr ovid es a t e s t s i z e o f about 0.0093
sum( t s t a t >=22)/ nloop
#NOTES:
# 1 . Simulations w i l l a l s o provide the power o f t e s t under various s c e n a r i o s .
# 2 . Suppose the f o l l o w i n g are the true p r o b a b i l i t i e s f o r the c o l o u r s :
#red i s 0 . 6 5 , blue i s 0 . 3 and gold i s 0 . 0 5 .
318
Electronic copy available at: https://ssrn.com/abstract=3863563
8.14. SIMULATING DISTRIBUTIONS FOR HYPOTHESIS TESTING
# 3 . This i s s u b s t a n t i a l l y d i f f e r e n t from a f a i r die ,
#and we would l i k e t o have a high p r o b a b i l i t y o f d e t e c t i n g the d i f f e r e n c e
#and r e j e c t i n g the n u l l hypothesis .
#2.3 Repeat the above simulations with prob=c ( 0 . 6 5 , 0 . 3 0 , 0 . 0 5 )
n=60; nloop =1000000
t s t a t =1: nloop
f o r ( i l o o p in 1 : nloop ) {
r o l l s =sample ( 1 : 3 , 6 0 ,
r e p l a c e =TRUE,
prob=c ( 0 . 6 5 , 0 . 3 , 0 . 0 5 ) )
nr=sum( r o l l s ==1)
nb=sum( r o l l s ==2)
ng=sum( r o l l s ==3)
t s t a t [ i l o o p ]= abs ( nr −30)+abs ( nb −20)+abs ( ng −10)
}
h i s t ( t s t a t , br =100 ,
main=" a l t e r n a t i v e d i s t r i b u t i o n " ,
f r e q =FALSE,
ylab =" P r o b a b i l i t y " )
#2.4 To f i n d the power f o r the d e c i s i o n r u l e with the smaller t e s t s i z e , we use :
sum( t s t a t >=22)/ nloop
#NOTES:
# 1 . And the power i s about 0.347
# 2 . Perhaps we need t o change the design o f our experiment
#and decide t o r o l l the d i e a l a r g e r number o f times
# so that we can have both a small t e s t s i z e and l a r g e power .
# 3 . i t i s not uncommon f o r s t a t i s t i c i a n s t o r e s o r t
# t o simulations t o provide t e s t s i z e s , p−values and power .
# 3 . BERNOULLI & BINOMIAL RANDOM VARIABLES
#3.1 CANCER TREATMENT EXAMPLE
319
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES:
# 1 . The f u n c t i o n pbinom ( q , n , p ) w i l l return the p r o b a b i l i t y
# that Y b i s l e s s than or equal t o q , where Y~Binom ( n , p )
# 3 . 1 . 1 Therefore , t o get the p r o b a b i l i t y o f
# at l e a s t 16 out o f 20 p a t i e n t s going i n t o remission when p=0.8 type
1−pbinom ( 1 5 , 2 0 , 0 . 8 )
#NOTES:
# 1 . Get 0.63 or 63%
# 2 . Suppose that we want the power t o be at l e a s t 80%,
#and a l s o want t o t e s t s i z e alpha at most 0.01
# 3 . Need t o i n c r e a s e the sample s i z e
# 4 . Can t r y n=40
# 3 . 1 . 2 The f i r s t step i s t o f i n d the d e c i s i o n rule ,
#and f o r now we w i l l do t h i s by t r a i l and error , using pbinom
1−pbinom ( 3 1 , 4 0 , 0 . 6 )
#NOTES:
# 1 . Returns 0 . 0 0 6 1 . That i s , i f Y~Binom ( 4 0 , 0 . 6 ) , then P (Y>=32)=0.0061
# 2 . In other words ,
# i f our d e c i s i o n r u l e i s t o r e j e c t H0
# i f at l e a s t 32 out o f 40 p a t i e n t s go i n t o remission ,
#then our t e s t s i z e i s alpha = 0.0061
# ( I f our d e c i s i o n r u l e i s t o r e j e c t H0
# i f at l e a s t 31 out o f 40 p a t i e n t s go i n t o remission ,
#then our t e s t s i z e w i l l be g r e a t e r than 0 . 0 1 .
#For a d i s c r e t e random v a r i a b l e ,
#we cannot always get the exact t e s t s i z e that we want )
# 3 . The corresponding power f o r p=0.8 i s 1−pbinom ( 3 1 , 4 0 , 0 . 8 ) ,
#which i s about 0.593
# 4 . I f want the power t o be at l e a s t 0 . 8
# ( a common t a r g e t f o r these types o f t e s t s ) ,
w i l l need a l a r g e r sample s i z e
320
Electronic copy available at: https://ssrn.com/abstract=3863563
8.14. SIMULATING DISTRIBUTIONS FOR HYPOTHESIS TESTING
# 5 . Can repeat t h i s procedure : f o r the c h o i c e o f n ,
# f i n d the d e c i s i o n r u l e ( using p = 0 . 6 )
# that g i v e s alpha c l o s e t o but not exceeding 0 . 0 1 ;
#then using the d e c i s i o n rule , f i n d the power ( using p = 0 . 8 )
# 6 . The goal i s t o f i n d the s m a l l e s t n , conclude that f o r n=55 ,
#can r e j e c t H0 when at l e a s t 42 p a t i e n t s go i n t o remission ,
#and when p=0.8 t h i s t e s t has power 0 . 8 0 3 .
# 7 . Other u s e f u l R f u n c t i o n s r e l a t e d t o
#the binomial d i s t r i b u t i o n i n c l u d e dbinom , rbinom , and qbinom .
# 8 . The command dbinom ( x , n , p ) w i l l return P (X=x ) when X~Binom ( n , p ) .
# 7 . Can use rbinom t o generate random values from a binomial d i s t r i b u t i o n ,
# u s e f u l f o r checking c a l c u l a t i o n s .
#3.2 For example , suppose want t o v e r i f y that E[X(X− 1)]=n(1 −p ) p ^2.
#For p=0.4 and n=10 , n(1 −p ) p ^2=14.4;
#we can randomly generate a l a r g e number o f values from t h i s d i s t r i b u t i o n
#and f i n d the average o f the value times one minus the value :
x=rbinom ( 1 0 0 0 0 0 , 1 0 , 0 . 4 )
mean ( x * ( x − 1))
#NOTES:
# 1 . This code w i l l return a number very c l o s e t o 1 4 . 4 ,
#with a l i t t l e " random e r r o r "
# 2 . The " q u a n t i l e f u n c t i o n " qbinom ( r , n , p )
# w i l l return the s m a l l e s t value x such that P (X<=x) >= r .
# 3 . This qbinom command can be used t o
# f i n d the d e c i s i o n r u l e f o r the above example , instead o f using t r i a l and e r r o r .
# 4 . THE GEOMETRIC RANDOM VARIABLE
321
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#4.1 Geometric Random Variable Game
1−pgeom ( 2 9 , 1 / 6 )
#4.2 Now t r y the other case
1−pgeom ( 1 4 , 0 . 0 2 )
#NOTES: Returns 0.739
# 5 . THE POISSON RANDOM VARIABLE
#NOTES: Automobile crash example using ppois ( )
#5.1 ppois ( k , lambda )
ppois ( 5 , 1 . 4 7 )
# 6 . THE HYPERGEOMETRIC RANDOM VARIABLE
#6.1 The f o l l o w i n g b i t o f code draws f i v e parts from the box ,
#100 ,000 times , and counts the draws f o r
#which there are e x a c t l y two d e f e c t i v e parts
cnt2def =0; nloop =100000
f o r ( i l o o p in 1 : nloop ) {
draw=sample ( 1 : 2 5 , 5 )
i f (sum( draw >21)==2){ cnt 2def=cnt 2def +1}
}
cnt2def / nloop
#6.2 Math Class example
1−phyper ( 6 , 1 4 , 1 7 , 1 0 )
# 7 . MULTINOMIAL DISTRIBUTION
#7.1 Metal Rods machine example
#NOTES: The R f u n c t i o n dmultinom ( x , prob ) i s u s e f u l f o r
#computing multinomial p r o b a b i l i t i e s ,
#where x i s a v e c t o r o f counts and prob i s a v e c t o r o f p r o b a b i l i t i e s .
322
Electronic copy available at: https://ssrn.com/abstract=3863563
8.14. SIMULATING DISTRIBUTIONS FOR HYPOTHESIS TESTING
#For example the command dmultinom ( c ( 1 , 5 , 1 ) , prob=c ( . 0 2 5 , . 9 5 , . 0 2 5 ) )
# returns the p r o b a b i l i t y that Y1=1 , Y2=5 , Y3=1 ,
#when n=7 , k=3 , and p1 =.025 , p2 = . 9 5 , and p3 = . 0 2 5 ; t h i s p r o b a b i l i t y i s 0.020
#7.2 Let us confirm now
dmultinom ( c ( 1 , 5 , 1 ) , prob=c ( . 0 2 5 , . 9 5 , . 0 2 5 ) )
#NOTES:
# 1 . Close enough
# 2 . The R f u n c t i o n rmultinom (N, n , prob ) generates N v e c t o r s o f
# length k from a multinomial d i s t r i b u t i o n with n t r i a l s ,
# according t o p r o b a b i l i t i e s in the v e c t o r prob ,
#which has length k .
#The p r o b a b i l i t i e s in prob must add t o one .
#The f u n c t i o n returns a matrix with k rows and n columns .
# 3 . The rmultinom f u n c t i o n i s u s e f u l f o r simulations .
# 4 . For example , suppose that ( Y1 . . . . Yk )
# i s a multinomial with n t r i a l s and p r o b a b i l i t i e s ( p1 . . . pk ) .
# 5 . Want t o f i n d P ( Y1 > Y2 ) .
# 6 . To compute t h i s by hand ,
#need t o f i n d a l l c o n f i g u r a t i o n s o f the v e c t o r f o r which
#Y1 > Y2 and add up the p r o b a b i l i t i e s .
# 7 . Instead can write code t o simulate many multinomial random v e c t o r s
#and count the p r o p o r t i o n that have Y1 > Y2 .
# 8 . Suppose n=10 , k=4 , and p = ( . 1 , . 2 , . 3 , . 4 ) .
#The f o l l o w i n g code approximates P ( Y1 > Y2 )
n=100000
y=rmultinom ( n , 1 0 , c ( . 1 , . 2 , . 3 , . 4 ) )
sum( y [ 1 , ] > y [ 2 , ] ) / n
323
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES: Thus about 18.4% o f samples have Y1 > Y2 . See how easy i t i s !
# 8 . THE MULTIVARIATE NORMAL DISTRIBUTION
#NOTES: Some simple R code can confirm these c a l c u l a t i o n s
#and develop our i n t u i t i o n
z1=rnorm ( 1 0 0 0 , 0 , 1 )
z2=rnorm ( 1 0 0 0 , 0 , 1 )
y1=4+z1+2 * z2
y2=3 * z1+z2
cov ( y1 , y2 )
p l o t ( z1 , z2 )
p l o t ( y1 , y2 )
#8.1 Second example
#NOTES: The f o l l o w i n g p i e c e o f code w i l l
# generate 1000 p a i r s o f normal random v a r i a b l e s
#with the d e s i r e d j o i n t m u l t i v a r i a t e normal d e n s i t y
smat=matrix ( c ( 4 , 2 , 2 , 9 ) , nrow=2)
z1=rnorm ( 1 0 0 0 )
z2=rnorm ( 1 0 0 0 )
y= t ( c h o l ( smat))%*% rbind ( z1 , z2 )
y1=y [ 1 , ] + 2
y2=y [ 2 , ] + 1
p l o t ( y1 , y2 , xlab=ex pre ss ion ( y [ 1 ] ) ,
ylab=e xpr ess ion ( y [ 2 ] ) , pch =20)
324
Electronic copy available at: https://ssrn.com/abstract=3863563
8.15. INTRODUCTION TIME SERIES ANALYSIS
8.15
Introduction time series analysis
#TIMES SERIES_EXERCISE
# 1 . FORECASTING
#1.2 Amtrak − Passenger Train I l l u s t r a t i o n
Amtrak . data <− read . csv ( f i l e . choose ( ) )
r i d e r s h i p . t s <− t s ( Amtrak . data$Ridership ,
start = c (1991 ,1) ,
end = c (2004 , 3 ) ,
f r e q = 12)
p l o t ( r i d e r s h i p . ts , xlab = " Time " ,
ylab = " Ridership " ,
ylim = c (1300 , 2300) ,
bty = " l " )
#NOTES: The f i r s t l i n e reads the csv f i l e i n t o the data frame c a l l e d Amtrak . data .
#The f u n c t i o n t s c r e a t e s a time s e r i e s o b j e c t out o f
#the data frame ’ s f i r s t column Amtrak . data$Ridership .
#Give the time s e r i e s the name r i d e r s h i p . t s .
#This time s e r i e s s t a r t s in January 1991 , ends in March 2004 ,
#and has a frequency o f 12 months per year .
#By d e f i n i n g i t s frequency as 12 ,
#can l a t e r use other f u n c t i o n s t o examine i t s seasonal pattern .
#The t h i r d l i n e above produces the a c t u a l p l o t o f the time s e r i e s .
#R g i v e s users c o n t r o l over the l a b e l s , axes l i m i t s , and the p l o t ’ s border type .
#1.2 Create 2 p l o t s : the f i r s t adds a polynomial trend l i n e ,
#the second zooms−in t o 4−year period 1997−2001
i n s t a l l . packages ( " f o r e c a s t " )
library ( forecast )
r i d e r s h i p . lm <− tslm ( r i d e r s h i p . t s ~ trend + I ( trend ^ 2 ) )
par ( mfrow = c ( 2 , 1 ) )
p l o t ( r i d e r s h i p . ts ,
xlab = " Time " ,
ylab = " Ridership " ,
325
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
ylim = c (1300 , 2300) ,
bty = " l " )
lines ( ridership . lm$fitted ,
lwd = 2 )
r i d e r s h i p . t s . zoom <− window ( r i d e r s h i p . ts ,
s t a r t = c (1997 , 1 ) ,
end = c (2000 , 1 2 ) )
p l o t ( r i d e r s h i p . t s . zoom ,
xlab = " Time " ,
ylab = " Ridership " ,
ylim = c (1300 , 2300) ,
bty = " l " )
#NOTES: I n s t a l l the package f o r e c a s t
#and load the package with the f u n c t i o n l i b r a r y .
#To f i t a quadratic trend ,
#run the f o r e c a s t package ’ s l i n e a r r e g r e s s i o n model f o r time s e r i e s ,
# c a l l e d tslm . The y v a r i a b l e i s r i d e r s h i p ,
#and the two x v a r i a b l e s are trend = ( 1 , 2 , . . . , 159) and the square o f trend .
#The f u n c t i o n I t r e a t s the square o f trend " as i s " .
# P l o t the time s e r i e s
#and i t s f i t t e d r e g r e s s i o n l i n e in the top panel o f the Figure .
#The bottom panel shows a zoomed
# in view o f the time s e r i e s from January 1997 t o December 2000.
#Use the window f u n c t i o n t o i d e n t i f y a subset o f a time s e r i e s .
# 2 . SOME REPRESENTATIVE TIME SERIES EXAMPLES
#2.1 Beveridge wheat p r i c e index s e r i e s
library ( tseries )
data ( bev )
p l o t ( bev ,
xlab = " " ,
ylab = " " ,
xaxt ="n " )
x . pos<−seq ( 1 5 0 0 , 1 8 6 9 ) [ c ( seq ( 1 , length ( bev ) − 60 ,60) , length ( bev ) ) ]
x . pos<−c (1500 , 1560 , 1620 , 1680 , 1740 , 1800 , 1869)
a x i s ( 1 , x . pos , x . pos )
326
Electronic copy available at: https://ssrn.com/abstract=3863563
8.15. INTRODUCTION TIME SERIES ANALYSIS
t i t l e ( xlab =" Year " ,
ylab ="Wheat p r i c e index " ,
l i n e =3 ,
cex . a x i s = 1 . 2 ,
cex . lab = 1 . 2 )
#2.1 SP500 c l o s e p r i c e and return s e r i e s
sp500<−read . csv ( f i l e . choose ( ) )
n<−nrow ( sp500 )
x . pos<−c ( seq ( 1 , n , 8 0 0 ) , n )
p l o t ( sp500$Return ,
type =" l " ,
xlab = " " ,
ylab = " " ,
xaxt ="n " )
axis (1 ,
x . pos ,
sp500$Date [ x . pos ] ,
cex . a x i s = 1 . 5 )
t i t l e ( xlab ="Day " ,
ylab =" Daily return " ,
l i n e =3 ,
cex . lab = 1 . 2 )
#2.2 Alaska monthly temperature
ala . temp<−read . csv ( f i l e . choose ( ) , header=T )
p l o t ( ala . temp$Cel [ 6 0 1 : 7 9 2 ] ,
type =" l " ,
xlab = " " ,
ylab = " " ,
xaxt ="n " )
x . pos<−c ( 6 0 1 , 649 , 697 , 745 , 792) −600
x . l a b e l <−c ( " 0 1 / 2 0 0 1 " ,
"01/2005" ,
"01/2009" ,
327
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
"01/2013" ,
"12/2016")
axis (1 ,
x . pos ,
x . label , c
ex . a x i s = 1 . 5 )
t i t l e ( xlab ="Month " ,
ylab =" Average temperature " ,
l i n e =3 ,
cex . lab = 1 . 2 )
#2.3 Domestic s a l e s o f f o r t i f i e d wine by winemakers
wine<−read . csv ( f i l e . choose ( ) , header=F )
p l o t ( wine [ , 2 ] , type =" l " , xlab = " " , ylab = " " , xaxt ="n " )
x . pos<−c ( seq ( 1 , 118 , 2 0 ) , 118)
x . l a b e l <−wine [ x . pos , 1 ]
a x i s ( 1 , x . pos , x . l a b e l , cex . a x i s = 1 . 2 )
t i t l e ( xlab ="Month " , ylab =" Sales in thousand l i t e r s " , l i n e =3 , cex . lab = 1 . 2 )
#2.4 US population and b i r t h r a t e s
pop<−read . csv ( f i l e . choose ( ) , header=T )
x . pos<−c ( seq ( 1 , 56 , 7 ) , 56)
x . l a b e l <−c ( seq (1960 , 2009 , by = 7 ) , 2015)
par ( mfrow=c ( 2 , 1 ) , mar=c ( 3 , 4 , 3 , 4 ) )
p l o t ( pop [ , 2 ] , type =" l " , xlab = " " , ylab = " " , xaxt ="n " )
p o i n t s ( pop [ , 2 ] )
a x i s ( 1 , x . pos , x . l a b e l , cex . a x i s = 1 . 2 )
t i t l e ( xlab =" Year " , ylab =" Population " , l i n e =2 , cex . lab = 1 . 2 )
p l o t ( pop [ , 3 ] , type =" l " , xlab = " " , ylab = " " , xaxt ="n " )
p o i n t s ( pop [ , 3 ] )
a x i s ( 1 , x . pos , x . l a b e l , cex . a x i s = 1 . 2 )
t i t l e ( xlab =" Year " , ylab =" Birth r a t e s " , l i n e =2 , cex . lab = 1 . 2 )
#2.5 Beveridge wheat p r i c e time s e r i e s r e v i s i t e d
# 2 . 5 . 1 Introducing the smoothing f u n c t i o n s
library ( tseries )
data ( bev )
smooth . sym<− f u n c t i o n (my. ts , window . q ) {
328
Electronic copy available at: https://ssrn.com/abstract=3863563
8.15. INTRODUCTION TIME SERIES ANALYSIS
window . s i z e <−2*window . q+1
my. t s . sm<−rep ( 0 , length (my. t s )−window . s i z e )
f o r ( i in 1 : length (my. t s . sm ) ) {
my. t s . sm[ i ]<−mean (my. t s [ i : ( i +window . s i z e − 1 ) ] )
}
my. t s . sm
}
smooth . spencer<− f u n c t i o n (my. t s ) {
weight<−c ( − 3 , − 6 , − 5 ,3 ,21 ,46 ,67 ,74 ,67 ,46 ,21 ,3 , − 5 , − 6 , − 3)/320
my. t s . sm<−rep ( 0 , length (my. t s ) − 15)
f o r ( i in 1 : length (my. t s . sm ) ) {
my. t s . sm[ i ]<−sum(my. t s [ i : ( i + 1 4 ) ] * weight )
}
my. t s . sm
}
bev . sm<−smooth . sym ( bev , 7 )
bev . spencer<−smooth . spencer ( bev )
x . pos<−c (1500 , 1560 , 1620 , 1680 , 1740 , 1800 , 1869)
par ( mfrow=c ( 3 , 1 ) , mar=c ( 4 , 4 , 4 , 4 ) )
p l o t ( bev , type =" l " , xlab =" Year " , ylab =" Index " , xaxt ="n " )
a x i s ( 1 , x . pos , x . pos , cex . a x i s = 1 . 2 )
p l o t ( c ( 1 , length ( bev ) ) , c ( 0 , max( bev ) ) , type ="n " ,
xlab =" Year " , ylab =" Index " , xaxt ="n " )
l i n e s ( seq ( 8 , length ( bev ) − 8) , bev . sm)
a x i s ( 1 , x . pos − 1500+1, x . pos , cex . a x i s = 1 . 2 )
t i t l e ( ylab ="Smoothed p r i c e index " , l i n e =2 , cex . lab = 1 . 2 )
p l o t ( c ( 1 , length ( bev ) ) , c ( 0 , max( bev ) ) , type ="n " ,
xlab =" Year " , ylab =" Index " , xaxt ="n " )
l i n e s ( seq ( 8 , length ( bev ) − 8) , bev . spencer )
a x i s ( 1 , x . pos − 1500+1, x . pos , cex . a x i s = 1 . 2 )
t i t l e ( ylab ="Smoothed p r i c e index " , l i n e =2 , cex . lab = 1 . 2 )
# 3 .DOMESTIC SALES OF FORTIFIED WINE BY WINEMAKERS
329
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#3.1 Use the decompose ( ) f u n c t i o n in R
wine<−read . csv ( f i l e . choose ( ) , header=F )
wine . ts <− t s ( wine [ , 2 ] , frequency =4 , s t a r t =c ( 1 9 8 5 , 1 ) )
wine . de<−decompose ( wine . ts , type =" a d d i t i v e " )
p l o t ( wine . de )
# 4 . CORRELOGRAM CODE
#4.1 F i r s t example
s e t . seed ( 1 )
x<−rnorm ( 4 0 0 )
par ( mfrow=c ( 2 , 1 ) , mar=c ( 3 , 4 , 3 , 4 ) )
p l o t ( x , type =" l " , xlab = " " , ylab = " " )
t i t l e ( xlab ="Time " , ylab =" S e r i e s " , l i n e =2 , cex . lab = 1 . 2 )
a c f ( x , ylab = " " , main = " " )
t i t l e ( xlab ="Lag " , ylab ="ACF" , l i n e =2)
#4.2 Fourth Example − Non−s t a t i o n a r y s e r i e s
s e t . seed ( 1 )
t s . sim3<−cumsum( rnorm ( 4 0 0 ) )
par ( mfrow=c ( 2 , 1 ) , mar=c ( 3 , 4 , 3 , 4 ) )
p l o t ( t s . sim3 , type =" l " , xlab = " " , ylab = " " )
t i t l e ( xlab ="Time " , ylab =" S e r i e s " , l i n e =2 , cex . lab = 1 . 2 )
a c f ( t s . sim3 , ylab = " " , main = " " )
t i t l e ( xlab ="Lag " , ylab ="ACF" , l i n e =2)
#4.3 Alaska monthly temperature
#NOTES: The correlograms o f monthly o b s e r v a t i o n s on a i r temperature in Anchorage ,
#Alaska f o r the raw data ( top ) and f o r the s e a s o n a l l y adjusted data ( bottom ) .
ala . temp<−read . csv ( f i l e . choose ( ) , header=T )
ala . ts <− t s ( ala . temp [ , 2 ] , frequency =12 , s t a r t =c ( 1 9 5 1 , 2 ) )
ala . de<−decompose ( ala . t s )
ala . nosea<−as . numeric ( ala . ts −ala . de$sea )
par ( mfrow=c ( 2 , 1 ) , mar=c ( 4 , 4 , 4 , 4 ) )
a c f ( ala . temp [ , 2 ] , ylab = " " , main = " " , lag =40)
t i t l e ( ylab ="ACF" , l i n e =2)
330
Electronic copy available at: https://ssrn.com/abstract=3863563
8.15. INTRODUCTION TIME SERIES ANALYSIS
a c f ( ala . nosea , ylab = " " , main = " " , lag =40)
t i t l e ( ylab ="ACF" , l i n e =2)
dev . o f f ( )
# 5 . IBM TAQ data
ibm<−read . t a b l e ( f i l e . choose ( ) , header=T , sep ="\ t " )
ibm . new<−ibm [ , c ( 1 , 2 , 7 ) ]
ibm[ ,2] < − as . numeric ( as . c h a r a c t e r ( ibm [ , 2 ] ) )
#5.1 take 9:35:00 − 9:37:59am trading r ecord
data<−ibm . new[ 1 4 5 8 : 2 3 7 1 , ]
newtime<−rep ( 0 , nrow ( data ) )
f o r ( i in 1 : nrow ( data ) ) {
min<−as . numeric ( substr ( as . c h a r a c t e r ( data$TIME [ i ] ) , 3 , 4 ) )
sec<−as . numeric ( substr ( as . c h a r a c t e r ( data$TIME [ i ] ) , 6 , 7 ) )
newtime [ i ]<− ( min− 30) * 60+ sec
}
x . l a b e l <−c ( " 9 : 3 5 : 0 0 " , " 9 : 3 5 : 3 1 " , " 9 : 3 6 : 0 0 " , " 9 : 3 6 : 3 0 " , " 9 : 3 7 : 0 0 " , " 9 : 3 7 : 3 0 " ,
"9:38:00")
x . pos<−c ( 1 , 139 , 249 , 485 , 619 , 776 , 914)
par ( mfrow=c ( 2 , 1 ) , mar=c ( 3 , 4 , 3 , 4 ) )
p l o t ( newtime , data [ , 2 ] , xlab = " " , ylab = " " , xaxt ="n " , type ="h " )
a x i s ( 1 , newtime [ x . pos ] , x . l a b e l , cex . a x i s = 1 . 2 )
t i t l e ( xlab ="Time " , ylab =" P r i c e " , l i n e =2 , cex . lab = 1 . 2 )
p l o t ( newtime , data [ , 3 ] , xlab = " " , ylab = " " , xaxt ="n " , type ="h " )
a x i s ( 1 , newtime [ x . pos ] , x . l a b e l , cex . a x i s = 1 . 2 )
t i t l e ( xlab ="Time " , ylab ="Volume " , l i n e =2 , cex . lab = 1 . 2 )
331
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.16
Autocorrelation & Evaluating Performance
# 1 . Performance Evaluation
#1.1 load the Amtrak data
Amtrak . data <− read . csv ( f i l e . choose ( ) )
r i d e r s h i p . t s <− t s ( Amtrak . data$Ridership ,
start = c (1991 ,1) ,
end = c (2004 , 3 ) ,
f r e q = 12)
#1.2 Then need t o use tslm function , so use l i b r a r y f o r e c a s t
library ( forecast )
#1.3 Following code c r e a t e s the f o r e c a s t
nValid <− 36
nTrain <− length ( r i d e r s h i p . t s ) − nValid
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , nTrain ) )
v a l i d . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , nTrain + 1 ) ,
end = c (1991 , nTrain + nValid ) )
r i d e r s h i p . lm <− tslm ( t r a i n . t s ~ trend + I ( trend ^ 2 ) )
r i d e r s h i p . lm . pred <− f o r e c a s t ( r i d e r s h i p . lm , h = nValid , l e v e l = 0 )
p l o t ( r i d e r s h i p . lm . pred ,
ylim = c (1300 , 2600) ,
ylab = " Ridership " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " , f l t y = 2 )
axis (1 ,
332
Electronic copy available at: https://ssrn.com/abstract=3863563
8.16. AUTOCORRELATION & EVALUATING PERFORMANCE
at = seq (1991 , 2006 , 1 ) ,
l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
lines ( ridership . lm$fitted ,
lwd = 2 )
lines ( valid . ts )
#NOTES:
# 1 . Define the numbers o f months in the t r a i n i n g and v a l i d a t i o n sets ,
#nValid and nTrain , r e s p e c t i v e l y .
# 2 . Use the time and window f u n c t i o n s t o c r e a t e the t r a i n i n g
#and v a l i d a t i o n s e t s .
# F i t a model ( here a l i n e a r r e g r e s s i o n model ) t o the t r a i n i n g s e t .
#Use the f o r e c a s t f u n c t i o n t o make p r e d i c t i o n s
# o f the time s e r i e s in the v a l i d a t i o n period . P l o t the p r e d i c t i o n s .
#1.4 Computing p r e d i c t i v e measures :
accuracy ( r i d e r s h i p . lm . pred$mean , v a l i d . t s )
#1.5 Time p l o t o f f o r e c a s t e r r o r s ( or r e s i d u a l s )
#from a quadratic trend model applied t o the Amtrak r i d e r s h i p data
names ( r i d e r s h i p . lm . pred )
r i d e r s h i p . lm . p r e d $ r e s i d u a l s
v a l i d . t s − r i d e r s h i p . lm . pred$mean
p l o t ( r i d e r s h i p . lm . pred$residuals ,
ylim = c ( − 400 , 4 0 0 ) ,
ylab = " Residuals " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " , f l t y = 2 )
a x i s ( 1 , at = seq (1991 , 2006 , 1 ) ,
l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
333
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES: Use the names f u n c t i o n t o f i n d out
#which elements are included with the f o r e c a s t o b j e c t c a l l e d r i d e r s h i p . lm . pred .
#The model ’ s r e s i d u a l s are the f o r e c a s t e r r o r s in the t r a i n i n g period .
# Subtracting the model ’ s mean ( or f o r e c a s t s ) from v a l i d . t s ( or a c t u a l s )
# in the v a l i d a t i o n period g i v e s the f o r e c a s t e r r o r s in the v a l i d a t i o n period .
#1.6 Histogram o f f o r e c a s t e r r o r s
h i s t ( r i d e r s h i p . lm . pred$residuals ,
ylab = " Frequency " ,
xlab = " Forecast Error " ,
bty = " l " , main = " " )
#NOTES: Put the model ’ s r e s i d u a l s i n t o the h i s t f u n c t i o n t o p l o t the histogram ,
# or frequency o f f o r e c a s t e r r o r s in each o f seven
# ( a u t o m a t i c a l l y and e q u a l l y s i z e d ) bins .
#1.7 Repeat the code from before ,
#but change Level = 95
nValid <− 36
nTrain <− length ( r i d e r s h i p . t s ) − nValid
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , nTrain ) )
v a l i d . t s <− window ( r i d e r s h i p . ts , s t a r t = c (1991 , nTrain + 1 ) ,
end = c (1991 , nTrain + nValid ) )
r i d e r s h i p . lm <− tslm ( t r a i n . t s ~ trend + I ( trend ^ 2 ) )
r i d e r s h i p . lm . pred <− f o r e c a s t ( r i d e r s h i p . lm , h = nValid , l e v e l = 95)
p l o t ( r i d e r s h i p . lm . pred ,
ylim = c (1300 , 2600) ,
ylab = " Ridership " ,
334
Electronic copy available at: https://ssrn.com/abstract=3863563
8.16. AUTOCORRELATION & EVALUATING PERFORMANCE
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " ,
f l t y = 2)
axis (1 ,
at = seq (1991 , 2006 , 1 ) ,
l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
lines ( ridership . lm$fitted ,
lwd = 2 )
lines ( valid . ts )
# 2 . TUMBLR EXAMPLE CO
tumblr . data <− read . csv ( f i l e . choose ( ) )
#2.1 Create a time s e r i e s out o f the data
people . t s <− t s ( tumblr . data$People . Worldwide ) / 1000000
#2.2 F i t Model 1 t o the time s e r i e s
people . e t s .AAN <− e t s ( people . ts , model = "AAN" )
#2.3 F i t Model 2
people . e t s .MMN <− e t s ( people . ts , model = "MMN" , damped = FALSE)
#2.4 F i t Model 3
people . e t s .MMdN <− e t s ( people . ts , model = "MMN" , damped = TRUE)
people . e t s .AAN. pred <− f o r e c a s t ( people . e t s .AAN,
h = 115 ,
level = c (0.2 , 0.4 , 0.6 , 0.8))
people . e t s .MMN. pred <− f o r e c a s t ( people . e t s .MMN,
h = 115 ,
level = c (0.2 , 0.4 , 0.6 , 0.8))
people . e t s .MMdN. pred <− f o r e c a s t ( people . e t s .MMdN,
335
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
h = 115 ,
level = c (0.2 , 0.4 , 0.6 , 0.8))
#2.5 P l o t
par ( mfrow = c ( 1 , 3 ) ) # This command s e t s the p l o t window t o show 1 row o f 3 p l o t s .
p l o t ( people . e t s .AAN. pred ,
xlab = "Month " ,
ylab = " People ( in m i l l i o n s ) " , ylim = c ( 0 , 1 0 0 0 ) )
p l o t ( people . e t s .MMN. pred ,
xlab = "Month " ,
ylab =" People ( in m i l l i o n s ) " ,
ylim = c ( 0 , 1 0 0 0 ) )
p l o t ( people . e t s .MMdN. pred ,
xlab = "Month " ,
ylab =" People ( in m i l l i o n s ) " ,
ylim = c ( 0 , 1 0 0 0 ) )
#NOTES:
# 1 . Use the e t s f u n c t i o n in the f o r e c a s t package t o
# f i t three e xpo nen tia l smoothing models t o the data
# 2 . Use the f o r e c a s t f u n c t i o n t o generate the 20%, 40%, 60%,
#and 80% p r e d i c t i o n i n t e r v a l s with the f i t t e d model
#and the argument l e v e l = c ( 0 . 2 , 0 . 4 , 0 . 6 , 0 . 8 ) .
# 3 . Also , s e t the number o f steps −ahead ,
#1 through h = 115 , t o c r e a t e the cones .
# 3 . ADVANCED DATA PARTIONING: ROLL−FORWARD VALIDATION
#3.1 Amtrak r i d e r s h i p i l l u s t r a t i o n
f i x e d . nValid <− 36
f i x e d . nTrain <− length ( r i d e r s h i p . t s ) − f i x e d . nValid
stepsAhead <− 1
336
Electronic copy available at: https://ssrn.com/abstract=3863563
8.16. AUTOCORRELATION & EVALUATING PERFORMANCE
e r r o r <− rep ( 0 ,
f i x e d . nValid − stepsAhead + 1 )
percent . e r r o r <− rep ( 0 ,
f i x e d . nValid − stepsAhead + 1 )
f o r ( j in f i x e d . nTrain : ( f i x e d . nTrain + f i x e d . nValid − stepsAhead ) ) {
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , j ) )
v a l i d . t s <− window ( r i d e r s h i p . ts , s t a r t = c (1991 ,
j + stepsAhead ) ,
end = c (1991 , j + stepsAhead ) )
naive . pred <− naive ( t r a i n . ts , h = stepsAhead )
e r r o r [ j − f i x e d . nTrain + 1 ] <− v a l i d . t s − naive . pred$mean [ stepsAhead ]
percent . e r r o r [ j − f i x e d . nTrain + 1 ] <− e r r o r [ j −
f i x e d . nTrain + 1 ] / v a l i d . t s
}
mean ( abs ( e r r o r ) )
s q r t ( mean ( e r r o r ^ 2 ) )
mean ( abs ( percent . e r r o r ) )
#NOTES:
# 1 . Use a f o r l o o p and the window f u n c t i o n t o step through the d i f f e r e n t l y
# s i z e d t r a i n i n g and v a l i d a t i o n s e t s .
# 2 . Use each t r a i n i n g s e t t o make a one−month−ahead naive f o r e c a s t
# in the v a l i d a t i o n s e t by s e t t i n g stepsAhead t o one .
# Define two v e c t o r s o f z e r o s t o s t o r e the f o r e c a s t e r r o r s and percentage e r r o r s .
# 3 . The l a s t three l i n e s c a l c u l a t e the p r e d i c t i v e measures
# 4 . EXAMPLES: COMPARING TWO MODELS
337
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#4.1 Code f o r computing naive and seasonal naive f o r e c a s t s
#and t h e i r p r e d i c t i v e measures
f i x e d . nValid <− 36
f i x e d . nTrain <− length ( r i d e r s h i p . t s ) − f i x e d . nValid
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 ,
f i x e d . nTrain ) )
v a l i d . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 ,
f i x e d . nTrain + 1 ) ,
end = c (1991 , f i x e d . nTrain + f i x e d . nValid ) )
naive . pred <− naive ( t r a i n . ts , h = f i x e d . nValid )
snaive . pred <− snaive ( t r a i n . ts , h = f i x e d . nValid )
accuracy ( naive . pred , v a l i d . t s )
accuracy ( snaive . pred , v a l i d . t s )
#NOTES:
# 1 . Use the naive and snaive f u n c t i o n s in the f o r e c a s t package
# t o c r e a t e the naive and seasonal naive f o r e c a s t s
# in the f i x e d v a l i d a t i o n period ( as shown in s l i d e s ) .
# 2 . Use the accuracy f u n c t i o n
# t o f i n d the p r e d i c t i v e measures in Table in PowerPoint
338
Electronic copy available at: https://ssrn.com/abstract=3863563
8.17. SMOOTHING METHODS & REGRESSION - BASED TS MODELS
8.17
Smoothing methods & Regression - Based TS models
# SMOOTHING METHODS & REGRESSION−BASED TS MODELS_EXERCISE
# 1 . Smoothing methods and r e g r e s s i o n
#1.1 F i r s t , load the Amtrak data
Amtrak . data <− read . csv ( f i l e . choose ( ) )
r i d e r s h i p . t s <− t s ( Amtrak . data$Ridership ,
start = c (1991 ,1) ,
end = c (2004 , 3 ) ,
f r e q = 12)
#1.2 Now i n s t a l l and load the zoo package
i n s t a l l . packages ( " zoo " )
l i b r a r y ( zoo )
#1.3 Need t o use the f o r e c a s t package − as t h i s i s where ma( ) f u n c t i o n l i v e s !
library ( forecast )
#1.4 The f u n c t i o n ma in the f o r e c a s t package c r e a t e s a centered moving average .
ma. t r a i l i n g <− rollmean ( r i d e r s h i p . ts , k = 12 , a l i g n = " r i g h t " )
ma. centered <− ma( r i d e r s h i p . ts , order = 12)
p l o t ( r i d e r s h i p . ts ,
ylim = c (1300 , 2200) ,
ylab = " Ridership " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 4 . 2 5 ) ,
main = " " )
axis (1 ,
at = seq (1991 , 2004.25 , 1 ) ,
l a b e l s = format ( seq (1991 , 2004.25 , 1 ) ) )
l i n e s (ma. centered , lwd = 2 )
339
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
l i n e s (ma. t r a i l i n g , lwd = 2 , l t y = 2 )
legend (1994 ,2200 , c ( " Ridership " ,
" Centered Moving Average " ,
" T r a i l i n g Moving Average " ) , l t y =c ( 1 , 1 , 2 ) ,
lwd=c ( 1 , 2 , 2 ) , bty = " n " )
#NOTES:
# 1 . The arguments k and order are the r e s p e c t i v e windows
# o f these moving averages .
#In the f u n c t i o n ma when order i s an even number ,
#the centered moving average i s the average o f two asymmetric moving averages .
# 2 . When order i s odd ,
#the centered moving average i s comprised o f one symmetric moving average .
# 2 . TRAILING MA
#NOTES: Set up the t r a i n i n g and v a l i d a t i o n s e t s using f i x e d p a r t i t i o n i n g .
#Use the rollmean f u n c t i o n t o c r e a t e
#the t r a i l i n g moving average over the t r a i n i n g period .
#With the t a i l function , f i n d the l a s t moving average in the t r a i n i n g period ,
#and use i t r e p e a t e d l y as the p r e d i c t i o n f o r each month in the v a l i d a t i o n period .
# P l o t the time s e r i e s in the t r a i n i n g and v a l i d a t i o n periods ,
#the moving average over the t r a i n i n g period ,
#and the p r e d i c t i o n s f o r the v a l i d a t i o n period .
nValid <− 36
nTrain <− length ( r i d e r s h i p . t s ) − nValid
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , nTrain ) )
v a l i d . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , nTrain + 1 ) ,
end = c (1991 , nTrain + nValid ) )
340
Electronic copy available at: https://ssrn.com/abstract=3863563
8.17. SMOOTHING METHODS & REGRESSION - BASED TS MODELS
ma. t r a i l i n g <− rollmean ( t r a i n . ts ,
k = 12 ,
align = " right " )
l a s t .ma <− t a i l (ma. t r a i l i n g , 1 )
ma. t r a i l i n g . pred <− t s ( rep ( l a s t . ma, nValid ) ,
s t a r t = c (1991 , nTrain + 1 ) ,
end = c (1991 , nTrain + nValid ) ,
f r e q = 12)
p l o t ( t r a i n . ts , ylim = c (1300 , 2600) ,
ylab = " Ridership " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " )
a x i s ( 1 , at = seq (1991 , 2006 , 1 ) ,
l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
l i n e s (ma. t r a i l i n g ,
lwd = 2 ,
c o l = " blue " )
l i n e s (ma. t r a i l i n g . pred ,
lwd = 2 ,
c o l = " blue " ,
l t y = 2)
lines ( valid . ts )
# 3 . EXPONENTIAL SMOOTHING
# Set up the t r a i n i n g and v a l i d a t i o n s e t s using f i x e d p a r t i t i o n i n g .
341
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#Use the e t s f u n c t i o n with the o p t i o n s model = "ANN"
#and alpha = 0 . 2 t o f i t a simple e xpo ne nti al smoothing model
# over the t r a i n i n g period . Use the f o r e c a s t f u n c t i o n t o make p r e d i c t i o n s
#using t h i s ses model f o r h = nValid steps −ahead in the v a l i d a t i o n period .
# Print these p r e d i c t i o n s in tabular form and then p l o t them .
d i f f . twice . t s <− d i f f ( d i f f ( r i d e r s h i p . ts , lag = 1 2 ) , lag = 1 )
nValid <− 36
nTrain <− length ( d i f f . twice . t s ) − nValid
t r a i n . t s <− window ( d i f f . twice . ts ,
s t a r t = c (1992 , 2 ) ,
end = c (1992 ,
nTrain + 1 ) )
v a l i d . t s <− window ( d i f f . twice . ts ,
s t a r t = c (1992 , nTrain + 2 ) ,
end = c (1992 ,
nTrain + 1 + nValid ) )
ses <− e t s ( t r a i n . ts , model = "ANN" ,
alpha = 0 . 2 )
ses . pred <− f o r e c a s t ( ses ,
h = nValid ,
l e v e l = 0)
ses . pred
p l o t ( ses . pred ,
ylim = c ( − 250 , 3 0 0 ) ,
ylab = " Ridership ( Twice−D i f f e r e n c e d ) " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " , f l t y = 2 )
342
Electronic copy available at: https://ssrn.com/abstract=3863563
8.17. SMOOTHING METHODS & REGRESSION - BASED TS MODELS
axis (1 ,
at = seq (1991 , 2006 , 1 ) ,
l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
l i n e s ( ses . p r e d $ f i t t e d ,
lwd = 2 ,
c o l = " blue " )
lines ( valid . ts )
#3.1 Code f o r Table 1
ses . opt <− e t s ( t r a i n . ts , model = "ANN" )
ses . opt . pred <− f o r e c a s t ( ses . opt , h = nValid , l e v e l = 0 )
accuracy ( ses . pred , v a l i d . t s )
accuracy ( ses . opt . pred , v a l i d . t s )
ses . opt
#NOTES: Use the e t s f u n c t i o n with the o p t i o n model = "ANN"
# t o f i t a simple e xpo nen tia l smoothing model where a i s chosen o p t i m a l l y .
#Use the f o r e c a s t f u n c t i o n t o make p r e d i c t i o n s using t h i s ses . opt model
# f o r h = nValid steps −ahead in the v a l i d a t i o n period .
#Compare the two models ’ performance metrics
#using the accuracy f u n c t i o n in the f o r e c a s t package .
#The l a s t l i n e p rovid es a summary o f the ses . opt model .
# 4 . SERIES WITH TREND + SEASONALITY_HOLT−WINTER’ S APPLIED TO AMTRAK ILLUSTRATION
#NOTES:
# 1 . Use the e t s f u n c t i o n with the o p t i o n model = "MAA"
# t o f i t Holt−Winter ’ s e xpo nen tia l smoothing model with m u l t i p l i c a t i v e error ,
# a d d i t i v e trend , and a d d i t i v e s e a s o n a l i t y .
# 2 . Use the f o r e c a s t f u n c t i o n t o make p r e d i c t i o n s using t h i s hwin model
# f o r h = nValid stepsahead in the v a l i d a t i o n period .
# P l o t these p r e d i c t i o n s and p r i n t out a summary o f the model .
#4.1 F i r s t , I had t o re−run these 4 l i n e s o f code again , t o avoid negative value
343
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
nValid <− 36
nTrain <− length ( r i d e r s h i p . t s ) − nValid
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , nTrain ) )
v a l i d . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , nTrain + 1 ) ,
end = c (1991 , nTrain + nValid ) )
hwin <− e t s ( t r a i n . ts , model = "MAA" )
hwin . pred <− f o r e c a s t ( hwin , h = nValid , l e v e l = 0 )
p l o t ( hwin . pred ,
ylim = c (1300 , 2600) ,
ylab = " Ridership " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " , f l t y = 2 )
a x i s ( 1 , at = seq (1991 , 2006 , 1 ) ,
l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
l i n e s ( hwin . p r e d $ f i t t e d , lwd = 2 , c o l = " blue " )
lines ( valid . ts )
#4.2 Finally , output hwin t o get Table 2
hwin
#4.3 And f o r i n i t i a l s t a t e s
hwin$states [ 1 , ]
344
Electronic copy available at: https://ssrn.com/abstract=3863563
8.17. SMOOTHING METHODS & REGRESSION - BASED TS MODELS
#4.4 Then f i n a l s t a t e s
hwin$states [ nrow ( hwin$states ) , ]
#4.5 Table 4 R code
e t s ( t r a i n . ts , r e s t r i c t = FALSE, allow . m u l t i p l i c a t i v e . trend = TRUE)
#4.6 Table 5 R code
accuracy ( hwin . pred , v a l i d . t s )
# 5 . EXTENSIONS OF EXPONENTIAL SMOOTHING
#5.1 Bike Share i l l u s t r a t i o n
# 5 . 1 . 1 HOURLY DATA
bike . hourly . df <− read . csv ( f i l e . choose ( ) )
nTotal <− length ( bike . hourly . df$cnt [ 1 3 0 0 4 : 1 3 7 4 7 ] )
# 31 days * 24 hours / day = 744 hours
bike . hourly . msts <− msts ( bike . hourly . df$cnt [ 1 3 0 0 4 : 1 3 7 4 7 ] ,
seasonal . p e r i o d s = c ( 2 4 , 1 6 8 ) ,
start = c (0 , 1))
nTrain <− 21 * 24 # 21 days o f hourly data
nValid <− nTotal − nTrain # 10 days o f hourly data
yTrain . msts <− window ( bike . hourly . msts , s t a r t = c ( 0 , 1 ) ,
end = c ( 0 , nTrain ) )
yValid . msts <− window ( bike . hourly . msts ,
s t a r t = c ( 0 , nTrain + 1 ) ,
end = c ( 0 , nTotal ) )
bike . hourly . dshw . pred <− dshw ( yTrain . msts , h = nValid )
bike . hourly . dshw . pred . mean <− msts ( bike . hourly . dshw . pred$mean ,
345
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
seasonal . p e r i o d s = c ( 2 4 , 1 6 8 ) ,
s t a r t = c ( 0 , nTrain + 1 ) )
accuracy ( bike . hourly . dshw . pred . mean , yValid . msts )
p l o t ( yTrain . msts ,
xlim = c ( 0 , 4 + 3 / 7 ) ,
xlab = "Week" ,
ylab = " Hourly Bike Rentals " )
l i n e s ( bike . hourly . dshw . pred . mean , lwd = 2 , c o l = " blue " )
#NOTES:
# 1 . Use the msts f u n c t i o n t o s e t up the time s e r i e s with two seasonal p e r i o d s :
#hours per day and hours per week .
#The two p e r i o d s must be nested .
#We s t a r t the time s e r i e s with mul tiple s e a s o n a l i t i e s at week 0 and hour 1 .
#Set up your t r a i n i n g and v a l i d a t i o n s e t s .
# 2 . Use the dshw f u n c t i o n t o make f i t and make f o r e c a s t s with the model .
# 3 . Restart the f o r e c a s t from dshw at the c o r r e c t s t a r t time .
# ( In the f o r e c a s t package v e r s i o n 7 . 1 ,
# there i s a minor e r r o r in the s t a r t time o f a f o r e c a s t from the f u n c t i o n dshw . )
# 4 . P l o t the f o r e c a s t s from dshw .
#5.2 2nd Bike Sharing I l l u s t r a t i o n
# 5 . 2 . 1 DAILY DATA
bike . d a i l y . df <− read . csv ( f i l e . choose ( ) )
bike . d a i l y . msts <− msts ( bike . d a i l y . df$cnt , seasonal . p e r i o d s = c ( 7 , 3 6 5 . 2 5 ) )
bike . d a i l y . t b a t s <− t b a t s ( bike . d a i l y . msts )
bike . d a i l y . t b a t s . pred <− f o r e c a s t ( bike . d a i l y . tbats , h = 365)
bike . d a i l y . stlm <− stlm ( bike . d a i l y . msts ,
346
Electronic copy available at: https://ssrn.com/abstract=3863563
8.17. SMOOTHING METHODS & REGRESSION - BASED TS MODELS
s . window = " p e r i o d i c " ,
method = " e t s " )
bike . d a i l y . stlm . pred <− f o r e c a s t ( bike . d a i l y . stlm , h = 365)
par ( mfrow = c ( 1 , 2 ) )
p l o t ( bike . d a i l y . t b a t s . pred , ylim = c ( 0 , 9000) ,
xlab = " Year " ,
ylab = " Daily Bike Rentals " ,
main = "TBATS" )
p l o t ( bike . d a i l y . stlm . pred ,
ylim = c ( 0 , 9000) ,
xlab = " Year " ,
ylab = " Daily Bike Rentals " ,
main = "STL + ETS " )
#NOTES:
# 1 . Use the msts f u n c t i o n t o s e t up the time s e r i e s with two seasonal p e r i o d s :
#days per week and days per year .
# 2 . Use the t b a t s and stlm t o apply TBATS and STL+ex pon ent ial smoothing (ETS) ,
#respectively .
# 3 . Use the f o r e c a s t f u n c t i o n t o make f o r e c a s t f o r the next 365 days .
# 4 . P l o t the f o r e c a s t s .
# 6 . REGRESSION−BASED MODELS
#6.1 LINEAR TREND
t r a i n . lm <− tslm ( t r a i n . t s ~ trend )
p l o t ( t r a i n . ts , xlab = " Time " ,
ylab = " Ridership " ,
ylim = c (1300 , 2300) ,
bty = " l " )
l i n e s ( t r a i n . l m $ f i t t e d , lwd = 2 )
#NOTES:
347
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 1 . Use the tslm f u n c t i o n ( which stands f o r time s e r i e s l i n e a r model )
#with the formula t r a i n . t s ~ trend t o produce a l i n e a r trend model .
# P l o t the time s e r i e s and the l i n e a r trend model ’ s values .
# 2 . Linear trend f i t t e d t o the Amtrak r i d e r s h i p data in the t r a i n i n g period
#and f o r e c a s t e d in the v a l i d a t i o n period
t r a i n . lm . pred <− f o r e c a s t ( t r a i n . lm , h = nValid , l e v e l = 0 )
p l o t ( t r a i n . lm . pred , ylim = c (1300 , 2600) , ylab = " Ridership " , xlab = " Time " ,
bty = " l " , xaxt = " n " , xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) , main = " " , f l t y = 2 )
a x i s ( 1 , at = seq (1991 , 2006 , 1 ) , l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
l i n e s ( t r a i n . lm . p r e d $ f i t t e d , lwd = 2 , c o l = " blue " )
lines ( valid . ts )
#6.2 Show the r e g r e s s i o n output
summary ( t r a i n . lm )
# 7 . EXPONENTIAL TREND
#NOTES: shows the p r e d i c t i o n s from a l i n e a r r e g r e s s i o n model
# f i t t o l o g ( Ridership ) .
#The l i n e a r trend model i s o v e r l a i d on top o f t h i s exp one nti al trend model .
# The two trends do not appear t o be very d i f f e r e n t
#because the time horizon i s short and the i n c r e a s e i s g e n t l e
t r a i n . lm . expo . trend <− tslm ( t r a i n . t s ~ trend , lambda = 0 )
t r a i n . lm . expo . trend . pred <− f o r e c a s t ( t r a i n . lm . expo . trend ,
h = nValid ,
l e v e l = 0)
t r a i n . lm . l i n e a r . trend <− tslm ( t r a i n . t s ~ trend , lambda = 1 )
t r a i n . lm . l i n e a r . trend . pred <− f o r e c a s t ( t r a i n . lm . l i n e a r . trend ,
h = nValid ,
l e v e l = 0)
p l o t ( t r a i n . lm . expo . trend . pred , ylim = c (1300 , 2600) ,
ylab = " Ridership " ,
xlab = " Time " ,
bty = " l " ,
xaxt = " n " ,
348
Electronic copy available at: https://ssrn.com/abstract=3863563
8.17. SMOOTHING METHODS & REGRESSION - BASED TS MODELS
xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) ,
main = " " , f l t y = 2 )
a x i s ( 1 , at = seq (1991 , 2006 , 1 ) , l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
l i n e s ( t r a i n . lm . expo . trend . p r e d $ f i t t e d , lwd = 2 , c o l = " blue " )
l i n e s ( t r a i n . lm . l i n e a r . trend . p r e d $ f i t t e d , lwd = 2 , c o l = " black " , l t y = 3 )
l i n e s ( t r a i n . lm . l i n e a r . trend . pred$mean , lwd = 2 , c o l = " black " , l t y = 3 )
lines ( valid . ts )
#NOTES: Use the tslm f u n c t i o n again , but with the argument lambda = 0 .
#With t h i s argument included ,
#the tslm f u n c t i o n a p p l i e s the Box−Cox transformation
# t o the r i d e r s h i p data in the t r a i n i n g period .
# 8 . POLYNOMIAL TREND
#8.1 Summary o f output from f i t t i n g a quadratic trend
# t o the Amtrak r i d e r s h i p data in the t r a i n i n g period
t r a i n . lm . poly . trend <− tslm ( t r a i n . t s ~ trend + I ( trend ^ 2 ) )
summary ( t r a i n . lm . poly . trend )
#NOTES:
# 1 . F i t the quadratic trend with tslm ( t r a i n . t s ~ trend + I ( trend ^ 2 ) .
# 2 . The f u n c t i o n I t r e a t s an o b j e c t " as i s " .
# 3 . The quadratic equation f i t by the tslm f u n c t i o n i s
#1888.88401 − 6.29780 t + 0.05362 t2 where t ( or trend )
# goes from 1 t o 123 in the t r a i n i n g period and from 124 t o 159
# in the v a l i d a t i o n period
#8.2 Code f o r t a b l e :
#Summary o f output from f i t t i n g a d d i t i v e s e a s o n a l i t y
# t o the Amtrak r i d e r s h i p data in the t r a i n i n g period
t r a i n . lm . season <− tslm ( t r a i n . t s ~ season )
summary ( t r a i n . lm . season )
349
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES: Include season as a p r e d i c t o r in the tslm f u n c t i o n .
#This p r e d i c t o r i s a c t u a l l y 11 dummy v a r i a b l e s ,
#one f o r each season−except f o r the f i r s t season , January .
# I f users ever f o r g e t which season i s season1 ,
#run the l i n e t r a i n . lm . season$data t o see which season number
# goes with which month ( or day or hour )
#8.3 Model with Trend + S e a s o n a l i t y
t r a i n . lm . trend . season <− tslm ( t r a i n . t s ~ trend + I ( trend ^2) + season )
summary ( t r a i n . lm . trend . season )
#8.4 Model with l i n e a r , quadratic , sine ,
#and c o s i n e terms f i t t o the Amtrak r i d e r s h i p data in the t r a i n i n g period .
tslm ( t r a i n . t s ~ trend + I ( trend ^2) + I ( s i n ( 2 * p i * trend / 1 2 ) )
+ I ( cos ( 2 * p i * trend / 1 2 ) ) )
350
Electronic copy available at: https://ssrn.com/abstract=3863563
8.18. AUTOREGRESSIVE (AR) & ARIMA MODELS, FORECASTING BINARY OUTCOMES
8.18
Autoregressive (AR) & ARIMA models, forecasting binary
outcomes
#AUTOREGRESSION (AR) ARIMA MODELS, FORECASTING BINARY OUTCOMES_EXERCISE
# load the Amtrak data
Amtrak . data <− read . csv ( f i l e . choose ( ) )
r i d e r s h i p . t s <− t s ( Amtrak . data$Ridership ,
start = c (1991 ,1) ,
end = c (2004 , 3 ) ,
f r e q = 12)
#Load the f o r e c a s t l i b r a r y f o r the a c f f u n c t i o n
library ( forecast )
#Create the a u t o c o r r e l a t i o n l a g s
r i d e r s h i p . 2 4 . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , 2 4 ) )
Acf ( r i d e r s h i p . 2 4 . ts , lag . max = 12 , main = " " )
# 1 . AUTOREGRESSIVE MODELS
#1.1 Create t r a i n . t s f o r Training Period
nValid <− 36
nTrain <− length ( r i d e r s h i p . t s ) − nValid
t r a i n . t s <− window ( r i d e r s h i p . ts ,
s t a r t = c (1991 , 1 ) ,
end = c (1991 , nTrain ) )
#1.2 F i t t i n g an AR1 model
t r a i n . lm . trend . season <− tslm ( t r a i n . t s ~ trend + I ( trend ^2) + season )
t r a i n . res . arima <− Arima ( t r a i n . lm . trend . season$residuals , order = c ( 1 , 0 , 0 ) )
t r a i n . res . arima . pred <− f o r e c a s t ( t r a i n . res . arima , h = nValid )
p l o t ( t r a i n . lm . trend . season$residuals , ylim = c ( − 250 , 2 5 0 ) , ylab = " Residuals " ,
xlab = " Time " , bty = " l " , xaxt = " n " , xlim = c ( 1 9 9 1 , 2 0 0 6 . 2 5 ) , main = " " )
351
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
a x i s ( 1 , at = seq (1991 , 2006 , 1 ) , l a b e l s = format ( seq (1991 , 2006 , 1 ) ) )
l i n e s ( t r a i n . res . arima . p r e d $ f i t t e d , lwd = 2 , c o l = " blue " )
#1.3 Use summary f o r the Training s e t e r r o r measures s t a t i s t i c a l output
summary ( t r a i n . res . arima )
# 2 . CORRELATED EXTERNAL SERIES
#2.1 Bike Sharing I l l u s t r a t i o n
#NOTES:
# 1 . I n s t a l l and load the l u b r i d a t e package .
#Use i t s month and wday f u n c t i o n s t o c r e a t e named Month and DOW v a r i a b l e s .
#Create named f a c t o r v a r i a b l e s WorkingDay and Weather .
# 2 . Use the f u n c t i o n model . matrix t o c r e a t e three s e t s o f dummy v a r i a b l e s :
#12 from Month , 7 from DOW,
#and 6 from the i n t e r a c t i o n between WorkingDay and Weather .
#2.1 I n s t a l l and load the package
i n s t a l l . packages ( " l u b r i d a t e " )
library ( lubridate )
#2.2 Load the data
bike . df <− read . csv ( f i l e . choose ( ) )
bike . df$Date <− as . Date ( bike . df$dteday , format = "%Y−%m−%d " )
bike . df$Month <− month ( bike . df$Date , l a b e l = TRUE)
bike . df$DOW <− wday ( bike . df$Date , l a b e l = TRUE)
bike . df$WorkingDay <− f a c t o r ( bike . df$workingday , l e v e l s = c ( 0 , 1 ) ,
l a b e l s = c ( " Not_Working " , " Working " ) )
bike . df$Weather <− f a c t o r ( bike . df$weathersit , l e v e l s = c ( 1 , 2 , 3 ) ,
l a b e l s = c ( " Clear " , " Mist " , " Rain_Snow " ) )
Month . dummies <− model . matrix (~ 0 + Month , data = bike . df )
DOW. dummies <− model . matrix (~ 0 + DOW, data = bike . df )
WorkingDay_Weather . dummies <− model . matrix (~ 0 +
WorkingDay : Weather , data = bike . df )
colnames ( Month . dummies ) <− gsub ( " Month " , " " , colnames ( Month . dummies ) )
352
Electronic copy available at: https://ssrn.com/abstract=3863563
8.18. AUTOREGRESSIVE (AR) & ARIMA MODELS, FORECASTING BINARY OUTCOMES
colnames (DOW. dummies ) <− gsub ( "DOW" , " " , colnames (DOW. dummies ) )
colnames ( WorkingDay_Weather . dummies ) <− gsub ( " WorkingDay " , " " ,
colnames ( WorkingDay_Weather . dummies ) )
colnames ( WorkingDay_Weather . dummies ) <− gsub ( " Weather " , " " ,
colnames ( WorkingDay_Weather . dummies ) )
colnames ( WorkingDay_Weather . dummies)<− gsub ( " : " ,
"_" ,
colnames ( WorkingDay_Weather . dummies ) )
x <− as . data . frame ( cbind ( Month . dummies [ , − 12] ,
DOW. dummies [ , − 7] ,
WorkingDay_Weather . dummies [ , − 6]))
y <− bike . df$cnt
nTotal <− length ( y )
nValid <− 90
nTrain <− nTotal − nValid
xTrain <− x [ 1 : nTrain , ]
yTrain <− y [ 1 : nTrain ]
xValid <− x [ ( nTrain + 1 ) : nTotal , ]
yValid <− y [ ( nTrain + 1 ) : nTotal ]
yTrain . t s <− t s ( yTrain )
( formula <− as . formula ( paste ( " yTrain . t s " , paste ( c ( " trend " ,
colnames ( xTrain ) ) ,
collapse = "+") ,
sep = " ~ " ) ) )
bike . tslm <− tslm ( formula , data = xTrain , lambda = 1 )
bike . tslm . pred <− f o r e c a s t ( bike . tslm , newdata = xValid )
p l o t ( bike . tslm . pred ,
ylim = c ( 0 , 9000) ,
xlab = " Days " ,
ylab = " Daily Bike Rentals " )
#NOTES:
# 1 . Use the f u n c t i o n gsub t o s u b s t i t u t e blanks
# in p l a c e o f elements in the v a r i a b l e s names .
# 2 . This s u b s t i t u t i o n w i l l make your output more readable .
# 3 . Set up t r a i n i n g and v a l i d a t i o n s e t s ( f o r a p r e d i c t i v e goal ) ,
353
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#dropping one dummy v a r i a b l e per s e t .
# 4 . Create a formula from your v a r i a b l e s in the t r a i n i n g set ,
#adding the v a r i a b l e trend in the p r o c e s s .
# 5 . F i t the tslm model t o the t r a i n i n g s e t .
#2.3 Output from the estimated model
o p t i o n s ( scipen = 999 , d i g i t s = 6 )
summary ( bike . tslm )
# 3 . EXAMPLE 4 : WALMART SALES BY STORE AND BY DEPARTMENT
one . p a i r <− read . csv ( f i l e . choose ( ) )
nTrain <− 143
yTrain . t s <− t s ( one . pair$Weekly_Sales [ 1 : nTrain ] , f r e q = 52 , s t a r t = c (2011 , 5 ) )
s t l . run <− s t l ( yTrain . ts , s . window = " p e r i o d i c " )
p l o t ( s t l . run )
#NOTES:
# 1 . Set up the data the from t r a i n i n g s e t as a time s e r i e s
#with weekly s e a s o n a l i t y ( f r e q = 52) s t a r t i n g from the f i f t h week in 2011.
# 2 . Use s t l t o decompose the time s e r i e s i n t o three components :
# seasonal , trend , and e r r o r ( " remainder " in R ) .
# 3 . Set the argument s . window t o " p e r i o d i c " ,
#which means the seasonal component w i l l be i d e n t i c a l over the years .
# ( To apply s t l , you must have at l e a s t two years
# in the t r a i n i n g s e t or 104 weeks in t h i s case . )
# 4 . P l o t the data ( or time s e r i e s ) and components .
#3.1 Summary output from f i t t i n g stlm
xTrain <− data . frame ( IsHoliday = one . pair$IsHoliday [ 1 : nTrain ] )
stlm . reg . f i t <− stlm ( yTrain . ts , s . window = " p e r i o d i c " , xreg = xTrain ,
method = " arima " )
stlm . reg . f i t $ m o d e l
#NOTES: Problem with data ! ! xreg= xTrain NOT WORKING? NA values creeping in ? ?
#3.2 Equivalent A l t e r n a t i v e Approach
354
Electronic copy available at: https://ssrn.com/abstract=3863563
8.18. AUTOREGRESSIVE (AR) & ARIMA MODELS, FORECASTING BINARY OUTCOMES
seasonal . comp <− s t l . run$time . s e r i e s [ , 1 ]
deasonalized . t s <− yTrain . t s − seasonal . comp
seasadj ( s t l . run ) # A l t e r n a t i v e l y , t h i s l i n e returns the d e s e a s o n l i z e d time s e r i e s .
arima . f i t . deas <− auto . arima ( deasonalized . ts , xreg = xTrain )
#NOTES: Problem with data ! ! xreg= xTrain NOT WORKING? NA values creeping in ? ?
arima . f i t . deas . pred <− f o r e c a s t ( arima . f i t . deas , xreg = xTest , h = nTest )
#NOTES: Problem with data ! ! o b j e c t ’ arima . f i t . deas ’ not found .
seasonal . comp . pred <− snaive ( seasonal . comp , h = nTest )
#NOTES: Problem with data ! ! o b j e c t ’ nTest ’ not found .
a l t . f o r e c a s t <− arima . f i t . deas . pred$mean + seasonal . comp . pred$mean
#NOTES: Problem with data ! ! o b j e c t ’ arima . f i t . deas . pred ’ not found .
stlm . reg . pred$mean − a l t . f o r e c a s t
#NOTES: Problem with data ! ! o b j e c t ’ stlm . reg . pred ’ not found .
#NOTES: Same problem with xreg
# Show the STL + ARIMA model that the stlm model f i t s .
#The stlm model produces f o r e c a s t s that are the same as those
# in a l t . f o r e c a s t from an arima model f i t t o the d e s e a s o n a l i z e time s e r i e s .
#3.3 Final Walmart Sales f i g u r e
xTrain <− data . frame ( IsHoliday = one . pair$IsHoliday [ 1 : nTrain ] )
nTest <− 39
xTest <− data . frame ( IsHoliday = one . pair$IsHoliday [ ( nTrain + 1 ) : ( nTrain + nTest ) ] )
stlm . reg . f i t <− stlm ( yTrain . ts ,
s . window = " p e r i o d i c " ,
xreg = xTrain ,
method = " arima " )
#NOTES: Problem with data ! !
stlm . reg . pred <− f o r e c a s t ( stlm . reg . f i t , xreg = xTest , h = nTest )
#NOTES: Problem with data ! ! o b j e c t ’ stlm . reg . f i t ’ not found .
p l o t ( stlm . reg . pred , xlab = " Year " , ylab = " Weekly Sales " )
#NOTES: Problem with data ! ! o b j e c t ’ stlm . reg . pred ’ not found .
#NOTES:
# 1 . Put the e x t e r n a l information f o r the t r a i n i n g
#and t e s t i n g p e r i o d s i n t o data frames .
# 2 . The e x t e r n a l information i n c l u d e s only one v a r i a b l e ( IsHoliday ) in t h i s case ,
#but i t can i n c l u d e more .
355
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 3 . Use stlm t o f i t an arima model
#with e x t e r n a l information t o the " data minus seasonal "
#from the s t l decomposition Figure .
# 4 . Set s . window t o " p e r i o d i c " , xreg t o xTrain ,
#and method t o " arima " .
#Use the f o r e c a s t f u n c t i o n t o make a f o r e c a s t from the stlm model .
# 5 . Finally , p l o t the f o r e c a s t s .
# 4 . FORECASTING BINARY OUTCOMES_EXAMPLE: RAINFALL IN MELBOURNE AUSTRALIA
#4.1 I n s t a l l and load the c a r e t package − t h i s w i l l
i n s t a l l . packages ( " c a r e t " )
library ( lattice )
l i b r a r y ( ggplot2 )
library ( caret )
rain . df <− read . csv ( f i l e . choose ( ) )
rain . df$Date <− as . Date ( rain . df$Date , format="%m/%d/%Y " )
rain . df$Rainy <− i f e l s e ( rain . d f $ R a i n f a l l > 0 , 1 , 0 )
nPeriods <− length ( rain . df$Rainy )
rain . df$Lag1 <− c (NA, rain . d f $ R a i n f a l l [ 1 : ( nPeriods − 1 ) ] )
rain . d f $ t <− seq ( 1 , nPeriods , 1 )
rain . df$Seasonal_sine = s i n ( 2 * p i * rain . d f $ t / 3 6 5 . 2 5 )
rain . df$Seasonal_cosine = cos ( 2 * p i * rain . d f $ t / 3 6 5 . 2 5 )
t r a i n . df <− rain . df [ rain . df$Date <= as . Date ( " 1 2 / 3 1 / 2 0 0 9 " , format="%m/%d/%Y " ) , ]
t r a i n . df <− t r a i n . df [ − 1 ,]
v a l i d . df <− rain . df [ rain . df$Date > as . Date ( " 1 2 / 3 1 / 2 0 0 9 " , format="%m/%d/%Y " ) , ]
x v a l i d <− v a l i d . df [ , c ( 4 , 6 , 7 ) ]
rainy . l r <− glm ( Rainy ~ Lag1 + Seasonal_sine + Seasonal_cosine ,
data = t r a i n . df ,
family = " binomial " )
#NOTES:
# 1 . Problem with mustart in glm ( ) function , but data i s complete ,
#no NA, and a l l values are numeric !
356
Electronic copy available at: https://ssrn.com/abstract=3863563
8.18. AUTOREGRESSIVE (AR) & ARIMA MODELS, FORECASTING BINARY OUTCOMES
# 2 . CODE IS TAKEN DIRECTLY FROM TEXT!
# 3 . Summary( t r a i n . df ) shows NA values .
summary ( rainy . l r )
#NOTES: Problem with data ! ! o b j e c t ’ rainy . l r ’ not found .
rainy . l r . pred <− p r e d i c t ( rainy . l r , xvalid , type = " response " )
#NOTES: Problem with data ! ! o b j e c t ’ rainy . l r ’ not found .
confusionMatrix ( i f e l s e ( rainy . l r $ f i t t e d > 0 . 5 , 1 , 0 ) , t r a i n . df$Rainy )
#NOTES: Problem with data ! ! o b j e c t ’ rainy . l r ’ not found .
confusionMatrix ( i f e l s e ( rainy . l r . pred > 0 . 5 , 1 , 0 ) , v a l i d . df$Rainy )
#NOTES: o b j e c t ’ rainy . l r . pred ’ not found .
357
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.19
Panel data regression analysis
#PANEL DATA REGRESSION ANALYSIS_EXERCISE
# 1 . Example 1 : i n d i v i d u a l h e t e r o g e n e i t y − F a t a l i t i e s data s e t
#NOTES:
# 1 . The F a t a l i t i e s dataset from Stock and Watson ( 2 0 0 7 )
# i s a good example o f the importance o f i n d i v i d u a l h e t e r o g e n e i t y
#and time e f f e c t s in a panel s e t t i n g .
# 2 . The research question i s whether taxing a l c o h o l i c s
#can reduce the road ’ s death t o l l .
# 1 . 1 . Load data
data ( " F a t a l i t i e s " , package ="AER" )
F a t a l i t i e s $ f r a t e <− with ( F a t a l i t i e s , f a t a l / pop * 10000)
fm <− f r a t e ~ beertax
#NOTES:
# 1 . The most b a s i c step
# i s a cross −s e c t i o n a l a n a l y s i s f o r one s i n g l e year ( here , 1 9 8 2 ) .
#One proceeds f i r s t c r e a t i n g a model o b j e c t through a c a l l t o lm ,
#then d i s p l a y i n g a summary . lm o f i t .
# 1 . 2 . P r i n t i n g t o screen occur s when i n t e r a c t i v e l y c a l l i n g an o b j e c t by name .
# Notice that s u b s e t t i n g can be done i n s i d e the c a l l t o lm
#by f e e d i n g an e xp res sio n that s o l v e s i n t o a l o g i c a l v e c t o r
# t o the subset argument :
#data p o i n t s corresponding t o TRUEs w i l l be s e l e c t e d , FALSEs discarded .
mod82 <− lm ( fm , F a t a l i t i e s ,
subset = year == 1982)
summary ( mod82 )
#1.3 The beer tax turns out s t a t i s t i c a l l y i n s i g n i f i c a n t .
#Turning t o the l a s t year in the sample ( and employing c o e f t e s t f o r compactness )
mod88 <− update ( mod82 , subset = year == 1988)
l i b r a r y ( " zoo " )
358
Electronic copy available at: https://ssrn.com/abstract=3863563
8.19. PANEL DATA REGRESSION ANALYSIS
l i b r a r y ( " lmtest " )
c o e f t e s t ( mod88 )
#NOTES:
# 1 . The c o e f f i c i e n t i s s i g n i f i c a n t and p o s i t i v e !
# Similar r e s u l t s appear f o r any s i n g l e year in the sample .
# 2 . Pooling a l l cross −s e c t i o n s together ,
#without c o n s i d e r i n g any form o f i n d i v i d u a l e f f e c t ,
#can be done using the reg ular lm f u n c t i o n or , e q u i v a l e n t l y , plm ;
# in t h i s second case , f o r reasons which w i l l be c l e a r e r in the f o l l o w i n g ,
# t h i s i s not the d e f a u l t behavior ,
# so the o p t i o n a l model argument has t o be s p e c i f i e d , s e t t i n g i t t o ’ pooling ’ .
#1.4 F i r s t we need the " plm " package , so please i n s t a l l
# i n s t a l l . packages ( " plm " ) −
i n s t a l l . packages ( " plm " )
l i b r a r y ( plm )
#1.5 Drawing on t h i s much enlarged dataset does not change the q u a l i t a t i v e r e s u l t
poolmod <− plm ( fm , F a t a l i t i e s , model =" p o o l i n g " )
c o e f t e s t ( poolmod )
#NOTES: Taxing beer would seem t o i n c r e a s e the number o f deaths
#from road a c c i d e n t s so that ,
# extending t h i s l i n e o f reasoning f a r beyond what the given evidence supports ,
# i . e . , f a r o u t s i d e the given sample ,
#one could even argue that f r e e beer might lead t o s a f e r d r i v i n g .
# Similar r e s u l t s , c o n t r a d i c t i n g the most b a s i c i n t u i t i o n ,
#appear f o r any s i n g l e year in the sample .
#1.6 Next , the si mplest way
# t o get r i d o f the i n d i v i d u a l i n t e r c e p t s i s t o estimate the model in d i f f e r e n c e s .
dmod <− plm ( d i f f ( f r a t e , 5 ) ~ d i f f ( beertax , 5 ) ,
Fatalities ,
model =" p o o l i n g " )
c o e f ( dmod )
359
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES: Estimation on f i v e −year d i f f e r e n c e s f i n a l l y y i e l d s a s e n s i b l e r e s u l t :
# a f t e r c o n t r o l l i n g f o r s t a t e heterogeneity ,
# higher t a x a t i o n on beer i s a s s o c i a t e d with a lower number o f f a t a l i t i e s .
#1.7 Another way t o c o n t r o l f o r time−i n v a r i a n t un−o b s e r v a b l e s
# i s t o estimate them out e x p l i c i t l y .
#Separate i n t e r c e p t s could be e a s i l y added in p l a i n R using the formula syntax :
l s d v . fm <− update ( fm , . ~ . + s t a t e − 1 )
lsdvmod <− lm ( l s d v . fm , F a t a l i t i e s )
c o e f ( lsdvmod ) [ 1 ]
#NOTES: The estimate i s numerically d i f f e r e n t but supports the same
# qualitative conclusions .
#1.8 Fixed e f f e c t s ( within ) estimation y i e l d s an e q u i v a l e n t
# r e s u l t in a more compact and e f f i c i e n t way .
# S p e c i f y i n g model = ’ within ’ in the c a l l t o plm i s not necessary
#because t h i s estimation method i s the d e f a u l t one .
l i b r a r y ( " plm " )
femod <− plm ( fm , F a t a l i t i e s )
c o e f t e s t ( femod )
#NOTES:
# 1 . The f i x e d e f f e c t s model ,
# r e q u i r i n g only minimal assumptions on the nature o f heterogeneity ,
# i s one o f the sim plest and most robust s p e c i f i c a t i o n s
# in panel data econometrics and o f t e n the benchmark against
#which more s o p h i s t i c a t e d , and p o s s i b l y e f f i c i e n t ,
#ones are compared and judged in applied p r a c t i c e .
# 2 . Therefore i t i s a l s o the d e f a u l t c h o i c e in the b a s i c estimating f u n c t i o n plm .
# 2 . Example 2 : no h e t e r o g e n e i t y − T i l e r i e s data s e t
#NOTES:
# 1 . There are cases when unobserved h e t e r o g e n e i t y i s not an i s s u e .
# 2 . The T i l e r i e s dataset c o n t a i n s data on output and l a b o r
#and c a p i t a l inputs f o r 25 t i l e r i e s in two r e g i o n s o f Egypt ,
# observed over 12 t o 22 years .
360
Electronic copy available at: https://ssrn.com/abstract=3863563
8.19. PANEL DATA REGRESSION ANALYSIS
#We estimate a production f u n c t i o n .
# 3 . The i n d i v i d u a l u n i t s are rather homogeneous ,
#and the technology i s standard ;
#hence , most o f the v a r i a t i o n in output i s explained by the observed inputs .
#Here , a p o o l i n g s p e c i f i c a t i o n and a f i x e d e f f e c t s one g i v e very s i m i l a r r e s u l t s ,
# e s p e c i a l l y i f r e s t r i c t i n g the sample t o one o f the two r e g i o n s considered :
#2.1 i n s t a l l . packages ( " pder " )
i n s t a l l . packages ( " pder " )
l i b r a r y ( pder )
data ( " T i l e r i e s " , package = " pder " )
c o e f ( summary ( plm ( l o g ( output ) ~ l o g ( l a b o r ) + machine , data = T i l e r i e s ,
subset = area == " fayoum " ) ) )
c o e f ( summary ( plm ( l o g ( output ) ~ l o g ( l a b o r ) + machine , data = T i l e r i e s ,
model = " p o o l i n g " , subset = area == " fayoum " ) ) )
#NOTES:
# 1 . Notice that users have employed yet another way o f
#compactly l o o k i n g at the c o e f f i c i e n t s ’ t a b l e only ,
# instead o f p r i n t i n g the whole model summary :
#the c o e f . plm e x t r a c t o r method , applied t o a summary . plm o b j e c t .
# 2 . By the o b j e c t o r i e n t a t i o n o f R,
# applying c o e f t o a model or t o the summary o f a model − in o b j e c t terms ,
# t o a plm or t o a summary . plm − w i l l y i e l d d i f f e r e n t r e s u l t s .
# 3 .THE ERROR COMPONENTS MODEL_ORDINARLY LEAST SQUARES ESTIMATOR_Example 3 :
# within estimator − TobinQ data s e t
#3.1 Load the data
data ( " TobinQ " , package = " pder " )
#3.2 Now users e x p l o r e the d i f f e r e n t p o s s i b i l i t i e s below
361
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
pTobinQ <− pdata . frame ( TobinQ )
pTobinQa <− pdata . frame ( TobinQ , index = 188)
pTobinQb <− pdata . frame ( TobinQ , index = c ( ’ cusip ’ ) )
pTobinQc <− pdata . frame ( TobinQ , index = c ( ’ cusip ’ ,
’ year ’ ) )
#NOTES:
# 1 . The pdim f u n c t i o n can be used t o i n s p e c t the i n d i v i d u a l
#and time dimensions o f the data .
# 2 . I t has a method f o r pdata . frame o b j e c t s ( without any f u r t h e r argument )
#and f o r data . frame .
#3.3 In the l a t t e r case , the index argument can be s e t ;
# i f not , i t i s once more assumed that the f i r s t two columns o f the data . frame
# contain the i n d i v i d u a l and the time index .
pdim ( pTobinQ )
pdim ( TobinQ , index = ’ cusip ’ )
pdim ( TobinQ )
#NOTES: A pdata . frame has an index a t t r i b u t e ,
#which i s a data . frame that c o n t a i n s the index .
#3.4 I t can be e x t r a c t e d using the index f u n c t i o n :
head ( index ( pTobinQ ) )
#3.5 TThen estimate the three models we have d e s c r i b e d :
Qeq <− ikn ~ qn
Q. p o o l i n g <− plm ( Qeq , pTobinQ , model = " p o o l i n g " )
Q. within <− update (Q. pooling , model = " within " )
Q. between <− update (Q. pooling , model = " between " )
#3.6 Either simple or extended p r i n t i n g o f
#the r e s u l t s i s obtained as usual
#with R applying the p r i n t . plm or summary . plm methods t o the o b j e c t
# c o n t a i n i n g the f i t t e d model_For example , f o r the within estimator , we get :
Q. within
summary (Q. within )
#3.7 Now f o r the within estimator , users apply the f o l l o w i n g :
head ( f i x e f (Q. within ) )
362
Electronic copy available at: https://ssrn.com/abstract=3863563
8.19. PANEL DATA REGRESSION ANALYSIS
head ( f i x e f (Q. within , type = " d f i r s t " ) )
head ( f i x e f (Q. within , type = "dmean " ) )
#3.8 The f i x e d e f f e c t s are then equal t o those obtained
#using the f i x e f . plm f u n c t i o n with the argument type equal t o ’ d f i r s t ’ .
head ( c o e f ( lm ( ikn ~ qn + f a c t o r ( cusip ) , pTobinQ ) ) )
# 4 . Example 4 − random e f f e c t s model − using TobinQ data s e t
#4.1 The f o l l o w i n g two commands estimate the same Swamy and Arora ( 1 9 7 2 ) model :
Q. swar <− plm ( Qeq , pTobinQ , model = " random " , random . method = " swar " )
Q. swar2 <− plm ( Qeq , pTobinQ , model = " random " ,
random . models = c ( " within " , " between " ) ,
random . d f c o r = c ( 2 , 2 ) )
summary (Q. swar )
#4.2 Estimation o f the e r r o r component using ercomp f u n c t i o n
ercomp ( Qeq , pTobinQ )
ercomp (Q. swar )
#4.3 Users then compare the r e s u l t s obtained
#with the 4 estimation methods we have presented :
Q. walhus <− update (Q. swar , random . method = " swar " )
Q. amemiya <− update (Q. swar , random . method = " amemiya " )
Q. nerlove <− update (Q. swar , random . method = " nerlove " )
Q. models <− l i s t ( swar = Q. swar , walhus = Q. walhus ,
amemiya = Q. amemiya , nerlove = Q. nerlove )
sapply (Q. models , f u n c t i o n ( x ) ercomp ( x ) $theta )
sapply (Q. models , c o e f )
#NOTES: The f i r s t sapply command e x t r a c t s from
#the ercomp o b j e c t the theta element ,
# i n d i c a t i n g the p r o p o r t i o n o f the i n d i v i d u a l mean
# that i s removed from the v a r i a b l e s .
#These are very c l o s e t o each other ,
#and consequently , the estimated c o e f f i c i e n t s
# f o r the 4 models are almost i d e n t i c a l .
363
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 5 . COMPARISONS OF THE ESTIMATORS_FIXED vs RANDOM EFFECTS_Example 5 :
#comparison o f the e stim ato rs − TobinQ data s e t
sapply ( l i s t ( p o o l i n g = Q. pooling , within = Q. within ,
between = Q. between , swar = Q. swar ) ,
f u n c t i o n ( x ) c o e f ( summary ( x ) ) [ " qn " , c ( " Estimate " , " Std . Error " ) ] )
summary ( pTobinQ$qn )
#5.1 Users can use the Within and the Between f u n c t i o n with
# t h i s s e r i e s in order t o compute i t s within and the between transformations ,
#and then the weights o f the within
#and the between e stim ato rs in the OLS estimator .
SxxW <− sum( Within ( pTobinQ$qn ) ^ 2 )
SxxB <− sum ( ( Between ( pTobinQ$qn ) − mean ( pTobinQ$qn ) ) ^ 2 )
SxxTot <− sum( ( pTobinQ$qn − mean ( pTobinQ$qn ) ) ^ 2 )
pondW <− SxxW / SxxTot
pondW
pondW * c o e f (Q. within ) [ [ " qn " ] ] +
( 1 − pondW) * c o e f (Q. between ) [ [ " qn " ] ]
#NOTES: The weight o f the within model i s 57%.
#The OLS estimator ( 0 . 0 0 4 4 ) i s then about halfway between
#the between estimator ( 0 . 0 0 5 2 ) and the within estimator ( 0 . 0 0 3 8 ) .
#5.2 To get the GLS estimator ,
#we f i r s t estimate the parameter Phi
#using the r e s i d u a l s o f the within and the between e sti mato rs :
T <− 35
N <− 188
smxt2 <− deviance (Q. between ) * T / (N − 2 )
s i d i o s 2 <− deviance (Q. within ) / (N * ( T − 1 ) − 1 )
phi <− s q r t ( s i d i o s 2 / smxt2 )
#5.3 The weights f o r the within
#and the between e stim ato rs and the g l s estimator are then computed :
pondW <− SxxW / (SxxW + phi ^ 2 * SxxB )
364
Electronic copy available at: https://ssrn.com/abstract=3863563
8.19. PANEL DATA REGRESSION ANALYSIS
pondW
pondW * c o e f (Q. within ) [ [ " qn " ] ] +
( 1 − pondW) * c o e f (Q. between ) [ [ " qn " ] ]
#NOTES:
# 1 . The weight o f the within estimator ( 0 . 9 5 )
# i s much l a r g e r f o r the g l s estimator than f o r the OLS estimator .
## # This i s mainly due t o the f a c t that T i s l a r g e (35 years ) .
# 2 . The GLS estimator ( 0 . 0 3 9 )
# i s t h e r e f o r e very c l o s e t o the within estimator ( 0 . 0 0 3 8 ) .
# 6 . Example 6
data ( " ForeignTrade " , package = " pder " )
FT <− pdata . frame ( ForeignTrade )
summary ( FT$gnp )
#6.1 Now use ercomp f u n c t i o n
ercomp ( imports ~ gnp , FT)
models <− c ( " within " , " random " , " p o o l i n g " , " between " )
sapply ( models , f u n c t i o n ( x ) c o e f ( plm ( imports ~ gnp , FT, model = x ) ) [ " gnp " ] )
#NOTES:
# 1 . For t h i s model , the variance o f the c o v a r i a t e
#and o f the e r r o r i s almost
# only due t o the i n t e r −i n d i v i d u a l v a r i a t i o n ( r e s p e c t i v e l y 98 and 93%).
# 2 . In t h i s case ,
#the g l s estimator c o n s i s t s in removing 94% o f the i n d i v i d u a l mean
#and i s t h e r e f o r e almost i d e n t i c a l t o the within
# 3 . model . Concerning the o l s estimator ,
#which takes i n t o account almost a l l the i n t e r −i n d i v i d u a l v a r i a t i o n ,
# i t i s very c l o s e t o the between estimator .
# Finally , the f i r s t two models g i v e r e s u l t s that are very d i f f e r e n t from
#the l a s t two models and return a much higher e l a s t i c i t y .
# 4 . The Figure ( see PowerPoint ) i n d i c a t e s
# that there i s a strong negative c o r r e l a t i o n between the i n d i v i d u a l e f f e c t s
#and the c o v a r i a t e .
#In t h i s case , the e stim ato rs that do not c o n t r o l
365
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# f o r the i n d i v i d u a l e f f e c t s are biased downward .
#This i s the case f o r the OLS and the between estimators ,
#and t o a much l e s s e r extent f o r the GLS estimator ,
#which uses only a very small part o f the i n t e r −i n d i v i d u a l v a r i a t i o n .
# 7 . Example 7 − Turkish Banks
data ( " TurkishBanks " , package = " pder " )
TurkishBanks <− na . omit ( TurkishBanks )
TB <− pdata . frame ( TurkishBanks )
summary ( l o g ( TB$output ) )
ercomp ( l o g ( c o s t ) ~ l o g ( output ) , TB)
sapply ( models , f u n c t i o n ( x )
c o e f ( plm ( l o g ( c o s t ) ~ l o g ( output ) , TB, model = x ) ) [ " l o g ( output ) " ] )
# 8 . Example 8 − Texas E l e c t r i c
data ( " TexasElectr " , package = " pder " )
T ex a sE l e c t r $ c o s t <− with ( TexasElectr , explab + e x p f u e l + expcap )
TE <− pdata . frame ( TexasElectr )
summary ( l o g ( TE$output ) )
ercomp ( l o g ( c o s t ) ~ l o g ( output ) , TE)
sapply ( models , f u n c t i o n ( x )
c o e f ( plm ( l o g ( c o s t ) ~ l o g ( output ) , TE, model = x ) ) [ " l o g ( output ) " ] )
# 9 . Example 9 : simple l i n e a r model − DemocracyIncome25 data s e t
data ( " DemocracyIncome25 " , package = " pder " )
DI <− pdata . frame ( DemocracyIncome25 )
summary ( lag ( DI$income ) )
ercomp ( democracy ~ lag ( income ) , DI )
sapply ( models , f u n c t i o n ( x )
c o e f ( plm ( democracy ~ lag ( income ) , DI , model = x ) ) [ " lag ( income ) " ] )
366
Electronic copy available at: https://ssrn.com/abstract=3863563
8.19. PANEL DATA REGRESSION ANALYSIS
#10. Example 1 0 ; two−ways on TobinQ data s e t
Q. models2 <− l a p p l y (Q. models , f u n c t i o n ( x ) update ( x , e f f e c t = " twoways " ) )
sapply (Q. models2 , f u n c t i o n ( x ) s q r t ( ercomp ( x ) $sigma2 ) )
#NOTES:
# 1 . The f i r s t l a p p l y command
# e x t r a c t s the standard d e v i a t i o n s o f the three components o f the e r r o r .
# 2 . As f o r the i n d i v i d u a l e f f e c t s model ,
#the estimates o f the variance components are very s i m i l a r .
# 3 . The standard d e v i a t i o n o f i n d i v i d u a l e f f e c t s i s more than
# twice the one o f time e f f e c t s .
# 4 . The second command e x t r a c t s the theta parameters .
# 5 . About 75% o f the i n d i v i d u a l and time means are removed from the v a r i a b l e s .
#11. ESTIMATION OF A WAGE EQUATION_Example 1 1 : UnionWage data s e t
#11.1 i n s t a l l and load packages
i n s t a l l . packages ( " pglm " )
library
l i b r a r y ( maxLik )
l i b r a r y ( miscTools )
l i b r a r y ( pglm )
data ( " UnionWage " , package = " pglm " )
pdim ( UnionWage )
UnionWage$exper2 <− with ( UnionWage , exper ^ 2 )
wages . within1 <− plm ( wage ~ union + s c h o o l + exper + exper2 +
com + r u r a l + married + health +
region + s e c t o r , UnionWage )
wages . within2 <− plm ( wage ~ union + s c h o o l + exper + exper2 +
com + r u r a l + married + health +
region + s e c t o r + occ , UnionWage )
wages . pooling1 <− update ( wages . within1 , model = " p o o l i n g " )
wages . pooling2 <− update ( wages . within2 , model = " p o o l i n g " )
#11.2 Users a l s o need t h i s package c a l l e d s t a r g a z e r
367
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
i n s t a l l . packages ( " s t a r g a z e r " )
library ( " stargazer " )
#NOTES: Package s t a r g a z e r can d i r e c t l y produces LaTeX code in RStudio .
s t a r g a z e r ( wages . pooling2 , wages . pooling1 , wages . within2 , wages . within1 ,
omit = c ( " region " , " s e c t o r " , " o c c " ) ,
omit . l a b e l s = c ( " region dummies " ,
" s e c t o r dummies " ,
" occupation dummies " ) ,
column . l a b e l s = c ( " p o o l i n g estimation " , " within estimation " ) ,
column . separate = c ( 2 , 2 ) ,
dep . var . l a b e l s = " l o g o f hourly wage " ,
c o v a r i a t e . l a b e l s = c ( " union membership " , " education years " ,
" experience years " , " experience years squared " ,
" black " , " h i s p a n i c " , " r u r a l r e s i d e n c e " ,
" married " , " health problems " ,
" Intercept " ) ,
omit . s t a t = c ( " adj . rsq " , " f " ) ,
t i t l e = "Wage equation " ,
l a b e l = " tab : wagesresult " ,
no . space = TRUE
)
#NOTES: Produces data shown in l a s t t a b l e on s l i d e s f o r Wage Equation
368
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
8.20
Advanced error component models
## ADVANCED ERROR COMPONENT MODELS & TESTS
# 1 . UNBALANCED PANELS_EXERCISE
#1.1 i n s t a l l and load package pder
i n s t a l l . packages ( " pder " )
#1.2 Example 1 − T i l e r i e s data s e t r e v i s i t e d
data ( " T i l e r i e s " , package = " pder " )
head ( T i l e r i e s , 3 )
#1.3 Users need the f o l l o w i n g plm l i b r a r y f o r pdim f u n c t i o n
l i b r a r y ( plm )
#1.4 Now users can t e s t f o r balanced
#vs unbalanced using panel dimension f u n c t i o n pdim
pdim ( T i l e r i e s )
#NOTES: Estimate a Cobb−Douglas production f u n c t i o n
#where the production ( output ) depends on the quantity o f two inputs ,
# l a b o r ( l a b o r ) and machines ( machine ) .
#1.5 F i r s t check that the same f i x e d e f f e c t s model
#can s t i l l be estimated e i t h e r
#by applying OLS on the within transformed v a r i a b l e s
# or using i n d i v i d u a l dummy v a r i a b l e s :
T i l e r i e s <− pdata . frame ( T i l e r i e s )
plm . within <− plm ( l o g ( output ) ~ l o g ( l a b o r ) + l o g ( machine ) , T i l e r i e s )
y <− l o g ( T i l e r i e s $ o u t p u t )
x1 <− l o g ( T i l e r i e s $ l a b o r )
x2 <− l o g ( Tileries$machine )
#1.6 Next measure Within v a r i a t i o n and Between v a r i a t i o n
lm . within <− lm ( I ( y − Between ( y ) ) ~ I ( x1 − Between ( x1 ) ) +
I ( x2 − Between ( x2 ) ) − 1 )
lm . l s d v <− lm ( l o g ( output ) ~ l o g ( l a b o r ) +
l o g ( machine ) +
factor ( id ) ,
369
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
Tileries )
c o e f ( lm . l s d v ) [ 2 : 3 ]
#1.7 The one−way random e f f e c t s model i s then estimated :
t i l e . r <− plm ( l o g ( output ) ~ l o g ( l a b o r ) + l o g ( machine ) ,
T i l e r i e s , model = " random " )
#1.8 Examine the output t a b l e :
summary ( t i l e . r )
#NOTES:
# 1 . The transformation parameter i s now i n d i v i d u a l s p e c i f i c ;
#more p r e c i s e l y ,
# i t depends on the number o f a v a i l a b l e o b s e r v a t i o n s
# f o r every i n d i v i d u a l . The theta parameter here v a r i e s from 0.49 t o 0 . 6 0 .
# 2 . The two−ways random e f f e c t model i s obtained
#by s e t t i n g the e f f e c t argument t o ’ twoways ’ .
# 3 . We check that the OLS model cannot be obtained any more by applying OLS
# t o v a r i a b l e s where the i n d i v i d u a l and time means have been removed .
plm . within <− plm ( l o g ( output ) ~ l o g ( l a b o r ) + l o g ( machine ) ,
T i l e r i e s , e f f e c t = " twoways " )
lm . l s d v <− lm ( l o g ( output ) ~ l o g ( l a b o r ) + l o g ( machine ) +
f a c t o r ( i d ) + f a c t o r ( week ) , T i l e r i e s )
y <− l o g ( T i l e r i e s $ o u t p u t )
x1 <− l o g ( T i l e r i e s $ l a b o r )
x2 <− l o g ( Tileries$machine )
y <− y − Between ( y , " i n d i v i d u a l " ) − Between ( y , " time " ) + mean ( y )
x1 <− x1 − Between ( x1 , " i n d i v i d u a l " ) − Between ( x1 , " time " ) + mean ( x1 )
x2 <− x2 − Between ( x2 , " i n d i v i d u a l " ) − Between ( x2 , " time " ) + mean ( x2 )
lm . within <− lm ( y ~ x1 + x2 − 1 )
#1.9 F i r s t check the plm f u n c t i o n output
370
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
c o e f ( plm . within )
#1.10 Second , we check the lm f u n c t i o n output
c o e f ( lm . within )
#1.11 Third , we look at the l e a s t squares dep var
c o e f ( lm . l s d v ) [ 2 : 3 ]
#1.12 F i n a l l y user estimate the time and i n d i v i d u a l random e f f e c t s model ,
#using the three methods o f estimation we have d e s c r i b e d :
wh <− plm ( l o g ( output ) ~ l o g ( l a b o r ) + l o g ( machine ) , T i l e r i e s ,
model = " random " , random . method = " walhus " ,
e f f e c t = " twoways " )
am <− update (wh, random . method = " amemiya " )
sa <− update (wh, random . method = " swar " )
ercomp ( sa )
#NOTES: The shares o f the i n d i v i d u a l and the time e f f e c t s
# in the t o t a l e r r o r variance are now about 19 percent ( 1 8 . 5 )
#and 5 percent ( 4 . 7 ) f o r the Swamy−Arora estimator .
#1.13 Comparing models using sapply f u n c t i o n
re . models <− l i s t ( walhus = wh, amemiya = am, swar = sa )
sapply ( re . models , f u n c t i o n ( x ) s q r t ( ercomp ( x ) $sigma2 ) )
#1.14 We can a l s o compare each models c o e f f i c i e n t s
sapply ( re . models , c o e f )
#NOTES: There i s not a huge d i f f e r e n c e between them .
#EXAMPLE 2 : SUR estimation using TexasElectr data s e t
#2.1 I n s t a l l the dplyr package
i n s t a l l . packages ( " dplyr " )
l i b r a r y ( dplyr )
#2.2 Now we can use the sample data
371
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
data ( " TexasElectr " , package = " pder " )
#2.3 Next we use the mutate ( ) f u n c t i o n
TexasElectr <− mutate ( TexasElectr ,
pf = l o g ( p f u e l / mean ( p f u e l ) ) ,
p l = l o g ( plab / mean ( plab ) ) − pf ,
pk = l o g ( pcap / mean ( pcap ) ) − pf )
#2.4 The production i s a l s o measured in logarithms
#and d i v i d e d by i t s sample mean .
TexasElectr <− mutate ( TexasElectr , q = l o g ( output / mean ( output ) ) )
#2.5 Then compute t o t a l production c o s t by summing the expenses
# f o r the three f a c t o r s and f a c t o r shares .
# Finally , we measure the c o s t in logarithms
#and d i v i d e i t by i t s sample mean and by the r e f e r e n c e p r i c e .
TexasElectr <− mutate ( TexasElectr ,
C = e x p f u e l + explab + expcap ,
s l = explab / C,
sk = expcap / C,
C = l o g (C / mean (C ) ) − pf )
#2.6 Finally , we compute the squares
#and the i n t e r a c t i o n terms f o r the v a r i a b l e s .
TexasElectr <− mutate ( TexasElectr ,
p l l = 1/2 * pl ^ 2 ,
plk = p l * pk ,
pkk = 1 / 2 * pk ^ 2 ,
qq = 1 / 2 * q ^ 2 )
#2.7 Define the three equations o f the system ,
#one f o r t o t a l c o s t and the other two f o r f a c t o r shares .
#NOTES: The f u e l share i s omitted t o avoid p e r f e c t c o l l i n e a r i t y ,
# given that the three shares sum t o one .
c o s t <− C ~ p l + pk + q + p l l + plk + pkk + qq
shlab <− s l ~ p l + pk
shcap <− sk ~ p l + pk
372
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
#2.8 We c o n s t r u c t f o r t h i s purpose a 6 ( number o f r e s t r i c t i o n s )
#by 14 ( number o f c o e f f i c i e n t s ) matrix .
R <− matrix ( 0 , nrow = 6 , n c o l = 14)
R[ 1 , 2 ] <− R[ 2 , 3 ] <− R[ 3 , 5 ] <− R[ 4 , 6 ] <− R[ 5 , 6 ] <− R[ 6 , 7 ] <− 1
R[ 1 , 9 ] <− R[ 2 , 12] <− R[ 3 , 10] <− R[ 4 , 11] <− R[ 5 , 13] <− R[ 6 , 14] <− −1
#NOTES:
# 1 . The f i r s t l i n e o f the matrix i n d i c a t e s that
#the second c o e f f i c i e n t ( the one a s s o c i a t e d t o p l in the c o s t equation )
#must be equal t o the ninth ( the constant term in the l a b o r share equation ) .
# 2 . The SUR model i s estimated pr ovid ing a l i s t o f formulae ,
# d e f i n i n g the system o f equations t o be estimated ,
#as the f i r s t argument t o plm .
# 3 . The d i f f e r e n t formulae in the l i s t can be named ,
#which makes the output more readable .
# 4 . The model argument i s s e t t o ’ random ’ in order t o
# estimate the SUR e r r o r components model .
# 5 . Lastly , the arguments r e s t r i c t . matrix
#and r e s t r i c t . rhs allow t o s p e c i f y the matrix R
#and the v e c t o r q d e f i n i n g the l i n e a r c o n s t r a i n t s o f the model .
#2.9 I f , as happens here , a l l elements o f q are zero ,
#the r e s t r i c t . rhs argument can be omitted .
z <− plm ( l i s t ( c o s t = C ~ p l + pk + q + p l l + plk + pkk + qq ,
shlab = s l ~ p l + pk ,
shcap = sk ~ p l + pk ) ,
TexasElectr , model = " random " ,
r e s t r i c t . matrix = R)
summary ( z )
#NOTES:
373
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 1 . The r e s u l t s i n d i c a t e the presence o f i n c r e a s i n g returns t o s c a l e ,
#as q i s s i g n i f i c a n t l y lower than 1 .
# 2 . Factor shares at the sample mean f o r l a b o r
#and c a p i t a l are r e s p e c t i v e l y 12 and 31 percent .
#THE MAXIMUM LIKELIHOOD ESTIMATOR_EXAMPLE 3 :
#maximum l i k e l i h o o d estimator RiceFarms
#3.1 F i r s t , i n s t a l l and l i b r a r y package splm
i n s t a l l . packages ( " splm " )
i n s t a l l . packages ( " u n i t s " )
i n s t a l l . packages ( " p r e t t y u n i t s " )
library ( prettyunits )
library ( units )
library ( sf )
l i b r a r y ( spdep )
l i b r a r y ( sp )
l i b r a r y ( spData )
l i b r a r y ( splm )
#NOTES: Some package can not be loaded ,
#but i t w i l l not have negative impacts on data loading and c a l c u l a t i n g .
#3.2 Load data
data ( " RiceFarms " , package = " splm " )
Rice <− pdata . frame ( RiceFarms , index = " i d " )
r i c e . ml <− pglm ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) ,
data = Rice , family = gaussian )
summary ( r i c e . ml )
#NOTES:
# 1 . The c o e f f i c i e n t s are very s i m i l a r t o those obtained with the GLS estimator .
# 2 . The two parameters c a l l e d sd . i d i o s and sd . i d
#are the estimated standard d e v i a t i o n s o f the i d i o s y n c r a t i c
#and o f the i n d i v i d u a l parts o f the e r r o r .
374
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
# 3 . These values are a l s o almost equal t o those obtained using the g l s estimator .
# 4 .THE NESTED ERROR COMPONENTS MODEL_EXAMPLE 4 : Produc data s e t
data ( " RiceFarms " , package = " plm " )
head ( RiceFarms , 2 )
#NOTES: Users can see the f i r s t 2 rows o f data
#4.1 The i n d i v i d u a l index i s i d ( the f i s t column ) ,
#we use region , the l a s t column o f the data frame as the group index .
#The three l i n e s below g i v e the same r e s u l t s :
R1 <− pdata . frame ( RiceFarms , index = c ( i d = " i d " ,
time = NULL,
group = " region " ) )
R2 <− pdata . frame ( RiceFarms , index = c ( i d = " i d " ,
group = " region " ) )
R3 <− pdata . frame ( RiceFarms , index = c ( " i d " ,
group = " region " ) )
head ( index ( R1 ) )
#NOTES:
# 1 . For the Produc data frame ,
# i t i s e a s i e r t o d e s c r i b e the s t r u c t u r e o f the sample as
#the f i r s t three columns are the i n d i v i d u a l , time , and group indexes .
# 2 . To estimate the nested e r r o r component model ,
#the model must be s e t t o ’ nested ’ .
#4.2 We f i r s t estimate the Swamy and Arora ( 1 9 7 2 ) model :
data ( " Produc " , package = " plm " )
nswar <− plm ( l o g ( gsp ) ~ l o g ( pc ) + l o g (emp) + l o g ( hwy ) + l o g ( water ) +
l o g ( u t i l ) + unemp , data = Produc ,
model = " random " , e f f e c t = " nested " ,
random . method = " swar " , index = c ( group = " re gion " ) )
summary ( nswar )
375
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES: Then update the model in order t o use the two other esti mato rs o f
#the variances o f the components o f the e r r o r .
#The r e s u l t s are summarized using the screenreg f u n c t i o n o f the texreg package .
#4.3 F i r s t , we need t o i n s t a l l t h i s package
i n s t a l l . packages ( " texreg " )
l i b r a r y ( " texreg " )
namem <− update ( nswar , random . method = " amemiya " )
nwalhus <− update ( nswar , random . method = " walhus " )
iswar <− update ( nswar , e f f e c t = " i n d i v i d u a l " )
iwith <− update ( nswar , model = " within " , e f f e c t = " i n d i v i d u a l " )
screenreg ( l i s t ( " fe −i d " = iwith , " re−i d " = iswar ,
" Swamy_Arora " = nswar , " Wallas−Hussein " = nwalhus ,
"Amemiya" = namem) , d i g i t s = 3 )
# 5 . TESTS FOR ERROR COMPONENTS_Example 5 : F and LM t e s t s − RiceFarms dat s e t
#5.1 Load the data
data ( " RiceFarms " , package = " splm " )
#5.2 Run the f o l l o w i n g command and IGNORE ’ time ’ ov er wri tt en
#by time index message !
Rice <− pdata . frame ( RiceFarms , index = " i d " )
r i c e .w <− plm ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) , Rice )
r i c e . p <− update ( r i c e . w, model = " p o o l i n g " )
r i c e . wd <− plm ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) , Rice ,
e f f e c t = " twoways " )
#5.3 Then supply the t e s t i n g f u n c t i o n with two f i t t e d models :
pFtest ( r i c e . w, r i c e . p )
#5.4 The formula−data syntax may a l s o be used :
pFtest ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) , Rice )
#NOTES: Both show the same r e s u l t . Unsurprisingly ,
376
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
#the absence o f i n d i v i d u a l e f f e c t s i s s t r o n g l y r e j e c t e d .
#5.5 To t e s t the absence o f i n d i v i d u a l and time e f f e c t s , one would use :
pFtest ( r i c e . wd, r i c e . p )
#5.6 Use r i c e . wd and not r i c e .w l i k e before , or we could use :
pFtest ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) , Rice ,
e f f e c t = " twoways " )
#NOTES: and get the same r e s u l t
#5.7 To t e s t the absence o f time e f f e c t s allowing f o r
#the presence o f i n d i v i d u a l e f f e c t s , compare the i n d i v i d u a l
#and the two−ways e f f e c t within models :
pFtest ( r i c e . wd, r i c e .w)
#NOTES:
# 1 . Once more , the n u l l hypothesis i s very s t r o n g l y r e j e c t e d
# 2 . The Breusch and Pagan ( 1 9 8 0 ) t e s t can be computed using the f u n c t i o n plmtest .
# 3 . The argument i s e i t h e r an OLS model or a formula−data p a i r .
# 4 . By d e f a u l t , the Honda ( 1 9 8 5 ) v e r s i o n i s computed .
# 5 . The d i r e c t i o n o f the e f f e c t s , as usual , i s determined by the e f f e c t argument .
plmtest ( r i c e . p )
plmtest ( l o g ( goutput ) ~ l o g ( seed )+ l o g ( t o t l a b o r )+ l o g ( s i z e ) , Rice )
plmtest ( r i c e . p , e f f e c t = " time " )
plmtest ( r i c e . p , e f f e c t = " twoways " )
#NOTES: The two u s e f u l e x t e n s i o n s proposed
#by B a l t a g i e t a l ( 1 9 9 2 ) can be applied t o t e s t the e x i s t e n c e o f i n d i v i d u a l
#and time e f f e c t s , s e t t i n g the argument type t o ’kw ’ or ’ghm ’
# t o use r e s p e c t i v e l y the techniques proposed by King and Wu ( 1 9 9 7 )
#and Gourieroux e t a l . ( 1 9 8 2 ) .
plmtest ( r i c e . p , e f f e c t = " twoways " , type = "kw " )
377
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
plmtest ( r i c e . p , e f f e c t = " twoways " , type = "ghm " )
# 6 . TESTS ON ERROR COMPONEMNT MODELS_Example 6 : Hausman t e s t − RiceFarms data
#6.1 Load the data
data ( " RiceFarms " , package = " splm " )
Rice <− pdata . frame ( RiceFarms , index = " i d " )
#6.2 Ignore message about column ’ time ’ being ov er wri tt en
r i c e .w <− plm ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) , Rice )
r i c e . r <− update ( r i c e . w, model = " random " )
#6.3 Now run the a c t u a l Hausman Test
phtest ( r i c e . w, r i c e . r )
#NOTES:
# 1 . Under the hypothesis o f no c o r r e l a t i o n between the r e g r e s s o r s
#and the i n d i v i d u a l e f f e c t s ,
#the s t a t i s t i c i s d i s t r i b u t e d as a chi −squared with three degrees o f freedom .
# 2 . This hypothesis , with a p−value o f 29%,
# i s not r e j e c t e d at the 5% c o n f i d e n c e l e v e l .
# 3 . One could get the same r e s u l t f o l l o w i n g the Mundlak ( 1 9 7 8 ) approach ,
#drawing on the d i f f e r e n c e between the within and between esti mato rs :
r i c e . b <− update ( r i c e . w, model = " between " )
cp <− i n t e r s e c t ( names ( c o e f ( r i c e . b ) ) , names ( c o e f ( r i c e .w ) ) )
d c o e f <− c o e f ( r i c e .w) [ cp ] − c o e f ( r i c e . b ) [ cp ]
V <− vcov ( r i c e .w) [ cp , cp ] + vcov ( r i c e . b ) [ cp , cp ]
as . numeric ( t ( d c o e f ) %*% s o l v e (V) %*% d c o e f )
#NOTES:
# 1 . So see the chi −square c r i t i c a l value i s very s i m i l a r .
# 2 . This r e s u l t i s confirmed by the c o r r e l a t i o n c o e f f i c i e n t
#between the i n d i v i d u a l e f f e c t s
# ( estimated by the f i x e d e f f e c t s o f the within model )
#and the i n d i v i d u a l means o f the explanatory v a r i a b l e ,
#which i s obtained by applying the between f u n c t i o n t o the s e r i e s .
#6.4 f i x e f f u n c t i o n e x t r a c t s the estimates o f
378
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
#the FE parameters o f a f i t t e d model
f i x e f ( r i c e .w)
#NOTES:
# 1 . c o r ( f i x e f ( r i c e .w) , between ( l o g ( Rice$goutput ) ) ) shows c o r r e l a t i o n = 0.448
# 2 . The c o r r e l a t i o n i s p o s i t i v e but moderate .
# 3 . The Chamberlain t e s t i s a v a i l a b l e in the f u n c t i o n p i e s t .
# 4 . I t i s computed using the usual formula−data i n t e r f a c e .
data ( " RiceFarms " , package = " splm " )
pdim ( RiceFarms , index = " i d " )
p i e s t ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) ,
RiceFarms , index = " i d " )
# Ignore ’ time ’ column ove rw rit te n message
# Note : the p−value = 0.03
#6.5 The variant o f the Chamberlain t e s t proposed
#by Angrist and Newey ( 1 9 9 1 ) i s a v a i l a b l e with
#the aneweytest f u n c t i o n which uses the same i n t e r f a c e .
aneweytest ( l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e ) ,
RiceFarms , index = " i d " )
#NOTES: The r e s t r i c t i o n s implied by the within model are r e j e c t e d by
#both t e s t s at the 5% l e v e l ,
#although they are not r e j e c t e d at the 1% l e v e l f o r
#Chamberlain ’ s v e r s i o n o f the t e s t . p−value o f Chi−Sq t e s t = 0.0002
# 7 . Example 7 : unobserved e f f e c t s t e s t − RiceFarms data s e t
data ( " RiceFarms " , package ="plm " )
Rice <− pdata . frame ( RiceFarms , index = " i d " )
fm <− l o g ( goutput ) ~ l o g ( seed ) + l o g ( t o t l a b o r ) + l o g ( s i z e )
pwtest ( fm , Rice )
#NOTES: Short and simple t e s t t o use .
# 8 . Example 8 : LM t e s t s f o r random e f f e c t s and / or s e r i a l c o r r e l a t i o n
379
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
i n s t a l l . packages ( " plm " )
l i b r a r y ( plm )
bsy .LM <− matrix ( n c o l =3 , nrow = 2 )
t e s t s <− c ( " J " , "RE" , "AR" )
dimnames ( bsy .LM) <− l i s t ( c ( "LM t e s t " , " p−value " ) , t e s t s )
f o r ( i in t e s t s ) {
mytest <− p b s y t e s t ( fm , data = Rice , t e s t = i )
bsy .LM[ 1 : 2 , i ] <− c ( m y t e s t $ s t a t i s t i c , mytest$p . value )
}
round ( bsy .LM, 6 )
#NOTES: The robust t e s t s allow us t o d i s c r i m i n a t e
#between time−i n v a r i a n t e r r o r p e r s i s t e n c e ( random e f f e c t s )
#and time−decaying p e r s i s t e n c e ( a u t o r e g r e s s i v e e r r o r s ) ,
# concluding in f a v o r o f the second .
# Finally , the optimal c o n d i t i o n a l t e s t o f B a l t a g i and
#Li f o r s e r i a l c o r r e l a t i o n ,
# allowing f o r random e f f e c t s o f any magnitude ,
# i s computed using the p b l t e s t function ,
#using the r e s i d u a l s o f the random e f f e c t s maximum l i k e l i h o o d estimator :
p b l t e s t ( fm , Rice , a l t e r n a t i v e = " onesided " )
# S e r i a l c o r r e l a t i o n i s detected , i . e . ,
#we conclude that Psi i s not equal t o Zero ( R e j e c t H0) in the encompassing model
# 9 . Example 9 : LR t e s t s f o r AR( 1 ) and I n d i v i d u a l E f f e c t s
#9.1 I n s t a l l and load the nlme package
i n s t a l l . packages ( " plm " )
l i b r a r y ( plm )
i n s t a l l . packages ( " nlme " )
l i b r a r y ( nlme )
#NOTES:
380
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
# 1 . Maximum l i k e l i h o o d estimation o f l i n e a r models with
# or without e i t h e r random i n d i v i d u a l e f f e c t s
# or s e r i a l l y c o r r e l a t e d e r r o r s can be estimated , e . g . ,
#with f u n c t i o n a l i t y from the nlme package .
# 2 . In i t s notation , the re s p e c i f i c a t i o n i s amodel
#with only one random e f f e c t s r e g r e s s o r : the i n t e r c e p t .
# 3 . Below we r e p o r t c o e f f i c i e n t s o f Grunfeld ’ s model estimated by g l s and then by ml .
#9.2 Load data
data ( Grunfeld , package = " plm " )
reGLS <− plm ( inv ~ value + c a p i t a l , data = Grunfeld , model = " random " )
#NOTE: there are 2 nlme package v e r s i o n s − USE THE OLDER ONE! Version 3.1 − 149
reML <− lme ( inv ~ value + c a p i t a l , data = Grunfeld , random = ~1 | firm )
rbind ( c o e f ( reGLS ) , f i x e f ( reML ) )
#9.3 Linear models with group wise s t r u c t u r e s o f time dependence
#may be f i t t e d by GLS, s p e c i f y i n g the c o r r e l a t i o n s t r u c t u r e
# in the c o r r e l a t i o n o p t i o n :
lmAR1ML <− g l s ( inv ~ value + c a p i t a l , data = Grunfeld ,
c o r r e l a t i o n = corAR1 ( 0 , form = ~ year | firm ) )
#9.4 and analogously the random e f f e c t s panel with ,
#e . g . , ar ( 1 ) e r r o r s ( see Baltagi , 2013 , Ch . 5 )
#may be f i t by lme s p e c i f y i n g an a d d i t i o n a l random i n t e r c e p t
reAR1ML <− lme ( inv ~ value + c a p i t a l , data = Grunfeld ,
random = ~ 1 | firm , c o r r e l a t i o n = corAR1 ( 0 ,
form = ~ year | firm ) )
#9.5 Now we examine output using summary ( ) f u n c t i o n
summary (reAR1ML)
#9.6 Compare e i t h e r with the r e s t r i c t e d a l t e r n a t i v e . T
#he g l s model without c o r r e l a t i o n in the r e s i d u a l s i s the same as OLS,
#and one could well use lm f o r the r e s t r i c t e d model .
#Here we estimate i t by GLS.
lmML <− g l s ( inv ~ value + c a p i t a l , data = Grunfeld )
anova (lmML, lmAR1ML)
381
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#9.7 The AR( 1 ) t e s t on the random e f f e c t s model i s t o be done
# in much the same way , using the random e f f e c t s model o b j e c t s estimated above :
anova ( reML , reAR1ML)
#9.8 A l i k e l i h o o d r a t i o t e s t f o r random e f f e c t s compares the s p e c i f i c a t i o n s
#with and without random e f f e c t s and s p h e r i c a l i d i o s y n c r a t i c e r r o r s :
anova (lmML, reML )
#9.9 The random e f f e c t s ,
#AR( 1 ) e r r o r s model in turn nests the AR( 1 ) p o o l i n g model ;
#therefore , a likelihood ratio test for
#random e f f e c t s sub ar ( 1 ) e r r o r s may be c a r r i e d out ,
#again , by comparing the two a u t o r e g r e s s i v e s p e c i f i c a t i o n s :
anova (lmAR1ML, reAR1ML)
#9.10 Once see that the Grunfeld model s p e c i f i c a t i o n
#does not seem t o need any random e f f e c t s once
#we c o n t r o l f o r s e r i a l c o r r e l a t i o n in the data . CONCLUSION.
#10. Example 1 0 : Breusch−Godfrey and Durbin Watson t e s t s − RiceFarms data s e t
#10.1 The f u n c t i o n s pbgtest
#and pdwtest re−estimate the r e l e v a n t quasi−demeaned model by o l s and apply ,
#respectively ,
#standard Breusch−Godfrey and Durbin−Watson t e s t s from package lmtest
# t o the r e s i d u a l s :
r i c e . re <− plm ( fm , Rice , model = ’ random ’ )
pbgtest ( r i c e . re , order = 2 )
pdwtest ( r i c e . re , order = 2 )
#NOTES:
# 1 . The t e s t s share the f e a t u r e s o f t h e i r o l s counterparts ,
# in p a r t i c u l a r the pbgtest allows t e s t i n g f o r higher−order s e r i a l c o r r e l a t i o n ,
#which can be o f p a r t i c u l a r i n t e r e s t f o r q u a r t e r l y data .
382
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
# 2 . As the f u n c t i o n s are simple wrappers toward b g t e s t and dwtest ,
# a l l arguments from the l a t t e r two apply
#and may be passed on through the
’ . ’ o pera tor
# 3 . As observed above ,
# applying the pbgtest and pdwtest f u n c t i o n s t o an FE model i s appropriate
# only i f the time dimension i s long enough .
# 4 . In the frequent case o f " short " panels ,
#one o f the two t e s t i n g procedures due toWooldridge ( 2 0 1 0 )
#and d e s c r i b e d in the next s e c t i o n should be used instead .
#11. Example 1 1 : s e r i a l c o r r e l a t i o n t e s t s f o r FE models − EmplUK data s e t
#11.1 Employing Wooldridge ’ s within
data ( "EmplUK" , package = " plm " )
pwartest ( l o g (emp) ~ l o g ( wage ) + l o g ( c a p i t a l ) , data = EmplUK)
#NOTES:
# Strongly r e j e c t the n u l l o f no s e r i a l c o r r e l a t i o n .
# I f the evidence o f p e r s i s t e n c e i s t o o strong ,
#one should wonder whether the r e s i d u a l s are s t a t i o n a r y at a l l ,
#and whether a s p e c i f i c a t i o n in d i f f e r e n c e s might be p r e f e r a b l e .
#In the next s e c t i o n we w i l l see a s i m i l a r t e s t
# that can be seen as a s p e c i f i c a t i o n d e v i c e in t h i s sense
#12. Example 1 2 : Wooldridge ’ s f i r s t d i f f e r e n c e t e s t − RiceFarm data s e t
# Results w i l l not always be so c l e a r −cut . For the RiceFarms example
W. fd <− matrix ( n c o l = 2 , nrow =2)
H0 <− c ( " fd " , " f e " )
dimnames (W. fd ) <− l i s t ( c ( " t e s t " , " p−value " ) , H0)
f o r ( i in H0) {
mytest <− pwfdtest ( fm , Rice , h0 = i )
W. fd [ 1 , i ] <− m y t e s t $ s t a t i s t i c
W. fd [ 2 , i ] <− mytest$p . value
}
383
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
round (W. fd , 6 )
# Thetruth c l e a r l y l i e s in themiddle
# ( both r e j e c t e d , although one more s t r o n g l y than the other ) ;
# in t h i s case ,
#whichever estimator i s chosen w i l l have s e r i a l l y c o r r e l a t e d e r r o r s :
# t h e r e f o r e i t w i l l be a d v i s a b l e t o use an a u t o c o r r e l a t i o n −robust c ov ari anc e matrix .
#13. TESTS FOR CROSS−SECTIONAL DEPENDENCE_Example 1 3 :
# t e s t s f o r cross −s e c t i o n a l dependence − RDSpillovers data s e t
data ( " RDSpillovers " , package = " pder " )
fm . rds <− lny ~ l n l + lnk + lnrd
# Pairwise −c o r r e l a t i o n s −based t e s t s are o r i g i n a l l y
#meant t o use the r e s i d u a l s o f separate estimation
# o f one time−s e r i e s r e g r e s s i o n f o r each cross −s e c t i o n a l unit .
#In the o r i g i n a l example o f Eberhardt e t a l . ( 2 0 1 3 ) ,
# t h i s i s done in the f i r s t o f the heterogeneous s t a t i c s p e c i f i c a t i o n s o f Table 7 ,
#the mean groups model . This i s a l s o the d e f a u l t behavior o f p c d t e s t .
# The d e f a u l t v e r s i o n o f the t e s t i s
’ cd ’ ,
#which i s appropriate in a l a r g e panel s e t t i n g l i k e t h i s
p c d t e s t ( fm . rds , RDSpillovers )
# The r e s i d u a l s from separate time s e r i e s r e g r e s s i o n s show
# strong evidence o f e r r o r cross −s e c t i o n a l dependence .
# I f a d i f f e r e n t model s p e c i f i c a t i o n ( within , re ,
. . . ) i s assumed c o n s i s t e n t ,
#one can r e s o r t t o i t s r e s i d u a l s f o r t e s t i n g 4 by s p e c i f y i n g the r e l e v a n t model type
#Themain argument o f t h i s f u n c t i o n may be e i t h e r plm or a formula and a data . frame ;
# in the second case , unless model i s s e t t o NULL,
# a l l usual parameters r e l a t i v e t o the estimation o f a plm model may be passed on .
#The t e s t i s compatible with any c o n s i s t e n t plm model f o r the data at hand ,
#with any s p e c i f i c a t i o n o f e f f e c t ; e . g . ,
# s p e c i f y i n g e f f e c t = ’ time ’ or e f f e c t = ’ twoways ’ allows t o t e s t f o r
# r e s i d u a l cross −s e c t i o n a l dependence a f t e r the i n t r o d u c t i o n o f
384
Electronic copy available at: https://ssrn.com/abstract=3863563
8.20. ADVANCED ERROR COMPONENT MODELS
#time f i x e d e f f e c t s t o account f o r common shocks .
#Let us c o n s i d e r the s t a t i c two−way f i x e d e f f e c t s s p e c i f i c a t i o n
# in Eberhardt e t a l . ( 2 0 1 3 )
rds . 2 f e <− plm ( fm . rds , RDSpillovers , model = " within " , e f f e c t = " twoways " )
p c d t e s t ( rds . 2 f e )
# As observed , the t e s t l o s e s i t s power i f time f i x e d e f f e c t s are
# included in the model s p e c i f i c a t i o n .
# Now the t e s t does not r e j e c t the hypothesis o f no cross −s e c t i o n a l c o r r e l a t i o n .
# One can get an idea o f what i s happening by comparing rho−hat and rho−hat ( abs )
cbind ( " rho " = p c d t e s t ( rds . 2 fe , t e s t = " rho " ) $ s t a t i s t i c ,
"| rho |"= p c d t e s t ( rds . 2 fe , t e s t = " absrho " ) $ s t a t i s t i c )
# whence i t can be seen how s u b s t a n t i a l cross −s e c t i o n a l dependence i s present ,
#but the a d d i t i o n o f time e f f e c t s has centered the mean o f
# c o r r e l a t i o n c o e f f i c i e n t s on zero so that p o s i t i v e
#and negative rho−hat (nm) s compensate
#15. Example 1 5 : s e c t i o n a l dependence t e s t f o r a p s e r i e s − HousePricesUS data s e t
data ( " HousePricesUS " , package = " pder " )
php <− pdata . frame ( HousePricesUS )
cbind ( " rho " = p c d t e s t ( d i f f ( l o g ( php$price ) ) , t e s t = " rho " ) $ s t a t i s t i c ,
"| rho |" = p c d t e s t ( d i f f ( l o g ( php$price ) ) , t e s t = " absrho " ) $ s t a t i s t i c )
# Text book code does not account f o r missing values !
r e g i o n s . names <− c ( "New Engl " , " Mideast " , " Southeast " , " Great Lks " ,
" Plains " , " Southwest " , " Rocky Mnt" , " Far West " )
c o r r . t a b l e . hp <− c o r t a b ( d i f f ( l o g ( php$price ) ) , grouping = php$region ,
groupnames = r e g i o n s . names )
colnames ( c o r r . t a b l e . hp ) <− substr ( rownames ( c o r r . t a b l e . hp ) , 1 , 5 )
round ( c o r r . t a b l e . hp , 2 )
#Can see the NA or Missing Values here in the t a b l e
385
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# The preliminary s p a t i a l dependence a n a l y s i s
# h i g h l i g h t s the c o r r e l a t i o n between neighboring r e g i o n s
#and a l s o some cases o f c o r r e l a t i o n with d i s t a n t ones ,
#as i s the case f o r C a l i f o r n i a and some more developed s t a t e s on the East Coast .
# According t o the authors ,
# t h i s i s evidence o f f a c t o r −r e l a t e d dependence :
#common shocks t o technology stimulate growth
# in the most advanced s t a t e s i r r e s p e c t i v e o f geographic proximity .
# The s i g n i f i c a n c e o f cross −s e c t i o n a l c o r r e l a t i o n in a p s e r i e s
#can a l s o be assessed through a formal t e s t ,
# e x a c t l y as done above f o r model r e s i d u a l s .
#Given that , beside s t a t i o n a r i t y
# ( which in our example was ensured by f i r s t d i f f e r e n c i n g the data ) ,
#the p r o p e r t i e s o f the CD t e s t r e s t on the hypothesis o f no s e r i a l c o r r e l a t i o n ,
#we f o l l o w Pesaran ( 2 0 0 4 ) ’ s suggestion t o remove any
#by s p e c i f y i n g a u n i v a r i a t e AR( 2 ) model o f the v a r i a b l e o f i n t e r e s t
#and proceed t e s t i n g the r e s i d u a l s o f the l a t t e r f o r cross −s e c t i o n a l dependence .
#This i s made easy by the lagging f u n c t i o n a l i t i e s o f plm .
#In the f o l l o w i n g , we t e s t cross −s e c t i o n a l c o r r e l a t i o n
# in l o g house p r i c e s drawing on the r e s i d u a l s o f an AR( 2 ) model
# in order t o c o n t r o l f o r any p e r s i s t e n c e in the data :
p c d t e s t ( d i f f ( l o g ( p r i c e ) ) ~ d i f f ( lag ( l o g ( p r i c e ) ) ) + d i f f ( lag ( l o g ( p r i c e ) , 2 ) ) ,
data = php )
# The t e s t s t r o n g l y r e j e c t s the n u l l hypothesis ,
# confirming s u b s t a n t i a l cross −s e c t i o n a l co−movement in house p r i c e s .
386
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
8.21
Robust inference & Endogeneity
#ROBUST INFERENCE & ENDOGENEITY_EXERCISE
# 1 . EXAMPLE 1 : Clustered standard e r r o r s f o r pooled models ‚ Ä ì Produc data s e t
#1.1 I n s t a l l and load the l i b r a r y and data
l i b r a r y ( " plm " )
data ( " Produc " , package = " plm " )
fm <− l o g ( gsp ) ~ l o g ( pcap ) + l o g ( pc ) + l o g (emp) + unemp
#1.2 Thefunction c o e f t e s t frompackage lmtest produces a compact c o e f f i c i e n t s
# t a b l e allowing f o r a f l e x i b l e c h o i c e o f the cov ari an ce matrix .
# Calculate a h e t e r o s c e d a s t i c i t y −robust d i a g n o s t i c t a b l e
# f o r two s t a t i s t i c a l l y e q u i v a l e n t models . F i r s t , pooled o l s by lm :
lmmod <− lm ( fm , Produc )
l i b r a r y ( zoo )
l i b r a r y ( lmtest )
l i b r a r y ( sandwich )
c o e f t e s t ( lmmod , vcov = vcovHC )
#1.3 Next , compute pooled o l s by plm .
#The c o e f t e s t f u n c t i o n complies with plm o b j e c t s ,
# so the same syntax as above can be employed .
#In turn , the summary . plm method i s i t s e l f compliant with
# providing a custom c ov ari anc e
# ( a note about using a nonstandard co var ian ce w i l l be issued ) :
plmmod <− plm ( fm , Produc , model = " p o o l i n g " )
summary ( plmmod , vcov = vcovHC )
#NOTES:
# 1 . C o e f f i c i e n t s are o b v i o u s l y the same ,
#but the estimated standard e r r o r s w i l l turn out d i f f e r e n t .
# 2 . In p a r t i c u l a r ,
#the standard e r r o r o f the c o e f f i c i e n t on pcap i s much l a r g e r ,
#and while s t i l l s i g n i f i c a n t at the 5% l e v e l ,
# i t i s not any more at the 1% l e v e l .
387
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 3 . This i s because the c l a s s e s o f the model o b j e c t s t o be t e s t e d are d i f f e r e n t ,
#and so are the d e f a u l t s e t t i n g s o f the vcovHC . lm and vcovHC . plm methods .
# 4 . Only i f one o v e r r i d e s the d e f a u l t s , here ,
# s p e c i f y i n g the method as ’white1’
#and the small sample c o r r e c t i o n as ’HC3’ , the lm r e s u l t s w i l l be r e p l i c a t e d .
# 5 . Therefore , thanks t o o b j e c t o r i e n t a t i o n ,
# i f applying the g e n e r i c robust method vcovHC t o a panelmodel o b j e c t ,
#one w i l l get a r e s u l t that i s l i k e l y t o be ‚ Ä ú s e n s i b l e ‚ Ä ù
# f o r the most common a p p l i c a t i o n s .
# 2 . EXAMPLE 2 : Clustered standard e r r o r s f o r non−panel data ‚ Ä ì Hedonic data s e t
data ( " Hedonic " , package = " plm " )
hfm <− mv ~ crim + zn + indus + chas + nox + rm + age + d i s +
rad + tax + p t r a t i o + blacks + l s t a t
hlmmod <− lm ( hfm , Hedonic )
c o e f t e s t ( hlmmod , vcov = vcovHC )
# and pooled o l s by plm ; then we compare White and c l u s t e r e d standard e r r o r s :
hplmmod <− plm ( hfm , Hedonic , model = " p o o l i n g " , index = " townid " )
sign . tab <− cbind ( c o e f ( hlmmod ) , c o e f t e s t ( hlmmod , vcov = vcovHC ) [ , 4 ] ,
c o e f t e s t ( hplmmod , vcov = vcovHC ) [ , 4 ] )
dimnames ( sign . tab ) [ [ 2 ] ] <− c ( " C o e f f i c i e n t " , " p−values , HC" , " p−val . , c l u s t e r " )
round ( sign . tab , 3 )
# CONCLUSION: Proximity t o the Charles River , average number o f rooms ,
#and the p r o p o r t i o n o f blacks in the population
#are not s i g n i f i c a n t any more a f t e r c l u s t e r i n g by town .
388
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
# 3 . EXAMPLE 3 : Newey−West and double−c l u s t e r i n g est imat ors − Produc data s e t
c o e f t e s t ( plmmod , vcov=vcovSCC )
# or p o s s i b l y , i f allowing f o r double c l u s t e r i n g :
c o e f t e s t ( plmmod , vcov=vcovDC )
# More complicated s t r u c t u r e s allowing f o r two−way c l u s t e r i n g
#and e r r o r p e r s i s t e n c e in the sense o f Thompson ( 2 0 1 1 ) can be obtained
#by combination , as i l l u s t r a t e d above .
#Below , the case o f double c l u s t e r i n g plus f o u r p e r i o d s o f p e r s i s t e n t ( unweighted )
#shocks a l a Thompson ( 2 0 1 1 )
# ( n o t i c e that the weighting f u n c t i o n wj has been defined as the constant 1
#but must s t i l l be a f u n c t i o n o f two arguments ) :
l i b r a r y ( " plm " )
myvcovDCS <− f u n c t i o n ( x , maxlag = NULL,
...) {
w1 <− f u n c t i o n ( j , maxlag ) 1
VsccL . 1 <− vcovSCC ( x , maxlag = maxlag , wj = w1,
...)
Vcx <− vcovHC ( x , c l u s t e r = " group " , method = " a r e l l a n o " , . . . )
VnwL. 1 <− vcovSCC ( x , maxlag = maxlag , inner = " white " , wj = w1,
...)
return ( VsccL . 1 + Vcx − VnwL. 1 )
}
c o e f t e s t ( plmmod , vcov= f u n c t i o n ( x ) myvcovDCS ( x , maxlag = 4 ) )
# 4 . EXAMPLE 4 : computing an array o f standard e r r o r s ‚ Ä ì Produc data s e t
#4.1 Need t o s e t up the v a r i a b l e s from Table 1
Vw <− f u n c t i o n ( x ) vcovHC ( x , method = " white1 " )
Vcx <− f u n c t i o n ( x ) vcovHC ( x , c l u s t e r = " group " , method = " a r e l l a n o " )
Vct <− f u n c t i o n ( x ) vcovHC ( x , c l u s t e r = " time " , method = " a r e l l a n o " )
Vcxt <− f u n c t i o n ( x ) Vcx ( x ) + Vct ( x ) − Vw( x )
Vct . L <− f u n c t i o n ( x ) vcovSCC ( x , wj = f u n c t i o n ( j , maxlag ) 1 )
389
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
Vnw. L <− f u n c t i o n ( x ) vcovNW ( x )
Vscc . L <− f u n c t i o n ( x ) vcovSCC ( x )
Vcxt . L <− f u n c t i o n ( x ) Vct . L ( x ) + Vcx ( x ) − vcovNW ( x , wj = f u n c t i o n ( j , maxlag ) 1 )
#4.2 Then b u i l d up a v e c t o r o f f u n c t i o n s on which t o l o o p :
vcovs <− c ( vcov , Vw, Vcx , Vct , Vcxt , Vct . L , Vnw. L , Vscc . L , Vcxt . L )
names ( vcovs ) <− c ( "OLS" , "Vw" , " Vcx " , " Vct " , " Vcxt " , " Vct . L " , "Vnw. L " ,
" Vscc . L " , " Vcxt . L " )
# in order t o c a l c u l a t e a comprehensive t a b l e o f p−values from robust es tima tor s .
# To t h i s end we d e f i n e a convenience f u n c t i o n :
c f r t a b <− f u n c t i o n ( mod, vcovs ,
...) {
c f r t a b <− matrix ( nrow = length ( c o e f (mod ) ) , n c o l = 1 + length ( vcovs ) )
dimnames ( c f r t a b ) <− l i s t ( names ( c o e f (mod ) ) ,
c ( " C o e f f i c i e n t " , paste ( " s . e . " , names ( vcovs ) ) ) )
c f r t a b [ , 1 ] <− c o e f (mod)
f o r ( i in 1 : length ( vcovs ) ) {
myvcov = vcovs [ [ i ] ]
c f r t a b [ , 1 + i ] <− s q r t ( diag ( myvcov (mod ) ) )
}
return ( t ( round ( c f r t a b , 4 ) ) )
}
# The a d d i t i v e nature o f the three b a s i c components V(WH) , V(CX) ,
#and V(CT) allows the res ea rc he r t o i n f e r on the r e l a t i v e importance o f
#each c l u s t e r i n g dimension by l o o k i n g at the c o n t r i b u t i o n o f each
# t o the standard e r r o r estimate , so that i f , e . g . ,
#V(CX) <= V(CT) approximates V(CXT)
#then t h i s i s evidence o f important cross −s e c t i o n a l c o r r e l a t i o n ( Petersen , 2 0 0 9 ) .
c f r t a b ( plmmod , vcovs )
# For t h i s pooled OLS model ,
#standard e r r o r s estimates assuming group c l u s t e r i n g are c o n s i s t e n t l y l a r g e r
# that the r e s t , i n c l u d i n g Newey−West and SCC,
390
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
# p o i n t i n g at non−decaying s e r i a l e r r o r dependence .
# 5 . EXAMPLE 5 : random e f f e c t s and robust c o v a r i a n c e s ‚ Ä ì Produc data
replmmod <− plm ( fm , Produc )
c f r t a b ( replmmod , vcovs )
# The cross −s e c t i o n a l dependence component becomes r e l a t i v e l y
#more important when accounting f o r time p e r s i s t e n c e
# in the model through random i n d i v i d u a l ( country ) e f f e c t s
# 6 . EXAMPLE 6 : time f i x e d e f f e c t s model − agl data s e t
# F i r s t we need t o i n s t a l l the pcse package
i n s t a l l . packages ( " pcse " )
l i b r a r y ( " pcse " )
data ( " agl " , package = " pcse " )
# In the f o l l o w i n g we estimate the model with time f i x e d e f f e c t s
#and produce the d i a g n o s t i c s t a b l e with pcse standard e r r o r s :
fm <− growth ~ lagg1 + opengdp + openex + openimp + c e n t r a l * l e f t c
aglmod <− plm ( fm , agl , model = "w" , e f f e c t = " time " )
c o e f t e s t ( aglmod , vcov=vcovBK )
# 7 . EXAMPLE 7 : t e s t i n g with robust co var ia nce matrices − Produc data s e t
c o e f t e s t ( plmmod , vcov = vcovHC ( plmmod , type = "HC3 " ) )
# or , rather , d e f i n e an appropriate f u n c t i o n i n s i d e the c a l l :
# in t h i s case , o p t i o n a l parameters are provided as shown below
# ( see a l s o Z e i l e i s , 2004 , p . 1 2 ) :
c o e f t e s t ( plmmod , vcov = f u n c t i o n ( x ) vcovHC ( x , type = "HC3 " ) )
391
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# For some t e s t s , e . g . , f o r mul tiple model comparisons by waldtest ,
#one should always provide a f u n c t i o n .
# 8 . EXAMPLE 8 : t e s t i n g with robust co va ria nce matrices ‚ Ä ì Parity data s e t
data ( " Parity " , package = " plm " )
fm <− l s ~ l d
pppmod <− plm ( fm , data = Parity , e f f e c t = " twoways " )
# The hypothesis o f i n t e r e s t i s Beta = 1 ,
#meaning that i n f l a t i o n d i f f e r e n t i a l s are f u l l y r e f l e c t e d in the exchange r a t e .
#We r e p o r t the corresponding robustWald t e s t from linearHypothesis
# in package car ( Fox andWeisberg , 2011) ,
#which would be done i n t e r a c t i v e l y as f o l l o w s :
l i b r a r y ( carData )
l i b r a r y ( car )
linearHypothesis ( pppmod , " l d = 1 " , vcov = vcov )
# ( output suppressed ) ,
# in a compact t a b l e supplying d i f f e r e n t cov ari anc e e sti mato rs
# t o each o f f o u r models : o l s , one−way time or country f i x e d e f f e c t s ,
#and two−way f i x e d e f f e c t
vcovs <− c ( vcov , Vw, Vcx , Vct , Vcxt , Vct . L , Vnw. L , Vscc . L , Vcxt . L )
names ( vcovs ) <− c ( "OLS" , "Vw" , " Vcx " , " Vct " , " Vcxt " , " Vct . L " , "Vnw. L " ,
" Vscc . L " , " Vcxt . L " )
t t t a b <− matrix ( nrow = 4 , n c o l = length ( vcovs ) )
dimnames ( t t t a b ) <− l i s t ( c ( " Pooled OLS" , " Time FE" , " Country FE" , " Two−way FE" ) ,
names ( vcovs ) )
pppmod . o l s <− plm ( fm , data = Parity , model = " p o o l i n g " )
392
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
f o r ( i in 1 : length ( vcovs ) ) {
t t t a b [ 1 , i ] <− linearHypothesis ( pppmod . o l s , " l d = 1 " ,
vcov = vcovs [ [ i ] ] ) [ 2 , 4 ]
}
pppmod . t f e <− plm ( fm , data = Parity , e f f e c t = " time " )
f o r ( i in 1 : length ( vcovs ) ) {
t t t a b [ 2 , i ] <− linearHypothesis ( pppmod . t f e , " l d = 1 " ,
vcov = vcovs [ [ i ] ] ) [ 2 , 4 ]
}
pppmod . c f e <− plm ( fm , data = Parity , e f f e c t = " i n d i v i d u a l " )
f o r ( i in 1 : length ( vcovs ) ) {
t t t a b [ 3 , i ] <− linearHypothesis ( pppmod . c f e , " l d = 1 " ,
vcov = vcovs [ [ i ] ] ) [ 2 , 4 ]
}
pppmod . 2 f e <− plm ( fm , data = Parity , e f f e c t = " twoways " )
f o r ( i in 1 : length ( vcovs ) ) {
t t t a b [ 4 , i ] <− linearHypothesis ( pppmod . 2 fe , " l d = 1 " ,
vcov = vcovs [ [ i ] ] ) [ 2 , 4 ]
}
p r i n t ( t ( round ( t t t a b , 6 ) ) )
# Conclusion : As i s apparent from the r e s u l t s ‚ Ä ô table ,
#the ppp hypothesis i s not r e j e c t e d any more once one c o n t r o l s f o r ,
# at a minimum, country f i x e d e f f e c t s and by−group c l u s t e r i n g
# 9 . EXAMPLE 9 : r e g r e s s i o n −based Hausman t e s t ‚ Ä ì Grunfeld data s e t
data ( " Grunfeld " , package = " plm " )
# Hausman Test
phtest ( inv ~ value + c a p i t a l , data = Grunfeld )
393
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# Now f o r the r e g r e s s i o n −based Hausman Test
phtest ( inv ~ value + c a p i t a l , data = Grunfeld , method = " aux " )
# Unsurprisingly ,
#the r e s u l t s from the r e g r e s s i o n −based
#and the o r i g i n a l Hausman t e s t are c o n s i s t e n t :
#both support the random e f f e c t s (RE) hypothesis .
## EXAMPLE 1 0 : robust Hausman t e s t ‚ Ä ì RDSpillovers data s e t
data ( " RDSpillovers " , package = " pder " )
pehs <− pdata . frame ( RDSpillovers , index = c ( " i d " , " year " ) )
ehsfm <− lny ~ l n l + lnk + lnrd
phtest ( ehsfm , pehs , method = " aux " )
phtest ( ehsfm , pehs , method = " aux " , vcov = vcovHC )
# CONCLUSION: The robust v e r s i o n o f the Hausman t e s t
#does not r e j e c t the random e f f e c t s hypothesis any more .
### UNRESTRICTED GENERALIZED LEAST SQUARES
#11. EXAMPLE 1 1 : g e n e r a l i z e d g l s estimator ‚ Ä ì EmplUK data s e t
# The ‚Äúrandom e f f e c t ‚ Ä ù equivalent , general GLS,
# i s estimated by s p e c i f y i n g the model argument as ‚ Ä ô p o o l i n g ‚ Ä ô :
data ( "EmplUK" , package = " plm " )
gglsmod <− pggls ( l o g (emp) ~ l o g ( wage ) + l o g ( c a p i t a l ) ,
data = EmplUK, model = " p o o l i n g " )
394
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
summary ( gglsmod )
# The pggls f u n c t i o n i s s i m i l a r t o plm in many r e s p e c t s .
#An e x c e p t i o n i s that the estimate o f the group cov ari an ce matrix
# o f e r r o r s ( sigma , a matrix ) i s reported in the model o b j e c t s instead o f
#the usual estimated variances o f the two e r r o r components .
# I t can be di splay ed as f o l l o w s :
round ( gglsmod$sigma , 3 )
# As can be seen , the c o r r e l a t i o n s between p a i r s o f r e s i d u a l s ( in time )
# f o r the same i n d i v i d u a l do not d i e out with the d i s t a n c e in time .
# The estimated e r r o r c ov ari anc e very much resembles
#the random e f f e c t s s t r u c t u r e ,
#with a strong prevalence o f the i n d i v i d u a l variance component sigma−squared eta
# over sigma−squared nu
# ( witness the small d i f f e r e n c e between values on and o u t s i d e the diagonal ) .
#12. EXAMPLE 1 2 : FEGLS estimator − EmplUK data s e t
feglsmod <− pggls ( l o g (emp) ~ l o g ( wage ) + l o g ( c a p i t a l ) , data = EmplUK,
model = " within " )
summary ( feglsmod )
# The phtest f u n c t i o n can be used t o assess the need f o r
# f i x e d e f f e c t s through a Hausman t e s t :
phtest ( feglsmod , gglsmod )
# CONCLUSION: The Hausman t e s t s t r o n g l y favours the f i x e d e f f e c t s model .
#13. EXAMPLE 1 3 : FDGLS estimator − EmplUK data s e t
fdglsmod <− pggls ( l o g (emp) ~ l o g ( wage ) + l o g ( c a p i t a l ) , data = EmplUK,
model = " fd " )
395
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
summary ( fdglsmod )
#14. EXAMPLE 1 4 : g e n e r a l i z e d GLS estimator − RiceFarms data s e t
data ( " RiceFarms " , package = " splm " )
RiceFarms <− transform ( RiceFarms ,
phosphate = phosphate / 1000 ,
p e s t i c i d e = as . numeric ( p e s t i c i d e > 0 ) )
fm <− l o g ( goutput ) ~ l o g ( seed ) + l o g ( urea ) + phosphate +
log ( totlabor ) + log ( size ) + pesticide + varieties +
+ region + time
gglsmodrice <− pggls ( fm , RiceFarms , model = " p o o l i n g " , index = " i d " )
# Ignore ’ time ’ column being ove rw rit te n message !
summary ( gglsmodrice )
# Regions do not seem t o be so important a f t e r a l l ,
# only Ciwangi being s i g n i f i c a n t l y d i f f e r e n t from the b a s e l i n e ;
#although a j o i n t r e s t r i c t i o n t e s t s t i l l r e j e c t s :
l i b r a r y ( " lmtest " )
waldtest ( gglsmodrice , " region " )
Model1 : l o g ( goutput ) ~ l o g ( seed ) +
l o g ( urea ) +
phosphate +
log ( totlabor ) +
log ( size ) +
pesticide +
varieties +
region +
time
396
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
Model2 : l o g ( goutput ) ~ l o g ( seed ) +
l o g ( urea ) +
phosphate +
log ( totlabor ) +
log ( size ) +
pesticide +
varieties +
time
f e g l s m o d r i c e <− pggls ( update ( fm , . ~ . − region ) , RiceFarms , index = " i d " )
# Q u a l i t a t i v e l y , the r e s u l t s do not seem t o change much
#when adding i n d i v i d u a l f i x e d e f f e c t s .
# The hypothesis that a f t e r c o n t r o l l i n g f o r the region ,
# a l l remaining i n d i v i d u a l h e t e r o g e n e i t y be o f the random e f f e c t s type
#can be t e s t e d f o r m a l l y by means o f a Hausman t e s t
# Hausman Test
phtest ( gglsmodrice , f e g l s m o d r i c e )
# The Hausman t e s t does in f a c t not r e j e c t .
# Given the low s i g n i f i c a n c e o f the r e g i o n a l e f f e c t s ,
#one might wonder whether a f u l l ‚Äúrandom e f f e c t s ‚ Ä ù s p e c i f i c a t i o n can be j u s t i f i e d .
# An updated GGLS s p e c i f i c a t i o n can r e a d i l y
#be compared t o the already estimated FEGLS model :
phtest ( pggls ( update ( fm , . ~ . − region ) , RiceFarms ,
model = " p o o l i n g " , index = " i d " ) , f e g l s m o d r i c e )
# Ignore warning message , t r y <CTRL>+<EMTER> a second time
# i f i t does not output t a b l e the f i r s t time !
# In f a c t , even omitting the r e g i o n a l f i x e d e f f e c t s ,
#the GGLS s p e c i f i c a t i o n s t i l l passes the Hausman t e s t .
# The 171 r i c e farms can a c t u a l l y be seen as random
#draws from the same population ,
#without the need f o r e i t h e r i n d i v i d u a l or r e g i o n a l f i x e d e f f e c t s .
397
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
## EXAMPLE 1 5 : g e n e r a l i z e d g l s estimator − RDSpillovers data s e t
fm <− lny ~ l n l + lnk + lnrd
gglsmodehs <− pggls ( fm , RDSpillovers , model = " p o o l i n g " )
c o e f t e s t ( gglsmodehs )
feglsmodehs <− pggls ( fm , RDSpillovers , model = " within " )
c o e f t e s t ( feglsmodehs )
phtest ( gglsmodehs , feglsmodehs )
# The Hausman t e s t r e j e c t s the ‚Äúrandom e f f e c t s ‚ Ä ù GGLS s p e c i f i c a t i o n .
# Given that c o r r e l a t e d h e t e r o g e n e i t y seems t o be present ,
#an a l t e r n a t i v e t o e l i m i n a t e i t i s the f i r s t d i f f e r e n c e transformation :
fdglsmodehs <− pggls ( fm , RDSpillovers , model = " fd " )
# Which one t o choose between FEGLS
#and FDGLS depends on the p r o p e r t i e s o f transformed r e s i d u a l s .
# FEGLS r e s i d u a l s show a high l e v e l o f p e r s i s t e n c e ,
#as a simple s e r i a l c o r r e l a t i o n t e s t ( Wooldridge , 2010 , 1 0 . 6 . 3 ) shows .
# Make a data . frame o f the r e s i d u a l s ,
#then estimate a ( pooled ) a u t o r e g r e s s i v e model :
f e e <− r e s i d ( feglsmodehs )
dbfee <− data . frame ( f e e =fee , i d = a t t r ( fee , " index " ) [ [ 1 ] ] )
c o e f t e s t ( plm ( f e e ~lag ( f e e )+ lag ( fee , 2 ) , dbfee , model = " p " , index =" i d " ) )
# The FEGLS r e s i d u a l s seem c l o s e t o being non−s t a t i o n a r y .
#The estimated a u t o c o r r e l a t i o n in FDGLS r e s i d u a l s i s instead much lower
fde <− r e s i d ( fdglsmodehs )
dbfde <− data . frame ( fde=fde , i d = a t t r ( fde , " index " ) [ [ 1 ] ] )
c o e f t e s t ( plm ( fde~lag ( fde )+ lag ( fde , 2 ) , dbfde , model = " p " , index =" i d " ) )
398
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
# hence i t i s a d v i s a b l e t o r e s o r t t o the FDGLS estimator :
c o e f t e s t ( fdglsmodehs )
# The r e s u l t ,
# d e s p i t e the l i m i t e d number o f degrees o f freedom in estimating Sigma ( T ) ,
# i s in l i n e with the more s o p h i s t i c a t e d analyses in the o r i g i n a l paper
#by Eberhardt e t a l . ( 2 0 1 3 ) and with the p r e f e r r e d FD s p e c i f i c a t i o n in Table .
# Moreover , d e s p i t e the expected downward bias ,
#standard e r r o r s are not t o o f a r from those o f the above−mentioned FD model .
### ENDOGENEITY ###
## EXAMPLE 1 6 : within 2SLS estimator ‚ Ä ì SeatBelt data s e t
data ( " SeatBelt " , package = " pder " )
S e a t B e l t $ o c c f a t <− with ( SeatBelt , l o g ( f a r s o c c / ( vmtrural + vmturban ) ) )
o l s <− plm ( o c c f a t ~ l o g ( usage ) + l o g ( percapin ) + l o g ( unemp ) + l o g ( meanage ) +
l o g ( precentb ) + l o g ( precenth )+ l o g ( densrur ) +
l o g ( densurb ) + l o g ( viopcap ) + l o g ( proppcap ) +
l o g ( vmtrural ) + l o g ( vmturban ) + l o g ( f u e l t a x ) +
lim65 + lim70p + mlda21 + bac08 , SeatBelt ,
e f f e c t = " time " )
f e <− update ( o l s , e f f e c t = " twoways " )
i v f e <− update ( fe , . ~ . | . − l o g ( usage ) + ds + dp +dsp )
rbind ( o l s = c o e f ( summary ( o l s ) ) [ 1 , ] ,
f e = c o e f ( summary ( f e ) ) [ 1 , ] ,
w2sls = c o e f ( summary ( i v f e ) ) [ 1 , ] )
# The r e s u l t s confirm that the endogeneity problem i s very important .
# For the f i r s t f i t t e d model ,
#the seat b e l t −use c o e f f i c i e n t i s s i g n i f i c a n t l y p o s i t i v e .
# I t becomes s i g n i f i c a n t l y negative f o r the f i x e d e f f e c t s model ,
399
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#which means that usage i s s t r o n g l y c o r r e l a t e d with the i n d i v i d u a l e f f e c t s .
# Finally , t h i s c o e f f i c i e n t i n c r e a s e s importantly ( in a b s o l u t e value )
# i f instrumental v a r i a b l e s are used ,
#which i n d i c a t e s that the i d i o s y n c r a t i c e r r o r i s a l s o c o r r e l a t e d with usage .
# In order t o t e s t the behavior compensation theory ,
#the authors estimate the samemodels ,
# t h i s time using the number o f non−occupants k i l l e d ( n o c c f a t ) as response .
S e a t B e l t $ n o c c f a t <− with ( SeatBelt , l o g ( f a r s n o c c / ( vmtrural + vmturban ) ) )
n i v f e <− update ( i v f e , n o c c f a t ~ . | . )
c o e f ( summary ( n i v f e ) ) [ 1 , ]
# CONCLUSION: The r e s u l t s i n d i c a t e that
# seat b e l t use has no i n f l u e n c e on out−o f − v e h i c l e mortality ,
# in c o n t r a d i c t i o n with Peltzman ( 1 9 7 5 ) ‚ Ä ô s theory o f behavior compensation .
## ENDOGENEITY
# EXAMPLE 1 7 : EC2SLS estimator ‚ Ä ì ForeignTrade data s e t
data ( " ForeignTrade " , package = " pder " )
w1 <− plm ( imports~pmcpi + gnp + lag ( imports ) + lag ( resimp ) |
lag ( consump ) + lag ( c p i ) + lag ( income ) + lag ( gnp ) + pm +
lag ( i n v e s t ) + lag ( money ) + gnpw + pw + lag ( r e s e r v e s ) +
lag ( exp orts ) + trend + pgnp + lag ( px ) ,
ForeignTrade , model = " within " )
r1 <− update ( w1, model = " random " , random . method = " nerlov e " ,
random . d f c o r = c ( 1 , 1 ) , i n s t . method = " b a l t a g i " )
# The hypothesis o f no c o r r e l a t i o n between the instruments
#and the i n d i v i d u a l e f f e c t s i m p l i e s that the within
#and the g l s models are c o n s i s t e n t ,
#the l a t t e r being more e f f i c i e n t .
#On the contrary , i f t h i s hypothesis i s r e j e c t e d ,
400
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
# only the within model i s c o n s i s t e n t .
# In order t o t e s t t h i s hypothesis the authors used the Hausman ( 1 9 7 8 ) t e s t :
phtest ( r1 , w1)
# The hypothesis o f no c o r r e l a t i o n between the instruments
#and the i n d i v i d u a l e f f e c t s i s r e j e c t e d at the 5% thre sho ld
# Kinal and Lahiri ( 1 9 9 3 ) f i n a l l y got the f o l l o w i n g s p e c i f i c a t i o n :
r1b <− plm ( imports ~ pmcpi + gnp + lag ( imports ) + lag ( resimp ) |
lag ( consump ) + lag ( c p i ) + lag ( income ) + lag ( px ) +
lag ( r e s e r v e s ) + lag ( exp orts ) | lag ( gnp ) + pm +
lag ( i n v e s t ) + lag ( money ) + gnpw + pw + trend + pgnp ,
ForeignTrade , model = " random " , i n s t . method = " b a l t a g i " ,
random . method = " n erlove " , random . d f c o r = c ( 1 , 1 ) )
phtest (w1, r1b )
# Based on the Hausman ( 1 9 7 8 ) t e s t ,
#the hypothesis o f c o n s i s t e n c y o f the GLS estimator i s no longer r e j e c t e d .
# Results are presented below ;
#the within and g l s e stim ato rs g i v e very s i m i l a r r e s u l t s .
rbind ( within = c o e f (w1 ) , e c 2 s l s = c o e f ( r1b ) [ − 1 ] )
# The short −term e l a s t i c i t y o f imports demand
# i s d i r e c t l y given by the p r i c e c o e f f i c i e n t .
# The long −term e l a s t i c i t y i s obtained
#by d i v i d i n g t h i s c o e f f i c i e n t
#by one minus the c o e f f i c i e n t o f the lagged response .
# Then have :
e l a s t <− sapply ( l i s t ( w1, r1 , r1b ) ,
f u n c t i o n ( x ) c ( c o e f ( x ) [ " pmcpi " ] ,
c o e f ( x ) [ " pmcpi " ] / ( 1 − c o e f ( x ) [ " lag ( imports ) " ] ) ) )
dimnames ( e l a s t ) <− l i s t ( c ( " ST " , "LT " ) , c ( " w1" , " r1 " , " r1b " ) )
elast
401
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# The use o f t h i s GLS estimator ,
#which e f f i c i e n t l y e x p l o i t s part o f the i n t e r −i n d i v i d u a l v a r i a t i o n ,
#has d r a m a t i c a l l y reduced the standard d e v i a t i o n s o f the c o e f f i c i e n t s
rbind ( within = c o e f ( summary (w1 ) ) [ , 2 ] ,
e c 2 s l s = c o e f ( summary ( r1b ) ) [ − 1 , 2 ] )
#18. EXAMPLE 1 8 : Hausman−Taylor estimator ‚ Ä ì TradeEU data s e t
data ( " TradeEU " , package = " pder " )
# Following the authors , we f i r s t estimate the OLS and the within model :
o l s <− plm ( trade ~ gdp + d i s t + r e r + r l f + sim + cee + emu + bor + lan , TradeEU ,
model = " p o o l i n g " , index = c ( " p a i r " , " year " ) )
f e <− update ( o l s , model = " within " )
fe
# As expected , c o e f f i c i e n t s a s s o c i a t e d t o d i s t , bor ,
#and lan are not estimated in the within model ,
#as these c o v a r i a t e s disappear with the within transformation .
#On the contrary ,
#the random e f f e c t s estimator produces estimates f o r t h e i r c o e f f i c i e n t s .
re <− update ( fe , model = " random " )
re
# The r e s u l t s o f the random e f f e c t s model
# i n d i c a t e a d i s t a n c e e l a s t i c i t y o f b i l a t e r a l trade o f about ‚ à í 0 . 6
#and that having a common border
# or a common language have a s i m i l a r e f f e c t ( an i n c r e a s e o f about 44%).
phtest ( re , f e )
# With the Hausman t e s t ,
#we r e j e c t the hypothesis o f no c o r r e l a t i o n at the 5% thre sho ld .
# Serlenga and Shin ( 2 0 0 7 ) c o n s i d e r that ,
#among the time−i n v a r i a n t v a r i a b l e s ,
# only lan i s c o r r e l a t e d with the i n d i v i d u a l e f f e c t s .
402
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
# Two Hausman and Taylor ( 1 9 8 1 ) models are then estimated .
# In the f i r s t one ,
#the only doubly exogenous v a r i a b l e i s the r e a l exchange r a t e r e r .
# In t h i s case , the instrumental v a r i a b l e s estimator i s j u s t i d e n t i f i e d ,
#as there i s only one instrument ( the between transformation o f r e r )
#and only one endogenous v a r i a b l e lan .
# In the second one , domestic product gdp
#and r e l a t i v e f a c t o r endowment r l f are a l s o used as instruments .
ht1 <− plm ( trade ~ gdp + d i s t + r e r + r l f + sim + cee + emu + bor + lan |
r e r + d i s t + bor | gdp + r l f + sim + cee + emu + lan ,
data = TradeEU , model = " random " , index = c ( " p a i r " , " year " ) ,
i n s t . method = " b a l t a g i " , random . method = " ht " )
ht2 <− update ( ht1 , trade ~ gdp + d i s t + r e r + r l f + sim + cee + emu + bor + lan |
r e r + gdp + r l f + d i s t + bor| sim + cee + emu + lan )
# Note than random . method i s s e t t o ‚ Ä ô h t ‚ Ä ô
# so that the within r e s i d u a l s used t o compute the variance o f the components o f
#the e r r o r are purged o f the i n f l u e n c e o f the time−i n v a r i a n t c o v a r i a t e s .
#The c o n s i s t e n c y o f e i t h e r s p e c i f i c a t i o n i s not r e j e c t e d by the Hausman t e s t
phtest ( ht1 , f e )
phtest ( ht2 , f e )
# The l a s t estimated model i s suggested by B a l t a g i ( 2 0 1 2 ) .
# I t i s s i m i l a r t o the second s p e c i f i c a t i o n but uses the instruments suggested
#by Amemiya and MaCurdy ( 1 9 8 6 ) instead .
# The r e s u l t s are presented in t a b l e ( see s l i d e s l a t e r )
#by using the texreg package ( see L e i f e l d , 2 0 1 3 ) .
ht2am <− update ( ht2 , i n s t . method = "am" )
l i b r a r y ( " texreg " )
texreg ( l i s t ( o l s , fe , re , ht1 , ht2 , ht2am ) ,
custom . model . names = c ( "OLS" , "FE" , "RE" , "HT1" , "HT2" , "AM2" ) ,
c a p t i o n = " Estimations o f the g r a v i t y model . " , l a b e l = " t a b l e : g r a v i t y " ,
403
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
custom . g o f . names = c ( " R$ ÃÇ 2$ " , " Adj . R$ ÃÇ 2$ " , "Num. obs . " , " s\\ _ i d i o s " ,
" s\\_ i d " ) ,
s c r i p t s i z e = FALSE)
# The r e s u l t s o f t a b l e show f i r s t
# that the c o e f f i c i e n t s o f the time−varying c o v a r i a t e s are i d e n t i c a l
# f o r the within and the j u s t i d e n t i f i e d Hausman and Taylor ( 1 9 8 1 ) estimator .
#This i s not the case with the ht2 model ,
#which i s o v e r i d e n t i f i e d , as noted by B a l t a g i ( 2 0 1 2 ) .
# Serlenga and Shin ( 2 0 0 7 ) i n s i s t on the f a c t
# that the Hausman and Taylor ( 1 9 8 1 ) e s t i m a t i o n s lead t o a great r e d u c t i o n o f
#the i n f l u e n c e o f the d i s t a n c e and an important i n c r e a s e o f the i n f l u e n c e o f
#common language and common border .
# This l a s t c o n c l u s i o n i s q u a l i f i e d by B a l t a g i ( 2 0 1 2 ) ,
#which uses the more e f f i c i e n t Amemiya and MaCurdy ( 1 9 8 6 ) estimator .
# The l a t t e r i n t r o d u c e s f u r t h e r o r t h o g o n a l i t y c o n d i t i o n s
#by imposing that doubly exogenous v a r i a b l e s be u n c o r r e l a t e d with
# i n d i v i d u a l e f f e c t s at any time ,
# while the Hausman and Taylor ( 1 9 8 1 ) estimator simply r e q u i r e s
#no c o r r e l a t i o n between i n d i v i d u a l e f f e c t s and the averages o f said v a r i a b l e s .
# I f these c o n d i t i o n s are v a l i d
# ( which can be t e s t e d through the Hausman procedure ) ,
# t h i s estimator i s n e c e s s a r i l y not l e s s e f f i c i e n t than that
# o f Hausman and Taylor ( 1 9 8 1 ) .
phtest ( ht2am , f e )
# The v a l i d i t y o f the supplementary instruments
#used f o r the Amemiya and MaCurdy ( 1 9 8 6 ) estimator
# i s not r e j e c t e d by the Hausman t e s t .
# The standard d e v i a t i o n o f the endogenous v a r i a b l e ( lan ) i s much lower
#than in the Hausman and Taylor ( 1 9 8 1 ) estimator ( 0 . 2 4 vs 0 . 6 8 ) .
# The c o e f f i c i e n t s o f the three time−i n v a r i a n t c o v a r i a t e s are c l o s e r t o
404
Electronic copy available at: https://ssrn.com/abstract=3863563
8.21. ROBUST INFERENCE & ENDOGENEITY
#the OLS c o e f f i c i e n t s than t o the Hausman and Taylor ( 1 9 8 1 ) c o e f f i c i e n t s
## END OF EXAMPLE ##
#19. EXAMPLE 1 9 : e r r o r components 3 s l s ‚ Ä ì ForeignTrade data s e t
eqimp <− imports ~ pmcpi + gnp + lag ( imports ) +
lag ( resimp ) | lag ( consump ) + lag ( c p i ) + lag ( income ) +
lag ( px ) + lag ( r e s e r v e s ) + lag ( exp orts ) | lag ( gnp ) + pm +
lag ( i n v e s t ) + lag ( money ) + gnpw + pw + trend + pgnp
eqexp <− expo rts ~ pxpw + gnpw + lag ( exp orts ) |
lag ( gnp ) + pw + lag ( consump ) + pm + lag ( px ) + lag ( c p i ) |
lag ( money ) + gnpw + pgnp + pop + lag ( i n v e s t ) +
lag ( income ) + lag ( r e s e r v e s ) + exrate
r12 <− plm ( l i s t ( import . demand = eqimp ,
export . demand = eqexp ) ,
data = ForeignTrade , index = 31 , model = " random " ,
i n s t . method = " b a l t a g i " , random . method = " nerlo ve " ,
random . d f c o r = c ( 1 , 1 ) )
summary ( r12 )
# The c o e f f i c i e n t s f o r the imports demand equation
#are very c l o s e t o those we obtained using the 2 s l s estimator .
# The c o r r e l a t i o n between the two components o f the e r r o r s o f
#the two equations i s about 10%.
# Taking i n t o account t h i s c o r r e l a t i o n s l i g h t l y reduces
#the standard e r r o r s o f the c o e f f i c i e n t s , as i l l u s t r a t e d below .
rbind ( e c 2 s l s = c o e f ( summary ( r1b ) ) [ − 1 , 2 ] ,
e c 3 s l s = c o e f ( summary ( r12 ) , " import . demand")[ − 1 , 2 ] )
# END OF EXAMPLES #
405
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.22
Dynamic models & Panel time series
# MODELS & PANEL TIME SERIES_EXERCISE
# 1 . ESTIMATION OF DYNAMIC PANELS_EXAMPLE 1 :
# d e s c r i p t i o n o f data − DemocracyIncome data s e t
#1.2 Ensure that you have package pder i n s t a l l e d on your computer
data ( " DemocracyIncome " , package = " pder " )
data ( " DemocracyIncome25 " , package = " pder " )
l i b r a r y ( " plm " )
#1.3 Can examine the dimension o f the data
pdim ( DemocracyIncome )
#1.4 Can then q u i c k l y look at f i r s t f o u r rows o f data
head ( DemocracyIncome , 4 )
#NOTES:
# 1 . The f i v e −year data c o n s t i t u t e s a balanced panel o f
#211 c o u n t r i e s observed over 11 p e r i o d s .
# 2 . However , such balance i s a r t i f i c i a l because
#many o b s e r v a t i o n s are a c t u a l l y missing , in p a r t i c u l a r as regards democracy l e v e l s .
# 3 . The data comprise the two i n d i v i d u a l and time indexes ( country and year ) ,
#the democracy index democracy ,
#the l o g o f per c a p i t a gross domestic product income ,
#and l a s t l y ,
#an i n d i c a t o r allowing t o s e l e c t the subset considered by the authors sample .
# 2 . EXAMPLE 2 : time e f f e c t s within model − DemocracyIncome data s e t
o l s <− plm ( democracy ~ lag ( democracy ) + lag ( income ) + year − 1 ,
DemocracyIncome , index = c ( " country " , " year " ) ,
model = " p o o l i n g " , subset = sample == 1 )
#2.1 The same model may be estimated by
# s e t t i n g the model t o ‚ Ä ô w i t h i n ‚ Ä ô and the e f f e c t t o ‚Äôtime‚Äô :
o l s <− plm ( democracy ~ lag ( democracy ) + lag ( income ) ,
406
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
DemocracyIncome , index = c ( " country " , " year " ) ,
model = " within " , e f f e c t = " time " ,
subset = sample == 1 )
#2.2 Can now take a look at the c o e f f i c i e n t s
c o e f ( summary ( o l s ) )
#NOTES:
# 1 . This f i r s t model h i g h l i g h t s two r e s u l t s .
# 2 . On one hand , the democracy v a r i a b l e shows high p e r s i s t e n c e ,
#with a c o e f f i c i e n t o f 0 . 7 1 .
# 3 . However , the o l s estimator s u f f e r s from a p o s i t i v e b i a s .
# 4 . On the other hand ,
# lagged income seems t o e x e r t a s i g n i f i c a n t l y p o s i t i v e
# i n f l u e n c e on the democracy index .
# 3 . EXAMPLE 3 : two−ways within Model − DemocracyIncome data s e t
within <− update ( o l s , e f f e c t = " twoways " )
c o e f ( summary ( within ) )
# With r e s p e c t t o the OLS model ,
#the a u t o r e g r e s s i v e c o e f f i c i e n t i s smaller ( 0 . 3 8 vs . 0 . 7 1 ) ,
#which was t o be expected as the within estimator i s biased downward
# while OLS i s biased upward .
# Notice a l s o that a f t e r i n t r o d u c i n g i n d i v i d u a l e f f e c t s ,
#the c o e f f i c i e n t o f income i s very c l o s e t o 0 and not s i g n i f i c a n t any more
# 3 . EXAMPLE 3 : two−ways within Model − DemocracyIncome data s e t
within <− update ( o l s , e f f e c t = " twoways " )
c o e f ( summary ( within ) )
# With r e s p e c t t o the OLS model ,
#the a u t o r e g r e s s i v e c o e f f i c i e n t i s smaller ( 0 . 3 8 vs . 0 . 7 1 ) ,
#which was t o be expected as the within estimator i s biased downward
# while OLS i s biased upward .
# Notice a l s o that a f t e r i n t r o d u c i n g i n d i v i d u a l e f f e c t s ,
#the c o e f f i c i e n t o f income i s very c l o s e t o 0 and not s i g n i f i c a n t any more
407
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 4 . EXAMPLE 4 : Anderson and Hsiao estimator ‚ Ä ì DemocracyIncome data s e t
ahsiao <− plm ( d i f f ( democracy ) ~ lag ( d i f f ( democracy ) ) +
lag ( d i f f ( income ) ) + year − 1 |
lag ( democracy , 2 ) + lag ( income , 2 ) + year − 1 ,
DemocracyIncome , index = c ( " country " , " year " ) ,
model = " p o o l i n g " , subset = sample == 1 )
c o e f ( summary ( ahsiao ) ) [ 1 : 2 , ]
# Anderson and Hsiao ( 1 9 8 2 ) ’ s model being c o n s i s t e n t ,
#one exp ects the estimated a u t o r e g r e s s i v e c o e f f i c i e n t t o be comprised
#between that o f the within model ( biased downward )
#and that o f the OLS model ( biased upward ) .
# This i s a c t u a l l y the case here ,
#the obtained value o f 0.47 f a l l i n g between 0.38 and 0.71
# 5 . EXAMPLE 5 : d i f f e r e n c e gmm estimator − DemocracyIncome data s e t
# F i r s t compute the one−step estimator :
d i f f 1 <− pgmm( democracy ~ lag ( democracy ) + lag ( income ) |
lag ( democracy , 2 : 9 9 ) | lag ( income , 2 ) ,
DemocracyIncome , index=c ( " country " , " year " ) ,
model =" onestep " , e f f e c t ="twoways " , subset = sample == 1 )
c o e f ( summary ( d i f f 1 ) )
# The two−step model i s obtained by s e t t i n g the model argument t o ‚Ä ôt wo st ep s‚ Äô :
d i f f 2 <− update ( d i f f 1 , model = " twosteps " )
c o e f ( summary ( d i f f 2 ) )
# A l l a v a i l a b l e l a g s having been used , the number o f instruments i s s i z a b l e .
# One a c t u a l l y has : 0 . 5
# 6 . EXAMPLE 6 : instruments p r o l i f e r a t i o n − DemocracyIncome25 data s e t
408
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
# Load the next data frame
data ( " DemocracyIncome25 " , package = " pder " )
# Next , check out the dimension ( s i z e ) o f the data
pdim ( DemocracyIncome25 )
# Estimate the gmmmodel in d i f f e r e n c e s ,
#using the two v a r i a b l e s democracy and income as
#GMM instruments with a l l the a v a i l a b l e lag
d i f f 2 5 <− pgmm( democracy ~ lag ( democracy ) + lag ( income ) |
lag ( democracy , 2 : 9 9 ) + lag ( income , 2 : 9 9 ) ,
DemocracyIncome25 , model = " twosteps " )
# For each gmm instrument ,
# there are 0 . 5 x 6 x 5 = 15 moment c o n d i t i o n s and hence a t o t a l o f
#30 GMM instruments plus the 5 time dummies , i . e . J = 35 ,
#when the number o f i n d i v i d u a l s i s N = 2 5 .
d i f f 2 5 l i m <− pgmm( democracy ~ lag ( democracy ) + lag ( income ) |
lag ( democracy , 2 : 4 ) + lag ( income , 2 : 4 ) ,
DemocracyIncome , index=c ( " country " , " year " ) ,
model =" twosteps " , e f f e c t ="twoways " , subset = sample == 1 )
d i f f 2 5 c o l l <− pgmm( democracy ~ lag ( democracy ) + lag ( income ) |
lag ( democracy , 2 : 9 9 ) + lag ( income , 2 : 9 9 ) ,
DemocracyIncome , index=c ( " country " , " year " ) ,
model =" twosteps " , e f f e c t ="twoways " , subset = sample == 1 ,
c o l l a p s e = TRUE)
# Now use the sapply ( ) f u n c t i o n
sapply ( l i s t ( d i f f 2 5 , d i f f 2 5 l i m , d i f f 2 5 c o l l ) , f u n c t i o n ( x ) c o e f ( x ) [ 1 : 2 ] )
# As can be r e a d i l y seen , the r e s u l t s o f the three models are q u i t e s i m i l a r ,
#which seems t o i n d i c a t e
# that the p r o l i f e r a t i o n o f instruments
# i s not an important i s s u e in t h i s p a r t i c u l a r c o n t e x t .
# 7 . EXAMPLE 7 : system gmm − DemocracyIncome data s e t
409
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
sys2 <− pgmm( democracy ~ lag ( democracy ) + lag ( income ) |
lag ( democracy , 2 : 9 9 ) | lag ( income , 2 ) ,
DemocracyIncome , index = c ( " country " , " year " ) ,
model = " twosteps " , e f f e c t = " twoways " ,
transformation = " l d " )
c o e f ( summary ( sys2 ) )
# The a u t o r e g r e s s i v e c o e f f i c i e n t obtained with
#the d i f f e r e n c e and the system models are c l o s e .
# The income c o e f f i c i e n t i s now s i g n i f i c a n t l y p o s i t i v e and
#much l a r g e r than p r e v i o u s l y .
# 8 . EXAMPLE 8 : robust estimation o f the co va ria nce matrix
# Below we e x t r a c t the standard e r r o r s o f
#the f i r s t two c o e f f i c i e n t s f o r the two−step d i f f e r e n c e model .
s q r t ( diag ( vcov ( d i f f 2 ) ) ) [ 1 : 2 ]
s q r t ( diag ( vcovHC ( d i f f 2 ) ) ) [ 1 : 2 ]
# One can a c t u a l l y see that in t h i s example
#the c l a s s i c a l variance formula seems t o be biased downward .
# In f a c t , " robust " standard e r r o r s are c l e a r l y s u p e r i o r .
# 9 . EXAMPLE 9 : Sargan−Hansen t e s t − DemocracyIncome data s e t
# The Sargan−Hansen t e s t can be performed through f u n c t i o n sargan .
# For example , f o r the one−step d i f f e r e n c e model , one has :
sargan ( d i f f 2 )
sargan ( sys2 )
# To i l l u s t r a t e t h i s r e s u l t ,
#compute Sargan’s t e s t on the model p r e v i o u s l y estimated on the dataset
#with 7 o b s e r v a t i o n s o f 25 c o u n t r i e s .
410
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
sapply ( l i s t ( d i f f 2 5 , d i f f 2 5 l i m , d i f f 2 5 c o l l ) ,
f u n c t i o n ( x ) sargan ( x ) [ [ " p . value " ] ] )
# The p−value f o r the model using a l l moment c o n d i t i o n s i s near 1 ,
# while those o f the other models are much lower ;
# in p a r t i c u l a r , f o r the model l i m i t i n g the number o f l a g s t o 3 ,
#the hypothesis o f instruments v a l i d i t y i s r e j e c t e d at the 10% s i g n i f i c a n c e l e v e l .
# EXAMPLE 1 0 : a u t o c o r r e l a t i o n t e s t − DemocracyIncome data s e t
# The e r r o r s e r i a l c o r r e l a t i o n t e s t o f Arellano and Bond ( 1 9 9 1 )
# i s obtained through the f u n c t i o n mtest .
mtest ( d i f f 2 , order = 2 )
# The hypothesis o f no s e r i a l c o r r e l a t i o n i s not r e j e c t e d .
#PANEL TIME SERIES
# EXAMPLE 1 1 : Random c o e f f i c i e n t model − D i a l y s i s data s e t
data ( " D i a l y s i s " , package = " pder " )
rndcoef <− pvcm ( l o g ( d i f f u s i o n / ( 1 − d i f f u s i o n ) ) ~ trend + trend : r e g u l a t i o n ,
D i a l y s i s , model ="random " )
summary ( rnd coef )
# The r e s u l t s i n d i c a t e that c e r t i f i c a t e −of −need r e g u l a t i o n
#has slowed the d i f f u s i o n o f haemodialyis technology ,
#as the c o e f f i c i e n t i s s i g n i f i c a n t l y ( at the 5% l e v e l ) negative .
#The estimated c ov ari anc e matrix o f the random c o e f f i c i e n t s i s an element o f
#the f i t t e d model c a l l e d " Delta " ;
# the f o l l o w i n g command e x t r a c t s the mean values o f
#the three c o e f f i c i e n t s and t h e i r standard d e v i a t i o n s .
411
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
cbind ( c o e f ( rnd coef ) , stdev = s q r t ( diag ( rndcoef$Delta ) ) )
# The random c o e f f i c i e n t s have l a r g e standard d e v i a t i o n s :
#about h a l f the mean f o r the trend c o e f f i c i e n t
#and about two times the mean f o r the r e g u l a t i o n c o e f f i c i e n t s .
#These l a r g e values j u s t i f y the use o f the random c o e f f i c i e n t model .
#12. EXAMPLE 1 2 : Heterogeneous c o e f f i c i e n t s ‚ Ä ì HousePricesUS data s e t
data ( " HousePricesUS " , package = " pder " )
swmod <− pvcm ( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS , model= " random " )
mgmod <− pmg( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS , model = "mg" )
c o e f s <− cbind ( c o e f ( swmod ) , c o e f (mgmod ) )
dimnames ( c o e f s ) [ [ 2 ] ] <− c ( "Swamy" , "MG" )
coefs
# One can see that f o r T = 29 ,
#the e f f i c i e n t Swamy estimator and the simpler MG are already very c l o s e ;
#moreover , both are s t a t i s t i c a l l y very f a r from one .
#12. EXAMPLE 1 2 : Heterogeneous c o e f f i c i e n t s ‚ Ä ì HousePricesUS data s e t
data ( " HousePricesUS " , package = " pder " )
swmod <− pvcm ( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS , model= " random " )
mgmod <− pmg( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS , model = "mg" )
c o e f s <− cbind ( c o e f ( swmod ) , c o e f (mgmod ) )
dimnames ( c o e f s ) [ [ 2 ] ] <− c ( "Swamy" , "MG" )
coefs
# One can see that f o r T = 29 ,
#the e f f i c i e n t Swamy estimator and the simpler MG are already very c l o s e ;
#moreover , both are s t a t i s t i c a l l y very f a r from one .
#13. EXAMPLE 1 3 : dynamic mg estimation ‚ Ä ì RDSpillovers data s e t
412
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
# In the f o l l o w i n g , estimate both the s t a t i c MG model ( see t h e i r Table 7 )
#and the dynamic mg.
# As in the o r i g i n a l paper , i n c l u d e i n d i v i d u a l trends by s p e c i f y i n g trend =TRUE:
l i b r a r y ( " texreg " )
data ( " RDSpillovers " , package = " pder " )
fm . rds <− lny ~ l n l + lnk + lnrd
mg. rds <− pmg( fm . rds , RDSpillovers , trend = TRUE)
dmg . rds <− update (mg. rds , . ~ lag ( lny ) + . )
screenreg ( l i s t ( ’ S t a t i c MG’ = mg. rds ,
’ Dynamic MG’ = dmg . rds ) , d i g i t s = 3 )
# The lagged dependent v a r i a b l e turns out s i g n i f i c a n t ,
#although the a u t o r e g r e s s i v e parameter’s magnitude i s modest .
# On the b a s i s o f the dynamic model ,
#the authors proceed t o c a l c u l a t e the long−run c o e f f i c i e n t s with or
#without common f a c t o r r e s t r i c t i o n s ( see comment t o t h e i r Table 8 ) .
# Here we only reproduce the computation o f the long− run e l a s t i c i t y
# o f production t o own R&D
# ( which i s the r a t i o o f the c o e f f i c i e n t o f R&D t o one minus
#the a u t o r e g r e s s i v e c o e f f i c i e n t ) ,
#and the estimation o f i t s standard error , through a Taylor approximation ,
#by the d e l t a method .
# With r e f e r e n c e t o a v e c t o r o f K random v a r i a t e s ,
#the f u n c t i o n deltamethod from package msm ( Jackson , 2 0 1 1 ) r e q u i r e s :
#a formula d e s c r i b i n g the transformation
# ( here , x5 /(1 − x2 ) as the c o e f f i c i e n t s on lag ( lny )
#and lnrd are r e s p e c t i v e l y 2nd and 5th ) ;
#a v e c t o r o f K estimates f o r the means ;
#and a K x K matrix o f c ov ari anc e estimates .
# For the l a t t e r two ,
#here we provide the c o e f . panelmodel and vcov . panelmodel o f the dynamic model :
# i n s t a l l . packages ( "msm" ) i f you have not already
413
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
l i b r a r y ( "msm" )
b . l r <− c o e f (dmg . rds ) [ " lnrd " ] / ( 1 − c o e f (dmg . rds ) [ " lag ( lny ) " ] )
SEb . l r <− deltamethod (~ x5 / ( 1 − x2 ) ,
mean = c o e f (dmg . rds ) , cov = vcov (dmg . rds ) )
z . l r <− b . l r / SEb . l r
pval . l r <− 2 * pnorm ( abs ( z . l r ) , lower . t a i l = FALSE)
l r . lnrd <− matrix ( c ( b . l r , SEb . l r , z . l r , pval . l r ) , nrow=1)
dimnames ( l r . lnrd ) <− l i s t ( " lnrd ( long run ) " , c ( " Est . " , "SE" , " z " , " p . val " ) )
round ( l r . lnrd , 3 )
# A f t e r o b t a i n i n g the p o i n t estimate
#and standard e r r o r o f the long −run c o e f f i c i e n t ,
#we compute the t− s t a t i s t i c
#and the corresponding asymptotic p−value f o r the two−t a i l e d t e s t .
#The long −run e l a s t i c i t y o f production t o own
#R&D from the dynamic mg model i s not s i g n i f i c a n t
# at any c o n v e n t i o n a l c o n f i d e n c e l e v e l .
#14. EXAMPLE 1 4 : P o o l a b i l i t y t e s t − HousePricesUS data s e t
# Estimating the competing models f o r the HousePricesUS data , we have :
housep . np <− pvcm ( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS ,
model = " within " )
housep . p o o l <− plm ( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS ,
model = " p o o l i n g " )
housep . within <− plm ( l o g ( p r i c e ) ~ l o g ( income ) , data = HousePricesUS ,
model = " within " )
# As usual ,
#the pvcm f u n c t i o n p rovid es a c o e f . pvcm method t o r e t r i e v e i n d i v i d u a l c o e f f i c i e n t s .
414
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
# As a f i r s t assessment o f t h e i r d i s p e r s i o n ,
# in Figure 4 ( see s l i d e s ) we d i s p l a y a histogram
# o f the d i s t r i b u t i o n o f e i t h e r c o e f f i c i e n t .
# The summary . pvcm method instead returns ,
# f o r each c o e f f i c i e n t ,
#the s y n t h e t i c s t a t i s t i c s usually produced
#by summary f o r a g e n e r i c numeric v e c t o r :
summary ( housep . np )
# The s t a b i l i t y t e s t can then be performed supplying housep . np
#and e i t h e r housep . p o o l or housep . within t o the t e s t function ,
#depending on whether we want t o assume absence o f i n d i v i d u a l e f f e c t s or not .
# Notice the d i f f e r e n t degrees o f freedom .
p o o l t e s t ( housep . pool , housep . np )
p o o l t e s t ( housep . within , housep . np )
# C o e f f i c i e n t s t a b i l i t y i s very s t r o n g l y r e j e c t e d ,
#even in i t s weakest form ( s p e c i f i c constants ) .
# The same t e s t s can be performed using a formula−data syntax ,
# s p e c i f y i n g the nature o f the r e s t r i c t e d model through the model argument .
## CROSS−SECTIONAL DEPENDENCE AND COMMON FACTORS
#15. EXAMPLE 1 5 : Common c o r r e l a t e d e f f e c t s mg − HousePricesUS data s e t
l i b r a r y ( " texreg " )
cmgmod <− pmg( l o g ( p r i c e ) ~ l o g ( income ) ,
data = HousePricesUS , model = "cmg " )
screenreg ( l i s t (mg = mgmod, ccemg = cmgmod ) ,
d i g i t s = 3)
415
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#16. EXAMPLE 1 6 : ccemg and ccep − HousePricesUS data s e t
ccemgmod <− pcce ( l o g ( p r i c e ) ~ l o g ( income ) , data=HousePricesUS , model ="mg" )
summary ( ccemgmod )
# Holly e t a l . ( 2 0 1 0 ) are i n t e r e s t e d in estimating the r e l a t i o n s h i p
#between house p r i c e s and income net o f the i n f l u e n c e o f common f a c t o r s
#under the pooled s p e c i f i c a t i o n as well .
#To t h i s end , they estimate a homogeneous ccep v e r s i o n o f the b a s e l i n e model :
ccepmod <− pcce ( l o g ( p r i c e ) ~ l o g ( income ) ,
data=HousePricesUS , model ="p " )
summary ( ccepmod )
# The r e s u l t s from the two s p e c i f i c a t i o n s are very c l o s e ,
#as regards both the c o e f f i c i e n t s
#and the standard e r r o r s t h e r e o f ,
#which speaks in f a v o r o f imposing the p o o l i n g r e s t r i c t i o n .
#17. EXAMPLE 1 7 : variance o f the ccep estimator − RDSpillovers data s e t
ccep . rds <− pcce ( fm . rds , RDSpillovers , model ="p " )
l i b r a r y ( zoo )
l i b r a r y ( lmtest )
ccep . tab <− cbind ( c o e f t e s t ( ccep . rds ) [ , 1 : 2 ] ,
c o e f t e s t ( ccep . rds , vcov = vcovNW ) [ , 2 ] ,
c o e f t e s t ( ccep . rds , vcov = vcovHC ) [ , 2 ] )
dimnames ( ccep . tab ) [ [ 2 ] ] [ 2 : 4 ] <− c ( " Nonparam . " , "vcovNW " , " vcovHC " )
round ( ccep . tab , 3 )
# A p r i o r i , homogeneous variance e sti mato rs are r e l a t i v e l y well −s u i t e d
# t o t h i s comparatively l a r g e and short dataset ,
# provided that the homogeneity assumption holds .
#From the r e s u l t s we can instead see
# that the nonparametric standard e r r o r s are much more c o n s e r v a t i v e ,
416
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
# h i n t i n g at p o o l i n g assumptions being t o o r e s t r i c t i v e .
### NONSTATIONARITY AND COINTEGRATION ###
#18. EXAMPLE 1 8 : Unit Root Test : G e n e r a l i t i e s
autoreg <− f u n c t i o n ( rho = 0 . 1 , T = 1 0 0 ) {
e <− rnorm ( T )
f o r ( t in 2 : ( T ) ) e [ t ] <− e [ t ] + rho * e [ t −1]
e
}
t s t a t <− f u n c t i o n ( rho = 0 . 1 , T = 1 0 0 ) {
y <− autoreg ( rho , T )
x <− autoreg ( rho , T )
z <− lm ( y ~ x )
c o e f ( z ) [ 2 ] / s q r t ( diag ( vcov ( z ) ) [ 2 ] )
}
r e s u l t <− c ( )
R <− 1000
f o r ( i in 1 :R) r e s u l t <− c ( r e s u l t , t s t a t ( rho = 0 . 2 , T = 4 0 ) )
quantile ( result , c (0.025 , 0.975))
prop . t a b l e ( t a b l e ( abs ( r e s u l t ) > 2 ) )
#Can see how the e m p i r i c a l q u a n t i l e s are very c l o s e t o t h e i r expected values
#and the share o f f a l s e p o s i t i v e s i s in the region o f 5%.
# Let us now do the same with two s e r i e s , each c o n t a i n i n g a unit r o o t :
r e s u l t <− c ( )
R <− 1000
f o r ( i in 1 :R) r e s u l t <− c ( r e s u l t , t s t a t ( rho = 1 , T = 4 0 ) )
quantile ( result , c (0.025 , 0.975))
prop . t a b l e ( t a b l e ( abs ( r e s u l t ) > 2 ) )
# Judging by the usual t− s t a t i s t i c ,
417
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# in two t h i r d s o f cases one would conclude
# in f a v o r o f a s i g n i f i c a n t r e l a t i o n s h i p between our
#two independently generated v a r i a b l e s .
R <− 1000
T <− 100
r e s u l t <− c ( )
f o r ( i in 1 :R ) {
y <− autoreg ( rho =1 , T=100)
Dy <− y [ 2 : T ] − y [ 1 : ( T− 1)]
Ly <− y [ 1 : ( T− 1)]
z <− lm (Dy ~ Ly )
r e s u l t <− c ( r e s u l t , c o e f ( z ) [ 2 ] / s q r t ( diag ( vcov ( z ) ) [ 2 ] ) )
}
# Depict a histogram o f the r e a l i z a t i o n s o f the t− s t a t i s t i c ,
# superposing a normal d e n s i t y curve :
#One can e a s i l y see that employing c l a s s i c i n f e r e n c e procedures
# t o d e t e c t the presence o f unit r o o t s i s unwarranted ,
#as the t− s t a t i s t i c f o l l o w s a d i s t r i b u t i o n
# that i s very f a r from the normal .
#Employing the usual c r i t i c a l value o f ‚ à í 1 . 6 4 , one has here :
prop . t a b l e ( t a b l e ( r e s u l t < − 1.64))
# which leads t o r e j e c t the true hypothesis o f a unit r o o t one h a l f o f the times .
# To perform the Dickey−F u l l e r t e s t ,
#one needs s p e c i f i c c r i t i c a l values
# that are not those o f the normal ( or the t ) d i s t r i b u t i o n .
# The t e s t can be performed augmenting the a u x i l i a r y model
#with a constant and / or a d e t e r m i n i s t i c trend ;
# l a g s o f Delta ( y ) can a l s o be added in order t o clean out
#any p o s s i b l e a u t o c o r r e l a t i o n o f e p s i l o n .
#19. EXAMPLE 1 9 : F i r s t generation unit r o o t t e s t i n g − HousePricesUS data s e t
i n s t a l l . packages ( " pder " )
i n s t a l l . packages ( " maxLik " )
i n s t a l l . packages ( " miscTools " )
418
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
l i b r a r y ( maxLik )
l i b r a r y ( miscTools )
l i b r a r y ( pder )
data ( " HousePricesUS " , package = " pder " )
p r i c e <− pdata . frame ( HousePricesUS ) $ p r i c e
p u r t e s t ( l o g ( p r i c e ) , t e s t = " l e v i n l i n " , l a g s = 2 , exo = " trend " )
p u r t e s t ( l o g ( p r i c e ) , t e s t = "madwu" , l a g s = 2 , exo = " trend " )
p u r t e s t ( l o g ( p r i c e ) , t e s t = " i p s " , l a g s = 2 , exo = " trend " )
# The three t e s t s s t r o n g l y don’t r e j e c t the n u l l hypothesis o f unit r o o t .
#20. EXAMPLE 2 0 : i p s and c i p s t e s t s − HousePricesUS data s e t
tab5a <− matrix (NA, n c o l = 4 , nrow = 2 )
tab5b <− matrix (NA, n c o l = 4 , nrow = 2 )
f o r ( i in 1 : 4 ) {
mymod <− pmg( d i f f ( l o g ( income ) ) ~ lag ( l o g ( income ) ) +
lag ( d i f f ( l o g ( income ) ) , 1 : i ) ,
data = HousePricesUS ,
model = "mg" , trend = TRUE)
tab5a [ 1 , i ] <− p c d t e s t (mymod, t e s t = " rho " ) $ s t a t i s t i c
tab5b [ 1 , i ] <− p c d t e s t (mymod, t e s t = " cd " ) $ s t a t i s t i c
}
f o r ( i in 1 : 4 ) {
mymod <− pmg( d i f f ( l o g ( p r i c e ) ) ~ lag ( l o g ( p r i c e ) ) +
lag ( d i f f ( l o g ( p r i c e ) ) , 1 : i ) ,
data=HousePricesUS ,
model ="mg" , trend = TRUE)
tab5a [ 2 , i ] <− p c d t e s t (mymod, t e s t = " rho " ) $ s t a t i s t i c
tab5b [ 2 , i ] <− p c d t e s t (mymod, t e s t = " cd " ) $ s t a t i s t i c
}
tab5a <− round ( tab5a , 3 )
419
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
tab5b <− round ( tab5b , 2 )
dimnames ( tab5a ) <− l i s t ( c ( " income " , " p r i c e " ) ,
paste ( "ADF( " , 1 : 4 , " ) " , sep = " " ) )
dimnames ( tab5b ) <− dimnames ( tab5a )
tab5a
tab5b
# Residual cross −c o r r e l a t i o n i s c l e a r l y apparent and motivates employing the c i p s t e
# In the f o l l o w i n g we assess the order o f i n t e g r a t i o n o f p r i c e s
#and income by t e s t i n g the o r i g i n a l s e r i e s
#and the d i f f e r e n c e d ones f o r unit r o o t s .
#To do so , the dataset now contained in the data . frame HousePricesUS
#has t o be converted i n t o a pdata . frame from which the t e s t i n g f u n c t i o n c i p s t e s t
# w i l l be able t o r e t r i e v e the panel i n d i c e s i t needs .
#The number o f l a g s i s l e f t at the d e f a u l t value o f 2 .
#As f o r the d e t e r m i n i s t i c component o f the cadf r e g r e s s i o n s ,
#we allow f o r an i n t e r c e p t ( type= ‚ Ä ô d r i f t ‚ Ä ô ) in the o r i g i n a l s e r i e s ;
# f o r the sake o f c o n s i s t e n c y ,
#we then exclude i t from the d i f f e r e n c e d one ( type=’none’ ) .
php <− pdata . frame ( HousePricesUS )
c i p s t e s t ( l o g ( php$price ) , type = " d r i f t " )
c i p s t e s t ( d i f f ( l o g ( php$price ) ) , type = " none " )
# The c i p s t e s t does not r e j e c t a unit r o o t f o r the o r i g i n a l s e r i e s ,
# while i t does f o r the d i f f e r e n c e d one .
# The c o n c l u s i o n i s that the p r i c e index i s i n t e g r a t e d o f order 1 .
#The same ( not reported ) happens f o r income ,
# at which p o i n t the c r u c i a l i s s u e i s whether house p r i c e s
#and income are c o i n t e g r a t e d ,
# or otherwise the r e g r e s s i o n o f i n t e r e s t i s spurious .
# A c i p s t e s t o f the r e g r e s s i o n r e s i d u a l s w i l l help shed l i g h t on the i s s u e :
# c u r r e n t l y the c i p s t e s t f u n c t i o n only a c c e p t s p s e r i e s o b j e c t s as arguments ;
420
Electronic copy available at: https://ssrn.com/abstract=3863563
8.22. DYNAMIC MODELS & PANEL TIME SERIES
#hence , we e x t r a c t r e s i d u a l s as a p s e r i e s through the usual r e s i d . ccep e x t r a c t o r
# f u n c t i o n p r i o r t o f e e d i n g them t o the unit r o o t t e s t .
#Given that i n d i v i d u a l trends have been c o n t r o l l e d f o r at the modeling stage
#and that by the very nature o f r e g r e s s i o n r e s i d u a l s ,
#the s e r i e s i s not expected t o contain a d r i f t ( i n t e r c e p t ) ,
#we e l i m i n a t e any d e t e r m i n i s t i c component from
#the cadf r e g r e s s i o n s by s p e c i f y i n g type=’none’ :
c i p s t e s t ( r e s i d ( ccemgmod ) , type =" none " )
c i p s t e s t ( r e s i d ( ccepmod ) , type =" none " )
# The unit r o o t hypothesis i s r e j e c t e d f o r both
#the r e s i d u a l s o f the ccemg and the ccepmodels .
# The c o n c l u s i o n i s that both models represent c o i n t e g r a t i n g r e g r e s s i o n s .
421
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.23
Confirmatory factor analysis
# CONFIRMATORY FACTOR ANALYSIS INTRODUCTION_EXERCISE
#1.1 I n s t a l l and load the sem package
i n s t a l l . packages ( " sem " )
l i b r a r y ( sem )
#1.2 Now load the data s e t c a l l e d Bollen
data ( " Bollen " )
dim ( Bollen )
summary ( Bollen )
#1.3 Now we estimate SEM
model <− ’
# measurement model
xi_1 =~ y1 + l 2 * y2 + l 3 * y3 + l 4 * y4
xi_2 =~ y5 + l 2 * y6 + l 3 * y7 + l 4 * y8
# residual correlations
y1 ~~ y5
y2 ~~ y4 + y6
y3 ~~ y7
y4 ~~ y8
y6 ~~ y8
’
#NOTES: The ~~ syntax s p e c i f i e s which e r r o r variances covary
#1.4 Need t o a l s o use the lavaan package as i t c o n t a i n s the c f a ( ) f u n c t i o n
l i b r a r y ( lavaan )
#1.5 Now use the model o b j e c t as the f i r s t argument t o the c f a ( ) f u n c t i o n
#and r e s u l t s obtained with the summary ( ) f u n c t i o n
f i t <− c f a ( model , data = Bollen )
summary ( f i t , f i t . measures = TRUE, standardized = TRUE)
#NOTES: Look at the output t a b l e − i t i s long , more than one page !
#NOTES: The lavaan package a l s o p rovid es a s e t o f e x t r a c t o r f u n c t i o n s
422
Electronic copy available at: https://ssrn.com/abstract=3863563
8.23. CONFIRMATORY FACTOR ANALYSIS
# t o p u l l s p e c i f i c p o r t i o n s o f the output t o f u r t h e r p r o c e s s or analyze .
#The parameterEstimates ( ) , st a nd ar d iz e dS ol u ti o n ( ) ,
#and fitMeasures ( ) f u n c t i o n s can be used t o return
# only the unstandardized estimates , standardized estimates ,
#and model f i t s t a t i s t i c s , r e s p e c t i v e l y .
#1.6 Use the f i r s t one parameterEstimates ( ) f u n c t i o n
#and apply t o our model c a l l e d f i t
parameterEstimates ( f i t )
#1.7 Next , t r y the second one c a l l e d s ta nd a rd i ze dS o lu t io n ( ) ,
# noting the c a p i t a l l e t t e r S (R i s case s e n s i t i v e )
standardizedS o lu t io n ( f i t )
#1.8 Lastly use the t h i r d o p t i o n c a l l e d fitMeasures ( )
fitMeasures ( f i t )
#1.9 V i s u a l i z e our model with semPlot : : semPaths ( ) ,
#So need t o i n s t a l l the package then load the semPlot l i b r a r y f i r s t
l i b r a r y ( semPlot )
#1.10 Now can use the semPaths ( ) f u n c t i o n with our model c a l l e d f i t
semPaths ( f i t , nCharNodes = 0 , s t y l e = " l i s r e l " , r o t a t i o n = 2 )
#NOTES:
# 1 . The semPaths ( ) f u n c t i o n takes the f i t t e d lavaan model o b j e c t
#as the main argument ,
#but has a number o f d i f f e r e n t o p t i o n s a v a i l a b l e t o customize the path diagram .
#Here , we s e t nCharNodes = 0 , so that the v a r i a b l e names are not abbreviated .
# 2 . Also s e t the s t y l i n g t o look l i k e the " l i s r e l " software output ,
#and s e t the r o t a t i o n so that the path diagram f l o w s h o r i z o n t a l l y .
#We can see in the output image that the s p e c i f i c a t i o n o f the model from lavaan
# i s r e f l e c t e d in the diagram
423
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.24
Analysis of confirmation factor analysis models kabc
# ANALYSIS OF CONFIRMATION FACTOR ANALYSIS MODELS kabc_EXERCISE
#1.1 Load the l i b r a r y
date ( )
l i b r a r y ( lavaan )
#1.2 input the c o r r e l a t i o n s in lower diagnonal form
kabcLower . c o r <− ’
1.00
. 3 9 1.00
.35
. 6 7 1.00
.21
.11
. 1 6 1.00
.32
.27
.29
. 3 8 1.00
.40
.29
.28
.30
. 4 7 1.00
.39
.32
.30
.31
.42
. 4 1 1.00
.39
.29
.37
.42
.58
.51
. 4 2 1.00 ’
#1.3 name the v a r i a b l e s and convert t o f u l l c o r r e l a t i o n matrix
kabcFull . c o r <− getCov ( kabcLower . cor , names = c ( "hm" ,
" nr " ,
"wo " ,
" gc " ,
" tr " ,
"sm" ,
"ma" ,
" ps " ) )
#1.4 d i s p l a y the c o r r e l a t i o n s
kabcFull . c o r
#1.5 add the standard d e v i a t i o n s and convert t o c o v a r i a n c e s
kabcFull . cov <− cor2cov ( kabcFull . cor , sds = c ( 3 . 4 0 ,
2.40 ,
2.90 ,
2.70 ,
2.70 ,
424
Electronic copy available at: https://ssrn.com/abstract=3863563
8.24. ANALYSIS OF CONFIRMATION FACTOR ANALYSIS MODELS KABC
4.20 ,
2.80 ,
3.00))
kabcFull . cov
#1.6 s p e c i f y c f a model
kabc . model <− ’
#1.7 l a t e n t v a r i a b l e s
Sequent =~ hm + nr + wo
Simultan =~ gc + t r + sm + ma + ps ’
# i n d i c a t o r s hm and gc a u t o m a t i c a l l y
# s p e c i f i e d as r e f e r e n c e v a r i a b l e s
# unanalyzed a s s o c i a t i o n between the two f a c t o r s
# automatically s p e c i f i e d
#1.8 f i t model t o data
model <− sem ( kabc . model ,
sample . cov=kabcFull . cov ,
sample . nobs =200)
summary ( model , f i t . measures = TRUE, standardized = TRUE, rsquare = TRUE)
f i t t e d ( model )
r e s i d u a l s ( model , type = " raw " )
r e s i d u a l s ( model , type = " standardized " )
r e s i d u a l s ( model , type = " c o r " )
modindices ( model )
425
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.25
Structural equation models illustration
#STRUCTURAL EQUATION MODELS ILLUSTRATION_EXERCISE
#1.1 Klein data model
l i b r a r y ( sem )
data ( Klein )
Klein
#1.2 The lagged v a r i a b l e s can be added t o the data frame as f o l l o w s ,
# p r i n t i n g the f i r s t three o b s e r v a t i o n s :
Klein$P . lag <− c (NA, Klein$P [ − 22])
Klein$X . lag <− c (NA, Klein$X [ − 22])
Klein [ 1 : 3 , ]
#NOTES: In S , NA ( not a v a i l a b l e ) r e p r e s e n t s missing data , and ,
# c o n s i s t e n t with standard s t a t i s t i c a l notation , a negative s u b s c r i p t ,
#such as −22, drops o b s e r v a t i o n s .
#Square brackets are used t o index o b j e c t s such as data frames
# ( e . g . , Klein [ 1 : 3 , ] ) , v e c t o r s ( e . g . , Klein$P [ − 22]) , matrices , arrays , and l i s t s .
#The d o l l a r sign ( $ ) can be used t o index elements o f data frames or l i s t s
#1.3 Using these instrumental v a r i a b l e s , the s t r u c t u r a l equations can be estimate
Klein . eqn1 <− t s l s (C ~ P + P . lag + I (Wp + Wg) ,
instruments=~G + T + Wg + I ( Year − 1931) +
K. lag +
P . lag +
X . lag ,
data=Klein )
Klein . eqn2 <− t s l s ( I ~ P + P . lag + K. lag ,
instruments=~G + T + Wg + I ( Year − 1931) +
K. lag +
P . lag +
X . lag ,
data=Klein )
Klein . eqn3 <− t s l s (Wp ~ X + X . lag + I ( Year − 1931) ,
426
Electronic copy available at: https://ssrn.com/abstract=3863563
8.25. STRUCTURAL EQUATION MODELS ILLUSTRATION
instruments=~G + T + Wg + I ( Year − 1931) +
K. lag +
P . lag +
X . lag ,
data=Klein )
#1.4 To produce a p rinted summary f o r the f i r s t s t r u c t u r a l equation :
summary ( Klein . eqn1 )
#1.5 Output equations 2 and 3
summary ( Klein . eqn2 )
summary ( Klein . eqn3 )
# Note : The c o e f f i c i e n t estimates are i d e n t i c a l t o those in Greene ( 1 9 9 3 ) ,
#but the c o e f f i c i e n t standard e r r o r s d i f f e r s l i g h t l y ,
#because the summary method f o r t s l s o b j e c t s yses r e s i d u a l
# degrees o f freedom ( 1 7 ) rather than the number o f o b s e r v a t i o n s ( 2 1 )
# t o estimate the e r r o r variance . To r e c o v e r Greene ’ s asymptotic standard e r r o r s ,
#the covarianc e matrix o f the c o e f f i c i e n t s can be e x t r a c t e d and adjusted ,
# i l l u s t r a t i n g a computation on a t s l s o b j e c t :
s q r t ( diag ( vcov ( Klein . eqn1 ) * 1 7 / 2 1 ) )
#1.6 Model s p e c i f i c a t i o n in the sem package i s handled most c o n v e n i e n t l y
# via the s p e c i f y . model f u n c t i o n :
mod .wh. 1 <− s p e c i f y . model ( )
Alienation67 −> Anomia67 , NA, 1
Alienation67 −> Powerless67 , lam1 , NA
Alienation71 −> Anomia71 , NA, 1
Alienation71 −> Powerless71 , lam2 , NA
SES −> Education , NA, 1
SES −> SEI , lam3 , NA
Alienation67 −> Alienation71 , beta , NA
SES −> Alienation67 , gam1 , NA
SES −> Alienation71 , gam2 , NA
SES <−> SES, phi , NA
Alienation67 <−> Alienation67 , psi1 , NA
Alienation71 <−> Alienation71 , psi2 , NA
Anomia67 <−> Anomia67 , the11 , NA
427
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
Powerless67 <−> Powerless67 , the22 , NA
Anomia71 <−> Anomia71 , the33 , NA
Powerless71 <−> Powerless71 , the44 , NA
Education <−> Education , thd1 , NA
SEI <−> SEI , thd2 , NA
#NOTES:
# 1 . Line number prompts are supplied
# 2 . Last l i n e i s blank + <CTRL>+<ENTER>
#1.7 To estimate the model , the cov ar ian ce or raw−moment matrix among
#the observed v a r i a b l e s has t o be computed .
#In the case o f the Wheaton data , the c o v a r i a n c e s rather than
#the o r i g i n a l data are a v a i l a b l e ,
#and consequently the c ov ari anc e matrix i s entered d i r e c t l y
S .wh <− matrix ( c (
+ 11.834 , 0 , 0 , 0 , 0 , 0 ,
+ 6.947 , 9.364 , 0 , 0 , 0 , 0 ,
+ 6 . 8 1 9 , 5 . 0 9 1 , 12.532 , 0 , 0 , 0 ,
+ 4.783 , 5.028 , 7.495 , 9.986 , 0 , 0 ,
+ − 3.839 , − 3.889 , − 3.841 , − 3.625 , 9 . 6 1 0 , 0 ,
+ − 21.899 , − 18.831 , − 21.748 , − 18.775 , 35.522 , 4 5 0 . 2 8 8 ) ,
+ 6 , 6 , byrow=TRUE)
rownames ( S .wh) <− colnames ( S .wh) <− c ( " Anomia67 " ,
" Powerless67 " ,
" Anomia71 " ,
" Powerless71 " ,
" Education " ,
" SEI " )
#NOTES: The c ov ari anc e matrix has been entered in lower−t r i a n g u l a r form :
#sem w i l l accept a lower t r i a n g u l a r , upper t r i a n g u l a r ,
# or symmetric c ov ari anc e matrix .
#One can e i t h e r assign the names o f the observed v a r i a b l e s t o the rows
#and columns o f the input c ov ari anc e matrix , as done here ,
# or pass these names d i r e c t l y t o sem ; in e i t h e r event ,
428
Electronic copy available at: https://ssrn.com/abstract=3863563
8.25. STRUCTURAL EQUATION MODELS ILLUSTRATION
# v a r i a b l e s in the model s p e c i f i c a t i o n that do not appear
# in the input c ov ari anc e matrix are assumed by sem t o be l a t e n t v a r i a b l e s .
#One must t h e r e f o r e be c a r e f u l in typing these names ,
#because the m issp ell ed name o f an observed v a r i a b l e i s i n t e r p r e t e d
#as a l a t e n t v a r i a b l e , producing an erroneous model .
#1.8 To estimate the model
sem .wh. 1 <− sem (mod .wh. 1 , S . wh, N=932)
summary ( sem .wh. 1 )
#1.9 The summary method f o r sem o b j e c t s produces the p r i n t o u t shown p r e v i o u s l y .
#One can perform a d d i t i o n a l computations on sem o b j e c t s ,
# f o r example , producing various kinds o f r e s i d u a l c o v a r i a n c e s
# or m o d i f i c a t i o n indexes
modIndices ( sem .wh. 1 )
#NOTES:
# 1 . Simply allow the o b j e c t returned by mod . i n d i c e s t o be printed ,
#which produces a b r i e f r e p o r t ;
#the summary method f o r these o b j e c t s produces a more complete report ,
#showing a l l m o d i f i c a t i o n indexes along with approximations t o the estimates
# that would r e s u l t were each omitted parameter included in the model .
# 2 . R e c a l l that the A matrix c o n t a i n s r e g r e s s i o n c o e f f i c i e n t s
#whereas the P matrix c o n t a i n s c o v a r i a n c e s .
#The m o d i f i c a t i o n indexes suggest that a b e t t e r f i t t o the data
#would be achieved by f r e e i n g one or more o f the c o v a r i a n c e s
#among the measurement e r r o r s o f the l a t e n t endogenous v a r i a b l e s ;
#the l a r g e s t m o d i f i c a t i o n index i s f o r the two anomia measures ,
# corresponding t o the broken arrow .
429
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.26
Analysis of structural equation models - roth
#ANALYSIS OF STRUCTURAL EQUATION MODELS ROTH_EXERCISE
#1.1 Load the l i b r a r y
date ( )
l i b r a r y ( lavaan )
#1.2 Input the c o r r e l a t i o n s in lower diagnonal form
rothLower . c o r <− ’
1.00
−.03 1.00
.39
. 0 7 1.00
−.05 −.23 −.13 1.00
−.08 −.16 −.29
. 3 4 1.00 ’
#1.3 name the v a r i a b l e s and convert t o f u l l c o r r e l a t i o n matrix
r o t h F u l l . c o r <− getCov ( rothLower . cor , names = c ( " e x e r c i s e " ,
" hardy " ,
" fitness " ,
" stress " ,
" illness "))
#1.4 d i s p l a y the c o r r e l a t i o n s
rothFull . cor
#1.5 add the standard d e v i a t i o n s and convert t o c o v a r i a n c e s
r o t h F u l l . cov <− cor2cov ( r o t h F u l l . cor , sds = c ( 6 6 . 5 0 , 3 8 . 0 0 , 1 8 . 4 0 , 3 3 . 5 0 , 6 2 . 4 8 ) )
#1.6 d i s p l a y the c o v a r i a n c e s
r o t h F u l l . cov
#1.7 s p e c i f y path model
roth . model <− ’
# regressions
fitness ~ exercise
s t r e s s ~ hardy
i l l n e s s ~ fitness + stress ’
430
Electronic copy available at: https://ssrn.com/abstract=3863563
8.27. ANALYSIS OF STRUCTURAL REGRESSION MODELS
#NOTES:
# 1 . unanalyzed a s s o c i a t i o n between e x e r c i s e and hardy
#2. automatically s p e c i f i e d
#1.8 f i t i n i t i a l model t o data ; variances and cov ar ian ce o f measured exogenous ;
# v a r i a b l e s are f r e e parameters ;
# variances c a l c u l a t e d with N − 1 in the denominator instead o f N
model <− sem ( roth . model ,
sample . cov= r o t h F u l l . cov ,
sample . nobs =373 , f i x e d . x = FALSE, sample . cov . r e s c a l e = FALSE)
summary ( model , f i t . measures = TRUE, standardized = TRUE, rsquare = TRUE)
f i t t e d ( model )
r e s i d u a l s ( model , type = " raw " )
r e s i d u a l s ( model , type = " standardized " )
r e s i d u a l s ( model , type = " c o r " )
modindices ( model )
8.27
Analysis of structural regression models
#ANALYSIS OF STRUCTURAL REGRESSION MODELS_EXERCISE
#1.1 Load and l i b r a r y the Data
date ( )
l i b r a r y ( lavaan )
#1.2 input the c o r r e l a t i o n s in lower diagnonal form
houghtonLower . c o r <− ’
1.000
.668 1.000
.635
.599 1.000
.263
.261
.164 1.000
.290
.315
.247
.486 1.000
.207
.245
.231
.251
.449 1.000
−.206 −.182
−.195 −.309 −.266 −.142 1.000
−.280 −.241
−.238 −.344 −.305 −.230
−.258 −.244 −.185
−.255 −.255 −.215
.080
.096
−.017
.061
.028 − .035
.094
.151
.753 1.000
.554
.587 1.000
.141 − .074 − .111
.016 1.000
−.058 −.051 −.003 −.040 −.040 −.018 .284 1.000
431
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
.113
.174
.059
.063
.138
.044 − .119 − .073 − .084 .563
.379 1.000 ’
#1.3 name the v a r i a b l e s and convert t o f u l l c o r r e l a t i o n matrix
houghtonFull . c o r <−
getCov ( houghtonLower . cor , names = c ( " wk1 " ,
"wk2 " ,
"wk3 " ,
" hap " ,
"md1" ,
"md2" ,
" pr1 " ,
" pr2 " ,
" app " ,
" bel " ,
" st " ,
" ima " ) )
#1.4 d i s p l a y the c o r r e l a t i o n s
houghtonFull . c o r
#1.5 add the standard d e v i a t i o n s and convert t o c o v a r i a n c e s
houghtonFull . cov <−
cor2cov ( houghtonFull . cor , sds = c ( . 9 3 9 ,
1.017 ,
.937 ,
.562 ,
.760 ,
.524 ,
.585 ,
.609 ,
.731 ,
.711 ,
1.124 ,
1.001))
houghtonFull . cov
#1.6 s p e c i f y c f a model
houghtonCFA . model <− ’
432
Electronic copy available at: https://ssrn.com/abstract=3863563
8.27. ANALYSIS OF STRUCTURAL REGRESSION MODELS
# measurement part
Constru =~ b e l + s t + ima
Dysfunc =~ pr1 + pr2 + app
WellBe =~ hap + md1 + md2
JobSat =~ wk1 + wk2 + wk3
# e r r o r covari anc e
hap ~~ md2 ’
#1.7 s p e c i f y sr model
houghtonSR . model <− ’
#1.8 measurement part
Construc =~ b e l + s t + ima
Dysfunc =~ pr1 + pr2 + app
WellBe =~ hap + md1 + md2
JobSat =~ wk1 + wk2 + wk3
# e r r o r covari anc e
hap ~~ md2
# s t r u c t u r a l part
Dysfunc ~ Construc
WellBe ~ Construc + Dysfunc
JobSat ~ Construc + Dysfunc + WellBe ’
#1.9 f i t c f a model t o data
cfamodel <− sem ( houghtonCFA . model ,
sample . cov=houghtonFull . cov ,
sample . nobs =263)
summary ( cfamodel , f i t . measures = TRUE, standardized = TRUE, rsquare = TRUE)
f i t t e d ( cfamodel )
r e s i d u a l s ( cfamodel , type = " raw " )
r e s i d u a l s ( cfamodel , type = " standardized " )
r e s i d u a l s ( cfamodel , type = " c o r " )
433
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
modindices ( cfamodel )
#1.10 f i t sr model t o data
srmodel <− sem ( houghtonSR . model ,
sample . cov=houghtonFull . cov ,
sample . nobs =263)
summary ( srmodel , f i t . measures = TRUE, standardized = TRUE, rsquare = TRUE)
f i t t e d ( srmodel )
r e s i d u a l s ( srmodel , type = " raw " )
r e s i d u a l s ( srmodel , type = " standardized " )
r e s i d u a l s ( srmodel , type = " c o r " )
modindices ( srmodel )
434
Electronic copy available at: https://ssrn.com/abstract=3863563
8.28. VECTOR AUTOREGRESSIVE MODELS
8.28
Vector autoregressive models
# VECTOR AUTOREGRESSIVE MODELS_EXERCISE
#NOTES:
# 1 . Estimating a VAR model using simulated data ( generated with known parameters )
# 2 . VAR model i s a system o f equations
#1.1 F i r s t , i n s t a l l some new R packages ,
# c a l l e d ’ cars ’ ,
’ t s e r i e s ’ and ’ t i d y v e r s e ’ and load them
i n s t a l l . packages ( " vars " )
i n s t a l l . packages ( " t s e r i e s " )
i n s t a l l . packages ( " t i d y v e r s e " )
i n s t a l l . packages ( " g g f o r t i f y " )
l i b r a r y (MASS)
l i b r a r y ( strucchange )
l i b r a r y ( zoo )
l i b r a r y ( sandwich )
l i b r a r y ( urca )
l i b r a r y ( lmtest )
l i b r a r y ( vars )
library ( tseries )
library ( tidyverse )
library ( stargazer )
library ( ggfortify )
#1.2 Read the sample simulated CSV data f i l e i n t o R
y <− read . csv ( f i l e . choose ( ) , header = TRUE)
head ( y , 1 0 )
tail (y ,10)
str ( y )
#2.1 c r e a t e time s e r i e s ( t s ) o b j e c t s f o r y1 and y2
# Let assume that we are using monthly observations , commencing in the year 2000
y1 <− t s ( y$y1 , s t a r t = c ( 2 0 0 0 , 1 ) , frequency = 12)
y2 <− t s ( y$y2 , s t a r t = c ( 2 0 0 0 , 1 ) , frequency = 12)
#2.1 Now simply p l o t the data using the a u t o p l o t ( ) f u n c t i o n from
435
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#the ’ g g f o r t i f y ’ package
a u t o p l o t ( y1 )
a u t o p l o t ( y2 )
# Both time s e r i e s appear t o be s t a t i o n a r y ones ( mean c l o s e t o zero )
#with constant variance
# 3 . UNIT−ROOT TESTING
#NOTES: t e s t f o r s t a t i o n a r i t y using the ADF ( Augmented Dickey F u l l e r ) t e s t
#found in ’ t s e r i e s ’
adfy1 <− adf . t e s t ( y1 )
p r i n t ( adfy1 )
#NOTES:
# 1 . view ( adfy1 ) or double− c l i c k on the l i s t o b j e c t in the Environment window
# ( on the r i g h t in RStudio )
# 2 . the n u l l hypothesis that there i s a unit −r o o t i s REJECTED,
#thus the s e r i e s i s s t a t i o n a r y
adfy2 <− adf . t e s t ( y2 )
p r i n t ( adfy2 )
#NOTES:
# 1 . So y2 i s s t a t i o n a r y t o o
# ( by design as t h i s sample data was created t o be s t a t i o n a r y )
# 2 . Remember though that f o r your data ,
#you w i l l always have t o c o n s t r u c t a " formal " t e s t f o r unit −r o o t
# 4 . OPTIMAL LAG LENGTH
#NOTES: Too few lags , we do not estimate the e r r o r term c o r r e c t l y ,
# t o o many l a g s could r e s u l t in o v e r f i t t i n g or t o o many parameters f o r the model
#4.1 Will use the VARselect ( ) f u n c t i o n from the ’ vars ’ package
? VARselect
#4.2 Can do t h i s f o r the data . frame ’ y ’
lag <− VARselect ( y )
436
Electronic copy available at: https://ssrn.com/abstract=3863563
8.28. VECTOR AUTOREGRESSIVE MODELS
lag
#NOTES: Or we can use p r i n t ( lag )
#4.3 The d e f a u l t number o f l a g s in R using t h i s f u n c t i o n = 10 ,
#but we could l i m i t t o say j u s t 3 l a g s
lag <− VARselect ( y , lag . max = 3 )
lag
#NOTES:
# 1 . O b j e c t i v e t o i s s e l e c t the lag length that minimizes these c r i t e r i a
# 2 . A l l 4 c r i t e r i a show that the optimal lag length i s 1 period
# ( by design from the raw data c r e a t i o n )
lag$selection
#NOTES: i s o l a t e the part o f the ’ lag ’ o b j e c t t e s t that we are i n t e r e s t e d in
# 5 . MODEL ESTIMATION
#5.1 going t o use the VAR( ) function ,
# so check out the f u n c t i o n arguments , syntax e t c with ? help
estim <− VAR( y , p = 1 , type = " none " )
#5.2 Results are contained in a new o b j e c t c a l l e d ’ estim ’ ,
#which we can now view with View ( )
View ( estim )
#5.3 Also examine the r e s u l t s using the summary ( ) f u n c t i o n
summary ( estim )
#5.4 Can use the s t a r g a z e r package t o make the output t a b l e more a e s t h e t i c a l l y
# p l e a s i n g ( f o r papers e t c )
s t a r g a z e r ( estim [ [ " v a r r e s u l t " ] ] , type = ’ text ’ )
s t a r g a z e r ( estim [ [ " v a r r e s u l t " ] ] , type = ’ l a t e x ’ )
#NOTES: This can generate l a t e x codes
#NOTES:
# 1 . Output i s c l e a n e r and e a s i e r t o read ,
437
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# plus c o n t a i n s the necessary content f o r a j o u r n a l t a b l e
# 2 . Note that model c o e f f i c i e n t i n t e r p r e t a t i o n may not be that useful ,
# e s p e c i a l l y i f you have l o t s o f v a r i a b l e s ( so mu ltiple equations )
#alongside l o t s of lags .
# 6 . STABILITY CHECK
#NOTES: Eigenvalues should be l e s s than 1
#6.1 can use the r o o t s ( ) f u n c t i o n l o c a t e d in the ’ vars ’ package
# t o perform t h i s s t a b i l i t y t e s t
r o o t s ( estim , modulus = TRUE)
#NOTES: Results are good , l e s s than 1 . 0 , so can proceed t o the next Step
# 7 . GRANGER CAUSALITY TEST
#QUESTIONS: I s one time s e r i e s u s e f u l f o r p r e d i c t i n g another time s e r i e s ?
grangery1 <− c a u s a l i t y ( estim , cause = " y1 " )
grangery1$Granger
#NOTES:
# 1 . Null hypothesis s t a t e s that y1 does not Granger−cause y2
# 2 . Our p−value i s almost zero ( very small ) so we r e j e c t the n u l l hypothesis
#7.1 See i f y2 causes y1 or not
grangery2 <− c a u s a l i t y ( estim , cause = " y2 " )
grangery2$Granger
#NOTES:
# 1 . Notice that the p−value i s not s t a t i s t i c a l l y s i g n i f i c a n t
# 2 . Thus , F a i l t o R e j e c t the n u l l hypothesis that y2 do not Granger−cause y1
# 8 . IRFs AND VARIANCE DECOMPOSITION
#8.1 Use the i r f ( ) f u n c t i o n b u i l t −i n t o the ’ vars ’ package
i r f 1 <− i r f ( estim ,
impulse = " y1 " ,
response = " y2 " ,
n . ahead = 20 ,
438
Electronic copy available at: https://ssrn.com/abstract=3863563
8.28. VECTOR AUTOREGRESSIVE MODELS
boot = TRUE,
run = 200 ,
ci = 0.95)
i r f 2 <− i r f ( estim ,
impulse = " y1 " ,
response = " y2 " ,
n . ahead = 20 ,
boot = TRUE,
run = 400 ,
ci = 0.90)
p l o t ( i r f 1 , ylab = " y2 " , main = " y2 response t o y1 shock " )
p l o t ( i r f 2 , ylab = " y2 " , main = " y2 response t o y1 shock " )
#NOTES:
# 1 . Notice the CI are o u t s i d e the zero l i n e , thus i n d i c a t i n g s i g n i f i c a n c e
# 2 . Also note that a f t e r 7 or 8 months , the impact becomes i n s i g n i f i c a n t
# 3 . Given the data s e r i e s are s t a t i o n a r y , the response should d i e out
# ( converge t o zero )
# 4 . Hence the decay pattern i s good .
#8.2 Do the same f o r the second endogenous v a r i a b l e
i r f 2 <− i r f ( estim ,
impulse = " y2 " ,
response = " y1 " ,
n . ahead = 20 ,
boot = TRUE,
run = 200 ,
ci = 0.95)
p l o t ( i r f 2 , ylab = " y1 " , main = " y1 response t o y2 shock " )
#NOTES:
# 1 . F i r s t thing we n o t i c e i s that the 95 percent CI are around the zero l i n e
#( implies i n s i g n i f i c a n c e )
# 2 . Thus , y1 response t o a shock t o y2 i s not s t a t i s t i c a l l y s i g n i f i c a n t
# 3 . Variance Decomposition : i f there i s a shock t o a v a r i a b l e from
439
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#the main v a r i a b l e versus other p o s s i b l e v a r i a b l e s ,
#how much can be a t t r i b u t e d t o t h i s main v a r i a b l e
# 4 . Remember ,
# in the r e a l world o f d o c t o r a l research
#you many have many v a r i a b l e s plus s e v e r a l l a g s
#8.3 Use the fevd ( ) f u n c t i o n in the ’ vars ’ package
#and there are few arguments required thankfully ,
# so j u s t use number o f months ahead equal t o 10
vard <− fevd ( estim , n . ahead = 10)
p l o t ( vard )
#NOTES: the dark c o l o u r i s y1 and the l i g h t e r one i s y2
440
Electronic copy available at: https://ssrn.com/abstract=3863563
8.29. VAR, SVAR, VECM, SVECM
8.29
VAR, SVAR, VECM, SVECM
#VAR, SVAR, VECM, SVEC MODELS_EXERCISE
# 1 . INSTALL AND LOAD THE PACKAGE & DATA
#1.1 Need t o use the urca package t o a c c e s s the ADF f u n c t i o n ur . df ( )
i n s t a l l . packages ( " urca " )
i n s t a l l . packages ( " vars " )
l i b r a r y ( urca )
l i b r a r y ( vars )
#1.2 Can a c c e s s the b u i l t −in Canada data in t h i s package
data ( " Canada " )
summary ( Canada )
dim ( Canada )
s t r ( Canada )
head ( Canada , n=10)
t a i l ( Canada , n=10)
#1.3 Can p l o t t h i s data very q u i c k l y using the f o l l o w i n g command
p l o t ( Canada , nc = 2 , xlab = " " )
#1.4 In a next step , the authors conducted unit r o o t t e s t s
#by applying the Augmented Dickey−F u l l e r t e s t r e g r e s s i o n s
# t o the s e r i e s ( henceforth : ADF t e s t ) .
#The ADF t e s t has been implemented in the package urca as f u n c t i o n ur . df ( )
adf1 <− summary ( ur . df ( Canada [ , " prod " ] , type = " trend " , l a g s = 2 ) )
adf1
#1.5 This p r o c e s s may be repeated f o r each endogenous v a r i a b l e ,
# f o r d r i f t , and / or d i f f e r e n t l a g s
adf2 <−summary ( ur . df ( d i f f ( Canada [ , " prod " ] ) , type = " d r i f t " , l a g s = 1 ) )
adf3 <−summary ( ur . df ( Canada [ , " e " ] , type = " trend " , l a g s = 2 ) )
adf4 <−summary ( ur . df ( d i f f ( Canada [ , " e " ] ) , type = " d r i f t " , l a g s = 1 ) )
adf5 <−summary ( ur . df ( Canada [ , "U" ] , type = " d r i f t " , l a g s = 1 ) )
adf6 <−summary ( ur . df ( d i f f ( Canada [ , "U" ] ) , type = " none " , l a g s = 0 ) )
adf7 <−summary ( ur . df ( Canada [ , "rw " ] , type = " trend " , l a g s = 4 ) )
441
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
adf8 <−summary ( ur . df ( d i f f ( Canada [ , "rw " ] ) , type = " d r i f t " , l a g s = 3 ) )
adf9 <−summary ( ur . df ( d i f f ( Canada [ , "rw " ] ) , type = " d r i f t " , l a g s = 0 ) )
#1.6 Results can be viewed i n d i v i d u a l l y
adf2
adf3
#NOTES:
# 1 . And so on , o f t e n t h i s i s how you conduct research ,
#running a t e s t mul tiple times , e x t r a c t i n g key s t a t i s t i c s
# 2 . I t can be concluded that a l l time s e r i e s are i n t e g r a t e d o f order one .
# Please note , that the reported c r i t i c a l values d i f f e r s l i g h t l y from the ones
# that are reported in Breitung e t a l . ( 2 0 0 4 ) . They used d i f f e r e n t software .
# 3 . The authors used JMulTi in which the c r i t i c a l values o f Davidson
#and MacKinnon ( 1 9 9 3 ) are used .
# 4 . Used the ur . df ( ) f u n c t i o n where the c r i t i c a l values are
#taken from Dickey and F u l l e r ( 1 9 8 1 ) and Hamilton ( 1 9 9 4 ) .
# 2 . LAG SELECTION STEP
#2.1 Moved on t o determine the optimal lag length f o r an u n r e s t r i c t e d VAR
# f o r a maximal lag length o f e i g h t .
VARselect ( Canada , lag . max = 8 , type = " both " )
#NOTES:
# 1 . According t o the AIC and FPE the optimal lag number i s p = 3 ,
#whereas the HQ c r i t e r i o n i n d i c a t e s p = 2
#and the SC c r i t e r i o n i n d i c a t e s an optimal lag length o f p = 1
# 2 . They estimated f o r a l l three lag orders a VAR i n c l u d i n g a constant
#and a trend as d e t e r m i n i s t i c r e g r e s s o r s and conducted d i a g n o s t i c t e s t s
#with r e s p e c t t o the r e s i d u a l s .
# 3 . F i r s t , the v a r i a b l e s have t o be reordered
# in the same sequence as in Breitung e t a l . ( 2 0 0 4 ) .
#This step i s necessary ,
#because otherwise the r e s u l t s o f the m u l t i v a r i a t e Jarque−Bera t e s t ,
# in which a Choleski decomposition i s employed ,
#would d i f f e r s l i g h t l y from the reported ones in Breitung e t a l . ( 2 0 0 4 ) .
442
Electronic copy available at: https://ssrn.com/abstract=3863563
8.29. VAR, SVAR, VECM, SVECM
#2.2 VAR( 1 ) Code
Canada <− Canada [ , c ( " prod " , " e " , "U" , "rw " ) ]
p1ct <− VAR( Canada , p = 1 , type = " both " )
p1ct
summary ( p1ct , equation = " e " )
p l o t ( p1ct , names = " e " )
#2.3 Show the p l o t o f VAR( 1 ) f o r equation " e " in the PowerPoint s l i d e s
# VAR( 2 ) and VAR( 3 ) are shown below
p2ct <− VAR( Canada , p = 2 , type = " both " )
p3ct <− VAR( Canada , p = 3 , type = " both " )
# 3 . DIAGNOSTIC TESTS 1
#3.1 Next , i t i s shown how the d i a g n o s t i c t e s t s are conducted f o r the VAR( 1 ) model .
ser11 <− s e r i a l . t e s t ( p1ct , l a g s . pt = 16 , type = "PT . asymptotic " )
ser11$serial
norm1 <− normality . t e s t ( p1ct )
norm1$jb . mul
arch1 <− arch . t e s t ( p1ct , l a g s . multi = 5 )
arch1$arch . mul
p l o t ( arch1 , names = " e " )
p l o t ( s t a b i l i t y ( p1ct ) , nc = 2 )
# 4 . DIAGNOSTIC TESTS 2 and 3
#4.1 S e r i a l
ser31 <− s e r i a l . t e s t ( p3ct , l a g s . pt = 16 , type = "PT . asymptotic " ) $ s e r i a l
ser21 <− s e r i a l . t e s t ( p2ct , l a g s . pt = 16 , type = "PT . asymptotic " ) $ s e r i a l
ser11 <− s e r i a l . t e s t ( p1ct , l a g s . pt = 16 , type = "PT . asymptotic " ) $ s e r i a l
ser32 <− s e r i a l . t e s t ( p3ct , l a g s . pt = 16 , type = "PT . adjusted " ) $ s e r i a l
ser22 <− s e r i a l . t e s t ( p2ct , l a g s . pt = 16 , type = "PT . adjusted " ) $ s e r i a l
ser12 <− s e r i a l . t e s t ( p1ct , l a g s . pt = 16 , type = "PT . adjusted " ) $ s e r i a l
443
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#4.2 JB
norm3 <− normality . t e s t ( p3ct ) $jb . mul$JB
norm2 <− normality . t e s t ( p2ct ) $jb . mul$JB
norm1 <− normality . t e s t ( p1ct ) $jb . mul$JB
#4.3 ARCH
arch3 <− arch . t e s t ( p3ct , l a g s . multi = 5 ) $arch . mul
arch2 <− arch . t e s t ( p2ct , l a g s . multi = 5 ) $arch . mul
arch1 <− arch . t e s t ( p1ct , l a g s . multi = 5 ) $arch . mul
#NOTES: The r e s u l t s o f a l l d i a g n o s t i c t e s t s ,
# f o r the VAR( 1 ) , VAR( 2 ) and VAR( 3 ) model ,
# 5 . DIAGNOSTIC TESTS ANALYSIS
#NOTES:
# 1 . Given the d i a g n o s t i c t e s t r e s u l t s the authors concluded
# that a VAR(1) − s p e c i f i c a t i o n might be t o o r e s t r i c t i v e .
# 2 . They argued further , that although some o f the s t a b i l i t y t e s t s do i n d i c a t e
# d e v i a t i o n s from parameter constancy ,
#the time−i n v a r i a n t s p e c i f i c a t i o n o f the VAR( 2 ) and VAR( 3 ) model
# w i l l be maintained as t e n t a t i v e candidates
# f o r the f o l l o w i n g c o i n t e g r a t i o n a n a l y s i s .
# 3 . The authors estimated a VECM whereby a d e t e r m i n i s t i c trend has been included
# in the c o i n t e g r a t i o n r e l a t i o n .
# 4 . The estimation o f these models as well as the s t a t i s t i c a l i n f e r e n c e with
# r e s p e c t t o the c o i n t e g r a t i o n rank can be
# s w i f t l y accomplished with the f u n c t i o n ca . j o ( )
# 5 . Although the f o l l o w i n g R code examples
#are using f u n c t i o n s contained in the package urca ,
# i t i s however b e n e f i c i a l t o reproduce these r e s u l t s f o r two reasons :
444
Electronic copy available at: https://ssrn.com/abstract=3863563
8.29. VAR, SVAR, VECM, SVECM
#the i n t e r p l a y between the f u n c t i o n s contained in package urca
#and vars i s e x h i b i t e d and i t p rovid es an understanding o f the then
# f o l l o w i n g SVEC s p e c i f i c a t i o n .
summary ( ca . j o ( Canada , type = " t r a c e " , ecdet = " trend " , K = 3 , spec = " t r a n s i t o r y " ) )
summary ( ca . j o ( Canada , type = " t r a c e " , ecdet = " trend " , K = 2 , spec = " t r a n s i t o r y " ) )
#NOTES: These r e s u l t s do i n d i c a t e one c o i n t e g r a t i o n r e l a t i o n s h i p .
# 6 . VECTOR ERROR CORRECTION MODEL (VECM)
vecm . p3 <− summary ( ca . j o ( Canada , type = " t r a c e " , ecdet = " trend " ,
K = 3 , spec = " t r a n s i t o r y " ) )
vecm . p2 <− summary ( ca . j o ( Canada , type = " t r a c e " ,
ecdet = " trend " , K = 2 , spec = " t r a n s i t o r y " ) )
#6.1 In the code below , the VECM i s re−estimated
#with t h i s r e s t r i c t i o n
#and a normalization o f the long −run r e l a t i o n s h i p with r e s p e c t t o r e a l wages .
vecm <− ca . j o ( Canada [ , c ( " rw " , " prod " , " e " , "U" ) ] ,
type = " t r a c e " , ecdet = " trend " , K = 3 , spec = " t r a n s i t o r y " )
vecm . r1 <− c a j o r l s ( vecm , r = 1 )
#6.2 C a l c u l a t i o n o f t−values f o r alpha and beta
alpha <− c o e f ( vecm . r1$rlm ) [ 1 , ]
names ( alpha ) <− c ( " rw " , " prod " , " e " , "U" )
alpha
beta <− vecm . r1$beta
beta
r e s i d s <− r e s i d ( vecm . r1$rlm )
N <− nrow ( r e s i d s )
sigma <− cros spr od ( r e s i d s ) / N
#6.3 t−s t a t s f o r alpha ( c a l c u l a t e d by hand )
alpha . se <− s q r t ( s o l v e ( cr oss pr od ( cbind (vecm@ZK %*% beta ,
vecm@Z1 ) ) ) [ 1 , 1 ] * diag ( sigma ) )
445
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
names ( alpha . se ) <−
c ( " rw " , " prod " , " e " , "U" )
alpha . t <− alpha / alpha . se
alpha . t
#6.4 D i f f e r s l i g h t l y from c o e f ( summary ( vecm . r1$rlm ) )
#due t o degrees o f freedom adjustment
c o e f ( summary ( vecm . r1$rlm ) )
#6.5 t−s t a t s f o r beta
beta . se <− s q r t ( diag ( kronecker ( s o l v e ( cr oss pr od (vecm@RK[ , − 1 ] ) ) ,
s o l v e ( t ( alpha ) %*% s o l v e ( sigma ) %*% alpha ) ) ) )
beta . t <− c (NA, beta [ − 1] / beta . se )
names ( beta . t ) <− rownames ( vecm . r1$beta )
beta . t
# 7 .STRUCTURAL VECTOR ERROR−CORRECTION (SVEC) MODEL
#NOTES:
# 1 . For a j u s t i d e n t i f i e d SVEC model o f type B
#one needs 0 . 5K(K−1)=6 l i n e a r independent r e s t r i c t i o n s .
# 2 . I t i s f u r t h e r reasoned from the Beveridge−Nelson decomposition that
# there are k−s t a r =r (K−r )=3 shocks with permanent e f f e c t s and only one shock
# that e x e r t s a temporary e f f e c t , due t o r = 1 .
# 3 . Because the c o i n t e g r a t i o n r e l a t i o n i s i n t e r p r e t e d
#as a s t a t i o n a r y wagesetting r e l a t i o n ,
#the temporary shock i s a s s o c i a t e d with the wage shock v a r i a b l e .
# 4 . Hence , the 4 e n t r i e s in the l a s t column o f the long −run impact
#matrix XiB are s e t t o 0 .
# 5 . Because t h i s matrix i s o f reduced rank ,
# only k−s t a r ( r )=3 l i n e a r independent r e s t r i c t i o n s are imposed .
# 6 . I t i s t h e r e f o r e necessary t o
# s e t 0 . 5 k−s t a r ( k−star −1)=3 a d d i t i o n a l elements t o zero .
446
Electronic copy available at: https://ssrn.com/abstract=3863563
8.29. VAR, SVAR, VECM, SVECM
# 7 . The authors assumed constant returns o f s c a l e
#and t h e r e f o r e p r o d u c t i v i t y i s only driven by output shocks .
# 8 . This reasoning i m p l i e s zero c o e f f i c i e n t s
# in the f i r s t row o f the long −run matrix f o r the v a r i a b l e s employment ,
#unemployment and r e a l wages , hence the elements XiB sub ( 1 , j ) f o r
# j =2 ,3 ,4 are s e t t o 0 .
# 9 . Because XiB sub ( 1 , 4 ) has already been s e t t o zero ,
# only 2 a d d i t i o n a l r e s t r i c t i o n s have been added .
#10. The l a s t r e s t r i c t i o n i s imposed on the element B sub ( 4 , 2 )
#11. Here i t i s assumed that labour demand shocks
#do not e x e r t an immediate e f f e c t on r e a l wages .
#7.1 In the f o l l o w i n g R code example below the matrix o b j e c t s LR
#and SR are s e t up a c c o r d i n g l y
#and the j u s t −i d e n t i f i e d SVEC i s estimated with f u n c t i o n SVEC ( ) .
#In the c a l l t o the f u n c t i o n SVEC( ) the argument boot = TRUE
#has been employed such that bootstrapped standard e r r o r s
#and hence t s t a t i s t i c s can be computed
# f o r the s t r u c t u r a l long −run and contemporaneous c o e f f i c i e n t s .
vecm <− ca . j o ( Canada [ , c ( " prod " , " e " , "U" , "rw " ) ] , type = " t r a c e " ,
ecdet = " trend " , K = 3 , spec = " t r a n s i t o r y " )
SR <− matrix (NA, nrow = 4 , n c o l = 4 )
SR[ 4 , 2 ] <− 0
LR <− matrix (NA, nrow = 4 , n c o l = 4 )
LR[ 1 , 2 : 4 ] <− 0
LR[ 2 : 4 , 4 ] <− 0
svec <− SVEC( vecm , LR = LR, SR = SR, r = 1 , l r t e s t = FALSE,
447
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
boot = TRUE, runs = 100)
summary ( svec )
#NOTES: The values o f the t s t a t i s t i c s d i f f e r s l i g h t l y from
#the reported ones in Breitung e t a l . ( 2 0 0 4 )
#which can be a t t r i b u t e d t o sampling .
#In the R code example above only 100 runs have been executed ,
#whereas in Breitung e t a l . ( 2 0 0 4 ) 2000 r e p e t i t i o n s have been used .
#7.2 SR Table
SR <− round ( svec$SR , 2 )
SRt <− round ( svec$SR / svec$SRse , 2 )
#7.3 LR Table
LR <− round ( svec$LR , 2 )
LRt <− round ( svec$LR / svec$LRse , 2 )
#NOTES:
# 1 . The authors i n v e s t i g a t e d f u r t h e r i f l a b o r supply shocks
#do have no long −run impact on unemployment .
# 2 . This hypothesis i s mirrored by s e t t i n g XiB sub ( 3 , 3 ) = 0 .
# 3 . Because one more zero r e s t r i c t i o n has been added
# t o the long −run impact matrix , the SVEC model i s now over−i d e n t i f i e d .
# 4 . The v a l i d i t y o f t h i s over− i d e n t i f i c a t i o n r e s t r i c t i o n
#can be t e s t e d with a LR t e s t .
#7.4 In the R code example below f i r s t the a d d i t i o n a l r e s t r i c t i o n
#has been s e t and then the SVEC i s re−estimated by using the update method .
#The r e s u l t o f the LR t e s t i s contained in the returned l i s t as named element LRove
LR[ 3 , 3 ] <− 0
svec . o i <− update ( svec , LR = LR, l r t e s t = TRUE, boot = FALSE)
svec . oi$LRover
448
Electronic copy available at: https://ssrn.com/abstract=3863563
8.29. VAR, SVAR, VECM, SVECM
#NOTES:
# 1 . The value o f the t e s t s t a t i s t i c i s 6.07
#and the p value o f t h i s Chi−squared (1) − d i s t r i b u t e d v a r i a b l e i s 0 . 0 1 4 .
# 2 . Therefore , the n u l l hypothesis that shocks t o the l a b o r supply
#do not e x e r t a long −run e f f e c t on unemployment
#has t o be r e j e c t e d f o r a s i g n i f i c a n c e l e v e l o f 5 percent .
# 8 . SVEC − Impusle Response Function ( IRF )
#NOTES:
# 1 . In order t o i n v e s t i g a t e the dynamic e f f e c t s on unemployment ,
#the authors applied an impulse response a n a l y s i s .
# 2 . The impulse response a n a l y s i s shows the e f f e c t s o f the d i f f e r e n t shocks ,
# i . e . , output , l a b o r demand , l a b o r supply and wage , t o unemployment .
#8.1 In the R code example below the i r f method f o r o b j e c t s
#with c l a s s a t t r i b u t e s v e c e s t i s employed
#and the argument boot = TRUE has been s e t
#such that c o n f i d e n c e bands around the impulse response t r a j e c t o r i e s
#can be c a l c u l a t e d .
svec . i r f <− i r f ( svec , response = "U" , n . ahead = 48 , boot = TRUE)
p l o t ( svec . i r f )
# 9 . SVEC − FEVD
#9.1 In a f i n a l step , a f o r e c a s t e r r o r variance decomposition i s conducted
#with r e s p e c t t o unemployment .
#This i s achieved by applying the fevd method
# t o the o b j e c t with c l a s s a t t r i b u t e s v e c e s t .
fevd .U <− fevd ( svec , n . ahead = 48)$U
449
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.30
Vector autoregression example
#VECTOR AUTOREGRESSION EXAMPLE − U. S . MACRO DATA
#0.1 Load the t i d y v e r s e
#and readxl l i b r a r y t o a c c e s s the read_xlsx f u n c t i o n a l i t y
library ( tidyverse )
l i b r a r y ( readxl )
# 1 . STAGE 1
#1.1 Load the U. S . macroeconomic data s e t
USMacroSWQ <− read_xlsx ( f i l e . choose ( ) ,
sheet = 1 ,
c o l _ t y p e s = c ( " t e x t " , rep ( " numeric " , 9 ) ) )
#1.2 Set the column names
colnames (USMacroSWQ) <− c ( " Date " , "GDPC96" , "JAPAN_IP " , "PCECTPI" , "GS10 " ,
"GS1" , "TB3MS" , "UNRATE" , "EXUSUK" , "CPIAUCSL" )
#1.3 Load the zoo l i b r a r y t o a c c e s s yearqtr ( ) and as . yearqtr ( ) f u n c t i o n s
l i b r a r y ( zoo )
#1.4 Format the date column
USMacroSWQ$Date <− as . yearqtr ( USMacroSWQ$Date , format = "%Y:0%q " )
#1.5 Define GDP as t s o b j e c t
GDP <− t s (USMacroSWQ$GDPC96,
s t a r t = c (1957 , 1 ) ,
end = c (2013 , 4 ) ,
frequency = 4 )
#1.6 Define GDP growth as a t s o b j e c t
GDPGrowth <− t s (400 * l o g (GDP[ − 1]/GDP[− length (GDP) ] ) ,
s t a r t = c (1957 , 2 ) ,
end = c (2013 , 4 ) ,
frequency = 4 )
#1.7 3−months Treasury b i l l i n t e r e s t r a t e as a ’ ts ’ o b j e c t
450
Electronic copy available at: https://ssrn.com/abstract=3863563
8.30. VECTOR AUTOREGRESSION EXAMPLE
TB3MS <− t s (USMacroSWQ$TB3MS,
s t a r t = c (1957 , 1 ) ,
end = c (2013 , 4 ) ,
frequency = 4 )
#1.8 10− years Treasury bonds i n t e r e s t r a t e as a ’ ts ’ o b j e c t
TB10YS <− t s (USMacroSWQ$GS10,
s t a r t = c (1957 , 1 ) ,
end = c (2013 , 4 ) ,
frequency = 4 )
#1.9 Generate the term spread s e r i e s
TSpread <− TB10YS − TB3MS
# 2 . STAGE 2
#NOTES: Estimate both equations s e p a r a t e l y by OLS
#and use c o e f t e s t ( ) t o obtain robust standard e r r o r s .
#2.1 F i r s t , i n s t a l l the dynlm package ( dynamic l i n e a r model )
i n s t a l l . packages ( " dynlm " )
l i b r a r y ( dynlm )
#2.2 Estimate both equations using ’ dynlm ( ) ’
VAR_EQ1 <− dynlm ( GDPGrowth ~ L ( GDPGrowth , 1 : 2 ) + L ( TSpread , 1 : 2 ) ,
s t a r t = c (1981 , 1 ) ,
end = c (2012 , 4 ) )
VAR_EQ2 <− dynlm ( TSpread ~ L ( GDPGrowth , 1 : 2 ) + L ( TSpread , 1 : 2 ) ,
s t a r t = c (1981 , 1 ) ,
end = c (2012 , 4 ) )
#2.3 Rename r e g r e s s o r s f o r b e t t e r r e a d a b i l i t y
names ( VAR_EQ1$coefficients ) <− c ( " I n t e r c e p t " , " Growth_t − 1" ,
" Growth_t − 2" , " TSpread_t − 1" , " TSpread_t − 2")
names ( VAR_EQ2$coefficients ) <− names ( VAR_EQ1$coefficients )
451
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#2.4 Robust c o e f f i c i e n t summaries
# 2 . 4 . 1 F i r s t invoke lmtest l i b r a r y f o r c o e f t e s t ( ) function ,
#and sandwich l i b r a r y f o r e x t r a c t i n g VAR−COV matrix
l i b r a r y ( lmtest )
l i b r a r y ( sandwich )
c o e f t e s t (VAR_EQ1, vcov . = sandwich )
c o e f t e s t (VAR_EQ2, vcov . = sandwich )
# 3 . STAGE 3
#NOTES: The f u n c t i o n VAR( ) can be used t o
# obtain the same c o e f f i c i e n t estimates as presented above
# s i n c e i t a p p l i e s OLS per equation , t o o .
#3.1 s e t up data f o r estimation using VAR( )
VAR_data <− window ( t s . union ( GDPGrowth , TSpread ) ,
s t a r t = c (1980 , 3 ) ,
end = c (2012 , 4 ) )
#3.2 Invoke the vars l i b r a r y t o use the VAR( ) f u n c t i o n
l i b r a r y (MASS)
l i b r a r y ( strucchange )
l i b r a r y ( urca )
l i b r a r y ( vars )
#3.3 estimate model c o e f f i c i e n t s using VAR( )
VAR_est <− VAR( y = VAR_data , p = 2 )
VAR_est
#NOTES: VAR( ) returns a l i s t o f lm o b j e c t s
#which can be passed t o the usual f u n c t i o n s ,
# f o r example summary ( ) and so i t i s s t r a i g h t f o r w a r d t o obtain model s t a t i s t i c s
# f o r the i n d i v i d u a l equations .
#3.4 Obtain the adjusted R−squared from the output o f VAR( )
summary ( VAR_est$varresult$GDPGrowth ) $adj . r . squared
summary ( VAR_est$varresult$TSpread ) $adj . r . squared
452
Electronic copy available at: https://ssrn.com/abstract=3863563
8.30. VECTOR AUTOREGRESSION EXAMPLE
# 4 . STAGE 4
#NOTES: We may use the i n d i v i d u a l model o b j e c t s t o conduct Granger−c a u s a l i t y t e s t s .
#4.1 F i r s t , we need the car package f o r the linearHypothesis ( ) f u n c t i o n
l i b r a r y ( carData )
l i b r a r y ( car )
#4.2 Granger−c a u s a l i t y t e s t s :
#4.3 Test i f term spread has no power in e x p l a i n i n g GDP growth
linearHypothesis (VAR_EQ1,
hypothesis . matrix = c ( " TSpread_t − 1" , " TSpread_t − 2") ,
vcov . = sandwich )
#4.4 Test i f GDP growth has no power in e x p l a i n i n g term spread
linearHypothesis (VAR_EQ2,
hypothesis . matrix = c ( " Growth_t − 1" , " Growth_t − 2") ,
vcov . = sandwich )
#NOTES:
# 1 . Both Granger−c a u s a l i t y t e s t s r e j e c t at the l e v e l o f 5 percent .
#This i s evidence in f a v o r o f the c o n j e c t u r e
# that the term spread has power in e x p l a i n i n g GDP growth and v i c e −versa .
#STAGE 5 ITERATED MULTIPERIOD FORECASTS
#NOTES: The f o l l o w i n g code shows how t o compute i t e r a t e d f o r e c a s t s
# f o r GDP growth and the term spread up t o period 2015:Q1,
# that i s h=10 , using the model o b j e c t VAR_est .
#5.1 Compute i t e r a t e d f o r e c a s t s f o r GDP growth
#and term spread f o r the next 10 quarters
f o r e c a s t s <− p r e d i c t ( VAR_est )
forecasts
453
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#NOTES:
# 1 . This r e v e a l s that the two−quarter −ahead f o r e c a s t o f GDP growth in 2013:Q2
#using data through 2012:Q4 i s 1 . 6 9 .
# 2 . For the same period , the i t e r a t e d VAR f o r e c a s t f o r the term spread i s 1 . 8 8 .
# 3 . The matrices returned by p r e d i c t ( VAR_est )
# a l s o i n c l u d e 95 percentage p r e d i c t i o n i n t e r v a l s
# ( however , the f u n c t i o n does not a d j u s t f o r a u t o c o r r e l a t i o n
# or h e t e r o s k e d a s t i c i t y o f the e r r o r s ! ) .
# 4 . Also p l o t the i t e r a t e d f o r e c a s t s
# f o r both v a r i a b l e s by c a l l i n g p l o t ( ) on the output o f p r e d i c t ( VAR_est )
#5.2 V i s u a l i z e the i t e r a t e d f o r e c a s t s
plot ( forecasts )
# 6 . STAGE 6 DIRECT MULTIPERIOD FORECASTS
#NOTES:
# To obtain two−quarter −ahead f o r e c a s t s o f GDP growth
#and the term spread we f i r s t estimate the equations
#and then s u b s t i t u t e the values GDPGR( 2 0 1 2 :Q4 ) , GDPGR( 2 0 1 2 :Q3 ) , TSpread ( 2 0 1 2 :Q4)
#and TSpread ( 2 0 1 2 :Q3) i n t o both equations .
#6.1 Estimate models f o r d i r e c t two−quarter −ahead f o r e c a s t s
VAR_EQ1_direct <− dynlm ( GDPGrowth ~ L ( GDPGrowth , 2 : 3 ) + L ( TSpread , 2 : 3 ) ,
s t a r t = c (1981 , 1 ) , end = c (2012 , 4 ) )
VAR_EQ2_direct <− dynlm ( TSpread ~ L ( GDPGrowth , 2 : 3 ) + L ( TSpread , 2 : 3 ) ,
s t a r t = c (1981 , 1 ) , end = c (2012 , 4 ) )
#6.2 Compute d i r e c t two−quarter −ahead f o r e c a s t s
c o e f ( VAR_EQ1_direct ) %*% c ( 1 , # i n t e r c e p t
window ( GDPGrowth , s t a r t = c (2012 , 3 ) ,
end = c (2012 , 4 ) ) ,
window ( TSpread , s t a r t = c (2012 , 3 ) ,
end = c (2012 , 4 ) ) )
454
Electronic copy available at: https://ssrn.com/abstract=3863563
8.30. VECTOR AUTOREGRESSION EXAMPLE
c o e f ( VAR_EQ2_direct ) %*% c ( 1 , # i n t e r c e p t
window ( GDPGrowth , s t a r t = c (2012 , 3 ) ,
end = c (2012 , 4 ) ) ,
window ( TSpread , s t a r t = c (2012 , 3 ) ,
end = c (2012 , 4 ) ) )
#NOTES: Applied economists o f t e n use the i t e r a t e d method s i n c e t h i s f o r e c a s t s
#are more r e l i a b l e in terms o f MSFE provided
# that the one−period −ahead model i s c o r r e c t l y s p e c i f i e d .
# I f t h i s i s not the case , f o r example because one equation in a VAR i s b e l i e v e d
# t o be mis−s p e c i f i e d , i t can be b e n e f i c i a l t o use d i r e c t f o r e c a s t s
# s i n c e the i t e r a t e d method w i l l then be biased
#and thus have a higher MSFE than the d i r e c t method .
# ( see Stock and Watson t e x t f o r more explanation ) .
# 7 . ORDERS OF INTEGRATION AND THE DF−GLS UNIT ROOT TEST
#7.1 Define t s o b j e c t o f the U. S . PCE P r i c e Index
PCECTPI <− t s ( l o g (USMacroSWQ$PCECTPI) ,
s t a r t = c (1957 , 1 ) ,
end = c (2012 , 4 ) ,
freq = 4)
#7.2 P l o t logarithm o f the PCE P r i c e Index
p l o t ( l o g (PCECTPI ) ,
main = " Log o f United States PCE P r i c e Index " ,
ylab = " Logarithm " ,
col = " steelblue " ,
lwd = 2 )
#NOTES:
# 1 . The logarithm o f the p r i c e l e v e l has a smoothly varying trend .
#This i s t y p i c a l f o r an I ( 2 ) s e r i e s .
# 2 . I f the p r i c e l e v e l i s indeed I ( 2 ) ,
#the f i r s t d i f f e r e n c e s o f t h i s s e r i e s should be I ( 1 ) .
# 3 . Since we are c o n s i d e r i n g the logarithm o f the p r i c e l e v e l ,
455
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#we obtain growth r a t e s by taking f i r s t d i f f e r e n c e s .
# 4 . Therefore , the d i f f e r e n c e d p r i c e l e v e l s e r i e s
# i s the s e r i e s o f q u a r t e r l y i n f l a t i o n r a t e s .
#7.3 This i s q u i c k l y done in R
#using the f u n c t i o n Delt ( ) from the package quantmod .
library ( xts )
l i b r a r y (TTR)
l i b r a r y ( quantmod )
# Note :
m ult ipl yin g the q u a r t e r l y i n f l a t i o n r a t e s
#by 400 y i e l d s the q u a r t e r l y r a t e o f i n f l a t i o n ,
#measured in percentage p o i n t s at an annual r a t e .
#7.4 P l o t U. S . PCE p r i c e i n f l a t i o n
p l o t (400 * Delt (PCECTPI ) ,
main = " United States PCE P r i c e Index " ,
ylab = " Percent per annum" ,
col = " steelblue " ,
lwd = 2 )
#7.5 Add a dashed l i n e at y =
0
abline (0 , 0 , l t y = 2)
#NOTES: The i n f l a t i o n r a t e behaves much more e r r a t i c a l l y
#than the smooth graph o f the logarithm o f the PCE p r i c e index .
# 8 . THE DF−GLS TEST FOR A UNIT ROOT.
#NOTES:
# 1 . The DF−GLS t e s t f o r a unit r o o t has been developed by E l l i o t t ,
#Rothenberg , and Stock ( 1 9 9 6 ) and has higher power than the ADF t e s t
#when the a u t o r e g r e s s i v e r o o t i s l a r g e but l e s s than one .
#That i s , the DF−GLS has a higher p r o b a b i l i t y o f r e j e c t i n g the f a l s e n u l l o
# f a s t o c h a s t i c trend when the sample data stems from a time s e r i e s that
# i s c l o s e t o being i n t e g r a t e d .
# 2 . The idea o f the DF−GLS t e s t i s t o t e s t f o r an a u t o r e g r e s s i v e unit r o o t
456
Electronic copy available at: https://ssrn.com/abstract=3863563
8.30. VECTOR AUTOREGRESSION EXAMPLE
# in the detrended s e r i e s ,
#whereby GLS estimates o f the d e t e r m i n i s t i c components
#are used t o obtain the detrended v e r s i o n o f the o r i g i n a l s e r i e s .
# 3 . A f u n c t i o n that performs the DF−GLS t e s t i s implemented in the package urca
# ( t h i s package i s a dependency o f the package vars
# so i t should be already loaded i f vars i s attached ) .
#The f u n c t i o n that computes the t e s t s t a t i s t i c i s ur . ers .
#8.1 DF−GLS t e s t f o r unit r o o t in GDP
summary ( ur . ers ( l o g ( window (GDP, s t a r t = c (1962 , 1 ) , end = c (2012 , 4 ) ) ) ,
model = " trend " ,
lag . max = 2 ) )
#NOTES:
# 1 . The summary o f the t e s t shows that the t e s t s t a t i s t i c i s about −1.2
# 2 . The 10 percent c r i t i c a l value f o r the DF−GLS i s − 2.57
# 3 . This i s , however , not the appropriate c r i t i c a l value
# f o r the ADF t e s t when an i n t e r c e p t and a time trend are included
# in the Dickey−F u l l e r r e g r e s s i o n : the asymptotic d i s t r i b u t i o n s
# o f both t e s t s t a t i s t i c s d i f f e r and so do t h e i r c r i t i c a l values
# 4 . The t e s t i s l e f t −sided so we cannot r e j e c t the n u l l hypothesis that U. S .
# i n f l a t i o n i s non−s t a t i o n a r y , using the DF−GLS t e s t .
# 9 . COINTEGRATION
#NOTES:
# 1 . As an example , r e c o n s i d e r the the r e l a t i o n between short −term
#and long −term i n t e r e s t r a t e s by the example o f US 3−month treasury b i l l s ,
#US 10− years treasury bonds and the spread in t h e i r i n t e r e s t r a t e s
#9.1 P l o t both i n t e r e s t s e r i e s
p l o t ( merge ( as . zoo (TB3MS) , as . zoo ( TB10YS ) ) ,
457
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
p l o t . type = " s i n g l e " ,
lty = c (2 , 1) ,
lwd = 2 ,
xlab = " Date " ,
ylab = " Percent per annum" ,
ylim = c ( − 5 , 1 7 ) ,
main = " I n t e r e s t Rates " )
#9.2 Add the term spread s e r i e s
l i n e s ( as . zoo ( TSpread ) ,
col = " steelblue " ,
lwd = 2 ,
xlab = " Date " ,
ylab = " Percent per annum" ,
main = "Term Spread " )
#9.3 Shade the term spread
polygon ( c ( time (TB3MS) , rev ( time (TB3MS ) ) ) ,
c ( TB10YS , rev (TB3MS) ) ,
c o l = alpha ( " s t e e l b l u e " , alpha = 0 . 3 ) ,
border = NA)
#9.4 Add h o r i z o n t a l l i n e add 0
abline (0 , 0)
#9.5 Add a legend
legend ( " t o p r i g h t " ,
legend = c ( "TB3MS" , "TB10YS " , "Term Spread " ) ,
c o l = c ( " black " , " black " , " s t e e l b l u e " ) ,
lwd = c ( 2 , 2 , 2 ) ,
lty = c (2 , 1 , 1))
#NOTES:
# The p l o t suggests that long −term and short −term i n t e r e s t r a t e s
#are c o i n t e g r a t e d : both i n t e r e s t s e r i e s seem t o have the same long−run behavior .
#They share a common s t o c h a s t i c trend . The term spread ,
#which i s obtained by taking the d i f f e r e n c e between long−term
#and short −term i n t e r e s t rates , seems t o be s t a t i o n a r y .
#In f a c t , the e x p e c t a t i o n s theory o f the term s t r u c t u r e suggests
458
Electronic copy available at: https://ssrn.com/abstract=3863563
8.30. VECTOR AUTOREGRESSION EXAMPLE
#the c o i n t e g r a t i n g c o e f f i c i e n t theta t o be 1 .
#This i s c o n s i s t e n t with the v i s u a l r e s u l t .
#10. TESTING FOR COINTEGRATION
#NOTES:
#1. Application to interest rates .
# 2 . The theory o f the term s t r u c t u r e suggests that long −term
#and short −term i n t e r e s t r a t e s are c o i n t e g r a t e d
#with a c o i n t e g r a t i o n c o e f f i c i e n t o f Theta = 1
# 3 . Above , there i s v i s u a l evidence f o r t h i s c o n j e c t u r e as the spread o f 10− year
#and 3−month i n t e r e s t r a t e s l o o k s s t a t i o n a r y .
# 4 . Continue by using formal t e s t s ( the ADF and the DF−GLS t e s t )
# t o see whether the i n d i v i d u a l i n t e r e s t r a t e s e r i e s are i n t e g r a t e d
#and i f t h e i r d i f f e r e n c e i s s t a t i o n a r y ( Theta = 1 ) .
# 5 . Both i s c o n v e n i e n t l y done by using the f u n c t i o n s ur . df ( ) f o r
#computation o f the ADF t e s t and ur . ers f o r conducting the DF−GLS t e s t
# 6 . Following the book we use data from 1962:Q1 t o 2012:Q4
#and employ models that i n c l u d e a d r i f t term .
#We s e t the maximum lag order t o 6 and use the AIC
# f o r s e l e c t i o n o f the optimal lag length .
#10.1 Test f o r non−s t a t i o n a r i t y o f 3−month treasury b i l l s using ADF t e s t
ur . df ( window (TB3MS, c (1962 , 1 ) , c (2012 , 4 ) ) ,
lags = 6 ,
s e l e c t l a g s = "AIC " ,
type = " d r i f t " )
#10.2 Test f o r non−s t a t i o n a r i t y o f 10− years treasury bonds using ADF t e s t
ur . df ( window ( TB10YS , c (1962 , 1 ) , c (2012 , 4 ) ) ,
lags = 6 ,
s e l e c t l a g s = "AIC " ,
459
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
type = " d r i f t " )
#10.3 Test f o r non−s t a t i o n a r i t y o f 3−month treasury b i l l s using DF−GLS t e s t
ur . ers ( window (TB3MS, c (1962 , 1 ) , c (2012 , 4 ) ) ,
model = " constant " ,
lag . max = 6 )
#10.4 Test f o r non−s t a t i o n a r i t y o f 10− years treasury bonds using DF−GLS t e s t
ur . ers ( window ( TB10YS , c (1962 , 1 ) , c (2012 , 4 ) ) ,
model = " constant " ,
lag . max = 6 )
#NOTES:
# 1 . The corresponding 10 percent c r i t i c a l value f o r both t e s t s i s − 2.57
# so we cannot r e j e c t the n u l l hypotheses o f non−s t a t i o n a r y
# f o r e i t h e r s e r i e s ( even at the 10 percent l e v e l o f s i g n i f i c a n c e ) .
# 2 . Conclude that i t i s p l a u s i b l e t o model both i n t e r e s t r a t e s e r i e s as I ( 1 ) .
# 3 . Next , apply the ADF and the DF−GLS t e s t t o t e s t f o r non−s t a t i o n a r i t y
# o f the term spread s e r i e s ,
#which means we t e s t f o r non−c o i n t e g r a t i o n
# o f long −term and short −term i n t e r e s t r a t e s .
#10.5 Test i f term spread i s s t a t i o n a r y
# ( c o i n t e g r a t i o n o f i n t e r e s t r a t e s ) using ADF
ur . df ( window ( TB10YS , c (1962 , 1 ) , c (2012 , 4 ) ) − window (TB3MS,
c (1962 , 1 ) ,
c (2012 , 4 ) ) ,
lags = 6 ,
s e l e c t l a g s = "AIC " ,
type = " d r i f t " )
#10.6 Test i f term spread i s s t a t i o n a r y ( c o i n t e g r a t i o n o f i n t e r e s t r a t e s )
#using DF−GLS t e s t
ur . ers ( window ( TB10YS , c (1962 , 1 ) , c (2012 , 4 ) ) − window (TB3MS,
c (1962 , 1 ) ,
c (2012 , 4 ) ) ,
460
Electronic copy available at: https://ssrn.com/abstract=3863563
8.30. VECTOR AUTOREGRESSION EXAMPLE
model = " constant " ,
lag . max = 6 )
#NOTES:
# 1 . Both t e s t s r e j e c t the hypothesis o f non−s t a t i o n a r i t y
# o f the term spread s e r i e s at the 1 percent l e v e l o f s i g n i f i c a n c e ,
#which i s strong evidence in favour o f the hypothesis
# that the term spread i s s t a t i o n a r y ,
# implying c o i n t e g r a t i o n o f long − and short −term i n t e r e s t r a t e s .
# 2 . Since theory suggests that theta =1 , there i s no need t o estimate theta ,
# so i t i s not necessary t o use the EG−ADF t e s t which allows theta t o be unknown .
#However , s i n c e i t i s i n s t r u c t i v e t o do so ,
#we f o l l o w the book and compute t h i s t e s t s t a t i s t i c .
#The f i r s t −stage OLS r e g r e s s i o n i s :
#10.7 Estimate f i r s t −stage r e g r e s s i o n o f EG−ADF t e s t
FS_EGADF <− dynlm ( window ( TB10YS , c (1962 , 1 ) , c (2012 , 4 ) ) ~ window (TB3MS,
c (1962 , 1 ) ,
c (2012 , 4 ) ) )
FS_EGADF
#NOTES: Thus we now have an estimate o f TB10YS ( t ) = 2.46 + 0.81 TB3MS( t )
#where theta estimate = 0.81
#10.8 Next take the r e s i d u a l s e r i e s and compute the ADF t e s t s t a t i s t i c .
# 1 0 . 8 . 1 Compute the r e s i d u a l s
z_hat <− r e s i d (FS_EGADF)
# 1 0 . 8 . 2 Compute the ADF t e s t s t a t i s t i c
ur . df ( z_hat , l a g s = 6 , type = " none " , s e l e c t l a g s = "AIC " )
#NOTES:
# 1 . The t e s t s t a t i s t i c i s − 3.19 which i s smaller than
#the 10 percent c r i t i c a l value but l a r g e r than the 5 percent c r i t i c a l value .
# 2 . Thus , the n u l l hypothesis o f no c o i n t e g r a t i o n can be r e j e c t e d
461
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# at the 10 percent l e v e l bu not at the 5 percent l e v e l .
# 3 . This i n d i c a t e s lower power o f the EG−ADF t e s t due t o the estimation o f theta ,
#when theta =1 i s the c o r r e c t value ,
#we expect the power o f the ADF t e s t f o r a unit r o o t in the r e s i d u a l s s e r i e s
#z−hat = TB10YS − TB3MS t o be higher than when some estimated theta i s used .
#11. VECTOR ERROR CORECTION MODEL (VECM) f o r TB10YS ( t ) and TB3MS
#NOTES: A VECM can be used t o model the two i n t e r e s t r a t e s considered
# in the previous s e c t i o n s .
#11.1 Following s p e c i f y the VECM t o i n c l u d e two l a g s o f both s e r i e s
#as r e g r e s s o r s and choose theta =1 , as theory suggests
TB10YS <− window ( TB10YS , c (1962 , 1 ) , c (2012 , 4 ) )
TB3MS <− window (TB3MS, c (1962 , 1 ) , c (2012 , 4 ) )
#11.2 Set up e r r o r c o r r e c t i o n term
VECM_ECT <− TB10YS − TB3MS
#11.3 Estimate both equations o f the VECM using ’ dynlm ( ) ’
VECM_EQ1 <− dynlm ( d ( TB10YS ) ~ L ( d (TB3MS) , 1 : 2 ) + L ( d ( TB10YS ) , 1 : 2 ) +
L (VECM_ECT) )
VECM_EQ2 <− dynlm ( d (TB3MS) ~ L ( d (TB3MS) , 1 : 2 ) + L ( d ( TB10YS ) , 1 : 2 ) +
L (VECM_ECT) )
#11.4 Rename r e g r e s s o r s f o r b e t t e r r e a d a b i l i t y
names ( VECM_EQ1$coefficients ) <− c ( " I n t e r c e p t " , "D_TB3MS_l1 " , "D_TB3MS_l2 " ,
" D_TB10YS_l1 " , " D_TB10YS_l2 " , " e c t _ l 1 " )
names ( VECM_EQ2$coefficients ) <− names ( VECM_EQ1$coefficients )
#11.5 C o e f f i c i e n t summaries using HAC standard e r r o r s
c o e f t e s t (VECM_EQ1, vcov . = NeweyWest (VECM_EQ1, prewhite = F , a d j u s t = T ) )
c o e f t e s t (VECM_EQ2, vcov . = NeweyWest (VECM_EQ2, prewhite = F , a d j u s t = T ) )
462
Electronic copy available at: https://ssrn.com/abstract=3863563
8.31. MULTILEVEL LINEAR MODELS (MLM)
8.31
Multilevel linear models (MLM)
###MULTILEVEL LINEAR MODELS (MLM)_EXERCISE
# 1 . SIMPLE (INTERCEPT−ONLY) MLM
#0.1 I n s t a l l and load the packages and data
library ( tidyverse )
l i b r a r y ( lme4 )
# Achieve <− read_csv ( "C : / R /mlm/ Achieve . csv " )
Achieve <− read . csv ( f i l e . choose ( ) )Mu
summary ( Achieve$geread )
#1.1 The lmer syntax necessary f o r estimating the n u l l model appears below .
Model3 . 0 <− lmer ( geread ~ 1 +(1| s c h o o l ) , data=Achieve )
#1.2 Can obtain output from t h i s model by typing summary ( Model3 . 0 )
summary ( Model3 . 0 )
#1.3 In order t o f i t a model with vocabulary t e s t s c o r e
#as the independent v a r i a b l e using lmer , submit the f o l l o w i n g syntax in R .
Model3 . 1 <− lmer ( geread ~ gevocab + (1| s c h o o l ) , data=Achieve )
#NOTES:
# 1 . In the f i r s t part o f the f u n c t i o n c a l l we d e f i n e the formula
# or the model f i x e d e f f e c t s ,
#which i s very s i m i l a r t o model d e f i n i t i o n o f l i n e a r r e g r e s s i o n using lm ( ) .
# 2 . The statement geread~gevocab e s s e n t i a l l y says
# that the reading s c o r e i s p r e d i c t e d with the vocabulary s c o r e f i x e d e f f e c t .
# 3 . The c a l l in parentheses d e f i n e s the random e f f e c t s and the nesting s t r u c t u r e .
# 4 . I f only a random i n t e r c e p t i s desired , the syntax f o r the i n t e r c e p t i s " 1 . "
# 5 . In t h i s example , (1 − bar−s c h o o l ) i n d i c a t e s
# that only a random i n t e r c e p t s model w i l l be used and
# that the random i n t e r c e p t v a r i e s within s c h o o l .
463
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#This corresponds t o the data s t r u c t u r e o f students nested within s c h o o l s .
#1.4 F i t t i n g t h i s model , which i s saved in the output o b j e c t Model3 . 1 ,
#we obtain the f o l l o w i n g
summary ( Model3 . 1 )
#1.5 Adding the l e v e l −2 p r e d i c t o r , s c h o o l enrollment ,
#would r e s u l t in the f o l l o w i n g
Model3 . 2 <− lmer ( geread ~ gevocab + s e n r o l l + (1| s c h o o l ) , data=Achieve )
summary ( Model3 . 2 )
# 2 . INTERACTIONS AND CROSS−LEVEL INTERACTIONS USING R
#2.1 Following are examples f o r f i t t i n g an i n t e r a c t i o n model
# f o r two l e v e l −1 v a r i a b l e s ( Model3 . 3 )
#and a cross −l e v e l i n t e r a c t i o n i n v o l v i n g l e v e l −1 and l e v e l −2 v a r i a b l e s ( Model3 . 4 )
Model3 . 3 <− lmer ( geread ~ gevocab + age + gevocab * age +
(1| s c h o o l ) , data=Achieve )
#2.2 Next , l e t ’ s examine a model that i n c l u d e s a cross −l e v e l i n t e r a c t i o n .
Model3 . 4 <− lmer ( geread ~ gevocab + s e n r o l l + gevocab * s e n r o l l +
(1| s c h o o l ) , data=Achieve )
#2.3 Model3 . 3 d e f i n e s a m u l t i l e v e l model
#where two l e v e l −1 ( student l e v e l ) p r e d i c t o r s i n t e r a c t with one another .
summary ( Model3 . 3 )
#2.4 Model3 . 4 d e f i n e s a m u l t i l e v e l model with a cross −l e v e l i n t e r a c t i o n :
#where a l e v e l −1 ( student l e v e l )
#and a l e v e l −2( s c h o o l l e v e l ) p r e d i c t o r are i n t e r a c t i n g .
summary ( Model3 . 4 )
#NOTES: As can be seen ,
# there i s no d i f f e r e n c e in the treatment o f v a r i a b l e s
# at d i f f e r e n t l e v e l s when computing i n t e r a c t i o n s
# 3 . RANDOM COEFFICIENT MODELS USING R
464
Electronic copy available at: https://ssrn.com/abstract=3863563
8.31. MULTILEVEL LINEAR MODELS (MLM)
#3.1 Modify Model3 . 1 with Model3 . 5 below
#and allow the e f f e c t o f gevocab t o be random
Model3 . 5 <− lmer ( geread ~ gevocab + ( gevocab| s c h o o l ) ,
data=Achieve )
summary ( Model3 . 5 )
#3.2 Models 3 . 6 and 3 . 7
Model3 . 6 <− lmer ( geread ~ gevocab + age +( gevocab + age| s c h o o l ) ,
data=Achieve )
summary ( Model3 . 6 )
#3.3 I f u n c o r r e l a t e d random e f f e c t s are d e s i r e d
#without the separate i n t e r c e p t estimated f o r each ,
#the f o l l o w i n g code can be used t o e l i m i n a t e the extra i n t e r c e p t s . MModel3 . 6 a
Model3 . 6 a <− lmer ( geread ~ gevocab + age + (1| s c h o o l ) + ( −1 +
gevocab| s c h o o l ) + ( −1 + age| s c h o o l ) , data=Achieve )
summary ( Model3 . 6 a )
# 4 . CENTERING PREDICTORS
#4.1 For example , returning t o Model3 . 3 ,
#grand mean centered gevocab
#and age v a r i a b l e s can be cr eated with the f o l l o w i n g syntax :
Cgevocab <− Achieve$gevocab−mean ( Achieve$gevocab )
Cage <− Achieve$age−mean ( Achieve$age )
#4.2 Once mean centered v e r s i o n s o f the p r e d i c t o r s have been created ,
#they can be i n c o r p o r a t e d i n t o the model in the same manner as b e f o r e .
Model3 . 3C <−lmer ( geread~Cgevocab + Cage + Cgevocab * Cage +
(1| s c h o o l ) , data=Achieve )
summary ( Model3 . 3C)
# 5 . ADDITIONAL OPTIONS
#5.1 Parameter estimation method
465
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
Model3 . 8 <− lmer ( geread~gevocab + (1| s c h o o l ) , data=Achieve ,
REML=FALSE)
summary ( Model3 . 8 )
#5.2 Estimation Controls , we can add new arguments
Model3 . 8 <− lmer ( geread ~ gevocab + (1| s c h o o l ) , data=Achieve ,
REML = FALSE, c o n t r o l = l i s t ( maxIter =100 , opt =" optim " ) )
#5.3 Comparing Model F i t
# 5 . 3 . 1 Use the anova ( ) f u n c t i o n on the f i r s t two models we did
anova ( Model3 . 0 , Model3 . 1 )
# 5 . 4 . 1 lme4 and Hypothesis Testing .
#These c o n f i d e n c e i n t e r v a l s can be obtained using the c o n f i n t function , as below .
Model3 . 5 <− lmer ( geread ~ gevocab + ( gevocab| s c h o o l ) ,
data=Achieve )
summary ( Model3 . 5 )
c o n f i n t ( Model3 . 5 , method=c ( " boot " ) , boot . type=c ( " perc " ) )
c o n f i n t ( Model3 . 5 , method=c ( " boot " ) , boot . type=c ( " b a s i c " ) )
c o n f i n t ( Model3 . 5 , method=c ( " boot " ) , boot . type=c ( " norm " ) )
c o n f i n t ( Model3 . 5 , method=c ( " Wald " ) )
c o n f i n t ( Model3 . 5 , method=c ( " p r o f i l e " ) )
466
Electronic copy available at: https://ssrn.com/abstract=3863563
8.32. THREE-LEVEL MULTILEVEL LINEAR MODELS
8.32
Three-level multilevel linear models
#THREE−LEVEL MLM_EXERCISE
# 0 . I n s t a l l and load the packages and data
l i b r a r y ( readr )
l i b r a r y ( Matrix )
l i b r a r y ( lme4 )
Achieve <− read . csv ( f i l e . choose ( ) )
# 1 . DEFINING SIMPLE THREE−LEVEL MODELS
#1.1 F i t a three −l e v e l model
Model4 . 0 <− lmer ( geread ~1+(1| s c h o o l / c l a s s ) , data=Achieve )
summary ( Model4 . 0 )
#NOTES:
# 1 . The c o n f i n t ( ) f u n c t i o n may be used t o obtain i n f e r e n c e information
#about f i x e d and random e f f e c t s .
# 2 . This may take a few seconds t o compute , depending on your computing device ,
# so be p a t i e n t and do not i n t e r r u p t .
c o n f i n t ( Model4 . 0 , method=c ( " p r o f i l e " ) )
#1.2 Use the f o l l o w i n g lmer command
Model4 . 1 <− lmer ( geread ~ gevocab + c l e n r o l l + c e n r o l l +(1| s c h o o l /
c l a s s ) , data=Achieve )
summary ( Model4 . 1 )
#1.3 Confidence i n t e r v a l s f o r the model e f f e c t s
#can be obtained with the c o n f i n t ( ) f u n c t i o n .
c o n f i n t ( Model4 . 1 , method = c ( " p r o f i l e " ) )
#1.4 Use the anova ( ) f u n c t i o n t o compare both models 4 . 0 and 4 . 1
anova ( Model4 . 0 , Model4 . 1 )
467
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#1.5 Test the hypothesis s t a t i n g that
#the impact o f vocabulary s c o r e on reading achievement v a r i e s depending upon
#the s i z e o f the s c h o o l that a student attends using Model4 . 2 below :
Model4 . 2 <− lmer ( geread ~ gevocab +
clenroll +
cenroll +
gevocab * c e n r o l l +
(1| s c h o o l / c l a s s ) ,
data=Achieve )
summary ( Model4 . 2 )
#1.6 Once again , use the anova ( ) f u n c t i o n t o compare models ,
# t h i s time 4 . 1 and 4 . 2
anova ( Model4 . 1 , Model4 . 2 )
# 2 . DEFINING SIMPLE MODELS WITH MORE THAN THREE LEVELS
#2.1 Run Model4 . 3 , t h i s time adding s c h o o l d i s t r i c t as another l e v e l
Model4 . 3 <− lmer ( geread ~ 1 + (1| corp / s c h o o l / c l a s s ) , data=Achieve )
summary ( Model4 . 3 )
#2.2 Run the c o n f i n t ( ) f u n c t i o n
c o n f i n t ( Model4 . 3 , method = c ( " p r o f i l e " ) )
# 3 . RANDOM COEFFICIENTS MODELS WITH THREE OR MORE LEVELS
# Now, the r e l a t i o n s h i p o f gender t o reading may d i f f e r a c r o s s s c h o o l s
#and a c r o s s classrooms , thus leading t o a model
#where the gender c o e f f i c i e n t i s allowed t o vary a c r o s s
#both random e f f e c t s in our three −l e v e l model .
#3.1 Below i s the lmer command sequence f o r f i t t i n g t h i s model
Model4 . 4 <− lmer ( geread ~ gevocab +
gender +
468
Electronic copy available at: https://ssrn.com/abstract=3863563
8.32. THREE-LEVEL MULTILEVEL LINEAR MODELS
( gender| s c h o o l / c l a s s ) ,
data=Achieve )
summary ( Model4 . 4 )
#3.2 Define the random c o e f f i c i e n t with ( gender| s c h o o l / c l a s s ) ,
#thus allowing both the i n t e r c e p t and s l o p e t o vary a c r o s s both classrooms
#and s c h o o l s
Model4 . 5 <− lmer ( geread ~ gevocab + gender + (1| s c h o o l ) +
( gender| c l a s s ) , data=Achieve )
summary ( Model4 . 5 )
#NOTES: Results are i d e n t i c a l t o the t e x t book , so ignore any warning messages
#3.3 Next model allows gender t o vary only a c r o s s s c h o o l s
#and the i n t e r c e p t t o vary only a c r o s s classrooms .
Model4 . 6 <− lmer ( geread~gevocab+gender + (−1+gender| s c h o o l ) +
(1| c l a s s ) , data=Achieve )
summary ( Model4 . 6 )
#3.4 The l a s t model d e p i c t s f o u r nested l e v e l s
#where the i n t e r c e p t i s allowed t o vary
# a c r o s s c o r p o r a t i o n , school , and classroom ,
#but the gender s l o p e i s only allowed t o vary a c r o s s classroom
Model4 . 7 <− lmer ( geread~gevocab+gender + (1| corp ) + (1| s c h o o l )
+ ( gender| c l a s s ) , data=Achieve )
summary ( Model4 . 7 )
469
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.33
Longitudinal data analysis using multilevel models
#LONGITUDINAL DATA ANALYSIS USING MULTILEVEL MODELS
#0 Load and l i b r a r y the data and packages
l i b r a r y ( readr )
l i b r a r y ( lme4 )
l i b r a r y ( dplyr )
l i b r a r y ( Matrix )
l i b r a r y ( lme4 )
Lang <− read . csv ( f i l e . choose ( ) )
#0.1 Now, take a look at t h i s sample data in the usual way
dim ( Lang )
s t r ( Lang )
summary ( Lang )
head ( Lang , 12)
t a i l ( Lang , 12)
#0.2 The f o l l o w i n g R commands w i l l rearrange the data i n t o the necessary format
LangPP <− data . frame ( ID=Lang$ID , s c h o o l =Lang$school ,
Process=Lang$Process ,
A p p l i c a t i o n =Lang$Application , Grammar=Lang$Grammar ,
Mechanics=Lang$Mechanics ,
stack ( Lang , s e l e c t =LangScore1 : LangScore6 ) )
#0.3 Rename the values v a r i a b l e t o Language .
#The values v a r i a b l e i s the seventh column
names ( LangPP ) [ c ( 7 ) ] <− c ( " Language " )
# 1 . FITTING LONGITUDINAL MODELS USING lme4
Model5 . 0 <− lmer ( Language~Time + (1|ID ) , data=LangPP )
summary ( Model5 . 0 )
#1.1 Simply add the 95 percent CI using c o n f i n t ( )
c o n f i n t ( Model5 . 0 , method=c ( " p r o f i l e " ) )
#1.2 Next , can add v a r i a b l e Grammar and r e v i s e our model
470
Electronic copy available at: https://ssrn.com/abstract=3863563
8.34. GRAPHING DATA IN MULTILEVEL CONTEXTS
Model5 . 1 <− lmer ( Language~Time + Grammar +(1|ID ) , LangPP )
summary ( Model5 . 1 )
#1.3 Allow growth
Model5 . 2 <− lmer ( Language~Time + Grammar+(Time|ID ) , LangPP )
summary ( Model5 . 2 )
#1.4 Running 95 percent CI as per usual
c o n f i n t ( Model5 . 2 , method=c ( " p r o f i l e " ) )
#1.5 Add a t h i r d l e v e l o f data s t r u c t u r e t o the model
#by i n c l u d i n g information regarding s c h o o l s ,
# within which examinees are nested
Model5 . 3 <− lmer ( Language~Time + (1| s c h o o l / ID ) , data=LangPP )
summary ( Model5 . 3 )
#1.6 Using the anova ( ) command ,
#we can compare the f i t o f the three −l e v e l and two−l e v e l v e r s i o n s o f t h i s model .
anova ( Model5 . 0 , Model5 . 3 )
#NOTES: Can conclude that i n c l u s i o n o f the s c h o o l l e v e l o f the data leads t o
# b e t t e r model f i t .
8.34
Graphing data in multilevel contexts
#GRAPHING DATA IN MULTILEVEL CONTEXTS_EXERCISE
#NOTES: The p l o t t i n g c a p a b i l i t i e s in R are t r u l y outstanding .
# I t i s capable o f producing high−q u a l i t y graphics
#with a great deal o f f l e x i b i l i t y .
#As a simple example , c o n s i d e r the Anscombe data ( included in R)
#0.1 I n s t a l l and load the data and package
data ( anscombe )
l i b r a r y ( Matrix )
l i b r a r y ( carData )
library ( effects )
471
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#0.2 Examination o f the data by c a l l i n g upon the dataset leads t o
anscombe
#0.3 The way in which one can p l o t the data f o r the f i r s t dataset
# ( i . e . x1 and y1 above ) , i s as f o l l o w s :
p l o t ( anscombe$y1 ~ anscombe$x1 )
#NOTES: The ~ symbol in the f u n c t i o n c a l l p o s i t i o n s the data t o
#the l e f t as the dependent v a r i a b l e p l o t t e d on the o r d i n a t e ( y−a x i s ) ,
#whereas the value t o the r i g h t i s t r e a t e d as an independent v a r i a b l e
#and p l o t t e d on the a b s c i s s a ( x−a x i s ) .
#0.4 A l t e r n a t i v e l y , one can rearrange the terms
# so that the independent v a r i a b l e comes f i r s t ,
#with a comma separating i t from the dependent v a r i a b l e :
p l o t ( anscombe$x1 , anscombe$y1 )
#NOTES:
# 1 . Both the former and the l a t t e r approaches lead t o the same p l o t .
# 2 . The p l o t f u n c t i o n has s i x base parameters ,
#with the o p t i o n o f c a l l i n g other g r a p h i c a l parameters including ,
#among others , par f u n c t i o n .
# 3 . The par f u n c t i o n has more than 70 g r a p h i c a l parameters
# that can be used t o modify a b a s i c p l o t .
#1.1 F i t the r e g r e s s i o n model .
data . 1 <− lm ( y1~x1 , data=anscombe )
#1.2 S c a t t e r p l o t o f the data .
p l o t ( anscombe$y1 ~ anscombe$x1 ,
ylab=e xpr ess ion ( i t a l i c (Y ) ) ,
ylim=c ( 2 , 1 2 ) ,
xlab=e xpr ess ion ( i t a l i c (X ) ) ,
main="Anscombe ’ s Data Set 1 " )
472
Electronic copy available at: https://ssrn.com/abstract=3863563
8.34. GRAPHING DATA IN MULTILEVEL CONTEXTS
#1.3 Add the f i t t e d r e g r e s s i o n l i n e .
a b l i n e ( data . 1 )
#1.4 Add the t e x t and e x p r e s s i o n s within the f i g u r e .
text (5.9 , 9.35 ,
express ion ( paste ( " The value o f " , i t a l i c (R) ^ 2 , " i s . 6 7 " ,
sep = " " ) ) )
text ( 5 . 9 , 10.15 ,
express ion ( i t a l i c ( hat (Y))==3+ i t a l i c (X ) * . 5 ) )
#1.5 Break the a x i s by adding a zigzag .
# 1 . 5 . 1 i n s t a l l . packages ( " p l o t r i x " )
require ( p l o t r i x )
a x i s . break ( a x i s =1 , s t y l e =" zigzag " )
a x i s . break ( a x i s =2 , s t y l e =" zigzag " )
# 2 . PLOTS FOR LINEAR MODELS
#2.1 Load and l i b r a r y the data and packages
l i b r a r y ( readr )
Cassidy <− read . csv ( f i l e . choose ( ) )
dim ( Cassidy )
s t r ( Cassidy )
summary ( Cassidy )
View ( Cassidy )
#NOTES:
# 1 . The Cassidy GPA, in which GPA was modeled by CTA. t o t and BStotal .
# 2 . Now d i s c u s s some p l o t s that are u s e f u l with s i n g l e −l e v e l data ,
#and which can be e a s i l y extended t o the m u l t i l e v e l case , with some caveats .
# 3 . F i r s t , Consider the p a i r s function , which p l o t s a l l p a i r s o f v a r i a b l e s in a dataset .
# 4 . The r e s u l t i n g graph i s sometimes r e f e r r e d t o as a s c a t t e r p l o t matrix ,
#because i t i s , in f a c t , a matrix o f s c a t t e r p l o t s .
473
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#2.2 As an example , c o n s i d e r the f o l l o w i n g c a l l t o p a i r s :
pairs (
cbind (GPA=Cassidy$GPA , CTA_Total=Cassidy$CTA . t o t ,
BS_Total=Cassidy$BStotal ) )
#NOTES:
# 1 . Such a p a i r s p l o t allows multipl e b i v a r i a t e r e l a t i o n s h i p s
# t o be v i s u a l i z e d simultaneously .
# 2 . Of course , one can quantif y the degree o f l i n e a r r e l a t i o n with a c o r r e l a t i o n .
#2.3 Code t o do t h i s can be given as f o l l o w s , using l i s t w i s e d e l e t i o n :
c o r ( na . omit (
cbind (GPA=Cassidy$GPA , CTA_Total=Cassidy$CTA . t o t ,
BS_Total=Cassidy$BStotal ) ) )
#NOTES:
# 1 . When using multipl e r e g r e s s i o n ,
#must make some assumptions regarding the d i s t r i b u t i o n o f the model r e s i d u a l s
# 2 . In p a r t i c u l a r , in order f o r the p−values and c o n f i d e n c e i n t e r v a l s t o be exact ,
#must assume that the d i s t r i b u t i o n o f r e s i d u a l s i s normal .
#2.4 Can obtain the r e s i d u a l s from a model using the r e s i d function ,
#which i s applied t o a f i t t e d lm o b j e c t
GPAmodel . 1 <− lm (GPA ~ CTA. t o t +
BStotal , data=Cassidy )
r e s i d . 1 <− r e s i d ( GPAmodel . 1 )
#2.5 One u s e f u l such p l o t i s a histogram with an o v e r l a i d normal d e n s i t y curve ,
#which can be obtained using the f o l l o w i n g R command :
hist ( resid .1 ,
f r e q =FALSE, main=" Density o f Residuals f o r Model 1 " ,
xlab =" Residuals " )
l i n e s ( density ( resid . 1 ) )
#NOTES:
474
Electronic copy available at: https://ssrn.com/abstract=3863563
8.34. GRAPHING DATA IN MULTILEVEL CONTEXTS
# 1 . This code f i r s t requests that a histogram be produced f o r the r e s i d u a l s .
# 2 . Note that f r e q =FALSE i s used ,
#which i n s t r u c t s R t o make the y−a x i s s c a l e d in terms o f p r o b a b i l i t y ,
#not the d e f a u l t , which i s frequency .
#The s o l i d l i n e r e p r e s e n t s the d e n s i t y estimate o f the r e s i d u a l s ,
# corresponding c l o s e l y t o the bars in the histogram
# 3 . Rather than a histogram and o v e r l a i d density ,
#a more s t r a i g h t f o r w a r d way o f evaluating the d i s t r i b u t i o n o f
#the e r r o r s t o the normal d i s t r i b u t i o n i s a quantile −q u a n t i l e p l o t (Q−Q p l o t ) .
#2.6 The f o l l o w i n g code w i l l produce the Q−Q p l o t ,
#based on the r e s i d u a l s from the GPA model
qqnorm ( s c a l e ( r e s i d . 1 ) , main="Normal Quantile−Quantile P l o t " )
qqline ( scale ( resid . 1 ) )
#NOTES:
# 1 . Notice that above we use the s c a l e function ,
#which standardizes the r e s i d u a l s t o have a mean o f 0
# ( already done due t o the r e g r e s s i o n model ) and a standard d e v i a t i o n o f 1 .
# 2 . Can see that p o i n t s in the Q−Q p l o t div erge from the l i n e
# f o r the higher end o f the d i s t r i b u t i o n .
# 3 . This i s c o n s i s t e n t with the histogram
#which shows a s h o r t e r t a i l on the high end
#as compared t o the low end o f the d i s t r i b u t i o n .
# 4 . This p l o t , along with the histogram ,
# r e v e a l s that the model r e s i d u a l s do di verge from a p e r f e c t l y normal d i s t r i b u t i o n
# 3 . PLOTTING NESTED DATA
#NOTES:
475
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 1 . Next , move on t o g r a p h i c a l t o o l s more s p e c i f i c a l l y u s e f u l with m u l t i l e v e l data .
# 2 . M u l t i l e v e l models are o f t e n applied t o r e l a t i v e l y large ,
#indeed sometimes huge , d a t a s e t s .
# 3 . Such d a t a s e t s provide a r i c h n e s s that
#cannot be r e a l i z e d in s t u d i e s with small samples .
# 4 . However , a c o m p l i c a t i o n that o f t e n a r i s e s from t h i s vastness o f
# m u l t i l e v e l data i s the d i f f i c u l t y o f c r e a t i n g p l o t s
# that can summarize l a r g e amounts o f information
#and thereby provide i n s i g h t s i n t o the nature o f r e l a t i o n s h i p s among the v a r i a b l e s
# 5 . For example , the Prime Time s c h o o l dataset c o n t a i n s
#more than 10 ,000 third −grade students .
# 6 . A s i n g l e p l o t with a l l 10 ,000 i n d i v i d u a l s
#can be overwhelming , and f a i r l y uninformative .
# 7 . Including a nesting s t r u c t u r e ( e . g . s c h o o l c o r p o r a t i o n s )
#can lead t o many " c o r p o r a t i o n s p e c i f i c " p l o t s ( the data contain 60 c o r p o r a t i o n s ) .
# 8 . Thus , when p l o t t i n g m u l t i l e v e l data ,
#the nested s t r u c t u r e should be e x p l i c i t l y considered .
# 9 . I f not , at minimum a r e a l i z a t i o n i s necessary
# that the r e l a t i o n s h i p s in the unstructured data may d i f f e r
#when the nesting s t r u c t u r e i s considered .
#10. Although dealing with the c o m p l e x i t i e s
# in p l o t t i n g nested data can be at times vexing ,
#the r i c h n e s s that nested data provide f a r outweighs
#any o f the c o m p l i c a t i o n s
# that a r i s e when a p p r o p r i a t e l y gaining an understanding from i t .
# 4 . USING THE LATTICE PACKAGE
476
Electronic copy available at: https://ssrn.com/abstract=3863563
8.34. GRAPHING DATA IN MULTILEVEL CONTEXTS
#4.1 i n s t a l l . packages ( " l a t t i c e " ) and load the data
library ( lattice )
Achieve <− read_csv ( f i l e . choose ( ) )
View ( Achieve )
#4.2 Next , need t o c r e a t e a new ( sub ) dataset c o n t a i n i n g data f o r
# j u s t t h i s i n d i v i d u a l school , using the f o l l o w i n g code :
Achieve .9 4 0 . 7 6 7 <− Achieve [ Achieve$corp ==940 &
Achieve$school ==767 , ]
#NOTES: Can see from the Global Environment window in RStudio
# that t h i s subset has 59 obs
#4.3 Then , can c r e a t e the graphic using the d o t p l o t ( ) f u n c t i o n
d o t p l o t ( c l a s s ~ geread ,
data=Achieve . 9 4 0 . 7 6 7 , j i t t e r . y = TRUE, ylab =" Classroom " ,
main=" Dotplot o f \ ’ geread \ ’ f o r Classrooms in School 767 ,
Which i s Within Corporation 9 4 0 " )
#4.4 Thus , modify Figure in order t o have the c l a s s e s
# placed in descending order by the mean o f geread .
dotplot (
reorder ( c l a s s , geread ) ~ geread ,
data=Achieve . 9 4 0 . 7 6 7 , j i t t e r . y = TRUE, ylab =" Classroom " ,
main=" Dotplot o f \ ’ geread \ ’ f o r Classrooms in School 767 ,
Which i s Within Corporation 9 4 0 " )
#4.5 Use the f o l l o w i n g code and produce Figure 8 :
d o t p l o t ( reord er ( corp , geread ) ~ geread , data=Achieve , j i t t e r . y
= TRUE, ylab =" Classroom " ,
main=" Dotplot o f \ ’ geread \ ’ f o r A l l Corporations " )
Achieve <− cbind ( Achieve , Classroom_Unique=paste ( Achieve$corp ,
Achieve$school ,
Achieve$class , sep = " " ) )
#4.6 Now, use the aggregate f u n c t i o n
Achieve . Class_Aggregated <− aggregate ( Achieve ,
477
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
by= l i s t ( Achieve$Classroom_Unique ) ,
FUN=mean )
#4.7 Now, c r e a t e the d o t p l o t with the f o l l o w i n g command
d o t p l o t ( r eord er ( corp , geread ) ~ geread , data=Achieve . Class_Aggregated ,
j i t t e r . y = TRUE,
ylab =" Corporation " ,
main=" Dotplot o f Classroom Mean \ ’ geread \ ’ Within the Corporations " )
#4.8 For example , the f o l l o w i n g code produces an x y p l o t
# f o r geread ( y−a x i s ) by gevocab ( x−a x i s ) , accounting f o r s c h o o l c o r p o r a t i o n
x y p l o t ( geread ~ gevocab | corp , data = Achieve )
#NOTES:
# 1 . Notice that here the v e r t i c a l bar symbol d e f i n e s the grouping / nesting s t r u c t u r e
#with the ~ symbol implying that geread i s p r e d i c t e d / modeled by gevocab .
# 2 . By d e f a u l t , the s p e c i f i c names o f the grouping s t r u c t u r e
# ( here the c o r p o r a t i o n numbers ) are not p l o t t e d on the " s t r i p . "
s t r i p = s t r i p . custom ( s t r i p . names=FALSE, s t r i p . l e v e l s =c (FALSE,TRUE) )
# 3 . For example , we might want t o p l o t the s c h o o l s in , say , c o r p o r a t i o n 940 ,
#which can be done by e x t r a c t i n g from the Achieve data only c o r p o r a t i o n 940 ,
#as i s done t o produce Figure 11
x y p l o t ( geread ~ gevocab | school , data = Achieve [ Achieve$corp ==940 ,] ,
s t r i p = s t r i p . custom ( s t r i p . names=FALSE, s t r i p . l e v e l s =c (FALSE, TRUE) ) ,
main=" Schools in Corporation 9 4 0 " )
#4.9 Code t o produce histogram
# 4 . 9 . 1 Need t o re−c r e a t e Model3 . 1 from t h i s morning ’ s l e c t u r e
l i b r a r y ( lme4 )
478
Electronic copy available at: https://ssrn.com/abstract=3863563
8.34. GRAPHING DATA IN MULTILEVEL CONTEXTS
Model3 . 1 <− lmer ( geread ~ gevocab + (1| s c h o o l ) , data=Achieve )
# 4 . 9 . 2 Now we can c r e a t e the histogram
h i s t ( s c a l e ( r e s i d ( Model3 . 1 ) ) ,
f r e q =FALSE, ylim=c ( 0 , . 7 ) , xlim=c ( − 4 , 5 ) ,
main=" Histogram o f Standardized Residuals from Model 3 . 1 " ,
xlab =" Standardized Residuals " )
l i n e s ( d e n s i t y ( s c a l e ( r e s i d ( Model3 . 1 ) ) ) )
box ( )
# 5 . PLOTTING MODEL RESULTS USING THE EFFECTS PACKAGE
i n s t a l l . packages ( " e f f e c t s " )
library ( effects )
#5.1 Import the Prime Time data f i l e i n t o a new data frame prime_time
prime_time <− read_csv ( f i l e . choose ( ) )
#5.2 F i t a model in which the dependent v a r i a b l e i s the reading score , geread ,
#and the p r e d i c t o r s are measures o f verbal ( npaverb )
#and nonverbal ( npanverb ) reasoning s k i l l s
Model6 . 1 <− lmer ( geread~npaverb + (1| s c h o o l ) ,
data=prime_time )
summary ( Model6 . 1 )
#NOTES:
# 1 . Note that the number o f o b s e r v a t i o n s i s d i f f e r e n t
# 2 . Achieve . csv data f i l e has 10 ,320 obs
# 3 . PrimeTime . csv data f i l e above shows number o f o b s e r v a t i o n s = 10 ,927
# 4 . Output t a b l e in the textbook shows Number o f obs : 10 ,765
# 5 . Thus , do not worry about d i s c r e p a n c i e s as the sample data f i l e s
#have been changed ( updated o n l i n e )
#5.3 Use the e f f e c t s package t o v i s u a l i z e the r e l a t i o n s h i p
479
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 1 ) )
#NOTES:
# 1 . The r e s u l t i n g l i n e p l o t r e f l e c t s the r e l a t i o n s h i p d e s c r i b e d numerically
#by the npaverb s l o p e o f 0 . 0 4 3 9 .
# 2 . In order t o p l o t the r e l a t i o n s h i p , simply employ the p r e d i c t o r E f f e c t s
#command in R .
# 3 . Multiple dependent v a r i a b l e s can be p l o t t e d at once
#using t h i s b a s i c command s t r u c t u r e .
#5.4 In Model6 . 2 , i n c l u d e s c h o o l socioeconomic s t a t u s (SES) in our model .
Model6 . 2 <− lmer ( geread~npaverb + ses +(1| s c h o o l ) ,
data=prime_time )
summary ( Model6 . 2 )
#NOTES:
# 1 . From these r e s u l t s , higher l e v e l s o f s c h o o l SES are a s s o c i a t e d
#with higher reading t e s t scores , as are higher s c o r e s on
#the verbal reasoning t e s t .
# 2 . p l o t these r e l a t i o n s h i p s and
#the a s s o c i a t e d c o n f i d e n c e i n t e r v a l s simultaneously as below .
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 2 ) )
#NOTES:
# 1 . I t i s a l s o p o s s i b l e t o p l o t only one o f these r e l a t i o n s h i p s per graph .
#5.5 Added a more d e s c r i p t i v e t i t l e f o r each o f the f o l l o w i n g p l o t s ,
#using the main subcommand
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 2 , ~ npaverb ) , main = " Reading by Verbal Reasoning " )
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 2 , ~ ses ) , main = " Reading by School SES " )
#NOTES:
480
Electronic copy available at: https://ssrn.com/abstract=3863563
8.34. GRAPHING DATA IN MULTILEVEL CONTEXTS
# 1 . One very u s e f u l aspect o f these p l o t s i s the i n c l u s i o n o f the 95% c o n f i d e n c e
# region around the l i n e .
# 2 . From t h i s , can see that our c o n f i d e n c e regarding
#the a c t u a l nature o f the r e l a t i o n s h i p s in the population i s much g r e a t e r f o r
# verbal reasoning than f o r s c h o o l SES .
# 3 . Of course , given the much l a r g e r l e v e l −1 sample s i z e ,
# t h i s makes p e r f e c t sense .
#5.6 I t i s a l s o p o s s i b l e t o i n c l u d e c a t e g o r i c a l independent v a r i a b l e s along
#with t h e i r p l o t s , as with student gender in Model6 . 3
Model6 . 3 <− lmer ( geread~npaverb + gender +(1| s c h o o l ) ,
data=prime_time )
summary ( Model6 . 3 )
#5.7 Gender i s not s t a t i s t i c a l l y r e l a t e d t o reading score ,
#as r e f l e c t e d in the p−value , and in the graph below .
#Note the very wide c o n f i d e n c e region around the l i n e .
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 3 , ~ gender ) )
#5.8 Finally , use the e f f e c t s package t o g r a p h i c a l l y
#probe i n t e r a c t i o n s among independent v a r i a b l e s .
## For example , c o n s i d e r Model6 . 4 ,
# in which the v a r i a b l e s npaverb and npanverb ( nonverbal reasoning ) are used as
# p r e d i c t o r s o f reading s c o r e .
#In addition , the i n t e r a c t i o n o f these two v a r i a b l e s i s a l s o included in the model .
Model6 . 4 <− lmer ( geread ~ npaverb + npanverb + npaverb * npanverb
+(1| s c h o o l ) , data=prime_time )
summary ( Model6 . 4 )
#5.9 These r e s u l t s i n d i c a t e that nonverbal reasoning i s not s t a t i s t i c a l l y
# r e l a t e d t o reading t e s t score , but that there i s an i n t e r a c t i o n between verbal
#and nonverbal reasoning .
#In order t o gain a b e t t e r understanding as t o the nature o f t h i s i n t e r a c t i o n ,
#we can p l o t i t using the e f f e c t s package .
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 4 , ~ npaverb * npanverb ) )
#NOTES:
# 1 . I t i s a l s o p o s s i b l e t o p l o t three −way i n t e r a c t i o n s using the e f f e c t s package .
481
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 2 . Model6 . 5 i n c l u d e s a measure o f memory ,
# in a d d i t i o n t o the verbal and nonverbal reasoning s c o r e s .
#5.10 The model summary information demonstrates
# that there i s a s t a t i s t i c a l l y s i g n i f i c a n t three −way i n t e r a c t i o n among
#the independent v a r i a b l e
Model6 . 5 <− lmer ( geread ~ npaverb * npanverb * npamem
+(1| s c h o o l ) , data=prime_time )
summary ( Model6 . 5 )
#5.11 The i n t e r a c t i o n can be p l o t t e d using e f f e c t s .
p l o t ( p r e d i c t o r E f f e c t s ( Model6 . 5 , ~ npaverb * npanverb *npamem ) )
482
Electronic copy available at: https://ssrn.com/abstract=3863563
8.35. MULTILEVEL GENERALIZED LINEAR MODELS (MGLM)
8.35
Multilevel generalized linear models (MGLM)
#MULTILEVEL GENERALIZED LINEAR MODELS (MGLM)_EXERCISE
# 1 . MGLMs FOR A DICHOTOMOUS OUTCOME VARIABLE_Random I n t e r c e p t L o g i s t i c Regression
#0.1 Load data and l i b r a r y
i n s t a l l . packages ( " s c a l e s " )
library ( scales )
l i b r a r y ( readr )
l i b r a r y ( lme4 )
attach ( mathfinal )
library ( scales )
mathfinal <− read . csv ( f i l e . choose ( ) )
View ( mathfinal )
#NOTES: I t i s q u i t e large , plus there are many columns with Null or missing data
#1.1 Run our f i r s t model c a l l e d Model8 . 1 with a f i x e d e f f e c t f o r i n t e r c e p t
#and the s l o p e o f the independent v a r i a b l e d c a l l e d numsense
#BUT DOES NOT EXIST IN THE SAMPLE DATA FILE .
#Attempt t o re−c r e a t e v a r i a b l e numsense
numsense <− r e s c a l e ( mathfinal$TestRITScoref10 , t o = c ( 0 , 1 0 0 ) )
h i s t ( numsense )
summary ( model8 . 1 <− glmer ( score2 ~ numsense + (1| s c h o o l ) ,
family=binomial ,
na . a c t i o n =na . omit ) )
#NOTES:
# 1 . Output r e s u l t s are d i f f e r e n t t o that in the text ,
# so we w i l l f o l l o w textbook output f o r i n t e r p r e t a t i o n
# 2 . I t i s a sample data a v a i l a b i l i t y issue ,
#but the above code works with the data we c u r r e n t l y have .
#1.2 P e r c e n t i l e Bootstrap
c o n f i n t ( model8 . 1 , method=c ( " boot " ) , boot . type=c ( " perc " ) )
#1.3 Random C o e f f i c i e n t s L o g i s t i c Regression
483
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#As with the l i n e a r m u l t i l e v e l models ,
# i t i s a l s o p o s s i b l e t o allow f o r random s l o p e s with m u l t i l e v e l GLMs,
#and t h i s can be done using glmer , as shown below
summary ( model8 . 2 <− glmer ( score2 ~ numsense + ( numsense| s c h o o l ) ,
family = binomial ) )
#NOTES:
# 1 . Raw data i s d i f f e r e n t , o b v i o u s l y updated data in second e d i t i o n .
# 2 . Rely upon the textbook output f o r a n a l y s i s
# ( use current o l d e r raw data purely
# f o r i l l u s t r a t i n g how t h i s s t a t i s t i c a l procedure works )
# 3 . Use the Wald p r o f i l e when computing c o n f i d e n c e i n t e r v a l s
# f o r the random and f i x e d e f f e c t s o f the random c o e f f i c i e n t s model
#1.4 Wald
c o n f i n t ( model8 . 2 , method=c ( " Wald " ) )
#1.5 P r o f i l e
c o n f i n t ( model8 . 2 , method=c ( " p r o f i l e " ) )
#1.6 P e r c e n t i l e Bootstrap
c o n f i n t ( model8 . 2 , method=c ( " boot " ) , boot . type=c ( " perc " ) )
# 2 . MGLM FOR COUNT DATA
#2.1 Random I n t e r c e p t Poisson Regression
# 2 . 1 . 0 Load and l i b r a r y the package and data
l i b r a r y ( readr )
rehab_data <− read_csv ( f i l e . choose ( ) )
View ( rehab_data )
# 2 . 1 . 1 The R commands and r e s u l t a n t output f o r
# f i t t i n g the Poisson r e g r e s s i o n model t o the data appear below
#using the glmer f u n c t i o n in the lme4 l i b r a r y
484
Electronic copy available at: https://ssrn.com/abstract=3863563
8.35. MULTILEVEL GENERALIZED LINEAR MODELS (MGLM)
# that was employed e a r l i e r t o f i t the dichotomous l o g i s t i c r e g r e s s i o n models
l i b r a r y ( lme4 )
summary ( model8.7< − glmer ( heart~ t r t +sex +(1| rehab ) ,
family=poisson ,
data=rehab_data ) )
#NOTES: Output data i s d i f f e r e n t t o that used in the t e x t
#2.2 Random C o e f f i c i e n t Poisson Regression
# 2 . 2 . 1 Model8 . 8
summary ( model8.8< − glmer ( heart~ t r t +sex +( t r t |rehab ) ,
family=poisson ,
data=rehab_data ) )
# 2 . 2 . 2 Compare both models 8 . 7 and 8 . 8 now using the anova ( ) f u n c t i o n
anova ( model8 . 7 , model8 . 8 )
# 2 . 2 . 3 I n c l u s i o n o f A d d i t i o n a l Level −2 E f f e c t s
# t o the Multinomial Poisson Regression Model
summary ( model8.9< − glmer ( heart~ t r t +sex+hours +(1| rehab ) ,
family=poisson ,
data=rehab_data ) )
# 2 . 2 . 4 I n s t a l l and load the nlme package
i n s t a l l . packages ( " nlme " )
l i b r a r y ( nlme )
l i b r a r y (MASS)
# 2 . 2 . 5 Now can estimate model8 . 1 0
summary ( rehab_data$heart )
#summary ( model8.10< −glmmPQL( heart~ t r t +sex , random=~1|rehab , family=quasipoisson ) )
# Model8 . 1 1 , but f i r s t we need t o load package lme4
l i b r a r y ( lme4 )
summary ( model8.11< − glmer . nb ( heart~ t r t +sex +(1| rehab ) , data=rehab_data ) )
485
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 2 . 2 . 6 The 95% p r o f i l e c o n f i d e n c e i n t e r v a l s
# f o r the f i x e d and random e f f e c t s in model8 . 1 1 appear below
c o n f i n t ( model8 . 1 1 , method=c ( " p r o f i l e " ) )
# 2 . 2 . 7 Now, f i t a random c o e f f i c i e n t f o r the t r t v a r i a b l e ,
#as we did f o r the Poisson r e g r e s s i o n model .
summary ( model8.12< − glmer . nb ( heart~ t r t +sex +( t r t |rehab ) , data=rehab_data ) )
c o n f i n t ( model8 . 1 2 , method=c ( " boot " ) , boot . type=c ( " perc " ) )
# 2 . 2 . 8 As with other models that have examined in t h i s book ,
# i t i s p o s s i b l e t o i n c l u d e l e v e l −2 independent v a r i a b l e s ,
#such as number o f hours the c e n t e r s are open ,
# in the model , and compare the r e l a t i v e f i t using the r e l a t i v e f i t i n d i c e s ,
#as in Model8 . 1 3 .
summary ( model8.13< − glmer . nb ( heart~ t r t +sex+hours +(1| rehab ) , data=rehab_data ) )
c o n f i n t ( model8 . 1 3 , method=c ( " p r o f i l e " ) )
# 2 . 2 . 9 In addition , compare the AIC and BIC values t o provide f u r t h e r
# evidence regarding the r e l a t i v e model f i t .
anova ( model8 . 1 1 , model8 . 1 2 )
anova ( model8 . 1 1 , model8 . 1 3 )
486
Electronic copy available at: https://ssrn.com/abstract=3863563
8.36. GENERALIZED LINEAR MODELS RECAP
8.36
Generalized linear models recap
#GENERALIZED LINEAR MODELS RECAP_EXERCISE
#0.1 I n s t a l l and load the package and data
l i b r a r y (MASS)
l i b r a r y ( readr )
l i b r a r y ( dplyr )
coronary <− read . csv ( f i l e . choose ( ) )
summary ( coronary )
View ( coronary )
# 1 . Recode the group v a r i a b l e t o 0 ’ s and 1 ’ s
coronary$group <− recode ( coronary$group , " 1 " = 0 , " 2 " = 1 )
# 2 . F i t a simple l o g i s t i c r e g r e s s i o n where group = outcome
#and time = number o f seconds walked on the t r e a d m i l l
coronary . l o g i s t i c <− glm ( coronary$group ~ coronary$time , family=binomial )
# 3 . Examine the output summary
summary ( coronary . l o g i s t i c )
# 4 . R a l s o prov ides the AIC f o r t h i s model
coronary . l o g i s t i c . n u l l <− glm ( coronary$group ~ 1 , family=binomial )
summary ( coronary . l o g i s t i c . n u l l )
# 5 . L o g i s t i c Regression f o r an Ordinal Outcome Variable
#5.1 Load the data
cooking <− read . csv ( f i l e . choose ( ) )
#5.2 view the input data
View ( cooking )
#5.3 The dependent v a r i a b l e , cook ,
#must be an R f a c t o r o b j e c t ,
#and the independent v a r i a b l e can be e i t h e r a f a c t o r or numeric .
#In t h i s case , treatment i s coded as 0 ( c o n t r o l ) and 1 ( treatment ) .
#To ensure that cook i s a f a c t o r
487
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
cook <− as . f a c t o r ( cooking$cook )
#5.4 In order t o f i t a cumulative LOGIT model t o t h i s data in R,
#use the p o l r ( ) f u n c t i o n in MASS
cooking . cum . l o g i t <− p o l r ( cook ~ cooking$treatment , method=c ( " l o g i s t i c " ) )
#5.5 Examine the output summary t a b l e
summary ( cooking . cum . l o g i t )
#5.6 Use the deviance , along with the appropriate degrees o f freedom ,
# t o obtain a t e s t o f the n u l l hypothesis that the model f i t s the data .
#The f o l l o w i n g command l i n e in R w i l l do t h i s f o r us .
1− pchisq ( deviance ( cooking . cum . l o g i t ) , df . r e s i d u a l ( cooking . cum . l o g i t ) )
# 6 . Multinomial L o g i s t i c Regression
#6.1 We need t o load i n t o R the data f i l e P o l i t i c s . csv
# i n t o a data frame ’ p o l i t i c s ’
p o l i t i c s <− read . csv ( f i l e . choose ( ) )
View ( p o l i t i c s )
#6.2 Wish t o use multinom ( ) f u n c t i o n from the nnet package ,
# so i n s t a l l t h i s package
i n s t a l l . packages ( " nnet " )
l i b r a r y ( nnet )
#6.3 To f i t the multinomial model
p o l i t i c s . multinom <− multinom ( viewpoint ~ age , data= p o l i t i c s )
# This message simply i n d i c a t e s the i n i t i a l
#and f i n a l values o f the maximum l i k e l i h o o d f i t t i n g function ,
# along with the information that the model converged .
# To look at the parameter estimates and standard e r r o r s ,
#we use the summary ( ) f u n c t i o n
summary ( p o l i t i c s . multinom )
488
Electronic copy available at: https://ssrn.com/abstract=3863563
8.36. GENERALIZED LINEAR MODELS RECAP
# 7 . MODELS FOR COUNT DATA
#NOTES: Look at 2 types ( 1 ) Poisson , and ( 2 ) Overdispersed data
#7.1 Poisson Regression
s e s _ b a b i e s <− read . csv ( f i l e . choose ( ) )
#7.2 Then attach i t using the attach ( ) f u n c t i o n
attach ( s e s _ b a b i e s )
#7.3 In order t o see the d i s t r i b u t i o n o f the number o f babies ,
#use the histogram command h i s t
h i s t ( babies )
#7.4 See that 0 was the most common response by i n d i v i d u a l s in the sample ,
#with the maximum number being 3 .
#In order t o f i t the model with the glm function ,
#we would use the f o l l o w i n g f u n c t i o n c a l l .
babies . poisson <−glm ( babies~ s e i , data= ses_babies ,
family=c ( " poisson " ) )
#7.5 In t h i s command sequence , we c r e a t e an o b j e c t c a l l e d babies . poisson ,
#which i n c l u d e s the output f o r the Poisson r e g r e s s i o n model
#7.6 Using the summary ( babies . poisson ) y i e l d s the f o l l o w i n g output .
summary ( babies . poisson )
#7.7 These r e s u l t s show that s e i did not have a s t a t i s t i c a l l y s i g n i f i c a n t
# r e l a t i o n s h i p with the number o f c h i l d r e n under s i x months o l d l i v i n g in the home .
#Can use the f o l l o w i n g command t o obtain the p−value f o r the t e s t o f
#the n u l l hypothesis that the model f i t s the data .
1− pchisq ( deviance ( babies . poisson ) ,
df . r e s i d u a l ( babies . poisson ) )
#NOTES: The r e s u l t i n g p i s c l e a r l y not s i g n i f i c a n t at alpha = 0 . 0 5 ,
# suggesting that the model does appear t o f i t the data adequately .
#The AIC w i l l be u s e f u l as we compare the r e l a t i v e f i t o f the Poisson r e g r e s s i o n
489
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#model with that o f other models f o r count data .
#7.8 Models f o r Overdispersed Count Data
# 7 . 8 . 1 The quasipoisson model can be f i t in R using the glm function ,
#with the family s e t t o quasipoisson : babies .
quasipoisson <− glm ( babies ~ s e i , data = ses_babies , family = c ( " quasipoisson " ) )
summary ( quasipoisson )
#NOTES: The c o e f f i c i e n t s themselves are the same in the quasipoisson
#and Poisson r e g r e s s i o n models .
#However , the standard e r r o r s in the former are somewhat
# l a r g e r than those in the l a t t e r .
# 7 . 8 . 2 Test f o r model f i t as we did with the Poisson r e g r e s s i o n ,
#using the command
1− pchisq ( deviance ( quasipoisson ) , df . r e s i d u a l ( quasipoisson ) )
#NOTES: And , as with the Poisson , the quasipoisson model
# a l s o f i t s the data adequately
# 7 . 8 . 3 The negative binomial d i s t r i b u t i o n
#can be f i t t o the data in R using the glm . nb f u n c t i o n that
# i s part o f the MASS l i b r a r y . For the current example ,
#the R commands t o f i t the negative binomial model and obtain the output
babies . nb <− glm . nb ( babies ~ s e i , data= s e s _ b a b i e s )
summary ( babies . nb )
#NOTES: Just as the quasipoisson r e g r e s s i o n ,
#the parameter estimates f o r the negative binomial r e g r e s s i o n are
# i d e n t i c a l t o those f o r the Poisson . In terms o f determining which model i s optimal ,
#we can compare the AIC from the negative binomial ( 1 8 2 9 . 5 ) t o
# that o f the Poisson ( 1 9 6 2 . 2 ) ,
# t o conclude that the former p rovid es somewhat b e t t e r f i t t o the data than the l a t t e
490
Electronic copy available at: https://ssrn.com/abstract=3863563
8.37. INTRODUCTION TO BAYESIAN PROBABILITY
8.37
Introduction to Bayesian probability
#INTRODUCTION TO BAYESIAN PROBABILITY_EXERCISE
# 1 . SOME SAMPLE TESTS
#1.1 Vampire t e s t
Pr_Positive_Vampire < −0.95
Pr_Positive_Mortal < −0.01
Pr_Vampire < −0.001
P r _ P o s i t i v e <− Pr_Positive_Vampire * Pr_Vampire+
Pr_Positive_Mortal *(1 − Pr_Vampire )
( Pr_Vampire_Positive<−Pr_Positive_Vampire * Pr_Vampire / P r _ P o s i t i v e )
#NOTES: Extremely simple procedure t o code in R,
# that shows true l i k e l i h o o d o f a p o s i t i v e vampire t e s t o f 8.7%
#1.2 Cancer t e s t
P_Cancer <− 1/100000
P_Positive_Cancer <− 0.999
P_No_Cancer <− 1 − P_Cancer
P_Positive_No_Cancer <− 1 − P_Positive_Cancer
P _ T e s t _ P o s i t i v e <− P_Positive_Cancer * P_Cancer / ( 2 * P_Positive_Cancer * P_Cancer +
1 − P_Positive_Cancer −
P_Cancer )
P_Test_Positive_Percent <− P _ T e s t _ P o s i t i v e * 100
round ( P_Test_Positive_Percent , d i g i t s =1)
#NOTES: Answer i s 1 percent
# 2 . THE BETA DISTRIBUTION
i n t e g r a t e ( f u n c t i o n ( p ) dbeta ( p , 1 4 , 2 7 ) , 0 , 0 . 5 )
491
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 2 . 1 . I l l u s t r a t i o n que stion s
# 2 . 1 . 1 Question 1
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x , 4 , 6 ) , 0 . 6 , 1 )
# 2 . 1 . 2 Question 2
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x , 9 , 1 1 ) , 0 . 4 5 , 0 . 5 5 )
# 2 . 1 . 3 Question 3
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x , 1 0 9 , 1 1 1 ) , 0 . 4 5 , 0 . 5 5 )
# 3 . BAYESIAN PRIORS WORKING WITH PROBABILITY DISTRIBUTIONS
# 3 . 1 . 1 Question 1
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x , 6 , 1 ) , 0 . 4 , 0 . 6 )
# 3 . 1 . 2 Question 2
p r i o r . val <− 10
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x ,6+ p r i o r . val ,1+ p r i o r . val ) , 0 . 4 , 0 . 6 )
p r i o r . val <− 55
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x ,6+ p r i o r . val ,1+ p r i o r . val ) , 0 . 4 , 0 . 6 )
# 3 . 1 . 3 Question 3
more . heads <− 5
i n t e g r a t e ( f u n c t i o n ( x ) dbeta ( x ,6+ p r i o r . val+more . heads ,1+ p r i o r . val ) , 0 . 4 , 0 . 6 )
# 4 . BAYESIAN RETHINKING
#4.0 I n s t a l l and load the l i b r a r y
i n s t a l l . packages ( c ( ’ coda ’ , ’ mvtnorm ’ ) )
o p t i o n s ( repos=c ( getOption ( ’ repos ’ ) , rethinking = ’ http : / / xcelab . net / R ’ ) )
i n s t a l l . packages ( ’ rethinking ’ , type = ’ source ’ )
492
Electronic copy available at: https://ssrn.com/abstract=3863563
8.37. INTRODUCTION TO BAYESIAN PROBABILITY
l i b r a r y ( rethinking )
# 4 . 3 Gaussian Model o f Height
data ( Howell1 )
d <−Howell1
# 4 . 3 . 1 Structure o f d
str (d)
# 4 . 3 . 2 P r e c i s summary f u n c t i o n o f d
p r e c i s ( d , h i s t = FALSE)
# 4 . 3 . 3 Can look at histograms q u i c k l y o f the f i r s t 3 v a r i a b l e s
h i s t ( d$height )
h i s t ( d$weight )
h i s t ( d$age )
d$height
# 4 . 3 . 4 F i l t e r with those age 18 and over , so we c r e a t e a new data frame d2
d2 <− d [ d$age >= 18 , ]
dim ( d2 )
#NOTES: Should have 352 rows ( i n d i v i d u a l s ) in i t
# 4 . 3 . 5 P l o t d e n s i t y o f height f o r d2 data frame
dens ( d2$height )
#4.3.6 Plot priors
curve ( dnorm ( x , 1 7 8 , 2 0 ) , from =100 , t o =250)
#NOTES:
# 1 . Here mu i s 178 cm
# 2 . Computer shows that most i n d i v i d u a l s are between 140 cm and 220 cm
# 4 . 3 . 7 The sigma p r i o r i s a t r u l y f l a t p r i o r , a uniform one ,
# that f u n c t i o n s j u s t t o c o n s t r a i n t o have p o s i t i v e p r o b a b i l i t y between zero
493
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#and 50cm . Can view
curve ( dunif ( x , 0 , 5 0 ) , from= −10, t o =60)
#NOTES: A standard d e v i a t i o n must be p o s i t i v e , so bounding at zero makes sense
#NOTES:
# 1 . P r i o r P r e d i c t i v e Simulation i s an e s s e n t i a l part o f the modeling p r o c e s s
# 2 . Quickly simulate heights by sampling from the p r i o r
# 4 . 3 . 8 Remember , every p o s t e r i o r i s a l s o p o t e n t i a l l y a p r i o r
# f o r a subsequent analysis , so can p r o c e s s p r i o r s j u s t l i k e p o s t e r i o r s
sample_mu <−rnorm ( 1 e4 , 1 7 8 , 2 0 )
sample_sigma <− r u n i f ( 1 e4 , 0 , 5 0 )
p r i o r _ h <−rnorm ( 1 e4 , sample_mu , sample_sigma )
dens ( p r i o r _ h )
# 4 . 3 . 9 Let use simulation t o see the implied heights
sample_mu <−rnorm ( 1 e4 , 1 7 8 , 1 0 0 )
p r i o r _ h <−rnorm ( 1 e4 , sample_mu , sample_sigma )
dens ( p r i o r _ h )
#4.4 Grid Approximation o f the P o s t e r i o r D i s t r i b u t i o n
# 4 . 4 . 1 Sample code without explanation
mu. l i s t <−seq ( from =150 , t o =160 , length . out =100)
sigma . l i s t <−seq ( from =7 , t o =9 , length . out =100)
post <−expand . g r i d (mu=mu. l i s t , sigma=sigma . l i s t )
post$LL <− sapply ( 1 : nrow ( post ) , f u n c t i o n ( i )sum(
dnorm ( d2$height , post$mu [ i ] , post$sigma [ i ] , l o g =TRUE) ) )
post$prod <−post$LL+dnorm ( post$mu , 1 7 8 , 2 0 ,TRUE)+
dunif ( post$sigma , 0 , 5 0 ,TRUE)
post$prob <−exp ( post$prod −max( post$prod ) )
# 4 . 4 . 2 I n s p e c t the p o s t e r i o r d i s t r i b u t i o n , now r e s i d i n g in post$prob ,
#using a v a r i e t y o f p l o t t i n g commands . Get a simple contour p l o t with :
contour_xyz ( post$mu , post$sigma , post$prob )
494
Electronic copy available at: https://ssrn.com/abstract=3863563
8.37. INTRODUCTION TO BAYESIAN PROBABILITY
# P l o t a simple heat map with
image_xyz ( post$mu , post$sigma , post$prob )
#NOTES: These f u n c t i o n s countour_xyz and image_xyz are both contained
# in the ’ rethinking ’ package
#4.5 Sampling From The P o s t e r i o r
#NOTES: The only t r i c k i s that there are two new parameters ,
#and t r y t o sample combinations o f them
# 4 . 5 . 1 Randomly sample row numbers in ’ post ’ in p r o p o r t i o n
# t o the values in ’ post$prob ’
sample . rows<−sample ( 1 : nrow ( post ) , s i z e =1e4 , r e p l a c e =TRUE,
prob=post$prob )
sample .mu <−post$mu [ sample . rows ]
sample . sigma <−post$sigma [ sample . rows ]
# 4 . 5 . 2 End up with 10 ,000 samples , with replacement ,
#from the p o s t e r i o r f o r the height data
#1. plot
p l o t ( sample .mu, sample . sigma , cex = 0 . 5 , pch =16 , c o l = c o l . alpha ( rangi2 , 0 . 1 ) )
# 2 . Think o f them l i k e data and c h a r a c t e r i z e the shapes
# o f the marginal p o s t e r i o r s l i k e mu and sigma
dens ( sample .mu)
dens ( sample . sigma )
# 3 . The term ’ marginal ’ in t h i s c o n t e x t means averaging over the other parameters .
#To summarize the widths o f these d e n s i t i e s with p o s t e r i o r c o m p a t a b i l i t y i n t e r v a l s
PI ( sample .mu)
PI ( sample . sigma )
# 4 . Quickly sample 20 o f the heights from the height data
d3<−sample ( d2$height , s i z e =20)
495
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
# 5 . Repeat the code from b e f o r e
mu. l i s t <−seq ( from =150 , t o =170 , length . out =200)
sigma . l i s t <−seq ( from =4 , t o =20 , length . out =200)
post2 <−expand . g r i d (mu=mu. l i s t , sigma=sigma . l i s t )
post2$LL <− sapply ( 1 : nrow ( post2 ) , f u n c t i o n ( i )
sum( dnorm ( d3 , mean=post2$mu [ i ] , sd=post2$sigma [ i ] ,
l o g =TRUE ) ) )
post2$prod <−post2$LL+dnorm ( post2$mu , 1 7 8 , 2 0 ,TRUE)+
dunif ( post2$sigma , 0 , 5 0 ,TRUE)
post2$prob <−exp ( post2$prod−max( post2$prod ) )
sample2 . rows <−sample ( 1 : nrow ( post2 ) , s i z e =1e4 , r e p l a c e =TRUE,
prob=post2$prob )
sample2 .mu <−post2$mu [ sample2 . rows ]
sample2 . sigma <−post2$sigma [ sample2 . rows ]
p l o t ( sample2 .mu, sample2 . sigma , cex = 0 . 5 ,
c o l = c o l . alpha ( rangi2 , 0 . 1 ) ,
xlab ="mu" , ylab =" sigma " , pch =16)
# 6 . With t h i s new p l o t , can i n s p e c t the marginal p o s t e r i o r d e n s i t y
dens ( sample2 . sigma , norm . comp=TRUE)
496
Electronic copy available at: https://ssrn.com/abstract=3863563
8.38. MARKOV CHAIN MONTE CARLO
8.38
Markov Chain Monte Carlo
#Markov Chain Monte Carlo (MCMC)_EXERCISE
#0.1 Load and l i b r a r y the packages
i n s t a l l . packages ( " mcmc " )
l i b r a r y (mcmc)
#0.2 Load the data
data ( l o g i t )
out <− glm ( y ~ x1 + x2 + x3 + x4 , data = l o g i t , family = binomial , x = TRUE)
summary ( out )
#0.3 Set up the f u n c t i o n
l u p o s t _ f a c t o r y <− f u n c t i o n ( x , y ) f u n c t i o n ( beta ) {
eta <− as . numeric ( x %*% beta )
logp <− i f e l s e ( eta < 0 , eta − log1p ( exp ( eta ) ) , − log1p ( exp(− eta ) ) )
logq <− i f e l s e ( eta < 0 , − log1p ( exp ( eta ) ) , − eta − log1p ( exp(− eta ) ) )
l o g l <− sum( logp [ y == 1 ] ) + sum( logq [ y == 0 ] )
return ( l o g l − sum( beta ^2) / 8 )
}
l u p o s t <− l u p o s t _ f a c t o r y ( out$x , out$y )
# 1 . Beginning MCMC
#NOTES: With those d e f i n i t i o n s in place ,
#the f o l l o w i n g code runs the Metropolis algorithm t o simulate the p o s t e r i o r .
s e t . seed (88866) #NOTES: t o get r e p r o d u c i b l e r e s u l t s
beta . i n i t <− as . numeric ( c o e f f i c i e n t s ( out ) )
out <− metrop ( lupost , beta . i n i t , 1e3 )
names ( out )
#1.1 Try f o r 20 percent
out <− metrop ( out , s c a l e = 0 . 1 )
out$accept
out <− metrop ( out , s c a l e = 0 . 3 )
out$accept
497
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
out <− metrop ( out , s c a l e = 0 . 5 )
out$accept
out <− metrop ( out , s c a l e = 0 . 4 )
out$accept
#2. Diagnostics
#2.1 That does i t f o r the acceptance r a t e .
#So do a longer run and look at the r e s u l t s .
out <− metrop ( out , nbatch = 1e4 )
t . t e s t ( out$accept . batch ) $conf . i n t
out$time
#2.2 Time s e r i e s p l o t
p l o t ( t s ( out$batch ) )
#2.3 Another way t o look at the output i s an a u t o c o r r e l a t i o n p l o t .
a c f ( out$batch )
#NOTES:
# 1 . As with any m u l t i p l o t p l o t , these are a b i t hard t o read .
# 2 . Students are i n v i t e d t o make the separate p l o t s t o get a b e t t e r p i c t u r e .
# 3 . As with a l l d i a g n o s t i c p l o t s in MCMC, these do not diagnose s u b t l e problems .
# 3 . Monte Carlo Estimates and Standard Errors
out <− metrop ( out , nbatch = 100 , blen = 100 , outfun = f u n c t i o n ( z ) c ( z , z ^ 2 ) )
t . t e s t ( out$accept . batch ) $conf . i n t
#NOTES: Have added an argument outfun
# that g i v e s the f u n c t i o n a l o f the Markov chain we want t o average
# 4 . Simple Means
#4.1 The grand means ( means o f batch means ) are
498
Electronic copy available at: https://ssrn.com/abstract=3863563
8.38. MARKOV CHAIN MONTE CARLO
apply ( out$batch , 2 , mean )
#NOTES:
# 1 . The f i r s t 5 numbers are the Monte Carlo estimates o f the p o s t e r i o r means .
# 2 . The second 5 numbers are the Monte Carlo estimates o f
#the p o s t e r i o r ordinary second moments .
#4.2 Get the p o s t e r i o r variances by
f o o <− apply ( out$batch , 2 , mean )
mu <− f o o [ 1 : 5 ]
sigmasq <− f o o [ 6 : 1 0 ] − mu^2
mu
#NOTES: Monte Carlo standard e r r o r s (MCSE) are c a l c u l a t e d from the batch means .
#This i s sim plest f o r the means .
# 5 . Functions o f Means
#NOTES:
# 1 . To get the MCSE f o r the p o s t e r i o r variances ,
# t r y t o apply the d e l t a method ( see PowerPoint )
#5.1 Now use the f o l l o w i n g code ,
# f o r parameters u and v , average , delta , sigmasq ( var ) e t c .
u <− out$batch [ , 1 : 5 ]
v <− out$batch [ , 6 : 1 0 ]
ubar <− apply ( u , 2 , mean )
vbar <− apply ( v , 2 , mean )
deltau <− sweep ( u , 2 , ubar )
d e l t a v <− sweep ( v , 2 , vbar )
f o o <− sweep ( deltau , 2 , ubar , " * " )
sigmasq . mcse <− s q r t ( apply ( ( d e l t a v − 2 * f o o ) ^ 2 , 2 , mean ) / out$nbatch )
sigmasq . mcse
#NOTES: Gives the MCSE f o r the p o s t e r i o r variance .
#5.2 Just check that t h i s complicated sweep
499
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
#and apply s t u f f does do the r i g h t thing .
s q r t ( mean ( ( ( v [ , 2 ] − vbar [ 2 ] ) − 2 * ubar [ 2 ] * ( u [ , 2 ] − ubar [ 2 ] ) ) ^ 2 ) / out$nbatch
#NOTES: The standard d e v i a t i o n o f the batch mean c a l l e d v here ,
#appears t o be very c o n s e r v a t i v e .
# 6 . Functions o f Functions o f Means
#6.1 About the p o s t e r i o r standard d e v i a t i o n ,
#the d e l t a method g i v e s i t s standard e r r o r in terms o f that f o r the variance .
sigma <− s q r t ( sigmasq )
sigma . mcse <− sigmasq . mcse / ( 2 * sigma )
sigma
sigma . mcse
# 7 . A FINAL RUN
#7.1 So that ’ s i t .
#The only thing l e f t t o do i s a l i t t l e more p r e c i s i o n
# ( the exam problem d i r e c t e d use a long enough run o f the Markov chain sampler
# so that the MCSE are l e s s than 0 . 0 1 )
out <− metrop ( out , nbatch = 500 , blen = 400)
t . t e s t ( out$accept . batch ) $conf . i n t
out$time
f o o <− apply ( out$batch , 2 , mean )
mu <− f o o [ 1 : 5 ]
sigmasq <− f o o [ 6 : 1 0 ] − mu^2
mu
sigmasq
mu. mcse <− apply ( out$batch [ , 1 : 5 ] , 2 , sd ) / s q r t ( out$nbatch )
mu. mcse
500
Electronic copy available at: https://ssrn.com/abstract=3863563
8.38. MARKOV CHAIN MONTE CARLO
#7.2 Run parameter estimation again , with standard e r r o r s
u <− out$batch [ , 1 : 5 ]
v <− out$batch [ , 6 : 1 0 ]
ubar <− apply ( u , 2 , mean )
vbar <− apply ( v , 2 , mean )
deltau <− sweep ( u , 2 , ubar )
d e l t a v <− sweep ( v , 2 , vbar )
f o o <− sweep ( deltau , 2 , ubar , " * " )
sigmasq . mcse <− s q r t ( apply ( ( d e l t a v − 2 * f o o ) ^ 2 , 2 , mean ) / out$nbatch )
sigmasq . mcse
sigma <− s q r t ( sigmasq )
sigma . mcse <− sigmasq . mcse / ( 2 * sigma )
sigma
sigma . mcse
# 8 . New Variance Estimation Functions
#NOTES:
# 1 . R f u n c t i o n i n i t s e q ( added in v e r s i o n 0 . 6 o f t h i s package ,
#now v e r s i o n 0 . 9 7 ) estimates variances
# in the Markov chain c e n t r a l l i m i t theorem (CLT) f o l l o w i n g the methodology
# introduced by Geyer ( 1 9 9 2 )
# 2 . These methods only apply t o s c a l a r −valued f u n c t i o n a l s
# o f r e v e r s i b l e Markov chains , but the Markov chains produced
#by the metrop f u n c t i o n s a t i s f y t h i s c o n d i t i o n ,
#even , as we s h a l l see below , when batching i s used .
# 3 . Rather than redo the Markov chains in the preceding material ,
#we j u s t look at a toy problem , an AR( 1 ) time s e r i e s ,
#which can be simulated in one l i n e o f R .
#This i s the example on the help page f o r i n i t s e q .
n <− 2e4
501
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
rho <− 0.99
x <− arima . sim ( model = l i s t ( ar = rho ) , n = n )
#NOTES: The time s e r i e s x i s a r e v e r s i b l e Markov chain
#and t r i v i a l l y a s c a l a r −valued f u n c t i o n a l o f a Markov chain .
#8.1 Time s e r i e s p l o t
out <− i n i t s e q ( x )
p l o t ( seq ( along = out$Gamma . pos ) − 1 , out$Gamma . pos ,
xlab = " k " ,
ylab = e xpr ess ion (Gamma[ k ] ) ,
type = " l " )
l i n e s ( seq ( along = out$Gamma . dec ) − 1 , out$Gamma . dec , l t y = " dotted " )
l i n e s ( seq ( along = out$Gamma . con ) − 1 , out$Gamma . con , l t y = " dashed " )
#NOTES:
# 1 . I f use View ( out ) and examine Gamma, can see i t has 3 s u f f i x e s ;
#pos , dec and con
# 2 . S o l i d l i n e , i n i t i a l p o s i t i v e sequence estimator . Dotted l i n e ,
# i n i t i a l monotone sequence estimator . Dashed l i n e ,
# i n i t i a l convex sequence estimator .
#8.2 What i s important , i s an estimate o f variance or sigma−squared ,
#which i s given by
out$var . con
( 1 + rho ) / ( 1 − rho ) * 1 / ( 1 − rho ^2)
#8.3 An example o f the exact t h e o r e t i c a l value o f sigmasq .
#The sequence o f batch means i s i t s e l f a s c a l a r − valued f u n c t i o n a l
# o f a r e v e r s i b l e Markov chain . Hence the i n i t i a l sequence esti mato rs
#can be applied t o i t .
blen <− 5
x . batch <− apply ( matrix ( x , nrow = blen ) , 2 , mean )
bout <− i n i t s e q ( x . batch )
502
Electronic copy available at: https://ssrn.com/abstract=3863563
8.38. MARKOV CHAIN MONTE CARLO
#NOTES: Because the batch length i s t o o short ,
#the variance o f the batch means does not estimate sigmasq
#8.4 Must account f o r the a u t o c o r r e l a t i o n o f the batches
p l o t ( seq ( along = bout$Gamma . con ) − 1 ,
bout$Gamma . con ,
xlab = " k " ,
ylab = e xpr ess ion (Gamma[ k ] ) ,
type = " l " )
#8.5 Because the the variance i s p r o p o r t i o n a l t o one over the batch length ,
#need t o multiply
#by the batch length t o estimate the sigmasq f o r the o r i g i n a l s e r i e s
out$var . com
bout$var . con * blen
#8.6 Another way t o look at t h i s i s that the MCMC estimator o f
#mu i s e i t h e r mean ( x ) or mean ( x . batch )
#8.7 And the variance must be d i v i d e d by
#the sample s i z e t o g i v e standard e r r o r s . So e i t h e r
mean ( x ) + c ( − 1 , 1 ) * qnorm ( 0 . 9 7 5 ) * s q r t ( out$var . con / length ( x ) )
mean ( x . batch ) + c ( − 1 , 1 ) * qnorm ( 0 . 9 7 5 ) * s q r t ( bout$var . con / length ( x . batch ) )
#NOTES:
# 1 . Result i s an asymptotic 95 percent c o n f i d e n c e i n t e r v a l f o r mu
# 2 . Just d i v i d e by the r e l e v a n t sample s i z e
# 9 . Only One Function Argument
#9.1 Previous v e r s i o n s o f t h i s v i g n e t t e used the dot−dot−dot mechanism everywhere
#9.2 In those v e r s i o n s the l o g un−normalized d e n s i t y f u n c t i o n was def ined by
l u p o s t <− f u n c t i o n ( beta , x , y ) {
eta <− as . numeric ( x %*% beta )
503
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
logp <− i f e l s e ( eta < 0 , eta − log1p ( exp ( eta ) ) , − log1p ( exp(− eta ) ) )
logq <− i f e l s e ( eta < 0 , − log1p ( exp ( eta ) ) , − eta − log1p ( exp(− eta ) ) )
l o g l <− sum( logp [ y == 1 ] ) + sum( logq [ y == 0 ] )
return ( l o g l − sum( beta ^2) / 8 )
}
#NOTES:
# 1 . Rather than the way i t i s now done in the main example b e f o r e .
# 2 . Note that everything i s the same in the two d e f i n i t i o n s except here we have
# 3 . l u p o s t <− f u n c t i o n ( beta , x , y ) {
# 4 . Where in the main t e x t we now have
# 5 . l u p o s t _ f a c t o r y <− f u n c t i o n ( x , y ) f u n c t i o n ( beta ) {
# 6 . Then we have t o execute the f u n c t i o n f a c t o r y l u p o s t _ f a c t o r y
# t o make the l u p o s t f u n c t i o n .
#9.3 So t o use t h i s l u p o s t function ,
#have t o add arguments x and y t o each c a l l t o R f u n c t i o n metrop .
# For example
out <− glm ( y ~ x1 + x2 + x3 + x4 , data = l o g i t , family = binomial , x = TRUE)
x <− out$x
y <− out$y
out <− metrop ( lupost , beta . i n i t , 1e3 , x = x , y = y )
out$accept
out <− metrop ( out , s c a l e = 0 . 1 , x = x , y = y )
out$accept
out <− metrop ( out , s c a l e = 0 . 3 , x = x , y = y )
out$accept
out <− metrop ( out , s c a l e = 0 . 5 , x = x , y = y )
out$accept
out <− metrop ( out , s c a l e = 0 . 4 , x = x , y = y )
out$accept
# and so on
504
Electronic copy available at: https://ssrn.com/abstract=3863563
8.38. MARKOV CHAIN MONTE CARLO
#NOTES: This method has the b e n e f i t that
#we do not have t o explain f u n c t i o n f a c t o r i e s
#and has the drawback that we need t o keep remembering t o
#add x = x , y = y t o each i n v o c a t i o n o f R f u n c t i o n metrop .
#10. More Than One Function Argument
#NOTES: S i t u a t i o n becomes more complicated when more than
#one f u n c t i o n argument i s passed t o the higher−order f u n c t i o n .
#Then they a l l must handle the same dot−dot−dot arguments whether or not
#they want them .
#10.1 So now must d e f i n e the output f u n c t i o n as :
outfun <− f u n c t i o n ( z ,
. . . ) c ( z , z ^2)
#NOTES: the . . . argument in the f u n c t i o n signature i s e s s e n t i a l
#because t h i s f u n c t i o n i s going t o be passed dot−dot−dot arguments x and y ,
#which i t does not need and does not want ,
# so i t has t o allow f o r them ( and then not use them ) .
#10.2 Then can continue
out <− metrop ( out , nbatch = 100 , blen = 100 , outfun = outfun , x = x , y = y )
out$accept
#NOTES: Have gotten an e r r o r about unused arguments
# i f we had def ined the output f u n c t i o n without
#the dot−dot−dot as we did in the main example at the s t a r t .
#11. Global Variables
#NOTES: Already in the preceding s e c t i o n de fined R o b j e c t s x and y as
# g l o b a l v a r i a b l e s ( in the R g l o b a l environment . GlobalEnv ) .
#They are g l o b a l v a r i a b l e s and use them as such , d e f i n i n g
l u p o s t <− f u n c t i o n ( beta ) {
505
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
eta <− as . numeric ( x %*% beta )
logp <− i f e l s e ( eta < 0 , eta − log1p ( exp ( eta ) ) , − log1p ( exp(− eta ) ) )
logq <− i f e l s e ( eta < 0 , − log1p ( exp ( eta ) ) , − eta − log1p ( exp(− eta ) ) )
l o g l <− sum( logp [ y == 1 ] ) + sum( logq [ y == 0 ] )
return ( l o g l − sum( beta ^2) / 8 )
}
outfun <− f u n c t i o n ( z ) c ( z , z ^2)
#11.1 Then the f o l l o w i n g works
out <− metrop ( lupost , beta . i n i t , 1e3 )
out$accept
out <− metrop ( out , s c a l e = 0 . 1 )
out$accept
out <− metrop ( out , s c a l e = 0 . 3 )
out$accept
out <− metrop ( out , s c a l e = 0 . 5 )
out$accept
out <− metrop ( out , s c a l e = 0 . 4 )
out$accept
#11.2 Set batch length t o 100 , and the same f o r batch number
out <− metrop ( out , nbatch = 100 , blen = 100 , outfun = outfun )
out$accept
#11.3 Get the best o f both worlds . We don ’ t need x = x , y = y and
#we don ’ t need a f u n c t i o n f a c t o r y .
out <− metrop ( out , s c a l e = 0 . 1 , x = modmat , y = resp )
#NOTES: See using x = modmat may not be de fined
#11.4 Using the f u n c t i o n f a c t o r y
l u p o s t <− l u p o s t _ f a c t o r y ( modmat , resp )
506
Electronic copy available at: https://ssrn.com/abstract=3863563
8.38. MARKOV CHAIN MONTE CARLO
#12. Function f a c t o r y
#12.1 Compare
f r e d <− f u n c t i o n ( y ) f u n c t i o n ( x ) x + y
fred ( 2 ) ( 3 )
l u p o s t _ f a c t o r y <− f u n c t i o n ( x , y ) f u n c t i o n ( beta ) {
eta <− as . numeric ( x %*% beta )
logp <− i f e l s e ( eta < 0 , eta − log1p ( exp ( eta ) ) , − log1p ( exp(− eta ) ) )
logq <− i f e l s e ( eta < 0 , − log1p ( exp ( eta ) ) , − eta − log1p ( exp(− eta ) ) )
l o g l <− sum( logp [ y == 1 ] ) + sum( logq [ y == 0 ] )
return ( l o g l − sum( beta ^2) / 8 )
}
l u p o s t <− l u p o s t _ f a c t o r y ( x , y )
l u p o s t ( beta . i n i t )
#NOTES: Could a l s o do the same c a l c u l a t i o n t r e a t i n g l u p o s t _ f a c t o r y as
# j u s t an ordinary c u r r i e d function , l i k e R f u n c t i o n f r e d in the preceding example ,
l u p o s t _ f a c t o r y ( x , y ) ( beta . i n i t )
507
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
8.39
LDA Topic Modelling
#0.1 I n s t a l l and load the Packages
i n s t a l l . packages ( " t e x t c l e a n " )
# 0 . 1 . 1 Data wrangling
l i b r a r y ( dplyr )
library ( tidyr )
library ( lubridate )
#0.1.2 Visualization
l i b r a r y ( ggplot2 )
# 0 . 1 . 3 Dealing with t e x t
library ( textclean )
l i b r a r y (NLP)
l i b r a r y ( tm )
l i b r a r y ( SnowballC )
library ( stringr )
# 0 . 1 . 4 Topic model
library ( tidytext )
l i b r a r y ( topicmodels )
l i b r a r y ( textmineR )
data <− read . csv ( f i l e . choose ( ) )
head ( data )
data <− data %>%
mutate ( o v e r a l l = as . f a c t o r ( o v e r a l l ) ,
reviewTime = s t r _ r e p l a c e _ a l l ( reviewTime , pattern = " " ,
replacement = " − ") ,
reviewTime = s t r _ r e p l a c e ( reviewTime , pattern = " , " ,
replacement = " " ) ,
reviewTime = mdy( reviewTime ) ) %>%
s e l e c t ( reviewText , o v e r a l l , reviewTime )
head ( data )
508
Electronic copy available at: https://ssrn.com/abstract=3863563
8.39. LDA TOPIC MODELLING
data <− data %>%
mutate ( o v e r a l l = as . f a c t o r ( o v e r a l l ) ,
reviewTime = s t r _ r e p l a c e _ a l l ( reviewTime , pattern = " " ,
replacement = " − ") ,
reviewTime = s t r _ r e p l a c e ( reviewTime , pattern = " , " ,
replacement = " " ) ,
reviewTime = mdy( reviewTime ) ) %>%
s e l e c t ( reviewText , o v e r a l l , reviewTime )
head ( data )
# build textcleaner function
t e x t c l e a n e r <− f u n c t i o n ( x ) {
x <− as . c h a r a c t e r ( x )
x <− x %>%
s t r _ t o _ l o w e r ( ) %>%
# convert a l l the s t r i n g t o low alphabet
r e p l a c e _ c o n t r a c t i o n ( ) %>% # r e p l a c e c o n t r a c t i o n t o t h e i r multi−word forms
r e p l a c e _ i n t e r n e t _ s l a n g ( ) %>% # r e p l a c e i n t e r n e t slang t o normal words
r e p l a c e _ e m o j i ( ) %>% # r e p l a c e emoji t o words
replace_emoticon ( ) %>% # r e p l a c e emoticon t o words
replace_hash ( replacement = " " ) %>% # remove hashtag
replace_word_elongation ( ) %>%
# r e p l a c e informal w r i t i n g with known semantic replacements
replace_number ( remove = T ) %>% # remove number
r e p l a c e _ d a t e ( replacement = " " ) %>% # remove date
replace_time ( replacement = " " ) %>% # remove time
s t r _ r e m o v e _ a l l ( pattern = " [ [ : punct : ] ] " ) %>% # remove punctuation
s t r _ r e m o v e _ a l l ( pattern = "[^\\ s ] * [0 − 9][^\\ s ] * " ) %>%
## remove mixed s t r i n g n number
s t r _ s q u i s h ( ) %>% # reduces repeated whitespace i n s i d e a s t r i n g .
s t r _ t r i m ( ) # removes whitespace from s t a r t and end o f s t r i n g
xdtm <− VCorpus ( VectorSource ( x ) ) %>%
tm_map ( removeWords , stopwords ( " en " ) )
# convert corpus t o document term matrix
return ( DocumentTermMatrix ( xdtm ) )
509
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
}
data_1 <− data %>% f i l t e r ( o v e r a l l == 1 )
data_2 <− data %>% f i l t e r ( o v e r a l l == 2 )
data_3 <− data %>% f i l t e r ( o v e r a l l == 3 )
data_4 <− data %>% f i l t e r ( o v e r a l l == 4 )
data_5 <− data %>% f i l t e r ( o v e r a l l == 5 )
table ( data$overall )
###Method 1_Text Cleaning
# apply t e x t c l e a n e r f u n c t i o n f o r review t e x t
dtm_5 <− t e x t c l e a n e r ( data_5$reviewText )
# f i n d most frequent terms . i choose words that at l e a s t appear in 50 reviews
freqterm_5 <− findFreqTerms ( dtm_5 , 5 0 )
# we have 981 words . subset the dtm t o only choose those s e l e c t e d words
dtm_5 <− dtm_5 [ , freqterm_5 ]
# only choose words that appear once in each rows
rownum_5 <− apply ( dtm_5 , 1 ,sum)
dtm_5 <− dtm_5 [ rownum_5> 0 , ]
# apply t o LDA f u n c t i o n . s e t the k = 6 , means we want t o b u i l d 6 t o p i c
lda_5 <− LDA( dtm_5 , k = 6 , c o n t r o l = l i s t ( seed = 1 5 0 2 ) )
# apply auto t i d y using t i d y and use beta as per−t o p i c −per−word p r o b a b i l i t i e s
t o p i c _ 5 <− t i d y ( lda_5 , matrix = " beta " )
# choose 15 words with highest beta from each t o p i c
top_terms_5 <− t o p i c _ 5 %>%
group_by ( t o p i c ) %>%
top_n ( 1 5 , beta ) %>%
ungroup ( ) %>%
arrange ( t o p i c , − beta )
# p l o t the t o p i c and words f o r easy i n t e r p r e t a t i o n
p l o t _ t o p i c _ 5 <− top_terms_5 %>%
mutate ( term = reorder_within ( term , beta , t o p i c ) ) %>%
g g p l o t ( aes ( term , beta , f i l l = f a c t o r ( t o p i c ) ) ) +
geom_col ( show . legend = FALSE) +
510
Electronic copy available at: https://ssrn.com/abstract=3863563
8.39. LDA TOPIC MODELLING
facet_wrap (~ t o p i c , s c a l e s = " f r e e " ) +
coord_flip ( ) +
scale_x_reordered ( )
plot_topic_5
###Method 2_TextmineR
t e x t c l e a n e r _ 2 <− f u n c t i o n ( x ) {
x <− as . c h a r a c t e r ( x )
x <− x %>%
s t r _ t o _ l o w e r ( ) %>%
# convert a l l the s t r i n g t o low alphabet
r e p l a c e _ c o n t r a c t i o n ( ) %>% # r e p l a c e c o n t r a c t i o n t o t h e i r multi−word forms
r e p l a c e _ i n t e r n e t _ s l a n g ( ) %>% # r e p l a c e i n t e r n e t slang t o normal words
r e p l a c e _ e m o j i ( ) %>% # r e p l a c e emoji t o words
replace_emoticon ( ) %>% # r e p l a c e emoticon t o words
replace_hash ( replacement = " " ) %>% # remove hashtag
replace_word_elongation ( ) %>%
## r e p l a c e informal w r i t i n g with known semantic replacements
replace_number ( remove = T ) %>% # remove number
r e p l a c e _ d a t e ( replacement = " " ) %>% # remove date
replace_time ( replacement = " " ) %>% # remove time
s t r _ r e m o v e _ a l l ( pattern = " [ [ : punct : ] ] " ) %>% # remove punctuation
s t r _ r e m o v e _ a l l ( pattern = "[^\\ s ] * [0 − 9][^\\ s ] * " ) %>%
## remove mixed s t r i n g n number
s t r _ s q u i s h ( ) %>% # reduces repeated whitespace i n s i d e a s t r i n g .
s t r _ t r i m ( ) # removes whitespace from s t a r t and end o f s t r i n g
return ( as . data . frame ( x ) )
}
# apply t e x t c l e a n e r f u n c t i o n .
# note : we only clean the t e x t without convert i t t o dtm
clean_5 <− t e x t c l e a n e r _ 2 ( data_5$reviewText )
clean_5 <− clean_5 %>% mutate ( i d = rownames ( clean_5 ) )
# c r e t e dtm
511
Electronic copy available at: https://ssrn.com/abstract=3863563
CHAPTER 8. FINANCIAL ECONOMETRICS R CODES DICTIONARY
s e t . seed ( 1 5 0 2 )
dtm_r_5 <− CreateDtm ( doc_vec = clean_5$x ,
doc_names = clean_5$id ,
ngram_window = c ( 1 , 2 ) ,
stopword_vec = stopwords ( " en " ) ,
verbose = F )
dtm_r_5 <− dtm_r_5 [ , colSums ( dtm_r_5 ) >2]
s e t . seed ( 1 5 0 2 )
mod_lda_5 <− FitLdaModel ( dtm = dtm_r_5 ,
k = 20 , # number o f t o p i c
i t e r a t i o n s = 500 ,
burnin = 180 ,
alpha = 0 . 1 , beta = 0 . 0 5 ,
optimize_alpha = T ,
calc_likelihood = T,
calc_coherence = T,
calc_r2 = T)
mod_lda_5$r2
p l o t ( mod_lda_5$log_likelihood , type = " l " )
mod_lda_5$top_terms <− GetTopTerms ( phi = mod_lda_5$phi ,M = 15)
data . frame ( mod_lda_5$top_terms )
mod_lda_5$coherence
mod_lda_5$prevalence <− colSums ( mod_lda_5$theta ) / sum( mod_lda_5$theta ) * 100
mod_lda_5$prevalence
mod_lda_5$summary <− data . frame ( t o p i c = rownames ( mod_lda_5$phi ) ,
coherence = round ( mod_lda_5$coherence , 3 ) ,
prevalence = round ( mod_lda_5$prevalence , 3 ) ,
top_terms = apply ( mod_lda_5$top_terms , 2 ,
f u n c t i o n ( x ) { paste ( x ,
collapse = " ,"
512
Electronic copy available at: https://ssrn.com/abstract=3863563
8.39. LDA TOPIC MODELLING
)}))
modsum_5 <− mod_lda_5$summary %>%
‘ rownames< − ‘(NULL)
modsum_5
modsum_5 %>% p i v o t _ l o n g e r ( c o l s = c ( coherence , prevalence ) ) %>%
g g p l o t ( aes ( x = f a c t o r ( t o p i c , l e v e l s = unique ( t o p i c ) ) , y = value , group = 1 ) ) +
geom_point ( ) + geom_line ( ) +
facet_wrap (~name , s c a l e s = " f r e e _ y " , nrow = 2 ) +
theme_minimal ( ) +
l a b s ( t i t l e = " Best t o p i c s by coherence and prevalence s c o r e " ,
s u b t i t l e = " Text review with 5 r a t i n g " ,
x = " Topics " , y = " Value " )
m o d _ l d a _ 5 $ l i n g u i s t i c <− C a l c H e l l i n g e r D i s t ( mod_lda_5$phi )
mod_lda_5$hclust <− h c l u s t ( as . d i s t ( m o d _ l d a _ 5 $ l i n g u i s t i c ) , " ward .D" )
mod_lda_5$hclust$labels <− paste ( mod_lda_5$hclust$labels , mod_lda_5$labels [ , 1 ] )
p l o t ( mod_lda_5$hclust )
513
Electronic copy available at: https://ssrn.com/abstract=3863563
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )