Time Series Analysis & Forecasting
Lecture 33
Estimation of MA(1) parameters
MLE
1
Gaussian MA(1)
MA(1) model
Xt = µ + ωt + εωt→1
Let us assume that ωt → N (0, ϑ 2 ).
2
Gaussian MA(1)
MA(1) model
Xt = µ + ωt + εωt→1
Let us assume that ωt → N (0, ϑ 2 ).
For given ωt→1
Xt |ωt→1 → N (µ + εωt→1 , ϑ 2 ).
2
Conditional Likelihood
Assume ω0 = 0, then we have:
X1 |ω0 = 0 → N (µ, ϑ 2 )
Xt µ
X
µ
Put
Et
Ott 1
E
to
4
1 in
d
X2 |X1 , ω0 = 0 = X2 |ω1 → N ((1 ↑ ε)µ + εx1 , ϑ 2 )
Puf
d
L
Eff
jeff
X3 |X2 , X1 , ω0 = 0 = X3 |ω2 → N (µ + ε(x2 ↑ µ) ↑ ε2 (x1 ↑ µ), ϑ 2 )
Iplease
..
..
verify
.
.
µ 0
1
3
Conditional Likelihood
Let us consider:
fXN ,...,X1 |ω0 =0 = fXN |XN →1 ,...,X1 ,ω0 =0 ↓ fXN →1 ,...,X1 |ω0 =0
fxn tn
Than.it
A
4
Conditional Likelihood
Therefore, we can write:
fXN ,...,X1 |ω0 =0 = fXN |ωN →1 . . . fX1 |ω0 =0
Some notation
ω = (µ, ε, ϑ 2 )↑
5
Jamie
angmax be y
2
Conditional likelihood function
LC (ω) =
N
!
t=1
Weknow that
fXt |ωt→1
t Et_
Mftott1,0
Conditional log-likelihood function
lC (ω) =
WE log 27
Eg logos
t
1,1092T
r
log
I
E
ᵗt
M 0Gt
6
Questions to think about!
• If the MA(1) model is invertible, then the e!ect of initialization will die. Why?
7
Exact MLE
Some notations
Let X = (X1 , . . . , XN )↑ , where Xt = µ + ωt + εωt→1 and ωt → N (0, ϑ 2 ). Let µ = µ1
and ! = cov(X)
in
Lin.mn
f
1
c
_on
8
Likelihood function
my
L(ϖ) = fX (x) =
7 40.02
T
1
i
(2ϱ)n/2 |!|1/2
exp
"
1
↑ (x ↑ µ)↑ !→1 (x ↑ µ)
2
#
9
Consider the decomposition:
! = ADA↑ .
Here,
1
ε
1+ε2
A= 0
.
..
0
*
0
1
ε(1+ε 2 )
1+ε 2 +ε 4
..
.
0
0
0
...
...
1
..
.
...
..
.
0
ε(1+ε 2 +...+ε 2(N →2) )
(1+ε 2 +...+ε 2(N →1) )
0
0
0
and,
..
.
1
2
4
1 + ε2 + . . . + ε2n ,
2 1+ε +ε
D = diag ϑ 1 + ε ,
,...,
1 + ε2
1 + ε2 + . . . + ε2(n→1)
+
2
10
Note that:
|A| =
|!| =
1
ADATI
AI IDI
ATI
IDI
dit
dit
where
Jitting
11
Likelihood function
We can rewrite the likelihood as follows:
L(ω) =
1
1
exp
.
1/2
(2ϱ)N/2 ( N
t=1 dtt )
"
1
→1
↑ (x ↑ µ)↑ (ADA↑ ) (x ↑ µ)
2
MAT D A 112.41
d
Define
A
µ
2 Then
attain
#
enpg yn.to
tj
tE
12
Summary
• CMLE:
• The conditional method simplifies likelihood by conditioning on initial value ω0 = 0.
• The conditional log-likelihood is given by:
N
/
↑N
N
1
lc (ω) =
log 2ϱ ↑
log ϑ 2 ↑ 2
ω2t .
2
2
2ϑ t=1
13
Summary
• EMLE:
• The joint distribution of Xt = µ + ωt + εωt→1 is multivariate normal distribution.
• Decompose ! = ADAT , where A is unit lower triangular matrix and D is diagonal
matrix - enabling fast, stable likelihood evaluation.
• Define x0 = A→1 (x ↑ µ), then the log-likelihood becomes:
N
N
↑1 /
1 / (x0t )2
l(ω) =
log dtt ↑
.
2 t=1
2 t=1 dtt
constant
• Numerically maximize the likelihood over ϖ, ensuring ε ↔ (↑1, 1) for invertibility.
14
Time Series Analysis & Forecasting
Lecture 34
Forecasting a stationary time series
Best Linear Predictor (BLP)
1
Basic Idea
Find the ‘best’ linear combination of (1, XN , XN →1 , . . . , X1 ) for forecasting XN +h .
h
o
2
Definition
BLP of XN +h based on (1, XN , XN →1 , . . . , X1 ) is the linear function
PN XN +h = a0 + a1 XN + . . . + aN X1 , if E(XN +h → PN XN +h )2 is minimum among
all such linear combinations.
3
Consider,
!
S(a) = E XN +h → a0 →
N
"
i=1
ai XN +1→i
#2
ad
al
ann
It
4
BLP equations
ωS(a)
=0
ωa0
ωS(a)
=0
ωa1
..
.
2E
Nth
90
2E
Xnn
ao
aixnti.ie
EF9j
ωS(a)
=0
ωaN
In
general
2
1
Xan
no
ai Xn i i
iii
Xn i i
n
5
On Solving
a0 = µ(1 +
N
"
ai )
i=1
εh+j→1 =
N
"
ai εi→j j = 1, . . . , N
i=1
6
In matrix notation
ω N (h) = !N a,
εh
a1
εh+1
a2
. , and !N = ((εi→j ))N .
where ω N (h) =
,
a
=
..
i,j=1
.
.
.
εh+N →1
aN
7
BLP: PN XN +h
PN XN +h = µ(1 →
=µ+
N
"
i=1
N
"
i=1
↑
ai ) +
N
"
ai XN +1→i
i=1
ai (XN +1→i → µ)
where h
= µ + a YN
Here,
a = !→1
N ω N (h)
where YN 1 i
f
NH
i
M
8
Questions to think about!
• !N will be full rank with probability 1.
9
Some notes
•
E(XN +h → PN XN +h ) = 0
In to
m
BLPequations
•
E(XN +h → PN XN +h )Xi = 0 ↑i = 1, . . . , N
10
Some notes
minimum mean squarepredictionend
•
ETXN.IM 9T
E[XN +h → PN XN +h ]2 =
Xn n
E
E
Vx o
In
at
µ
Xn n µ
Tn a
E
2
Xn n
µ
atyn
T
at Y Ya
2 a vow h
nth
M
2 10
Yn
h
In
at
11
Some notes
a
• If µ = 0, then a0 = 0.
p
11
Efi
• (A special case) One step ahead prediction:
PN XN +1 = µ + a↑ YN ,
where a is such that ω N (1) = !N a
12
Some examples
X(t): covariance stationary AR(2) process
Predict X2 based on X1 using BLP.
where a is such that
aX1
Px X2
a
not
minimzed
is
ax
X
E
gla
48
ax
E X2
2
ETX
ECKX
2
ax
a
x
0
et
13
Some examples
C S'T
X(t): covariance stationary AR(2) process
Predict X3 based on X2 and X1 using BLP.
91
Px x
we
2
4
x
a
3
a
X2
where
1
argmin
G 927
sdÉ
92
0 192
anti
X2
a
x
0
E
Xz
9,72 92 1
V1
a Po
apr
a 81
9
12
14
Homework
• Suppose X(t) is a covariance stationary AR(2) process. Predict X4 based on X3 ,
X2 , and X1 using BLP.
• Suppose X(t) is a covariance stationary AR(2) process. Predict XN +1 based on
XN , XN →1 , . . ., X1 using BLP.
15
Summary
• The goal is to find the ‘best’ predictor of XN +h using a linear combination of
(1, XN , XN →1 , . . . , X1 ).
• BLP of XN +h based on (1, XN , XN →1 , . . . , X1 ) is the linear function
PN XN +h = a0 + a1 XN + . . . + aN X1 , if E(XN +h → PN XN +h )2 is minimum
among all such linear combinations.
16
Summary
• Coeeficients aj s are obtained from:
a0 = µ(1 +
N
"
ai )
i=1
εh+j→1 =
N
"
ai εi→j j = 1, . . . , N
i=1
• BLP is model-free (no AR/MA assumption) and provides the best linear forecast
in the least squares sense.
• Special case: For AR processes, BLP coe!cients coincide with Yule-Walker
estimates.
17
Time Series Analysis & Forecasting
Lecture 35
Best Linear Predictor (BLP)
Estimation of Missing values
1
Prediction of Second-Order Random Variables
Suppose that Y and Wn , . . . , W1 are any random variables with finite second
moments. Also,
E(Y ) = µY ,
E(Wi ) = µWi
,
cov(Y, Y ), cov(Y, Wi ), and cov(Wi , Wj )
are all known.
2
Suppose we want to predict Y based on Wn , . . . , W1 . Then, define
1 Wn
Wi
PWn ,...,W1 Y = ω0 + ω1 Wn + . . . + ωn W1
such that
E(Y → PWn ,...,W1 Y )2
is minimum with respect to ω0 , ω1 , . . . , ωn .
B
in
E
Y β β Wn
E Y Po β Wn
argmin
Po
pi
R
Pnw
3
Let g(ω) = E(Y → PW Y )2 .
BLP equations:
A
β
EfYWn i i
E
Pi
εg(ω)
=0
εω0
ELWHI I
εg(ω)
=0
εω
E wnti i.wnti.it .1 .
..
..
neqns
ELY po
εg(ω)
=0
BTMw
Munte
εωn
Mx
Pi Elwnie i writing
E
AI
Po
j
EY
β
PIE
Wnt i
β
Bn
piWn 1 i
ca
2ff
ElY Bo IPi
Wnt i Writing
o.LA
My BT.ME
Mw
Man
Ww
4
On solving:
ω0 = µY → ω → µ W
ε = !ω
1
T
cov
fcouCY.nl
cou Y Wi
P
wnte
i
wn i
j
corf4
5
BLP of Y
The best linear predictor of Y in terms of 1, Wn , . . . , W1 is given by:
PW Y = µY + ω → (W → µW )
where ω is a solution of ε = !ω.
6
Minimum mean squared prediction error
2
!
→
"2
E(Y → PW Y ) = E Y → µY → ω (W → µW )
!
"!
"→
→
→
= E Y → µY → ω (W → µW ) Y → µY → ω (W → µW )
= V(Y ) + ω → E(W → µW )(W → µW )→ ω → 2ω → E(Y → µY )(W → µW )
7
Tp 8
E(Y → PW Y )2 = V(Y ) + ω → !ω → 2ω → ϑ
= V(Y ) → ω → ϑ
8
The case of missing value!
Consider the stationary series
Xt = ϖXt↑1 + ϱt t = 0, ±1, . . .
where |ϖ| < 1 and ϱt ↑ W N (0, ς 2 ).
Suppose that we observe the series at times 1 and 3 and wish to use these observations
to find the linear combination of 1, X1 , and X3 that estimates X2 with minimum mean
squared error.
9
1
p
Cor
an
X2
T 1.8
Set Y = X2 and W = (X1 , X3 )→
x4
T
Cory
4
22
3
conf
13
1
E.it
alie
10
α
The best estimator of X2 is
PW X 2 =
22
ϖ
(X1 + X3 )
1 + ϖ2
with the mean squared error:
ς2
E(X2 → PW X2 ) =
1 + ϖ2
2
11
Time Series Analysis & Forecasting
Lecture 36
Best Linear Predictor (BLP)
Partial autocorrelation function
1
Partial autocorrelation function (PACF)
Definition
The PACF of a covariance stationary time series {Xt } is given by:
ω(1) = corr(X1 , X2 )
!
"
ω(k) = corr Xk+1 → PX2 ,...,Xk Xk+1 , X1 → PX2 ,...,Xk X1 ,
where PX2 ,...,Xk Xk+1 and PX2 ,...,Xk X1 are obtained using BLP technique.
2
Example 1: AR(1)
Xt = εXt→1 + ϑt
If
C 1
Gt WN 0,5
Let us first calculate ω(1):
11
corr
X1 X2
11
For an AR process
en
9ᵗʰ
3
Let us calculate ω(2):
PX
Px
Px
2
x
E
st
Pex
con
X1 Px
Px
corr
E X1
s
1
9
2
X
X
9 2
β
2
BeX2
is minimum
is
2.1
1
colts
minimm
p
92
4
Let us calculate ω(3):
Px
3
corr
Xy Px x
B X2 β
4
Xy
car
car
Pax
PY.BY
3
4 3
p xp
β
2 x
X
93
Ey
X
13 2
Pux
X1
P3X2
β4
3
0
5
Example 2: AR(2)
Xt = ε1 Xt→1 + ε2 Xt→2 + ϑt
Let us first calculate ω(1):
α
1
6
Let us calculate ω(2):
com
corr
Xs
Xs
P
Y
f X2
Px 1
X
X
p
2
0
7
Let us calculate ω(3):
com
Xu
corr
Xi
PpaxY
Ey
X
Pan
β X2
β2 3
0
8
Questions to think about!
• What can we say about the PACF of an AR(p) process?
• Compute PACF for an MA(1) process at lag k.
9