Bayesian Econometrics
Lecture Notes
Prof. Doron Avramov
The Jerusalem School of Business Administration
The Hebrew University of Jerusalem
Bayes Rule
ο΄ Let x and π¦ be two random variables
ο΄ Let π π₯ and π π¦ be the two marginal probability distribution functions of x and y
ο΄ Let π π₯ π¦ and π π¦ π₯ denote the corresponding conditional pdfs
ο΄ Let π π₯, π¦ denote the joint pdf of x and π¦
ο΄ It is known from the law of total probability that the joint pdf can be decomposed as
π π₯, π¦ = π π₯ π π¦ π₯ = π π¦ π π₯ π¦
ο΄ Therefore
π π¦ π π₯π¦
π π₯
= ππ π¦ π π₯ π¦
π π¦π₯ =
where c is the constant of integration (see next page)
ο΄ The Bayes Rule is described by the following proportion
π π¦π₯ ∝π π¦ π π₯π¦
2
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayes Rule
ο΄ Notice that the right hand side retains only factors related to y, thereby excluding π π₯
ο΄ π π₯ , termed the marginal likelihood function, is
π π₯ =
=
π π¦ π π₯ π¦ ππ¦
π π₯, π¦ ππ¦
as the conditional distribution π π¦ π₯ integrates to unity.
ο΄ The marginal likelihood π π₯ is an essential ingredient in computing an important quantity model posterior probability.
ο΄ Notice from the second equation above that the marginal likelihood obtains by integrating
out y from the joint density π π₯, π¦ .
ο΄ Similarly, if the joint distribution is π π₯, π¦, π§ and the pdf of interest is π π₯, π¦ one integrates
π π₯, π¦, π§ with respect to z.
3
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayes Rule
ο΄ The essence of Bayesian econometrics is the Bayes Rule.
ο΄ Ingredients of Bayesian econometrics are parameters underlying a given model, the sample
data, the prior density of the parameters, the likelihood function describing the data, and the
posterior distribution of the parameters.
ο΄ A predictive distribution could also be involved.
ο΄ In the Bayesian setup, parameters are stochastic while in the classical (non Bayesian)
approach parameters are unknown constants.
ο΄ Decision making is based on the posterior distribution of the parameters or the predictive
distribution of next period quantities as described below.
ο΄ On the basis of the Bayes rule, in what follows, y stands for unknown stochastic parameters,
x for the data, π π¦ π₯ for the posetior distribution, π π¦ for the prior, and π π₯ π¦ for the likelihood.
ο΄ The Bayes rule describes the relation between the prior, the likelihood, and the posterior, or
put differently it shows how prior beliefs are updated to produce posterior beliefs:
π π¦π₯ ∝π π¦ π π₯π¦
ο΄ Zellner (1971) is an excellent source of reference.
4
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayes Econometrics in Financial Economics
ο΄ You observe the returns on the market index over T months: π1 , … , ππ
ο΄ Let π
: π1 , … , ππ ’ denote the π × 1 vector of all return realizations
ο΄ Assume that ππ‘ ~π π, π0 2 for π‘ = 1, … , π
where
µ is a stochastic random variable denoting the mean return
π0 2 is the variance which, at this stage, is assumed to be a known constant
and returns are IID (independently and identically distributed) through time.
ο΄ By Bayes rule
π π π
, π0 2 ∝ π π π π
π, π0 2
where
π π π
, π0 2 is the posterior distribution of µ
π π is the prior distribution of µ
and π π
π, π0 2 is the joint likelihood of all return realizations.
5
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayes Econometrics: Likelihood
ο§ The likelihood function of a normally distributed return realization is given by
1
π ππ‘ π, π0 2 =
2ππ0
ππ₯π −
2
1
2π0 2
ππ‘ − π 2
ο΄ Since returns are assumed to be IID, the joint likelihood of all realized returns is
π π
π, π0
2
= 2ππ0
π
2 −2
1
ππ₯π − 2π 2
0
π
2
π‘=1 ππ‘ − π
ο΄ Notice:
ππ‘ − π 2 =
ππ‘ − π + π − π
2
= νπ 2 + π π − π 2
since the cross product is zero, and
ν=π−1
1
2
π =
π−1
1
6
π=π
ππ‘ − π 2
ππ‘
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Prior
ο΄ The prior is specified by the researcher based on economic theory, past experience, past data,
current similar data, etc. Often, the prior is diffuse or non-informative
ο΄ For the next illustration, it is assumed that π π ∝ π, that is, the prior is diffuse, noninformative, in that it apparently conveys no information on the parameters of interest.
ο΄ I emphasize “apparently” since innocent diffuse priors could exert substantial amount of
information about quantities of interest which are non-linear functions of the parameters.
ο΄ Informative priors with sound economic appeal are well perceived in financial economics.
ο΄ For instance, Kandel and Stambaugh (1996), who study asset allocation when stock returns
are predictable, entertain informative prior beliefs weighted against predictability. Pastor
and Stambaugh (1999) introduce prior beliefs about expected stock returns which consider
factor model restrictions. Avramov, Cederburg, and Kvasnakova (2017) study prior beliefs
about predictive regression parameters which are disciplined by consumption based asset
pricing models including habit formation, prospect theory, and long run risk.
ο΄ Computing posterior probabilities (as opposed to posterior densities) of competing models
(e.g., Avramov (2002)) necessitates the use of informative priors. Diffuse priors won’t fit.
7
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Posterior Distribution of Mean Return
ο΄ With diffuse prior and normal likelihood, the posterior is proportional to the likelihood
function:
1
2
2
2
π π π
, π0 ∝ ππ₯π −
νπ
+
π
π
−
π
2π0 2
π
2
∝ ππ₯π −
π
−
π
2π0 2
ο΄ The bottom relation follows since only factors related to µ are retained
ο΄ The posterior distribution of the mean return is given by
2
π|π
, π0 ~π π,
π0 2
π
ο΄ In classical econometrics:
π|π
, π0 2 ~π π, π0
2
π
ο΄ That is, in classical econometrics, the sample estimate of µ is stochastic while µ itself is an
unknown non-stochastic parameter.
8
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Informative Prior
ο΄ The prior on the mean return is often modeled as
π π ∝ ππ
−
1
2 ππ₯π
−
1
π − ππ 2
2
2ππ
where ππ and ππ are prior parameters to be specified by the researcher
ο΄ The posterior obtains by combining the prior and the likelihood:
π π π
, π0 2 ∝ π π π π
π, π0 2
∝ ππ₯π
∝ ππ₯π
π−ππ 2
π π−π 2
− 2π 2 + 2π 2
π
0
1 π−π 2
− 2 π2
ο΄ The bottom relation obtains by completing the square on µ
ο΄ Notice, in particular,
9
π2
π 2 π2
+
π = 2
ππ 2 π0 2
π
1
π
1
+
=
ππ 2 π0 2
π2
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Posterior Mean
ο΄ Hence, the posterior variance of the mean is
2
π =
1
1
+
π0 2
ππ 2
−1
π
= (prior precision + likelihood precision) −1
ο΄ Similarly, the posterior mean of π is
π0
ππ
+
ππ 2 π0 2
= π€1 π0 + π€2 π
π = π2
where
π€1 =
1
ππ 2
1
1
+
ππ 2 π0 2
prior precision
=
prior precision + likelihood precision
π
π€2 = 1 − π€1
10
ο΄ Intuitively, the posterior mean of µ is the weighted average of the prior mean and the sample
mean with weights depending on prior and likelihood precisions, respectively.
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
What if σ is unknown? – The case of Diffuse Prior
ο΄ Bayes: π π, π π
∝ π π, π π π
π, π
ο΄ The non-informative prior is typically modeled as
π π, π ∝ π π π π
π π ∝π
π π ∝ π −1
ο΄ Thus, the joint posterior of µ and σ is
π π, π π
∝ π
− π+1
1
ππ₯π − 2 νπ 2 + π π − π
2π
2
ο΄ The conditional distribution of the mean follows straightforwardly
2
π
π π π, π
is π π,
π
ο΄ More challenging is to uncover the marginal distributions, which are obtained as
11
π ππ
=
π π, π π
ππ
π ππ
=
π π, π π
ππ
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Solving the Integrals: Posterior of π
ο΄ Let α = νπ 2 + π π − π 2
ο΄ Then,
∞
π − π+1 ππ₯π −
π ππ
∝
π=0
ο΄ We do a change of variable
πΌ
ππ
2π 2
πΌ
π₯= 2
ππ 2π −11 1 −11
= −2 2 πΌ 2 π₯ 2
ππ₯
π+1
−
πΌ
2
π −π+1 =
2π₯
π−2
2
π
∞
−1
2
π ππ
∝2 πΌ
π₯
ππ₯π −π₯ ππ₯
π₯=0
π
∞
π
−1
2
π₯
ππ₯π
−π₯
ππ₯
=
Γ
π₯=0
2
ο΄ Then
Notice
ο΄ Therefore,
π
2
π −π
π ππ
πΌ 2
2
ν+1
2
2 − 2
∝ νπ + π π − π
π−2
∝2 2 Γ
π−π
ο΄ We get π‘ = π
12
−
~π‘ ν , corresponding to the Student t distribution with ν degrees of freedom.
π
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Marginal Posterior of σ
ο΄ The posterior on π
π ππ
∝
∝π
ο΄ Let π§ =
− π+1
ππ₯π
νπ 2
− 2π2
2
ππ
π
ππ₯π − 2π2 π − π 2 ππ
π π−π
, then
π
ππ§
=
ππ
ο΄
1
π − π+1 ππ₯π − 2π2 νπ 2 + π π − π
π ππ
∝π
1
ππ
−π
ππ₯π
νπ 2
− 2π2
2
∝ π −π ππ₯π −
∝π
− ν+1
ππ₯π
ππ₯π − π§
2
2
ππ§
νπ
2π 2
νπ 2
− 2π2
which corresponds to the inverted gamma distribution with ν degrees of freedom and
parameter s
ο΄ The explicit form (with constant of integration) of the inverted gamma is given by
13
2
π π ν, π =
ν
Γ 2
2
νπ
2
ν
2
2
νπ
π − ν+1 ππ₯π − 2
2π
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Multiple Regression Model
ο΄ The regression model is given by
π¦ = ππ½ + π’
where
y is a π × 1 vector of the dependent variables
X is a π × π matrix with the first column being a π × 1 vector of ones
β is an M × 1 vector containing the intercept and M-1 slope coefficients
and u is a π × 1 vector of residuals.
ο΄ We assume that π’π‘ ~π 0, π 2 ∀ π‘ = 1, … , π and IID through time
ο΄ The likelihood function is
1
π π¦ π, π½, π ∝ π ππ₯π − 2 π¦ − ππ½ ′ π¦ − ππ½
2π
1
′
−π
∝ π ππ₯π − 2 νπ 2 + π½ − π½ π ′ π π½ − π½
2π
−π
14
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Multiple Regression Model
where
ν=π−π
π½ = π ′ π −1 π ′ π¦
π 2 =
1
ν
π¦ − ππ½
′
π¦ − ππ½
ο΄ It follows since
π¦ − ππ½ ′ π¦ − ππ½ = π¦ − ππ½ − π π½ − π½
= π¦ − ππ½
′
′
π¦ − ππ½ − π π½ − π½
′
π¦ − ππ½ + π½ − π½ π ′ π π½ − π½
while the cross product is zero
15
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Assuming Diffuse Prior
ο΄ The prior is modeled as
π π½, π ∝
1
π
ο΄ Then the joint posterior of β and σ is
π π½, π π¦, π ∝ π
− π+1
ππ₯π
1
− 2π2
2
′
νπ + π½ − π½ π ′ π π½ − π½
ο΄ The conditional posterior of β is
π π½ π, π¦, π ∝ ππ₯π
1
− 2π2
′
π½ − π½ π′π π½ − π½
which obeys the multivariable normal distribution
π π½, π ′ π −1 π 2
16
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Assuming Diffuse Prior
ο΄ What about the marginal posterior for β ?
π π½ π¦, π =
π π½, π π¦, π ππ
2
′
′
∝ νπ + π½ − π½ π π π½ − π½
−π 2
which pertains to the multivariate student t with mean π½ and T-M degrees of freedom
ο΄ What about the marginal posterior for π?
π π π¦, π =
∝π
π π½, π π¦, π ππ½
− ν+1
νπ 2
ππ₯π − 2
2π
which stands for the inverted gamma with T-M degrees of freedom and parameter s
ο§ You can simulate the distribution of β in two steps without solving analytically the integral,
drawing first π from its inverted gamma distribution and then drawing from the conditional
of β given π which is normal as shown earlier. This mechanism generates draws from the
Student t distribution.
17
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Updating/Learning
ο΄ Suppose the initial sample consists of π1 observations of π1 and π¦1 .
ο΄ Suppose further that the posterior distribution of π½, π based on those observations is given
by:
1
π π½, π π¦1 , π1 ∝ π −(π1+1) ππ₯π − 2 π¦1 − π1 π½ ′ π¦1 − π1 π½
2π
1
−(π1 +1)
∝π
ππ₯π − 2 ν1 π 1 2 + π½ − π½1 ′π1 ′π1 π½ − π½1
2π
where
ν1 = π1 − π
π½1 = π1 ′π1 −1 π1 π¦1
′
ν1 π 1 2 = π¦1 − π1 π½1 π¦1 − π1 π½1
ο΄ You now observe one additional sample π2 and π¦2 of length π2 observations
ο΄ The likelihood based on that sample is
1
π π¦2 , π2 π½, π ∝ π −π2 ππ₯π − 2π2 π¦2 − π2 π½ ′ π¦2 − π2 π½
18
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Updating/Learning
ο΄ Combining the posterior based on the first sample (which becomes the prior for the second
sample) and the likelihood based on the second sample yields:
1
− π1 +π2 +1
π π½, π π¦1 , π¦2 , π1 , π2 ∝ π
ππ₯π − 2 π¦1 − π1 π½ ′ π¦1 − π1 π½ + π¦2 − π2 π½ ′ π¦2 − π2 π½
2π
1
′
− π1 +π2 +1
∝π
ππ₯π − 2 νπ 2 + π½ − π½ Ο» π½ − π½
2π
where
Ο» = π1′ π1 + π2′ π2
π½ = Ο»−1 π1′ π¦1 + π2′ π¦2
′
′
νπ 2 = π¦1 − π1 π½ π¦1 − π1 π½ + π¦2 − π2 π½ π¦2 − π2 π½
ν = π1 + π2 − π
ο΄ Then the posterior distributions for β and σ follow using steps outlined earlier
ο΄ With more observations realized you follow the same updating procedure
ο΄ Notice that the same posterior would have been obtained starting with diffuse priors and
then observing the two samples consecutively Y=[π¦1 ′, π¦2 ′]′ and π = [π1 ′, π2′ ]′.
19
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Example: Predictive Regressions
ο΄ In finance and economics you often use predictive regressions of the form
π¦π‘+1 = π + π ′ π§π‘ + π’π‘+1
π’π‘+1 ~π 0, π 2 ∀ π‘ = 1, … , π − 1 and IID
where π¦π‘+1 is an economic quantity of interest, be it stock or bond return, inflation, interest
rate, exchange rate, and π§π‘ is a collection of π − 1 predictive variables, e.g., the term spread
ο΄ At this stage, the initial observation of the predictors, π§0 , is assumed to be non stochastic.
ο΄ Stambaugh (1999) considers stochastic π§0 . Then, some complexities emerge as shown later.
ο΄ The predictive regression can be written more compactly as
π¦π‘+1 = π₯π‘ ′π½ + π’π‘+1
where
π₯π‘ = 1, π§π‘ ′
π½ = π, π ′
20
ο΄ In a matrix form, comprising all time-series observations, the normal regression model
obtains
π¦ = ππ½ + π’
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Predictive Distribution
ο΄ You are interested to uncover the predictive distribution of the unobserved π¦π+1
ο΄ Let Π€ denote the observed data and let θ denote the set of parameters β and π 2
ο΄ The predictive distribution is:
π π¦π+1 Π€ =
θ
π π¦π+1 Π€,θ π θ Π€ πθ
where
π π¦π+1 Π€,θ is the conditional or classical predictive distribution
π θ Π€ is the joint posterior of β and π 2
ο΄ Notice that the predictive distribution integrates out β and π from the joint distribution
π π¦π+1 , π½, π Π€
since
π π¦π+1 , π½, π Π€ = π π¦π+1 Π€,θ π θ Π€
21
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Predictive Distribution
ο΄ The conditional distribution of the next period realization is
1
π π¦π+1 θ,Π€ ∝ π −1 ππ₯π − 2 π¦π+1 − π₯π ′π½ 2
2π
1
ο΄ Thus π π¦π+1 , π½, π Π€ is proportional to π − π+2 ππ₯π − 2π2 [ π¦ − ππ½ ′ π¦ − ππ½ + π¦π+1 − π₯π ′π½ 2 ]
ο΄ On integrating π π¦π+1 , π½, π Π€ with respect to π we obtain
π π¦π+1 , π½ Π€ ∝
π¦ − ππ½
′
π¦ − ππ½ + π¦π+1 − π₯π ′π½
2
− π+1
2
ο΄ Now we have to integrate with respect to the π elements of π½
ο΄ On completing the square on π½ we get
π¦ − ππ½ ′ π¦ − ππ½ + π¦π+1 − π₯π ′π½ 2
= π¦ ′ π¦ + π¦π+1 2 + π½ ′ Ο»π½ − 2π½′ π ′ π¦ + π₯π ′π¦π+1
= π¦ ′ π¦ + π¦π+1 2 − π¦ ′ X + π¦π+1 ′π₯π Ο»−1 π ′ π¦ + π₯π ′π¦π+1
+ π½ − Ο»−1 π ′ π¦ + π₯π ′π¦π+1 ′Ο» π½ − Ο»−1 π ′ π¦ + π₯π ′π¦π+1
where
Ο» = π ′ X + π₯π ′π₯π
22
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Predictive Distribution
ο΄ Integrating with respect to β yields
′
2
′
−1
π π¦π+1 Π€ ∝ π¦ π¦ + π¦π+1 − π¦ π + π¦π+1 ′π₯π Ο»
′
π π¦ + π₯π ′π¦π+1
−
ν+1
2
where
ν=π−π
ο΄ With some further algebra it can be shown that the predictive distribution is
π π¦π+1 Π€ ∝ ν + π¦π+1 − π₯π ′π½ π» π¦π+1 − π₯π ′π½
−
ν+1
2
where
1
π» = π 2 1 − π₯π′ Ο»−1 π₯π
νπ 2 = π¦ − ππ½ ′ π¦ − ππ½
π½ = π ′ π −1 π ′ π¦
23
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Predictive Moments
ο΄ The first and second predictive moments, based on the t-distribution, are
ππ¦ = πΈ π¦π+1 Π€ = π₯π ′π½
πΈ π¦π+1 − ππ¦
2
ν
= ν−2 π» −1
νπ 2
=
ν−2
1 − π₯π′ Ο»−1 π₯π −1
νπ 2
= ν−2
1 + π₯π′ π ′ π π₯π
o With diffuse prior, the predictive mean coincides with the classical (non Bayesian) mean.
o The predictive variance is slightly higher due to estimation risk.
o Kandel and Stambaugh (JF 1996) provide more economic intuition about the predictive density
o The estimation risk effect on the predictive variance is analytically derived by Avramov and
Chordia (JFE 2006) in a multi-asset (asset pricing) context.
o Later, we will use the predictive distribution to recover asset allocation under estimation risk
and even under model uncertainty considering informative priors.
24
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Multivariate Regression Models
ο΄ Consider the multivariate form (N dependent variables) of the predictive regression
π
= ππ΅ + π
where R and U are both a π × π matrix , X is a π × M matrix, B is an M× π matrix
π£ππ(π) ∼ π(0, Σ ⊗ πΌπ )
and where vec denotes the vectorization operator and ⊗ is the kronecker product
ο΄ The priors for π΅ and Σ are assumed to be the normal inverted Wishart (conjugate priors)
π(π|Σ) ∝ |Σ|
−
1
2
π(Σ) ∝ |Σ|
25
1
exp − 2 (π − π0 )′ [Σ −1 ⊗ Ψ0 ](π − π0 )
π +π+1
2
− 0
1
exp − 2 tr[π0 Σ −1 ]
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Multivariate regression
where
π = π£ππ(π΅)
and π0 , Ψ0 , and π0 are prior parameters to be specified by the researcher.
ο΄ The likelihood function of normally distributed data constituting the actual sample obeys the
form
π(π
|π΅, Σ, π) ∝ |Σ|
−
π
2
1
exp − 2 tr (π
− ππ΅)′ (π
− ππ΅) Σ −1
where tr stands for the trace operator. This can be rewritten in a more convenient form as
π(π
|π΅, Σ, π) ∝ |Σ|
−
π
2
1
exp − 2 tr π + (π΅ − π΅)′ π ′ π(π΅ − π΅) Σ −1
where
π = (π
− ππ΅)′ (π
− ππ΅)
π΅ = π ′ π −1 π ′ π
26
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Multivariate regression
ο΄ An equivalent representation for the likelihood function is given by
π(π
|π, Σ, π) ∝ |Σ|
−
π
2
1
2
exp − (π − π)′ [Σ −1 ⊗ (π ′ π)](π − π)
1
× exp − 2 tr[πΣ −1 ]
where
π = vec(π΅)
ο΄ Combining the likelihood with the prior and completing the square on π yield
1
1
−
π(π|Σ, π
, π) ∝ |Σ| 2 exp − (π − π)′ [Σ −1 ⊗ Ψ](π − π)
2
π(Σ|π
, π) ∝ |Σ|
27
−
π+π+1
2
exp
1
− tr πΣ −1
2
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Multivariate regression
where
Ψ = Ψ0 + π ′ π
π = vec π΅
π΅ = Ψ −1 π ′ ππ΅ + Ψ −1 Ψ0 π΅0
π = π + π0 + π΅ ′ π ′ ππ΅ + π΅0 ′ Ψ0 π΅0 − π΅′ Ψπ΅
π = π0 + π
ο΄ So the posterior for π΅ is normal and for Σ is inverted Wishart.
ο΄ Again,, that is the conjugate prior idea - the prior and posterior have the same distributions but
with different parameters.
ο΄ Not surprisingly, π΅ is a weighted average of π΅0 and π΅:
π΅ = ππ΅0 + (πΌ − π)π΅
where π = πΌ − Ψ−1 π ′ π. Notice, the weights are represented by matrices.
28
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
What if the posterior does not obey a well-known
expression?
ο΄ Thus far, the posterior densities can readily be identified.
ο΄ However, what if the posterior does not obey a well known expression?
ο΄ Markov Chain Monte Carlo (MCMC) methods can be employed to simulate from the posterior.
ο΄ The basic intuition behind MCMC is straightforward.
ο΄ Suppose the distribution is π π₯ which is unrecognized.
ο΄ The MCMC idea is to define a Markov chain over possible values of x π₯0 , π₯1 , π₯2 , … such that as
π → ∞, we can guarantee that π₯π ~π π₯ , that is, that we have a draw from the posterior.
ο΄ As the number of draws (each draw pertains to a distinct chain) gets larger you can simulate
the posterior density.
ο΄ The simulation gets more precise with increasing number of draws.
ο΄ There are various ways to set up such Markov chains
ο΄ Here, we cover two MCMC methods: the Gibbs Sampling and the Metropolis Hastings.
29
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ The Gibbs Sampling analysis is based on “Measuring the Pricing Error of the Arbitrage
Pricing Theory” by Geweke and Zhou (RFS 1996).
ο΄ This paper advocates a Bayesian method in which to test the APT of Ross (1976).
ο΄ Both APT and ICAPM motivate multiple factors – extending the CAPM.
ο΄ While APT motivates statistical based factors, as shown below, the ICAPM motivates
economic factors related to the marginal utility of the investor – such as consumption growth.
ο΄ The basic APT model assumes that returns on π risky portfolios are related to πΎ pervasive
unknown factors (K<N).
ο΄ The relation is described by the πΎ factor model
ππ‘ = π + π½ππ‘ + ππ‘
where ππ‘ denotes returns (not excess returns) on π assets and ππ‘ is a set of πΎ factor
innovations (factors are not pre-specified, rather, they are latent).
30
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ Specifically,
πΈ ππ‘ = 0
πΈ{ππ‘ π ′ π‘ } = πΌπΎ
πΈ{ππ‘ |ππ‘ } = 0
πΈ{ππ‘ π ′ π‘ |ππ‘ } = Σ = ππππ(π12 , … , ππ2 )
π½ = [π½1 , … , π½πΎ ]
ο΄ Moreover, under exact APT, the π vector satisfies the restriction
π = π0 + π½1 π1 + β― , +π½πΎ ππΎ
ο΄ Notice that π0 is the component of expected return unrelated to factor exposures.
ο΄ The original APT model is about an approximated relation.
ο΄ An exact version is derived by Huberman (1982) among others.
ο΄ The objective throughout is to explore a measure that summarizes the deviation from exact
pricing.
31
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ That measure is denoted by π 2 and is given by
1 ′
∗
2
π¬ = π [πΌπ − π½ ∗ (π½ ′ π½ ∗ )−1 π½∗ ′ ]π
π
where
π½ ∗ = [πΌπ , π½]
ο΄ Recovering the sampling distribution of π 2 is hopeless.
ο΄ Notice that one cannot even recover an analytic expression for the posterior density of model
parameters π(π©|π
). π π© π
, πΉ is something known – but this is not the posterior.
ο΄ However, using Gibbs sampling, we can simulate the posterior distribution of π 2 as well as
simulate the posterior density of all parameters and latent factors.
ο΄ In what follows, we assume that observed returns and latent factors are jointly normally
distributed:
ππ‘
∼π
ππ‘
32
πΌ
0
, πΎ
π
π½
π½′
π½π½ ′ + Σ
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ Here are some additional notations:
ο΄ Data: π
= [π1 ′, … , πT ′]′
ο΄ Parameters: Θ = [π ′ , π£ππ(π½)′ , π£ππβ(Σ)′ ]′ where vech denotes the distinct elements of the matrix
ο΄ Latent factors: π = [π1 ′, … , ππ ′]′
ο΄ To evaluate the pricing error we need to simulate draws from the posterior distribution
π(π©|π
).
ο΄ We draw from the joint posterior in a slightly different manner than that suggested in the
Geweke-Zhou paper.
ο΄ First, we employ a multivariate regression setting. Moreover, the well-known identification
(of the factors) problem is not accounted for to simplify the analysis.
ο΄ The prior on the diagonal covariance matrix is assumed to be non-informative
π0 (Θ) ∝ |Σ|
33
−
1
2 = (π1 … ππ )−1
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ Re-expressing the arbitrage pricing equation, we obtain:
π ′ π‘ = πΉπ‘ ′ π΅′ + ππ‘ ′
where
πΉπ‘ ′ = [1, ππ‘ ′ ]
π΅ = [π, π½]
ο΄ Rewriting the system in a matrix notation, we get
π
= πΉπ΅′ + πΈ
ο΄ Why do we need to use the Gibbs sampling technique?
ο΄ Because the likelihood function π(π
|Θ) (and the posterior density) cannot be expressed
analytically.
34
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ However, π(π
|Θ, πΉ) does obey an analytical form:
π
1
−2
π(π
|Θ, πΉ) ∝ |Σ| exp [ − tr [π
− πΉπ΅′ ]′ [π
− πΉπ΅′ ]Σ −1
2
ο΄ Therefore, we can compute the full conditional posterior densities:
π(π΅|Σ, πΉ, π
)
π(Σ|π΅, πΉ, π
)
π(πΉ|π΅, Σ, π
)
ο΄ The Gibbs sampling chain is formed as follows:
1.
Specify starting values π΅(0) , Σ (0) , and πΉ (0) and set π = 1.
2.
Draw from the full conditional distributions:
ο΄ Draw π΅(π) from π(π΅|Σ (π−1) , πΉ (π−1) , π
)
ο΄ Draw Σ (π) from π(Σ|π΅(π) , πΉ (π−1) , π
)
ο΄ Draw πΉ (π) from π(F|π΅(π) , Σ (π) , π
)
3.
35
Set π = π + 1 and go to step 2.
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ After π iterations the sample π΅(π) , Σ (π) , πΉ (π) is obtained.
ο΄ Under mild regularity conditions (see, for example, Tierney ,1994), (π΅(π) , Σ (π) , πΉ (π) )
converges in distribution to the relevant marginal and joint distributions:
π(π΅(π) |π
) → π(π΅|π
)
π(Σ (π) |π
) → π(Σ|π
)
π(πΉ (π) |π
) → π(πΉ|π
)
π(π΅(π) , Σ (π) , πΉ (π) |π
) → π(π΅, Σ, πΉ|π
)
ο΄ For π (burn-in draws) large enough, the πΊ values
(π΅(π) , Σ (π) , πΉ (π) )π+πΊ
π=π+1
are a sample from the joint posterior.
ο΄ What are the full conditional posterior densities?
ο΄ Note:
1
π(π΅|Σ, πΉ, π
) ∝ exp − tr[π
− πΉπ΅′ ]′ [π
− πΉπ΅′ ]Σ −1
2
36
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
ο΄ Let π = vec(π΅′ ). Then
1
π(π|Σ, πΉ, π
) ∝ exp − [π − π]′ Σ −1 ⊗ (πΉ ′ πΉ) [π − π]
2
where
π = vec[(πΉ ′ πΉ)−1 πΉ ′ π
]
ο΄ Therefore,
π|Σ, πΉ, π
∼ π π, Σ ⊗ (πΉ ′ πΉ)−1
ο΄ Also note:
−(π+1)
π(ππ |π΅, πΉ, π
) ∝ ππ
exp
πππ2
− 2
2ππ
where πππ2 is the i-th diagonal element of the π × π matrix
[π
− πΉπ΅′ ]′ [π
− πΉπ΅′ ]
37
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Gibbs Sampling
suggesting that:
πππ2
2
2 ∼ π (π)
ππ
ο΄ Finally,
ππ‘ |π, π½, Σ, ππ‘ ∼ π(ππ‘ , π»π‘ )
where
ππ‘ = π½ ′ (π½π½ ′ + Σ)−1 (ππ‘ − π)
π»π‘ = πΌπΎ − π½ ′ (π½π½ ′ + Σ)−1 π½
38
•
Here, we are basically done with the GS implementation.
•
Having all the essential draws form the joint posterior at hand, we can analyze the simulated
distribution of the pricing errors and make a call about model’s pricing abilities.
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Metropolis Hastings (MH)
ο΄ Indeed, Gibbs Sampling is intuitive, easy to implement, and convergence to the true posterior
density is accomplished relatively fast and with mild regularity conditions.
ο΄ Often, however, integrable expressions for the full conditional densities are infeasible.
ο΄ Then it is essential to resort to more complex methods in which draws from the posterior
distributions could be highly correlated and convergence could be rather slow. Still, such
methods are useful.
ο΄ One example is the Metropolis Hastings (MH) algorithm, a MCMC procedure introduced by
Metropolis et al (1953) and later generalized by Hastings (1970).
ο΄ The basic idea in MH is to make draws from a candidate distribution which seems to be
related to the target (unknown) distribution.
ο΄ The candidate draw from the posterior is accepted with some probability – the Metropolis
rule. Otherwise, it is rejected and the previous draw is retained. As in the Gibbs Sampling,
the Markov Chain starts with some initial value set by the researcher.
ο΄ The Gibbs Sampling is a special case of MH in which all draws are accepted with probability
one.
39
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Metropolis Hastings (MH)
ο΄ I will display two applications of the MH method in the context of financial economics.
ο΄ The first application goes to the seminal work of Jacquier, Polson, and Rossi (1994) on
estimating a stochastic volatility (SVOL) model. Coming up on the next page.
ο΄ The second application is based on Stambaugh (1999) who analyzes predictive regressions
when the first observation of the predictive variable is stochastic.
ο΄ Mostly, analyses of predictive regressions are conducted based on the assumption that the
first observation is fixed non-stochastic.
ο΄ While analytically tractable this assumption does not seem to hold true.
ο΄ Relaxing that assumption entertains several complexities and the need to use MH to draw
from the joint posterior distribution of the predictive regression parameters.
ο΄ I will discuss that application in the section on asset allocation when stock returns are
predictable.
40
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stochastic Volatility (SVOL)
ο΄ The SVOL model is given by
ππ βπ‘
π¦π‘ = βπ‘ π’π‘
= πΌ + πΏππ βπ‘−1 + ππ£ π£π‘
π’π‘
π£π‘ ~π 0, πΌ2
ο΄ Notice that volatility varies through time rather than being constant. While in ARCH,
GARCH, EGARCH models there is no stochastic innovation, here volatility is stochastic.
ο΄ Now let
β′ = β1 … β π
π½ ′ = πΌ, πΏ
π€ ′ = πΌ, πΏ, ππ£
π¦ ′ = π¦1 … π¦π
ο΄ The posterior of w given values of h is available from the standard regression model described
earlier: π½ has the multivariable normal and ππ£ has the inverted gamma distribution.
ο΄ However, drawing from β|π€, π¦ requires more efforts.
41
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stochastic Volatility (SVOL)
ο΄ We cannot really draw h at once, rather, we have to break down the joint posterior of the
entire h vector by considering the series of uni-variate conditional densities.
π βπ‘ βπ‘−1 , βπ‘+1 , π€, π¦π‘ for π‘ = 1, … , π
ο΄ If it were possible to draw directly these uni-variate densities, the algorithm would reduce to
a Gibbs sampler in which we would draw successively from π π€ β, π¦ and then each of the T
univariate conditionals in turn to form one step in the Markov chain.
ο΄ The uni-variate conditional densities, however, exhibit an unusual form:
π βπ‘ βπ‘−1 , βπ‘+1 , π€, π¦π‘
∝ π π¦π‘ βπ‘ π βπ‘ βπ‘−1 π βπ‘+1 βπ‘
1
1 π¦π‘ 2 1
−
∝ βπ‘ 2 ππ₯π −
ππ₯π − ln(βπ‘ ) − ππ‘ 2 2π 2
2 β π‘ βπ‘
where
42
ππ‘ = πΌ 1 − πΏ + πΏ ln(βπ‘+1 ) + ππ(βπ‘−1 )
2
π
π£
π2 =
1 + πΏ2
1 + πΏ2
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stochastic Volatility (SVOL)
ο΄ The result follows by combining two likelihood normal terms and completing the square on
ππ(βπ‘ ).
ο΄ Notice, the density is not of a standard form. It is proportional to:
π¦π‘ 2
1
ππ₯π −
− 2 ln(βπ‘ ) − π 2
2βπ‘ 2π
βπ‘ 3/2
=
π¦π‘ 2
1
ππ₯π −
− 2 ππ(βπ‘ ) − π − π 2 2
2βπ‘ 2π
2
βπ‘
ο΄ A good proposal here can be the lognormal density given by
1
1
ππ₯π − 2 ππ(βπ‘ ) − π + π 2 2 2
2π
2ππβπ‘
43
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stochastic Volatility (SVOL)
ο΄ From here we compute the ratio of the target to the proposal as:
π¦π‘ 2
1
2
2
ππ₯π
−
−
ππ(β
)
−
π
−
π
2
βπ‘
π‘
π‘πππππ‘
2βπ‘
2π 2
∝
1
ππππππ ππ
ππ₯π − 2 ln(βπ‘ ) − π − π 2 2 2 βπ‘
2π
π¦π‘ 2
∝ ππ₯π −
2βπ‘
ο΄ For the MH algorithm, the relevant ratio is
π¦π‘ 2
π¦π‘ 2
ππ₯π
−
2βπ‘ 2βπ‘ ∗
where βπ‘ ∗ is the new proposed draw and βπ‘ is the current state (or previously accepted
draw).
44
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Portfolio Analysis
ο΄ We next present topics on Bayesian portfolio analysis based upon a review paper of Avramov
and Zhou (2010) that came up in the Annual Review of Financial Economics
ο΄ We first study asset allocation when stock returns are assumed to be IID
ο΄ We then incorporate potential return predictability based on macro economy variables.
ο΄ What are the benefits of using the Bayesian approach?
ο΄ There are at least three important benefits including (i) the ability to account for estimation
risk and model uncertainty, (ii) the feasibility of powerful and tractable simulation methods,
and (iii) the ability to elicit economically meaningful prior beliefs about the distribution of
future returns.
45
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ We start with the mean variance framework
ο΄ Assume there are π + 1 assets, one of which is riskless and others are risky.
ο΄ Denote by πππ‘ and ππ‘ the rates of returns on the riskless asset and the risky assets at time π‘,
respectively.
ο΄ Then
π
π‘ ≡ ππ‘ − πππ‘ 1π
are excess returns on the π risky assets, where 1π is an π × 1 vector of ones.
ο΄ Assume that the joint distribution of π
π‘ is IID over time, with mean π and covariance matrix
π.
ο΄ In the static mean-variance framework an investor at time π chooses his/her portfolio weights
π€, so as to maximize the quadratic objective function
πΎ
πΎ ′
′
π(π€) = πΈ[π
π ] − Var[π
π ] = π€ π − π€ ππ€
2
2
46
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
where π
π = π€ ′ π
π+1 is the future uncertain portfolio return at time π + 1 and πΎ is the
coefficient of relative risk aversion.
ο΄ When both π and π are assumed to be known, the optimal portfolio weights are
1
π€ ∗ = π −1 π
πΎ
and the maximized expected utility is
2
1
π
π(π€ ∗ ) =
π′ π −1 π =
2πΎ
2πΎ
where π 2 = π′ π −1 π is the squared Sharpe ratio of the ex ante tangency portfolio of the risky
assets.
ο΄ This is the well known mean-variance theory pioneered by Markowitz (1952).
ο΄ In practice, the problem is that π€ ∗ is not computable because π and π are unknown. As a
result, the above mean-variance theory is usually applied in two steps.
ο΄ In the first step, the mean and covariance matrix of the asset returns are estimated based on
the observed data.
47
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Given a sample size of π, the standard maximum likelihood estimators are
π=
π=
1
π
π
1
π
π
π
π‘
π‘=1
( π
π‘ − π)(π
π‘ − π)′
π‘=1
ο΄ Then, in the second step, these sample estimates are treated as if they were the true
parameters, and are simply plugged in to compute the estimated optimal portfolio weights,
1 −1
ML
π€ = π π
πΎ
ο΄ The two-step procedure gives rise to a parameter uncertainty problem because it is the
estimated parameters, not the true ones, that are used to compute the optimal portfolio
weights.
ο΄ Consequently, the utility associated with the plug-in portfolio weights can be substantially
different from π(π€ ∗ ).
48
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Denote by π the vector of all the parameters (both π and π).
ο΄ Mathematically, the two-step procedure maximizes the expected utility conditional on the
estimated parameters, denoted by π, being equal to the true ones,
max[ π(π€) | π = π]
π€
and the uncertainty or estimation errors are ignored.
ο΄ To account for estimation risk, let us specify the posterior distribution of the parameters as
π(π, π|π·π ) = π(π | π, Φ π ) × π(π | Φ π )
with
1
π(π | π, Φ π ) ∝ |π|
ππ₯π{− π‘π[π(π − π)(π − π)′ π −1 ]}
2
π
1
π(π) ∝ |π|−2 ππ₯π{− π‘π π −1 (ππ)}
2
−1/2
where π = π + π.
49
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ The predictive distribution is:
π π
π+1 Φ π ∝ V + π
π+1 − π π
π+1 − π ′ / π + 1
−π 2
which is a multivariate π‘-distribution with π − π degrees of freedom.
ο΄ While the problem of estimation error is recognized by Markowitz (1952), it is only in the 70s
that this problem receives serious attention.
ο΄ Winkler (1973) and Winkle and Barry (1975) are earlier examples of Bayesian studies on
portfolio choice.
ο΄ Brown (1976, 1978) and Klein and Bawa (1976) lay out independently and clearly the
Bayesian predictive density approach, especially Brown (1976) who explains thoroughly the
estimation error problem and the associated Bayesian approach.
50
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Later, Bawa, Brown, and Klein (1979) provide an excellent review of the early literature.
ο΄ Under the diffuse prior, it is known that the Bayesian optimal portfolio weights are
1 π − π − 2 −1
π€ Bayes =
π π
πΎ
π+1
ο΄ In contrast with the classical weights π€ ML , the Bayesian portfolio is proportion to π −1 π, but
the proportional coefficient is (π − π − 2)/(π + 1) instead of 1.
ο΄ The coefficient can be substantially smaller when π is large relative to π.
ο΄ Intuitively, the assets are riskier in the Bayesian framework since parameter uncertainty is
an additional source of risk and this risk is now factored into the portfolio decision.
ο΄ As a result, the overall position in the risky assets are generally smaller than before.
51
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ However, in the classical framework, π€ ML is a biased estimator of the true weights since,
under the normality assumption,
πΈ π€ ML =
π−π−2 ∗
π€ ≠ π€∗
π
ο΄ Let
π
−1
π − N − 2 −1
=
π
π
ο΄ Then π −1 is an unbiased estimator of π −1 .
ο΄ The unbiased estimator of π€ ∗ is
1 π − π − 2 −1
π€=
π π
πΎ
π
which is a scalar adjustment of π€ ML .
52
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ The unbiased classical weights differ from their Bayesian counterparts by a scalar π/(π + 1).
ο΄ The difference is independent of π, and is negligible for all practical sample sizes π.
ο΄ Hence, parameter uncertainty makes little difference between Bayesian and classical
approaches if the diffuse prior is used.
ο΄ Therefore, to provide new insights, it is important for a Bayesian to use informative priors,
which is a decisive advantage of the Bayesian approach that can incorporate useful
information easily into portfolio analysis.
53
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Conjugate Prior
ο΄ The conjugate prior, which retains the same class of distributions, is a natural and common
information prior on any problem.
ο΄ This prior in our context assumes a normal prior for the mean and inverted Wishart prior for
π,
1
π | π ∼ π(π0 , π)
π
π ∼ πΌπ(π0 , π0 )
where π0 is the prior mean, and π is a prior parameter reflecting the prior precision of π0 , and
π0 is a similar parameter on π.
ο΄ Under this prior, the posterior distribution of π and π are of the same form as the case for the
diffuse prior, except that now the posterior mean of π is given by
π
π
π=
π0 +
π
π+π
π+π
ο΄ This says that the posterior mean is simply a weighted average of the prior and sample.
54
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Conjugate Prior
ο΄ Similarly, π0 can be updated by
π+1
ππ
π=
π0 + ππ +
(π0 − π)(π0 − π)′
π(π0 + π − 1)
π+π
a weighted average of the prior variance, sample variance, and deviations of π from π0 .
ο΄ Frost and Savarino (1986) provide an interesting case of the conjugate prior by assuming the
assets are with identical means, variances, and patterned co-variances, a priori.
ο΄ They find that such a prior improves performance.
ο΄ This prior is related the well known 1/π rule that invests equally across π assets.
55
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Hyper-Parameter Prior
ο΄ By introducing hyper-parameters π and π, Jorion (1986) uses the following prior on π,
1
π0 (π | π, π) ∝ |π|−1 exp{− (π − π1π )′ (ππ)−1 (π − π1π )}
2
in which π and π govern the prior distribution of π.
ο΄ Using diffuse priors on both π and π, and integrating them out from a suitable distribution,
the predictive distribution of the future asset return can be obtained as usual.
56
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Hyper-Parameter Prior
ο΄ Based on the hyper parameter prior, Jorion (1986) obtains
1
π€ PJ = (π PJ )−1 π PJ
πΎ
where
π PJ = (1 − π£)π + π£ ππ 1π
′
1
1
π
π
π PJ = 1 +
π+
′
π+π
π(π + 1 + π) 1π π −1 1π
1
π
π+2
π£=
(π + 2) + π(π − ππ 1π )′ π −1 (π − ππ 1π )
π = (π + 2)/[(π − ππ 1π )′ π −1 (π − ππ 1π )]
with π = ππ/(π − π − 2) an adjusted sample covariance matrix, and ππ = 1π ′ π −1 π/1π ′ π −1 1π the
average excess return on the sample global minimum-variance portfolio.
57
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Shrinkage Target
ο΄ An alternative motivation of Jorion (1986) portfolio rule is from a shrinkage perspective by
considering the Bayes-Stein estimator of the expected return, π, with
πBS = (1 − π£)π + π£ππ 1π
where ππ 1π is the shrinkage target with
ππ = 1π ′ π −1 π/1π ′ π −1 1π
and π£ is the weight given to the target.
ο΄ Jorion (1986) as well as subsequent studies find that π€ PJ improves π€ ML substantially,
implying that it does so also for the Bayesian strategy under the diffuse prior.
58
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ The Black-Litterman (BL) approach attempts to propose new estimates for expected returns.
ο΄ Indeed, the sample means are simply too noisy.
ο΄ Asset pricing models - even if misspecified - could potentially deliver a good guidance.
ο΄ To illustrate, you consider a K-factor model (factors are portfolio spreads) and run the time
series regression
ππ‘π = α + β1 π1π‘ + β2 π2π‘ + β― + βπΎ ππΎπ‘ + ππ‘
π×1
π×1
π×1
π×1
π×1
π×1
ο΄ Then the estimated excess mean return is given by
ππ = π½1 ππ1 + π½2 ππ2 + β― + π½πΎ πππΎ
where π½1 , π½2 … π½πΎ are the sample estimates of the factor loadings, and ππ1 , ππ2 … πππΎ are the
sample estimates of the factor mean returns.
59
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ The BL approach combines a model (CAPM) with some views, either relative or absolute,
about expected returns.
ο΄ The BL vector of mean returns is given by
μπ΅πΏ =
π×1
Σ
1×1 π×π
−1
−1
+ π′ Ω−1 π
π×πΎ πΎ×πΎ πΎ×π
Σ
1×1 π×π
−1
μππ + π′ Ω−1 μπ£
π×1
π×πΎ πΎ×πΎ πΎ×1
ο΄ We need to understand the essence of the following parameters, which characterize the mean
return vector: , π ππ , π, π, Ω, ππ£
ο΄ Starting from the Σ matrix - you can choose any feasible specification either the sample
covariance matrix, or the equal correlation, or an asset pricing based covariance.
60
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ The π ππ , which is the equilibrium vector of expected return, is constructed as follows.
ο΄ Generate πππΎπ , the N ×1 vector denoting the weights of any of the N securities in the market
portfolio based on market capitalization.
ο΄ Of course, the sum of weights must be unity.
ο΄ Then, the price of risk is πΎ =
the market portfolio.
ππ −π
π
2
where ππ and ππ
are the expected return and variance of
2
ππ
ο΄ Later, we will justify this choice for that price of risk.
ο΄ One could pick a range of values for πΎ and examine performance for each choice.
ο΄ If you work with monthly observations, then switching to the annual frequency does not
change πΎ as both the numerator and denominator are multiplied by 12 under the IID
assumption.
ο΄ It does change the Sharpe ratio, however, as the standard deviation grows with the horizon
by the square root of the period, while the expected return grows linearly.
61
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ Having at hand both πππΎπ and πΎ, the equilibrium return vector is given by
π ππ = πΎΣπππΎπ
ο΄ This vector is called neutral mean or equilibrium expected return.
ο΄ To understand why, notice that if you have a utility function that generates the tangency
portfolio of the form
π€ππ = π′
−1 ππ
−1 π π
ο΄ Then using π ππ as the vector of excess returns on the N assets would deliver πππΎπ as the
tangency portfolio.
62
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ The question being – would you get the same vector of equilibrium mean return if you
directly use the CAPM?
ο΄ Yes, if…
ο΄ Under the CAPM the vector of excess returns is given by
π
μπ = β ππ
π
πππ£ π , πππ
π½=
2
ππ
π×1
π×1
π
π ′
πππ£ π , π π€ππΎπ
π€ππΎπ
=
2
2
ππ
ππ
π€ππΎπ π
π
πΆπ΄ππ: μ =
ππ = πΎ
π€ππΎπ
2
ππ
π×1
63
=
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ Since
π
ππ
= ππ ′ π€ππΎπ and πππ = π π ′ π€ππΎπ
Then
π
π
π
ππ
= 2
ππ
π€ππΎπ = π ππ
ο΄ So indeed, if you use (i) the sample covariance matrix, rather than any other specification, as
well as (ii)
ππ − π
π
πΎ=
2
ππ
then the BL equilibrium expected returns and expected returns based on the CAPM are
identical.
64
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ In the BL approach the investor/econometrician forms some views about expected returns as
described below.
ο΄ P is defined as that matrix which identifies the assets involved in the views.
ο΄ To illustrate, consider two "absolute" views only.
ο΄ The first view says that stock 3 has an expected return of 5% while the second says that stock
5 will deliver 12%.
ο΄ In general the number of views is K.
ο΄ In our case K=2.
ο΄ Then P is a 2 ×N matrix.
ο΄ The first row is all zero except for the third entry which is one.
ο΄ Likewise, the second row is all zero except for the fifth entry which is one.
65
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ Let us consider now two "relative views".
ο΄ Here we could incorporate market anomalies into the BL paradigm.
ο΄ Anomalies are cross sectional patterns in stock returns unexplained by the CAPM.
ο΄ Example: price momentum, earnings momentum, value, size, accruals, credit risk, dispersion,
and volatility.
ο΄ Let us focus on price momentum and the value effects.
ο΄ Assume that both momentum and value investing outperform.
ο΄ The first row of P corresponds to momentum investing.
ο΄ The second row corresponds to value investing.
ο΄ Both the first and second rows contain N elements.
66
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ Winner stocks are the top 10% performers during the past six months.
ο΄ Loser stocks are the bottom 10% performers during the past six months.
ο΄ Value stocks are 10% of the stocks having the highest book-to-market ratio.
ο΄ Growth stocks are 10% of the stocks having the lowest book-to-market ratios.
ο΄ The momentum payoff is a return spread – return on an equal weighted portfolio of winner
stocks minus return on equal weighted portfolio of loser stocks.
ο΄ The value payoff is also a return spread – the return differential between equal weighted
portfolios of value and growth stocks.
67
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ Suppose that the investment universe consists of 100 stocks
ο΄ The first row gets the value 0.1 if the corresponding stock is a winner (there are 10 winners
in a universe of 100 stocks).
ο΄ It gets the value -0.1 if the corresponding stock is a loser (there are 10 losers).
ο΄ Otherwise, it gets the value zero.
ο΄ The same idea applies to value investing.
ο΄ Of course, since we have relative views here (e.g., return on winners minus return on losers)
then the sums of the first row and the sum of the second row are both zero.
ο΄ The same applies to value versus growth stocks.
68
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ Rule: the sum of the row corresponding to absolute views is one, while the sum of the row
corresponding to relative views is zero.
ο΄ ππ£ is the K ×1 vector of K views on expected returns.
ο΄ Using the absolute views above
ππ£ = 0.05,0.12 ′
ο΄ Using the relative views above, the first element is the payoff to momentum trading strategy
(sample mean); the second element is the payoff to value investing (sample mean).
ο΄ πΊ is a K ×K covariance matrix expressing uncertainty about views.
ο΄ It is typically assumed to be diagonal.
ο΄ In the absolute views case described above πΊ 1,1 denotes uncertainty about the first view
while πΊ 2,2 denotes uncertainty about the second view – both are at the discretion of the
econometrician/investor.
69
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Black-Litterman Model
ο΄ In the relative views described above: πΊ 1,1 denotes uncertainty about momentum. This
could be the sample variance of the momentum payoff.
ο΄ πΊ 2,2 denotes uncertainty about the value payoff. This is the could be the sample variance of
the value payoff.
ο΄ There are many debates among professionals about the right value of π.
ο΄ From a conceptual perspective it should be 1/T where T denotes the sample size.
ο΄ You can pick π = 0.1
ο΄ You can also use other values and examine how they perform in real-time investment
decisions.
70
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Extending BL to Incorporate Sample Moments
ο΄ Consider a sample of size T, e.g., T=60 monthly observations.
ο΄ Let us estimate the mean and covariance of our N assets based on the sample.
ο΄ Then the vector of expected return that serves as an input for asset allocation is given by
−1
π = Δ−1 + (ππ πππππ /π)−1
β Δ−1 ππ΅πΏ + (ππ πππππ /π)−1 ππ πππππ
where
Δ = (πΣ)−1 +π′ Ω−1 π −1
71
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Asset Pricing Prior
ο΄ Pástor (2000) and Pástor and Stambaugh (1999) introduce interesting priors that reflect an
investor’s degree of belief in an asset pricing model.
ο΄ To see how this class of priors is formed, assume π
π‘ = (π¦π‘ , π₯π‘ ), where π¦π‘ contains the excess
returns of π non-benchmark positions and π₯π‘ contains the excess returns of
πΎ (= π − π) benchmark positions.
ο΄ Consider a factor model multivariate regression
π¦π‘ = πΌ + π΅π₯π‘ + π’π‘
where π’π‘ is an π × 1 vector of residuals with zero means and a non-singular covariance
matrix Σ = π11 − π΅π22 π΅′ , and πΌ and π΅ are related to π and π through
−1
πΌ = π1 − π΅π2 ,
π΅ = π12 π22
where ππ and πππ (π, π = 1,2) are the corresponding partitions of π and π,
π1
π
π12
π = π , π = 11
π21 π22
2
72
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Asset Pricing Prior
ο΄ For a factor-based asset pricing model, such as the three-factor model of Fama and French
(1993), the restriction is πΌ = 0.
ο΄ To allow for mispricing uncertainty, Pástor (2000), and Pástor and Stambaugh (2000) specify
the prior distribution of πΌ as a normal distribution conditional on Σ,
1
2
πΌ|Σ ∼ π 0, ππΌ 2 Σ
π π΄
where π Σ2 is a suitable prior estimate for the average diagonal elements of Σ. The above alphaSigma link is also explored by MacKinlay and Pástor (2000) in the classical framework.
ο΄ The magnitude of ππΌ represents an investor’s level of uncertainty about the pricing ability of
a given model.
ο΄ When ππΌ = 0, the investor believes dogmatically in the model and there is no mispricing
uncertainty.
ο΄ On the other hand, when ππΌ = ∞, the investor believes that the pricing model is entirely
useless.
73
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Objective Prior
ο΄ Previous priors are placed on π and π, not on the solution to the problem, the portfolio
weights.
ο΄ In many applications, supposedly innocuous diffuse priors on some basic model parameters
can actually imply rather strong prior convictions about particular economic dimensions of
the problem.
ο΄ For example, in the context of testing portfolio efficiency, Kandel, McCulloch, and Stambaugh
(1995) find that the diffuse prior in fact implies a strong prior on inefficiency of a given
portfolio.
ο΄ Tu and Zhou (2009) show that the diffuse prior implies a large prior differences in portfolios
weights across assets.
ο΄ In short, diffuse priors can be unreasonable in an economic sense in some applications.
74
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Objective Prior
ο΄ As a result, it is important to use informative priors on the model parameters that can imply
reasonable priors on functions of interest.
ο΄ Tu and Zhou (2009) advocate a method of constructing priors based on a prior on the solution
of an economic objective.
ο΄ In maximizing an economic objective, even before the Bayesian investor observes any data,
he/she is likely to have some idea about the range of the solution.
ο΄ This will allow the investor to form a prior on the solution, from which the prior on the
parameters can be backed out.
ο΄ In other words, to maximizing the mean-variance utility here, an investor may have a prior
on the portfolio weights, such as an equal- or value-weighted portfolios of the underlying
assets.
75
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Objective Prior
ο΄ This prior can then be transformed into a prior on π and π.
ο΄ This prior on π and π implies a reasonable prior on the portfolio weights by construction.
ο΄ Because of the way such priors on the primitive parameters are motivated, they are called
objective-based priors.
ο΄ Formally, for the current portfolio choice problem, the objective-based prior starts from a
prior on π€,
π€ ∼ π(π€0 , π0 π −1 /πΎ)
where π€0 and π0 are suitable prior constants with known values, and then back out a prior on
π,
1
2
π ∼ π πΎππ€0 , ππ 2 π
π
and where π 2 is the average of the diagonal elements of π.
ο΄ The prior on π can be taken as the usual inverted Wishart distribution.
76
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Investing in Mutual Funds
ο΄ Baks, Metrick, and Wachter (JF 2001) (henceforth BMW) and Pastor and Stambaugh (JFE
2002) have explored the role of prior information about fund performance in making
investment decisions.
ο΄ BMW consider a mean variance optimizing investor who is largely skeptical about the ability
of a fund manager to pick stocks and time the market.
ο΄ They find that even with a high degree of skepticism about fund performance the investor
would allocate considerable amounts to actively managed funds.
ο΄ Pastor and Stambaugh nicely extend the BMW methodology to the case where prior
uncertainly is not only about managerial skills but also about model pricing abilities.
ο΄ In particular, starting from Jensen (1965), mutual fund performance is typically defined as
the intercept in the regression of the fund’s excess returns on excess return of one or more
benchmark assets.
ο΄ However, the intercept in such time series regressions could reflect a mix of fund
performance as well as model mispricing.
77
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: Mutual Funds
ο΄ In particular, consider the case wherein benchmark assets used to define fund performance
are unable to explain the cross section dispersion of passive assets, that is, the sample alpha
in the regression of nonbenchmark passive assets on benchmarks assets is nonzero.
ο΄ Then model mispricing emerges in the performance regression.
ο΄ Thus, Pastor and Stambaugh formulate prior beliefs on both performance and mispricing.
ο΄ Geczy, Stambaugh, and Levin (2005) apply the Pastor Stambaugh methodology to study the
cost of investing in socially responsible mutual funds.
ο΄ Comparing portfolios of these funds to those constructed from the broader fund universe
reveals the cost of imposing the socially responsible investment (SRI) constraint on investors
seeking the highest Sharpe ratio.
ο΄ This SRI cost depends crucially on the investor’s views about asset pricing models and stockpicking skill by fund managers.
78
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Mutual Funds: Prior dependence across funds
ο΄ BMW and Pastor and Stambaugh assume that the prior on alpha is independent across
funds.
ο΄ Jones and Shanken (JFE 2002) show that under such an independence assumption, the
maximum posterior mean alpha increases without bound as the number of funds increases
and "extremely large" estimates could randomly be generated.
ο΄ This is true even when fund managers display no skill.
ο΄ Thus they propose incorporating prior dependence across funds.
ο΄ Then, investors aggregate information across funds to form a general belief about the
potential for abnormal performance.
ο΄ Each fund’s alpha estimate is shrunk towards the aggregate estimate, mitigating extreme
views
79
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stock Return Predictability Based on Macro Variables
ο΄ Prior to analyzing asset allocation in the presence of return predictability, here is some quick
background about return predictability. The literature is vast.
ο΄ Indeed, a question of long-standing interest to both academics and practitioners is whether
returns on risky assets are predictable by aggregate variables such as the dividend yield, the
default spread, the term spread, the consumption-to-wealth ratio, etc.
ο΄ Fama and Schwert (1977), Keim and Stambaugh (1986), Fama and French (1989), and
Lettau and Ludvigson (2001), among others, identify ex ante observable variables that
predict future returns on stocks and bonds.
ο΄ The evidence on predictability is typically based upon the joint system
ππ‘ = π + π½ ′ π§π‘−1 + π’π‘
π§π‘ = π + ππ§π‘−1 + π£π‘
ο΄ Statistically, predictability means that at least one of the π½ coefficients is significant at
conventional levels.
ο΄ Economically, predictability means that you can properly time the market, switching between
an equity fund and a money market fund, based on expected stock return.
80
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stock Return Predictability - Is The Evidence on
Predictability Robust?
ο΄ Predictability based on macro variables is still a research controversy:
— Asset pricing theories do not specify which ex ante variables predict asset returns;
— Menzly, Santos, and Veronesi (2006) among others provide some theoretical validity for the predictive
power of variables like the dividend yield – but that is an ex post justification.
— Statistical biases in slope coefficients of a predictive regression;
— Potential data mining in choosing the macro variables;
— Poor out-of-sample performance of predictive regressions;
ο΄ Schwert (2003) shows that time-series predictability based on the dividend yield tends to
attenuate and even disappears after its discovery.
ο΄ Indeed, the power of macro variables to predict the equity premium substantially
deteriorates during the post-discovery period.
81
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stock Return Predictability - Potential Data Mining
ο΄ Repeated visits of the same database lead to a problem that statisticians refer to as data
mining (also model over-fitting or data snooping).
ο΄ It reflects the tendency to discover spurious relationships by applying tests inspired by
evidence coming up from prior visits to the same database.
ο΄ Merton (1987) and Lo and MacKinlay (1990), among others, discuss the problems of overfitting data in tests of financial models.
ο΄ In the context of predictability, data mining has been investigated by Foster, Smith, and
Whaley (FSW - 1997).
82
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stock Return Predictability - Foster, Smith, and Whaley
(1997)
ο΄ FSW adjust the test size for potential over-fitting, using theoretical approximations as well as
simulation studies.
ο΄ They assume that
1.
M potential predictors are available.
2.
All possible regression combinations are tried.
3.
Only π < π predictors with the highest π
2 are reported.
ο΄ Their evidence shows:
83
1.
Inference about predictability could be erroneous when potential specification search is
not accounted for in the test statistic.
2.
Using other industry, size, or country data as a control to guard against variableselection biases can be misleading.
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stock Return Predictability - The Poor Out-of-Sample
Performance of Predictive Regressions
ο΄ A good reference here is Bossaerts and Hillion (1999) and a follow up work by Goyal and
Welch (2006).
ο΄ The adjustment of the test size merely help in correctly rejecting the null of no relationship.
ο΄ It would, however, provide little information if the empiricist is asked to discriminate
between competing models under the alternative of existing relation.
ο΄ BH and GW propose examining predictability using model selection criteria.
ο΄ Suppose there are M potential predictors.
ο΄ There are 2π competing specifications.
ο΄ Select one winning specification based on the adjusted π
2 , the AIC, the SIC, and other criteria.
ο΄ The selected model (regardless of the criterion used) always retains predictors – not a big
surprise – indicating in-sample predictability.
84
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stock Return Predictability
1
ο΄ Implicitly you assign a π probability that the IID model is correct - so you are biased in favor
2
of detecting predictability.
ο΄ The out-of-sample performance of the selected model is always a disaster.
ο΄ The Bayesian approach of model averaging improves out-of-sample performance (see
Avramov (2002)).
ο΄ BMA can be extended to account for time varying parameters.
ο΄ There are three major papers responding to the apparently nonexistent out of sample
predictability of the equity premium.
ο΄ Cochrane (2008) points out that we should detect predictability either in returns or dividend
growth rates.
ο΄ Campbell and Thompson (2008) document predictability after restricting the equity premium
to be positive.
ο΄ Rapach, Strauss, and Zhou (2010) combine model forecast, similar to the Bayesian Model
Averaging concept, but using equal weights and considering a small subset of models.
85
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Predictive regressions- Finite sample bias in the slope
coefficients
ο΄ Recall, we work with (assuming one predictor for simplicity) the predictive system
ππ‘ = π + π½π§π‘−1 + π’π‘
π§π‘ = π + ππ§π‘−1 + π£π‘
ο΄ Now, let ππ£2 denote the variance of π£π‘ , and let ππ’π£ denote the covariance between π’π‘ and π£π‘ .
ο΄ We know from Kandall (1954) that the OLS estimate of the persistence parameter π is
biased, and that the bias is −1(1 + 3π)/π.
ο΄ Stambaugh (1999) shows that under the normality assumption, the finite sample bias in π½,
the slope coefficient in a predictive regression, is
ππ’π£ 1 + 3π
π΅πππ = πΌ π½ − π½ = − 2
π
ππ£
ο΄ The bias can easily be derived.
86
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Finite Sample Bias
ο΄ Note that the OLS estimates of π½ and π are
π½ = (π ′ π)−1 π ′ π
= π½ + (π ′ π)−1 π ′ π
π = (π ′ π)−1 π ′ π = π + (π ′ π)−1 π ′ π
where
π
= [π1 , π2 , … , ππ ]′ , π = [π§1 , π§2 , … , π§π ]′
π = [π’1 , π’2 , … , π’ π ]′ , π = [π£1 , π£2 , … , π£π ]′
π = [π π , π−1 ], π π is a π-dimension vector of ones
π−1 = [π§0 , π§1 , … , π§π−1 ]′
87
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Finite Sample Bias
ο΄ Also note that π’π‘ can be decomposed into two orthogonal components
ππ’π£
π’π‘ = 2 π£π‘ + ππ‘
ππ£
where ππ‘ is uncorrelated with π§π‘−1 and π£π‘ .
ο΄ Hence, the predictive regression slope can be rewritten as
ππ’π£
π½ = π½ + 2 (π − π) + (π ′ π)−1 π ′ πΈ
ππ£
where πΈ = [π1 , π2 , … , ππ ]′ .
ο΄ Amihud and Hurvich (2004, 2009) advocate a nice approach to deal with statistical inference
in the presence of a small sample bias.
88
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation: With Predictive Variables
ο΄ Consider a one-period optimizing investor who must allocate at time π funds between the
value-weighted NYSE index and one-month Treasury bills.
ο΄ The investor makes portfolio decisions based on estimating the predictive system
ππ‘ = π + π ′ π§π‘−1 + π’π‘
π§π‘ = π + ππ§π‘−1 + π£π‘
where ππ‘ is the continuously compounded NYSE return in month π‘ in excess of the
continuously compounded T-bill rate for that month, π§π‘−1 is a vector of π predictive variables
observed at the end of month π‘ − 1, π is a vector of slope coefficients, and π’π‘ is the regression
disturbance in month π‘.
ο΄ The evolution of the predictive variables is essentially stochastic, as shown earlier. Here, the
evolution is crucial for understanding expected return and risk over long horizons.
ο΄ The regression residuals are assumed to obey the normal distribution and are IID.
89
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ In particular, let ππ‘ = [π’π‘ , π£π‘ ′ ]′ then ππ‘ ∼ π(0, Σ)
where
ππ’2
Σ=
ππ£π’
ππ’π£
Σπ£
ο΄ The distribution of ππ+1 , e.g., the time π + 1 NYSE excess return, conditional on data and
model parameters is π(π + π ′ π§π , ππ’2 ).
ο΄ Assuming the inverted Wishart prior distribution for Σ and multivariate normal prior for the
intercept and slope coefficients in the predictive system, the Bayesian predictive distribution
π ππ+1 |Φ π obeys the Student t density.
ο΄ Then, considering a power utility investor with parameter of relative risk aversion denoted
by πΎ the optimization formulation is
1−πΎ
(1
−
π)
ππ₯π
(
π
)
+
π
ππ₯π
(
π
+
π
)
π
π
π+1
π∗ = arg max
π ππ+1 |Φ π πππ+1
π
1
−
πΎ
ππ+1
subject to π being nonnegative.
90
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ It is infeasible to have analytic solution for the optimal portfolio.
ο΄ Then, considering a power utility investor with parameter of relative risk aversion denoted
by πΎ the optimization formulation is
1−πΎ
(1
−
π)
ππ₯π
(
π
)
+
π
ππ₯π
(
π
+
π
)
π
π
π+1
π∗ = arg max
π ππ+1 |Φ π πππ+1
π
1
−
πΎ
ππ+1
subject to π being nonnegative.
ο΄ Then, given πΊ independent draws for π
π+1 from the predictive distribution, the optimal
portfolio is found by implementing a constrained optimization code to maximize the quantity
1
πΊ
πΊ
1−πΎ
(π)
(1 − π) ππ₯π ( ππ ) + π ππ₯π ( ππ + π
π+1 )
π=1
1−πΎ
subject to π being nonnegative.
91
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Predictive Regressions with Stochastic Initial Observation
ο΄ Let us revisit the predictive regression with one predictor only
ππ‘ = π + ππ§π‘−1 + π’π‘
π§π‘ = π + ππ§π‘−1 + π£π‘
π’π‘
ππ’2 ππ’π£
0
π£π‘ ~ 0 , ππ’π£ ππ£2
ο΄ The system of two equations can be re-expressed in the normal multivariate form
π
= ππ½ + π’
where R is a π × 2 matrix with the first (second) column consisting of excess stock return
(current values of the predictors) and X is a π × 2 matrix with the first (second) column
consisting of a π × 1 vector of ones (lagged values of the predictors).
ο΄ Previously we analyzed such multivariable regressions, assuming the initial observation is
non stochastic.
ο΄ In particular, the posterior is given by (recall from the section on multivariate regression)
π+π+1
1
−
2
Σ
ππ₯π − π½ − π½ Σ −1 ⊗ Ψ π½ − π½ + π‘π πΣ −1
2
92
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Stochastic Initial Observation
ο΄ In the case of stochastic initial observation we multiply this posterior by π π§0 π, Σ which is
1
2 2
1−π
2πππ£2
1 − π2
π
ππ₯π −
π§0 −
1−π
2ππ£2
2
ο΄ Notice that this is the normal distribution with unconditional mean and variance of the
predictive variables.
ο΄ Integrating the posterior analytically to obtain the marginal posterior density of the
parameters does not appear to be feasible
ο΄ Instead, the posterior density can be obtained using the MH algorithm.
ο΄ The candidate distribution is normal/inverted Wishart
ο΄ See Stambaugh (1999) for a detailed discussion.
93
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Kandel and Stambaugh (1996) show that even when the statistical evidence on predictability,
as reflected through the π
2 is the predictive regression, is weak, the current values of the
predictive variables, π§π , can exert a substantial influence on the optimal portfolio.
ο΄ Whereas Kandel and Stambaugh (1996) study asset allocation in a single-period framework,
Barberis (2000) analyzes multi-period investment decisions, considering both a buy-and-hold
investor as well as an investor who dynamically rebalances the optimal stock-bond allocation.
ο΄ Implementing long horizon asset allocation in a buy-and-hold setup is quite straightforward.
ο΄ In particular, let πΎ denote the investment horizon, and let π
π+πΎ = πΎ
π=1 ππ+π be the
cumulative (continuously compounded) return over the investment horizon.
94
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Avramov (2002) shows that the distribution for π
π+πΎ conditional on the data (denoted Φ π )
and set of parameters (denoted Θ) is given by
π
π+πΎ |, Θ, Φ π ∼ π π, ϒ
where
π = πΎπ + π ′ (ππΎ − πΌπ )(π − πΌπ )−1 π§π + π ′ π ππΎ−1 − πΌπ (π − πΌπ )−1 − (πΎ − 1)πΌπ (π − πΌπ )−1 π
πΎ
ϒ = πΎππ’2 +
πΎ
πΏ (π)Σπ£ πΏ(π)′ +
π=1
πΏ(π) = π
′
π
π−1
πΎ
ππ’π£ πΏ(π)′ +
π=1
− πΌπ (π − πΌπ )
−1
πΏ (π)ππ£π’
π=1
ο΄ Drawing from the Bayesian predictive distribution is done in two steps.
ο΄ First, draw the model parameters Θ from their posterior distribution with either fixed or stochastic
first observation.
ο΄ Second, conditional on model parameters, draw π
π+πΎ from the normal distribution.
ο΄ The optimal portfolio is accomplished by numerically maximizing
95
1
πΊ
πΊ
1−πΎ
(π)
(1 − π) ππ₯π ( ππ ) + π ππ₯π ( ππ + π
π+π )
1−πΎ
π=1
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Investment Opportunities: Risk for the Long Run
ο΄ The mean and variance in an IID world increase linearly with the
investment horizon.
ο΄ There is no horizon effect when (i) returns are IID and (ii) estimation risk
is not accounted for, as indeed shown by Samuelson (1969) and Merton
(1969) in an equilibrium framework.
ο΄ Incorporating estimation risk in an IID setup, Barberis (2000) shows that
stocks appear riskier in longer horizons.
ο΄ Incorporating return predictability and estimation risk, Barberis (2000)
shows that investors allocate considerably more heavily to equity the
longer their horizon.
ο΄ This is not clear ex ante - there is a tradeoff between mean reversion and
estimation risk. The mean reversion effect appears to be stronger.
96
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Investment Opportunities: Risk for the Long Run
ο΄ Pastor and Stambaugh (2012) implement a predictive system to show that
stocks may be more risky over long horizons from an investment perspective.
ο΄ They exhibit one additional source of uncertainty about current mean return,
which increases with the horizon.
ο΄ Avramov, Cederbug, and Kvasnakova (2017) suggest model based priors on
the return dynamics and show that per-period variance can be either higher
or lower with the horizon.
ο΄ In particular, prospect theory and habit formation (Long Run Risk) investors
perceive less (more) risky equities due to strong (weak) mean reversion.
ο΄ Bottom line: Economic theory could give important guidance about
investments for the long run as the sample is not informative enough.
ο΄ The next slide gives more intuition about mean reversion.
97
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Risk for the Long Run – Mean Reversion
ο΄ Let ππ‘ be the cc return in time π‘ and ππ‘+1 be the cc return in time π‘ + 1.
ο΄ Assume that π£ππ(ππ‘ ) = π£ππ(ππ‘+1 ).
ο΄ Then the cumulative two period return is π
(π‘, π‘ + 1) = ππ‘ + ππ‘+1 .
ο΄ The question of interest: is π£ππ(π
(π‘, π‘ + 1)) greater than equal to or smaller than 2π£ππ ππ‘ .
ο΄ Of course if stock returns are iid the variance grows linearly with the investment horizon, as
long as estimation risk is overlooked.
ο΄ However, let us assume that stock returns can be predictable by the dividend yield:
ππ‘ = πΌ + π½πππ£π‘−1 + ππ‘
πππ£π‘ = π + πΏπππ£π‘−1 + ππ‘
where π£ππ(ππ‘ ) = π12 , π£ππ(ππ‘ ) = π22 , and πππ£(ππ‘ , ππ‘ ) = π12 and the residuals are uncorrelated
in all leads and lags.
98
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Risk for the Long Run: Mean Reversion
ο΄ It follows that
π£ππ(ππ‘ + ππ‘+1 ) = 2π12 + π½2 π22 + 2π½π12
ο΄ Thus, if π½2 π22 + 2π½π12 < 0 the conditional variance of two period return is
less than the twice conditional variance of one period return, which is
indeed the case based on the empirical evidence.
ο΄ This is the mean-reversion property.
ο΄ While mean reversion makes stocks appear less risky with the investment
horizon, there are other effects (estimation risk, uncertainty about
current and future mean return) which make stocks appear riskier.
ο΄ In the next slide, we decompose the predictive variance of long horizon
return to all these effects.
99
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Risk for the Long Run: The Predictive Variance of Long Horizon
Cumulative Return
ο΄ Based on Avramov, Cederbug, and Kvasnakova (2017), in the traditional predictive
regression setup, the predictive variance consists of four parts:
πππ ππ,π+π |π·π = πΈ πππ 2 |π·π + πΈ
+πΈ
π−1
π=1
ππ πΌ − π΅π₯ −1 πΌ − π΅π₯π
+πππ πππ + ππ πΌ − π΅π₯ −1
100
π−1
−1
π=1 2ππ πΌ − π΅π₯
πΌ − π΅π₯π
−1 πΌ − π΅ π
π₯ ππ πΌ − π΅π₯
π₯
ππΌ − πΌ − π΅π₯ −1 πΌ − π΅π₯π
′
π₯π |π·π
|π·π
ππ₯ + πΌ − π΅π₯π π₯π |π·π
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
The Predictive Variance of Long Horizon Cumulative
Return
ο΄ The first is the IID component
ο΄ The second is the mean-reversion component
ο΄ The third reflects the uncertainty about future mean return
ο΄ The fourth is the estimation risk component
ο΄ In the predictive system of Pastor and Stambaugh (2012) there is a fifth component
describing uncertainty about current mean return.
ο΄ Accounting for model uncertainty induces one more component of the predictive variance.
ο΄ Avramov, Cederbug, and Kvasnakova (2017) also decompose the predictive variance
taking account of the present formula of Campbell and Shiller.
ο΄ They show that the different implications of consumption based models for mean
reversion, as noted earlier, are due to uncertainty on the dividend growth component in
the overall return dynamics.
101
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Model Uncertainly
ο΄ Financial economists have identified variables that predict future stock returns, as noted
earlier.
ο΄ However, the “correct” predictive regression specification has remained an open issue for
several reasons.
ο΄ For one, existing equilibrium pricing theories are not explicit about which variables should
enter the predictive regression.
ο΄ This aspect is undesirable, as it renders the empirical evidence subject to data over-fitting
concerns.
ο΄ Indeed, Bossaerts and Hillion (1999) confirm in-sample return predictability, but fail to
demonstrate out-of-sample predictability.
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Model Uncertainty
ο΄ Moreover, the multiplicity of potential predictors also makes the empirical evidence difficult
to interpret.
ο΄ For example, one may find an economic variable statistically significant based on a particular
collection of explanatory variables, but often not based on a competing specification.
ο΄ Given that the true set of predictive variables is virtually unknown, the Bayesian
methodology of model averaging is attractive, as it explicitly incorporates model uncertainty.
ο΄ The idea is to compute posterior probability for each candidate return forecasting model –
then predicted return is the weighted average of return forecasting models with weights
being the posterior model probabilities.
ο΄ Avramov (2002) derives analytic expressions for model posterior probability. Often numerical
methods are proposed.
ο΄ Assuming equal prior model probability, the posterior probability is a normalized version of
the marginal likelihood, which in turn is a mix of two components standing for model
complexity and goodness-of-fit.
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation under Model Uncertainty
ο΄ In the context of asset allocation, the Bayesian weighted predictive distribution of cumulative
excess continuously compounded returns averages over the model space, and integrates over
the posterior distribution that summarizes the within-model uncertainty about Θπ where π is
the model identifier. It is given by
2π
π π
π+πΎ |Φ π =
π β³π |Φ π
π=1
π Θπ |β³π , Φ π π π
π+πΎ |β³π , Θπ , Φ π πΘπ
Θπ
where π β³π |Φ π is the posterior probability that model β³π is the correct one.
ο΄ Drawing from the weighted predictive distribution is done in three steps.
ο΄ First draw the correct model from the distribution of models.
ο΄ Then conditional upon the model implement the two steps, noted above, of drawing future
stock returns from the model specific Bayesian predictive distribution.
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation - Stock return predictability
and asset pricing models - informative priors
ο΄ The classical approach has examined whether return predictability is explained by rational
pricing or whether it is due to asset pricing misspecification [see, e.g., Campbell (1987),
Ferson and Korajczyk (1995), and Kirby (1998)].
ο΄ Studies such as these approach finance theory by focusing on two polar viewpoints: rejecting
or not rejecting a pricing model based on hypothesis tests.
ο΄ The Bayesian approach incorporates pricing restrictions on predictive regression parameters
as a reference point for a hypothetical investor’s prior belief.
ο΄ The investor uses the sample evidence about the extent of predictability to update various
degrees of belief in a pricing model and then allocates funds across cash and stocks.
ο΄ Pricing models are expected to exert stronger influence on asset allocation when the prior
confidence in their validity is stronger and when they explain much of the sample evidence on
predictability.
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ In particular, Avramov (2004) models excess returns on π investable assets as
rt = α(zt−1 ) + βft + urt ,
α(zt−1 ) = α0 + α1 zt−1
ft = λ(zt−1 ) + uft
λ(zt−1 ) = λ0 + λ1 zt−1
where ππ‘ is a set of πΎ monthly excess returns on portfolio based factors, πΌ0 stands for an πvector of the fixed component of asset mispricing, πΌ1 is an π × π matrix of the time varying
component, and π½ is an π × πΎmatrix of factor loadings.
ο΄ Now, a conditional version of an asset pricing model (with fixed beta) implies the relation
πΌ(rt | zt−1 ) = βλ(zt−1 )
for all π‘, where πΌ stands for the expected value operator.
ο΄ The model imposes restrictions on parameters and goodness of fit in the multivariate
predictive regression
ππ‘ = π0 + π1 π§π‘−1 + π£π‘
where π0 is an π-vector and π1 is an π × π matrix of slope coefficients.
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
rt = (μ0 − βλ0 ) + (μ1 − βλ1 )zt−1 + βft + urt
μ0 = α0 + βλ0
μ1 = α1 + βλ1 .
ο΄ That is, under pricing model restrictions where πΌ0 = πΌ1 = 0 it follows that:
μ0 = βλ0
μ1 = βλ1
ο΄ This means that stock returns are predictable (π1 ≠ 0) iff factors are predictable.
ο΄ Makes sense; after all, under pricing restrictions the systematic component of returns on π
stocks is captured by πΎ common factors.
ο΄ Of course, if we relax the fixed beta assumption - time varying beta could also be a source of
predictability. More later!
ο΄ Is return predictability explained by asset pricing models? Probably not!
ο΄ Kirby (1998) shows that returns are too predictable to be explained by asset pricing models.
ο΄ Ferson and Harvey (1999) show that πΌ1 ≠ 0.
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Avramov and Chordia (2006) show that strategies that invest in individual stocks
conditioning on time varying alpha perform extremely well. More later!
ο΄ So, should we disregard asset pricing restrictions? Not necessarily!
ο΄ The notion of rejecting or not rejecting pricing restrictions on predictability reflects extreme
polar views.
ο΄ What if you are a Bayesian investor who believes pricing models could be useful albeit not
perfect?
ο΄ As discussed earlier, such an approach has been formalized by Black and Litterman (1992)
and Pastor (2000) in the context of IID returns and by Avramov (2004) who accounts for
predictability.
ο΄ The idea is to mix the model and data.
ο΄ This is shrinkage approach to asset allocation.
ο΄ Let ππ and π΄π (ππ and π΄π ) be the expected return vector and variance covariance matrix
based on the data (model).
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Simplistically speaking, moments used for asset allocation are
μ = ωμd + (1 − ω)μm
Σ = ωΣd + (1 − ω)Σm
where ω is the shrinkage factor
ο΄ In particular,
ο΄ If you completely believe in the model you set ω = 0.
ο΄ If you completely disregard the model you set ω = 1.
ο΄ Going with the shrinkage approach means that 0 < ω < 1.
ο΄ The shrinkage of π΄ is quite meaningless in this context.
ο΄ There are other quite useful shrinkage methods of π΄ - see, for example, Jagannathan and Ma
(2005).
10
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ Avramov (2004) derives asset allocation under the pricing restrictions alone, the data alone,
and pricing restrictions and data combined.
ο΄ He shows that
— Optimal portfolios based on the pricing restrictions deliver the lowest Sharpe ratios.
— Completely disregarding pricing restrictions results in the second lowest Sharpe ratios.
— Much higher Sharpe ratios are obtained when asset allocation is based on the shrinkage approach.
11
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Could one exploit predictability to design
outperforming trading strategies?
ο΄ Avramov and Chordia (2006), Avramov and Wermers (2006), and Avramov, Kosowski, Naik,
and Teo (2009) are good references here.
ο΄ Let us start with Avramov and Chordia (2006).
ο΄ They study predictability through the out of sample performance of trading strategies that
invest in individual stocks conditioning on macro variables.
ο΄ They focus on the largest NYSE-AMEX firms by excluding the smallest quartile of firms from
the sample.
ο΄ They capture 3123 such firms during the July 1972 through November 2003 investment
period.
ο΄ The investment universe contains 973 stocks, on average, per month.
11
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation - The Evolution of Stock
Returns
ο΄ The underlying statistical models for excess stock returns, the market premium, and macro
variables are
rt = α(zt−1 ) + β(zt−1 )mkt t + vt
α zt−1 = α0 + α1 zt−1
β(zt−1 ) = β0 + β1 zt−1
mkt t = a + b′ zt−1 + ηt
zt = c + dzt−1 + et
ο΄ Stock level predictability could come up from:
112
1.
Model mispricing that varies with changing economic conditions (α1 ≠ 0);
2.
Factor sensitivities are predictable (β1 ≠ 0);
3.
The equity premium is predictable (b ≠ 0).
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation
ο΄ In the end, time varying model alpha is the major source of predictability and investment
profitability focusing on individual stocks, portfolios, mutual funds, and hedge funds.
ο΄ In the mutual fund and hedge fund context alpha reflects skill (but can also entails
mispricing).
ο΄ Indeed, alpha reflects skill only if the benchmarks used to measure performance are able to
price all passive payoffs.
113
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation - The Proposed Strategy
ο΄ Avramov and Chordia (2006) form optimal portfolios from the universe of AMEX-NYSE stocks
over the period 1972 through 2003 with monthly rebalancing on the basis of various models
for stock returns.
ο΄ For instance, when predictability in alpha, beta, and the equity premium is permissible, the
mean and variance used to form optimal portfolios are
μt−1 = α0 + α1 zt−1 + β(zt−1 ) a + bzt−1
Σt−1 = β(zt−1 )β(zt−1 )′ σ2mkt + Ψ
+ δ1 β(zt−1 )β(zt−1 )′ σ2mkt + δ2 Ψ .
Estimation risk
ο΄ The trading strategy is obtained by maximizing
1
wt = arg max wt ′ μt −
wt ′ Σt−1 + μt−1 μt−1 ′ wt
wt
2(1/γt − rft )
where πΎπ‘ is the risk aversion level.
ο΄ We do not permit short selling of stocks but we do allow buying on margin.
114
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation - Performance evaluation
ο΄ We implement a recursive scheme:
ο΄ The first optimal portfolio is based on the first 120 months of data on excess returns,
market premium, and predictors. (That is, the first estimation window is July 1962
through June 1972.)
ο΄ The second optimal portfolio is based on the first 121 months of data.
ο΄ Altogether, we form 377 optimal portfolios on a monthly basis for each model under
consideration.
ο΄ We record the realized excess return on any strategy
rp,t+1 = ωt ′ rt+1 .
ο΄ We evaluate the ex-post out-of-sample performance of the trading strategies based on the
realized returns.
ο΄ Ultimately, we are able to assess the (quite large) economic value of predictability as well as
show that our strategies successfully rotate across the size, value, and momentum styles
during changing business conditions.
115
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation - So far, we have shown
ο΄ Over the 1972-2003 investment period, portfolio strategies that condition on macro variables
outperform the market by about 2% per month.
ο΄ Such strategies generate positive performance even when adjusted by the size, value, and
momentum factors as well as by the size, book-to-market, and past return characteristics.
ο΄ In the period prior to the discovery of the macro variables, investment profitability is
primarily attributable to the predictability in the equity premium.
ο΄ In the post-discovery period, the relation between the macro variables and the equity
premium is attenuated considerably.
ο΄ Nevertheless, incorporating macro variables is beneficial because such variables drive stocklevel alpha and beta variations.
ο΄ Predictability based strategies hold small, growth, and momentum stocks and load less
(more) heavily on momentum (small) stocks during recessions.
ο΄ Such style rotation has turned out to be successful ex post.
116
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
Bayesian Asset Allocation - Exploiting Predictability in
Mutual Fund Returns
ο΄ Can we use our methodology to generate positive performance based on the universe of
actively managed no-load equity mutual funds?
ο΄ What do we know about equity mutual funds?
— In 2015 about $6 trillion is currently invested in U.S. equity mutual funds, making
them a fundamental part of the portfolio of a domestic investor.
— Active fund management underperforms, on average, passive benchmarks.
— Strategies that attempt to identify subsets of funds using information variables such
as past returns or new money inflows (“hot hands” or “smart money” strategies)
underperform when investment payoffs are adjusted the Fama-French and
momentum benchmarks.
ο΄ Avramov and Wermers (2006) show that strategies that invest in no-load equity funds
conditioning on macro variables generate substantial positive performance.
117
Prof. Doron Avramov, The Jerusalem School of Business Administration, The Hebrew University of Jerusalem, Bayesian Econometrics
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )