QUANTITATIVE METHODS LEVEL 1 1 READING 1: RATES AND RETURNS READING 2: THE TIME VALUE OF MONEY IN FINANCE READING 3: STATISTICAL MEASURES OF ASSET RETURNS READING 4: PROBABILITY TREES AND CONDITIONAL EXPECTATIONS READING 5: PORTFOLIO MATHEMATICS READING 6: SIMULATION METHODS READING 7: ESTIMATION AND INFERENCE READING 8: HYPOTHESIS TESTING READING 9: PARAMETRIC AND NON-PARAMETRIC TESTS OF INDEPENDENCE READING 10: SIMPLE LINEAR REGRESSION READING 11: INTRODUCTION TO BIG DATA TECHNIQUES 2 READING 1: RATES AND RETURNS 3 LOS 1.a: Interpret interest rates as required rates of return, discount rates, or opportunity costs and explain an interest rate as the sum of a real risk‐free rate and premiums that compensate investors for bearing distinct types of risk No Term Definition 1 Required Rate of Return Equilibrium interest rate that investors demand for a particular investment. 2 Discount Rate = Required Rate of Return. 3 Opportunity Cost Interest rate earned on current consumption. 4 Real Risk-Free Rate Theoretical rate of return on a risk-free investment without inflation. 5 Nominal Risk-Free Rate Real risk-free rate adjusted for expected inflation. 6 Default Risk Risk that a borrower may not fulfill their payment obligations. 7 Liquidity Risk Risk of incurring losses when selling an investment quickly. 8 Maturity Risk Risk associated with the length of time until an investment matures. 9 Required Interest Rate on a Security Sum of the nominal risk-free rate, default risk premium, liquidity risk premium, and maturity risk premium. 10 Holding Period Return (HPR) Return earned from holding an asset over a specific period. 4 LOS 1.b: Calculate and interpret different approaches to return measurement over time and describe their appropriate uses No Statistical Measure 1 Arithmetic Mean Return 2 Geometric Mean Return 3 Note Example Calculating investment return over multiple periods or when measuring compound growth rates For the last three years, the returns for Acme Corporation common stock have been -9.34%, 23.45%, and 8.92%. Calculate the compound annual rate of return over the three-year period. Harmonic Mean Is used for certain computations, such as the average cost of shares purchased over time. An investor purchases $1,000 of mutual fund shares each month, and over the last three months, the prices paid per share were $8, $9, and $10. What is the average cost per share? 5 LOS 1.b: Calculate and interpret different approaches to return measurement over time and describe their appropriate uses No Statistical Measure Note Example 4 Trimmed Mean Excludes a stated %age of the most extreme observations A 5% trimmed mean discards the lowest 2.5% and the highest 2.5% of values and computes the mean of the remaining 95% of values. 5 Winsorized Mean Is calculated by assigning a stated %age of the lowest values equal to one specified low value and a stated %age of the highest values equal to one specified high value, and then it computes a mean from the restated data A 95% winsorized mean sets the bottom 2.5% of values equal to the value at or below which 2.5% of all the values lie (this is called the “2.5th %ile” value) and the top 2.5% of values equal to the value at or below which 97.5% of all the values lie (the “97.5th %ile” value) 6 LOS 1.c: Compare the money-weighted and time-weighted rates of return and evaluate the performance of portfolios based on these measures 1. The money-weighted return ● Accounts for the money invested and provides the investor with information on the actual return she earns on her investment. ● The money-weighted return and its calculation are similar to the internal rate of return and a bond’s yield to maturity. ○ Amounts invested from the investor’s perspective and amounts returned or withdrawn → cash outflows ○ The money that remains at the end of an investment cycle → cash inflow. 7 LOS 1.c: Compare the money-weighted and time-weighted rates of return and evaluate the performance of portfolios based on these measures 2. The time-weighted rate of return ● Measures the compound rate of growth of USD1 initially invested in the portfolio over a stated measurement period. ● The time-weighted rate of return is the preferred performance measure as it neutralizes the effect of cash withdrawals or additions to the portfolio, which are generally outside of the control of the portfolio manager. The annual timeweighted return for an investment may be computed by performing the following steps: ○ Step 1: Value the portfolio immediately preceding significant additions or withdrawals. Form subperiods over the evaluation period that correspond to the dates of deposits and withdrawals. ○ Step 2: Compute the holding period return (HPR) of the portfolio for each subperiod. ○ Step 3: Compute the product of (1 + HPR) for each subperiod to obtain a total return for the entire measurement period [i.e., (1 + HPR1) × (1 + HPR2) … (1 + HPRn)] - 1. If the total investment period is greater than one year, you must take the geometric mean of the measurement period return to find the annual time-weighted rate of return. Example: Assume an investor buys a share of stock for $100 at t = 0, and at the end of the year (t = 1), she buys an additional share for $120. At the end of Year 2, the investor sells both shares for $130 each. At the end of each year in the holding period, the stock paid a $2 per share dividend. What is the money-weighted rate of return and time-weighted rate of return? 8 LOS 1.d: Calculate and interpret annualized return measures and continuously compounded returns, and describe their appropriate uses 1. Annualized rate of return: annual rate of return actually being earned after adjustments have been made for different compounding periods. Example: A saver deposits $100 into a bank account. After 90 days, the account balance is $100.75. What is the saver’s annualized rate of return? Example: An investor buys a 500-day government bill for $970 and redeems it at maturity for $1,000. What is the investor’s annualized return? ● One major limitation of annualizing returns is the implicit assumption that returns can be repeated precisely, that is, money can be reinvested repeatedly while earning a similar return 9 LOS 1.d: Calculate and interpret annualized return measures and continuously compounded returns, and describe their appropriate uses 2. Continuously compounded return: return with the number of compounding periods per year goes to infinity: Effective annual rate = eRcc ‐1 ● The price relative is just the end‐of‐period value divided by the beginning‐of period value. ● The continuously compounded rate of return is: Example: A stock was purchased for $100 and sold one year later for $120. Calculate the investor’s annual rate of return on a continuously compounded basis. 10 LOS 1.e: Calculate and interpret major return measures and describe their appropriate uses No Term Definition 1 Gross return The return earned by an asset manager prior to deductions for management expenses, custodial fees, taxes, or any other expenses that are not directly related to the generation of returns but rather related to the management and administration of an investment. 2 Net return A measure of what the investment vehicle (e.g., mutual fund) has earned for the investor. 3 After-tax nominal return Computed as the total return minus any allowance for taxes on dividends, interest, and realized gains. 4 Real Returns Nominal return adjusted for inflation 5 Leveraged Return A return to an investor that is a multiple of the return on the underlying asset. 11 READING 2: THE TIME VALUE OF MONEY IN FINANCE 12 LOS 2.a: Calculate and interpret the present value (PV) of fixed-income and equity instruments based on expected future cash flows 1. Overview ● Future value ○ Amount to which a current deposit will grow over time when it is placed in an account paying compound interest ○ FVt = PV*(1 + r)t ● Present value ○ The amount of money that must be invested today, at a given rate of return over a given period of time, in order to end up with a specified FV ○ PV= FVt/(1 + r)t Where: PV: amount of money invested today (present value) r: rate of return per period t= total number periods (1 + r)t is future value factor or future value investment factor 13 LOS 2.a: Calculate and interpret the present value (PV) of fixed-income and equity instruments based on expected future cash flows 2. Fixed-income instruments ● Are debt instruments, such as a bond or a loan, that represent contracts under which an issuer borrows money from an investor in exchange for a promise of future repayment. ● The discount rate for fixed-income instruments is an interest rate, and the rate of return on a bond or loan is often referred to as its yield-to-maturity (YTM). ● 3 general patterns of fixed-income instruments: ○ Discount: An investor pays an initial price (PV) for a bond or loan and receives a single principal cash flow (FV) at maturity. The difference (FV − PV) represents the interest earned over the life of the instrument. PV(Discount Bond) = FVt / (1 + r)t ○ Periodic Interest: An investor pays an initial price (PV) for a bond or loan and receives interest cash flows (PMT) at pre-determined intervals over the life of the instrument, with the final interest payment and the principal (FV) paid at maturity. PV(Coupon Bond) = PMT1 / (1 + r)1 + PMT2 / (1 + r)2 + … + (PMTN + FVN) / (1 + r)N ○ Level Payments: An investor pays an initial price (PV) and receives uniform cash flows at pre-determined intervals (A) through maturity which represent both interest and principal repayment ● A perpetual bond is a less common type of coupon bond with no stated maturity date PV(Perpetual Bond) = PMT/r 14 LOS 2.a: Calculate and interpret the present value (PV) of fixed-income and equity instruments based on expected future cash flows 2. Fixed-income instruments Example: A zero-coupon bond with a face value of $1,000 will mature 15 years from today. The bond has a yield to maturity of 4%. Assuming annual compounding, what is the bond’s price? If the bond has a yield to maturity of -0.5%, what is its price, assuming annual compounding? Example: Consider a 10-year, $1,000 par value, 10% coupon, annual-pay bond. What is the value of this bond if its yield to maturity is 8%? Example: Suppose you are considering applying for a $2,000 loan that will be repaid with equal end-of-year payments over the next 13 years. If the annual interest rate for the loan is 6%, how much are your payments? 15 LOS 2.a: Calculate and interpret the present value (PV) of fixed-income and equity instruments based on expected future cash flows 3. Equity investments ● Such as preferred or common stock, represent ownership shares in a company which entitle investors to receive any discretionary cash flows in the form of dividends. ● 3 general approaches of equity instruments: ○ Constant Dividends: An investor pays an initial price (PV) for a preferred or common share of stock and receives a fixed periodic dividend (D) PVt = Dt/r ○ Constant Dividend Growth Rate: An investor pays an initial price (PV) for a share of stock and receives an initial dividend in one period (Dt+1), which is expected to grow over time at a constant rate of g PVt = Dt+1/(r-g) ○ Changing Dividend Growth Rate: An investor pays an initial price (PV) for a share of stock and receives an initial dividend in one period (Dt+1). The dividend is expected to grow at a rate that changes over time as a company moves from an initial period of high growth to slower growth as it reaches maturity. 16 LOS 2.b: Calculate and interpret the implied return of fixed-income instruments and required return and implied growth of equity instruments given the present value (PV) and cash flows 1. Fixed-Income: If we observe the present value (or price) and assume that all future cash flows occur as promised, then the discount rate (r) or yield-to-maturity (YTM) is a measure of implied return under these assumptions for the cash flow pattern. Example: A zero-coupon bond with a face value of $1,000 will mature 15 years from today. The bond’s price is $650. Assuming annual compounding, what is the investor’s annualized return? Example: Consider the 10-year, $1,000 par value, 10% coupon, annual-pay bond we examined in an earlier example, when its price was $1,134.20 at a yield to maturity of 8%. What is its yield to maturity if its price decreases to $1,085.00? 2. Equity: the price of a share of stock reflects not only the required return but also the growth of cash flows 17 LOS 2.c: Explain the cash flow additivity principle, its importance for the no arbitrage condition, and its use in calculating implied forward interest rates, forward exchange rates, and option values. ● The cash flow additivity principle: present value of any stream of cash flows equal the sum of the present value of the cash flows. Example: A security will make the following payments at the end of the next four years: $100, $100, $400, and $100. Calculate the PV of these cash flows using the concept of the PV of an annuity when the appropriate discount rate is 10%. ● No-arbitrage principle: if two sets of future cash flows are identical under all conditions, they will have the same price today ● Implied Forward Rates Example: The 2-period spot rate, S2, is 8%, and the 1-period spot rate, S1, is 4%. Calculate the forward rate for one period, one period from now, 1y1y ● Forward Exchange Rates Example: Consider two currencies, the ABE and the DUB. The spot ABE/DUB exchange rate is 4.5671, the 1-year riskless ABE rate is 5%, and the 1-year riskless DUB rate is 3%. What is the 1-year forward exchange rate that will prevent arbitrage profits? ● Option Pricing Example: Assume an asset has a current price of 40 Chinese yuan (i.e., CNY40). The asset is risky in that its price may rise 40 % to CNY56 during the next time period or its price may fall 20% to CNY32 during the next time period. Calculate the price for a contract on the asset in which the buyer of the contract has the right, but not obligation, to buy the noted asset for CNY50 at the end of the next time period. The risk free rate is 5%. 18 READING 3: STATISTICAL MEASURES OF ASSET RETURNS 19 LOS 3.a: Calculate, interpret, and evaluate measures of central tendency and location to address an investment problem ● Measures of central tendency identify the center, or average, of a data set. This central point can then be used to represent the typical, or expected, value in the data set. ● Population mean vs Sample mean: ● 1. Arithmetic means ○ The population mean and sample mean are both examples of arithmetic means. It is the most widely used measure of central tendency and has the following properties: ■ All interval and ratio data sets have an arithmetic mean. ■ All data values are considered and included in the arithmetic mean computation. ■ A data set has only one arithmetic mean (i.e., the arithmetic mean is unique). ■ The sum of the deviations of each observation in the data set from the mean is always zero. ○ Note: The arithmetic mean of a sample from a population is the best estimate of both the true mean of the 20 sample and the value of the next observation. LOS 3.a: Calculate, interpret, and evaluate measures of central tendency and location to address an investment problem ● 2. Weighted mean: different observations may have a disproportionate influence on the mean. ● 3. Median: midpoint of data set when data is arranged in ascending or descending order. Example: What is the median return for five portfolio managers with 10 year annualized total records of: 30%, 15%, 25%, 21% and 23%. Example: Suppose we add a sixth manager to the previous example with a return of 28%. What is the median return? ● 4. Mode: value that occurs most frequently in data set. It may have more than one mode or even no mode. ○ Unimodal: one value appears most frequently ○ Bimodal and trimodal: set of data has two or three values Example: What is the mode of the following data set: 30%, 28%, 25%, 23%, 28%, 15% and 5%. ● 5. Geometric mean: calculating investment return over multiple periods or when measuring compound growth rates. ○ This equation has a solution only if the product under the radical sign is non‐negative. ○ Note: When calculating the geometric mean for a returns data set, it is necessary to add 1 to each value under 21 the radical and then subtract 1 from the result. LOS 3.a: Calculate, interpret, and evaluate measures of central tendency and location to address an investment problem ● 6. Quantile (known as measure of location): value at or below which stated proportion of the data in a distribution lies. Quantiles and measures of central tendency are known collectively as measures of location ○ Notes: ■ Quartiles‐the distribution is divided into quarters; ■ Quintiles‐the distribution is divided into fifths; ■ Decile‐the distribution is divided into tenths; ■ Percentile‐the distribution is divided into hundredths (%) ■ Note: Any quantile may be expressed as a percentile ○ The formula for the position of the observation at a given percentile, y, with n data points sorted in ascending order: Example: What is the third quartile for the following distribution of returns? 8%, 10%, 12%, 13%, 15%, 17%, 17%, 18%, 19%, 23% ○ Note: The difference between the third quartile and the first quartile is known as the interquartile range ○ To visualize a data set based on quantiles, we can create a box and whisker plot 22 LOS 3.b: Calculate, interpret, and evaluate measures of dispersion to address an investment problem Dispersion is defined as the variability around the central tendency. The common theme in finance and investments is the tradeoff between reward and variability, where the central tendency is the measure of the reward and dispersion is a measure of risk Example: 5‐year annualized total returns for 5 investment managers; 30%, 12%, 25%,20% and 23% ● 1. Range: is the distance between the largest and the smallest value (max‐min) ● 2. Mean absolute deviation (MAD): average of the absolute values of the deviations of individual observations from the arithmetic mean. ● 3. Sample variance: s2 measure of dispersion applying when we are evaluating a sample of n observation from a population. ● 4. Sample standard deviation: calculated by taking the square root of the sample variance 23 LOS 3.b: Calculate, interpret, and evaluate measures of dispersion to address an investment problem ● 5. Relative dispersion: amount of variability in a distribution relative to a reference point benchmark, and measured with coefficient of variation (CV) Example: You have just been presented with a report that indicates that the mean monthly return on T‐bills is 0.25% with a standard deviation of 0.36%, and the mean monthly return for the S&P 500 is 1.09% with a standard deviation of 7.30%. Your unit manager has asked you to compute the CV for these two investments and to interpret your results. ● 6. Target downside deviation, also referred to as the target semideviation, is a measure of dispersion of the observations (here, returns) below the target. Example: Calculate the target downside deviation for a target return equal to the mean (22%) and for a target return of 24% 24 LOS 3.c: Interpret and evaluate measures of skewness and kurtosis to address an investment problem Symmetrical: Distribution is shaped identically on both sides of the mean, implying that intervals of losses and gains will exhibit the same frequency. Note: For symmetrical distribution, unimodal distribution, the mean, median and mode are equal ● 1. Skewness: distribution is not symmetrical, may be either positively or negatively skewed and result from the occurrence of outliers in the data set. A symmetric distribution has skewness of 0. ○ Positively skewed: distribution with many outliers in the upper region, or right tail. positive skew has frequent small losses and a few extreme gains ○ Negative skewed: distribution has a disproportionately large amount of outliers that fall within its lower tail. ○ Skewed distribution shown has a long tail on its left side ○ Outliers: observations with extraordinary large values, either positive or negative. ○ For unimodal distribution 25 LOS 3.c: Interpret and evaluate measures of skewness and kurtosis to address an investment problem ● 2. Kurtosis: measure of combined weight of the tails of distribution relative to the rest of distribution – that is, proportion of the total probability in the tail; ○ Leptokurtic: a distribution has fatter tails than normal distribution tends to generate more‐frequent extremely large deviation from the mean than normal distribution; ○ Platykurtic: a distribution has thinner tails than normal distribution; ○ Mesokurtic: same kurtosis as a normal distribution ○ Excess kurtosis: characterizes kurtosis relative to the normal distribution, is defined as kurtosis minus three → Normal distribution, computed kurtosis is 3 → Normal distribution: excess kurtosis equal 0; leptokurtic: excess kurtosis greater than 0; platykurtic: excess kurtosis less than 0 26 ○ Note: In general, greater positive kurtosis and negative skew indicates increased risk. LOS 3.c: Interpret and evaluate measures of skewness and kurtosis to address an investment problem ● 3. Measures of sample skew and kurtosis ○ Sample skewness: ■ When a distribution is right skewed, sample skewness is positive; ■ When a distribution is left skewed, sample skewness is negative; ■ Value of sample skewness in excess of 0.5 in absolute value are considered significant ○ Sample kurtosis: 27 LOS 3.d: Interpret correlation between two variables to address and investment problem. ● The sample covariance (sXY) is a measure of how two variables in a sample move together: ● The sample correlation coefficient is a standardized measure of how two variables in a sample move together. ● Properties of correlation of two random variables Ri and Rj: ○ Correlation measures the strength of the linear relationship between two random variables. ○ Correlation has no units. ○ The correlation ranges from –1 to +1. That is, –1 ≤ rXY ≤ +1. ○ If rXY = 1.0, the random variables have perfect positive correlation. This means that a movement in one random variable results in a proportional positive movement in the other relative to its mean. ○ If rXY = –1.0, the random variables have perfect negative correlation. This means that a movement in one random variable results in an exact opposite proportional movement in the other relative to its mean. ○ If rXY = 0, there is no linear relationship between the variables, indicating that prediction of Ri cannot be made on the basis of Rj using linear methods. Example: The variance of returns on Stock A is 0.0028, the variance of returns on Stock B is 0.0124, and their covariance of returns is 0.0058. Calculate and interpret the correlation of the returns for Stocks A and B. 28 LOS 3.d: Interpret correlation between two variables to address and investment problem. ● Scatterplots are a method for displaying the relationship between two variables. With one variable on the vertical axis and the other on the horizontal axis, their paired observations can each be plotted as a single point ○ A key advantage of creating scatter plots is that they can reveal non‐linear relationships, which are not described by the correlation coefficient. ● Limitations of Correlation Analysis: Correlation measures the linear association between two variables, but it may not always be reliable. Two variables can have a strong nonlinear relation and still have a very low correlation. Causation isn’t implied just from significant correlation. ● Role of outliers (extreme values) in the correlation of two variables: If removing the outliers significantly reduces the calculated correlation, further inquiry is necessary into whether the outliers provide information or are caused by noise (randomness) in the data used ● Spurious correlation refers to correlation that is either the result of chance relationships in a particular data set29 READING 4: PROBABILITY TREES AND CONDITIONAL EXPECTATIONS 30 LOS 4.a: Calculate expected values, variances, and standard deviations and demonstrate their application to investment problems ● Expected value: weighted average of the possible outcomes for the variable ● Variance and standard deviation measure the dispersion of a random variable around its expected value, sometimes referred to as the volatility of a random variable. Example: Using the probabilities given in the following table, calculate the expected return on Stock A, the variance of returns on Stock A, and the standard deviation of returns on Stock A Probability 30% 50% 20% R(A) 20% 12% 5% 31 LOS 4.b: Formulate an investment problem as a probability tree and explain the use of conditional expectations in investment application ● A probability tree is used to show the probabilities of various outcomes ● Expected values or returns can be calculated using conditional probabilities. Conditional expected values are contingent upon the outcome of some other event. An analyst would use a conditional expected value to revise his expectations when new information arrives. 32 LOS 4.c: Calculate and interpret an updated probability in an investment setting using Bayes’ formula ● Bayes’ formula is used to update a given set of prior probabilities for a given event in response to the arrival of new information Example: There is a 60% probability the economy will outperform, and if it does, there is a 70% chance a stock will go up and a 30% chance the stock will go down. There is a 40% chance the economy will underperform, and if it does, there is a 20% chance the stock in question will increase in value (have gains) and an 80% chance it will not. Given that the stock increased in value, calculate the probability that the economy outperformed. 33 READING 5: PORTFOLIO MATHEMATICS 34 LOS 5.a: Calculate and interpret the expected value, variance, standard deviation, covariances, and correlations of portfolio returns LOS 5.b: Calculate and interpret the covariance and correlation of portfolio returns using a joint probability function for returns ● Weight of portfolio asset: ● Portfolio expected return: ● Covariance: measure how two assets move together: ● Portfolio variance: ● The variance of a portfolio composed of risky asset A and risky asset B can be expressed as: Var(Rp) = wAwACov(RA,RA) + wAwBCov(RA,RB) + wBwACov(RB,RA) + wBwBCov(RB,RB) = wA2σ2(RA)+ 2wAwBσ(RA)σ(RB)ρ(RA,RB) + wB2σ2(RB) ● Correlation coefficient: 35 LOS 5.a: Calculate and interpret the expected value, variance, standard deviation, covariances, and correlations of portfolio returns LOS 5.b: Calculate and interpret the covariance and correlation of portfolio returns using a joint probability function for returns Example: Consider a portfolio with three assets: an index of domestic stocks (60%), an index of domestic bonds (30%), and an index of international equities (10%). Calculate portfolio standard deviation, correlation. Example: The joint probabilities of the returns of Asset A and Asset B are given in the following figure. Calculate the covariance of returns for Asset A and Asset B. 36 LOS 5.c: Define shortfall risk, calculate the safety-first ratio, and identify an optimal portfolio using Roy’s safetyfirst criterion ● Shortfall risk: the probability that a portfolio value or return will fall below a particular (target) value or return over a given time period. ● Roy’s safety‐first criterion: the optimal port minimizes the prob that return of the port falls below some minimum acceptable level (Threshold level). Roy’s safety‐first criterion can be stated as: ○ minimize P(Rp < RL) where: Rp = portfolio return, RL = threshold level return ● If portfolio returns are normally distributed, then Roy’s safety‐first criterion can be stated as: maximize the SFRatio, where SFRatio = ● When choosing among portfolios with normally distributed returns using Roy’s safety‐first criterion, there are two steps: ○ Step 1: Calculate the SFRatio ○ Step 2: Choose the portfolio that has the largest SFRatio. 37 READING 6: SIMULATION METHODS 38 LOS 6.a: Explain the relationship between normal and lognormal distributions and why the lognormal distribution is used to model asset prices when using continuously compounded asset returns ● The lognormal distribution is generated by the function ex, where x is normally distributed. Since the natural logarithm, ln, of ex is x, the logarithms of lognormally distributed random variables are normally distributed, thus the name. ● Properties of lognormal distribution: ○ The lognormal distribution is skewed to the right. ○ The lognormal distribution is bounded from below by zero so that it is useful for modeling asset prices which never take negative values. 39 LOS 6.b: Describe Monte Carlo simulation and explain how it can be used in investment applications LOS 6.c: Describe the use of bootstrap resampling in conducting a simulation based on observed data in investment applications ● Monte Carlo simulation: a technique based on the repeated generation of one or more risk factors that affect security values to generate a distribution of security values. ○ For each of the risk factors, the analyst must specify the parameters of the probability distribution that the risk factor is assumed to follow ○ A computer is then used to generate random values for each risk factor based on its assumed probability distributions ○ Each set of randomly generated risk factors is used with a pricing model to value the security ○ This procedure is repeated many times (100s, 1,000s, or 10,000s), ○ The distribution of simulated asset values is used to draw inferences about the expected (mean) value of the security and possibly the variance of security values about the mean as well ● The simulation procedure for valuation of stock option: ○ 1. Specify the probability distributions of stock prices and of the relevant interest rate, as well as the parameters (mean, variance, possibly skewness) of the distributions. ○ 2. Randomly generate values for both stock prices and interest rates. ○ 3. Value the options for each pair of risk factor values. ○ 4. After many iterations, calculate the mean option value and use that as your estimate of the option’s 40 LOS 6.b: Describe Monte Carlo simulation and explain how it can be used in investment applications LOS 6.c: Describe the use of bootstrap resampling in conducting a simulation based on observed data in investment applications ● Monte Carlo simulation is used to: ○ Value complex securities. ○ Simulate the profits/losses from a trading strategy. ○ Calculate estimates of value at risk (VaR) to determine the riskiness of a portfolio of assets and liabilities. ○ Simulate pension fund assets and liabilities over time to examine the variability of the difference between the two. ○ Value portfolios of assets that have nonnormal returns distributions. ● The limitation of MC: ○ fairly complex; ○ answer depends on the distribution of the risk factors and pricing/valuation model used; ○ it is statistical not analytic ; ○ cannot provide the insights that analytic methods can ● Bootstrap: Calculate standard error of sample mean from sample means of repeated samples of size n => computational demanding, improve accuracy 41 READING 7: ESTIMATION AND INFERENCE 42 LOS 7.a: Compare and contrast simple random, stratified random, cluster, convenience, and judgmental sampling and their implications for sampling error in an investment problem ● Sampling error: The difference between a sample statistic (the mean, variance, or standard deviation of the sample) and its corresponding population parameter (the true mean, variance, or standard deviation of the population) ● The sampling distribution of the sample statistic is a probability distribution of all possible sample statistics computed from a set of equal‐size samples that were randomly drawn from the same population. 43 LOS 7.b: Explain the central limit theorem and its importance for the distribution and standard error of the sample mean LOS 7.c: Describe the use of resampling (bootstrap, jackknife) to estimate the sampling distribution of a statistic ● The central limit theorem states that for simple random samples of size n from a population with a mean μ and a finite variance σ2, the sampling distribution of the sample mean x approaches a normal probability distribution with mean μ and a variance equal to σ2/n as the sample size becomes large. ○ Important properties of the central limit theorem: ■ If the sample size n is sufficiently large (n≥30), the sampling distribution of the sample mean will approximately normal. ■ The mean of the population, μ, and the mean of the distribution of all possible sample means are equal. ■ The variance of the distribution of sample mean is σ2/n, the population variance divided by the sample size. ● Standard error of the sample mean: standard distribution of the distribution of the sample means ○ When do not know the population standard deviation and need to use sample standard deviation ● Jacknife: Calculate standard error of sample mean from multiple sample mean, each with one of the observations 44 removed from the sample => simple, can be used when the number of observations available is small READING 8: HYPOTHESIS TESTING 45 LOS 8.a: Explain hypothesis testing and its components, including statistical significance, Type I and Type II errors, and the power of a test. LOS 8.b: Construct hypothesis tests and determine their statistical significance, the associated Type I and Type II errors, and power of the test given a significance level ● Hypothesis: A statement about the value of a population parameter developed for the purpose of testing a theory or belief. ● Hypothesis testing procedures ● The null hypothesis: the hypothesis that the researcher wants to reject. ● The alternative hypothesis: what is concluded if there is sufficient evidence to reject the null hypothesis. It is the hypothesis that are really trying to assess. 46 LOS 8.a: Explain hypothesis testing and its components, including statistical significance, Type I and Type II errors, and the power of a test. LOS 8.b: Construct hypothesis tests and determine their statistical significance, the associated Type I and Type II errors, and power of the test given a significance level ● Hypothesis testing involves two statistics: the test statistic calculated from the sample data and the critical value of the test statistic. The value of the computed test statistic relative to the critical value is a key step in assessing the validity of a hypothesis. ● The decision for a hypothesis test: either to reject the null hypothesis or fail to reject the null hypothesis. ● The decision rule for rejecting or failing to reject the null hypothesis is based on the distribution of the test statistic ● The p‐value is the probability of obtaining a test statistic leading to a rejection of the null hypothesis, assuming the null hypothesis is true. It is the smallest level of significance for which the null hypothesis can be rejected 47 LOS 8.a: Explain hypothesis testing and its components, including statistical significance, Type I and Type II errors, and the power of a test. LOS 8.b: Construct hypothesis tests and determine their statistical significance, the associated Type I and Type II errors, and power of the test given a significance level ● When drawing inferences from a hypothesis test, there are two types of error: ○ Type I: the rejection of null hypothesis when it is actually true. The significance level (α) is the probability of making a Type I error. ○ Type II: the failure to reject the null hypothesis when it is false. ● The power of the test: the probability of correctly rejecting the null hypothesis when it is false or 1 – P(type II error). ● Decreasing the significance level from 5% to 1 %, will increase the probability of failing to reject the false null (Type II error), reduce the power of the test. ● For given of sample size: increase the power of the test with the cost that the probability of rejecting the true null increases (Type 1 error); ● For given significance level: decrease the probability of type II error and increase the power of the test, by increasing sample size. 48 LOS 8.a: Explain hypothesis testing and its components, including statistical significance, Type I and Type II errors, and the power of a test. LOS 8.b: Construct hypothesis tests and determine their statistical significance, the associated Type I and Type II errors, and power of the test given a significance level ● A test statistic is calculated by comparing the point estimate of the population parameter with the hypothesized value of the parameter ● The focal point of our statistical decision is the value of the test statistic. The test statistic that we use depends on what we are testing 49 1. Test of a single mean 1.1. T‐Test: employs a test statistic distributed according to a t‐distribution ● Using the t‐test if the population variance is unknown and either of the following conditions exist: ● The sample is large (n≥30); ● The sample is small (less than 30), but the distribution of the population is normal or approximately normal. ● If the sample is small and distribution is nonnormal, no reliable statistical test. ● t‐statistic with n‐1 degrees of freedom: 1.2. Z‐test: ● Appropriate hypothesis test of the population mean when the population is normally distributed with known variance: ● When the sample size is large and the population variance is unknown: Example: When your company’s gizmo machine is working properly, the mean length of gizmos is 2.5 inches. However, from time to time the machine gets out of alignment and produces gizmos that are either too long or too short. When this happens, production is stopped and the machine is adjusted. To check the machine, the quality control department takes a gizmo sample each day. Today, a random sample of 49 gizmos showed a mean length of 2.49 inches. The population standard deviation is known to be 0.021 inches. Using a 5% significance level, determine 50 if the machine should be shut down and adjusted. 2. Difference Between Means (Independent Samples) ● A pooled variance is used with the t‐test for testing the hypothesis that the means of two normally distributed populations are equal, when the variances of the populations are unknown but assumed to be equal. Assuming independent samples: ● Note: The degrees of freedom, df, is (n1 + n2 − 2). Example: Sue Smith is investigating whether the abnormal returns for acquiring firms during merger announcement periods differ for horizontal and vertical mergers. She estimates the abnormal returns for a sample of acquiring firms associated with horizontal mergers and a sample of acquiring firms involved in vertical mergers. Smith finds that abnormal returns from horizontal mergers have a mean of 1.0% and a standard deviation of 1.0%, while abnormal returns from vertical mergers have a mean of 2.5% and a standard deviation of 2.0%. Smith assumes the samples are independent, the population means are normally distributed, and the population variances are equal. Smith calculates the t-statistic as −5.474 and the degrees of freedom as 120. Using a 5% significance level, should Smith reject or fail to reject the null hypothesis that the abnormal returns to acquiring firms during the announcement period are the same 51 for horizontal and vertical mergers? 3. Paired Comparisons (Means of Dependent Samples) ● If the observations in the two samples both depend on some other factor, using “paired comparison” ○ “paired comparison”: test of whether the means of the differences between observations for the two samples are different, ○ “paired comparison”: requires that the sample data be normally distributed. ● The general form of the test for any hypothesized mean difference, μdz is as follows: 52 4. Value of a Population Variance ● The chi‐square test is used for hypothesis tests concerning the variance of a normally distributed population. ● The hypotheses for a two‐tailed test of a single population variance are structured as: H0: σ2 = σ02 versus Ha: σ2 ≠ σ02 ● The hypotheses for one‐tailed tests are structured as: H0: σ2 ≤ σ02 versus Ha: σ2 > σ02 or H0: σ2 ≥ σ02 versus Ha: σ2 < σ02 ● Note: The chi‐square distribution is asymmetrical and approaches the normal distribution in shape as the degrees of freedom increase ● The chi‐square test statistic, χ2, with n − 1 degrees of freedom, is computed as: Example: Historically, High‐Return Equity Fund has advertised that its monthly returns have a standard deviation equal to 4%. This was based on estimates from the 2005–2013 period. High‐Return wants to verify whether this claim still adequately describes the standard deviation of the fund’s returns. High‐Return collected monthly returns for the 24‐month period between 2013 and 2015 and measured a standard deviation of monthly returns of 3.8%. High‐Return calculates a test statistic of 20.76. Using a 5% significance level, determine if the more recent standard deviation is 53 different from the advertised standard deviation. 5. Testing the Equality of the Variances of Two Normally Distributed Populations, Based on Two Independent Random Samples ● The hypotheses concerned with the equality of the variances of two populations are tested with an F-distributed test statistic (F‐test) ● The F‐test is used under the assumption that the populations from which samples are drawn are normally distributed and that the samples are independent. ● Let σ12 and σ22 represent the variances of normal Population 1 and Population 2, respectively, ○ The hypotheses for the two‐tailed F‐test of differences in the variances can be structured as: H0:σ12= σ22 versus Ha:σ12≠ σ22 ○ The one‐sided test structures can be specified as: H0: σ12 ≤ σ22 versus Ha: σ12 > σ22 or H0: σ12 ≥ σ22 versus Ha: σ12 < σ22 ● F‐statistic: ● Note: n1 − 1 and n2 − 1 are the degrees of freedom used to identify the appropriate critical value from the Ftable 54 5. Testing the Equality of the Variances of Two Normally Distributed Populations, Based on Two Independent Random Samples ● The upper critical value is always greater than one; ● The lower critical value is always less than one Example: Annie Cower is examining the earnings for two different industries. Cower suspects that the variance of earnings in the textile industry is different from the variance of earnings in the paper industry. To confirm this suspicion, Cower has looked at a sample of 31 textile manufacturers and a sample of 41 paper companies. She measured the sample standard deviation of earnings across the textile industry to be $4.30 and that of the paper industry companies to be $3.80. Cower calculates a test statistic of 1.2805. Using a 5% significance level, determine if the earnings of the textile industry have a different standard deviation than those of the paper industry. 55 LOS 8.c: Compare and contrast parametric and nonparametric tests, and describe situations where each is the more appropriate type of test. ● Parametric tests rely on assumptions regarding the distribution of the population and are specific to population parameters. ● Nonparametric tests either do not consider a particular population parameter or have few assumptions about the population that is sampled. Nonparametric tests are used: ○ There is concern about quantities other than the parameters of a distribution; ○ The assumptions of parametric tests can’t be supported; ● Situations where a nonparametric test is called for are the following: ○ The assumptions about the distribution of the random variable that support a parametric test are not met; ○ When data are ranks (an ordinal measurement scale) rather than values; ○ The hypothesis does not involve the parameters of the distribution, such as testing whether a variable is normally distributed; 56 READING 9: PARAMETRIC AND NON-PARAMETRIC TESTS OF INDEPENDENCE 57 LOS 9.a: Explain parametric and nonparametric tests of the hypothesis that the population correlation coefficient equals zero, and determine whether the hypothesis is rejected at a given level of significance ● The appropriate test statistic for the hypothesis that the population correlation equals zero, when the two variables are normally distributed, is: where r = sample correlation and n = sample size. This test statistic follows a t‐distribution with n – 2 degrees of freedom. Example: A researcher computes the sample correlation coefficient for two normally distributed random variables as 0.35, based on a sample size of 42. Determine whether to reject the hypothesis that the population correlation coefficient is equal to zero at a 5% significance level. ● The Spearman rank correlation test can be used when the data are not normally distributed. ● The Spearman rank correlation coefficient is essentially equivalent to the usual correlation coefficient but is calculated on the ranks of the two variables D=difference in ranks 58 LOS 9.b: Explain tests of independence based on contingency table data ● If we want to test whether there is a relationship between the size and investment type, we can perform a test of independence using a nonparametric test statistic that is chi‐square distributed with (r‐1)(c‐1) degrees of freedom 59 LOS 9.b: Explain tests of independence based on contingency table data 60 LOS 9.b: Explain tests of independence based on contingency table data 61 READING 10: SIMPLE LINEAR REGRESSION 62 LOS 10.a: Describe a simple linear regression model, how the least squares criterion is used to estimate regression coefficients, and the interpretation of these coefficients ● The purpose of simple linear regression is to explain the variation in a dependent variable in terms of the variation in a single independent variation. ○ Note: Variation is interpreted as the degree to which a variable differs from its mean value ● The dependent variable is the variable whose variation is explained by the independent variable. We are interested in answering the question, “What explains fluctuations in the dependent variable?” The dependent variable is also referred to as the terms explained variable, endogenous variable, or predicted variable. ● The independent variable is the variable used to explain the variation of the dependent variable. The independent variable is also referred to as the terms explanatory variable, exogenous variable, or predicting variable. ○ Suppose we want to use excess returns on the S&P500 (the independent variable) to explain the variation in excess returns on ABC common stock (the dependent variable). 63 LOS 10.a: Describe a simple linear regression model, how the least squares criterion is used to estimate regression coefficients, and the interpretation of these coefficients ● Linear regression assumes a linear relationship between the dependent and the independent variables. ○ bo: intercept ○ b1: slope coefficient ○ εi: Error term (residual) for the ith observation ● In simple linear regression, the estimated intercept, and slope, are such that the sum of the squared vertical distances from the observations to the fitted line is minimized. This is the least squares criterion. Because of its common use, linear regression is often referred to as ordinary least squares (OLS) regression. ● The linear equation, often called the line of best fit or regression line, takes the following form: ● Note: The error term refers to the true underlying population relationship, whereas the residual refers to the fitted linear relation based on the sample. 64 LOS 10.a: Describe a simple linear regression model, how the least squares criterion is used to estimate regression coefficients, and the interpretation of these coefficients ● The estimated slope coefficient for the regression line describes the change in Y for a one-unit change in X. It can be positive, negative, or zero, depending on the relationship between the regression variables. The slope term is calculated as follows: ● The intercept term is the line’s intersection with the Y-axis at X = 0. It can be positive, negative, or zero. A property of the least squares method is that the intercept term may be expressed as follows: ● The intercept equation highlights the fact that the regression line passes through a point with coordinates equal to the mean of the independent and dependent variables (i.e., the point X, Y). Example: Compute the slope coefficient and intercept term using the following information, and then interpret each coefficient estimate: Cov(S&P 500, ABC) = 0.000336, Var(S&P 500) = 0.000522 Mean return (S&P 500) = -2.70% Mean return (ABC) = -4.05% 65 LOS 10.a: Describe a simple linear regression model, how the least squares criterion is used to estimate regression coefficients, and the interpretation of these coefficients Interpreting the Regression Coefficients Keep in mind, however, that any conclusions regarding the importance of an independent variable in explaining a dependent variable are based on the statistical significance of the slope coefficient. The magnitude of the slope coefficient tells us nothing about the strength of the linear relationship between the dependent and independent variables. A hypothesis test must be conducted, or a confidence interval must be formed, to assess the explanatory power of the independent variable. Later in this reading we will perform these hypothesis tests. 66 LOS 10.b: Explain the assumptions underlying the simple linear regression model, and describe how residuals and residual plots indicate if these assumptions may have been violated Linear regression assumes the following: ● 1. Linearity: The relationship between the dependent variable, Y, and the independent variable, X, is linear. ● 2. Homoskedasticity: The variance of the regression residuals is the same for all observations. ● 3. Independence: The observations, pairs of Ys and Xs, are independent of one another. This implies the regression residuals are uncorrelated across observations. ● 4. Normality: The regression residuals are normally distributed. 1. Linearity ● If the relationship between the independent and dependent variables is nonlinear in the parameters, estimating that relation with a simple linear regression model will produce invalid results. ● One way of checking for linearity is to examine the model residuals in relation to the independent variable 67 LOS 10.b: Explain the assumptions underlying the simple linear regression model, and describe how residuals and residual plots indicate if these assumptions may have been violated 2. Homoskedasticity • In terms of notation, this assumption relates to the squared residuals: • If the residuals are not homoscedastic, that is, if the variance of residuals differs across observations, then we refer to this as heteroskedasticity 68 LOS 10.b: Explain the assumptions underlying the simple linear regression model, and describe how residuals and residual plots indicate if these assumptions may have been violated 3. Independence and 4. Normality ● The assumption of normality requires that the residuals be normally distributed ● For large sample sizes, we may be able to drop the assumption of normality by appealing to the central limit theorem 69 LOS 10.c: Calculate and interpret measures of fit and formulate and evaluate tests of fit and of regression coefficients in a simple linear regression LOS 10.d: Describe the use of analysis of variance (ANOVA) in regression analysis, interpret ANOVA results, and calculate and interpret the standard error of estimate in a simple linear regression LOS 10.e: Calculate and interpret the predicted value for the dependent variable, and a prediction interval for it, given an estimated linear regression model and a value for the independent variable ● Sum of squares total (SST): The total variation in Y ● Sum of squares error (SSE): The unexplained variation in Y ● Sum of squares regression (SSR): The explained variation in Y. ● Coefficient of determination (the R‐squared or R2), is the %age of the variation of the dependent variable that is explained by the independent variable 70 Note: In a simple linear regression, the square of the pairwise correlation is equal to the coefficient of determination. ● ANOVA table summarizes the variation in the dependent variable ● The standard error of the estimate (se)is a measure of the distance between the observed values of the dependent variable and those predicted from the estimated regression; the smaller the se, the better the fit of the model. 71 ● Whereas the coefficient of determination descriptive, it is not a statistical test. To see if our regression model is likely to be statistically meaningful, we will need to construct an F‐distributed test statistic. ○ Note: This is always a one‐tailed test ○ Note: The se, along with the coefficient of determination and the F‐statistic, is a measure of the goodness of the fit of the estimated regression line. 72 ● To test a hypothesis about a slope, we calculate the test statistic as follow: ○ : standard error of the slope coefficient ○ This test statistic is t‐distributed with n − k − 1 or n − 2 degrees of freedom because two parameters (an intercept and a slope) were estimated in the regression. ○ Note: the t‐test for a simple linear regression is equivalent to a t‐test for the correlation coefficient between x and y ● A forecasted value of the dependent variable, , is determined using the estimated intercept and slope, as well as the expected or forecasted independent variable, Xf : ● Confidence intervals for predicted values (sf is the standard error of the forecast) 73 LOS 10.f: Describe different functional forms of simple linear regressions ● The log‐lin model: The slope coefficient in this model is the relative change in the dependent variable for an absolute change in the independent variable lnYi = b0 + b1Xi ● The lin‐log model: The slope coefficient in this model is the absolute change in the dependent variable for a relative change in the independent variable Yi = b0 + b1lnXi ● The log‐log model: The slope coefficient in this model is the relative change in the dependent variable for a relative change in the independent variable LnYi = b0 + b1lnXi ● Selecting the correct functional form involve determining the nature of the variables and evaluation of the goodness of fit measures (R2, SEE, F‐test). 74 READING 11: INTRODUCTION TO BIG DATA TECHNIQUES 75 LOS 11.a: Describe aspects of “fintech” that are directly relevant for the gathering and analyzing of financial data LOS 11.b: Describe Big Data, artificial intelligence, and machine learning ● Fintech generally refers to technology-driven innovation occurring in the financial services industry ● The term Big Data typically refers to datasets that have the following characteristics: No Characteristics 1 Volume Description The amount of data collected in files, records, and tables is very large, representing many millions, or even billions, of data points 2 Velocity The speed and frequency with which the data are recorded and transmitted has accelerated. Real-time or near-real-time data have become the norm in many areas 3 Variety The data are collected from many different sources and in a variety of formats, including structured data (e.g., SQL tables), semistructured data (e.g., HTML code), and unstructured data (e.g., video messages). 4 credibility andorreliability of different ● DataVeracity can be structured, The semi-structured, unstructured data data sources 76 LOS 11.a: Describe aspects of “fintech” that are directly relevant for the gathering and analyzing of financial data LOS 11.b: Describe Big Data, artificial intelligence, and machine learning ● Sources of Big Data: Data Source Examples Financial Markets Equity, fixed income, futures, options, derivatives Businesses Corporate financials, commercial transactions, credit card purchases Governments Trade, economic, employment, payroll data Individuals Credit card purchases, product reviews, internet search logs, social media posts Sensors Satellite imagery, shipping cargo information, traffic patterns Internet of Things (IoT) Data from "smart" buildings (climate control, energy consumption, security, etc.) ● The three main sources of alternative data Data Source Structure Individuals (text, video, photo, audio) Unstructured Business Processes (corporate and public entity information flows) Structured Sensors (smartphones, cameras, RFID chips, satellites) Unstructured ● Big Data Challenges: quality, volume, and appropriateness of the data. 77 LOS 11.c: Describe Big Data, artificial intelligence, and machine learning ● Artificial intelligence (AI) computer systems are capable of performing tasks that traditionally have required human intelligence. AI technology has enabled the development of computer systems that exhibit cognitive and decision-making ability comparable or superior to that of human beings. ○ The expert system is a type of computer programming that attempted to simulate the knowledge base and analytical abilities of human experts in specific problem-solving contexts. ○ Neural networks is based on how our brain learns and processes information ● Machine learning (ML) involves computer-based techniques that seek to extract knowledge from large amounts of data without making any assumptions on the data’s underlying probability distribution. ○ ML involves splitting the dataset into three distinct subsets: a training dataset, a validation dataset, and a test dataset. ○ Notes: ML techniques can appear to be opaque or “black box” approaches, which arrive at outcomes that might not be entirely understood or explainable. ML approaches can help identify relationships ○ ML can be divided broadly into three distinct classes of techniques: supervised learning, unsupervised learning, and deep learning ■ In supervised learning, computers learn to model relationships based on labeled training data. ■ In unsupervised learning, computers are not given labeled data but instead are given only data from which the algorithm seeks to describe the data and their structure. ■ In deep learning (or deep learning nets), computers use neural networks, often with many hidden 78 layers, to perform multistage, non-linear data processing to identify patterns. LOS 11.d: Describe applications of Big Data and Data Science to investment management ● Data science can be defined as an interdisciplinary field that harnesses advances in computer science (including ML), statistics, and other disciplines for the purpose of extracting information from Big Data (or data in general ● Data Processing Methods No Method Description 1 Capture Data capture refers to how the data are collected and transformed into a format that can be used by the analytical process. Low-latency systems— systems that operate on networks that communicate high volumes of data with minimal delay (latency)—are essential for automated trading applications that make decisions based on real-time prices and market events. In contrast, high-latency systems do not require access to real-time data and calculations. 2 Curation Data curation refers to the process of ensuring data quality and accuracy through a data cleaning exercise. This process consists of reviewing all data to detect and uncover data errors—bad or inaccurate data—and making adjustments for missing data when appropriate. 3 Storage Data storage refers to how the data will be recorded, archived, and accessed and the underlying database design. An important consideration for data storage is whether the data are structured or unstructured and whether analytical needs require low-latency solutions. 4 Search Search refers to how to query data. Big Data has created the need for advanced applications capable of examining and reviewing large quantities of data to locate requested data content. 5 Transfer 79 Transfer refers to how the data will move from the underlying data source or storage location to the underlying LOS 11.d: Describe applications of Big Data and Data Science to investment management ● Text analytics involves the use of computer programs to analyze and derive meaning typically from large, unstructured text- or voice-based datasets, such as company filings, written reports, quarterly earnings calls, social media, email, internet postings, and surveys. ● Natural language processing (NLP) is a field of research at the intersection of computer science, AI, and linguistics that focuses on developing computer programs to analyze and interpret human language 80
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )