Basic Statistics Review Statistical Inference Refresher Population Sampling Sample Statistical Inference • Methods and procedures for drawing Inference about the population Descriptive Statistics • Summarizing and visualizing variables and relationships between two variables • Depends on the type of variable being measured / studied Sample vs Population • Population • The set of all possible cases or units under study • For example • All freshmen undergrads at Rutgers • All single family homes in Middlesex County • Sample • The subset of cases or units used for a statistical study • How to collect the sample • Simple random sample - each case/unit has an equal chance of being selected • Haphazardly collecting a subset of measurements is not a random sample • Bootstrapping - we will discuss this later Organization, Presentation, and Description of Data hist_score <- read.table("~/Desktop/Spring 2019 401/Lecture Material/Hist1.txt", header=T) hist1 <- hist(hist_score$Scores, main="Histogram of Scores", xlab="Scores", ylab="Frequency") Responses Histogram of Scores 4 5 Yes 0 1 2 Frequency 3 No 40 50 60 70 Scores 90 80 70 60 50 Showing Median, Min and Max, and Quartiles BoxPlot of Scores 80 90 100 Central Tendency, Location, Shape, Outliers, Spread • For quantitative variables, we can meaningfully talk about different types of data visualization and associated measures • Center or central tendency or location - measured by mean and median • Spread - measured by variance, range, percentiles, Inter-quartile range • Skewness factor - Measured by tail behavior ((Symmetric, long tails to the right, long tails to the left) • Outliers - those data points that appear to not belong to the data set z-score The z-score for a data value, x, is 𝑥−𝑥 𝑧= 𝑠 • For a population, 𝑥 is replaced with µ and s is replaced with • This is a way of standardizing the data. z-scores are unit less • z-score puts the data on a common scale. Allows apples and oranges comparison • Hence they are called standardized scores • Values farther from 0 are more extreme • A z-score is the number of standard deviations a value falls from the mean • For bell-shaped distributions, 95% of all z-scores fall between -2 and 2 • Exercise: What is the mean and st dev of the standardized scores? Shape of a Distribution Long right tail Symmetric Right-Skewed Long Left tail Left-Skewed 150 0 50 -10 -5 0 5 10 15 -15 -10 -5 0 5 10 15 150 -15 0 50 Frequency Frequency Bell-Shaped Most of the data is centered around the mean / median. Tails to the right and left are rare The 95% Rule for Normal Distributions Normal Probability Distributions X ~ N(m, ) Z = (X - m) / is N(0,1) Normal Probability Distributions X ~ N(m, ) Z = (X - m) / is N(0,1) Dealing With Bivariate / Multivariate Data • A large part of statistical analysis deals with understanding the relationship between two or more quantitative variables associated with a sampling unit • Bivariate Data • Effect of Texting and driving on young adults • Starting pay as a function of the number of years of education • Multivariate Data • Crop yield as a function of soil moisture, rainfall, soil temperature, and so on Scatter Diagram - Examining Bivariate Relationship Population vs Garbage from 1960 to 2007 250 2007 150 200 1990 1980 1970 100 Garbage Generated in Millions of Tonnes 2000 A scatterplot is the graph of the relationship between two quantitative variables. 1960 180 200 220 240 260 US Population in Millions 280 300 Correlation & Direction of Association • A positive association means that values of one variable tend to be higher when values of the other variable are higher • A negative association means that values of one variable tend to be lower when values of the other variable are higher • Two variables are not associated if knowing the value of one variable does not give you any information about the value of the other variable Correlation The correlation is a measure of the strength and direction of linear association between two quantitative variables • Sample correlation: r • Population correlation: (“rho”) Calculating Sample Correlation Coefficient R Simpler formulas for calculating Sxx, Sxy, and Syy Sxx, = Sum(x2) - n * xbar2 Sxy = Sum(x y) - n * xbar * ybar Syy = Sum(y2) - n * ybar2 Population Sampling Sample Statistical Inference Sampling Distribution is Key for Stat Inference A sampling distribution is the distribution of sample statistics computed for different samples of the same size from the same population. • A sampling distribution shows us how the sample statistic varies from sample to sample • Sample statistics vary from sample to sample. They will not match the parameter exactly, but they will have a central tendency and spread • Spread is a measure of the error in estimation • If you take random samples, the sampling distribution will be centered around the true population parameter • If sampling bias exists (if you do not take random samples), your sampling distribution may give you bad information about the true parameter Thinking About Sampling Distribution Sample Population Sample Sample ... Sample Sample Sample Sampling Distribution Calculate statistic for each sample (e.g., mean, proportion, etc.) Standard Error (SE): standard deviation of sampling distribution Center and Shape Center: If samples are randomly selected, the sampling distribution will be centered . around the population parameter Shape: For most of the statistics we consider, if the sample size is large enough the sampling distribution will be symmetric and bell-shaped. Above is essentially what’s known in Statistics as the Central Limit Theorem, a very fundamental concept The standard error of a statistic, SE, is the standard deviation of the sample statistic Margin of Error and Confidence Interval for Pop Mean m • According to the Central Limit Theorem, the sampling distribution of the sample mean is approximately normal for large samples with mean m and st dev / sqrt(n). (𝑋𝑏𝑎𝑟 − 𝜇) • Hence P( -1.96 < < 1.96) = 0.95 𝜎/ 𝑛 • Therefore this interval will contain the true mean m with 95% probability x 1.96 x = x 1.96 n • Thus 95% margin of error for estimating m is 1.96* / sqrt(n) • Similarly 99% margin of error for estimating m is 2.58* / sqrt(n) Example of a CI Using Formulas Hypothesis Testing - Summary • A statistical test is used to determine whether results from a sample are convincing enough to allow us to conclude something about the population. • Null Hypothesis (H0): Claim that there is no effect or difference. • Alternative Hypothesis (Ha): Claim for which we seek evidence. • Type 1 Error : Reject Ho when Ho is true (False Positive). • Type 2 Error : Do Not Reject Ho when H1 is true (False Negative). • How to test - p-value • How to test - critical region • Significance Level p-Value p-value, for a specific statistical test is the probability (assuming H0 is true) of observing a value of the test statistic that is at least as contradictory to the null hypothesis, and supportive of the alternative hypothesis, as the actual one computed from the sample data. p-Value • Probability of obtaining a test statistic more extreme ( or ) than actual sample value, given H0 is true • Called observed level of significance • • Smallest value of for which H0 can be rejected Used to make rejection decision • • If p-value , do not reject H0 If p-value < , reject H0 Steps for Calculating the p-Value for a Test of Hypothesis One Tailed Tests • Determine the value of the test statistic z corresponding to the result of the sampling experiment. • If the test is one-tailed, the p-value is equal to the tail area beyond z in the same direction as the alternative hypothesis. • Thus, if the alternative hypothesis is of the form > , the p-value is the area to the right of, or above, the observed z-value. • Conversely, if the alternative is of the form < , the p-value is the area to the left of, or below, the observed z-value. Steps for Calculating the p-Value for a Test of Hypothesis - Two Tailed Tests • If the test is two-tailed, the p-value is equal to twice the tail area beyond the observed z-value in the direction of the sign of z • That is, if z is positive, the p-value is twice the area to the right of, or above, the observed z-value. • Conversely, if z is negative, the p-value is twice the area to the left of, or below, the observed z-value. Rejection Regions Alternative Hypotheses LowerTailed = .10 z < –1.282 UpperTwo-Tailed Tailed z > 1.282 z < –1.645 or z > 1.645 = .05 z < –1.645 z > 1.645 z < –1.96 or z > 1.96 = .01 z < –2.326 z > 2.326 z < –2.575 or z > 2.575 Testing Hypothesis Example ? • • • 𝑥 = 1261.6 chips s = 117.6 chips n = 42 bags Can we test the Nabisco claim that there are more than 1000 chips in each bag? A group of Air Force cadets bought bags of Chips Ahoy! cookies from all over the country to verify this claim. They hand-counted the number of chips in 42 bags. Source: Warner, B. & Rutledge, J. (1999). “Checking the Chips Ahoy! Guarantee,” Chance, 12(1). Review of Methodologies for CI and Hyp Testing • Population Mean • Population Proportion • Difference in Population Means • Difference in Population Proportions • Paired Difference Inference for Single Population Mean Inference for Difference in Population Means Inference for Difference in Population Prop • If n is large enough for n1p1 ≥ 10, n1(1 – p1) ≥ 10, n2p2 ≥ 10, n2(1 – p2) ≥ 10, then a confidence interval for p1 – p2 can be computed by • 𝑝1 −𝑝2 ± 𝑧* 𝑝1 (1−𝑝1 ) 𝑝 (1−𝑝2 ) + 2 𝑛1 𝑛2 𝑠𝑡𝑎𝑡𝑖𝑠𝑡𝑖𝑐 − 𝑛𝑢𝑙𝑙 𝑧= 𝑆𝐸 𝑧= 𝑝1 −𝑝2 𝑝(1 − 𝑝) 𝑝(1 − 𝑝) + 𝑛1 𝑛2 p^ = Combined estimate of pop prop under Ho Quick ANOVA Review Analysis of Variance • Analysis of Variance (ANOVA) compares the variability between groups to the variability within groups Total Variability = Variability Between Groups + Variability Within Groups Analysis of Variance If the groups are actually different, then which of these is more accurate? a) the variability between groups should be higher than the variability within groups b) the variability within groups should be higher than the variability between groups If the groups are different, there will be high variability between the groups. Questions • How to measure variability between groups? • How to measure variability within groups? • How to compare the two measures? • How to determine significance? Questions • How to measure variability between groups? • How to measure variability within groups? • How to compare the two measures? • How to determine significance? Sums of Squares • We will measure variability as sums of squared deviations (aka sums of squares) • Should be a familiar because that’s how we computed variance and standard deviation Sums of Squares Total Variability = Variability Between Groups + Variability Within Groups ( x x ) = ni ( xi x ) + ( x xi ) 2 2 data value overall mean Sum over all data values group mean overall mean Sum over all groups 2 data value group mean Sum over all data values Sums of Squares Total Variability = Variability Between Groups Variability Within Groups ( x x ) = ni ( xi x ) + ( x xi ) 2 2 SSTotal (Total sum of squares) 2 SSG SSE (sum of squares (“Error” sum = due to groups) + of squares) ANOVA Example - Cuckoo Birds • Cuckoo birds lay their eggs in the nests of other birds • Do cuckoo birds found in nests of different species differ in size? http://opinionator.blogs.nytimes.com/2010/06/01/cuck oo-cuckoo/ Length of Cuckoo Eggs Cuckoo Eggs s1 to s5 k=5 Bird Sample Mean Sample SD Sample Size Pied Wagtail Pipit Robin 22.90 22.50 22.58 1.07 0.97 0.68 15 60 16 Sparrow Wren Overall 23.12 21.13 22.46 1.07 0.74 1.07 14 15 120 𝑥1 to 𝑥5 𝑥 s n1 to n5 n Cuckoo Birds Sums of Squares: SSG = ni ( xi x ) = 35.90 2 + SSE = (x x ) = 101.29 2 i SSTotal = ( x x ) = 137 .19 2 ANOVA Table The “mean square” is the sum of squares divided by the degrees of freedom Source df Groups k-1 Sum of Squares SSG Error n-k SSE Total n-1 SSTotal variability Mean Square MSG = SSG/(k-1) MSE = SSE/(n-k) average variability ANOVA Table • Fill in the beginnings of the ANOVA table based on the Cuckoo birds data. Source df Groups k-1 Error Total n-k Bird Sample Mean Sample SD Sample Size Sum of Squares Mean Square Pied Wagtail 22.90 1.07 15 Pipit 22.50 0.97 60 SSG MSG = SSG/(k-1) Robin 22.58 0.68 16 Sparrow 23.12 1.07 14 Wren 21.13 0.74 15 Overall 22.46 1.07 120 SSE n-1 SSTotal MSE = SSE/(n-k) SSG = 35.90 SSE = 101.29 ANOVA Table Source df Groups 4 Sum of Squares 35.90 Error 115 101.29 Total 119 137.19 Mean Square 35.9/4 = 8.97 101.29/115 = 0.88 • MSS Between Groups / MSS within Groups = 8.97 / 0.88 ~ 10 • If there is no difference between the groups, then this ratio should be small. High ratios shows significant difference between groups • Use the F-statistic (with degrees of freedom = (4,115)) for p-value • p-value is small. Hence reject the null hypthosesis that the pop mean of the different groups are the same Equal Variance • The F-distribution assumes equal within group variability for each group • As a rough rule of thumb, this assumption is violated if the standard deviation of one group is more than double the standard deviation of another group • Homoscedasticity assumption Testing Equal Variance To test for equality of variances among different groups, use the Bartlett’s or Levene’s test (needs a statistical package)
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )