Summary Data Analysis Table of Content Introduction to scientific writing ......................................................................................................... 3 Title ..................................................................................................................................................... 3 Abstract ............................................................................................................................................... 3 Introduction + literature review ........................................................................................................... 3 Methodology + measurement .............................................................................................................. 4 Results ................................................................................................................................................. 4 Discussion and conclusion .................................................................................................................. 5 Limitations & Further Research .......................................................................................................... 5 Research process .................................................................................................................................... 6 Hypotheses ............................................................................................................................................. 7 Directional (one tailed) hypothesis...................................................................................................... 7 Non-directional (one tailed) hypothesis .............................................................................................. 7 Experimental approach ........................................................................................................................ 7 Correlational approach ........................................................................................................................ 7 Variables .............................................................................................................................................. 7 Hypothesis Formulation ...................................................................................................................... 8 Errors Hypothesis Testing ................................................................................................................... 8 Type I error .......................................................................................................................................... 9 Statistical Significance (ποΏ½) ........................................................................................................... 9 Type II error......................................................................................................................................... 9 Statistical Power (ποΏ½) ..................................................................................................................... 9 Avoid the Errors ................................................................................................................................ 10 Type I and Type II relationship.......................................................................................................... 11 Effect sizes ........................................................................................................................................ 11 Normal Distribution ............................................................................................................................ 13 Skewness ........................................................................................................................................... 13 Ceiling effect.................................................................................................................................. 14 Floor effect .................................................................................................................................... 14 Kurtosis ............................................................................................................................................. 14 Outliers .............................................................................................................................................. 15 How to deal with outliers?............................................................................................................. 15 Trimming ....................................................................................................................................... 15 Winsorizing ................................................................................................................................... 15 Analyze with robust methods ........................................................................................................ 15 Transform the data ......................................................................................................................... 16 Population.......................................................................................................................................... 16 Probability ......................................................................................................................................... 16 Sampling Distribution ......................................................................................................................... 17 Samples and possibilities................................................................................................................... 17 Standard Error of the Mean ............................................................................................................... 17 Plan Sample Size ............................................................................................................................... 18 t-Distribution ....................................................................................................................................... 19 Degrees of Freedom .......................................................................................................................... 19 Confidence Intervals.......................................................................................................................... 19 Measures of Central Tendency ........................................................................................................... 21 Mean .................................................................................................................................................. 21 Median ............................................................................................................................................... 21 Mode.................................................................................................................................................. 21 Dispersion/spread .............................................................................................................................. 22 Range ............................................................................................................................................. 22 Variance and Standard Deviation .................................................................................................. 22 Introduction to scientific writing Title ο· Concise, informative & clear ο· No jargon or abbreviation ο· Accurately reflects the content of the paper ο· Begin the title with the main subject of your paper ο· Use words that might be used to search for your title Abstract The abstract involves the following, in this order: ο· Context or background ο· Method ο· Main results ο· Discussion/conclusion ο· Implications ο· Prepare it after the rest of the writing is completed Introduction + literature review The introduction includes the following: ο· Overview/summary of previous research ο· Rationale for your ideas ο· Few studies on your topic? o Overcome limitations with your study ο· Conflicting results? o Shed more on the subject o Help resolve the controversy Explore, resolve or investigate in more detail or in a novel way! Methodology + measurement Methodology involves: ο· Details of how you conducted your research ο· What you did and how you did it ο· Sample ο· Design ο· Measurements ο· Procedure ο· Dry section; No big room for creativity Measurement involves: ο· Name of your variable (e.g., scale) ο· Operationalization ο· What was exactly asked to participants? ο· In which scale was it measured? ο· How reliable is this scale in your sample? Results ο· Report descriptive statistics ο· Central tendency and measures of variability ο· If you have several scales, report the correlations of interest ο· Report in text only the most important results; use tables to report numerical data ο· This section is to report, not to explain! Discussion and conclusion ο· Summarize your findings ο· Make sense of your results ο· Put your research into perspective ο· Cite literature that helps you explain your findings ο· Compare your data with that reported by previous authors; explain differences. ο· Get creative! Limitations & Further Research ο· What could have been done differently? ο· What are the potential shortcomings of your study? ο· How can future research address these shortcomings? ο· What other ideas do you think are worth following up? Research process The world is not constant, but variable. Nobody is the same. That is why predications have to be made based on data and models. Research question Do your predictions have support? Theory/ predictions Collect data Hypotheses Variables/measures Hypotheses Null hypotheses (H0): claim or statement being made/no effect on population o Null hypothesis (H0) assumes that there is no mean differences between groups or no association between variables Alternative hypotheses (H1): try to prove/effect on population o Alternative hypotheses (H1) assume that the differences between groups or the associations between variables exist and are not due to chance Statistically written: H0: ρ = 0; H1: ρ ≠ 0 Directional (one tailed) hypothesis H0 : ποΏ½1 = ποΏ½2 , H1 : ποΏ½1 > ποΏ½2 | H1 : ποΏ½1 < ποΏ½2 mean of the sample is only less than or only greater than the mean of the control group, but not both Non-directional (one tailed) hypothesis H0 : ποΏ½1 = ποΏ½2 , H1 : ποΏ½1 ≠ ποΏ½2 mean of the sample is different than the mean of the control group Experimental approach ο· Independent variable = the cause; its value does not depend on other variables ο· Dependent variable = the effect; its value depends on the cause ο· Causality, effect, influence Correlational approach ο· Predictor variable (independent variable) = a variable thought to predict an outcome ο· Outcome variable (dependent variable) = a variable thought to change as a function of changes in a predictor ο· Non-Causal, relationship, link Variables ο· Anything that varies and can be measured ο· Things that can change or vary between people, locations or even time ο· Most hypotheses are expressed in terms of two variables: o Proposed cause o Proposed outcome ο· Examples: o People: IQ, personality traits, preferences, emotions, behavior, gender etc. o Locations: unemployment, temperature, available resources, etc. o Time: mood, profit, pollution, cancerous cells ο· Dancing is an effective calorie burning Hypothesis Formulation a)… state the expected relationship between variables b)… make sure they are testable c)… simple and concise as possible d)… founded in the problem statement e)… supported by the literature f)… aligned with your research question ο· Every research question should have at least one corresponding hypothesis ο· Number of hypotheses needed is based upon the number of variables under investigation Errors Hypothesis Testing What we want: 1) Reject the null hypothesis when it is false 2) Fail to reject the null hypothesis when it is true What we DO NOT want: 1) Reject the null hypothesis when it is true 2) Fail to reject the null hypothesis when it is false ποΏ½ (alpha) represents the probability of making type I error o Rejecting the null hypothesis when it is true (significance level, p -value) ποΏ½ (Beta) represents the probability of making type II error o Fail to reject the null hypothesis when it is false (statistical power) For example, 80% of statistical power (ποΏ½) means that if the same study were to be repeated 100 times, there was a probability of replicating the same findings 80 times Type I error False positives ο· See things that do not exist ο· Consequences are serious: Associating phenomena that are unrelated Statistical Significance (ποΏ½) ο· Cutoff value representing the probability of the result occurrence if the null hypothesis is true ο· It is represented by the Greek letter alpha ο· Cutoff: 0.05 that is 1 in 20, or 5% o Merely a convention: One in 20 is rare enough to be trusted; but not so stringed that it is impossible to find o When p < .05 it is described as statistically significant with a 5% chance of making the wrong decision o Meaning: there is less than 1 in 20 probability that the results would have occurred if the null hypothesis were correct Type II error False negatives ο· Fail to see things that do exist ο· Consequences are less serious: Take longer time to associate phenomena that are related Statistical Power (ποΏ½) ο· Studies with low statistical power can lead to erroneous conclusions about the meaning of the results ο· To plan and to make diagnosis, results of power analysis help to: o Determine how large a sample should be o decide what criterion should be used to define “statistical significance” o Determine whether a specific study has adequate power for specific purposes o Identify the effects that can be reliably detected ο· Probability of rejecting a false null hypothesis (1-ποΏ½) ο probability of detecting an effect that is really there ο· The bigger the sample size, the bigger the power to detect an effect ο· Main output analysis: o Estimation of an appropriate sample size o Too big: waste of resources o Too small: may miss the effect o Research related: increasingly asked for justification Avoid the Errors Restrict the statistical significance ο· Significance level < .01 (1 in 100) or <.001 (1 in 1000) ο· Paradox: Increases the probability of Type II error, or of not detecting an effect that exists. ποΏ½ and ποΏ½ influence each other o Setting a lower ποΏ½ decreases Type I error risk, but increases Type II error risk o Increasing the power of a test decreases a Type II error risk but increases a Type I error risk Increase number of observations ο· If sample size is larger, tests have a higher statistical power ο· How many people do you need? ο· Paradox: Too many participants increase the Type I error, or detecting an effect that does not exist Type I and Type II relationship If Type I error is more serious than type II error, the rule is to treat Type I 4 times more seriously than Type II EXAMPLE: For a p-value of 5% (ποΏ½ = .05), ποΏ½ should be 4 times higher (ποΏ½ = ποΏ½*4 = .05*4 = .20) ποΏ½ = .20 Statistical power = 1-ποΏ½ statistical power = 1- ποΏ½ = 1- .20 = .80 (or 80%) 1-ποΏ½ = .80 EXAMPLE: For a p-value of 1% (ποΏ½ = .01), ποΏ½ should be 4% (ποΏ½ = (.01*4) = .04), with a corresponding statistical power of 96% (1-ποΏ½ = 1 - 0.04 = 0.96) Effect sizes ο· The result of something, a consequence, a reaction, a change ο· It is about magnitude ο· How big is the difference? ο· Although the effects are generally researched in the lab, the effect size exists in the “real world” Effect size as a d Concern differences between groups (e.g., Psych vs. Math students; treatment vs. control trials; males vs. females) Effect size as a r Concern association measures relating 2 or more variables with a correlation coefficient – it quantifies the relational strength and direction between two variables Methodological and measurement procedures – The higher the reliability of an instrument, the better Parametric tests have higher statistical power than nonparametric tests One-tailed hypotheses have more statistical power (vs. two-tailed) Exact: Correlation/Regression The H0 is that in the correlation (ποΏ½) between two variables has the fixed value of ποΏ½0 (ποΏ½ = ποΏ½0 ) The H1 is that the correlation coefficient has a different value (ποΏ½ ≠ ποΏ½0 ) Small ποΏ½ = 0.1, Medium ποΏ½ = 0.3, Large ποΏ½ = 0.5 F test: One Way ANOVA The H0 is that all k means are identical The H1 is that at least two of the k means differ Small f = 0.10, Medium f = 0.25, Large f = 0.40 T test: Mean differences between 2 independent groups The H0 : ποΏ½1 - ποΏ½2 = 0 The H1 : ποΏ½1 - ποΏ½2 ≠ 0 Small d = 0.2, Medium d = 0.5, Large d = 0.8 Normal Distribution Skewness Non-symmetrical distribution = Skewed Positive skew → pile-up of scores on the left of the distribution Negative skew → pile-up of scores on the right of the distribution The further the value is from zero, the more likely it is that the data are not normally distributed. ο· Values are considered normal if they are between -3 and 3 Ceiling effect ο· A ceiling effect associated with statistics in social sciences refers to the phenomenon in which the majority of the data are close to the upper limit or highest possible score of a test. This means that (almost) all of the test participants achieved the highest (or very near to the highest) score (https://www.scribbr.com/researchbias/ceiling-effect/). Floor effect A floor effect occurs when a high proportion of individuals endorse the minimum score on the observed variable (https://taylorandfrancis.com/knowledge/Medicine_and_healthcare/Psychiatry/Floor_effect/#: ~:text=A%20floor%20effect%20occurs%20when,score%20on%20the%20observed%20varia ble.) Kurtosis Too many people at the extremes OR not enough people at the extremes Negatively kurtosed → too many people in the tails (flat and light-tailed distribution) Positively kurtosed → insufficient people in the tails (pointy and heavy-tailed distribution) The further the value is from zero, the more likely it is that the data are not normally distributed. ο· Values are considered normal if they are between -3 and 3 Outliers ο· Data points that lie outside the distribution ο· They bias estimates of parameters (e.g., mean) ο· Easy to spot them in a histogram or boxplot How to deal with outliers? ο· Check if it is not a mistake on data entry, equipment malfunction, or similar Trimming: deleting extreme scores, but only if you believe that these cases are not from the population you intend to sample ο· Trimmed mean o Percentage based rule: deleting the e.g., 5%, 10%, 15%... Of highest and lowest scores in the data o Standard-deviation based rule: deleting the highest and lowest values that are e.g., 2SD, 3SDs away from the mean ο· M-estimator o Weighted averages with heavier weight to the observations close to the median and less weight to the observations in the tails o The amount of trimming is determined empirically; M-estimator determines the optimal amount of trimming necessary to give a robust estimate of the parameter of interest (e.g., the mean) Winsorizing: replace the outlier with the next highest score that is not an outlier Analyze with robust methods: bootstrapping ο· Bootstap o Takes random samples from the data you collected o Simulates the calculation of the parameter of interest (e.g., mean) in each bootstrap sample o Repeats the process several times (1000, 2000, 5000, 10.000…) o End-result: several parameters estimates (e.g., 1000), one from each bootstrap sample o Use these values to calculate the 95% confidence intervals of the parameters Transform the data: apply a mathematical function to the original scores Population ο· It can be any group of people (Females, Students, Extroverted, Unemployed…) ο· Study everyone is impossible and needless – subset of these people ο· Findings that are applicable to a large numbers of people ο· Generalize them to the population – Or making inferences Probability ο· All comes down to: o Mean and standard deviation o Find the prob of any value o Use the area under the curve ο· π‘οΏ½ = standardizing a score (z-score) o Number of standard deviations above and below mean the Sampling Distribution Samples and possibilities ο· Measure a sample of people to find the estimate of a value in the population ο· Sample from the population is likely to have a value close to the ‘true’ mean ο· Randomly chosen from the population ο should be selected through randomisation (everyone same change to be selected for sample) ο· Unlikely scenarios: 1. Get the ‘true’ population mean 2. Be too far from the ‘true’ population mean ο· We are as likely to underestimate the population mean as we are to overestimate the population mean ο· Distribution is symmetrical and mean is unbiased ο· Distribution = Sampling distribution of the mean (normal) Standard Error of the Mean ο· Measures how much a sample mean is likely to vary from the true population mean ο· It is essentially a measure of the precision or reliability of the sample mean as an estimate of the population mean ο· Sample Mean: When you take a sample from a population and calculate the mean of that sample, it's an estimate of the population mean ο· Standard Error of the Mean: The SEM tells you how much your sample mean might differ from the true population mean if you were to take many different samples from the same population: ο· The larger the standard error of the mean, the less confident you can be that your sample mean accurately reflects the true population mean. A smaller standard error indicates that your sample mean is likely to be close to the true population mean. Plan Sample Size ο· Aim for 80% of statistical power ο· According to Cohen (1988) it is the level representing 20% of probability of making type II errors o 50% of statistical power means 50-50 chances of correctly detecting an effect o 90% of statistical power means higher chances of correctly detecting the effect, but means a very high N t-Distribution If your sample is not large, the distribution is not normal, and it should follow the t distribution. But this t distribution is very closely related to the normal distribution. π‘οΏ½ = π ππππ−ππππ π π Degrees of Freedom Number of values free to vary Used to make inferences about population parameters based on sample data When calculating the mean of a sample of 30 numbers, the first 29 are free to vary but the 30th number would be determined as the value needed to achieve the given sample mean Degrees of freedom ο specific t-distribution used to calculate p-values, t-values and t-tests When t distribution has an infinity degree of freedom, there will be a z-distribution. Depending on the statistical significance, there is a critical t value for each degree of freedom. Besides, take a look if there is a one tail hypothesis or a two tails hypothesis. Confidence Intervals More than knowing a pre-specified value, we want to know how large a value is likely to be in the population - Confidence Intervals: range of the population value - Confidence Limits contain the largest and smallest values in the interval The most common is to calculate 95% confidence limits - Matches the significance level used in hypothesis testing (5%, p-value < .05) πΆοΏ½πΌοΏ½ = xΜ ± π‘οΏ½πΌοΏ½ × π οΏ½ποΏ½οΏ½with Lower Confidence Limit (LCL) and Upper Confidence Limit (UCL) Interpretation: 95% of studies, the true (population) mean will be contained within the confidence limits 5% of studies the true value is not contained (H0 is true) If the confidence intervals contain H0 , the result is not statistically significant ο· If 0 does fall inside the interval (between LCI and UCI) it is not statistically significant. EXAMPLE [-2.34; 6.45] (p > .05) ο not statistically significant. If the confidence intervals DO NOT contain H0 , the result is statistically significant (p <.05) ο· If 0 does not fall inside the interval (between LCI and UCI) it is statistically significant. Measures of Central Tendency Mean Mean, average, arithmetic mean Is computed by adding up all scores and dividing by the number of individual scores ASSUMPTIONS The distribution is symmetrical: data are not skewed and no outliers Data measured interval or ratio level: you cannot calculate the mean of gender or hair color Median ο· Second most common measure of central tendency ο· Represents the middle score in a set of scores ο· It can be used when the mean is not valid (data not symmetric or not normally distributed or measured at an ordinal level) ο· Scores must be placed in ascending order of size, from smallest to largest UNEVEN NUMBER OF SCORES Find the middle score of the distribution and take the next whole number; that value corresponds π+1 to the median: ππ = οΏ½ 2 EVEN NUMBER OF SCORES Find the two position scores that are placed in the middle; find the average of the two values to get the median: ππ = οΏ½ π 2 π( )+π₯( π+1 ) 2 2 Mode Not very common in quantitative psychological research The most frequent score in the distribution It is the best measure of central tendency for categorical data Dispersion/spread ο· To understand central tendency, a measure of dispersion or spread is needed ο· Reporting the mean without reporting its dispersion measure is useless ο· Takes all the values in the dataset into account ο· ASSUMPTION: data normally distributed ο· Range, Variance and Standard Deviation are measures of dispersion Range ο· Range is the simplest measure of dispersion: o Distance between the highest and the lowest score o Range = βποΏ½ποΏ½βποΏ½π οΏ½π‘οΏ½ π£οΏ½ποΏ½ποΏ½π’οΏ½ποΏ½ − ποΏ½ποΏ½π€οΏ½ποΏ½π οΏ½π‘οΏ½ π£οΏ½ποΏ½ποΏ½π’οΏ½e ο· It can be a single number or expressed as the highest and lowest score ο· Hugely affected by any outliers Variance and Standard Deviation The extent to which every score differs from the mean score The average amount of deviation from the mean (ignoring the signs) is known as the mean deviation Variance (ποΏ½2 ) is calculated by squaring each deviation from the mean before summing up the the total: Standard Deviation (ποΏ½) is the square root of variance (most common measure): Correlation - Correlation is a statistical measure that expresses the extent to which two variables are linearly related - For example, the height and weight of a person are related, and taller people tend to be heavier than shorter people - One can expect, for example, that a child's IQ is related to his/her academic performance Pearson Correlation r ο· Reflects the linear association/relation/link between two variables ο· It is expressed as an r of Pearson ο· Varies between +1 (perfect positive linear relation) and -1 (perfect negative linear relation) ο· It is commonly used to test association hypotheses Strength: ο· How good is the relation (values closer to 1 or -1 indicate better relations) - r = 0.10-0.30 ο Small/weak effect - r = 0.30-0.50 ο Moderate effect - r > 0.50 ο Large effect Direction: ο· Positive or negative relation o A positive r coefficient means that the variables are associated in the same direction ο§ Taller people have larger shoe sizes ο§ The longer your hair grows, the more shampoo you will need o A negative r coefficient means that the variables are associated in the opposite direction ο§ The more one studies, the less free time one has ο§ If a car decreases in speed, travel time to destination increases ο§ The r is negative, means that when the independent variable increases one unit, the dependent decreases in the same amount ππ Coefficient of Determination ο Percentage of variance explained by the association/relation between the two variables o It is represented by π 2 because it is the value of r squared o Varies between 0 and 1 o The closer to 1 (0), the greater (lesser) percentage of explained variance o Ex.: r = 0.40 is 16% of variance explained AND r = 0.80 is 64% of variance explained Point Biserial Correlation ο· Quantifies the relation between a continuous variable and a discrete dichotomous variable ο· There is no continuum underlying the two categories, such as being dead or alive ο· Used when one of the variables is dichotomous (categorical with only two levels) ο· (rpb) Ordinal Data ο· ο· ο· Non-parametric tests Spearman: useful for ranked data ο§ Useful to minimize the effects of extreme values ο§ Useful when two variables do not assume the same scale Kendall’s tau-b: Similar to Spearman’s coefficient, however, applied to small sample sizes Pearson’s Chi Square (test) ο· Examines whether there is an association between two categorical variables ο· Based on the comparison of frequencies observed in certain categories to the frequencies expected to get in those categories by chance ο· If the p-value is small (significant), we reject the hypothesis that the variables are independent and gain confidence assuming they are somehow related The fact that counts have different subscript letters (a & b) means that the column proportions are significantly different Conclusion: “The proportion of cats that danced after affection was significantly less than the proportion that did not dance after affection” First line indicates the Pearson’s ChiSquare test and its associated p-value Symmetric Measures (Phi and Cramer’s V) are measures of association between variables, it can be interpreted as a correlation In terms of magnitude, it would be a moderate association or a moderate effect size Odds-ratio are similar to the effect size measure. Example: odds ratio of 6.67 means that the likelihood of dancing after food is 6.67 times higher than after affection Correlation and regression in one picture: Regression (analysis) ο§ Regression analysis is a statistical test used to fit a linear model to your data ο§ Predicts values of an outcome variable from one or more predictor variables ο§ Statistical analysis used to predict or explain an outcome variable (continuous) o One predictor: Simple linear regression (ποΏ½ποΏ½ = ποΏ½0 + ποΏ½1ποΏ½1 + επ ) o Two or more predictors: Multiple linear regression (ποΏ½ποΏ½ = ποΏ½0 + ποΏ½1ποΏ½1 + ποΏ½2ποΏ½2 + επ ) ο§ Mathematical model associating the values of an independent variable to the values of an outcome/dependent variable ο§ Find the slope that best fit our data Model Fit ο· Variables: measured constructs that vary across entities in the sample ο· Parameters: estimated from the data (not measured) and are normally constants believed to represent a relation between variables in the model (e.g., mean, median, coeff. Regression) Outcome is the dependent variable, Model includes the independent variables. It is like f(x)=ax+b a math function: The b0 is the coefficient, with + giving a line going up and – line going down ο§ An outcome (ποΏ½ποΏ½ ) is predicted by a model, with some error ο§ The outcome (ποΏ½ποΏ½ ) is predicted using a predictor (ποΏ½1 ) and a parameter (b1) ο§ The parameter (ποΏ½1) shows the relationship between ποΏ½1and ποΏ½ποΏ½ ο§ Another parameter (ποΏ½0) tells us the outcome's value when the predictor is zero Statistical Models ο· Models predict outcome variables ο· Parameters inform about the shape/form of the model ο· To understand a model, we have to estimate these parameters (or the value of b) ο· The form of the models can change ο· You always have to account for the error Error ο· The error, or the deviance for a particular entity, is the score predicted by the model for that person, subtracted from the observed score for that entity ο· We can use SSE (sum of squared error) and MSE (mean of squared error) to assess the fit of a model ο· Larger values of SSE or MSE indicate a poorer fit – larger dispersion from the mean Parameter Estimates ο· Hypothetical value used to summarize the data ο· We use the mean computed in our sample to estimate the value in the population (parameter estimate) Regression Coefficients ο· b0 : Estimates of the unknown population parameter (e.g., mean) that indicate the numeric value of the outcome, in the absence of predictor(s) ο· In a plot, it represents the point at which the line meets the vertical axis ο This is also called the intercept ο· Imagine that the outcome is the number of cases diagnosed with lung cancer associated with smoking. The predictor b0 concerns how many people are diagnosed with lung cancer who never smoked ο Outcome = b0 + b1X1 + errori ο Yi = b0 + b1X1 + εi Least Squares Method ο· Model fit: deviations between the model and the actual data ο· Deviations (errors) are the distances between the model predicted (e.g., the line) and each data point ο· In a perfect fit, it would predict the exact same value of the observed outcome (no deviation) ο· Models are rarely perfect there is always some error in our predictions (overestimation or underestimation) Total Error: ο· Deviations between the model (e.g., the line) and the actual data are called residuals ο· To calculate the total error you sum all the squares of these deviations; this is called the Sum of Squared Residuals OR Residual Sum of Squares (SSR) ο· When squared differences are small, the line has a better fit to the model, it represents better the actual data ο This method is known as Ordinary Least Squares (OLS) Sum of Squares (Goodness of Fit) ο· Goodness of fit is important to understand how much our selected model represents the reality ο· Some models may be selected as the best models, but they might have a poor fit to the data ο· SSR helps us to have an idea of the deviation of the data ο π 2 : Represents the amount of variance in the outcome explained by the model o If you take the square root of this value, you obtain the Pearson’s correlation coefficient between the predicted and observed values o Correlation coefficient is a good estimate of the overall fit of the model o π 2 οΏ½tells us about the size of this model fit ο explain 19% (πΉπ οΏ½= .19) ο F Ratio: F is the amount of systematic variance (the model) divided by the amount of unsystematic variance (the error) o Ratio of the improvement of the fit due the model (SSM) and the difference between the model and the observed data (SSR) o Large F-ratio means that a model is good o The F-ratio should be at least greater than 1 o The F-statistic can be used to calculate the significance of the π 2 οΏ½ ο§ tests the null hypothesis that π 2 = 0 o Sum of squares: associated to the three sources of variance (total, regression model and residual) o Total variance: sum of the variance that is explained by the independent variable (regression) + error o df: degrees of freedom associated to the sources of variance o Mean Square: Sum of squares ÷ respective df o F & sig. – F-value and p-value associated to the regression model. Tests the H1 vs. H0 (where all the coefficients of the model are zero). F(1, 101) = 24.38, p < .001 Individual Predictors ο· All predictors have a coefficient (b1) that represents the gradient/slope of the regression line ο· Values of b = change in the outcome resulting from a one-unit increase in the predictor, holding all the other predictors constant ο· If a variable significantly predicts an outcome, the coefficient (b) should be different from zero → the hypothesis that each predictor is significant is tested using a t-test variance (ß = -.44, p < .001). ο· Predictors: first row is the constant; it is the predictor of the exam score when anxiety = 0 ο· Unstandardized B: regression equation to predict the dependent variable through the independent variable ο· Std error: standard error of each coefficient ο· Beta (ß): standardized coefficient that allows comparing coefficients of different variables ο· t e sig.: tests two-tailed if each of the coefficients are statistically different from zero ο· When anxiety = 0 exam score = 106.07 (b0= 106.071) For each unit of anxiety that increases, the exam score decreases 0,67 (b1 = --0,666) REPORTING A simple regression analysis was conducted to test whether anxiety predicts exam scores of students. Results showed that the regression model was statistically significant, F(1, 101) = 24.38, p < .001. Proportionally, the anxiety levels explain 19% (πΉπ οΏ½= .19) of the exam score variance (ß = -.44, p < .001). t-statistic ο· Ratio of explained to unexplained variance ο· Specifically, it tests if the coefficient is big when compared to the amount of error in that estimate ο· Uses the standard error (SE) to indicate how different would the b-values be across different samples o Small SE = little variation, meaning that most samples are likely to have a similar bvalue ο· t-test checks if the b-value is different from zero relative to the variation in b-values across samples
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )