Data Visualization, 2e Chapter 9: Telling the Truth with Data Visualization Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 1 Chapter Objectives (1 of 2) After completing this chapter, you will be able to: LO 9.1 Identify missing data and data errors using Excel. LO 9.2 Define the meaning of biased data and explain the concepts of selection bias and survivor bias. LO 9.3 Define Simpson’s paradox and explain how a scatter chart can be used to identify some instances of Simpson’s paradox. LO 9.4 Explain the importance of adjusting for inflation in time series data that represent long time periods and use a price index to adjust nominal values to account for inflation. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 2 Chapter Objectives (2 of 2) LO 9.5 Identify deceptive design practices related to the axes used in charts and suggest ways to improve the axes in these charts to communicate insights to the audience more clearly. LO 9.6 Explain why dual-axis charts are often confusing and misleading for audiences and suggest an alternative to using a dual-axis chart that is less confusing for the audience. LO 9.7 Explain how the range of data and the temporal frequency of data included in a chart affects the insights conveyed to the audience. LO 9.8 Explain why some geographic maps can result in misleading data visualizations and provide recommendations for how to improve these types of maps. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 3 Data Visualization Makeover Scatter charts showing the effect of inflation adjustment on gross revenues for the lifetime North America box office top-10 movies of all time In 1951, the average price of a movie ticket was $0.53. In 2020, it was $9.37, or 17.7 times more expensive. Source: https://help.imdb.com/article/imdbpro/industry-research/box-office-mojo-by-imdbprofaq/GCWTV4MQKGWRAUAP?ref_=mojo_cso_md#inflation. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 4 9.1 Missing Data and Data Errors A portion of an Excel spreadsheet showing Blakely Tires data Data file: BlakelyTires Data set: it consists of information about the four tires from 116 automobiles. Company: Blakely Tires is a producer of automobile tires located in the U.S. Issue: assess data quality. Study objective: to learn about the conditions of its tires on automobiles in Texas. Reported variables: tire position, age, tread depth, and mileage. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 5 9.1 Identifying Missing Data by Counting Data Blanks A portion of an Excel spreadsheet showing number of missing values for variables in Blakely Tires data To count the missing observations for each variable in the file BlakelyTires, follow the step-by-step instructions included in the notes. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 6 9.1 Identifying Missing Data Using Conditional Formatting A portion of an Excel spreadsheet showing Blakely Tires data sorted from lowest to highest by ID number and missing value highlighted To highlight the missing observations for each variable in the file BlakelyTires, using the Excel Conditional Formatting tool, follow the step-bystep instructions included in the notes. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 7 9.1 Identifying Extreme Data Errors A portion of an Excel spreadsheet showing Blakely Tires data sorted on Life of Tires (Months) and calculated summary statistics We can use summary statistics to detect suspicious values. • Life of Tire: a maximum value of 601 months for a left rear tire. Comparison with the other tires suggests a value of 60.1 months instead. • Tread Depth: a minimum value of 0.0/12ths (undrivable) and a maximum value of 16.7/12ths, which exceeds tread depth on new tires. *See notes for step-by-step instructions on how to calculate summary statistics. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 8 9.1 Identifying Non-Extreme Data Errors Scatter chart for the Blakely Tires data used to identify possible data errors Not all erroneous values in a data set are extreme. To identify non-extreme erroneous values, we can investigate relationships between quantitative variables using data visualization tools such as scatter charts. Here, tire mileage has a strong negative relationship with tread depth because the tire tread slowly erodes while driving. See the notes for additional information about identifying potential non-extreme data errors in the data set for further analysis by exploring data points lying outside of the expected red ellipse. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 9 9.2 Biased Data Biased data occurs when representing an intended population by drawing data from a non-random sample. There are two many types of bias: • Selection bias occurs when the data drawn from a non-random sample to represent the intended population. • Example of selection bias: political polling conducted on landline phone numbers may result in a sample biased in terms of age of the respondents. • Simpson’s paradox is a type of selection bias. It occurs when a specific trend that appears in subsets of data disappears or reverses when the subsets are aggregated. • Survivor bias occurs when sample data consists of a disproportionately large number of observations corresponding to outcomes for a particular event. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 10 9.2 Example of Simpson’s Paradox A data portion used to examine the relationship between age and income Data file: AgeIncome Study objective: to analyze factors that affect annual incomes in the U.S. Reported variables: age, income, and location (Dallas, San Francisco, and Naples) for 106 respondents. Bias issue: the three locations are not necessarily representative of the U.S. population. Approach: build a scatter chart of the data aggregated by location and compare it to scatter charts in which the data were disaggregated by location. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 11 9.2 Visualization of Simpson’s Paradox Scatter charts illustrating how Simpson’s paradox may lead to erroneous conclusions when considering the relationship between age and income on aggregated data Simpson’s paradox: the trend that appears in the aggregated data reverses when the locations are plotted separately. Here, age and income display • a negative relationship when aggregated. • a positive relationships when disaggregated by location. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 12 9.2 Origin of Survivor Bias Depiction of common areas of damage to Just as for selection bias, with survivor bias, the sample data are not represent-tative of the returning planes in World War II population under study. Examined data on surviving World War II planes showed consistent damage to the indicated areas (X). The initial suggested action was to armor tails and wings. However, the mathematician Abraham Wald successfully argued that the correct action was the opposite: because damage to tails and wings were least harmful to the aircraft, armor should be added instead to areas not showing damage. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 13 9.2 Example of Survivor Bias Comparing average risk tolerance of entrepreneurs to a control group of nonentrepreneurs Research hypothesis: entrepreneurs are more likely to have a greater risk tolerance than non-entrepreneurs. Data set: results of a survey that measures the tolerance to take on financial risk for 87 entrepreneurs, compared to a control group of 100 randomly selected nonentrepreneurs. The conclusions are affected by survival bias, as the only data we have on entrepreneurs is successful ones. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 14 9.3 Adjusting for Inflation: Example A portion of data with nominal prices for a gallon of gasoline and price index in the United States between 1978 and 2017 A price index is the aggregated price of a basket of products and services. Inflation refers to the general increase in nominal prices over time for the price index. Data file: PriceGasoline Study objective: to analyze the effect of inflation on gas prices between 1978 and 2017. Reported variables: nominal gas price per gallon and price index. Calculated variable: inflation-adjusted gas price per gallon. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 15 9.3 Adjusting for Inflation: Calculations Adjusting Nominal gasoline prices for Inflation in Excel Inflation-adjusted prices are calculated for a base year. In this example, we used 2017: Infl.adjusted G as Price 1978 Price Index 2017 G as Price 1978 Price Index 1978 211.8 Infl.adjusted G as Price 1978 $0.65 $2.66 51.9 In cell D2 of the Excel spreadsheet, we can implement the above equation as: =B2*($C$41/C2) And paste it to cells D3:D41. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 16 9.3 Adjusting for Inflation: Visualization The inflation-adjusted price per gallon of gasoline in the United States has decreased between 1978 and 2017 We can adjust nominal values to account for inflation using any year in the time series. The selection of the base year generally depends on the comparisons to be made Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 17 9.4 Improper Chart Axes: Exaggerated Difference A column chart that exaggerates the difference in the proportion of likely voters Redesigned column chart showing just a small difference in the proportion of likely voters Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 18 9.4 Improper Chart Axes: Hidden Insight A line chart with a wide range showing no increase in the earth’s surface temperature A line chart with a revised range showing an increase in the earth's surface temperature *See notes for additional information on deceptive chart designs. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 19 9.4 Improper Chart Axes: Incorrect Aspect Ratio Two different line charts for the average annual global surface temperature data showing the effect of line chart aspect ratio on the insights conveyed to the audience The aspect ratio refers to the width to height ratio of a chart. The line chart in (a), tall but not very wide, has a smaller aspect ratio than the line chart in (b), which is short but wide. *See notes for additional information on the aspect ratio of charts. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 20 9.4 Dual-Axis Charts: Deceptive Design Effect of using a dual-axis chart to display the gross domestic product and unemployment rate in the United States since 2000 A dual-axis chart uses a secondary axis to represent one of the variables to show on the same chart the different vertical ranges. Dual-axis charts may be difficult to interpret and may mislead the audience, such as by causing to erroneously interpret the points in which the lines appear to intersect. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 21 9.4 Dual-Axis Charts: Remedy Using two different charts to display the gross domestic product and unemployment rate in the United States since 2000 A simpler alternative to a dual-axis chart consists of using two separate charts with each variable shown individually. *See notes for instructions on how to build dual-axis charts. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 22 9.4 Data Selection and Temporal Frequency Comparing stock prices with the date range but different temporal frequencies The temporal frequency refers to the rate at which time-series data are plotted: a) Share price plotted with a daily temporal frequency shows peaks and valleys. b) Share price plotted with a bi-monthly temporal frequency appears smoothed. Viewing time series data at different temporal frequencies can reveal different patterns. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 23 9.4 Issues Related to Geographic Maps Choropleth map showing number of people living in poverty by state Choropleth map showing poverty rate by state The poverty rate by state is a better choropleth map for communicating which states have the highest number of people living in poverty relative to the overall state population. *See notes on how to calculate poverty rate by state and adding data labels to maps. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 24 Discussion Activity • Consider these three studies: 1. In the 1930s, Dr. Joseph Rhine tested 500 participants to guess the order in a shuffled deck of cards. Only those participants who guessed correctly moved to the next round. The last participant to guess right every round was declared to have telepathic abilities. 2. A policy group lobbying for the alcohol industry conducted an exit poll outside of a popular restaurant late at night. The restaurant patrons were asked whether driving after drinking was a serious problem. The policy group concluded that driving after drinking was not a serious problem. 3. A research group followed for 40 years a cohort of individuals. The study’s main conclusion was that those who drank red wine were richer, better educated, had lower levels of cholesterol ratio, and consequently a lower risk of heart attacks than those who did not. • Associate each of the three studies to one of the following forms of bias: a. Selection bias b. Simpson’s Paradox c. Survivor Bias Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 25 Check Your Knowledge • Due to their cause and effect on the insights drawn from data, outliers should __________. a. only be removed after careful consideration b. immediately be removed for accurate analysis c. never be removed d. be replaced with incorrect data • Which of these statements describes how survivor bias and selection bias are related? a. The sample data set consists of a disproportionately large number of observations corresponding to positive outcomes for a particular event. b. The sample data are not representative of the population that is being studied. c. The sample has not been properly randomized to represent the intended population. d. They are not related in any way. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 26 Summary In this chapter, you should have learned how: • To identify missing data and data errors using different Excel methods that depend on the nature of the type of data and problem settings being analyzed. • Biased data may mislead insight from data visualization and how to identify it. • To adjust for inflation when dealing with time series data that have been collected over long time periods. • Different ranges of values for the axes can greatly affect the insights conveyed to the audience for the same data. • To remedy dual-axis charts that may often be misleading for the audience. • The selection and frequency of time series data can change the insights conveyed to the audience. • To prevent misleading the audience when using geographic maps. Camm, Cochran, Fry & Ohlmann, Data Visualization- Exploring and Explaining with Data, 2nd Edition. © 2025 Cengage. All Rights Reserved. May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part. 27
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )