Business Statistics - Quiz Prep
Notes
Business Statistics — Data
Organization & Visualization
Quiz Prep Notes (Prof. Alok Kumar Singh, IIM
Nagpur)
1. What is Statistics?
Statistics = a way of thinking that leads to better decisions.
It is the science of gathering, presenting, analyzing, and
interpreting data.
Uses mathematics and probability.
Requires analytical skills — core part of business education.
Modern IT lets businesses apply statistics to large data using
analytical tools.
Business statistics helps to: 1. Summarize and visualize business
data 2. Reach conclusions from business data 3. Make reliable
predictions about business activities 4. Improve business processes
Cartoon takeaway from the slides: “AI” in business is often
really just statistics + a bit of machine learning — a
reminder that stats underlies most analytics/ML work.
2. Basic Definitions (memorize exactly —
these are classic quiz fodder)
Term
Definition
Variable
A characteristic of an item or
individual (e.g., Gender)
Data
Set of individual values
associated with one or more
variables (e.g., Gender, Age,
Qualification, Earning)
Population
The whole collection of all
persons, objects, or items under
study
Sample
A subset of the population
Census
Gathering data from the entire
population
Statistic
A value that summarizes data of
a variable for a sample
Descriptive Statistics
Methods that summarize and
present data
Inferential Statistics
Uses data from a sample to
reach conclusions about a larger
group (population)
⚠ Common trap: “Statistic” (summary of a sample) vs. “Parameter”
(summary of a population) — don’t mix them up.
3. Parameter vs. Statistic
Parameter
Statistic
Describes
Population
Sample
Notation
Greek letters Roman letters
Mean
μ (mu)
x̄ (x-bar)
Variance
σ²
s²
Std. Deviation σ
s
Mnemonic: Population → Parameter → Greek. Sample → Statistic →
Roman.
4. Process of Inferential Statistics (4-step
cycle — could appear as a sequencing
question)
1. Population has true parameter μ (unknown)
2. Select a random sample from the population
3. Compute the sample statistic x̄
4. Use x̄ to estimate μ
This is the whole logic of inference: we can’t measure everyone, so we
estimate the population parameter from a sample statistic.
5. Classifying Variables by Type
Categorical (Qualitative): values are categories → “yes/no”,
“blue/brown/green”, “Easy/Normal/Tough”
Numerical (Quantitative): values represent counted or measured
quantities - Discrete — arises from a counting process (whole
numbers, countable) - Continuous — arises from a measuring
process (can take any value in a range)
Levels of Qualitative Data Measurement
Nominal - Numbers/labels just name the attribute — no order
implied - Examples: Gender (M/F), Religion
(Hindu/Muslim/Sikh/Christian/Jain), Employment classification codes
(1=Educator, 2=Construction Worker, 3=Manufacturing Worker)
Ordinal - Ranking is possible, but differences between values
are NOT meaningful/comparable - Classic example: Olympic
medals — Gold > Silver > Bronze, but Gold + Bronze ≠ 2×Silver Organizational position ranks (President=1 … Employee=5) Preference/Likert scales (Like it a lot → Dislike it a lot)
Levels of Quantitative Data Measurement
Discrete — countable values only - e.g., Number of customers
entering a restaurant, Number of defects in a product
Continuous — can take any value (measured) - e.g., Temperature,
Weight of a chips packet, Time taken to travel
Full Hierarchy (this diagram is a favorite quiz source)
Variables
/
\
Categorical
Numerical (Interval/Ratio)
/
\
/
\
Nominal
Ordinal
Discrete
Continuous
(Marital
(Ratings:
(Number of
(Weight,
Status,
Good/Better/
Children,
Voltage)
Political
Best; Low/
Defects/hr)
Party,
Med/High)
Eye Color)
⚠ Practice this classification skill — quizzes love giving you a
scenario (“Do you have a Tinder account?”, “How many WhatsApp
messages sent in last hour?”, “Colour of your eyes?”, “Rate the new
Netflix series”, “Age in days”, “Temperature today”) and asking you to
identify: Categorical/Numerical →
Nominal/Ordinal/Discrete/Continuous.
Quick self-test (from the slide’s example table) — classify each: 1. Do
you have a Tinder account? → Categorical – Nominal (Yes/No) 2.
WhatsApp messages sent in past hour? → Numerical – Discrete
(countable) 3. Time taken for app update to download? → Numerical
– Continuous (measured) 4. Colour of your eyes? → Categorical –
Nominal 5. Your weight? → Numerical – Continuous 6. Rating of a
Netflix series (poor/average/good)? → Categorical – Ordinal 7.
Gender? → Categorical – Nominal 8. Age (in days)? → Numerical –
Discrete (counted in whole days) — note: age is often treated as
continuous when measured precisely, but here it’s counted in wholeday units, so watch how your instructor frames it 9. Temperature in °C
today? → Numerical – Continuous
6. Sources of Data
Type
Description
Examples
Primary
The data collector
uses the data
themselves
Political survey data,
experiment data,
observed data
Secondary
Analyst is not the
original data
collector
Census data analysis,
print journal data,
internet-published
data
7. Organizing Data
Organization of Categorical Data
Categorical Data
├── One Categorical Variable
└── Two/More Categorical Vars
→ Summary Table
→ Contingency Table
Organization of Numerical Data
Numerical Data
├── Ordered Array
├── Frequency Distribution
└── Cumulative Distribution
8. Visualization of Categorical Data
Data Structure
Chart(s)
Summary Table (1 variable)
Bar Chart, Pie Chart, Pareto
Chart
Contingency Table (2 variables)
Side-by-Side Bar Chart
Pareto Chart: a bar chart with categories sorted by frequency
(descending) + a cumulative % line — useful to identify the “vital
few” causes (80/20 rule).
9. Visualization of Numerical Data
One Variable (built from Frequency/Cumulative Distributions): Histogram — bars for frequency distribution of a continuous variable
- Polygon — line version of a frequency distribution - Ogive — line
graph of a cumulative distribution
Two Variables: - Scatter Plot — relationship between two numerical
variables - Time Series (Line Chart) — a numerical variable plotted
over time
10. Organizing Many Variables — Pivot
Table/Chart
Summarizes variables as a multidimensional summary table
Allows interactive changes to level of summarization/formatting
Allows “slicing” data to see subsets meeting specific criteria
Helps discover patterns/relationships that simple tables/charts
miss
11. Best Practices for Constructing
Visualizations (very quizzable — list these
exactly)
1. Use the simplest possible visualization
2. Include a title and label all axes
3. Include a scale for each axis if the chart has axes
4. Begin the vertical axis scale at zero and use a constant scale
5. Avoid 3D or “exploded” effects
6. Use consistent colors across charts meant for comparison
7. Avoid uncommon chart types: radar, surface, bubble, cone,
pyramid charts
Practice Quiz Questions
A. Definitions / Concept Recall
1. Define “Statistics” as presented in the session, and name the four
things business statistics allows an organization to do.
2. Distinguish between a population and a sample with an example.
3. What is the difference between a census and a sample survey?
4. Define descriptive statistics and inferential statistics. Give one
example of each in a business context.
5. Differentiate between a parameter and a statistic. Which uses
Greek letters and which uses Roman letters?
6. Outline the four-step process of inferential statistics, starting from
the population.
B. Variable Classification (apply the concept — most
likely quiz format)
7. Classify each of the following as Categorical (Nominal/Ordinal) or
Numerical (Discrete/Continuous):
a. Number of defective items in a shipment
b. Customer satisfaction rating (Poor/Fair/Good/Excellent)
c. Monthly salary of an employee
d. Blood group of a patient
e. Number of calls received by a call center in an hour
f. Time taken to complete an online transaction
8. Why is “Olympic medal type” (Gold/Silver/Bronze) an ordinal
variable and not an interval variable? Explain using the averaging
argument from class.
9. Give one nominal and one ordinal example from an
HR/organizational context (not from the slides).
10. Explain the difference between discrete and continuous numerical
data with two original examples.
C. Data Organization & Visualization Matching
11. Match the data type to the correct summarization tool: (i) One
categorical variable (ii) Two categorical variables (iii) One
numerical variable — [Summary Table / Contingency Table /
Frequency Distribution]
12. Which chart(s) are appropriate for visualizing one categorical
variable? Which for two?
13. What is a Pareto chart, and why is it useful in quality/business
analysis?
14. Differentiate between a Histogram, a Polygon, and an Ogive.
15. Which visualization would you use to show the relationship
between advertising spend and sales revenue? Why?
16. Which visualization is best suited to show a company’s monthly
revenue trend over the last 3 years?
17. What is a Pivot Table/Chart, and what advantage does it offer over
a standard summary table when working with many variables?
D. Best Practices / Applied Judgement
18. List any five best practices for constructing effective business data
visualizations.
19. Why should the vertical axis of a bar/column chart usually start at
zero? What problem occurs if it doesn’t?
20. Why are 3D and “exploded” chart effects discouraged in
professional reporting?
21. Name three chart types considered “uncommon” and generally
discouraged in business reporting.
E. Short Case/Application (typical MBA-style twist)
22. A retail chain wants to study “customer preference for payment
mode: Cash / Card / UPI / Wallet” across 5 store locations. What
data organization tool and chart would you recommend, and why?
23. An analyst has sample data on 200 customers’ ages and wants to
both see the shape of the distribution and the cumulative % of
customers below a certain age. Name the two charts they should
draw.
24. A company collects data from its own customer satisfaction survey
vs. downloading government census data for market sizing.
Classify each data source as primary or secondary.
Quick-Fire Answer Key Hints (self-check before your
quiz)
Q7: a) Discrete, b) Ordinal, c) Continuous, d) Nominal, e) Discrete,
f) Continuous
Q11: (i) Summary Table (ii) Contingency Table (iii) Frequency
Distribution
Q15: Scatter Plot (two numerical variables, relationship)
Q16: Time Series/Line Chart
Q23: Histogram (shape) + Ogive (cumulative)
Q24: Own survey = Primary; Government census = Secondary
Good luck on the quiz tomorrow!