UGBA104: Introduction to Business Analytics
Week 07 Predictive Models
Classification Trees
Professor Thomas Y. Lee
Operations and Information Technology Management
Recitation Exercises: BlueOrRed
Problem 1.
Campaign organizers for both the Republican and Democratic parties are
interested in identifying individual undecided voters who would consider voting
for their party in an upcoming election. The workbook W07DataBlueOrRed.xlsx contains data on a sample of voters with tracked
variables, including whether or not they are undecided regarding their
candidate preference, age, whether they own a home, gender, marital status,
household size, income, years of education, and whether they attend church.
Question 1.1. Based upon datatype, which of the following attributes would
create an ASP error when configuring a Classification Tree (Choose all that
apply)?
a. Age
b. HomeOwner
c. Female
d. Married
e. Income
f. Household Size
g. Education
h. Church
i. None of these
Adapted From Camm Problem 11.44
thomasyl@berkeley.edu
UC Berkeley, Haas School
Slide 2
Recitation Exercises: BlueOrRed
Question 1.2. How would you expect results to differ when training your
classification tree if you standardize (Z-scores) v. do not standardize?
a. Increase overall accuracy on training data but not the validation data (i.e.
overfitting)
b. Increase overall accuracy (both training and validation data)
c. Decrease overall accuracy (both training and validation data)
d. Makes no difference
e. Cannot tell
Question 1.3. You may set the number of records in a Terminal Node to 50 or
100. Which Tree is more likely to overfit?
a. 50 records in a Terminal Node
b. 100 records in a Terminal Node
c. no differences
thomasyl@berkeley.edu
UC Berkeley, Haas School
Slide 3
Recitation Exercises: BlueOrRed
Question 1.4. Assume you wish to use Undersampling to correct for imbalance
in the data set. If a 1 in the data set represents Undecided voters and a 0 in
the data set represents Decided voters, which class will you undersample in the
case of the Blue or Red data data set?
a. Undecided
b. Decided
Question 1.5. Continue with your undersampling. Suppose 50% of the
observations from the under-represented class are used for Training, 30% for
Validation and 20% for Testing. For the Blue or Red data set, how many
observations from the over-represented class do you expect to be used in
Validation?
thomasyl@berkeley.edu
UC Berkeley, Haas School
Slide 4
Recitation Exercises: BlueOrRed
Problem 2. Data is pre-partioned with a designated variable named Partition in
the following proportions: 50% of observations in the training set, 30% in the
validation set, and 20% in the test set. Fit a single classification tree using Age,
HomeOwner, Female, Married, HouseholdSize, Income, Education, and
Church as input variables and Undecided as the output variable. In Step 2 of
ASP’s Classification Tree procedure, do not Rescale Input Data
(standardize). Set the Minimum # records in a terminal node to 100.
Display the Full tree, Best Pruned Tree, and Minimum Error Tree with a
maximum of 7 levels. Validate on the Best Pruned Tree (use best pruned tree
for scoring).
Question 2.1. From the CT_TrainingScore worksheet, what is the overall error
rate of the full tree on the training set? Answer to two decimal places.
thomasyl@berkeley.edu
UC Berkeley, Haas School
Slide 5
Recitation Exercises: BlueOrRed (Cont’d)
Question 2.2. Why not use the full tree for predicting the status of future
voters?
a. The full tree is most likely underfit for future voters.
b. The full tree is most likely overfit for future voters.
c. Future voters will have different variables and/or different variable values
than current and past voters.
d. We should use the full tree for predicting the status of future voters.
e. None of these.
Question 2.3. Consider a 50-year-old man who attends church, has 15 years
of education, owns a home, is married, lives in a household of four people, and
has an annual income of $150,000. How does the best pruned tree classify this
observation?
a. Undecided
b. Decided
Question 2.4. For the best pruned tree, what is the lift on the top 30% of the
test set deemed most likely to be undecided? Answer to two decimal places.
thomasyl@berkeley.edu
UC Berkeley, Haas School
Slide 6
Recitation Exercises: BlueOrRed (Cont’d)
Question 2.5. In most cases, which of the following increases your risk of
overfitting when training your classification tree? Select all that apply.
a. Adding more observations for training
b. Adding more input variables
c. Increasing the maximum depth of the tree
d. Increasing the number of records per node
e. None of these
thomasyl@berkeley.edu
UC Berkeley, Haas School
Slide 7