DATA SCI 415: Overview Snigdha Panigrahi Department of Statistics, University of Michigan Snigdha Panigrahi (University of Michigan) Overview 1 / 35 What’s in the name? • Our course is called “Data mining and Statistical Learning” • Our book is called “Statistical Learning” • Other names you hear: machine learning, artificial intelligence, data science, analytics... and of course statistics! • The differences are mostly historical and cultural; by and large, these fields solve the same problems, but may sometimes differ in their focus. Snigdha Panigrahi (University of Michigan) Overview 2 / 35 Statistical Learning from Big Data • Fact: The amount of data collected and stored is exponentially increasing, due to advances in data collection, computerization of many aspects of life and breakthroughs in technology. Last time: data generated by GAI • Consequence: Data analysis problems have dramatically increased in size and complexity. • Your future job: know how to make sense of all these data! Snigdha Panigrahi (University of Michigan) Overview 3 / 35 Statistical Learning from Big Data • Your future job: make sense of all these data! • Identify patterns and trends: uncover “interesting” relationships among the variables and/or the observations Replication crisis: Not all of them are real!! • Predict future behavior • Attach uncertainties to patterns and make inferences • Your future job: know how to make sense of all these data! Convert uncovered patterns to knowledge Snigdha Panigrahi (University of Michigan) Overview 4 / 35 • Technology helps • Faster computers, more storage ⇒ more flexible and thus more powerful techniques ⇒ fewer modeling assumptions • New visualisation capabilities (a picture is worth a thousand words...) • But not always • Some problems are inherently computationally intractable • “Easy” black-box data analysis can lead to flexible modeling: a lot of misuse and misunderstanding A famous example: Google Flu (predict flu prevalence) GF failed at the peak of the 2013 flu season by 140 percent Why?? Snigdha Panigrahi (University of Michigan) Overview 5 / 35 • Technology helps • Faster computers, more storage ⇒ more flexible and thus more powerful techniques ⇒ fewer modeling assumptions • New visualisation capabilities (a picture is worth a thousand words...) • But not always • Some problems are inherently computationally intractable • “Easy” black-box data analysis can lead to flexible modeling: a lot of misuse and misunderstanding A famous example: Google Flu (predict flu prevalence) GF failed at the peak of the 2013 flu season by 140 percent Why?? Flexible models can overfit (too much of a good thing) Understanding underlying assumptions and interpreting conclusions correctly remains as important as ever Snigdha Panigrahi (University of Michigan) Overview 6 / 35 Statistical Learning from Big Data: Motivation • There is often “hidden” information in the data that is not readily evident • Human analysts without large-scale algorithms may never discover that useful information • Much of the data available are never analyzed at all Snigdha Panigrahi (University of Michigan) Overview 7 / 35 Statistical Learning from data: benefits to science Data are collected and stored at enormous speeds • remote sensors on a satellite (NASA) • telescopes scanning the skies (SDSS) • microarrays generating gene expression data (MEDLINE) • synthetic data by generative AI Snigdha Panigrahi (University of Michigan) Overview 8 / 35 Statistical Learning from data: benefits to science Statistical learning helps scientists with • classifying and clustering data • formulating hypotheses/theories • validating hypotheses/theories • predicting future behavior Snigdha Panigrahi (University of Michigan) Overview 9 / 35 Statistical Learning from data: benefits to business Almost any commercial transaction generates data • Web searches, social connections (Google, Facebook, Twitter, etc) • Purchases both online and at stores (Amazon, eBay, Walmart) • Bank/credit card transactions (Bank of America, Visa, Mastercard) Snigdha Panigrahi (University of Michigan) Overview 10 / 35 Statistical Learning from data: benefits to business • Customized services and successful advertizing give competitive edge • Have to balance useful services vs annoyance and privacy concerns: a fine line! • Fair to customers: ensure inclusive search and recommendations Snigdha Panigrahi (University of Michigan) Overview 11 / 35 Statistical Learning from data: benefits to YOU • We see/read massive amounts of data • We generate massive amounts of data ourselves (everything on your phone is data!) • Human brains are great at spotting and recognizing patterns, but also easy to trick • Big data can help you (fitness trackers, movie recommendations) • Big data can harm you (link between social media and depression, fake news) Snigdha Panigrahi (University of Michigan) Overview 12 / 35 Statistical Learning from data: benefits to YOU Learning to critically think about data is one of the most imporant skills in the modern world Snigdha Panigrahi (University of Michigan) Overview 13 / 35 The process of learning from data Snigdha Panigrahi (University of Michigan) Overview 14 / 35 Some standard notation • n: the number of observations (cases, data points) • p: the number of variables (features, predictors, predictor variables) • Variables can be quantitative, ordinal, categorical, or a mix • Data matrix: n × p matrix X x11 . . . x21 . . . X = . .. .. . xn1 . . . x1p x2p .. . xnp • (Optional) Response: Y , which can be one or several variables for each observation – a special variable of interest. Typically, with one response, Y is stored as an n × 1 vector. Snigdha Panigrahi (University of Michigan) Overview 15 / 35 Two main types of Statistical Learning Supervised learning: X and Y are observed • Goal: understand/summarize/visualize the relationships between X and Y , and/or learn to predict Y from X Unsupervised learning: only X is observed • Goal: understand/summarize/visualize the relationships between the variables in X • Examples: dimension reduction, e.g., principal components analysis, clustering Snigdha Panigrahi (University of Michigan) Overview 16 / 35 Examples • Visualization (applicable to both supervised and unsupervised tasks, often with different plots) • Classification (supervised; Y is a categorical variable) • Regression (supervised; Y is a continuous variable) • ANOVA (supervised; categorical X , continuous Y ) • Clustering (unsupervised) Snigdha Panigrahi (University of Michigan) Overview 17 / 35 Classification: Definition • Given a collection of data points (training set) (x1 , y1 ), . . . , (xn , yn ), where y is a class label (categorical) • Find a model or an algorithm that outputs the class label y as a function of the values of variables x • Goal: previously unseen data points x should be assigned a class label y as accurately as possible • A test set (previously unseen) is used to determine the accuracy of the model Snigdha Panigrahi (University of Michigan) Overview 18 / 35 Classification example: Customer scoring • A bank has a database of 1M past customers, 10% of whom took out mortgages with the bank • Task: predict whether a current customer will take out a mortgage or not, based on the customer’s data • History of transactions with the bank • Other credit data • Demographic data Snigdha Panigrahi (University of Michigan) Overview 19 / 35 Classification example: Spam filter • Task: customize an email spam detection system for an individual user. • Features: relative frequencies of words and punctuation • Most commonly occuring words: spam email george 0.01 1.27 Snigdha Panigrahi (University of Michigan) you 2.26 1.27 your 1.38 0.44 hp 0.02 0.90 Overview free 0.52 0.07 re 0.13 0.42 remove 0.28 0.01 20 / 35 Classification example: Microarray gene expression data Task: predict risk of developing cancer based on genetic profile BREAST RENAL MELANOMA MELANOMA MCF7D-repro COLON COLON K562B-repro COLON NSCLC LEUKEMIA RENAL MELANOMA BREAST CNS CNS RENAL MCF7A-repro NSCLC K562A-repro COLON CNS NSCLC NSCLC LEUKEMIA CNS OVARIAN BREAST LEUKEMIA MELANOMA MELANOMA OVARIAN OVARIAN NSCLC RENAL BREAST MELANOMA OVARIAN OVARIAN NSCLC RENAL BREAST MELANOMA LEUKEMIA COLON BREAST LEUKEMIA COLON CNS MELANOMA NSCLC PROSTATE NSCLC RENAL RENAL NSCLC RENAL LEUKEMIA OVARIAN PROSTATE COLON BREAST RENAL UNKNOWN SIDW299104 SIDW380102 SID73161 GNAL H.sapiensmRN SID325394 RASGTPASE SID207172 ESTs SIDW377402 HumanmRNA SIDW469884 ESTs SID471915 MYBPROTO ESTsChr.1 SID377451 DNAPOLYME SID375812 SIDW31489 SID167117 SIDW470459 SIDW487261 Homosapiens SIDW376586 Chr MITOCHONDR SID47116 ESTsChr.6 SIDW296310 SID488017 SID305167 ESTsChr.3 SID127504 SID289414 PTPRC SIDW298203 SIDW310141 SIDW376928 ESTsCh31 SID114241 SID377419 SID297117 SIDW201620 SIDW279664 SIDW510534 HLACLASSI SIDW203464 SID239012 SIDW205716 SIDW376776 HYPOTHETIC WASWiskott SIDW321854 ESTsChr.15 SIDW376394 SID280066 ESTsChr.5 SIDW488221 SID46536 SIDW257915 ESTsChr.2 SIDW322806 SID200394 ESTsChr.15 SID284853 SID485148 SID297905 ESTs SIDW486740 SMALLNUC ESTs SIDW366311 SIDW357197 SID52979 ESTs SID43609 SIDW416621 ERLUMEN TUPLE1TUP1 SIDW428642 SID381079 SIDW298052 SIDW417270 SIDW362471 ESTsChr.15 SIDW321925 SID380265 SIDW308182 SID381508 SID377133 SIDW365099 ESTsChr.10 SIDW325120 SID360097 SID375990 SIDW128368 SID301902 SID31984 SID42354 21 / 35 Overview Snigdha Panigrahi (University of Michigan) Classification example: Anomaly detection • Detect significant deviations from normal behavior • Applications: • Credit card fraud • Network intrusion Snigdha Panigrahi (University of Michigan) Overview 22 / 35 Classification example: Credit card fraud detection • Credit card losses in the US are over 1 billion $ per year • Roughly 1 in 50k transactions are fraudulent • Fair-Issac’s fraud detection software based on neural networks, led to reported fraud decreases of 30-50% • Challenge: false alarm rate vs missed detection Snigdha Panigrahi (University of Michigan) Overview 23 / 35 Regression: Definition • Predict a value of a continuous-valued response variable based on the values of other variables • Linear regression: predict Y from a linear combination of X • There are many other tools for regression: nonlinear functions, trees, neural networks, etc Snigdha Panigrahi (University of Michigan) Overview 24 / 35 Regression: Examples • Predicting sales amounts of new product based on advertising expenditure and product characteristics • Predicting wind velocity as a function of temperature, humidity, air pressure, etc • Predict a student’s freshman year GPA based on high school grades and SATs Snigdha Panigrahi (University of Michigan) Overview 25 / 35 Universal trade-offs • Occam’s razor (“Less is More”): accuracy vs. interpretability • Bias (accuracy) vs variance (replicability) • Overfitting vs underfitting Snigdha Panigrahi (University of Michigan) Overview 26 / 35 Prediction vs inference • Prediction: the goal is to predict Y from X . The predictor could be a black box as long as it’s accurate. • Inference: the goal is to understand how Y is connected to X and find explanatory value in X Get good predictions, but fail to make inferences Snigdha Panigrahi (University of Michigan) Overview 27 / 35 Clustering: Definition • Given a set of data points, each having a set of variables, find clusters such that • data points in the same cluster are “ more similar” to one another, and • data points in different clusters are “less similar” to one another. • Similarity measures • Euclidean distance if variables are continuous • Other problem-specific measures Snigdha Panigrahi (University of Michigan) Overview 28 / 35 Importance of similarity measures Snigdha Panigrahi (University of Michigan) Overview 29 / 35 Clustering example: Market segmentation • Goal: subdivide a market into distinct subsets of customers for targeted marketing • Collect different variables on customers (age, gender, marital status, education; geographical location; lifestyle, hobbies, etc) • Find clusters of similar customers • Here we rely on the assumption similar customers will like the same kind of marketing, but have no previous data on how they actually respond to particular marketing strategies Snigdha Panigrahi (University of Michigan) Overview 30 / 35 Clustering example: Images Snigdha Panigrahi (University of Michigan) Overview 31 / 35 More clustering examples • Cluster documents that are similar to each other based on the important terms appearing in them, for example to organize news articles • Cluster patients by symptoms and medical history, to develop personalized treatments • Cluster stocks based on their movements every day, to find patterns in the market Snigdha Panigrahi (University of Michigan) Overview 32 / 35 Modern challenges in Statistical Learning • Hype: people often expect more than is realistic • Data snooping and fishing: finding spurious structure that is not replicable (Topic that is close to my heart!!) • Irreplicable analysis • Trade-offs: • Prediction vs inference: may get great performance from a black box, and a more interpretable simpler model may not predict as well • Bias vs variance: flexibility vs overfitting • Balancing false alarms against missed detection (Type 1 vs Type 2 error) Snigdha Panigrahi (University of Michigan) Overview 33 / 35 Data science/ Statistics in the age of AI The future of data science is intimately linked to the future of AI. As AI advances, it will create new opportunities and challenges for data scientists. Snigdha Panigrahi (University of Michigan) Overview 34 / 35 Data science/ Statistics in the age of AI Google’s chief economist Hal Varian, 2009: The ability to take data - to be able to understand it, to process it, to extract value from it, to visualize it, to communicate it - that’s going to be a hugely important skill in the next decades, not only at the professional level but even at educational levels... Because now we really do have essentially free and ubiquitous data. So the complementary scarce factor is the ability to understand that data and extract value from it. Snigdha Panigrahi (University of Michigan) Overview 35 / 35
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )