Data Analysis Post Graduate Program Data Analytics Post Graduate Program in Data Analytics Data Analysis: Participant Manual Contents Overview of Analytics ............................................................................................................................................................. 4 Data Overload .............................................................................................................................................................................................. 4 What is Analytics? ...................................................................................................................................................................................... 4 Introduction to Data Analytics ............................................................................................................................................................. 4 Enter Data Scientists ................................................................................................................................................................................. 4 Growing Need for Analytics ................................................................................................................................................................... 5 The Case for Analytics .............................................................................................................................................................................. 5 Types of Analytics ...................................................................................................................................................................................... 5 Applications of Analytics ......................................................................................................................................................................... 5 Careers in Analytics ................................................................................................................................................................................... 8 What is Data Analysis? .......................................................................................................................................................................... 10 Difference between Data Analysis and Data Analytics ........................................................................................................... 11 Similarity between Data Analysis and Data Analytics ............................................................................................................ 11 Types of Data ............................................................................................................................................................................................. 11 Data Collection Types: What it Mean?............................................................................................................................................ 14 Forms of Data ............................................................................................................................................................................................ 15 Sources of Data ......................................................................................................................................................................................... 15 Sources of Errors ..................................................................................................................................................................................... 18 Applied Statistics ..................................................................................................................................................................................... 19 Statistics ...................................................................................................................................................................................................... 19 Inferential Statistics ............................................................................................................................................................................... 20 Hypothesis Testing ................................................................................................................................................................................. 20 Process of Hypothesis Testing ........................................................................................................................................................... 20 Possible Scenarios in Hypothesis Testing .................................................................................................................................... 20 Data Mining ................................................................................................................................................................................................ 21 Application of Data Mining .................................................................................................................................................................. 21 Characteristics of Data Mining .......................................................................................................................................................... 22 Data Mining Activities ........................................................................................................................................................................... 22 Analytics across Industries ................................................................................................................................................................. 25 Analytics across Functions .................................................................................................................................................................. 27 Data Analysis Process ............................................................................................................................................................................ 27 Data Analysis & Architecture .............................................................................................................................................30 Exploratory Data Analysis (EDA) ..................................................................................................................................................... 30 Introduction ............................................................................................................................................................................................... 30 Why Do You Need EDA? ....................................................................................................................................................................... 30 Goals of EDA............................................................................................................................................................................................... 30 Confidential and restricted. Do not distribute. (c) Imarticus Learning 2 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Categories of EDA .................................................................................................................................................................................... 30 EDA: Univariate Analysis ..................................................................................................................................................................... 31 Univariate Analysis ................................................................................................................................................................................. 31 Multivariate Analysis ............................................................................................................................................................................. 31 Data Treatment ........................................................................................................................................................................................ 33 Outliers......................................................................................................................................................................................................... 33 Default Values ........................................................................................................................................................................................... 34 Missing Values .......................................................................................................................................................................................... 34 Skewed Distributions ............................................................................................................................................................................ 34 Case Study - Transforming Skewed Data ...................................................................................................................................... 35 Univariate Analysis ................................................................................................................................................................................. 35 Transformed Data ................................................................................................................................................................................... 35 Derived Variable ...................................................................................................................................................................................... 36 Summary ..................................................................................................................................................................................................... 36 Data Architecture .................................................................................................................................................................................... 36 Components of Data Architecture.................................................................................................................................................... 36 Enterprise Data Architecture ............................................................................................................................................................. 37 OLTP and OLAP Systems ...................................................................................................................................................................... 38 OLTP Characteristics ............................................................................................................................................................................. 38 OLAP Characteristics ............................................................................................................................................................................. 38 OLTP vs. OLAP .......................................................................................................................................................................................... 39 How is Data Stored? ............................................................................................................................................................................... 39 Database Management System.......................................................................................................................................................... 39 A Database Schema ................................................................................................................................................................................. 40 Database Types ........................................................................................................................................................................................ 40 Subject Orientation of a Data Warehouse .................................................................................................................................... 40 Non-Volatility of a Data Warehouse................................................................................................................................................ 41 Time Variant Data Warehouse .......................................................................................................................................................... 41 Architectural Components of MDM ................................................................................................................................................. 41 Information Architecture ..................................................................................................................................................................... 42 Key Questions for MDM Design ......................................................................................................................................................... 42 MDM Domains........................................................................................................................................................................................... 42 Master Data Management .................................................................................................................................................................... 42 MDM Vs. DW: Different Goals ............................................................................................................................................................ 43 Confidential and restricted. Do not distribute. (c) Imarticus Learning 3 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Overview of Analytics Data Overload Big Data is the data that is TOO LARGE & TOO COMPLEX for conventional data tools to capture, store and analyze. 7 Billion shares traded on US Stock Markets each day; 10 Terabytes data generated in one flight from NY to London; 400 Million number of tweets per day on Twitter and 3 Billion ‘Likes’ each day on Facebook. 90% of the world’s data was generated in the last two years. The 3V’s of Big Data 1. Volume 2. Variety 3. Velocity What is Analytics? The scientific process of transforming data into insight for making better decisions, offering new opportunities for a competitive advantage. Analytics is not so much about tools or technologies – It is a way of thinking that uses knowledge, tools and techniques to extract valuable insights from unstructured data, which then leads to a business strategy. Business Analytics refers to the practice of investigation of past business performance, using data and statistical models in order to develop new insights and understanding of future business performance. It makes extensive use of statistical and quantitative analysis, explanatory and predictive modeling and fact-based management to drive decision-making. Introduction to Data Analytics Data Analytics is the science of analyzing data to convert information to useful knowledge. This knowledge could help us understand our world better and enable us to make better decisions. Analytics refers to the collection of: • Tools • Techniques • Skills This aid the investigation of past business performance. Enter Data Scientists A Business analyst is not able to discover insights from huge sets of data of different domains. That’s where data scientists are hired. Data scientists can work in co-ordination with different verticals of an organization and find useful patterns/insights for a company to make tangible business decisions. A business analyst never knew that data would become highly unstructured and difficult to analyze with the 3 v’s. Confidential and restricted. Do not distribute. (c) Imarticus Learning 4 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Growing Need for Analytics Generation of large amount of data from business transactions. Availability of the large data storage system at lower cost. Availability of better tools and technology to analyze the large data sets. The Case for Analytics • Business Need: The Business environment today is more complex than ever before. Businesses are expected to be diligently responsive to the increasing demands of customers, various stakeholders and even regulators. • Solution: Organizations have been turning to the use of analytics. More than 83% of Global CIOs surveyed by IBM in 2010 singled out Business Intelligence and Analytics as one of their visionary plans for enhancing competitiveness. • Goal: Most cases the primary objective of an organization that seeks to turn to analytics is Revenue/Profit growth and Optimize expenditure. Types of Analytics 1. Prescriptive analytics 2. Predictive analytics 3. Descriptive analytics Applications of Analytics Uses of Analytics for Recommendations One of the most common problem solved by various websites is that of “Recommendations”. Several business have emerged that track customer data and solve various business problems. Few of the key factors impacting success: • Opposition • Venue Confidential and restricted. Do not distribute. (c) Imarticus Learning 5 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • • • Temperature Player Performance Weather Conditions Two tech firms - Cognizant and Clarabridge claim they know which film will win the Best Picture Oscar this year. They have reached their conclusions not by watching the films and applying artistic criticism, but simply by crunching data - lots of data. They looked at 150 variables, from film genre to box office takings, from review ratings to the percentage of female viewers under 18 and they applied their algorithm to data going back 15 years to work out which of these variables were the most important. Can we predict Oscar winners using data analytics alone? Interestingly, they also measured sentiment - the emotional reactions each film elicited - on popular movie review sites such as IMDB and Rotten Tomatoes. Just to give some sense of the amount of data, the tech firms looked at 150,000 text reviews and more than 38 million star ratings from IMDB alone. The tech firms say they're 64% confident that the Revenant will win. Google Suggest Google Suggest leverages predictive analytics to give out real time search query recommendations. This feature was the work of Kevin Gibbs, who started out trying to build a URL predictor that auto-completes the URL as a user types it in. This idea was then ported over to Google search itself, driving a lot of adoption, and saving users a lot of time while searching. Amazon Product Recommendations We’ve all gotten lost in product research on Amazon.com, jumping from one product to another, comparing one list with another, in a seemingly endless maze of products. This addictive shopping experience is due to the smart work of Amazon in the background where they’ve perfected the art of cross-selling. The contexts in which Amazon’s Recommendations are made include: • Item page • Checkout page • Search results page • Customized home page • Email promotions These recommendations are based on aggregate data for tons of transactions, and user behavior. Facebook’s News Feed In recent times, have you noticed that Facebook increasingly prioritizes updates from your close friends and family more than others? In the early days of Facebook, the newsfeed used to have all updates from your friends in chronological order. The recent changes in prioritisation is due to its predictive analytics algorithms at play, guessing which stories Confidential and restricted. Do not distribute. (c) Imarticus Learning 6 Post Graduate Program in Data Analytics Data Analysis: Participant Manual you’d be most interested in based on who you speak with most frequently, and your profile information. Video Games Gaming companies like EA and Zynga generate a lot of structured and unstructured data from their games. This data is in the form of gameplay data, micro-transactions, time stamps, in-game advertising, multi-player information, and much more. These companies see the huge opportunity to customize gameplay, find new ways of monetizing games, and even enriching the gaming experience by making it social. FICO Credit Score FICO score is used to assess the creditworthiness of a person, and the level of risk a finance company incurs when lending to that person. It uses factors such as on-time payments, credit capacity used, and duration of credit history to assign each person a score. This process of assessing risk when lending used to take banks weeks before, but now can be done in just a few hours, thanks to the FICO score. Unfortunately, as with all predictive models, accuracy is a concern, and even an established model like the FICO score can go badly wrong as in the sub-prime mortgage crisis of 2008. Email Spam Filtering Spam filtering uses predictive analytics. When Gmail launched back in 2004 it was lauded for its exceptionally better spam filtering than the then popular Yahoo Mail. Spam is a major concern for email users. When dealing with spam, email service providers need to not only filter out spam email, but ensure the spam filters don’t catch false positives. USPS Zip Code Recognition We don’t often use snail mail these days, but the digit recognition system used by the US post office is an example of predictive analytics. The postal service needs to recognize the numbers in a zip code, and sort its mail by zip code. It uses algorithms that have been trained to recognize hand-written digits. It has a large set of pictures that are labeled by hand as being in 0 to 9, and uses them to accurately identify zip code numbers on mail. Behavioral Advertising Advertising platforms are ever in the search of a better way to target ads to their user base. Be it Google with AdWords, Facebook with their ads within the News Feed, or advertising startups like UberAds. Election Campaigns Election campaign planners have long recognized the role of social media in deciding the outcome of an election. The next step is to leverage data in making more accurate predictions on the impact of their messaging. Confidential and restricted. Do not distribute. (c) Imarticus Learning 7 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Careers in Analytics New Analytics Jobs by Industry Confidential and restricted. Do not distribute. (c) Imarticus Learning 8 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Employment Landscape in India Global Clients Confidential and restricted. Do not distribute. (c) Imarticus Learning 9 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Career Options in Analytics What is Data Analysis? Analysis refers to breaking a whole into its separate components for individual examination. Data analysis is a process for obtaining raw data & converting it into information useful for decision-making by users. Data is collected and analysed to answer questions, test hypotheses or disprove theories. Definition Statistician John Tukey (Born 1915) defined data analysis as: "Procedures for analyzing data, techniques for interpreting the results of such procedures, ways of planning the gathering of data to make its analysis easier, more precise or more accurate, and all the machinery and results of (mathematical) statistics which apply to analyzing data.” Process There are several phases that can be distinguished and these phases are iterative, in that feedback from later phases may result in additional work in earlier phases. • Step 1: Define your questions • Step 2: Set Clear Measurement Priorities • Step 3: Collect Data • Step 4: Data Cleaning • Step 5: Analyse Data • Step 6: Interpret Results Confidential and restricted. Do not distribute. (c) Imarticus Learning 10 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Difference between Data Analysis and Data Analytics • Data Analysis: Analysis looks backwards over time, providing marketers with a historical view of what has happened. • Data Analytics: Analytics look forward to model the future or predict a result. Similarity between Data Analysis and Data Analytics Both analysis and analytics provide the insight marketers use to: • Value customers accurately • Target the right audience • Improve the effectiveness of their marketing budget. Both help marketers transform customer data by exploring and analyzing that data to help: • Uncover unknown patterns • Opportunities • Insights that can drive proactive & evidence-based decision making. Types of Data Data is a set of values of qualitative or quantitative variables. Basis of Data Categorization • Source of Generation • Size • Type • Processing • Form Confidential and restricted. Do not distribute. (c) Imarticus Learning 11 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Quantitative Data Types Examples of quantitative data type: Confidential and restricted. Do not distribute. (c) Imarticus Learning 12 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Classification Based on Processing Confidential and restricted. Do not distribute. (c) Imarticus Learning 13 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Categorization Based on Data Collection Type Data Collection Types: What it Mean? • Census: Systematic collection of data about all members of population. • Observational study: Collection of data to draw inference of outcome of a treatment on subjects. It is not in control of the investigator to assign the subjects either to the test or the control groups. • Convenience sample: Collection of data from a sample where the subjects are selected because of their convenient accessibility and proximity to the researcher. • Randomized trial: Collection of data to draw inference of outcome of a treatment on subjects. The investigator randomly allocates the subjects to either the test or the control group. • Predictive study: Training data: A set of data that is used for discovering potential relationships. Testing data: A set of data that is used for assessing the strength and robustness of a predictive relationship. • Studies over time: Cross-sectional: Data collected by observing many subjects at the same time. Snapshot view of the subjects/cross-section of the population. Longitudinal: Data collected on repeated observation of the same set of variables over a period of time. • Retrospective: Collection of data of events that have already occurred. • Time series: It is a sequence of data points, measured typically at successive points in time spaced at uniform time intervals. One dimensional data and has a natural temporal ordering. Confidential and restricted. Do not distribute. (c) Imarticus Learning 14 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • • • Spatial: Data that have a spatial (geographical) component, which means, they have a location on earth. Cross-section: Cross-sectional data refers to data collected by observing many subjects (such as individuals, organizations, or countries/regions) at the same time. One dimensional data. Panel data: Panel data contains observations of multiple phenomena obtained over multiple time periods for the same organizations or individuals. Forms of Data Structured Data: • Data can be organized in well-defined structures • Structures include arrays, vectors, or tables • Relationships in data defined within the structure Unstructured Data: Data is organized in an arbitrary manner with no pre-defined structure. The types of content includes free text, documents, images, and videos. Example: Resume document of a student with free text and images. Semi-structured Data Semi-structured data does not conform to a formal defined structure, but entities belong to classes with attributes. Data cannot be processed as effectively as the structured data. Example: Information stored as XML. Batch Data Data is collected in a batch mode at periodic intervals. There is a delay in the availability of data as a certain periodicity is maintained for its collection. Real-time Data Data is collected in real time; the data is delivered as it gets generated. There is no delay in timeliness of data provided. Sources of Data • • User Generated: Blogs and Documents System Generated: Web Logs and Network event logs Confidential and restricted. Do not distribute. (c) Imarticus Learning 15 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • • • Device Generated: Surveillance cameras capturing traffic patterns which is device generated Internal: Generated internally in an organization across business processes External: Generated by external bodies or data aggregators or credit bureaus Demographic Data Demographics or Demographic Data are the characteristics of a human population. Demographic profiling makes generalization about groups of people in a population; it describes a typical member of the group. Example of a demographic profile: “Single, male, age 18 to 24, college educated”. Common demographic variables are gender, race, age, home ownership, employment status, and location. Psychographic Data Psychographic variables are attributes, which are related to personality, values, attitudes, interests, or lifestyles. Psycho-graphic data represents behavioural data under categories of Interests, Activities, and Opinions (IAO). They can be used to segment or divide the market into groups based on customer’s lifestyle or behavior. Internal Data Data generated inside an organization. Example: In a bank’s credit card business data is generated by multiple processes. Derived Data: Data derived from attributes available • Months since last sale • Number of transactions • Number of payments missed in the last 3 months • Number of times credit limit increased in the last 6 months Confidential and restricted. Do not distribute. (c) Imarticus Learning 16 Post Graduate Program in Data Analytics Data Analysis: Participant Manual External Data Data available with external organizations, such as market research companies, data aggregators, benchmark studies, regulatory authorities, and so on. Example: Credit Bureaus aggregate consumer data from creditors, lenders, utilities, debt collection agencies, and public records from courts to provide an overall view of an individual consumer’s credit performance. Data Quality Data Quality Issues • • Missing Data: Input process not capturing all the data or not mandatory in the input process Junk Values: Lack of validation Confidential and restricted. Do not distribute. (c) Imarticus Learning 17 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • • • • • • Definitions: Inaccurate or incomplete definition Completeness: Incomplete data, left blank Validity: Invalid data (does not follow the expected structure) Accuracy: Inaccurate data due to problems with either the measurement system or the operator Timeliness: Delay in availability Consistency: Inconsistent definition across systems, difference in scales of measurements. Sources of Errors Many of the sources of error in databases fall into one or more of the following categories: • Data entry errors • Measurement errors • Data integration errors • Distillation errors Data Entry Errors It remains common in many settings for data entry to be done by humans, who typically extract information from speech (e.g., in telephone call centers) or by keying in data from written or printed sources. In these settings, data is often corrupted at entry time by typographic errors or misunderstanding of the data source. Another very common reason that humans enter “dirty” data into forms is to provide what we call spurious integrity: many forms require certain fields to be filled out, and when a data-entry user does not have access to values for one of those fields, they will often invent a default value that is easy to type, or that seems to them to be a typical value. This often passes the crude data integrity tests of the data entry system, while leaving no trace in the database that the data is in fact meaningless or misleading. Many forms require certain fields to be filled out, and when a data-entry user does not have access to values for one of those fields, they will often invent a default value that is easy to type, or that seems to them to be a typical value. Measurement Errors In many cases data is intended to measure some physical process in the world: the speed of a vehicle, the size of a population, the growth of an economy, etc. In some cases these measurements are undertaken by human processes that can have errors in: • Their design (e.g., improper surveys or sampling strategies) • Execution (e.g. misuse of instruments) In the measurement of physical properties, the increasing proliferation of sensor technology has led to large volumes of data that is never manipulated via human intervention. While this avoids various human errors in data acquisition and entry, data errors are still quite common. The human design of a sensor deployment (e.g., selection and placement of Confidential and restricted. Do not distribute. (c) Imarticus Learning 18 Post Graduate Program in Data Analytics Data Analysis: Participant Manual sensors) often affects data quality, and many sensors are subject to errors including miscalibration and interference from unintended signals. Distillation Errors In many settings, raw data are preprocessed and summarized before they are entered into a database. This data distillation is done for a variety of reasons: • To reduce the complexity or noise in the raw data (e.g., many sensors perform smoothing in their hardware). • To perform domain-specific statistical analyses not understood by the database manager. • To emphasize aggregate properties of the raw data (often with some editorial bias). • In some cases simply to reduce the volume of data being stored. All these processes have the potential to produce errors in the distilled data, or in the way that the distillation technique interacts with the final analysis. Data Integration Errors It is actually quite rare for a database of significant size and age to contain data from a single source, collected and entered in the same way over time. In almost all settings, a database contains information collected from multiple sources via multiple methods over time. Moreover, in practice many databases evolve by merging in other pre-existing databases; this merging task almost always requires some attempt to resolve inconsistencies across the databases involving data representations, units, measurement periods, and so on. Any procedure that integrates data from multiple sources can lead to errors. Applied Statistics Why Statistics? Analysis of data starts with statistical analysis. Describing the data using Statistics is the first step in understanding the data. • Make decisions based on the data collected • Read, evaluate, and interpret the results • Develop systematic approach to analyze data • Evaluate the information you have been given Statistics Statistics is collection, organization, analysis, and interpretation of data. Statistics includes Design of experiments, Sampling, Descriptive Statistics, Inferential Statistics and Probability theory. In statistical analysis, the three fundamental concepts associated with describing data are: • Location or Central Tendency • Dispersion or Spread • Shape or Distribution Confidential and restricted. Do not distribute. (c) Imarticus Learning 19 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Descriptive Statistics helps in: • Identifying potential data problems such as errors, Outliers, and extreme values • Identifying process issues • Selection of appropriate statistical test for understanding the underlying relationships/patterns Inferential Statistics Hypothesis: Business Examples No difference in performance of the sales team across geographies and product lines. Change in gas price will have no impact on losses in automotive finance. Change in CEO will have no impact on the stock price. Real estate yields are the same in all metros. Compensation changes will not impact attrition. Hypothesis Testing Is a method of making an inference about a population parameter based on sample data. Is statistical analysis used to determine if the difference observed in samples is not a random occurrence but a true difference. Process of Hypothesis Testing The process of hypothesis testing consists of four steps: • • • • Step 1: Formulate the Null Hypothesis and the alternative hypothesis. Please remember Null Hypothesis is always status-quo, which means “no difference”. Data is gathered as evidence to either reject or not reject the Null Hypothesis. Step 2: Identify a test statistic that can be used to assess the truth of the Null Hypothesis. Step 3: Compute the P-value; the smaller the P-value, the stronger the evidence against the Null Hypothesis. Step 4: Compare the P-value to an acceptable significance value. If P-value is less than the significance level, the Null Hypothesis is rejected. Possible Scenarios in Hypothesis Testing Four possible scenarios: Confidential and restricted. Do not distribute. (c) Imarticus Learning 20 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • • Type I Error (α): Reject the Null Hypothesis when it is true Type II Error (β): Accept the Null Hypothesis when it is false Null Hypothesis = “Person is innocent” Points to Remember Important points to note regarding Hypothesis Testing are: • It is always “Reject” or “Do Not Reject” the Null Hypothesis. • Rejecting the Null Hypothesis means there is evidence that there is a difference (based on the sample data). • Failure to reject the Null Hypothesis means data is insufficient to conclude that there is a difference. Data Mining What is Data Mining? Data mining is the process of exploration and analysis, by automatic or semi-automatic means, of large quantities of data in order to discover meaningful patterns and rules. Application of Data Mining • Clustering • Finding dependency networks • Data summarization • Analyzing changes • Learning classification rules • Detecting anomalies Summarizing It All Analysis of data and the use of software techniques for finding patterns and regularities in sets of data. The computer is responsible for finding patterns by identifying underlying rules and features. It is possible to ‘strike gold’ in unexpected places as the data mining software extracts patterns not previously discernible or so obvious that no-one has noticed them before. Confidential and restricted. Do not distribute. (c) Imarticus Learning 21 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Mining analogy: Large volumes of data are sifted in an attempt to find something worthwhile (For example: In a mining operation, large amounts of low grade materials are sifted through in order to find something of value). Characteristics of Data Mining • Large quantities of data: Volume of data so great it has to be analyzed by automated techniques. E.g. POS, satellite information, credit card transactions etc. • Noisy, incomplete data: Imprecise data is characteristic of all data collection databases usually contaminated by errors, cannot assume that the data they contain is error free. E.g. some attributes rely on subjective or measurement judgments. • Complex data structure: Conventional statistical analysis not possible. • Heterogeneous data: Stored in legacy systems. Scope Wherever there are large quantities of data and something worth learning. Example: 1. Medicine: drug side effects, hospital cost analysis, genetic sequence analysis, prediction etc. 2. Finance: stock market prediction, credit assessment, fraud detection etc. 3. Marketing/sales: product analysis, buying patterns, sales prediction, target mailing, identifying `unusual behaviour. 4. Scientific discovery: superconductivity research, etc. 5. Engineering: automotive diagnostic expert systems, fault detection etc. Data Mining: Business Context Useful wherever the resulting knowledge is worth more than the cost to discover. • As a research tool Pharmacy – NPD • For process improvement Manufacturing: optimal bound combinations • Specific applications like CRM, ECR, SCM, BPR, etc. • For marketing, all the decisions based on market information Data Mining Activities Directed Data Mining • Classification (assigning to a predefined class) • Estimation (continuously valued outcomes) • Prediction (completing incomplete values) Undirected Data Mining • Affinity grouping or association rules (what go together) • Clustering (similar subgroups) • Description and visualization (descriptive outline) Confidential and restricted. Do not distribute. (c) Imarticus Learning 22 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Classification DM system learns from examples or the data how to partition or classify the data i.e. it formulates classification rules. Example: Customer database in a bank. Typical rule formulated, if STATUS = married and INCOME > 10000 and HOUSE_OWNER = yes then INVESTMENT_TYPE = good. Estimation Estimation essentially deals with continuously valued outcomes (as against classification) to come up with a value for some unknown continuous variable. Example: Who should we offer our new home equity loan? Typical Estimation Formulation Generate a score between 0 and 1 by running all its customer through a model giving the probability that the customer would respond favorably. Prediction Prediction same as estimation or classification only that in prediction the emphasis is on future behavior and only way to check is to wait and see. Historical data is used to build a model that explains the current behavior. When this model is applied to current inputs, the result is a prediction of future behavior. Examples: Predicting which customers will leave in the next six months. Association Rules that associate one attribute of a relation to another. Set oriented approaches are the most efficient means of discovering such rules. Example - Supermarket Database 72% of all the records that contain items A and B also contain item C. Clustering Segmenting a diverse group into a number of more similar subgroups or clusters. It does not depend on pre-defined class. The meaning to the emerging cluster defined by the miner. Description & Visualization Descriptive rules of association as emergent from the data. Very helpful and has critical implications for the actual process. Example: Gender gap in religiousness. Women more religious than men: has implication for all sorts of journalists, cultural studies, sociologists, religious cults etc. Confidential and restricted. Do not distribute. (c) Imarticus Learning 23 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Sequence/Temporal Sequential pattern functions analyse collections of related records and detect frequently occurring patterns over a period of time. Difference between sequence rules and other rules is the temporal factor. Example Retailers Database: Can be used to discover the set of purchases that frequently precedes the purchase of a microwave oven Example Natural Disasters Database: Discovery could be that when there is an earthquake in Los Angeles the next day mount Kilimanjaro erupts. D Mining: Technical Context Three main areas: 1. Algorithms and techniques 2. Data 3. Modeling practices Antecedents to data mining • Machine learning: Computer science and AI: iterative learning • Statistics: Predictive algorithms or regressions • Decision support: Data warehouse, OLAP, data marts, multidimensional databases Transforming Data to Results Confidential and restricted. Do not distribute. (c) Imarticus Learning 24 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Analytics Value Chain Analytics: Data-Intelligence-Insight Analytics is the transformation of data to insight. The transformation involves: • Understanding the past and current performance to predict future performance • Understanding the relations, identifying patterns and translating them to meaningful, useful and relevant business insights, and intelligent strategies • Laying the foundation for a data driven decision making process in an enterprise Analytics across Industries Retail • Market Basket Analysis • Store segmentation • Product mix and promotions at stores Telecommunications • Subscriber churn analysis • Customer experience Analytics • Asset utilization and productivity Healthcare • Clinical Data Analytics • Healthcare Administration • Finance Analytics Confidential and restricted. Do not distribute. (c) Imarticus Learning 25 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Banking and Financial Services • Risk management • Loss mitigation • Optimize operations • Revenue growth • Customer Analytics Sports • Player Analytics • Fan Analytics • Business of sports Learning and Education • Social learning Analytics • Student performance and learning path • Education policy Manufacturing • Inventory Analytics • Manufacturing Process Analytics • Supply chain Analytics Energy Management • Energy price and risk management • Demand forecast and capacity planning • Smart Energy solutions eCommerce and Social Media • Web Analytics • Social Media Analytics: Sentiment analysis • Search Engine Analytics Government • eGovernance • Improve public policy goals • Cyber security Law Enforcement • Tactical and behavioral crime analysis • Risk and threat assessment and prediction • Resource optimization Entertainment • Customer Analytics • Social sentiment analysis Confidential and restricted. Do not distribute. (c) Imarticus Learning 26 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • Event Driven Analytics Analytics across Functions Customer Relationship Management (CRM) → Assess • • • Acquire → Grow → Retain Assess: Segmentation, Survey Analytics, market sizing Acquire: Models for response, market mix, pricing Grow and Retain: Models for cross-sell, attrition, profitability Risk Management • Value at risk • Credit scoring • Churn and exception modeling • Fraud detection and modeling; models for anti-money laundering Operations • Capacity planning • Asset optimization • Cost of service optimization • Customer service and experience Infrastructure Management • License management • Warranty management • Incidence (issue) management and failure analysis Supply Chain and Logistics • Demand forecasting • Inventory analysis • Logistics engineering • Scheduling and Planning Human Resource Management • Hiring and attrition analysis • Employee satisfaction surveys • Compensation and benefits modeling Data Analysis Process Step 1: Define Your Questions In your organizational or business data analysis, you must begin with the right question(s). Questions should be measurable, clear and concise. Design your questions to either qualify or disqualify potential solutions to your specific problem or opportunity. Start with a clearly defined problem. A government contractor is experiencing rising costs and is no longer able Confidential and restricted. Do not distribute. (c) Imarticus Learning 27 Post Graduate Program in Data Analytics Data Analysis: Participant Manual to submit competitive contract proposals. One of many questions to solve this business problem might include: ‘Can the company reduce its staff without compromising quality?’ Step 2: Set Clear Measurement Priorities This step breaks down into two sub-steps: 1. Decide what to measure? Using the government contractor example, consider what kind of data is required to answer your question? Need to know the number and cost of current staff and the percentage of time they spend on necessary business functions. In answering this question, you likely need to answer many sub-questions: Are staff currently underutilized? If so, what process improvements would help? Finally, in your decision on what to measure, be sure to include any reasonable objections any stakeholders might have. Example: If staff are reduced, how would the company respond to surges in demand? 2. Decide how to measure it? Thinking about how you measure your data is just as important, especially before the data collection phase. Because you’re measuring process either backs up or discredits your analysis later on. Key questions to ask for this step include: • What is your time frame? (e.g., annual versus quarterly costs) • What is your unit of measure? (e.g., USD versus Euro) • What factors should be included? (e.g., just annual salary versus annual salary plus cost of staff benefits) Step 3: Collect Data With your question clearly defined and your measurement priorities set, now it’s time to collect your data. As you collect and organize your data, remember to keep these important points in mind: • Before you collect new data, determine what information could be collected from existing databases or sources on hand. Collect this data first. • Collect this data first. • Determine a file storing and naming system ahead of time to help all tasked team members collaborate. • This process saves time and prevents team members from collecting the same information twice. If you need to gather data via observation or interviews, then develop an interview template ahead of time to ensure consistency and save time. Keep your collected data organized in a log with collection dates and add any source notes as you go (including any data normalization performed). This practice validates your conclusions down the road. Step 4: Data Cleaning Data cleaning is also called data cleansing or scrubbing. It deals with detecting and removing errors & inconsistencies from data in order to improve the quality of data. Data quality Confidential and restricted. Do not distribute. (c) Imarticus Learning 28 Post Graduate Program in Data Analytics Data Analysis: Participant Manual problems are present in single data collections, such as files and databases. E.g. Misspellings during data entry, missing information or other invalid data. A variable called GENDER would be expected to have only two values; a variable representing HEIGHT in inches would be expected to be within reasonable limits. Some critical applications require a double entry and verification process of data entry. Whether this is done or not, it is still useful to run your data through a series of data checking operations. Step 5: Analyse Data After collecting the right data from Step 1, it’s time for deeper data analysis. Begin by manipulating your data in a number of different ways, such as plotting it out and finding correlations or by creating a pivot table in Excel. A pivot table lets you sort and filter data by different variables and lets you calculate the mean, maximum, minimum and standard deviation of your data. Just be sure to avoid the five pitfalls of statistical data analysis. As you manipulate data, you may find you have the exact data you need, but more likely, you might need to revise your original question or collect more data. Either way, this initial analysis of trends, correlations, variations and outliers helps you focus your data analysis on better answering your question and any objections others might have. During this step, data analysis tools and software are extremely helpful. Visio, Minitab and Stata are all good software packages for advanced statistical data analysis. In most cases, nothing quite compares to Microsoft Excel in terms of decision-making tools. If you need a review or a primer on all the functions Excel accomplishes for your data analysis, Harvard Business Review class is recommended. Step 6: Interpret Results After analyzing your data and possibly conducting further research, it’s finally time to interpret your results. As you interpret your analysis, keep in mind that you cannot ever prove a hypothesis true: rather, you can only fail to reject the hypothesis. Meaning that no matter how much data you collect, chance could always interfere with your results. As you interpret the results of your data, ask yourself these key questions: • Does the data answer your original question? How? • Does the data help you defend against any objections? How? • Are there any limitation on your conclusions, any angles you haven’t considered? If your interpretation of the data holds up under all of these questions and considerations, then you likely have come to a productive conclusion. The only remaining step is to use the results of your data analysis process to decide your best course of action. Confidential and restricted. Do not distribute. (c) Imarticus Learning 29 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Data Analysis & Architecture Exploratory Data Analysis (EDA) Exploring data is a critical step in analyzing data. Exploratory Data Analysis (EDA) is used to systematically identify relations between variables when there are no a prior expectations about those relations. Introduction It is a method of looking at data that does not include formal statistical modelling and inference. EDA does not involve predictive modelling. Focus of EDA is not a set of techniques but the attitude or philosophy about how data analysis should be carried out. EDA is not just about creating data summaries, which are passive but EDA uses the data as a “window” to look into the heart of the business process that generated the data. Why Do You Need EDA? • Detect mistakes • Check assumptions • Perform preliminary selection of appropriate models • Determine relationships among the explanatory variables • Assess the direction and rough size of relationships between explanatory and outcome variables Goals of EDA • Assist in looking at large data by hiding certain aspects of the data while making other aspects more clear. • Identify systematic relations between variables when there are no (or not complete) a prior expectations as to the nature of those relations. • Maximize insight into a data set. • Insight implies detecting and uncovering underlying structure in the data; it is also important to know what’s not in the data. • Extract important variables • Detect Outliers and anomalies • Test underlying assumptions • Determine optimal factor settings Categories of EDA 1. Univariate graphical 2. Univariate Non-graphical 3. Multivariate graphical 4. Multivariate Non-graphical Confidential and restricted. Do not distribute. (c) Imarticus Learning 30 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Exploratory data analysis was promoted by John Tukey to encourage statisticians to explore the data, and possibly formulate hypotheses that could lead to new data collection and experiments. EDA is different from initial data analysis (IDA), which focuses more narrowly on checking assumptions required for model fitting and hypothesis testing. Also checks while handling missing values and making transformations of variables as needed. EDA encompasses IDA. EDA: Univariate Analysis For each variable, you can look at Mean, Median, and Quartiles, Frequency Distribution (Distribution %), Distribution/Graphs (Scatter Plot and Box Plot), Check for missing values & default values, Check for Outliers and Maximum and Minimum values. Case Study – Home Prices The data from a random sample of records of resale of homes from Feb 15 to Apr 30, 1993 is captured from the files maintained by the Albuquerque Board of Realtors. Data captured includes: • PRICE: Selling price ($hundreds) • SQFT: Square feet of living space • AGE: Age of home (years) • COR: Corner location (1) or not (0) • NE: Located in northeast sector of city (1) or not (0) • TAX: Annual taxes ($) • FEATS: Number out of 11 features Univariate Analysis The Univariate analysis of the variable price gives the below result: Multivariate Analysis Exploring data requires you to consider Cross-tab between two variables or more, Comparisons with portfolio or business metrics and Comparisons across groups. Confidential and restricted. Do not distribute. (c) Imarticus Learning 31 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Exploring data requires you to consider: • Slicing and Dicing of data, Data roll-up, and Drill-down: Multidimensional view of data to understand the multi-directional relationships. • Variable transformation: Data in the current form may not give the insight, but transformation could magnify the relations for our understanding. Creation of new variables or relationships that may not be captured within the data, but can be derived from the available data. This is an important question of concern to citizens as a policy matter as well as a personal financial concern. Transformation A linear relation is now better explained with the log transformed variables. Confidential and restricted. Do not distribute. (c) Imarticus Learning 32 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Data Treatment EDA will reveal data issues and need for better explanatory variables. The data treatment to prepare the data for further analysis includes: • • • • Missing values and outlier treatment Default values Variable transformations Derived variable creation Outliers Outliers are extreme values that can occur on both sides (minimum or maximum). In a normal distribution, 0.4% are Outliers (>2.7 SD) and 1 in a million is an extreme Outlier (>4.72 SD). Outlier treatment can be done by: Deleting the Outliers and Trimming or capping the Outliers. Average public teacher pay and spending on public schools per pupil in 1985 for 50 states and the District of Columbia were reported by the Albuquerque Tribune. Outlier Capping Change the Outlier to another value (Trimming) • Value can be 3SD from the mean • Value can be the 99th percentile and so on Make the Outlier equal to the next highest value • Requires a graphical representation to see what Outliers are • Which value should it be trimmed to Confidential and restricted. Do not distribute. (c) Imarticus Learning 33 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Default Values Default values are numbers used in a variable that have some other meaning. For example, the credit score of a person might vary from 50 to 100, but, if the card was inactive, the card might be assigned different values, such as: • 10 to denote inactive for 2 years • 20 to denote inactive for 12 months • 30 to denote inactive for 6 months and so on It is important to know that these are default values and also to treat these values appropriately before beginning any analysis. Default Values: Treatment Treatment of default values can vary by the variable and what the default values mean. Defaults can be treated as missing. Can also be treated as the lowest or highest values. In some cases, leave the values as is and in any analysis, treat these observations differently (such as, creating a subclass of the inactive accounts). Missing Values What to do if certain values are missing? • Find the variables with missing values • Determine the % of missing values for each variable • Understand the variable - Does the missing value have a meaning or is it a mistake? • Treat the missing values Missing Value: Treatment If it is a mistake and over 50% of the variable has missing values, delete the variable. If the missing values have a meaning (if missing means zero, treat the missing as zero). The missing value can be treated to minimum, maximum, or Mean or Median depending on the variable and our understanding of the variable. Skewed Distributions Confidential and restricted. Do not distribute. (c) Imarticus Learning 34 Post Graduate Program in Data Analytics Data Analysis: Participant Manual The waiting time at the teller in a bank is a skewed distribution. The transaction volumes are typically very high in the first 10 days of the month when most people come to collect their salaries, make payments, and so on. It tapers off towards the end of the month. Treatment time in the hospital emergency ward is also a skewed distribution. It is not a data issue, but the nature of the process at the emergency room makes the treatment time distribution skewed. Case Study - Transforming Skewed Data • Background: Cloud seeding is done to increase rainfall. Clouds are seeded with Silver Nitrate to see the clouds. • Experiment: An experiment was conducted to see if indeed cloud seeding would increase rainfall. • Methodology: Clouds were randomly selected for seeding and rainfall data was captured in acre-feet. Univariate Analysis The histogram of the two data sets show that data is skewed in both the cases. Transformed Data A log transformation is used to reduce the huge variance in the data. The histogram of the log transformed data. Confidential and restricted. Do not distribute. (c) Imarticus Learning 35 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Derived Variable A derived variable is a variable, which is calculated from information from one or more other fields. The derived variable should have a meaning and relevance in the context of the analysis. It cannot be created out of any set of variables. Example 1. Retail: Based on the date and the customer visited a store, derived variables could be the days since last visit or months since last visit (This uses the current date and the date of last visit.) 2. Credit Card: Percentage utilization on a customer’s credit card based on available balance. Sales to Balance or Sales to Payment ratio. Summary • EDA is an approach to data analysis. • EDA involves inspecting data without any assumptions. • EDA build a strong understanding of the data, issues related to either the data or the process. It is a systematic approach to discover the story of the data. • The four categories of EDA are: 1. Univariate (One variable) non-graphical 2. Univariate graphical 3. Multivariate (Multiple Variables) non-graphical 4. Multivariate graphical • Outliers are extreme values that can occur on both sides (minimum or maximum) • Default values are numbers used in a variable that have some other meaning • A derived variable is a variable, which is calculated from information from one or more other fields Data Architecture What is Data Architecture? Data architecture comprises of models, policies, rules, or standards that govern: What (and how) data is collected? • How data is stored? • How data is arranged? • How data is integrated? • How data is put to use? Data is usually one of the key architecture domains that form the pillars of an enterprise architecture. Components of Data Architecture • Management Information Systems: Decision Support, Business Intelligence, Reporting • Aggregation: Databases Confidential and restricted. Do not distribute. (c) Imarticus Learning 36 Post Graduate Program in Data Analytics Data Analysis: Participant Manual • • • Analytical Processing: Dimensional Modeling Transformation: Online Analytical Processing (OLAP) Generation: Online Transaction Processing(OLTP), ATM Machines, Billing at Retail Enterprise Data Architecture An integrated view of architecture for data in an enterprise including structured and unstructured data. Data Architecture Confidential and restricted. Do not distribute. (c) Imarticus Learning 37 Post Graduate Program in Data Analytics Data Analysis: Participant Manual OLTP and OLAP Systems OLTP Characteristics OLAP Characteristics Confidential and restricted. Do not distribute. (c) Imarticus Learning 38 Post Graduate Program in Data Analytics Data Analysis: Participant Manual OLTP vs. OLAP How is Data Stored? Database • Collection of relevant data • Persistent, logically coherent collection of inherently meaningful data, relevant to some aspects of the real world Database Management System • A database management system is a collection of programs that enable users to create and maintain a database • A relational database management system (RDBMS )is a database that treats all of its data as a collection of relations Database Management System Confidential and restricted. Do not distribute. (c) Imarticus Learning 39 Post Graduate Program in Data Analytics Data Analysis: Participant Manual A Database Schema Database Types • Flat-file • Hierarchical • Network • Relational • Object-oriented • Object-relational Subject Orientation of a Data Warehouse Confidential and restricted. Do not distribute. (c) Imarticus Learning 40 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Non-Volatility of a Data Warehouse Time Variant Data Warehouse Multi-Dimensional Data Data is interpreted as measurement along multiple dimensions. A central measurement is chosen (such as value of sales) and projected along all possible dimensions. The multidimensional data model is designed to solve complex queries in real time. The multidimensional data model is an integral part of OLAP. Architectural Components of MDM Technology to profile, consolidate and synchronize the master data across the enterprise. Applications to manage, cleanse, and enrich the structured and unstructured master data. Confidential and restricted. Do not distribute. (c) Imarticus Learning 41 Post Graduate Program in Data Analytics Data Analysis: Participant Manual Information Architecture Key Questions for MDM Design 1. Who are the business clients? 2. What are their business objectives? 3. What business processes are performed to achieve those business objectives? 4. What data is necessary to perform the business processes? 5. Is the absence of a high quality unified view of that data impeding business success? 6. If so, what are the quality, consistency, completeness, and synchronization characteristics that will close any perceived gap? MDM Domains • Product Master Data • Material Master Data • Customer Master Data • Supplier Master Data • Contract Master Data • Financial Master Data • Employee Master Data Master Data Management Master Data Management vs. Data Warehousing Master Data Management and Data Warehousing have a lot in common. For instance, the effort of data transformation and cleansing is very similar to an ETL process in data warehousing. They can use the same ETL tools. MDM and data warehousing fall into the same project. Confidential and restricted. Do not distribute. (c) Imarticus Learning 42 Post Graduate Program in Data Analytics Data Analysis: Participant Manual What is different between the two? • Different Goals • Different Types of Data • Different Reporting Needs • Data Usage Location MDM Vs. DW: Different Goals Master Data Management The main purpose of MDM is to create and maintain a single source of truth for a particular dimension within the organization. MDM requires solving the root cause of the inconsistent metadata, because master data needs to be propagated back to the source system. Master Data Management is only applied to entities and not transactional data. The easiest way to think about this is that MDM only affects data that exists in dimensional tables and not in fact tables. Data Warehouse The main purpose of a data warehouse is to analyze data in a multidimensional fashion. Solving the root cause is not always needed, as it may be enough just to have a consistent view at the data warehousing level rather than having to ensure consistency at the data source level. Data warehouse includes data that are both transactional and nontransactional in nature. In a data warehousing, environment includes both dimensional tables and fact tables. MDM Vs. DW: Different Reporting Needs & Data Location Master Data Management The reporting needs are very different. It is far more important to be able to provide reports on data governance, data quality, and compliance, rather than reports based on analytical needs. In master data management, on the other hand, we often need to have a strategy to get a copy of the master data back to the source system. Confidential and restricted. Do not distribute. (c) Imarticus Learning Data Warehouse It is important to deliver to end users the proper types of reports using the proper type of reporting tool to facilitate analysis. Usually the only usage of this "single source of truth" is for applications that access the data warehouse directly, or applications that access systems that source their data straight from the data warehouse. Most of the time, the original data sources are not affected 43 Post Graduate Program in Data Analytics Data Analysis: Participant Manual This poses challenges that do not exist in a data warehousing environment Confidential and restricted. Do not distribute. (c) Imarticus Learning 44
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )