198 IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, VOL. 12, NO. 2, APRIL-JUNE 2019 Interpretable Multiview Early Warning System Adapted to Underrepresented Student Populations Alberto Cano , Member, IEEE and John D. Leonard, Member, IEEE Abstract—Early warning systems have been progressively implemented in higher education institutions to predict student performance. However, they usually fail at effectively integrating the many information sources available at universities to make more accurate and timely predictions, they often lack decision-making reasoning to motivate the reasons behind the predictions, and they are generally biased toward the general student body, ignoring the idiosyncrasies of underrepresented student populations (determined by socio-demographic factors such as race, gender, residency, or status as a freshmen, transfer, adult, or first-generation students) that traditionally have greater difficulties and performance gaps. This paper presents a multiview early warning system built with comprehensible Genetic Programming classification rules adapted to specifically target underrepresented and underperforming student populations. The system integrates many student information repositories using multiview learning to improve the accuracy and timing of the predictions. Three interfaces have been developed to provide personalized and aggregated comprehensible feedback to students, instructors, and staff to facilitate early intervention and student support. Experimental results, validated with statistical analysis, indicate that this multiview learning approach outperforms traditional classifiers. Learning outcomes will help instructors and policy-makers to deploy strategies to increase retention and improve academics. Index Terms—Educational data mining, early prediction, student performance, multi-view learning, genetic programming I. INTRODUCTION E ARLY prediction of student performance is a challenging but crucial task in modern higher education [1], [2]. Early warning systems have been demonstrated to be powerful tools for the early identification of students at risk of failure or dropout [3], [4]. However, there is a large number of factors that influence the learning and performance of students [5], [6]. Socio-demographic factors play an important role as predictors of academic success [7]. Studies have identified associations between school grades, socio-economic deprivation, neighborhood, gender, race, and parental education, among others, and at-risk students. Manuscript received September 10, 2018; revised March 17, 2019; accepted April 9, 2019. Date of publication April 15, 2019; date of current version June 17, 2019. This work was supported in part by an Amazon AWS Machine Learning Research Award and in part by the VCU Presidential Research Quest Fund under the project “Interpretable Data Mining Models for Early Prediction of Student Performance and Dropout.” (Corresponding author: Alberto Cano.) The authors are with the Department Computer Science, Virginia Commonwealth University, Richmond, VA 23284, USA (e-mail: acano@vcu.edu; jleonard@vcu.edu). Digital Object Identifier 10.1109/TLT.2019.2911079 Learning Management Systems (LMS), such as Blackboard or Moodle, are widespread e-learning platforms for teaching and learning, and comprise large collections of student data [8], [9]. Many research works have demonstrated the power of data from such systems to predict student performance [10]–[12]. However, there is much criticism on how these systems are built and employed by educators, especially the lack of information provided about the motivation of the performance alert and the best way to provide early intervention. There are three main critiques to existing early warning systems in education: 1) Learning from data is often limited to a single view of the e-learning system, many times conducting courseindependent analysis, thus missing data contextualization, multi-course concurrency, and multi-year student tracking, as well as not being able to access student records allocated in other isolated information systems [13]. 2) Decision-making reasoning is not explained. Early warning systems often perform as a black-box, simply emitting performance alerts but providing no means to understand or trace the motivation of the predictions [11]. Without knowing the risk factors, educators cannot adapt or improve the curricula, nor can they provide a personalized intervention for the student. 3) Idiosyncratic student subpopulations are not considered. Early warning systems traditionally view the student body as a single homogeneous group, ignoring the particularities of underrepresented heterogeneous subgroups [14]. Specifically, the systems do not account for the nature of different student subpopulations, e.g. minorities, first-generation students, freshmen, transfer students, adult learners, and international students, each of which face their own challenges and have documented persistent performance gaps. The algorithms and systems are frequently biased towards the general student body (majority class) and struggle to achieve accurate predictions on underrepresented student populations (minority classes) [15]. While the impact of small minority groups may not be quantitatively relevant in the overall performance statistics, the unique challenges facing these underrepresented groups prompt an increasing need to focus early warning analysis on them, especially in the diverse student bodies of highereducation institutions. 1939-1382 ß 2019 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See http://www.ieee.org/publications_standards/publications/rights/index.html for more information. Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS This paper addresses the technical components of the aforementioned issues and introduces a multi-view early warning system empowered by an interpretable rule-based Genetic Programming classifier, aimed at the early identification of students at risk of failure, course withdrawal, and dropout. Specifically, we address student populations that traditionally having greater difficulties and performance gaps, such as underrepresented minorities, first-generation students, freshmen, transfer students, adults, and international students. Genetic Programming has been widely employed for building interpretable classification models in education [16]–[19]. Our approach builds an incremental model on a weekly basis as more data becomes available in the e-learning platforms, updating the learning model rather than rebuilding from scratch. The system has been implemented and evaluated on student data from the Virginia Commonwealth University College of Engineering, renowned for its large and diverse student population. This paper addresses the design and experimental evaluation of the early warning system compared to traditional and multiview classifiers. The data and results will allow education experts from the Virginia Commonwealth University Center for Teaching and Learning Excellence to resolve a problem of vital societal importance: improving learning, performance, and retention of students to foster higher graduation rates among the diverse student population. It seeks to understand how students learn in higher education, what motivates performance gaps and dropout rates, and how early warning systems can contribute to improving student performance, and to increasing retention and graduation rates. This research develops next-generation interpretable multi-view educational data mining algorithms for modeling student learning in technologically-rich environments where massive collections of interrelated and distributed data are available. Interpretable models in the form of classification rules are assessed and validated by educational experts, extracting previously unexplored findings from university databases that will support policy-makers’ decision making. A thorough experimental study evaluates and compares the performance of the multi-view approach with other traditional and multi-view interpretable classifiers. The performance evaluation considers imbalanced metrics, taking into account the imbalanced data distribution of underrepresented groups. Results are validated using non-parametric statistical analysis. The rest of the paper is organized as follows. Section II reviews related works on early warning systems and predicting student performance. Section III describes the system’s architecture and methodology. Section IV details the experimental study and discusses the results. Finally, Section V collects the lessons learned and concluding remarks. II. BACKGROUND Educational data mining [20]–[24] and learning analytics [25]–[28] are multidisciplinary research areas where investigators develop accurate data mining models for predicting student performance. Early warning systems are intelligent information systems based on student data that assist in the early identification of students who exhibit behavior or academic 199 performance that puts them at risk of failure, course withdrawal, or dropout [29]–[31]. These systems utilize available real-time and historical student data to identify students at risk of missing key educational milestones, to diagnose the needs of at-risk students, and to identify early interventions that may help at-risk students get back on track [15], [32]. Research studies have identified socio-demographics, attendance, behavior, class participation, activity in e-learning systems, and course performance as powerful predictors [33]–[36]. Institutions must ensure that intervention and academic support, such as tutoring and advising, are in place for any students identified as being atrisk by an early alert program [37]. The use of early warning systems in higher education, in a systematic fashion, is relatively recent. However, academic performance prediction systems are perceived to be moderately effective, black-box (no comprehensible feedback or explanations are provided), and only useful for administrators conducting posterior analysis [11], [38]. In recent years, Deep Learning techniques have been incorporated into early warning systems, due to their high accuracy [39]. However, this only makes understanding the motivation of the performance alerts more complicated. The challenge, from the data mining perspective, is to develop systems that provide not only early and accurate alerts, but also meaningful justification and explanation of the reasoning behind the system’s decision, making the decisionmaking comprehensible to humans. Consequently, interpretable models in the form of classification rules or decision trees are preferred [16], [19], [40]. This falls within the recent trend of explainable artificial intelligence [41], which seeks an understanding of how the machine learning methods learn and reach their predictions. The early warning system implemented at Purdue University is well known and simplifies performance alerts to traffic light colors [42], [43]. Green indicates that a student is likely to succeed. Predictions are based on all available student background and LMS activity. Students identified as at-risk are referred to academic advisors and counseling. However, as an uninterpretable system, the reason why such alerts are triggered remains unknown to instructors. Interpretability and accuracy are unfortunately conflicting objectives [44], [45]. The more detailed the variables are, the more accurate the created prediction model will be. However, models become so complex that simpler, easier to interpret models are often preferred. Considering the large number of variables represented within an LMS, including the number of visits, downloads, contributions to discussion threads, etc, data mining models which perform feature selection and reduce the complexity of the model to a small number of truly important variables are recommended [46]–[48]. Evolutionary computation, and more specifically Genetic Programming, has been widely used to extract classification rules that predict student performance while providing interpretable models to educators [16]–[19]. Genetic Programming allows for the extraction of rules that provide feedback to instructors on quizzes and courseware [49], [50], the prediction of academic achievement [51], and the extraction of rules from web-based educational systems [52]. More recently, educational process mining aims at making unexpressed Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. 200 IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, knowledge explicit and to facilitate a better understanding of the educational process. Educational process mining uses log data gathered specifically from educational environments to analyze and provide a visual representation of the complete educational process [53]. One of the key findings of a recent survey on early warning systems [54] is that they are most frequently used to address the general student population, but they struggle to achieve accurate predictions in underperforming subpopulations. This undesired behavior is due to the imbalance of the data classes (the percentage of examples belonging to each data class is not proportional). Naturally, the number of students belonging to such underperforming subpopulations is significantly smaller than the general student body. This causes bias in the predictions, especially in the groups that this project focuses on. Consequently, early warning systems have also been used to selectively predict within individual at-risk subgroups, rather than within the entire student population. Freshmen have traditionally been the most focused upon subgroup in early alert programs [55]. However, traditional metrics for evaluating prediction accuracy may be misleading under the presence of imbalanced subpopulations [56], [57]. Therefore, there is a need to develop models specifically addressing student subpopulations that have performance gaps. Predictive models for student performance are often applied to a single course using the local information of the course, rather than simultaneously using the concurrent information of the many courses in which a student is enrolled [58]. The different internal structure of courses makes it difficult to integrate data uniformly [59]. Therefore, multi-view learners are a better approach for integrating the concurrent student information [60], along with the historical data of the course in previous years, to build more accurate and timely predictions. Simultaneous monitoring of student activity in multiple courses not only helps to identify student failure in a given course, but triggers alerts of dropout risk whenever it detects a lack of activity in all courses in which a student is enrolled. Multi-view learning [61] is a recent paradigm that exploits data represented by multiple distinct feature sets, called views, which are obtained from different data sources. The objective is to learn functions to model the views and jointly exploit the redundancy and complementarity of data in the views [62]. The multi-view foundations are based on the consensus and complementary principles of data. The consensus principle maximizes the agreement of distinct views, i.e. the relationships between different subsets of features. The complementary principle advocates for the distribution of the information among the views, i.e., each view contains partial data that may not provide interesting information when analyzed separately, but when merged together provide meaningful and comprehensive knowledge to the learner. Learning classification models from the fusion of multiple views and information sources increases the strength of the classification predictions as compared with the predictions on independent views [63]. However, accomplishing such a task is not straightforward and data from multiple heterogeneous sources cannot be easily integrated into multi-view data. Heterogeneity refers to the definition of the VOL. 12, NO. 2, APRIL-JUNE 2019 domains in each of the views and their distinct cardinalities. The integration of multi-view data is not much different than the data warehousing in an enterprise system, the key here though, is the ability to develop machine learning algorithms that directly learn from the multiple views in real time, extracting meaningful predictions from the whole [60]. Moreover, the robustness against missing values is one of the advantages of multi-view learners since they overcome the limitation of a missing value from one view by gathering information from the other views [64]. Genetic programming has been successfully adapted to multi-view spaces for extracting classification rules [65]. Multi-view learning is formally defined by regularization terms (left) where the first term enforces the agreement of two distinct views on unlabeled examples (x), and the second term evaluates the empirical loss of the labeled data (y) with respect to a loss function V ð; Þ such as (Eq. 1): min X ½f 1 ðxi Þ f 2 ðxi Þ2 þ X V ðyi ; fðxi ÞÞ: (1) i2L i2U III. SYSTEM ARCHITECTURE This section presents the system’s architecture and the multi-view methodology for learning from student data. The system overview is depicted in Figure 1 which represents the agents, information systems, and the multi-view Genetic Programming algorithm involved in the analysis of student data, predicting academic performance, and triggering early alerts in the form of comprehensible classification rules. There are three main reasons for using Genetic Programming in this work. First, it plays very nicely with multi-view learning, as the multi-view construction of the rules is implicit in the definition of the context-free grammar used to build the classification rules. Second, the model learned and evolved is directly interpretable without any further modification. Third, the classification rules evolve naturally with time and the progression of the semester as more data becomes available, i.e., the rules evolve to the new data, rather than building a new classifier from scratch whenever the data changes. A. Multi-View Student Data The first step is to identify information sources and develop a system capable of retrieving live data automatically from the university databases and e-learning sites. The objective is to gather collections of data in existing repositories, such as personal data, GPA, course grades, year of enrollment, etc, allocated in the e-Services platform; course information and performance indicators (attendance to theory and practice lectures, attitude and class participation, polls, grades of different assignments and tasks performed during the course, etc); activity of students in e-learning platforms (Blackboard), which comprise the student records and usage activity in each course (number of clicks, total time in the system and in each of the activities, number of activities, number of visited resources and downloaded, scores on the assessed activities, successes and failures in the questionnaires, number of messages posted and read in forums, grades, the timestamps, etc); and other potential Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS Fig. 1. 201 System overview: data, models, and agents involved. sources of data with more student information (university admission notes, transcripts for transfer students, family data, socioeconomic data, SAT, GRE and TOEFL scores, etc). This information is extremely valuable; from it one may extract knowledge about the students (individually and collectively) that could be used to improve teaching. However, the volume and heterogeneity of the information generated by these systems is so large that it is impossible to perform this information extraction process manually. Therefore, it is essential to have multiple views accessing the university databases and collecting student information in real-time. While most of the data is invariant, data from e-learning systems must be retrieved in real time to access the most up to date student records in courses. Rather than creating a static dataset with the join of the views that would become outdated very soon, our approach uses multi-view genetic programming to build the classification rules, which pull the real time data from each of the views whenever they are evaluated. The next section details the technical implementation of the construction of the multi-view rules using a context-free grammar. B. Multi-View Genetic Programming This section describes the multi-view genetic programming algorithm for learning classification rules. There are three major differences as compared with traditional genetic programming rule learners. First, learning from multiple views implies that the evolutionary process must learn and identify the most relevant attributes in each of the views to maximize the classification accuracy with the data available at a given time. Second, this is not a static problem and therefore more data becomes available on e-learning systems as the semester progresses. In such a scenario, there are two main approaches. The former rebuilds the prediction model from scratch periodically, assuming there is data independence and behaving like a regular static learner. The latter updates the prediction model as more data becomes available, referred to as the data stream mining problem [66]. Data is no longer static but arrives as a stream of real-time data which is processed as soon as it is produced. In this paper, we consider the latter approach as it is more natural to this spatio-temporal problem, and helps us to analyze the evolution and changes of the prediction model through time. Third, there are two sets of rules to be learned, one set focuses on the general student population whereas the other set prefixes the context and specifically focuses on each of the underrepresented student populations to extract rules particularly accurate for such groups. Specifically, the rules for underrepresented populations are built according to race, gender, residency, or status as a freshman, transfer, adult or first-generation student. These groups overlap, e.g., a freshman student who is an African American woman will be covered by the general student population rule set, plus the rules for race, gender, and freshman status, respectively. This section is organized as follows. First, the individual representation for multi-view classification rules is introduced. Second, the genetic operators for evolving the rules are detailed. Third, the iterative rule learning evolutionary model is described. Finally, the fitness function adapted to imbalanced underrepresented groups is described. 1) Individual Representation: Genetic programming evolves a population of individuals represented by a genotypephenotype pair. The genotype encodes a syntax tree, also known as derivation tree, generated through the production rules of a context-free grammar. The phenotype encodes the expression tree function resulting from the syntax tree, and it is represented in the form of an IF-THEN classification rule. Context-free grammars provide a formal definition of the syntactical restrictions in the generation of the classification rules. They provide high flexibility to generate any combination of linear, non-linear, and personalized functions. The use of grammars to generate genetic programming individuals is known as grammar-guided genetic programming, and it has been widely applied to data mining [67]. Figure 2 shows the context-free grammar employed to create the rules in a multiview environment. The grammar G is defined by a tuple (VN , VT , P , S) where VN represents the non-terminal symbols, VT the alphabet of terminal symbols, P the production rules, and S the root symbol. A production rule is defined as a ! b where a 2 VN and b 2 fVT [ VN g . The generation of individuals is a stochastic process in which the production rules Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. 202 Fig. 2. IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, Context-free grammar to generate multi-view classification rules. are randomly selected to produce an initial set of classification rules with varied length (limited by a maximum number of derivations). This guarantees a wide and random exploration of the search space, exploring different sets of features in each of the views. Production rules are applied to obtain a syntax tree (derivation tree) for each individual, in which internal nodes contain only non-terminal symbols, whereas leaf nodes contain only terminals. During the evolution process, genetic operators will employ the grammar to update the components of the classification rule. The grammar provides fast adaptability to new concepts on new more recent data, e.g. concept inversion may be modeled by adding a NOT symbol at the root; the surge of new concepts may be modeled by adding new branches to an existing derivation combined with OR, and old concepts may be forgotten by removing the branch. Example individuals generated using the context-free grammar are illustrated in Figures 3 and 4, showing the tree representation of a classification rule antecedent. The consequent (at-risk, not-at-risk) and the context of the target population (race, gender, residency, or status as a freshman, transfer, adult or first-generation student) are fixed. The genetic programming focuses on the evolution of the most appropriate features identified in the antecedent of the rule, given the consequent and target population. 2) Genetic Operators: Individuals are crossed and mutated to create new solutions along the evolutionary process. Crossover recombines the genetic information from parents in hopes of producing offspring with improved fitness. The crossover operator creates new syntax trees by mixing two random branches from the parent trees. The selective crossover chooses randomly with uniform probability a non-terminal symbol in each of the parents and swaps their subtrees creating two new offspring individuals. Nevertheless, in order to avoid bloating, individuals are not crossed if the resulting individual exceeds a maximum size predefined by the user. Figure 3 illustrates the crossover operation for two parents on a compatible node to produce an offspring by switching the tree branches. Mutation maintains genetic diversity in the population by randomly altering an individual with a relatively low probability. The mutation operator randomly chooses a node (terminal or VOL. 12, NO. 2, Fig. 3. Example of crossover in genetic programming. Fig. 4. Example of mutation in genetic programming. APRIL-JUNE 2019 non-terminal) in the individual’s syntax tree. If the node is terminal, the mutation will replace the current node with another compatible terminal. If the node is non-terminal, the subtree underneath will be replaced by a new randomly generated derivation using the grammar production rules. Figure 4 illustrates the mutation operation on a non-terminal node; the parent branch is removed and replaced by a new subtree randomly generated using the production rules from the grammar. The mutation operator guarantees that the offspring derivation size does not exceed the maximum size. Mutation is essential for forgetting non-coding genetic information and introducing new useful information. Removing tree branches from classification rules modeling old concepts is a useful mechanism for forgetting knowledge that is no longer interesting and is not represented in the current data. On the other hand, introducing new mutated tree branches into a classification rule is useful for modeling new data concepts. Thus, genetic programming rules evolve dynamically according to their adaptation data [68], [69]. 3) Iterative Rule Learning: Genetic programming algorithms usually follow the genetic cooperative-competitive learning approach, in which the rules for the classifier are selected among the best individuals obtained after running the evolutionary process once. However, this approach requires introducing complex mechanisms for maintaining the diversity of the population in order to avoid having all individuals converge to the same area of search space [70]. Conversely, the iterative rule learning approach runs the evolutionary process multiple times, and every time selects the single best individual to become a rule of the classifier. Iterative rule learning has the advantage of preventing the convergence of rules evolving in different runs, and provides diversity by exploiting the search space. Consequently, the rule set for a class includes diverse components which will improve the quality of the predictions. Our approach for Multi-View Genetic Programming (MVGP) follows the iterative rule learning approach and its pseudo-code is introduced in Algorithm 1. The algorithm simultaneously learns two sets of rules, one for the general student body and another for each of the underrepresented populations (categorized according to race, gender, residency, or status as a freshman, transfer, adult or first-generation Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS student), as detailed in Section III-C. These rules are preserved and evolved on a weekly basis throughout the semester as more information becomes available. The algorithm learns and outputs a number of user-parameterized classification rules per class (at-risk, not-at-risk) to provide a diverse voting schema for instances. The rules are initialized according to the context-free grammar defined in Figure 2. In every iteration to obtain a rule, a population of individuals is evolved along a number of generations. This approach intrinsically matches the incremental spatio-temporal flow of e-learning systems, making it easy to learn rules in the very beginning of the course, and adapting them as more information becomes available. In every iteration, new data is received for training and updating the classifier. The algorithm adapts and evolves the classification rules to learn from the current data. The advantage of evolving rules is that they maintain the previously learned knowledge represented in the rules’ axioms, and they adapt to the new incremental data characteristics, possibly reflecting a concept drift in the behavior of students. Algorithm 1: MVGP algorithm for mining rule learning. 1: elitistRules f;g 2: trainData f;g 3: while views:newDataðÞ do 4: ruleBase f;g 5: trainData updateðtrainData [ views:newDataÞ 6: iteration 0 7: while iteration < numberRules do 8: population initializeðsizeÞ [ elitistRules 9: evaluateðpopulation; trainDataÞ 10: generation 0 11: while generation < numberGenerations do 12: parents parentSelectorðpopulationÞ 13: crossed crossoverðparentsÞ 14: offspring mutatorðcrossedÞ 15: evaluateðoffspring; trainDataÞ 16: population selectðpopulation [ offspringÞ 17: generation generation þ 1 18: end while 19: ruleBase ruleBase [ bestRuleðpopulationÞ 20: iteration iteration þ 1 21: end while 22: elitistRules ruleBase 23: end while The population is reinitialized randomly in every iteration, but the single best individual from the previous iteration is maintained to preserve its genetic information, encoding the model learned in previous iterations. On the one hand, the elitist solution from the previous iteration will spread very quickly because it is already adapted to model the same data distribution, avoiding any accuracy penalty due to reinitialization. On the other hand, the new randomly generated individuals will provide genetic diversity to quickly adapt to new data characteristics. This way, in every iteration the population will comprise both sufficient historical and new genetic material, adapting to the new data and preserving previous knowledge. This approach is much more effective and efficient than 203 rebuilding the model from scratch in every iteration every time new data becomes available. Moreover, it is very interesting to analyze the changes in the rules’ conditions when new features appear in the dataset. Importantly, the difficulty of predicting each of the data classes may vary (e.g. some classes may comprise noisy, overlapped or sparse data). Therefore, the voting of the class prediction for each triggered rule for a test instance is weighted by the fitness of the rule. This way, we minimize the possibility that inaccurate rules with low fitness will raise false positives. 4) Fitness Function: The fitness function evaluates the quality of the individuals in the population by checking the coverage of the classification rule across the data instances in the dataset, then comparing the predicted and ground-truth data class. Specifically, it computes the confusion matrix for the true positives (TP ), true negatives (TN ), false positives (FP ), and false negatives (FN ) to maximize the sensitivity, also known as the true positive rate (TPR) (Eq. 2), and specificity, also known as the true negative rate (TNR) (Eq. 3), of the classification rules. The positive class identifies at-risk students. The goal is to maximize the true positives, students correctly classified as at-risk, and minimize the false negatives, students who are at-risk but are not identified by the system, while still not having too many false positives, students incorrectly identified as at-risk. The single-objective fitness function (Eq. 4) is defined as the weighted harmonic mean (F-measure) of sensitivity and specificity. The b parameter (default 1) balances the relevance of the sensitivity and specificity in the weight of the fitness. A value 0 < b < 1 gives more weight to specificity while b > 1 gives more weight to sensitivity. The traditional F-measure (F1) is a private case of the weighted F-measure with b ¼ 1. This function is commonplace in evolutionary algorithms for classification and has been shown to perform well under the presence of imbalanced data [71]. Robustness of algorithms for imbalanced data is a new challenging area. This issue is particularly relevant for our student data. The problem of analyzing subpopulations of students is due to their representation within the data. The collected data is highly imbalanced (the number of students belonging to each subpopulation is low), which introduces additional difficulties in the data processing and evaluation. The problem with imbalanced data is that classical classification algorithms and metrics aim to maximize the overall accuracy, which introduces a bias towards maximizing the accuracy for the majority class (general student body). This bias causes a low sensitivity for the prediction of the minority class (underrepresented student populations at risk). Therefore, algorithms developed must be aware of the imbalanced data distribution and be capable of providing an accurate prediction for both majority and minority classes. Genetic Programming for rule learning has been demonstrated to perform accurate classifications on imbalanced datasets [72]. Therefore, it is an ideal methodology for obtaining accurate rule models within specific student subpopulations. sensitivity ¼ TP T P þ FN Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. (2) 204 IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, TN T N þ FP (3) ð1 þ b2 Þ sensitivity specificity sensitivity þ b2 specificity (4) specificity ¼ fitness ¼ VOL. 12, NO. 2, APRIL-JUNE 2019 C. Addressing the General Student Body and Underrepresented Student Populations One of the major contributions of this paper is to specifically address underrepresented student populations by learning classification rules particularly targeting such groups. This way, educators may be able to analyze and understand the motivation for the performance gaps within these groups and provide personalized intervention on time. The MVGP algorithm is executed in two ways following the iterative rule learning model detailed in Section III-B3. First, general classification rules are obtained for the general student body, making no distinctions on the population categorization. The idea is to extract general rules applicable to all students. Second, specific classification rules are obtained for each of the underrepresented population groups, according to race, gender, residency, or status as a freshman, transfer, adult or first-generation student. The objective is to extract particular rules applicable to the minority classes. Therefore, educators may analyze the similarities and differences between general and specific rules, helping them understand any performance differences. General and specific rules may be triggered when a student is identified as at-risk of failure or dropout. Whenever any of the two sets are triggered the student and instructor will be informed about the risk. If both rule types are triggered, it increases the confidence of the prediction. Both rule sets are evolved and updated along the semester as more information becomes available. Moreover, every new section (a new instance of a course on a new semester or year) inherits the prediction model from the previous section, making the learned knowledge transferable to new students from the very first day as behavior and performance patterns tend to repeat. D. System Interfaces for Student, Instructor, and Staff The multi-view genetic programming model described in previous sections is the core of the early prediction system. However, in real-world environments, it is necessary to provide a user-friendly interface to facilitate the daily use of the system, by end-users, in a comprehensible manner. In our educational early warning system there are three main agents involved: students, instructors, and staff. Each requires a particular view of the data and the predictions. We implemented a software module in the Blackboard e-learning system to present valuable feedback to the agents and facilitate intervention. Nevertheless, the usefulness of this tool to intervene is up to the students and instructors. Figure 5 illustrates the student interface on Blackboard. Each student may see their performance and compare with the aggregated performance of their peers in the course. The ranking system based on the current grade encourages students to Fig. 5. Blackboard module: student interface. improve, while keeping this information private from other students. Students identified as at-risk are provided with a message to contact the instructor for additional support. Moreover, whenever such a performance alert is triggered, both the student and the instructor receive an email. Students are not presented with the rules, so as to prevent them from cheating the system’s reasoning. Figure 6 illustrates the instructor interface on Blackboard. This view allows an instructor to analyze the details of each student, particularly those identified as at-risk, as well as the average performance of the course, course statistics, population statistics, and the comparison with the latest section of the same course. This allows the instructor to evaluate the performance of the students at the same week of class between different sections/semesters to identify significant performance deviations and provide early intervention. Moreover, the classification rules for identifying at-risk students are shown to instructors, both the general rules and per subpopulation, so that they can understand the at-risk factors, and support the students accordingly. Figure 7 illustrates the staff interface on Blackboard. Administrators, and particularly academic advisors, have full access to the individual and collective results of students (under their academic unit) identified as at-risk of failure for a given course, or labeled as at-risk of dropping out of college due to having multiple alerts on multiple courses simultaneously. This allows school counselors to provide timely intervention and personalized support to assist each student with the courses involved. IV. EXPERIMENTAL STUDY This section presents the experiments carried out to evaluate the performance of the early warning system. First, the dataset description, statistics, and structure are detailed. Second, the experimental setup including compared algorithms, performance metrics, and non-parametric statistical analysis are presented. Third, experimental results are provided, analyzed, and validated through statistical tests. Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS Fig. 7. 205 Blackboard module: staff interface. TABLE I VCU COLLEGE OF ENGINEERING STUDENT POPULATION Fig. 6. Blackboard module: instructor interface. A. Dataset Description The early warning system has been implemented and evaluated on student data from the Virginia Commonwealth University (VCU) College of Engineering. The college comprises a large and diverse student body in terms of programs, majors, race/ethnicity, socio-demographics, and in-state/ out-of-state/international students. Table I summarizes the student count for each of the programs at the College of Engineering. Dropout rates are not only significant in undergraduate but also in graduate studies at VCU. Tables II and III show student demographics. According to the VCU Office for Planning & Decision Support statistics, the 6-year graduation rate is 63% and the first year retention is 83%. However, there are persistent performance gaps among these underrepresented groups. First-generation students represent a large group who encounter specific challenges and have low graduation rates. These challenges include the lack of family members who graduated and who could offer advice or support about college education, in turn making the transition to college life more difficult. These hardships contribute to a lower rate of college completion than students who have at least one parent with a four-year degree. Freshman may have a good understanding of what they need to do to be accepted into college, but may not fully understand what it takes to successfully transition from high school to higher education and successfully graduate. Transfer students at VCU are primarily coming from community colleges, which often encounter higher failure and dropout rates, and they represent a large portion of our undergraduate population (13% overall and 30% in computer science). Statistics reflect the persistent performance gap between students who transferred from a community college and those who successfully completed the first two years at a four-year college [73]. The retention of adult students is an issue of growing concern in many educational institutions. This population of students, enrolled at the age of 25 or older, is usually demographically different than the general student body (e.g. veterans, full-time young professionals) and often Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. 206 IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, TABLE II RACE / ETHNICITY POPULATION DISTRIBUTION apply to distance education and online learning degrees, which in turn, have greater difficulties and higher dropout rates. Finally, international students face additional challenges such as language, social differences, and cultural barriers, which handicap their success in higher education abroad. Therefore, students belonging to any of these underrepresented groups according to race, gender, residency, or status as a freshman, transfer, adult or first-generation student, are analyzed with further specific rules. The most relevant attributes included in the dataset come from the multi-view data collection described in Section III-A. The system has been evaluated using student data from the Fall 2017 and Spring 2018 semesters, while having access to students’ historical data from the previous five years. B. Experimental Setup This section presents the algorithms evaluated, performance metrics, and non-parametric statistical tests employed to compare the results of the algorithms. 1) Algorithms: The algorithms evaluated can be classified into two categories: single-view and multi-view methods. Single-view methods comprise traditional learners. In order to perform a fair comparison, we employ only interpretable methods (rule-based or tree-based). Specifically, these singleview methods implemented on Weka [74] are compared: JRip [75]: a propositional rule learner, Repeated Incremental Pruning to Produce Error Reduction (RIPPER) proposed by W. Cohen as an optimized version of IREP. J48 [76]: C4.5 decision tree by R. Quinlan, an extension of the ID3 algorithm, using information entropy. ICRM [45]: An Interpretable Classification Rule Mining Algorithm for extracting interpretable classification rules from educational data. We also compared our multi-view genetic programming approach with another multi-view method based on ensemble classification: MVMI [64]: multi-view ensemble of classifiers. It selects the most appropriate base learner for each of the views and combines the predictions by majority voting. MVGP: multi-view genetic programming introduced in this paper. Main parameters: Population size: 25, Number of generations: 50, Rules per class: 5, b: 1, Crossover probability: 0.9, Mutation probability: 0.1. VOL. 12, NO. 2, APRIL-JUNE 2019 TABLE III STUDENT DEMOGRAPHICS Parameters for all algorithms were the recommended values reported by their authors and other studies in this area. MVGP has been implemented in the JCLEC software [77] for classification rule learning. Experiments were run on an Intel i7-4790 CPU at 3.6 GHz with 32 GB-DDR3 RAM on a Linux CentOS 7 system. 2) Performance Metrics: There are many performance metrics for evaluating classification algorithms. Different metrics allow us to observe different behaviors, which increases the strength of the empirical study in such a way that more complete conclusions can be obtained from different (complementary, not opposite) deductions [78], [79]. These measures are based on the values of the confusion matrix, where each column of the matrix represents the count of instances in a predicted class, while each row represents the number of instances in an actual class. The standard performance measure for classification is accuracy, which is the number of successful predictions relative to the total number of classifications. However, accuracy may be misleading when data classes are strongly imbalanced since the all-positive or all-negative classifier may achieve high accuracy. Real-world problems frequently deal with imbalanced data, and predicting the performance of student populations is one of them. Therefore, other metrics are preferred. Sensitivity and specificity were introduced in Eqs. 2 and 3. The geometric mean (GM) in Eq. (5) attempts to maximize and balance the accuracy of the two classes by correlating both objectives: pffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi Geometric mean ðGMÞ ¼ sensitivity specificity: (5) The kappa statistic evaluates the competence of a classifier by contrasting the successful predictions and the statistical distribution of the data classes. It ranges from 1 (total disagreement) through 0 (default probabilistic classification) to 1 (total agreement), computed as in Eq. (6): Kappa ¼ N Xk Xk x x x ii i¼1 i¼1 i: :i Xk N2 x x i¼1 i: :i Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. (6) CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS Fig. 8. 207 Classification performance of algorithms (weekly). where xii is the count of cases in the main diagonal of the confusion matrix (successful predictions), N is the number of examples, and x:i , xi: are the column and row total counts, respectively. Kappa penalizes all-positive or all-negative predictions, which is especially useful in problems with multiclass imbalanced data, like our student populations. 3) Statistical Analysis: The area under the receiver operating characteristic curve (AUC) is a commonly used evaluation measure for imbalanced classification. The curve presents the trade-off between the true positive rate and the false positive rate. A classifier generally misclassifies more negative examples as positive the more true positive examples are captured. AUC is computed by means of the confusion matrix values: true positives (TP ), false positives (FP ), true negatives (TN ), and false negatives (FN ), and it is calculated relating the true positive and false positive ratio as in Eq. (7): TP FP 1 þ TPrate FPrate 1 þ TP þFN FP þTN ¼ AUC ¼ 2 2 (7) Non-parametric statistical tests are used to analyze and validate the results [80]. First, the algorithms are ranked according to the best results. To evaluate whether there are significant differences in the results, the Bonferroni–Dunn test is used to find the significant differences occurring between algorithms. The Wilcoxon rank-sum test is used to perform multiple pairwise comparisons among the algorithms [80]. C. Results This section is organized as follows. First, the overall performance of the system is presented and evaluated with comparable methods in the literature. Second, a detailed analysis of specific student populations is provided. Third, some example rules learned by the multi-view genetic programming are shown. Figure 8 shows the performance metrics (accuracy, sensitivity, specificity, geometric mean, Kappa, and AUC) for each of the five algorithms (Jrip, J48, ICRM, MVMI, and MVGP). The analysis is conducted on a weekly basis based on the current information contained in the system. The results show the overall performance of the classification system on the whole student population, considering both the generic rules for the general student body and the specific rules for the underrepresented minorities. Specifically, there are three main points in time that are interesting to analyze in detail. First, on week 1, the information system comprises only information about the students’ past performance in previous courses, their background information, and the history of the course in previous sections. There is no student activity on Blackboard so far; the course just started. However, it is interesting to point out how algorithms, especially multi-view based ones, are able to perform a relatively accurate prediction of the students’ performance. Importantly, due to the imbalanced nature of the data classes, it is necessary to analyze the sensitivity and specificity separately. While specificity values are relatively high (it is easy to predict that students who previously succeeded will continue to succeed), it is significantly more challenging to improve the sensitivity (at-risk students not being identified by the system) at the beginning of the course. This means that it is difficult to extract reasons why students struggle. Second, on weeks 6-7 there is a blip in the accuracy caused by a drop in the specificity. The reason for this drop is that midterms take place at that time, but the grades are not input Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. 208 IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, VOL. 12, NO. 2, APRIL-JUNE 2019 TABLE IV AVERAGE AND RANK OF CLASSIFICATION ALGORITHMS Fig. 9. into the system until week 8. These weeks are a period of stress and instability for students, their behavior and activities in the e-learning system are unstable. Some students choose to download all materials and study offline for their midterms, not logging in the platform while studying. On the other hand, other students do the opposite and are continuously accessing the materials online, typically downloading the slides over and over again. This instability of the pattern of student behavior introduces instability in the predictions, particularly for the specificity, i.e., the system triggers a larger than expected number of false positives of students at risk during these weeks of instability because it cannot distinguish between students actively studying with offline materials and students who have stopped studying and interacting with the e-learning system. Conversely, the sensitivity is relatively stable, especially for our multi-view approach, correctly identifying students at risk. After midterm grades are input in week 8, the overall performance of all methods increases, as it is one of the most reliable factors in predicting the students’ final performance. When such data is included in the system, the accuracy of all algorithms significantly increases, as does the specificity. Midterm results allow students to make the final decision on whether to continue the course or withdraw. Therefore, by taking into account the system’s input, instructors must intervene whenever a student is identified as at-risk of withdrawal. However, the timing here is critical, as students are barely given one week to make such a decision. Students achieving poor performance but not withdrawing are highly likely to fail the course unless proper intervention is provided at this time by the instructor. Third, on weeks 16 and 17 students are preparing for the final exams. The system comprises all data from students activities and can make a much more accurate prediction of the students’ final performance. However, from the point of view of early prediction, it is too late to prevent students dropping out of courses. They often decide (without informing instructors) they’ll drop a course, usually due to burnout or an overflowing workload from a large number of courses. Due to the imbalanced nature of the data classes, it is necessary to analyze the performance of Kappa and AUC (Figures 8e Bonferroni-Dunn test. and 8f). Performance differences between multi-view and single-view learners are especially identified on Kappa. This proves the superiority, from integrating many information sources, of the multi-view learners. While MVMI and MVGP achieve similar performance for Kappa, superior performance of MVGP is identified on AUC. Table IV shows the average and rank of the classification algorithms for each of the metrics evaluated. There is a clear performance gap between multi-view learners and traditional single-view methods, especially in the context of imbalanced metrics. MVGP achieves similar accuracy than MVMI but consistently better Kappa and AUC. The ranks (the lower the better) indicate similar results. The meta-rank evaluates the algorithms’ average performance across all metrics, giving a perspective of the overall performance. Jrip obtains significantly worse results across all metrics, and is not recommended for extracting rules from educational data systems. The Bonferroni-Dunn test indicates that there are statistically significant performance differences whenever the rank of the methods, as compared with the control algorithm, differ by more than 1.3547 for an a ¼ 0:05, i.e., 95% statistical confidence. Figure 9 illustrates the rank of the algorithms. The critical difference area is represented by the gray rectangle. Algorithms right of the critical difference rectangle are the ones with significantly worse results. Table V shows the p-values for the Wilcoxon pairwise nonparametric test when comparing the results of MVGP vs the other approaches. p-values < 0.05 (95% confidence) indicate statistically significant differences. According to the test, there are remarkable differences, including for sensitivity to at-risk students, but there are none for specificity on J48 and MVMI, and Kappa on MVMI. Table VI shows the detailed results for the three most representative underrepresented subpopulations categorized according to (a) race, (b) residency status, and (c) freshman status. The performance of all methods across all metrics is generally lower than in the general student body shown in Table IV. This is due to the increased complexity of classifying these underrepresented groups. While the classification according to race and residency is generally stable, there is a Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS 209 TABLE V WILCOXON TEST (p-VALUES) TABLE VI Fig. 10. Example rules obtained by MVGP. RESULTS FOR UNDERREPRESENTED SUBPOPULATIONS for race/ethnic minorities, socio-demographics play an important role in the decision making. For these students, having not been admitted to the honors college, being labeled as firstgeneration, and residing in a zip code associated with a lower income and large African American population indicate they are at-risk. These rules are a sample of a larger set of rules learned weekly both for at-risk and not-at-risk students. V. CONCLUSIONS significant gap in the classification of at-risk freshman students, as indicated by the very low sensitivity of methods in Table VIc. This poor performance may be explained by two main factors. First, the lack of historic information, i.e., since freshmen students have just started college, there is no information on their performance in past courses. This may be improved if high-school transcripts were included as part of the students’ historic data. Second, freshman, especially in engineering, are known for high dropout rates as they encounter core courses on math, statistics, and physics, which are challenging for freshmen. Nevertheless, the performance of multi-view learners is superior to traditional classifiers. Figure 10 shows two example rules obtained with the data available at week 6 (before the midterm grades are input into the databases) for students identified as at-risk, for the general student body and the subpopulation categorization according to race, respectively. Rules not only take into account the historic performance of the student, but also the current weighted average grade of the student in the current course. In the case In this paper, we have introduced an early warning system aimed at triggering early performance alerts for students at risk of failure, course withdrawal, and dropout. The contribution included developing a multi-view approach to collect data from multiple student data repositories and combine the information to make more accurate predictions. The machine learning approach was based on multi-view genetic programming to automatically extract and learn interpretable classification rules, making it comprehensible for instructors and facilitating early intervention. The system specifically addressed student populations traditionally underrepresented in higher education. Specifically, underrepresented minorities, freshman, first-generation, transfer, and international students are of special consideration. The system learns classification rules targeting these imbalanced groups to learn specific rules to improve predictions, in contrast to traditional systems that are often biased towards the general student body. A multi-view genetic programming algorithm (MVGP) was introduced, which extracts classification rules from multiple heterogeneous information sources. The algorithm was compared with other interpretable (tree and rule-based) singleview and multi-view methods. Six different performance metrics, taking into account the imbalanced nature of data classes, were evaluated. Classification performance was analyzed on a weekly basis according to the information available in e-learning systems (Blackboard) and other data sources using student data from the Virginia Commonwealth University College of Engineering. MVGP demonstrably achieves consistently better performance than the other methods, especially considering the imbalanced metrics. Results were validated using nonparametric statistical analysis, which supports the better performance of our method. Learning management systems have been persistently shown to provide meaningful information, allowing early alert systems to provide accurate student performance predictions. However, the collaboration of instructors and administrators is Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. 210 IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, absolutely necessary to provide early intervention and student support. The implementation of personalized user interfaces for students, instructors, and administrators helps to convey the information in a friendly way. For future work, we plan to extend the system to include additional information, such as high-school grades, type of device used when accessing the elearning systems, and IP address (to infer location), which could provide a better understanding of student behavior, such as: are students working together in the library vs studying independently at home? are there differences among students living on campus in residences vs off campus? ACKNOWLEDGEMENT This research complies with all approvals of ethics and data privacy required to conduct this work. REFERENCES [1] A. Elbadrawy, A. Polyzou, Z. Ren, M. Sweeney, G. Karypis, and H. Rangwala, “Predicting student performance using personalized analytics,” Computer, vol. 49, no. 4, pp. 61–69, 2016. [2] F. Marbouti, H. A. Diefes-Dux, and K. Madhavan, “Models for early prediction of at-risk students in a course using standards-based grading,” Comput. Educ., vol. 103, pp. 1–15, 2016. [3] Y.-H. Hu, C.-L. Lo, and S.-P. Shih, “Developing early warning systems to predict students online learning performance,” Comput. Human Behav., vol. 36, pp. 469–478, 2014. [4] W. Chen, C. G. Brinton, D. Cao, A. Mason-singh, C. Lu, and M. Chiang, “Early detection prediction of learning outcomes in online short-courses via learning behaviors,” IEEE Trans. Learn. Technol., vol. 12, no. 1, pp. 44–58, Jan.–Mar. 2019. [5] S. S. Jaggars and D. Xu, “How do online course design features influence student performance?” Comput. Educ., vol. 95, pp. 270–284, 2016. [6] P. M. Moreno-Marcos, C. Alario-Hoyos, P. J. Munoz-Merino, and C. D. Kloos, “Prediction in MOOCs: A review and future research directions,” IEEE Trans. Learn. Technol., to be published, doi: 10.1109/ TLT.2018.2856808. [7] T. Thiele, A. Singleton, D. Pope, and D. Stanistreet, “Predicting students’ academic performance based on school and socio-demographic characteristics,” Stud. High. Educ., vol. 41, no. 8, pp. 1424–1446, 2016. [8] L. P. Macfadyen and S. Dawson, “Mining LMS data to develop an early warning system for educators: A proof of concept,” Comput. Educ., vol. 54, no. 2, pp. 588–599, 2010. [9] R. Conijn, C. Snijders, A. Kleingeld, and U. Matzat, “Predicting student performance from LMS data: A comparison of 17 blended courses using Moodle LMS,” IEEE Trans. Learn. Technol., vol. 10, no. 1, pp. 17–29, Jan.–Mar. 2017. [10] S. Helal, J. Li, L. Liu, E. Ebrahimie, S. Dawson, and D. J. Murray, “Identifying key factors of student academic performance by subgroup discovery,” Int. J. Data Sci. Anal., vol. 7, no. 3, pp. 227–245, 2019. [11] A. Jokhan, B. Sharma, and S. Singh, “Early warning system as a predictor for student performance in higher education blended courses,” Stud. High. Educ., pp. 1–12, 2018. [12] R. C. Kushwaha, A. Singhal, and P. K. Chaurasia, “Study of studentperformance in learning management system,” Int. J. Contemporary Res. Comput. Sci. Technol., vol. 1, no. 6, pp. 213–217, 2015. [13] K. Verbert, et al.“Context-aware recommender systems for learning: A survey and future challenges,” IEEE Trans. Learn. Technol., vol. 5, no. 4, pp. 318–335, Oct.–Dec. 2012. [14] N. Bosch, et al., “Modeling key differences in underrepresented students’ interactions with an online STEM course,” in Proc. APA Conf. Technol., Mind, Soc., 2018, Art. no. 6. [15] R. Villano, S. Harrison, G. Lynch, and G. Chen, “Linking early alert systems and student retention: a survival analysis approach,” Higher Educ., vol. 76, no. 5, pp. 903–920, 2018. [16] C. Marquez-Vera, A. Cano, C. Romero, and S. Ventura, “Predicting student failure at school using genetic programming and different data mining approaches with high dimensional and imbalanced data,” Appl. Intell., vol. 38, no. 3, pp. 315–330, 2013. VOL. 12, NO. 2, APRIL-JUNE 2019 [17] W. Xing, R. Guo, E. Petakovic, and S. Goggins, “Participation-based student final performance prediction model through interpretable genetic programming: Integrating learning analytics, educational data mining and theory,” Comput. Hum. Behav., vol. 47, pp. 168–181, 2015. [18] J. Luna, C. Romero, J. Romero, and S. Ventura, “An evolutionary algorithm for the discovery of rare class association rules in learning management systems,” Appl. Intell., vol. 42, no. 3, pp. 501–513, 2015. [19] C. Marquez-Vera, A. Cano, C. Romero, A. Y. Mohammad, H. M. Fardoun, and S. Ventura, “Early dropout prediction using data mining: A case study with high school students,” Expert Syst., vol. 33, no. 1, pp. 107–124, 2016. [20] C. Romero and S. Ventura, “Educational data mining: A review of the state of the art,” IEEE Trans. Syst., Man, Cybern., C, vol. 40, no. 6, pp. 601–618, Nov. 2010. [21] C. Romero, S. Ventura, M. Pechenizkiy, and R. Baker, Handbook of Educational Data Mining, (ser. Mining and Knowledge Discovery). Boca Raton, FL, USA: CRC Press, 2011. [22] C. Romero and S. Ventura, “Educational data science in massive open online courses,” WIRES Data Min. Knowl., vol. 7, no. 1, 2017, Art. no. p. e1187. [23] S. Slater, S. Joksimovic, V. Kovanovic, R. S. Baker, and D. Gasevic, “Tools for educational data mining: A review,” J. Educ. Behav. Stat., vol. 42, no. 1, pp. 85–106, 2017. [24] K. R. Koedinger, S. D’Mello, E. A. McLaughlin, Z. A. Pardos, and C. P. Rose, “Data mining and education,” WIRES Cognit. Sci., vol. 6, no. 4, pp. 333–353, 2015. [25] R. Baker, “Educational data mining: An advance for intelligent systems in education,” IEEE Intell. Syst., vol. 29, no. 3, pp. 78–82, May–Jun. 2014. [26] R. S. Baker and P. S. Inventado, “Educational data mining and learning analytics,” in Design of Learning Analytics Experiences. New York, NY, USA: Springer, 2014, pp. 61–75. [27] M. Bienkowski, M. Feng, and B. Means, “Enhancing teaching and learning through educational data mining and learning analytics,” US Dept. Educ., Washington, DC, USA, Tech. Rep. no. 20121, pp. 1–57, 2012. [28] D. Azcona, O. Corrigan, P. Scanlon, and A. F. Smeaton, “Innovative learning analytics research at a data-driven HEI,” in Proc. 3rd Int. Conf. Higher Educ. Adv., 2017, pp. 435–443. [29] Y. Hu, C. Lo, and S. Shih, “Developing early warning systems to predict students’ online learning performance,” Comput. Human Behav., vol. 36, pp. 469–478, 2014. [30] E. Howard, M. Meehan, and A. Parnell, “Contrasting prediction methods for early warning systems at undergraduate level,” Internet Higher Education, vol. 37, pp. 66–75, 2018. [31] J. Whitehill, K. Mohan, D. Seaton, Y. Rosen, and D. Tingley, “MOOC dropout prediction: How to measure accuracy?” in Proc. 4th ACM Conf. Learn. Scale, 2017, pp. 161–164. [32] S. M. Jayaprakash and E. J. M. Laurıa, “Open academic early alert system: Technical demonstration,” in Proc. 4th Int. Conf. Learn. Analytics Knowl., 2014, pp. 267–268. [33] M. C. O’Connor and S. V. Paunonen, “Big five personality predictors of post-secondary academic performance,” Personality Individual Differences, vol. 43, no. 5, pp. 971–990, 2007. [34] G. Kennedy, C. Coffrin, P. de Barba, and L. Corrin, “Predicting success: How learners’ prior knowledge, skills and activities predict MOOC performance,” in Proc. 5th Int. Conf. Learn. Analytics Knowl., 2015, pp. 136–140. [35] D. Azcona and A. F. Smeaton, “Targeting at-risk students using engagement and effort predictors in an introductory computer programming course,” in Proc. Eur. Conf. Technol. Enhanced Learn., 2017, pp. 361–366. [36] D. Azcona, I.-H. Hsiao, and A. F. Smeaton, “An exploratory study on student engagement with adaptive notifications in programming courses,” in Proc. Eur. Conf. Technol. Enhanced Learn., 2018, pp. 644–647. [37] J. Dong, W. Hwang, R. Shadiev, and G. Chen, “Implementing on-calltutor system for facilitating peer-help activities,” IEEE Trans. Learn. Technol., vol. 12, no. 1, pp. 73–86, Jan.–Mar. 2019. [38] J. Vandamme, N. Meskens, and J. Superby, “Predicting academic performance by data mining methods,” Educ. Econ., vol. 15, no. 4, pp. 405–419, 2007. [39] W. Xing and D. Du, “Dropout prediction in MOOCs: Using deep learning for personalized intervention,” J. Educ. Comput. Res., pp. 1–24, 2018. [40] B. K. Pursel, L. Zhang, K. W. Jablokow, G. Choi, and D. Velegol, “Understanding MOOC students: Motivations and behaviours indicative of MOOC completion,” J. Comput. Assist. Learn., vol. 32, no. 3, pp. 202–217, 2016. [41] D. Gunning, “Explainable artificial intelligence (XAI),” DARPA, pp. 1–36, 2017. Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply. CANO AND LEONARD: INTERPRETABLE MULTIVIEW EARLY WARNING SYSTEM ADAPTED TO UNDERREPRESENTED STUDENT POPULATIONS [42] R. Ferguson, “Learning analytics: Drivers, developments and challenges,” Int. J. Technol. Enhanced Learn., vol. 4, no. 5/6, pp. 304–317, 2012. [43] Y.-S. Tsai and D. Gasevic, “Learning analytics in higher education: Challenges and policies: A review of eight learning analytics policies,” in Proc. 7th Int. Conf. Learn. Analytics Knowl., 2017, pp. 233–242. [44] D. Azcona and K. Casey, “Micro-analytics for student performance prediction,” Int. J. Comput. Sci. Softw. Eng., vol. 4, no. 8, pp. 218–223, 2015. [45] A. Cano, A. Zafra, and S. Ventura, “An interpretable classification rule mining algorithm,” Inf. Sci., vol. 240, pp. 1–20, 2013. [46] J. W. You, “Identifying significant indicators using LMS data to predict course achievement in online learning,” Internet Higher Educ., vol. 29, pp. 23–30, 2016. [47] O. Corrigan, A. F. Smeaton, M. Glynn, and S. Smyth, “Using educational analytics to improve test performance,” in Proc. 10th Eur. Conf. Technol. Enhanced Learn., 2015, pp. 42–55. [48] Y. Vance Paredes, D. Azcona, I.-H. Hsiao, and A. F. Smeaton, “Predictive modelling of student reviewing behaviors in an introductory programming course,” in Proc. Int. Conf. Educ. Data Mining Comput. Sci. Educ., 2018, pp. 1–5. [49] C. Romero, A. Zafra, J. M. Luna, and S. Ventura, “Association rule mining using genetic programming to provide feedback to instructors from multiple-choice quiz data,” Expert Syst., vol. 30, no. 2, pp. 162–172, 2013. [50] C. Romero, S. Ventura, and P. D. Bra, “Knowledge discovery with genetic programming for providing feedback to courseware authors,” User Model. User-Adap. Interact., vol. 14, no. 5, pp. 425–464, 2004. [51] A. Zafra, C. Romero, and S. Ventura, “Predicting academic achievement using multiple instance genetic programming,” in Proc. 9th Int. Conf. Intell. Syst. Des. Appl., 2009, pp. 1120–1125. [52] C. Romero, S. Ventura, C. Hervas, and P. Gonzalez, “Rule discovery in web-based educational systems using grammar-based genetic programming,” WIT Trans. Inf. Commun. Technol., vol. 35, pp. 205–214, 2005. [53] A. Bogarın, R. Cerezo, and C. Romero, “A survey on educational process mining,” WIRES Data Min. Knowl., vol. 8, no. 1, 2018, Art. no. e1230. [54] H. Research, “Early alert systems in higher education,” Acad. Admin., pp. 1–22, 2014. [55] J. E. Knowles, “Of needles and haystacks: Building an accurate statewide dropout early warning system in Wisconsin,” J. Educ. Data Min., vol. 7, no. 3, pp. 18–67, 2015. [56] O. S. Patil and P. M. Dhere, “Predicting dropout students using data mining techniques,” Int. J. Res., vol. 2, no. 1, pp. 369–375, 2015. [57] A. Pradeep, S. Das, and J. J. Kizhekkethottam, “Students dropout factor prediction using EDM techniques,” in Proc. Int. Conf. Soft-Comput. Netw. Secur., 2015, pp. 1–7. [58] A. Wolff, Z. Zdrahal, A. Nikolov, and M. Pantucek, “Improving retention: Predicting at-risk students by analysing clicking behaviour in a virtual learning environment,” in Proc. 3rd Int. Conf. Learn. Analytics Knowl., 2013, pp. 145–149. [59] D. Gasevic, S. Dawson, T. Rogers, and D. Gasevic, “Learning analytics should not promote one size fits all: The effects of instructional conditions in predicting academic success,” Internet Higher Educ., vol. 28, pp. 68–84, 2016. [60] J. Zhao, X. Xie, X. Xu, and S. Sun, “Multi-view learning overview: Recent progress and new challenges,” Inf. Fusion, vol. 38, pp. 43–54, 2017. [61] S. Sun, “A survey of multi-view machine learning,” Neural Comput. Appl., vol. 23, no. 7/8, pp. 2031–2038, 2013. [62] H. Xue, S. Chen, J. Liu, and J. Huang, “Multi-view classification method based on cross-view constraints,” Pattern Recognit. Artif. Intell., vol. 27, no. 2, pp. 97–102, 2014. [63] F. Wu, Y. Huang, and Z. Yuan, “Domain-specific sentiment classification via fusing sentiment knowledge from multiple sources,” Inf. Fusion, vol. 35, pp. 26–37, 2017. [64] A. Cano, “An ensemble approach to multi-view multi-instance learning,” Knowl.-Based Syst., vol. 136, pp. 46–57, 2017. [65] C. Garcia-Martinez and S. Ventura, “Multi-view semi-supervised learning using genetic programming interpretable classification rules,” in Proc. IEEE Conf. Evol. Comput., 2017, pp. 573–579. [66] M. M. Gaber, A. Zaslavsky, and S. Krishnaswamy, “Mining data streams: a review,” Sigmod Rec., vol. 34, no. 2, pp. 18–26, 2005. [67] K. Nag and N. Pal, “A multiobjective genetic programming-based ensemble for simultaneous feature selection and classification,” IEEE Trans. Cybern., vol. 46, no. 2, pp. 499–510, Feb. 2016. 211 [68] A. Shaker and E. Lughofer, “Resolving global and local drifts in data stream regression using evolving rule-based models,” in Proc. IEEE Conf. Evolving Adaptive Intell. Syst., 2013, pp. 9–16. [69] E. Lughofer and P. P. Angelov, “Handling drifts and shifts in on-line data streams with evolving fuzzy systems,” Appl. Soft Comput., vol. 11, no. 2, pp. 2057–2068, 2011. [70] M. O’Neill, L. Vanneschi, S. Gustafson, and W. Banzhaf, “Open issues in genetic programming,” Genet. Program. Evol. M., vol. 11, no. 3, pp. 339–363, 2010. [71] S. Khanchi, M. Heywood, and N. Zincir-Heywood, “On the impact of class imbalance in GP streaming classification with label budgets,” in Proc. 19th Eur. Conf. Genetic Program., 2016, pp. 35–50. [72] U. Bhowan, M. Zhang, and M. Johnston, “Genetic programming for classification with unbalanced data,” in Proc. 13th Eur. Conf. Genetic Program., 2010, pp. 1–13. [73] G. Lassibille, “Student progress in higher education: What we have learned from large-scale studies,” Open Educ. J., vol. 4, no. 4, pp. 1–8, 2011. [74] I. H. Witten, E. Frank, M. A. Hall, and C. J. Pal, Data Mining, Fourth Edition: Practical Machine Learning Tools and Techniques, 4th ed. San Mateo, CA, USA: Morgan Kaufmann, 2016. [75] W. W. Cohen, “Fast effective rule induction,” in Proc. 12th Int. Conf. Mach. Learn., 1995, pp. 115–123. [76] R. Quinlan, C4.5: Programs for Machine Learning. San Mateo, CA, USA: Morgan Kaufmann, 1993. [77] A. Cano, J. M. Luna, A. Zafra, and S. Ventura, “A classification module for genetic programming algorithms in JCLEC,” J. Mach. Learn. Res., vol. 16, pp. 491–494, 2015. [78] M. Fatourechi, R. K. Ward, S. G. Mason, J. Huggins, A. Schl€ogl, and G. E. Birch, “Comparison of evaluation metrics in classification applications with imbalanced datasets,” in Proc. 7th Int. Conf. Mach. Learn. Appl., 2008, pp. 777–782. [79] V. Lopez, A. Fernandez, S. Garcıa, V. Palade, and F. Herrera, “An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics,” Inf. Sci., vol. 250, pp. 113–141, 2013. [80] S. Garcıa, A. Fernandez, J. Luengo, and F. Herrera, “Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining,” Inf. Sci., vol. 180, no. 10, pp. 2044–2064, 2010. Alberto Cano received the B.Sc. degrees in computer engineering and in computer science from the University of Cordoba, Spain, in 2008 and 2010, respectively, and the M.Sc.and Ph.D. degrees in intelligent systems and computer science from the University of Granada, Granada, Spain, in 2011 and 2014, respectively. He is an Assistant Professor with the Department of Computer Science, Virginia Commonwealth University, Richmond, VA, USA, where he heads the High-Performance Data Mining Lab. His research is focused on machine learning, data mining, parallel, distributed, and GPU computing. He has published more than 35 articles in high-impact factor journals, 39 contributions to international conferences, two book chapters, and one book in the areas of machine learning, data mining, and parallel, distributed, and GPU computing. His research is supported by an Amazon AWS Machine Learning award (2018) and the VCU Presidential Research Quest Fund (2018). He is Associate Editor of IEEE ACCESS. John D. Leonard received the M.Sc. and Ph.D. degrees from the University of California, Irvine, CA, USA, in 1986 and 1991, respectively. He is a Full Professor with the Department of Computer Science, Virginia Commonwealth University, Richmond, VA, USA, where he is Executive Associate Dean of the VCU College of Engineering. His research is focused on data-oriented approaches to resource allocation, decision-making, and operations. Authorized licensed use limited to: Rowan University Libraries. Downloaded on June 21,2025 at 20:48:42 UTC from IEEE Xplore. Restrictions apply.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )