Optimal cricket team selection using machine learning Abstract: Building a team in any sport requires a careful assessment process, and the team's overall performance in a particular match depends on which individuals are chosen to be on the team. A range of criteria and algorithms are employed to selectively choose players. The strengths and weaknesses of the opposition, pitch conditions, player analytics, and climatic data are all carefully considered while selecting a cricket squad. A key component of player performance prediction is machine learning algorithms, which use past data on team dynamics and individual player statistics to predict the success of the team as a whole. The process of forming a team is streamlined by this predictive analysis of individual player performance, which promotes the best possible team chemistry. Even though there have been recent academic attempts to build prediction models for cricket player performance, previous studies frequently ignore important variables like weather dynamics and field characteristics, which have a significant influence on player production. By completing a thorough evaluation of the literature, our study closes this gap and develops a solid framework for predicting player performance in cricket. By using our technique, team managers are able to make well-informed judgments that lead to the selection of a lineup that is most suited to improve the team's performance as a whole. Keywords— Cricket, Performance Prediction, Machine Learning, Decision Tree, Random Forest I. INTRODUCTION The ultimate goal of a cricket match is to win. Right now, cricket draws more viewers than soccer [1]. Sports need a professional approach as they grow more and more competitive [2]. Numerous teams started using scientific approaches to determine their game plan and plan of attack. A machine learning technology aids in the analysis of opponent team members' performances by the team management. A machine learning method may be used to identify the strengths and weaknesses of both teammates and the adversary [5]. A lot of researchers have been working on developing algorithms that can forecast players' performances in recent times [1] [2]. Predicting player performance in cricket assists team management in making the final team selection prior to the toss. Victory in cricket is determined by several variables, including previous results, playing on home or away ground, match experience, venue-specific performance, performance against a particular opposition, and the team's and player's form. As technology advances more quickly and the public's desire for cricket grows, so does the need for machine learning calculations to predict cricket match results. Cricket is one of the most well-known and popular sports in the world today, particularly in the Indian subcontinent. Strong incentives to model and train the game from several viewpoints have been provided by a large betting market, a multiplicity of natural elements impacting the game, and extensive media coverage [5]. Every element of life is made simpler by the use of artificial intelligence, machine learning, deep learning, and data science. The coaches and players will be able to assess the areas for growth through the use of machine learning and result prediction prior to the real game [5]. Machine learning is a rapidly expanding field that is closely related to computational insights and often overlaps with it. It is centered on using technology to make predictions. It is closely related to numerical improvement, which is hypothesized to provide the field with methods and application areas. Sometimes, data mining—a subset known as supervised learning—is confused with machine learning. Data mining focuses more on exploratory information processing. Assessing players' performances is a difficult undertaking. The team managers and team selectors may find it useful to have an intelligent system that forecasts player performance based on historical data [1] [2]. Numerous studies have been conducted in an attempt to forecast player performance [1] [2] [5]. Nevertheless, none of them take advantage of weather-related elements, which might have an impact on players' performance. As a result, we are working on a literature review to provide the best approach for predicting each player's performance throughout a cricket match. II. BASIC INTRODUCTION OF CRICKET GAME Three forms are available for playing cricket: 1. ODIs (50 overs matches), 2. T20 (20 overs matches), and 3. Tests (five day games). Confined overs cricket refers to both Twenty20 and One-Day Internationals, whereas Test matches are confined to play for 5 days. Every format has two innings, with one bowling and one batting opportunity for each team. One inning in an ODI consists of a maximum of 50 bowling overs. In T20 cricket, a bowling innings consists of a maximum of 20 overs. When the bowling team bowls 50 overs or more in an ODI or T20 match, or before the batting team loses 10 wickets, the inning is over. Teams switch roles after the first inning; for instance, the batting team bowls, and the bowling team bowls. In ODI and T20 cricket, the bowling team chases down a target set by the opposing side within 50 overs or more after the batting team bowls. Similarly, the bowling team attempts to stop the other side from chasing the target in less than 50 overs, 20 overs, or 10 wickets via bowling. Each side can bat and bowl in the two innings in ODIs and T20s. In ODIs, an inning may consist of up to 50 overs, but in T20s, an inning may consist of 20 overs or fewer. Test matches last up to five days each, and single game, each side may play up to two innings. Both teams in limited cricket have eleven players apiece, and each will bat and bowl once in an inning. Each team switches roles after the first inning. The bowling team attempts to reach the opposite side's target in 50 overs or less than 10 goals conceded. In addition to bowling, the attacking team attempts to prevent the other side from pursuing the target or scoring ten goals in 50 overs or fewer. Once the batting team has lost 10 wickets or 50 overs to the opposing team, the batting innings comes to an end. In the same manner, the opposing team will play after the fielding team attacks. To win the game, the batting team must pursue the goal established by the first betting team. To win, the opposition must limit the batting team's ability to chase the objective. For both bowlers and batsmen, limited-overs cricket presents greater challenges. Batsmen need to get runs as fast as they can, while bowlers need to limit them by conceding the fewest runs and picking up the fewest wickets. The most widely used format in modern international cricket is the One-Day International (ODI) format. A. CRICKET FOR STATISTICS The two primary types of statistics are bowling and batting. The match's ultimate result is mostly determined by these two statistics. STATS ON BATTING [4] The following are all of the batting statistics: 1. Innings: The total number of innings a player has played. 2. Not Outs: The quantity of times the batter is still in the game. 3. Runs: The total amount of runs a batsman has scored. 4. Highest Score: The greatest score that the batsman has ever achieved in the past. 5. Batting Average: The batsman’s batting average. 6. Centuries: How many times a batsman reaches 100 runs in a game. 7. Half-Centuries: The frequency with which a batsman reaches 50 runs. 8. Balls Faced: The total quantity of balls that the batter has faced. 9. Strike Rate: Strike Rate = (100 * Runs) / Ball Faced 10. 10. Run Rate: The typical amount of runs a team scores per over. STATS ON BOWLING [4] Statistics related to bowlers are: 1. Over: A bowler's six balls bowled constitutes an over. 2. Maiden Overs: The overs in which the bowler gave up no runs in an over. 3. Wickets: The quantity of wickets a bowler has claimed. 4. No-Balls: How many no-balls a bowler has bowled. 5. Wide: The quantity of wide balls that the bowler bowls. 6. Bowling Average: The mean quantity of runs lost per wicket. 7. Strike Rate: The typical quantity of balls bowled. III Correlated Works Only a few documents on player performance prediction in cricket were found using search [5]. Based on player data and attributes, Passi and Pandey [5] reconstructed hitting and bowling records. Other elements that could not be taken into consideration yet have an impact on player performance include the kind of wickets and the weather. To forecast player performance, Siripurapu [5] created an adaptive Neuro-fuzzy model, although it only takes into account team members based on certain bowling and batting metrics. Therefore, there's a chance that using every participant at once will raise the rating. In a one-day cricket match, Barr et al. [7] suggested criteria for comparing and choosing batsmen. When choosing a hitter, authors simply take into account the strike rate and batting average. Nonetheless, weather and wicket types have a big impact on how well players perform . Bhattacharjee et al. [8] conducted a study in 2012 to forecast the performance of bowlers. The pace at which a bowler bowls is one crucial factor that writers overlook when analyzing a player's performance. This research aims to forecast bowler performance exclusively. Bowlers and batters are divided into three groups by Iyer and Sharda [9]: performance, failure, and middling. The authors' method for predicting a player's performance is based on neural networks. This method was used to choose the 2007 World Cup team based on each player's performance rating. Rising stars in cricket were forecasted by Hasseb Ahmed et al. [10] in 2017. Based on performance development and weighted average, they compiled a list of the top 10 emerging cricket players using a variety of machine learning and mathematical techniques. Lastly, they made a comparison between the International Cricket Council ranks and rising star scores. . Predict player performance in an ODI cricket match using the SVM algorithm. For their study, they only use six players, though. SVM requires greater processing time when there is non-linear data separation. Additionally, for best performance, the right kernel function must be chosen. A. MODEL OF GENERIC PREDICTION The machine learning technique for predicting players' performance was proposed by several academics. Prediction is done using machine learning techniques such as random forest, SVM, naïve Bayesian, and decision tree. With the exception of random forest, the majority of algorithms operate below optimally. Figure 1 displays the generalized prediction model derived from previous studies together with the output of the corresponding machine-learning method. FIG 1: PREDICTION MODEL FOR OPTIMAL CRICKET B. PROBLEM IDENTIFICATION Predicting a player's performance is becoming increasingly desirable in a competitive setting. In this field, machine learning has a significant potential impact. The majority of previous studies did not include the weather in their analysis. Weather conditions may have an impact on players' and teams' performances [5]. Also, a limited collection of datasets is used in the majority of current studies [3]. They focus just on a small aspect of the cricket game and restrict their study efforts. C. PROPOSED SYSTEM FLOW The majority of previous research solely includes statistics pertaining to cricket matches and leaves out weather-related information [1, 3, 7, 10]. The weather has a significant impact on players' performances. The wind that blows in the bowler's direction enhances their performance and hinders that of the batsman. The following factors might affect the outcome of a day-night match: humidity, wind direction, rain, and cold. We suggested a method that predicts a player's performance by utilizing meteorological data and statistics from cricket matches. . We discovered that a machine learning strategy using meteorological attributes has the potential to enhance prediction performance after reviewing a survey of the literature on players' performance prediction. We suggested a technique that combines meteorological and cricket datasets. After that, feature extraction will be used to extract the most valuable traits. In order to ensure that every class has an equal impact on the final forecast, we additionally balance the uneven cricket dataset. The suggested system flow is provided in Figure 2. FIG 3: Hyper Plane of SVM The SVM hyperplane is depicted in figure 2. The type of issue statement and the application we are designing will determine how the hyperplane's value is set up. FIG 2: FLOW MODEL OPTIMAL TEAM B. Random Forest Classifier. IV. Classification Steps . The challenge of classification is to determine, given a training set of data that includes observations whose category is known, which of a set of categories a new observation belongs to. Random Forest is a classifier that contains a number of decision trees on various of the given dataset and takes the average to improve the predictive accuracy of that dataset. The greater number of trees in the forest leads to higher accuracy an prevents the problem of overfitting. A. C. Decision Tree Support Vector Machines (SVM) Classifier The foundation of SVM (Support Vector Machine) is supervised learning. SVM needs training sets and labels to go along with them. If test data is added after training, the model classifies it into one of the two categories. It functions effectively in linear classification. By translating the inputs into a high dimensional feature space and employing a kernel method, it can even do well on a nonlinear classification. It builds the classification's hyper plane. The hyper plane is selected to maximize the separation between the closest data points on either side. Formally, SVM is a machine learning technique that uses a hyperplane to categorize datasets according to a given criteria. The data is divided into 2-D space using a hyperplane based on input variables. Decision tree algorithms segment data by feature values, optimizing criteria like Gini impurity or entropy. They provide transparent insights into fraud indicators but may overfit complex datasets. Techniques like pruning mitigate this. Despite limitations, decision trees are foundational in fraud detection, aiding in feature significance understanding and serving as building blocks for advanced models. Linear Classification: In its simplest form, SVM performs linear classification by finding the hyperplane that best separates th data into two classes. This hyperplane can be expressed as: W.x+b=0 (1) FIG 4: DECISION TREE OF RANDOM FOREST V. RESULT AND DISCUSSION Where w is the weight vector perpendicular to the hyperplane, x is the input feature vector, and b is the bias term. This model emphasizes skill alignment, clear goals, effective communication, feedback, adaptive leadership, and celebrating achievements to optimize team performance and collaboration within project environments. Accuracy table :- Algorithm Accuracy Decision Random Tree Forest 86.50 92.25 Weighted SVM Random Forest 68.78 93.73 TABLE 1: COMPARSION OF ACCURACY. Table 1 illustrates that our weighted random forest technique yields the highest accuracy of 93.73 when compared to other conventional classification algorithms. FIG 6: ACCURACY GRAPH VI. CONCLUSION In order to create an intelligent system that can anticipate players' performances for one-day international cricket matches, it is necessary to take into account factors other than bowling and batting. The weather aspect is a crucial consideration that greatly affects how well players perform. We suggested a machine learning approach that makes use of weather-related datasets and cricket statistics. We suggest data balancing as a preprocessing step before using a machine learning method, as combined data is inherently unbalanced. We evaluated our suggested model on a combined balanced dataset of meteorological and cricket match data, and it employs a novel weighted random forest classifier with hyperparameter tweaking. When compared to other algorithms, we discovered that the accuracy of our suggested model was good. FIG 5: COMPARISON ANALYSIS FOR ALGORITHMS. REFERENCES The performance of many machine learning models for choosing cricket teams is displayed in the pie chart: 1. M. G. Jhanwar and V. Pudi, "A team composition based approach for predicting the outcome of One-day International cricket matches," CEUR Workshop Proc., vol. 1842, no. September, 2016. - 86.50% in Decision Tree - 92.25% in Random Forest 2. "Improved Prediction Accuracy in the Game of Cricket Using Machine Learning," by K. Passi and N. Pandey, International Journal of Data Mining, Knowledge Management, and Process, vol. 8, no. 2, pp. 19–36, 2018. - Support Vector Machine (SVM): 68.78 percent - 93.73% in Weighted Random Forest The accuracy % attained by a particular model is represented by each slice of the pie. As shown, Random Forest and Weighted Random Forest were the two models that performed the best in terms of accuracy. Among the models examined, Decision Tree fared very well, although SVM had the lowest accuracy, shown in figure 4. 3. Int. J. Trend Res. Dev., vol. 5, no. 1, pp. 2394–9333; R. Lokhande and A. Professor, "Live Cricket Score and Winning Prediction." 4. M. G. Jhanwar and V. Pudi, "A team composition based approach for predicting the outcome of One-day International cricket matches," CEUR Workshop Proc., vol. 1842, no. September, 2016. 5. "Intelligent system for team selection and decision making in the game of cricket," N. Siripurapu, A. Mittal, R. P. Mukku, and R. Tiwari, Smart Innov. Syst. Technol., vol. 77, pp. 467–474, 2018.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )