Mechanical Systems and Signal Processing 143 (2020) 106832 Contents lists available at ScienceDirect Mechanical Systems and Signal Processing journal homepage: www.elsevier.com/locate/ymssp A novel approach for predicting tool remaining useful life using limited data Hai Li, Wei Wang ⇑, Ziwei Li, Liyi Dong, Qingzhao Li School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China, Chengdu 611731, China a r t i c l e i n f o Article history: Received 15 October 2019 Received in revised form 26 January 2020 Accepted 20 March 2020 Available online 30 March 2020 Keywords: Tool Remaining useful life prediction Limited data Adaptive time window Deep bidirectional long-short term memory a b s t r a c t Wear, fracture, and other tool faults affect the quality of a machined workpiece and can even damage machine tools. The accurate prediction of remaining useful life (RUL) can prevent a tool from suddenly failing, an ability of significance for ensuring machining quality and providing effective predictive maintenance strategies. Most current approaches for predicting tool RUL are based on historical failure and truncation data. However, for new types of tools or when a similar tool has just launched, such failure and truncation data are limited or even unavailable, making RUL prediction a challenge when using previously proposed methods. To address this problem, a novel method for the prediction of tool RUL using limited data is proposed in this study. A time window is constructed to track the tool condition using sensor data, and its size can be dynamically adjusted according to the wear factor and increase rate. Then, a deep bidirectional long short-term memory (DBiLSTM) neural network in which sequential data are predicted and smoothed by forwards and backwards directions, respectively, is developed to encode temporal information and identify long-term dependencies. On this basis, multi-step ahead rolling predictions are then employed to predict tool RUL. Finally, the effectiveness of the proposed method is verified using the results of milling experiments. These results show that the proposed method is able to predict tool RUL with high accuracy using only limited data. Ó 2020 Elsevier Ltd. All rights reserved. 1. Introduction A tool is a critical component in manufacturing processes such as grinding, milling, turning, and so on. Tool failures such as wear and breakage impact the machining performance of a tool and can be harmful to high-value machinery. Thus, condition-based maintenance (CBM) has been proposed as a feasible strategy for avoiding unexpected tool failure [1,2]. The prediction of the remaining useful life (RUL) of a tool, defined as the length of time from the current moment to the end of its useful life, is a key step in implementing CBM [3–5]. The primary objective of RUL is to predict the remaining useful time before a tool loses its capacity based on condition monitoring information. Due to its importance, many scholars have begun to pay careful attention to RUL in recent decades, and several published papers have reviewed different approaches to RUL prediction [6]. Generally, tool RUL prediction methods are classified into physics-based, reliability function, data-driven, and hybrid models. Physics-based models describe the tool wear process by establishing mathematical models based on failure mechanisms. The tool life model and tool wear rate model have been widely studied [7–10]. The Wiener process and particle filtering have ⇑ Corresponding author. E-mail addresses: wangwhit@163.com (W. Wang), qingzhaoli@yeah.net (Q. Li). https://doi.org/10.1016/j.ymssp.2020.106832 0888-3270/Ó 2020 Elsevier Ltd. All rights reserved. 2 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 also been employed to perform RUL prediction [11,12]. A more comprehensive summary of physics-based approaches for the prediction of tool RUL can be found in the literature [13]. Although a physics-based approach can provide accurate RUL prediction, in practical application, it can be difficult or even impossible to acquire the necessary extensive offline measurements like tool wear width, restricting the scope of application of these approaches. Reliability function models, also called statistical model-based approaches, predict tool RUL by establishing a reliability function on the basis of empirical knowledge [3]. These approaches assume that the tool degradation process is subject to a certain distribution, and the parameters of this distribution are estimated by employing the entire degradation data set [14]. The Gaussian process regression [15], hidden Markov [16], Bayesian [17], and adaptive hidden Markov models have all been proposed to predict tool RUL [18]. In these methods, the most critical steps are the acquisition of failure or truncation data and the assumption of a proper distribution. If failure data is unavailable, it is difficult to predict RUL using reliability functions, and if the assumed distribution is improper, the corresponding prediction error will be considerable. Data-driven models are effective for condition monitoring and RUL prediction [19]. These methods employ vibration sensors, acoustic emission sensors, current sensors, or other kinds of sensors to monitor tool working conditions. Signals are extracted from the sensor data by signal processing technology to characterize the actual physical state of the tool [20– 22]. Because data-driven methods have advantages including easy implementation and non-interruption of the manufacturing process, these methods have been a research hotspot for tool condition monitoring and RUL prediction. The logistic regression [23], artificial neural network (ANN) [24–27], sensor fusion [28], machine learning [29], support vector regression (SVM) [30], and least square SVM approaches have been all utilized to predict tool life and tool wear [31]. Generally, these approaches require failure data to establish their models. If such failure data are unavailable, it is difficult for these models to fully describe the tool wear process, resulting in large RUL prediction errors. There is no universally superior model because each model type has advantages and disadvantages [11,32]. The combination of multiple models to create hybrid models attempts to capitalize on the advantages of different approaches by integrating them [33]. Some research has combined the data-driven and physics-based methods to predict tool wear [11,34]. Additionally, both the hidden Markov and polynomial regression models have been employed to predict the RUL and health state of tools [16]. However, these methods also require failure data to establish an accurate prediction model. There has been little research on tool RUL prediction using limited data [11,32]. A review of the current literature suggests that most tool RUL prediction models are based on failure or truncation data [3,35]. However, actually, it is difficult to acquire failure and truncation data for new types of tools or similar tools that have just been put into use, and tool failure is not acceptable when machining some high-value parts to avoid damaging them, resulting in limited data that contains little event data or no failure and truncation data [32]. Similarly, the need for failure and truncation data means that the reliability function, data-driven, and hybrid models waste resources required to construct training samples [36]. For these reasons, it would be beneficial to develop a method for predicting tool RUL with only limited data. In addition to the data requirement limitations, other inherent shortcomings of existing methods present major difficulties in accurately predicting tool RUL. Some widely employed methods, such as the SVM and ANN, can accurately predict in the short term, but because of the value range of the activation function or kernel function, these algorithms cannot capture the inter-relationships of sequential data acquired from multiple sensors and therefore they obtain poor prediction results in the long term. Therefore, the prediction of tool RUL under limited data is even more challenging [32]. To overcome the difficulties of limited data and improve the accuracy of RUL predictions, a new approach for tool RUL prediction using limited data is presented and evaluated in this study. This novel approach first uses an adaptive time window to track the tool wear process. A deep bidirectional long short-term memory (DBiLSTM) neural network is then proposed to capture the relationships between past and future contexts. Finally, the effectiveness of the proposed method is verified using the results of milling experiments. This paper is organized as follows. Section 2 presents the framework of proposed method, including adaptive time window, time window construction, deep bidirectional LSTM and multi-step ahead rolling prediction. Milling experiments are carried out in Section 3. Finally, section 4 gives a conclusion. 2. Proposed approach for tool RUL prediction An illustration of RUL prediction based on limited data is shown in Fig. 1. There are two limited data scenarios: (1) the data in both Region I and II are known; (2) the data in Regions II is known, and those in Region I is unknown. In the figure, Ft is the value of a sensitive feature at a sampling time t, which can strongly represent tool condition. When Ft exceeds the failure threshold value FFT, the RUL is obtained. In this paper, the proposed method utilizes this limited data to predict tool RUL based on an adaptive time window, DBiLSTM, and multi-step ahead rolling prediction. 2.1. Adaptive time window 2.1.1. Tool wear factor determination In the process of material cutting, the cutting area of a tool is under high temperature, high pressure, high speed, and mechanical or thermal shock, so the tool wear is complex and increases with increasing temperature [9,10]. According to Takeyama and Murata’s model and Usui’s tool wear rate model, Ft increases nonlinearly as defined by H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 3 Fig. 1. Illustration of limited data for RUL prediction. F t ¼ b0 expðHðt ÞÞ ð1Þ where b0 is a constant and equal to F1. H(t) represents a value determined by tool working condition at sampling time t, where H(1) is equal to zero. pðtÞ ¼ Ft ¼ expð g ðt ÞÞ; where g ðt Þ ¼ HðtÞ Hðt 1Þ F t1 ð2Þ where pt denotes the increase factor at sampling time t and should not be less than 1, where p1 is equal to 1. g(t) denotes the wear factor at sampling time t. Note that in Eqs. (1) and (2), g(t) has a considerable influence on p(t) and then Ft. Therefore, it is critical to accurately determine g(t). In particular, g(t) is typically very small and can even be close to zero when the hardness of tool is much higher than that of workpiece. According to the Taylor expansion [37], the relationship between g(t) and p(t) can be written as lim pðt Þ ¼ expðgðtÞÞ / ð1 þ g ðt ÞÞ ð3Þ g ðt Þ!0 The feature threshold ratio (FTR) at at a sampling time t, which measures the proximity of the current value to the failure threshold, is shown in Fig. 2 and defined as: at ¼ Qt Q Qt1 t Y Ft pm pt t1 pm ð1 þ g ðt ÞÞ m¼1 pm pt ¼ Qm¼1 ¼ Qm¼1 ¼ P ¼ b pt ; where F Q t 0 n n1 F FT pn ð1 þ g max Þ pn n1 l¼1 pl m¼1 m¼1 pu l¼1 pl ð4Þ where gmax is the given maximum wear factor. pn,pm and pu denote the increase factor at sampling time n, m and u, respectQ -1 nQ -1 tively. pm = pl is in the interval (0,1] and at is in the interval [0,1]. m¼1 l¼1 Then, g(t) can be further calculated by g ðt Þ 6 at þ at g max 1 6 at g max ð5Þ Fig. 2. Illustration of feature threshold ratio. 4 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 2.1.2. Time window construction 2.1.2.1. Time window without outliers. To monitor the tool condition, the time window concept, which refers to the change of the value of a selected feature over time, is introduced. The length of time window nl is equal to the total numbers of features in a feature matrix M ¼ ½F 1 ; F 2 ; ; F nl , where M represents a time window. To simplify the calculation, it is assumed that the sampling frequency is kept constant, so the same number of features is present in any two adjacent time windows. Under this approach, SP is defined as the prediction start time. A requirement that the total length of two adjacent time window sets is smaller than the length of limited data should be satisfied to perform RUL prediction. To quantify the increase in characteristic features between two adjacent time windows, the increase rate r is proposed to represent the ratio of the geometric means of a feature in two adjacent time windows. For the initial phase, the relationships among nl0 , M01, M02, and SP are illustrated in Fig. 3. Feature matrixes M01 and M02, r0, and a0 are defined as, respectively. h i M01 ¼ F SP2nl þ1 ; F SP2nl þ2 ; ; F SPnl 0 h 0 i0 M02 ¼ F SPnl þ1 ; F SPnl þ2 ; ; F SP 0 vffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi u nl uY 0 t F SPnl þi =F SP2nl þi r0 ¼ nl 0 0 i¼1 ð6Þ 0 ð7Þ 0 Pnl0 i¼1 F SPnl0 þi a0 ðSPÞ ¼ ð8Þ nl0 where M01 and M02 are the two adjacent feature matrixes of initial time windows. r0 and a0(SP) represent the initial increase rate and FTR, respectively, and nl0 is the length of the initial time window. Note that nl0 is critical in the determination of M01, M02, r0, and a0. If nl0 is too big, the calculation of the feature matrix will be complex, and the predicted tool condition will be over-dependent on the global trend, which cannot reflect the local trend. On the contrary, if nl0 is too small, the feature matrix will contain few features, and although this simplifies the calculation, it will result in too few training samples and thus poor prediction accuracy. The adaptive time window is proposed to solve this problem: nl is dynamically determined by the wear factor and increase rate. If r0 is smaller than (1 + g0(SP)), the time window will be compressed step by step until rk is no less than (1 + g0(SP)) or reaches the minimum compression limit nlkl , where rk is the increase rate after the kth compression. Otherwise, the initial time window does not need to be compressed. The compression of feature matrixes Mk1 and Mk2, rk, and ak(SP) are shown in Fig. 4 and expressed as follows, respectively. Mk1 ¼ ½F SP2nl þ1 ; F SP2nl þ2 ; ; F SPnl k k k Mk2 ¼ ½F SPnl þ1 ; F SPnl þ2 ; ; F SP k vffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi u nl uY k t rk ¼ F SPnl þi =F SP2nl þi nl k k i¼1 ð9Þ k ð10Þ k Pnlk ak ðSPÞ ¼ i¼1 F SPnlk þi ð11Þ nlk Fig. 3. Illustration of nl0 , M01, M02 and SP. H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 5 Fig. 4. Time window compression. where Mk1 and Mk2 denote two adjacent feature matrixes after the kth compression, and nlk and ak(SP) are the time window length and FTR, respectively, after the kth compression. 2.1.2.2. Time window including outliers. The increase rate r has a considerable impact on the determined time window. In Section 2.1.2.1, if r0 is less than (1 + g0(SP)), the time window should be compressed. Otherwise, the initial time window does not need to be compressed. However, when outliers, which refers to some points where their values rapidly increase like the outlier region Ms1 in Fig. 5, appear in the initial time window, r0 may be larger than (1 + gmax). In this case, if the time window including outliers is not properly compressed, the RUL prediction will be inaccurate. To solve the problem of outliers, an approach for constructing a time window containing outliers is presented. Here, the outlier region should satisfy both of the following conditions: Condition (1): Pnlk RðkÞ ¼ 2 2 P 1 F SPnl þi F W k2 i¼1 F SP2nl þi F W k1 Pnlk i¼1 k ð12Þ k where F Mk1 and F Mk2 are the mean of features in Mk1 and Mk2, respectively. Condition (2): MinðMk2 Þ > MaxðMk1 Þ ð13Þ If r0 is larger than (1 + gmax), the initial time window should be compressed step by step until rk is first in the interval [1 + g0(SP),1 + gmax] or the length of the compressed time window reaches the given minimum nlkl . To further identify whether there is an outlier region, the time window should continue to be compressed until both Condition (1) and Condition (2) are satisfied, or the length of the compressed time window reaches the given minimum nlsl . If an outlier region exists, the outlier process will be performed, that is, Mk2 will be divided into Ms1 and Ms2. The relationships among Ms1, Ms2, Mk,nlk , and nls are illustrated in Fig. 6, where nls ¼ nlk =2 . The compression of feature matrixes Ms1 and Ms2,nls , rs, and as(SP) with the outlier process are expressed as, respectively. h i Ms1 ¼ F SP2nls þ1 ; F SP2nls þ2 ; ; F SPnls h i Ms2 ¼ F SPnls þ1 ; F SPnls þ2 ; ; F SP vffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi u nl u s nls Y rs ¼ t F SPnls þi =F SP2nls þi ð14Þ ð15Þ i¼1 Pnls as ðSPÞ ¼ i¼1 F SPnls þi nls ð16Þ where Ms1 and Ms2 are the two adjacent feature matrixes of outlier time windows. rs and as(SP) represent the increase rate and FTR, respectively, after the outlier process, and nls is the length of the outlier time window. Based on this analysis, the final length of time window nlf , feature matrix Mf, increase rate rf, and wear factor gf can be determined by 6 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 Fig. 5. Time windows including outliers. nlf ¼ 8 l n0 r 0 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > > > > < nl without outlier; r 0 2 ð1; 1 þ g ðSPÞÞ or r0 2 ð1 þ g ; þ1Þ 0 max k > > and rk 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > > : l ns with outlier 8 M0 r 0 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > > > > < Mk without outlier; r 0 2 ð1; 1 þ g 0 ðSPÞÞ or r0 2 ð1 þ g max ; þ1Þ Mf ¼ > and r k 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > > > : Ms with outlier 8 r > 0 r 0 2 ½1 þ g 0 ðSP Þ; 1 þ g max > > > > < rk without outlier; r 0 2 ð1; 1 þ g 0 ðSPÞÞ or r0 2 ð1 þ g max ; þ1Þ rf ¼ > > > and r k 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > : rs with outlier 8 a0 ðSPÞ g max r0 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > > > > < ak ðSPÞ g max without outlier; r0 2 ð1; 1 þ g 0 ðSPÞÞ or r 0 2 ð1 þ g max ; þ1Þ gf ¼ > > > and rk 2 ½1 þ g 0 ðSPÞ; 1 þ g max > > : as ðSPÞ g max with outlier ð17Þ 2.2. Deep bidirectional LSTM 2.2.1. Basic LSTM The core idea of a basic long short-term memory (LSTM) neural network is that several gates, such as the input gate, forget gate, and output gate, are employed to control and select the information passing along the sequence to more accurately capture the contents of long-term memory. A popular LSTM framework found in [38] was used as a basis in this study and is shown in Fig. 7. At each time step t, the hidden state ht is updated by the hidden state ht-1 at time step t-1 with the current data xt, input gate f t, output gate ot, and memory cell ct. Here, the layer output zt is equal to ht. The updating process can be expressed as Fig. 6. Relationships among Ms1, Ms2, Mk,nlk , and nls . H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 7 Fig. 7. Basic LSTM framework. t1 i t i ¼ r W i xt þ V i h þ b t t1 f f ¼ r W f xt þ V f h þ b t1 o ot ¼ r W o xt þ V o h þ b t t1 c t ct ¼ f ct1 þ i tanh W c xt þ V c h þ b ð18Þ t zt ¼ h ¼ ot tanhðct Þ where W and V‧ are the weight matrixes and b‧ is the bias, both of which are shared by all time steps and obtained during model training. In this study, the LSTM model is trained by the backpropagation through time (BPTT) method, more details of which can be found in [39]. 2.2.2. Bidirectional LSTM Each layer in a traditional neural network represents the network inputs in a specific dimensional space, so a deeper network with more layers has a greater chance of discovering the complex relationship between the inputs and outputs in multi-dimensional space. The core idea of deep architecture is to employ a deeper network to represent input data in order to better describe the embedded patterns [40]. Motived by this reasoning, more LSTM layers are included in the proposed network to more effectively deal with system nonlinearity. Notably, a basic LSTM can only access the previous sensor data at each time step. However, for tool condition monitoring, the sensor data acquired from various sensors have strong temporal dependencies [41]. As a result, it is necessary to smooth the current data considering the future information. Therefore, a deep bidirectional LSTM (DBiLSTM) network is proposed as shown in Fig. 8, in which multiple LSTM layers are stacked together. The following equations define the DBiLSTM network, in which the labels ? and indicate the forward and backward paths, respectively. ! it ¼ r ! W i ! xt þ ! V i ! ht1 þ ! bi ! f t ¼ r ! W f ! xt þ ! V f ! ht1 þ ! bf ! ot ¼ r ! W o ! xt þ ! V o ! ht1 þ ! bo ! ct ¼ ! f t ! ct1 þ ! it tanh ! W c ! xt þ ! V c ! ht1 þ ! bc ð19Þ ! zt ¼ ! ht ¼ ! ot tanhð! ct Þ it ¼ r ft ¼ r ot ¼ r Wi xt þ Vi ht1 þ bi Wf xt þ Vf ht1 þ bf Wo xt þ Vo c ¼ f c zt ¼ ht ¼ ot tanhð t t zt ¼ h ¼ ! h t1 h t þ ht1 þ bo it tanh W c xt þ t t V c h t1 þ b c ð20Þ ct Þ ð21Þ The proposed DBiLSTM network is able to process sequential data in both the forward and backward directions, each of which are computed independently. The forward path is employed to discover the future trend of the information and the backward path is used to smooth the forward prediction result. The outputs of each path are fed forward to the same output 8 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 Fig. 8. Architecture of DBiLSTM. layer, and then concatenated and fed into the next layer. Combined with the feature matrix, in the DBiLSTM network, h i h i F SP2nl þ1 ; F SP2nl þ2 ; :::; F SPnl is the input sequence and F SPnl þ1 ; F SPnl þ2 ; :::; F SP is regarded as the output sequence, which f f f f f is obtained from both directional paths at each time step. 2.3. Multi-step ahead rolling prediction After determining the time window size and constructing the DBiLSTM network, the network must be trained using Mf1 and Mf2 as the training sample. Before being input into the DBiLSTM network, z-score standardization is employed to normalize Mf1 and Mf2, expressed as NFt = (Ft - mean(Mf)) / std.dev(Mf), where mean(Mf) and std.dev(Mf) are the mean and standard deviation, respectively, of feature matrix Mf. Even though there is just one training sample, there are many elements in Mf1 and Mf2. After training the DBiLSTM network, rolling prediction is performed. Generally, one-step and multi-step ahead prediction are used to predict system degradation [42,43]. However, when using one-step ahead prediction, only one point is obtained in each iteration, as shown in ½F 1 ; F 2 ; F 3 ; F 4 ! F 5 ½F 2 ; F 3 ; F 4 ; F 5 ! F 6 .. . ð22Þ ½F mþ1 ; F mþ2 ; F mþ3 ; F mþ4 ! F mþ5 ½F mþ2 ; F mþ3 ; F mþ4 ; F mþ5 ! F mþ6 As a result, training a one-step ahead prediction model is time-consuming. Therefore, in this study, the multi-step ahead prediction method is utilized as follows. h i h i NF SP2nl þ1 ; NF SP2nl þ2 ; NF SPnl ! NF SPnl þ1 ; NF SPnl þ2 ; ; NF SP f f f f f h i h i NF SPnl þ1 ; NF SPnl þ2 ; ; NF SP ! NF SPþ1 ; NF SPþ2 ; ; NF SPþnl f f f h i h i ð23Þ NF SPþðk1Þnl þ1 ; ; NF SPþknl ! NF SPþknl þ1 ; ; NF SPþðkþ1Þnl f f f f After being combined with the wear factor, the predicted result of the multi-step ahead process should be denormalized by DFt = NFt std.dev(Mf) + mean(Mf) and post-processed as follows H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 DF t ¼ 1 þ g f ðSPÞ DF t nl 9 ð24Þ f where DF t nl is larger than DFt. If DF t nl is smaller than DFt, DF t nl is equal to DFt. f f f When the rolling prediction value exceeds the given failure threshold FFT, the RUL is obtained as RUL ¼ t SP if DF t r f P F FT rolling prediction if DF t rf < F FT ð25Þ 2.4. Framework of proposed method The objective of this study is to predict the tool RUL under limited data once the predicted value Ft exceeds the predetermined threshold of tool failure. A flowchart of the procedures in the proposed method is presented in Fig. 9. Step 1: Collect the limited sensor signals and select the most sensitive features. Step 2: Calculate the FTR of initial time windows and compute the wear factor. Step 3: Construct the time window according to increase rate and wear factor. Step 4: Train DBiLSTM model using adjacent time windows. Step 5: Perform multi-step ahead rolling prediction until the predicted value exceeds the predetermined threshold. Otherwise, continue rolling prediction. Step 6: Output the RUL. 3. Experimental validation 3.1. Experimental setup To verify the effectiveness of the proposed tool RUL prediction method, several experiments are performed on the stainless-steel workpieces using the milling machine tool [31]. The schematic diagram of the experimental platform is shown in Fig. 10. A microscope is employed to provide offline measurements of the tool wear width. The key cutting parameters are listed as follows: two different feeds (0.5 mm/rev and 0.25 mm/rev), two different cut depths (1.5 mm and 0.75 mm), and constant cutting speed of 200 m/min (826 rev/min). All experiments are conducted by milling until the useful life of the tool inserts had been reached. This is done twice under each condition using a different set of inserts. After each insert is removed from the tool, a milling run using only the tool is completed. Table 1 presents the machining conditions and milling runs of the eight tool wear cases evaluated. Six sensors are installed on the machine tools as shown in Fig. 10: two vibration sensors (on the spindle and table respectively), two AE sensors (on the spindle and table respectively), and two motor current sensors (AC spindle motor current and DC spindle motor current). The sensor data are collected by a data acquisition board (MIO-16 from National Instrument Com- Fig. 9. Flowchart of the proposed tool RUL prediction method. 10 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 Fig. 10. Milling machine tool experiment setup. Table 1 Cutting parameters and milling runs. Case Cut depth Feed Run Total samples 1 2 3 4 5 6 7 8 0.75 1.5 0.75 1.5 1.5 1.5 0.75 0.75 0.5 0.5 0.25 0.25 0.5 0.25 0.25 0.5 14 17 14 7 9 10 23 15 1120 1360 1120 560 720 800 1840 1200 (a) Good condition (b) Medium wear (c) Severe wear Fig. 11. Spindle motor current of three working conditions for one run (Entry-Milling-Exit). 11 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 Table 2 List of extracted features and formulas. Domain Extracted feature Formula Time Maximum Mean F 1 ¼ maxðF i Þ n P Fi F 2 ¼ l ¼ 1n i¼1 sffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi n P F 3 ¼ 1n Fi2 Root mean square i¼1 Standard deviation Kurtosis Frequency sffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi n P 1 ðF i - lÞ2 n - 1 F4 ¼ r ¼ F 5 ¼ 1n Peak-to-Peak Maximum power spectrum Mean power spectrum Standard deviation power spectrum n P i¼1 i¼1 ðF i lÞ4 =r4 F 6 ¼ maxðF i Þ - minðF i Þ F 7 ¼ max Sð f Þi n P Sð f Þi F 8 ¼ lS ¼ 1n i¼1 sffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi n 2 P F 9 ¼ rS ¼ n 1- 1 Sð f Þi - l i¼1 Power spectrum kurtosis F 10 ¼ 1n n P i¼1 Sð f Þi lS 4 =r4S Table 3 Ranking result of LASSO for ten features. Feature F1 F2 F3 F4 F5 F6 F7 F8 F9 F10 R2 Ranking Case1 Case2 Case3 Case4 Case5 Case6 Case7 Case8 2 9 8 4 10 1 7 5 6 3 0.998 4 5 7 3 9 2 10 1 6 8 0.896 6 5 2 10 7 1 8 4 3 9 0.997 7 5 4 10 6 1 8 2 3 9 0.984 9 1 8 2 7 3 4 10 6 5 0.963 2 10 6 3 9 1 8 7 4 5 0.981 4 7 3 2 9 1 10 5 6 8 0.998 10 1 3 7 8 2 5 4 6 9 0.997 Fig. 12. Kurtosis of eight cases. pany) connected to a computer. As these data are acquired from a real milling experiment, not a simulation or collected from an experimental platform, the tool wear process can be considered realistic, providing important information for investigating the performance of the proposed model when linked to sensor data in a tool condition monitoring application [44]. Generally, cutting force is the most sensitive indicator in tool condition monitoring [45,46], which increases with the wear of tool. However, force-based measuring instruments are expensive and comparatively hard to install. The spindle motor current is closely related to cutting force and thus sensitive to tool wear condition. Therefore, employing spindle motor current to monitoring tool condition is an economical and effective online monitoring solution [34]. 12 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 Fig. 11 shows the motor current signal in good condition, medium wear and severe wear of Case1. Each milling run of spindle motor current signal contains approximately 9000 points and it includes three processes like tool entry, milling and tool exit, as shown in Fig. 11. Actually, only the motor current signals of milling process can well monitor the tool working condition, because the tool entry and tool exit represent the motor current signals when the spindle is idling without machining, which cannot reflect the tool condition. Therefore, to better monitor tool working condition, 4000 points from 3001 to 7000 of each fun are recorded during the steady milling stage and these points are divided into 80 data samples (each data sample contains 50 points with time constant 8 ms according to the analysis the experimental data). Thus, the total samples of each machining condition are illustrated in Table 1. Here, the data samples represent the useful life of tool under each working condition. The more data samples, the longer the useful life of tool. 3.2. Experimental results and discussion 3.2.1. Feature selection Before constructing time windows, we should extract features first from spindle motor current sensor and Table 2 illustrates ten features and their formulas. Feature selections are performed using Least Absolute Shrinkage and Selection Operator (LASSO) algorithm [47], which finally provides a ranking result for all the extracted features. LASSO is a data dimension reduction method, which is not only suitable for linear cases, but also for non-linear cases. It is based on penalty method to select variables from sample data. By compressing the original coefficients, the small coefficients are directly compressed to Table 4 Typical maximum wear factor selection under different v. v [1,2] (2, 5] (5, 10] (10, +1) gmax 0.5 0.3 0.2 0.1 Fig. 13. Distribution of wear factor. Table 5 Optimal DBiLSTM network structure parameters. Batch size Epoch Learning rate Neurons 32 200 0.0015 [140,150] Fig. 14. Proposed model using all historical data. 13 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 zero, thus the variables corresponding to these coefficients are regarded as non-significant variables, and these variables are discarded directly. Therefore, this is suitable for selecting the sensitive features in tool condition monitoring. Here, the Least Angle Regression (LAR) is employed to solve LASSO model [48]. By analyzing the ranking results of all features and correlation coefficient (R2) of eight cases in Table 3, the Kurtosis is ranked first and this means that the Kurtosis is the most sensitive feature to reflect the tool wear in all experiment. Therefore, we select the Kurtosis as the most sensitive feature Ft for construing time windows. Fig. 12 illustrates the Kurtosis of eight machining conditions and their values increases with the wear of tool. The failure threshold value of Kurtosis in eight cases are 14.2912, 7.6566, 4.8440, 6.4387, 15.1966, 10.1779, 4.7645 and 8.1258, respectively. 3.2.2. Result and discussion The maximum wear factor gmax is important to determine whether the initial time window should be compressed and to obtain the feature matrixes Mk1 and Mk2, increase rate rk, and feature threshold ratio ak(SP) after kth compression. The selection of wear factor should consider the characteristics of the tool and workpiece, such as material, geometric angle and coating. Tool material is a fundamental and critical indicator to determine the cutting performance of the tool, which has a great influence on the machining efficiency, quality, cost and tool life [49]. The harder the tool material is, the better its wear resistance and hardness are. Similarly, the material is also important to determine the workpiece performance. The workpiece with the higher hardness has strong resistance to permanent deformation. Thus, for one tool, the gmax changes with the change of the workpiece material. Therefore, the selection of gmax should consider both hardness of tool material and workpiece material. To quantify the hardness difference between tool and workpiece, the hardness ratio v is introduced in Eq. (26). The larger the v, the greater the hardness gap between tool and workpiece. Generally, the hardness of tool material should be larger than that of workpiece material. Therefore, the v is larger than 1. Based on this, the typical selection method of gmax is listed in Table 4. v¼ HRC ðToolÞ HRC ðWorkpieceÞ ð26Þ where HRC is Rockwell C hardness, which is a hardness evaluation index for material. In experiments, the material of cutting tool and workpiece are tungsten carbide and stainless steel, respectively, and v is 2.12. Therefore, combined with the wear factor distribution of the eight evaluated cases illustrated in Fig. 13, most wear fac- Table 6 Comparison of RUL prediction results for Case 1. SP RRUL 200 300 400 500 600 700 800 900 1000 MAE RMSE 920 820 720 620 520 420 320 220 120 PRUL (eight approaches) 1 2 3 4 5 6 7 8 936 784 442 492 590 398 303 225 108 0.1130 0.1592 723 619 487 645 916 678 629 520 267 0.6393 0.7790 366 312 247 333 514 342 326 265 117 0.3097 0.4041 1015 1007 595 412 564 326 415 103 168 0.2642 0.2969 Inf Inf Inf Inf Inf Inf Inf Inf Inf Inf Inf 2753 2362 857 494 379 279 192 176 114 0.6137 0.9412 930 976 447 736 458 380 259 201 58 0.1973 0.2470 595 789 925 163 470 381 52 234 34 0.3577 0.4699 Table 7 Comparison of RUL prediction results for Case 4. SP RRUL 200 300 400 500 MAE RMSE 360 260 160 60 PRUL (eight approaches) 1 2 3 4 5 6 7 8 297 269 151 58 0.0748 0.0950 386 409 251 116 0.5445 0.6192 201 249 151 77 0.2052 0.2636 425 325 133 81 0.2719 0.2835 Inf Inf Inf Inf Inf Inf 362 311 310 86 0.3931 0.5256 224 193 201 51 0.2604 0.2726 124 220 123 46 0.3182 0.2835 14 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 tors are less than 0.3 and thus, we define the maximum wear factor gmax = 0.3. Additionally, to simplify the calculation, we define g(t) = at gmax. The initial time window size is 100, and nlkl and nlsl are 50 and 30, respectively. In this study, Case 1 and Case 4 are utilized as the representative tool wear cases to validate the effectiveness of the proposed tool RUL prediction method as these cases both contain outliers, and their full lives are relatively long (1120 data samples) and short (560 data samples), respectively. It is therefore expected that if the proposed method is effective for these cases, it will also be effective for the remaining six cases. h i h i In the proposed DBiLSTM network, the input is F SP2nl þ1 ; F SP2nl þ2 ; :::; F SPnl and the output is F SPnl þ1 ; F SPnl þ2 ; :::; F SP , f f f f f where Ft = [Ft1, Ft2, Ft3, Ft4, Ft5]T and Fti (i = 1, 2, 3, 4, 5) denote the mean in interval [(i-1) 10 + 1, i 10] at time step t. In post-processing, DFt is the maximum of Fti. Using the multi-step ahead rolling prediction, if DFt exceeds FFT, the RUL is SP = 200 SP = 300 SP = 400 SP = 500 SP = 600 SP = 700 Fig. 15. Results using Approaches 1–3 in Case 1. H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 15 SP = 900 SP = 800 SP = 1000 Fig. 15 (continued) obtained. Additionally, the training works are conducted on a Windows server with Inter Xeon E5-2650, 2.30 GHz and a RAM of 112 GB. The tunable parameters contain the number of neurons in each layer, batch size, learning rate, and number of epochs. It should be noted here that the definitions of the parameters used in the DBiLSTM network are same as those in used the basic LSTM network [38]. Here, the k-fold cross-validation method is employed to obtain the optimal network structure with the final parameters provided in Table 5. Note that among the neurons, the forward and backward paths share the same [140, 150] neurons in the two layers. To demonstrate the performance of the proposed model, the following approaches are used to predict the tool RUL and the results are compared. Approach 1: The proposed approach in this paper consisting of the adaptive time window, DBiLSTM, and multi-step ahead rolling prediction. Approach 2: The proposed model with a fixed wear factor of 0.1 in the whole prediction process and the other parameters the same as in Approach 1. This approach is evaluated because several publications have assumed that the tool wear process is linear [50], so the wear factor is fixed. Additionally, as the gmax is 0.3, it indicates that the wear factor ranges from 0 to 0.3. In the incipient wear process, the wear factor is less than 0.1, so the fixed wear factor is 0.1. Approach 3: The proposed model with a fixed wear factor of 0.2 in the whole prediction process and the other parameters the same as in Approach 1, also evaluated to replicate a linear tool wear process [50]. Combined with the gmax = 0.3, the wear factor of the severe wear process is larger than 0.2, so the fixed wear factor is 0.2. Approach 4: The proposed model using all historical data and the other parameters the same as in Approach 1, as shown in Fig. 14. This approach is evaluated because most tool RUL prediction research is based on all historical data [31]. Approach 5: The proposed model without post-processing, but the prediction results must satisfy DF t P DF tnl f ð27Þ Approach 6: An exponential regression model using all condition monitoring data before SP [51]. F t ¼ q1 expðq2 tÞ ð28Þ 16 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 Approach 7: The proposed method using a basic LSTM network instead of the DBiLSTM network and the other parameters the same as in Approach 1 [41]. Approach 8: The proposed method using an ANN network instead of the DBiLSTM network [52]. As the parameters of an ANN network are tunable, the final epochs, learning rate, and goal are set to 5000, 0.01, and 1 10-7, respectively. The number of input layers, hidden layers, and output layers are nlf ,nlf + 1, and nlf , respectively, and the activation function of the hidden layer and output layer are tansig and logsig, respectively. The mean absolute error (MAE) and root mean square error (RMSE) are adopted to evaluate the performance of the different RUL prediction approaches as follows. MAE ¼ n 1X jPRULi RRULi j n i¼1 RRULi ð29Þ vffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi u n u1 X PRULi RRULi 2 RMSE ¼ t n i¼1 RRULi ð30Þ where PRUL and RRUL denote the predicted RUL and real RUL, respectively. Tables 6 and 7 list the RUL predictions for Case 1 and Case 4 under these different approaches. Note that the MAE and RMSE for the proposed method (0.1130 and 0.1592, respectively, in Case 1; 0.0748 and 0.0950, respectively, in Case 4) are smaller than for the other approaches, verifying the feasibility and accuracy of the proposed method. In particular, comparing Approaches 1, 2, and 3 in Tables 6 and 7, the proposed method (Approach 1) achieve more accurate prediction results and track the tool condition well when using limited data. As can be observed in Figs. 15 and 16, Approach 2 exhibits better performance than Approach 3 in the incipient wear phase (as seen at an SP from 200 to 500 in Case 1 and of 200 in Case 4) and worse performance in the later phase due to its smaller fixed wear factor. This smaller wear factor requires more multi-step ahead prediction iterations to reach the failure threshold value, which is suitable for SP = 200 SP = 300 SP = 400 SP = 500 Fig. 16. Results using Approaches 1–3 in Case 4. H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 17 RUL prediction in the incipient phase. Conversely, the larger fixed wear factor in Approach 3 exhibits better performance than Approach 2 in the later phase (as seen at an SP from 600 to 1000 in Case 1 and from 300 to 500 in Case 4) and worse performance in the incipient phase. The larger wear factor reaches the failure threshold value with fewer iterations, thus the PRUL is closer to the RRUL in the later phase. Notably, the proposed method (Approach 1) can dynamically adjust the wear factor according to the changing FTR, therefore, it exhibits better performance than Approaches 2 or 3 in both the incipient and later phases, validating the effectiveness of the adaptive time window in monitoring tool condition. The associated results of these methods are shown in Figs. 15 and 16, and Tables 6 and 7. The comparison results in Tables 6 and 7 of Approach 1 and Approaches 4, 5, and 6 validate the effectiveness of applying the time window and multi-step ahead rolling prediction. Approach 4 uses all historical data to construct time window, but prediction results are not as accurate as the proposed method. Results from Cases 1 and 4 using Approach 4 illustrate that too large of a time window is inappropriate as it results in a prediction that is over-dependent on the global trend, ignoring the SP = 200 SP = 300 SP = 400 SP = 500 SP = 600 SP = 700 Fig. 17. Case 1 results using Approach 5. 18 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 SP = 900 SP = 800 SP = 1000 Fig. 17 (continued) local trend of tool wear process. Especially in the later phase, for SP from 800, 900, 1000 in Case 1, the MAE values of Approach 4 (0.3125, 0.2045 and 0.4) are larger than those of Approach 1 (0.0531, 0.0227, and 0.1) respectively. Approach 5 performs its predictions without post-processing. Figs. 17 and 18 show that for both Cases 1 and 4, as the prediction results cannot exceed the failure threshold, this approach cannot accurately predict the tool RUL because it ignores the increase of wear factor in later phase due to the ongoing wear of tool. Though some researchers have utilized all historical data including failure data to train an exponential model and achieve better RUL prediction results [51], the predictions of Approach 6 indicate that the exponential model using all the historical data is inaccurate, as it is also overly dependent on global trends. Therefore, the proposed method is not only shown to trace the tool wear process well, but avoids wasting the resources required to construct the training samples [36]. Approaches 7 and 8 employ the basic LSTM and ANN, respectively, to predict tool RUL [41,25]. As can be observed in Tables 6 and 7, the results of the ANN method are inferior to those of the two LSTM-based methods in long-term prediction. The main reason for this inferiority is the inherent shortcomings of the ANN network, which performs well for only shortterm prediction; in the long-term, the prediction will converge to a certain value or fluctuate in certain range due to the poor extrapolability caused by the range of activation functions in the ANN. In this study, the stacking of the forward path, backward path, and multiple layers into the basic LSTM clearly allows the proposed DBiLSTM model to more accurately track the wear process and thus provide superior prediction performance. Changes in tool working conditions can also change the tool wear process and cause rapid increases or decreases in the spindle motor current. This behavior is often manifested as outliers in the data that present a challenge when predicting tool RUL. Figs. 19 and 20 show the kurtosis of the predicted RUL with and without outliers in the time window. In these figures, the kurtosis near the SP rapidly increases, causing a jump in the increase rate. For Case 1, as shown in Fig. 19 (SP = 400), it should be noted that the initial time windows are compressed according to the increase rate (see in Fig. 19 (a)–(b)) and that r0 is 1.3018, which can be regarded as an outlier region. However, after compression, the time windows cannot fully reflect the local increase trend, and this will affect the FTR, as can be observed in Section I of Fig. 19 (b). If these time windows are employed to train the DBiLSTM model directly, the trained model will add the information of Section I to the memory, where the kurtosis is lower than that of Section II, so the prediction will deviate from the real wear process. After the outlier process as detailed in Section 2.1.2.2, the FTR increases from 0.3392 to 0.3497 and the MAE decreases from 0.3827 to 0.2608. Thus, the use of compressed time windows is able to more accurately trace local trends and obtain a better RUL prediction, as H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 SP = 200 SP = 300 SP = 400 SP = 500 19 Fig. 18. Case 4 results using Approach 5. shown in Fig. 19 (c). Similarly, for Case 4, the outlier region is SP = 500 and r0 is 1.3211, as shown in Fig. 20. After processing this outlier, the FTR also increases from 0.8018 to 0.8333 and the MAE decreases from 1.0330 to 0.0333. Most of the existing approaches evaluated in this study utilize the failure and truncation data of the same or similar tools to train models and then employ these models to predict tool RUL. These approaches are centered on group-based inference. However, research has shown that the wear processes of the same tool can be different, which is called individual-based inference. The proposed method using limited data is strongly associated with the current condition of the tool and weakly correlated with the overall wear process of the same or similar tools. Even though some degree of error in the RUL predicted by the proposed method is unavoidable, it is still meaningful for addressing individual-based inference problems in cases of limited data. 4. Conclusions Most existing approaches for tool remaining useful life (RUL) prediction are based on truncation and failure data. However, it is difficult to acquire these data for a new type of tool or for a similar tool that has just launched. Therefore, the prediction of tool RUL using limited data is a considerable challenge. A novel tool RUL prediction method is therefore proposed in this paper based on an adaptive time window and deep bidirectional long short-term memory (DBiLSTM) neural network. In this method, the wear factor is first determined by the current feature threshold ratio (FTR). Then, time windows are constructed to trace tool conditions well, and adjacent time windows are adaptively compressed while accounting for any outliers. On this basis, a DBiLSTM is trained to capture and discover meaningful features using both forward and backward paths. Finally, multi-step ahead rolling prediction is performed to obtain RUL results. The proposed method is validated using the results of milling experiments. The mean absolute error and root mean square error of the proposed method are 0.1130 and 0.1592, respectively, for milling Case 1 (with a long tool life), and 0.0748 and 0.0950, respectively, for milling Case 4 (with a short tool life). In all cases, the errors of the proposed method are lower than those for the previously proposed approaches evaluated. Indeed, a comparison of results using eight different RUL approaches verify the advantages of the proposed RUL prediction method using the adaptive time window, DBiLSTM, and multi-step ahead rolling prediction. In particular, the proposed method has been demonstrated to effectively address 20 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 (b) Compressed time window (a) Initial time window (c) Outlier region Fig. 19. Outlier processing for Case 1. Fig. 20. Outlier region in Case 4. the effects of outlier regions of the predicted RUL. Therefore, our proposed method is able to accurately trace tool condition using limited data and offers a promising approach for the prediction of remaining useful life in other systems. In future work, we intend to investigate methods for feature extraction that could improve the accuracy of RUL prediction. Additionally, the integration of the RUL predicted using the proposed method into production scheduling represents a promising area of research. CRediT authorship contribution statement Hai Li: Writing - review & editing, Writing - original draft. Wei Wang: Supervision. Ziwei Li: Software, Validation. Liyi Dong: Software. Qingzhao Li: Validation. H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 21 Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgement The authors would like to thank the Major Project of National Science and Technology (No. 2017ZX04002001), the Fundamental Research Funds for the Central Universities (ZYGX2019J032), NSAF (U1830110) and Foundation of key laboratory of ultra-precision machining technology in Chinese academy of engineering physics (K1126-17-Y). References [1] P. Mehta, A. Werner, L. Mears, Condition based maintenance-systems integration and intelligence using Bayesian classification and sensor fusion, J. Intell. Manuf. 26 (2) (2015) 331–346. [2] A.K.S. Jardine, D. Lin, D. Banjevic, A review on machinery diagnostics and prognostics implementing condition-based maintenance, Mech. Syst. Signal Process. 20 (7) (2006) 1483–1510. [3] S.X. Si, W. Wang, C.H. Hu, H.Z. Dong, Remaining useful life estimation–A review on the statistical data driven approaches, Eur. J. Oper. Res. 213 (1) (2011) 1–14. [4] H.Z. Sikorska, M. Hodkiewicz, L. Ma, Prognostic modelling options for remaining useful life estimation by industry, Mech. Syst. Signal Process. 25 (5) (2011) 1803–1836. [5] A. Heng, S. Zhang, A.C.C. Tan, J. Mathew, Rotating machinery prognostics: State of the art, challenges and opportunities, Mech. Syst. Signal Process. 23 (3) (2009) 724–739. [6] J. Lee, F. Wu, W. Zhao, M. Ghaffari, L.X. Liao, D. Siegel, Prognostics and health management design for rotary machinery systems—Reviews, methodology and applications, Mech. Syst. Signal Process. 42 (1–2) (2014) 314–334. [7] I.S. Jawahir, R. Ghosh, X.D. Fang, P.X. Li, An investigation of the effects of chip flow on tool-wear in machining with complex grooved tools, Wear 184 (2) (1995) 145–154. [8] P. Mathew, Use of predicted cutting temperatures in determining tool performance, Int. J. Mach. Tool. Manu. 29 (4) (1989) 481–497. [9] H. Takeyama, R. Murata, Basic Investigation of Tool Wear, J. Eng. Industry. 85 (1) (1963) 33–37. [10] E. Usui, A. Hirota, M. Masuko, Analytical prediction of three dimensional cutting process-Part 1: basic cutting model and energy approach, J. Eng. Industry. 100 (2) (1978) 222–228. [11] H. Sun, D. Cao, Z. Zhao, X. Kang, A hybrid approach to cutting tool remaining useful life prediction based on the Wiener process, IEEE Trans. Reliab. 67 (3) (2018) 1294–1303. [12] J. Wang, P. Wang, R.X. Gao, Enhanced particle filter for tool wear prediction, J. Manuf. Syst. 36 (2015) 35–45. [13] P.W. Marksberry, I.S. Jawahir, A comprehensive tool-wear/tool-life performance model in the evaluation of NDM (near dry machining) for sustainable manufacturing, Int. J. Mach. Tool. Manu. 48 (7–8) (2008) 878–886. [14] Y. Wang, C. Deng, J. Wu, Y. Xiong, Failure time prediction for mechanical device based on the degradation sequence, J. Intell. Manuf. 26 (6) (2015) 1181–1199. [15] D. Kong, Y. Chen, N. Li, Gaussian process regression for tool wear prediction, Mech. Syst. Signal Process. 104 (2018) 556–574. [16] A. Kumar, R.B. Chinnam, F. Tseng, An HMM and polynomial regression based approach for remaining useful life and health state estimation of cutting tools, Comput. Ind. Eng. 128 (2019) 1008–1014. [17] P. Wang, R.X. Gao, Adaptive resampling-based particle filtering for tool life prediction, J. Manuf. SysT. 37 (2015) 528–534. [18] W. Li, T. Liu, Time varying and condition adaptive hidden Markov model for tool wear state estimation and remaining useful life prediction in micromilling, Mech. Syst. Signal Process. 131 (2019) 689–702. [19] W. Li, S. Zhang, S. Rakheja, Feature Denoising and Nearest-Farthest Distance Preserving Projection for Machine Fault Diagnosis, IEEE Trans. Ind. Inform. 12 (1) (2016) 393–404. [20] C. Sun, P. Wang, R. Yan, R.X. Gao, X. Chen, Machine health monitoring based on locally linear embedding with kernel sparse representation for neighborhood optimization, Mech. Syst. Signal Process. 114 (2019) 25–34. [21] G. He, K. Ding, W. Li, Y. Li, Frequency response model and mechanism for wind turbine planetary gear train vibration analysis, IET Renew. Power. Gen. 11 (4) (2017) 425–432. [22] X. Jiang, S.L.Y. Wang, Study on nature of crossover phenomena with application to gearbox fault diagnosis, Mech. Syst. Signal Process. 83 (2017) 272– 295. [23] B. Chen, X. Chen, B. Li, Z. He, H. Cao, G. Cai, Reliability estimation for cutting tools based on logistic regression model using vibration signals, Mech. Syst. Signal Process. 25 (7) (2011) 2526–2537. [24] R.K. Venkata, B.S.N. Murthy, R.N. Mohan, Prediction of cutting tool wear, surface roughness and vibration of work piece in boring of AISI 316 steel with artificial neural network, Measurement 51 (2014) 63–70. [25] D.M. D’Addona, A.M.M.S. Ullah, D. Matarazzo, Tool-wear prediction and pattern-recognition using artificial neural network and DNA-based computing, J. Intell. Manuf. 28 (6) (2017) 1285–1301. [26] T. Mikołajczyk, K. Nowicki, A. Bustillo, D.Y. Pimenov, Predicting tool life in turning operations using neural networks and image processing, Mech. Syst. Signal Process. 104 (2018) 503–513. [27] C. Drouillet, J. Karandikar, C. Nath, A.C. Journeaux, Tool life predictions in milling using spindle power with the neural network technique, J. Manuf. Processes. 22 (2016) 161–168. [28] A.P. Kene, S.K. Choudhury, Analytical modeling of tool health monitoring system using multiple sensor data fusion approach in hard machining, Measurement 145 (2019) 118–129. [29] J. Karandikar, Machine learning classification for tool life modeling using production shop-floor tool wear data, Procedia Manuf. 34 (2019) 446–454. [30] Y. Yang, Y. Guo, Z. Huang, N. Chen, L. Li, Y. Jiang, N. He, Research on the milling tool wear and life prediction by establishing an integrated predictive model, Measurement 145 (2019) 178–189. [31] C. Zhang, H. Zhang, Modelling and prediction of tool wear using LS-SVM in milling operation, Int. J. Comput. Integ. Manuf. 29 (1) (2016) 76–91. [32] Y. Lei, N. Li, L. Guo, N. Li, T. Yan, J. Lin, Machinery health prognostics: A systematic review from data acquisition to RUL prediction, Mech. Syst. Signal Process. 104 (2018) 799–834. [33] K. Medjaher, N. Zerhouni, Framework for a hybrid prognostics, Chem. Eng. Trans. 33 (2013) 91–96. [34] H. Hanachi, W. Yu, I.Y. Kim, J. Liu, C.K. Mechefske, Hybrid data-driven physics-based model fusion framework for tool wear prediction, Int. J. Adv. Manuf. Tech. 101 (2019) 2861–2872. [35] Z. Tian, M.J. Zuo, Health condition prediction of gears using a recurrent neural network approach, IEEE Trans. Reliab. 59 (4) (2010) 700–705. 22 H. Li et al. / Mechanical Systems and Signal Processing 143 (2020) 106832 [36] C. Hu, B.D. Youn, P. Wang, J.T. Yoon, Ensemble of data-driven prognostic algorithms for robust prediction of remaining useful life, Reliab. Eng. Syst. Safe. 103 (2012) 120–135. [37] R. Nazir, Taylor series expansion based repetitive controllers for power converters, subject to fractional delays, Control Eng. Pract. 64 (2017) 140–147. [38] F.A. Gers, J. Schmidhuber, F. Cummins, Learning to forget: Continual prediction with LSTM, Neur. Comput. 12 (2000) 2451–2471. [39] K. Greff, R.K. Srivastava, J. Koutník, B.R. Steunebrink, J. Schmidhuber, LSTM: A search space odyssey, IEEE Trans. Neur. Net. Lear. Syst. 28 (10) (2016) 2222–2232. [40] G.E. Hinton, Learning multiple layers of representation, Trends Cogn. Sci. 11 (10) (2007) 428–434. [41] R. Zhao, J. Wang, R. Yan, K. Mao, Machine health monitoring with LSTM networks, in: in: 2016 10th International Conference on Sensing Technology (ICST), IEEE, 2016, pp. 1–6. [42] Z. Tian, L. Wong, N. Safaei, A neural network approach for remaining useful life prediction utilizing both failure and suspension histories, Mech. Syst. Signal Process. 24 (5) (2010) 1542–1555. [43] V.T. Tran, B.S. Yang, A.C.C. Tan, Multi-step ahead direct prediction for the machine condition prognosis using regression trees and neuro-fuzzy systems, Expert Syst. Appl. 36 (5) (2009) 9378–9387. [44] J. Yu, Tool condition prognostics using logistic regression with penalization and manifold regularization, Appl. Soft. Comput. 64 (2018) 454–467. [45] V. Jain, T. Raj, Tool life management of unmanned production system based on surface roughness by ANFIS, Int. J. Syst. Assu. Eng. Manage. 8 (2) (2017) 458–467. [46] N. Ghosh, Y.B. Ravi, A. Patra, S. Mukhopadhyay, A.R. Mohanty, A.B. Chattopadhay, Estimation of tool wear during CNC milling using neural networkbased sensor fusion, Mech. Syst. Signal Process. 21 (1) (2007) 466–479. [47] C.Y. Lee, J.Y. Cai, LASSO variable selection in data envelopment analysis with small datasets, Omega. (2018), https://doi.org/10.1016/j. omega.2018.12.008. [48] B. Efron, T. Hastie, I. Johnstone, R. Tibshirani, Least angle regression, Annals Statis. 32 (2) (2004) 407–499. [49] M. Verma, S.K. Pradhan, Experimental and numerical investigations in CNC turning for different combinations of tool inserts and workpiece material [J], Mater. Today. Proc. (2020), https://doi.org/10.1016/j.matpr.2019.12.193. [50] Z. Yang, Y.X. Chen, Y.F. Li, E. Zio, R. Kang, Smart electricity meter reliability prediction based on accelerated degradation testing and modeling, Int. J. Elec. Power. Ener. Syst. 56 (2014) 209–219. [51] H. Hanachi, W. Yu, I.Y. Kim, J. Liu, C.K. Mechefske, Hybrid data-driven physics-based model fusion framework for tool wear prediction, Int. J. Adv. Manuf. Tech. 101 (9–12) (2019) 2861–2872. [52] C. Drouillet, J. Karandikar, C. Nath, A.C. Joureaux, M.E.I. Mansori, T. Kurfess, Tool life predictions in milling using spindle power with the neural network technique, J. Manuf. Process. 22 (2016) 161–168.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )