IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 719 Estimating Mean Residual Life for a Case Study of Rail Wagon Bearings Alireza Ghasemi and Melinda R. Hodkiewicz Abstract—This paper develops a prognostics model to estimate the Mean Residual Life of Rail Wagon Bearings within certain confidence intervals. The prognostics model is constructed using a Proportional Hazards approach, informed by imperfect data from a bearing acoustic monitoring system, and a failure database. The model supports prediction within a defined maintenance planning window from the time of receipt of the latest acoustic condition monitoring information. We use the model to decide whether to replace a bearing, or leave it until collection of the next condition monitoring indicators. The model is tested on a limited number of cases, and demonstrates good predictive capability. Opportunities to improve the performance of the model are identified, and the processes necessary and time required to build the model are described. Lessons learned from dealing with real field data will assist those interested in using prognostics to support maintenance planning activities. Probability of getting condition indicator value while the equipment is in state ; an element of the observation probability matrix Index Terms—Acoustic monitoring, condition monitoring, imperfect information, prognostics, proportional hazards model. The hazard function of the proportional hazards model at time while the system state is Current value of the indirect indicator of the system’s degradation state History of the indirect indicator values up to time Probability of going from state to state between two consecutive observations (inspections), knowing that the equipment has not failed during that interval; an element of the transition matrix Baseline hazard function ABBREVIATIONS MRL Mean Residual Life PH Model Proportional Hazards Model HMM Hidden Markov Model CBM Condition Based Maintenance MLE Maximum Likelihood Estimation CI Confidence Interval State effect function Conditional probability distribution of the equipment’s state at observation moment , Probability of being in state at observation where is an element moment , of the vector Equipment’s state conditional probability distribution history up to time Survival function at time when the history of the indicator values up to time is NOMENCLATURE Failure time of the equipment Observation interval Probability distribution function of the time-to-failure at time when the history of the indicator values up to time is Equipment’s degradation state at time Conditional reliability at observation moment for a period of , knowing that the state is Equipment’s degradation state after the -th observation interval; Manuscript received February 06, 2012; revised April 10, 2012; accepted April 11, 2012. Date of publication July 31, 2012; date of current version August 28, 2012. This work was supported by Australian Government Endeavour Post-Doctoral Fellowship Award. Associate Editor: E. Zio. A. Ghasemi is with Department of Industrial Engineering, Dalhousie University, Halifax, NS B3H 4R2 Canada (e-mail: alireza.ghasemi@dal.ca). M. R. Hodkiewicz is with University of Western Australia, M050, Crawley, WA 6009 Australia (e-mail: Melinda.Hodkiewicz@uwa.edu.au). Color versions of one or more of the figures in this paper are available online at http://ieeexplore.ieee.org. Digital Object Identifier 10.1109/TR.2012.2209251 0018-9529/$31.00 © 2012 IEEE Conditional mean sojourn time at period knowing that the state is Conditional reliability of the equipment for a period of , at observation moment , while the conditional probability distribution of the equipment’s state is Mean residual life at observation moment while the state is Mean residual life at observation moment while the conditional probability distribution of the equipment’s state is 720 IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 I. INTRODUCTION II. LITERATURE REVIEW HIS RESEARCH project develops a prognostics algorithm to determine the Mean Residual Life (MRL) of rail wagon bearings for use in maintenance planning. While the development of prognostic algorithms is a well-established field, academically successful implementations by industry have only had limited success [1]. The ability of those responsible for asset maintenance planning to 1) perform the prognostic modeling, and 2) deliver the outputs to those responsible for the maintenance action via existing business systems as a routine part of the maintenance cycle, is dependent on a number of factors. These factors include the availability of appropriate condition indicators, skilled modeling personnel (in house or contracted), appropriate computing and data exchange infrastructure, and a disciplined work management processes. This list is in addition to the organizational processes involved in managing a routine condition-monitoring program cost-effectively [2]. The rail industry uses a variety of monitoring tools to assess the condition of the rail wagon bearings. These tools include, but are not limited to, thermal (hot bearing detection and hot wheel detection), acoustic (rail bearing acoustic monitor), and weight-based techniques (wheel impact detector). These tools provide condition indicators to be used for diagnosis. Diagnostic interpretations and synthesis of the different indicators is generally done by rail staff, is predominantly intuitive, and success is correlated with the experience of personnel familiar with the equipment. One of the main challenges is the number of bearings monitored. In this case study, there are tens of thousands of rail car wheel bearing assemblies in service. In recent years, there has been increased expectation of availability for rail systems, and hence on rail cars, with little spare capacity in the system for unplanned failures. This requirement has increased pressure of the rail staff to ensure that incipient failures are detected, correctly diagnosed, and that interventions are scheduled to minimize risk of unplanned failure, and hence disruption to the trains’ schedule. The aim of the project is to develop and test a prognostics model to identify bearings that are going to fail in a specific schedule window (15 pass-bys) so they can plan maintenance intervention. A pass-by is defined as the action of the train passing by a specific point in the rail network. This pass-by definition will allow maintenance planners to schedule the rail cars on which these bearings are mounted to be brought into the maintenance facility in a planned way. This approach reduces the risk of unplanned rail bearing failures, which can result in derailment or expensive repairs if the failure is in an isolated location with associated impacts on train schedule compliance. The challenge with this project is that the data come from a number of systems that have evolved separately, have different coding to identify the assets, and possess the usual issues with maintenance data quality [3]–[5]. This report describes the steps involved in developing the prognostics model, the technical and organizational assistance and barriers, the data requirements, and the factors that impact on confidence in the results. Condition monitoring is the act of observing and collecting information concerning the degradation condition of equipment to schedule maintenance actions prior to failure. Data concerning one or more indicators of degradation are collected periodically to establish a diagnosis of the equipment’s condition, and a prognosis for future performance. A combination of human expertise, signal processing, and modeling is used through this process. Considerable work has been done in the last 20 years to develop models to assist with the prognosis processes [6]–[8]. Jardine et al. [6] explains that the approaches to prognosis fall into three main categories: statistical approaches, artificial intelligent approaches, and model-based approaches. Statistical approaches are usually applied to predict the chance that a machine operates without failure up to a future time, given the current machine condition, and past operational profile. These approaches require considerable data and information on historical performance. A well-used statistical approach for prognostics with condition monitoring data is proportional hazards (PH) modeling. PH Models have been used to address many practical (e.g. [9]–[11]) and theoretical or experimental (e.g. [12]–[14]) health monitoring and prognostics problems. Artificial intelligence approaches, also known as data-driven approaches, are derived directly from routine condition monitoring (CM) data (e.g. temperature, vibration, oil debris, current, etc.). These methods predict the selected features that correlate with the failure progression based on the learning or training process. In general, data-driven approaches adopt a one-step or multi-step ahead prediction technique to predict the future state. Several data-driven prognosis methods have been developed and published [15]–[17] and [12]. A main drawback of data-driven approaches is their dependency on the quality of the monitored data, and there is limited physical understanding usually provided. Realistically, information may contain noise due to errors of measurement, of interpretation, or due to the limited accuracy of the measurement’s instruments, and it may not reveal the exact state of the equipment. The information is, however, stochastically correlated with the degradation state. In this case, the information collected may be referring to more than one possible state with different probabilities. This category of condition monitoring is referred to as “Indirect Monitoring” [18], “Partial Observation” [19], or Imperfect Observation [20]. This imperfection of information limits the application of data-driven and statistical approaches in many cases. Model-based approaches (also known as physics-based approaches) assume that accurate mathematical or physical models are available, and provide a technically comprehensive approach that has been used traditionally to understand failure mode progression, e.g. the crack propagation model. The drawback of this approach is its specificity, which means it can only be applied to specific types of components and failure modes limiting its flexibility. A common challenge with all of these approaches is the reliance for development or validation or both, on historical field failure and operational context data. Field data are rarely T GHASEMI AND HODKIEWICZ: ESTIMATING MEAN RESIDUAL LIFE FOR A CASE STUDY OF RAIL WAGON BEARINGS “clean,” often suffering from the effects of environmental conditions and other factors. Field data often come as a by-product of maintenance records, and are not originally aimed to help reliability practitioners. The data collected are usually the times of events which may be the failure or replacement of components. These data have to be matched with the condition monitoring data gathered throughout the equipment’s life. This task is usually a challenge, as the data for condition monitoring and failures are usually stored in separate systems, data arrives at different times, and the data may not always be structured in the same data hierarchy. Hence, considerable time and expert input are often required to create a common database of failure and health data for a specific asset or component. This paper is concerned with prognostics of mechanical rotating components. It is focused on the identification of potential for a catastrophic failure, or the occurrence of bearing faults such as cup or ball defects. It draws on two important developments that have enhanced the survival analysis method: 1) construction of a survival curve from censored data, derived from a nonparametric method introduced by Kaplan and Meier [21], and 2) construction of the degradation curve using a Proportional Hazard Models, introduced by Cox [22], [23], with censored data to estimate the survival function, and investigate the effect of covariates on the failure time distribution. However, one of the main on-going challenges with the use of the PH Model is that field data are assumed to be perfect, and can be directly used as the covariates of the PH Model. In [9]–[28], the authors are fortunate to have “perfect” data, or have assumed that the data are perfect, and reflects the exact condition of the system state. This assumption has allowed them to use the data directly in the PH Model. Where field data have not been available, researchers have used simulated or laboratory experimental data [13], [29]. In the PH Model process, it is often necessary to describe the behavior of the covariates. Early work on this area was done by [30], who used a non-homogenous discrete Markov process to predict future development of covariates, and failure times. This work results in the failure model containing two sub models: one distribution model with covariates, and the second a degradation model representing the behavior of covariates with respect to various health states. Further developments combining PH Models and HMM include [31], and [32]. The issues of determining the health state relative to other states, or as an overall health level, was explored in [33]. The combination of PH Model and HMM can also be used to address the imperfect observation problem, and to manage one of the challenges in the time-dependent PH Model, which is the inclusion of only the latest condition monitoring information in the model. This technique was demonstrated in and tested on simulated data in [20]. This paper extends the work by testing the model proposed in [20] using imperfect acoustic data from condition monitoring, and failure and repair data for rail wagon bearings from maintenance databases in a heavy haulage rail system. There is very little published work on the prognostics of bearing failure in heavy haulage rail wagons using either or both vibration and acoustics as condition indicators. There have been developments in systems to measure different condition 721 Fig. 1. Key phases in the model development and testing. indices including acoustic emission [34], [35], and work has been done to diagnose relevant faults [36], [37]. Work to understand failure distributions for rail wagon bearings is reported in [38]. This paper therefore contributes both data and a model for prognostics specifically for rail wagon bearings. The paper is organized as follows. Section III describes the mathematics underpinning the development of the model, the identification and preparation of the performance and failure data, and the modeling approach to estimating residual life. Section IV presents the results including the challenges faced with “real” data, and the outcome of testing a replacement policy based on a specified maintenance planning window. Section V discusses the results of the parameter estimation, importance of bearing type, model performance and confidence, and recommendations for dealing with “real” data from multiple systems. The conclusions are presented in Section VI. III. APPROACH This section describes the two phases of the project: 1) model development and parameter estimation, and 2) use of the model to estimate mean residual life. The main phases in this development are shown in Fig. 1. A. Proportional Hazards Models The prognostics model is based on the PH Model proposed by Cox [22], but assumes that the condition monitoring is indirect; i.e., at each observation moment, an indicator value of the underlying degradation state is available. Observations are collected at a constant (or a near constant) interval . In this study, represents the degradation state of the equipment at time , which will be used as the diagnostic covariate in the PH Model, and is a value from the set of all the possible indicator values . In other words, the set of indicator values is discretized into a finite set of possible values. The equipment’s condition is described as follows. 722 IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 • The equipment has a finite, known number of degradation states . is the set of all possible degradation states. • The degradation state transition is modeled by a HMM. The transition matrix is , where is the probability of going from state to state , during one observation interval, knowing that the equipment does not fail before the end of the interval. • The value of the indicator is stochastically related to the equipment’s state through the observation probability matrix, , , . is the probability of getting the indicator value , while the equipment is in degradation state . • The indicator is collected periodically at fixed or near fixed intervals . • Failure is not a degradation state. It is a non-working or dangerous condition of the equipment that can happen at any time, while the system is in any degradation state. To deal with unobservable degradation states, a new transition rule is introduced. This rule allows for all observations from the last renewal moment up to current moment to be included when calculating the conditional probability of being in degradation state at the th observation moment [39]: history into the state conditional probability distribution history , , , , up to time , accordingly. The elements of are calculated by (1) and (2) at corresponding observation moments, when a new indicator value is available [39]. By assuming that the equipment’s unobservable degradation state transition follows a stationary Markov Process, the transition probability at time , from state to state , knowing that the equipment has survived at least until the next observation moment, is The survival function of the assumed model is [40] where (3) The conditional probability distribution of the equipment’s degradation state at period is defined as is the PH Model’s hazard function at time while and the equipment’s degradation state is . Based on the MLE, one adopts an estimate of the unknown set of parameters , i.e. , which gives the maximum of the expected likelihood of the available information [29], . The likelihood of the set of parameters is defined by (1) where meaning that the equipment is in its best possible state at period zero. After obtaining an indicator value via an inspection at an observation moment, the prior conditional probability is updated to . By using Bayes’ formula, and knowing that the indicator has occurred at the -st observation moment, is determined as The method of MLE estimates by finding the value of that maximizes the likelihood of the set of parameters. We consider a parametric PH Model with a baseline two parameter Weibull hazard function, which is also known as the Weibull parametric regression model [41], and is given as (2) (4) B. Parameter Estimation To construct the parameter estimation model, we assume that , , is the history of the indicator values up to time . We map this And then [29], GHASEMI AND HODKIEWICZ: ESTIMATING MEAN RESIDUAL LIFE FOR A CASE STUDY OF RAIL WAGON BEARINGS C. Data Requirements for Parameter Estimations Suitable data for the parameter estimation process of a prognostics model is derived from two sources: (a) Failure related data are based on the working age at which the equipment has failed. This information comes from installation and failure dates. (b) Performance data are based on the condition indicators that relate to the performance condition of the equipment from its installation date up to its failure date obtained in fixed or near fixed observation intervals. In this research, the number of pass-bys is considered as the working age of a bearing. The railway under study is mainly between one major mine and one port station. Although, there are other sub railways in the railway network, the data recorded from wayside monitoring systems shows that more than 80% of the wagons’ movement is back-and-forth between these two points in the railway, and do not change their service path frequently. In this study, as it is explained later in the data cleansing section, we exclude the failure data of bearings with a long period of absence on the main railway. By applying this approach, the bearings that have been used elsewhere in the network (other than the main railway) for a long period, or been removed from one wagon and not replaced immediately on another wagon, are not taken into account. By considering the pass-by as the working age of the equipment, and by excluding bearings with a long period of missing acoustic data, we address two issues in the modeling. First, the assumption of fixed or near fixed observation intervals is satisfied. Secondly, short idle periods (when a wagon may be parked up, or in the workshop) during a bearing’s lifetime are not considered in their aging process. Failure Related Data: To develop the prognostics model, the following specific failure and historical data are necessary. • Exact location of the bearing during its lifetime (wagon number, wheel set ID, left or right bearing on wheel set) • Installation date • Removal date • Failure or suspension • Bearing type (new, overhauled or second-hand) • Failure codes There are 50 000 rail car wheel bearing assemblies in service. Within this population, there are three distinct bearing categories: new, overhauled, and second-hand. Bearings are recalled from service due to 1) hazardous condition monitoring values, or 2) a routine 2 year service. Part of this recall process is a hand-qualification and inspection test by the tradesman; following this check, the bearing is sent to an external shop for overhauling, or returned to service. The former is described here as “overhauled,” and the latter as “second-hand.” To collect failure information, bearing failure has to be clearly defined. We have defined failure as follows. A bearing is failed if it was removed due to: • an acoustic alarm (roller, cup, cone), • a hot bearing detected alarm, • hand verification, or • a catastrophic failure. According to current practice guidelines, only the wagons creating acoustic alarms are recalled for further investigation. 723 But roller, cup, and cone have much higher weight in the decision making process. After a recall, all bearings on the recalled wagon are examined manually for defects, and replaced as required. We have included all the recalls initiated by acoustic alarms of roller, cup, and cone as failures. However, in practice, some of the recalls identified by the acoustic alarms are not actually failures. To better identify these alarms in the future, we discuss ways to improve data collection to better distinguish between failures and suspensions in the recommendations section. Failure data were collected on failures on a specific, single series of bearings over a three month period. This produced an initial 328 records of failure. Performance Data: The bearing acoustic monitoring system (RailBAM) records acoustic data of bearings using arrays of microphones inside the trackside cabinet assemblies, as well as wheel sensors on the track. A process of vehicle and wheel-set tagging allows the system to attribute data to each individual bearing. The data are analysed by signal processing techniques to extract acoustic signatures for each individual bearing. This acoustic signature is then used to calculate the acoustic energy level for several classes of bearing defects: roller, cup, cone, looseness and fretting, noisy, clipped or shrieking wheels, and wheel flats [36]. The energy levels are then used to produce alarm scores using a preset threshold. An alarm is triggered if a measured acoustic energy level by RailBAM is higher than a preset threshold. In this work, the energy levels are used as condition monitoring indicators (covariates), not the alarms. A description of the development of the RailBAM system is provided in [36]. In the current practice, a heuristic method is used to calculate an alarm score by giving a severity weight (1 to 7) to different possible alarms, and summing the weights over the last 40 passes of the bearing. Bearings with high score are prioritized in a recall list, and removed from wagon. A wayside Hot Bearing Detection (HBD) system also sends alarms to the operator of the train if a bearings temperature is higher than a pre-set threshold. The operator then stops the train immediately, and investigates the situation. We have only used these alarms to identify the historical failures, and they are not used as performance data in the proposed model. A suggestion for inclusion of this information in the PH Model is discussed later in the discussion section. A PH-Model models the degradation due to two sources: the aging process, provided by failure data, and represented by a ; and a diagnostic process represented baseline failure rate that by a vector of -independent variables or covariates reflects the conditions of utilization of the system [20]. The output of the first phase is used to estimate the parameters of the model, and to test its efficiency. Once the model is constructed, current performance data of the asset are fed to the model to calculate the MRL of the asset, and make the maintenance intervention or replacement decision. D. Modeling the Residual Life The two most used measures of equipment health are the hazard function, and the MRL. These two measures are calculated from the reliability function. 724 IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 In reliability analysis, two reliability functions are of interest. The first is the unconditional reliability function given by the probability , which is the probability that the failure time , of a piece of equipment that has not yet been put into operation, is bigger than a certain time . The second is the conditional reliability function calculated by , which is the probability that the time to failure is bigger than , knowing that the equipment has already survived until time , where . In this later case, the MRL is [42], which is equal to where has distribution function . The hazard function is obtained from . In this equation, is a small time interval less than the time to next inspection. In some reliability analysis, it is assumed that every piece of equipment is used in the same environment, and under the same conditions. This assumption allows the calculation of the MRL, and the hazard function prior to the actual use of the equipment. In real-life, the environments in which the equipment is performing, and the conditions of utilization affect the process of degradation. Consequently, the conditional reliability, the residual life, and the failure rate of the equipment are affected. Taking this fact into consideration improves the diagnosis of the equipment’s degradation state, and the prognosis for future performance. Many researchers have proposed different reliability models incorporating the information gathered periodically regarding the equipment’s condition. These models are used to calculate an adjusted hazard function, and the corresponding MRL. In this study, the PH Model with imperfect information is used. In the initial PH Model with perfect information, the conditional reliability is given by [43] (5) The conditional reliability indicates the probability of survival until time , , knowing that the failure has not happened until time , and the states of the equipment have been , at . is the random variable indicating the time to failure. Equation (5) is not valid for because may change at any subsequent interval. The conditional reliability for the case of perfect information for all is [39] (6) Fig. 2. Approximate removed records during data cleansing. In the case of imperfect information, the conditional reliability can be calculated by [40] (7) , calculated at the th obAnd consequently, the MRL, servation moment, while the state conditional probability is , is calculated by (8) The steps for calculating the MRL at each observation moment , where the indicator obtained is , are as follows. At any observation moment , when an indicator value is obtained, follow this procedure. (a) Calculate the conditional probability distribution , at period using (1). (b) Calculate the conditional reliability of the equipment, , at the -th observation moment by using (7). (c) Calculate the MRL of the equipment, , by applying (8). IV. RESULTS A. Performance Data Acquisition, and Cleansing Outcomes The original data set contained 328 failure records. However, in the cleansing process, a large number of records had to be removed due to numerous reasons: • unknown age of the failed bearing, • unknown type of the bearing at installation moment, • absence of historical acoustic data for a long time period (lost data or indication of wheels being used in sub sections of railway network), and • inconsistency of dates in the databases. GHASEMI AND HODKIEWICZ: ESTIMATING MEAN RESIDUAL LIFE FOR A CASE STUDY OF RAIL WAGON BEARINGS 725 Fig. 3. TABLE I AVERAGE LIFE OF DIFFERENT BEARING TYPE TABLE II MODEL FITNESS WITHOUT CM INDICATORS WITH 2 PARAMETER WEIBULL MODEL Fig. 2 shows the approximate number of records removed due to these reasons. The outcome of the data acquisition and cleansing stage is a set of 63 records of failure: 57 for modeling, and the six remaining records (almost 10%) for testing the model result. The six records are selected by choosing equipment numbers ending with 1 to maintain a random selection. These were removed from the data set used to develop the model parameters. The distribution of bearing types for the 57 records is shown in Table I. The estimation of mean life for the bearing data set is shown in Table II. This table shows that the mean life of the bearings is 343 pass-bys. TABLE III MODEL FITNESS WITH ACOUSTIC BEARING INDICATORS AND BEARING TYPE B. Parameter Estimation Results Table III shows the result of the parameter estimation method applied to the collected data. The current heuristic practice used by the practitioners uses only the final alarm of the system to make a decision. The alarm is initiated if one of the energy levels is higher than a pre-set threshold level. In this study, we use the energy level obtained from the acoustic condition monitoring system. These energy levels are condition monitoring indicators (covariates) used for the PH-Model. Based on statistical tests, cup, cone, roller, and noisy energy levels are not significant in the PH Model. However, looseness, and bearing type 726 IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 TABLE IV 15 DAY ALARM RECALL POLICY ANALYSIS FOR SELECTED CONFIDENCE INTERVALS that have been advised to be replaced could survive on average 11 days more than predicted. V. DISCUSSION In this part, we explore the significance of the results in detail in each of the following five sub sections. A. Outcomes of the Parameter Estimation are of great meaningful significance in the model. Bearing type is almost ten times more significant than looseness. C. Testing a 15 Pass-By Replacement Policy Although the current heuristic replacement policy considers alarms levels obtained from RailBAM system for cone, cup, and roller in decision making, and practically ignores the bearing type, and looseness alarms (and consequently energy level), while latter pieces of information carry much more meaningful data related to the bearing failure. A replacement policy is criterion-based. For example, we decide whether to replace a bearing, or leave it until the collection of the next indicators. In this research, we have introduced and tested following replacement policy: “Replace the bearing if the lower confidence interval of the MRL is less than 15 pass-bys.” We found the best value of x that suits this replacement objective. This criterion has been tested for confidence intervals of 68.3%, 86.4%, and 89.3%. The results are summarized in Table IV. A 68.3% CI demonstrates 50% True Negative (TN), and 50% False Negative (FN) results. The FN results are 31 pass-bys longer than real failure pass-by of the bearings. This result means that, in 50% of the cases, the policy has estimated a longer residual life for the bearings. On average, the estimation has been 31 pass-bys longer than the real residual pass-bys. At 89.3% CI, the results are satisfactory. Out of 12 simulated decisions, 50% have been correctly advised to replace before 15 pass-bys (True Positive). 17% of the cases did not need to be replaced, and are correctly advised not to be replaced before next 15 pass-bys (True Negative). The remaining 33% are the cases in which the policy advises to replace the bearing before the next 15 pass-bys, while in reality the bearing could survive more than 15 pass-bys. However, in the latter case, the average difference of the estimated remaining pass-bys and the actual failure time is 11 pass-bys. This means that the 33% of the cases Table III shows the result of the parameter estimation method applied on collected data. Based on statistical tests, cup, cone, roller, and noisy energy bonds are not significant in the model. However, looseness, and bearing type have meaningful significance in the model. Insignificant behavior of these bearing acoustic indicators is contrary to expectations, and may be due to one or a combination of the following reasons. • Cone, roller, and cup were the indicators with the most missing data, having values of zero in the field. These missing data were replaced with the moving average. • The looseness energy level correlates with some of the other indicators, and is reflecting their effect. • There is much more significant information in the looseness indicator than imagined, which is not being used currently. • There is an insufficient number of failure data (samples). • There is error in data (as discussed in the following section). The significance of bearing type in the model is an important finding, and indicates that the maintenance staff should consider the consequences of selecting a second-hand bearing over a new or overhauled bearing. The rest of the indicators (cone, roller, cup, noisy) do not have significant effect on the hazard (failure) rate. This result can be explained through the indicator-state (health condition) matrix Q, used in the modeling stage. The structure of the matrix Q is shown in Table V(a). Element , for example, is the probability of seeing indicator level 2 when the bearing is in health condition 2 (example: 0.10 in Table V(b)). Table V(b) and (c) show two possible examples of this matrix. As can be seen in (b), any possible health condition can have diverse probabilities for different levels of the indicator. However, in case (c), the probabilities in each line are more or less equal. If the estimated values of matrix Q show statistically equal values in each line, then the indicator is assumed to be insignificant in the model. B. Sample Population Characteristics Table I summarizes the mean working age (pass-by) of bearings grouped by different bearing types. The last column is an approximation of each bearing type percentage used in the fleet. This approximation was calculated using notification entries in the maintenance management system, where the type of bearing installed each time is recorded. The samples for all three bearing types have large standard deviations. The large standard deviation of two first types (new and overhauled) can be explained by the small sample sizes available (nine and five). However, second-hand bearings with 43 samples also show a large standard deviation value. The authors postulate that this result is mostly due to the diversity of second-hand bearings. In this case study, there is no record of GHASEMI AND HODKIEWICZ: ESTIMATING MEAN RESIDUAL LIFE FOR A CASE STUDY OF RAIL WAGON BEARINGS TABLE V INDICATOR OR STATE MATRIX STRUCTURE, AND EXAMPLES the times a bearing has been used, removed, and reinstalled as a second-hand bearing. Also, it may have been on other wheels as little as few weeks to several months. It also can be a second hand bearing of an originally new or overhauled bearing. Improvements in data collection, discussed in the recommendation section, should assist in accounting for this fact in future models. C. Significance of the MRL Model Results MRL is one of the major tools used by maintenance practitioners in environments where cost is not the major issue. In these systems, usually the availability, reliability, or the risk are of interest. In this research, we have been able to calculate the MRL of the bearing based on available condition monitoring data. By comparing the MRL to the next planned inspection of the equipment, while taking into account the CI, one will be able to decide whether to recall the equipment for further investigation before the next planed inspection, or leave it to work until the next updated MRL calculation result. Table VI summarizes the result of the MRL calculation for the six test failure records. However, the standard deviations are relatively large, which results in the cumulative distribution function (CDF) curve having a large spread. This result is also due to one or a combination of issues: 1 a lack of sufficient records of failure; 2 an imprecise definition of second-hand bearings, and a lack of information on their history; or 3 a combination of different types of failures into just one failure type. Fig. 4 shows more detailed steps of the data acquisition and modeling phase, and also demonstrates the processes performed by the research group during this project. A key part of this project involves developing and understanding the time required to complete different stages. The most time consuming stage, 40–50% of the total project time, was preparing data. This preparation includes data sorting, 727 TABLE VI CALCULATION OF MRL FOR TEST DATA cleansing, matching, and validation. This work is a key consideration in the development of data-based models. To improve the efficiency of model development and testing, the following steps would help with future models. • Data should ideally be recorded in the maintenance management system for individual bearings. Bearings carry a unique serial number that should be possible to track in the system. However, there can be practical challenges. Ideally, the CM system records, which are collected for both sides of each wheel-set should be matched with the maintenance records of bearings. • Improvement in the tag reading hardware and software will facilitate the data acquisition and cleansing part of the project. • Improve the consistency of data in the databases. There were many instances of inconsistent data due to data entry error, tag code change, and different coding systems in different databases. Improve these point to increase the amount of reliable data available, and reduce the time to cleans the data. D. Recommendations and Future Work Recommendations and future work can be discussed in short, and long terms. We suggest improving and applying this model by considering the following tasks. • Extend the data set. Currently the model is constructed based on 57 failures, and their related performance records. A first step to improve the output of the model is to extend the data to a larger set of failure data. At the same time, better cleansing techniques and assumptions should be applied to the data. This approach will assure the integrity and correctness of the data fed to the model. • Extend the use of the acoustic energy levels available for modeling. In this work, we used 10 (five for each bearing side) acoustic energy levels to feed the model. Ideally, one may consider using some or all 124 fields of acoustic signature produced by the acoustic monitoring system to construct the model. 728 IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 Fig. 4. • Extend the use of other degradation indicators such as the HBD, and the Wheel Impact Detector (WID) systems from the data acquisition, data cleansing, and modeling phases in this project. One of the ongoing challenges for equipment prognostics is the availability of clean, appropriate data. As mentioned earlier, 40–50% of the time spent on this project was on cleaning data, on the order of months. Although 328 records were originally identified, only 63 records were complete and appropriate for the final analysis. The cleansing process required access to multiple systems, and was labor intensive. For this work to be done cost-effectively, more thought will need to be given to what data are collected, how they are stored, and how they are coded across multiple systems. The ability to align data between systems, for example the computerized maintenance management system, and the condition monitoring system(s), continues to be a challenge, and stands as a barrier to routine use and updating of prognostic models. VI. CONCLUSIONS We have presented a model for estimating the mean residual life of wagon bearings to support maintenance interventions. Data for the model have been collected from a rail industry case study. Data collected include acoustic signatures of the bearing GHASEMI AND HODKIEWICZ: ESTIMATING MEAN RESIDUAL LIFE FOR A CASE STUDY OF RAIL WAGON BEARINGS at each pass-by. Acoustic data in this study are inherently imperfect due to the noisy environment, and the nature of the data monitoring technique. A prognostics model based on selected bearing acoustic data levels, bearing type, and its age was developed. This work provides an ability to detect failure in a 15 pass-bys window with acceptable levels of -confidence. Significant indicators that can be effectively used in the calculation of MRL were distinguished, and the MRL model was constructed based on these indicators. This case study has shown that a) bearing type is of high significant impact in the MRL, and b) some indicators like looseness that have not been considered previously can carry very indicative information regarding the bearing MRL. We have also explored the frustrations and time consumption of the data acquisition and cleansing process, and have suggested ways to address these problems. The positive impact of improving the data acquisition and cleansing process on getting reliable MRL result was discussed. Future work includes testing the significance of other condition monitoring indicators (like Hot Wheel Detector and Wheel Impact Detector), and also a higher level of information from the existing acoustic system data. ACKNOWLEDGMENT The authors would like to thank the staff at Rio Tinto Technology & Innovation and Rio Tinto Railways Division who supplied data, information, time, and expertise to this project. REFERENCES [1] J. S. Sikorska, M. R. Hodkiewicz, and L. Ma, “Prognostic Modelling options for remaining useful life estimation by industry,” Mechanical Systems and Signal Processing, vol. 25, no. 5, pp. 1803–1836, 2011. [2] M. Carnero, “An evaluation system of the setting up of predictive maintenance programmes,” Reliability Engineering and System Safety, vol. 91, no. 8, pp. 945–963, 2006. [3] M. R. Hodkiewicz, P. Kelly, J. Sikorska, and L. Gouws, “A framework to assess data quality for reliability variables,” in Proceedings of the World Congress of Engineering Asset Management, Gold Coast, Australia. [4] H. A. Sandtorv, P. Hokstad, and D. W. Thompson, “Practical experiences with a data collection project: The OREDA project,” Reliability Engineering and System Safety, vol. 51, pp. 159–167, 1996. [5] K. Unsworth, E. Adriasola, A. Johnston-Billings, A. Dmitrieva, and M. Hodkiewicz, “Goal hierarchy: Improving asset data quality by improving motivation,” Reliability Engineering and System Safety, vol. 96, pp. 1474–1481, 2011. [6] A. K. S. Jardine, D. Lin, and D. Banjevic, “A review on machinery diagnostics and prognostics implementing condition-based maintenance,” Mechanical Systems and Signal Processing, vol. 20, pp. 1483–1510, 2006. [7] A. Heng, S. Zhang, A. C. C. Tan, and J. Mathew, “Rotating machinery prognostics: state of the art, challenges and opportunities,” Mechanical Systems and Signal Processing, vol. 23, pp. 724–739, 2009. [8] R. Khotamasu, S. H. Huang, and W. H. Ver Dui, “System health monitoring and prognostics—A review of current paradigms and practices,” International Journal of Advanced Manufacturing Technology, vol. 28, pp. 1012–1024, 1996. [9] L. Zhiguo and G. Kott, “Predicting remaining useful life based on the failure time data with heavy-tailed behavior and user usage patterns using proportional hazards model,” in Ninth International Conference on Machine Learning and Application 623–628, 2010. [10] S. Park, C. L. Choi, J. H. Kim, and C. H. Bae, “Evaluating the economic residual life of water pipes using the proportional hazards model,” Water Resources Management, vol. 24, no. 12, pp. 3195–3217, 2010. 729 [11] X. Gao, J. Barabady, and T. Markeset, “An approach for prediction of petroleum production facility performance considering arctic influence factors,” Reliability Engineering & System Safety, vol. 95, no. 8, pp. 837–846, 2010. [12] W. Caesarendra, A. Widodo, and B. S. Yang, “Application of relevance vector machine and logistic regression for machine degradation assessment,” Mechanical Systems and Signal Processing, vol. 24, pp. 1161–1171, 2010. [13] E. A. Elsayed and C. K. Chan, “Estimation of thin-oxide reliability using proportional hazards models,” IEEE Trans. Reliability, vol. 39, no. 3, pp. 329–335, 1990. [14] R. O. Ansell and J. I. Ansell, “Modelling the reliability of sodium sulphur cells,” Reliability Engineering, vol. 17, no. 2, pp. 127–137, 1987. [15] G. Niu and B. S. Yang, “Dempster-Shafer regression for multistep-ahead time-series prediction towards data-driven machinery prognosis,” Mechanical Systems and Signal Processing, vol. 23, pp. 740–751, 2009. [16] V. T. Tran, B. S. Yang, M. S. Oh, and A. C. C. Tan, “Machine condition prognosis based on regression tress and one-step-ahead prediction,” Mechanical Systems and Signal Processing, vol. 22, pp. 1179–1193, 2008. [17] V. T. Tran, B. S. Yang, and A. C. C. Tan, “Multi-step ahead direct prediction for the machine condition prognosis using regression trees and neuro-fuzzy systems,” Expert Systems with Applications, vol. 36, pp. 9378–9387, 2009. [18] W. Wang and A. H. Christer, “Towards a general condition based maintenance model for a stochastic dynamic system,” Journal of the Operational Research Society, vol. 51, pp. 145–155, 2000. [19] V. Makis and X. Jiang, “Optimal replacement under partial observations,” Mathematics of Operations Research, vol. 28, pp. 382–394, 2003. [20] A. Ghasemi, S. Yacout, and M. S. Ouali, “Optimal condition based maintenance with imperfect information and the proportional hazards model,” International Journal of Production Research, vol. 45, pp. 989–1112, 2007. [21] E. L. Kaplan and P. Meier, “Nonparametric estimation from incomplete observations,” Journal of the American Statistical Association, vol. 53, pp. 457–481, 1958. [22] D. R. Cox, “Regression models and life-tables (with discussion),” Journal of the Royal Statistical Society, vol. 34, no. 2, pp. 187–220, 1972. [23] I. K. Omurlu, K. Ozdamar, and M. Ture, “Comparison of Bayesian survival analysis and Cox regression analysis in simulated and breast cancer data sets,” Expert Systems with Applications, vol. 36, pp. 11341–11346, 2009. [24] C. J. Lu and W. Q. Meeker, “Using degradation measures to estimate a time-to-failure distribution,” Technometrics, vol. 35, no. 2, pp. 161–173, 1993. [25] J. Hilden, J. Dik, and F. Habbema, “Prognosis in medicine: an analysis of its meaning and roles,” Theoretical Medicine and Bioethics, vol. 8, pp. 349–365, 1987. [26] P. J. F. Lucas and A. Abu-Hanna, “Prognostic methods in medicine,” Artificial Intelligence in Medicine, vol. 15, pp. 105–119, 1999. [27] V. Makis, J. Wu, and Y. Gao, “An application of DPCA to oil data for CBM modeling,” European Journal of Operational Research, vol. 174, no. 1, pp. 112–123, 2006. [28] S. Gasmi, C. E. Love, and W. Kahle, “A general repair, proportional-hazards, framework to model complex repairable systems,” IEEE Trans. Reliability, vol. 52, no. 1, pp. 26–32, 2003. [29] A. Ghasemi, S. Yacout, and M. S. Ouali, “Parameter estimation methods for condition based maintenance with indirect observations,” IEEE Trans. Reliability, vol. 59, no. 2, pp. 426–439, 2010. [30] V. Makis and A. K. S. Jardine, “Optimal replacement in the proportional hazards model,” INFOR, vol. 30, pp. 172–183, 1991. [31] D. Banjevic and A. K. S. Jardine, “Calculation of reliability function and remaining useful life for a Markov failure time process,” IMA Journal of Management Mathematics, vol. 17, p. 115−130, 2006. [32] P. Vlok, J. L. Coetzee, D. Banjevic, A. K. S. Jardine, and V. Makis, “Optimal component replacement decisions using vibration analysis and the proportional hazards model,” Journal of the Operational Research Society, vol. 53, pp. 193–202, 2002. 730 IEEE TRANSACTIONS ON RELIABILITY, VOL. 61, NO. 3, SEPTEMBER 2012 [33] R. Jiang and A. K. S. jardine, “Health state evaluation of an item: A general framework and graphical representation,” Reliability Engineering and System Safety, vol. 93, pp. 89–99, 2008. [34] G. B. Anderson and R. S. McWilliams, “Vehicle health monitoring system development and deployment,” ASME International Mechanical Engineering Congress, Nov. 15–21, 2003. [35] G. B. Anderson, “Acoustic detection of distressed freight car roller bearings,” in Proceedings of the ASME/IEEE Joint Rail Conference and the ASME Internal Combustion Engine Division, Mar. 13–17, 2007. [36] O. D. Snell, “Acoustic bearing monitoring—The future,” in RCM 2008, IEEE International Conference on Railway Condition Monitoring, June 1–5, 2008, pp. 18–20. [37] W. Kirchner, S. Southward, and M. Ahmadian, “Ultrasonic acoustic health monitoring of ball bearings using neural network pattern classification of power spectral density,” in Proceedings of the ASME Joint Rail Conference, 2010, vol. 2, pp. 255–265. [38] J. L. A. Ferreira, J. C. Balthazar, and A. P. N. Araujo, “An investigation of rail bearing reliability under real conditions of use,” Engineering Failure Analysis, vol. 310, no. 6, pp. 745–758, 2003. [39] A. Ghasemi, S. Yacout, and M. S. Ouali, “Optimal strategies for non-costly and costly observations in condition based maintenance,” IAENG International Journal of Applied Mathematics, vol. 38, no. 2, pp. 99–107, 2008. [40] A. Ghasemi, S. Yacout, and M. S. Ouali, “Evaluating the reliability function and the mean residual life for equipment with unobservable states,” IEEE Trans. Reliability, vol. 59, no. 1, pp. 45–54, 2010. [41] D. Banjevic, A. K. S. Jardine, V. Makis, and M. Ennis, “A control-limit policy and software for condition-based maintenance optimization,” INFOR Journal, vol. 39, pp. 32–50, 2001. [42] A. K. S. Jardine, D. Banjevic, M. Wiseman, S. Buck, and T. Joseph, “Optimizing a mine haul truck wheel motors’ condition monitoring program,” Journal of Quality in Maintenance Engineering, vol. 7, pp. 286–301, 2001. [43] V. Makis and A. K. S. Jardine, “Computation of optimal policies in replacement models,” IMA Journal of Mathematics Applied in Business & Industry, vol. 3, pp. 169–175, 1992. Alireza Ghasemi is an Assistant Professor at the Department of Industrial Engineering of Dalhousie University, Halifax, Canada. He holds a Ph.D. in Industrial Engineering from École Polytechnique de Montréal in Canada. He also holds a M.Sc. degree from École Polytechnique de Montréal, a M.Sc. Degree from Sharif University of Technology in Iran, and a B.Sc degree from Isfahan University of Technology in Iran, all in Industrial Engineering. His research field is Condition Based Maintenance with Imperfect Observations. He is a member of the Institute of Industrial Engineers, and a member of the Canadian Operational Research Society. Melinda R. Hodkiewicz is a Professor in the School of Mechanical and Chemical Engineering at the University of Western Australia (UWA) in Perth. She holds a Ph.D. in Mechanical Engineering from UWA, and a BA (Honors) in Metallurgy and Materials Science from Oxford University. She is responsible for teaching and research programs in asset management at UWA, and sits on the ISO committee PC251 for asset management.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )