Reliability Engineer Study Notes
Week 1: Reliability Concepts
� Day 1: Introduction to Reliability and Availability
Definition
Reliability:
The probability that a system or component performs its required functions under stated
conditions for a specified period of time without failure.
Availability:
The proportion of time a system is in a functioning condition compared to the total planned
operational time.
Formula
𝑀𝑇𝐵𝐹
Availability=𝑀𝑇𝐵𝐹+𝑀𝑇𝑇𝑅
Formula Explanation
MTBF (Mean Time Between Failures) represents the average time between system failures.
MTTR (Mean Time To Repair) represents the average time needed to repair the system after
a failure.
A high MTBF and a low MTTR yield a high availability.
Practical Example
Given:
MTBF = 400 hours
MTTR = 20 hours
Calculation:
400
Availability=400+20=0.952 or 95.2%
Daily Key Takeaway
Maximizing asset availability requires increasing MTBF (reliability) and reducing MTTR
(maintainability).
References
Gulati, R. (2013). Maintenance and Reliability Best Practices (2nd ed.). Industrial Press.
Moubray, J. (1997). RCM II: Reliability-Centered Maintenance (2nd ed.). ButterworthHeinemann.
IEC 60050-191: Dependability and Quality of Service Vocabulary.
� Day 2: Mean Time Between Failures (MTBF)
Definition
MTBF is the average operational time between two consecutive failures of a repairable system.
Formula
𝑇𝑜𝑡𝑎𝑙 𝑈𝑝𝑡𝑖𝑚𝑒
MTBF=𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝐹𝑎𝑖𝑙𝑢𝑟𝑒𝑠
Formula Explanation
Measures the time between inherent failures of a system during operation.
A higher MTBF indicates better system reliability.
Practical Example
A machine operated for 2000 hours and failed 5 times:
2000
MTBF= 5 =400 hours
Daily Key Takeaway
Monitoring MTBF helps in scheduling maintenance and predicting asset lifespan.
References
Mobley, R. K. (2002). Maintenance Fundamentals (2nd ed.). Butterworth-Heinemann.
SMRP. (2021). Body of Knowledge and Best Practices (6th ed.).
� Day 3: Mean Time To Repair (MTTR)
Definition
MTTR is the average time required to repair a system or component and restore it to operational
condition after a failure.
Formula
𝑇𝑜𝑡𝑎𝑙 𝐷𝑜𝑤𝑛𝑡𝑖𝑚𝑒
MTTR=𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝐹𝑎𝑖𝑙𝑢𝑟𝑒𝑠
Formula Explanation
Indicates how quickly equipment can be restored.
Lower MTTR means better maintainability.
Practical Example
A total downtime of 15 hours over 3 failures:
15
MTTR= =5 hours
3
Daily Key Takeaway
Reducing MTTR improves system availability and reduces production loss.
References
Mobley, R. K. (1999). Root Cause Failure Analysis. Butterworth-Heinemann.
Palmer, R. D. (1999). The Maintenance Management Framework. McGraw-Hill.
� Day 4: Failure Rate (λ) and Reliability Function (R(t))
Definition
Failure Rate (λ):
The number of failures per unit of time.
Reliability Function (R(t)):
The probability that a system will perform without failure up to a certain time (t).
Formulas
1
λ=𝑀𝑇𝐵𝐹
R(t)=𝑒 −𝜆𝑡
Formula Explanation
λ gives the expected frequency of failures.
R(t) measures the survival probability over time.
Practical Example
MTBF = 500 hours:
1
λ=500=0.002 failures/hour
Probability of surviving 100 hours:
R(100)=𝑒 −0.002×100 =𝑒 −0.2≈0.8187 or 81.87%
Daily Key Takeaway
Understanding failure rate and reliability over time helps plan maintenance strategies and spares
stocking.
References
Ebeling, C. E. (1997). An Introduction to Reliability and Maintainability Engineering.
Waveland Press.
MIL-HDBK-338B: Electronic Reliability Design Handbook.
� Day 5: Asset Hierarchy and System Definition
Definition
Asset:
A resource owned or controlled by an organization that delivers value through operational
functionality.
Asset Hierarchy:
Structured breakdown of physical assets into logical levels:
Plant → System → Equipment → Component.
Practical Example
Plant: Oil Refinery
System: Cooling Water System
Equipment: Centrifugal Pump
Component: Pump Shaft
Daily Key Takeaway
Clear asset hierarchies enable effective planning, analysis, and management of maintenance
activities.
References
ISO 14224: Collection and Exchange of Reliability and Maintenance Data.
Palmer, R. D. (1999). The Maintenance Management Framework.