Performance-Telemetry-Based Anomaly Detection for Early Failure Identification in Enterprise Storage Systems
Keywords:
performance telemetry, failure prediction, enterprise storage, anomaly detectionAbstract
Background: Enterprise storage failures rarely emerge as a single binary event. They are often preceded by weak, heterogeneous changes in device-health counters, latency, error rates, queueing behavior, retry activity, media condition, and workload-dependent performance. Static thresholds can identify only a subset of these precursors and frequently fail to generalize across drive models, firmware generations, and deployment environments.
Objective: To synthesize evidence on telemetry-driven anomaly detection and failure prediction in large-scale storage systems and to develop an operational framework for early failure identification using device-health, performance, temporal, and contextual signals.
Methods: A structured evidence synthesis was conducted across peer-reviewed studies of hard-disk-drive (HDD) and solid-state-drive (SSD) reliability, Self-Monitoring, Analysis and Reporting Technology (SMART)-based prediction, anomaly detection, temporal deep learning, feature selection, and large-scale field failure characterization. Studies were prioritized when they used production or real-world storage populations, evaluated low-prevalence failure prediction, or reported operationally relevant metrics such as failure detection rate, false-alarm rate, precision, recall, F1 score, lead time, or generalization across device models.
Results: Field studies show that real storage failures are heterogeneous, temporally correlated, and only partially explained by isolated health attributes. Predictive performance improves when telemetry is treated as a multivariate and temporal signal rather than as independent threshold crossings. Decision-tree, gradient-boosted, recurrent, attention-based, multi-view, semi-supervised, and cost-sensitive methods all demonstrate advantages under particular data conditions. Large production studies further show that device model, wear state, physical location, workload, and correlated node/rack failures materially affect interpretation. The main barriers to enterprise deployment are severe class imbalance, model drift, heterogeneous telemetry semantics, inconsistent labels, false-alarm cost, and the difference between high classification accuracy and actionable early warning.
Conclusion: Performance-telemetry-based anomaly detection is most effective when implemented as a multi-scale, context-aware, closed-loop reliability service rather than a stand-alone classifier. Enterprise systems should combine device-health indicators with performance drift, temporal persistence, topology and workload context, and explicit operational cost functions. Evaluation should emphasize precision-recall behavior, false alarms per device-time, warning lead time, and the proportion of alerts that enable preventive migration or replacement before service impact.
References
Schroeder B, Gibson GA. Disk failures in the real world: what does an MTTF of 1,000,000 hours mean to you? In: 5th USENIX Conference on File and Storage Technologies (FAST 07). Berkeley (CA): USENIX Association; 2007. p. 1-16.
Pinheiro E, Weber WD, Barroso LA. Failure trends in a large disk drive population. In: 5th USENIX Conference on File and Storage Technologies (FAST 07). Berkeley (CA): USENIX Association; 2007. p. 17-29.
Hughes GF, Murray JF, Kreutz-Delgado K, Elkan C. Improved disk-drive failure warnings. IEEE Trans Reliab. 2002;51(3):350-357. doi:10.1109/TR.2002.802886.
Murray JF, Hughes GF, Kreutz-Delgado K. Machine learning methods for predicting failures in hard drives: a multiple-instance application. J Mach Learn Res. 2005;6:783-816.
Xu C, Wang G, Liu X, Guo D, Liu TY. Health status assessment and failure prediction for hard drives with recurrent neural networks. IEEE Trans Comput. 2016;65(11):3502-3508. doi:10.1109/TC.2016.2538237.
Li J, Stones RJ, Wang G, Liu X, Li Z, Xu M. Hard drive failure prediction using decision trees. Reliab Eng Syst Saf. 2017;164:55-65. doi:10.1016/j.ress.2017.03.004.
Wang G, Zhang L, Xu W. What can we learn from four years of data center hardware failures? In: 47th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). Piscataway (NJ): IEEE; 2017. p. 25-36. doi:10.1109/DSN.2017.26.
Schroeder B, Lagisetty R, Merchant A. Flash reliability in production: the expected and the unexpected. In: 14th USENIX Conference on File and Storage Technologies (FAST 16). Santa Clara (CA): USENIX Association; 2016.
Han S, Lee PPC, Xu F, Liu Y, He C, Liu J. An in-depth study of correlated failures in production SSD-based data centers. In: 19th USENIX Conference on File and Storage Technologies (FAST 21). Berkeley (CA): USENIX Association; 2021. p. 417-429.
Yang Q, Jia X, Li X, Feng J, Li W, Lee J. Evaluating feature selection and anomaly detection methods of hard drive failure prediction. IEEE Trans Reliab. 2021;70(2):749-760. doi:10.1109/TR.2020.2995724.
Lu S, Luo B, Patel T, Yao Y, Tiwari D, Shi W. Making disk failure predictions SMARTer! In: 18th USENIX Conference on File and Storage Technologies (FAST 20). Berkeley (CA): USENIX Association; 2020. p. 151-167.
Xu F, Han S, Lee PPC, Liu Y, He C, Liu J. General feature selection for failure prediction in large-scale SSD deployment. In: 51st Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN). Piscataway (NJ): IEEE; 2021. p. 263-270. doi:10.1109/DSN48987.2021.00039.
Wang G, Wang Y, Sun X. Multi-instance deep learning based on attention mechanism for failure prediction of unlabeled hard disk drives. IEEE Trans Instrum Meas. 2021;70:1-9. doi:10.1109/TIM.2021.3068180.
Zhang X, Shan K, Tan Z, Feng D. CSLE: a cost-sensitive learning engine for disk failure prediction in large data centers. In: Design, Automation & Test in Europe Conference & Exhibition (DATE). Leuven: EDAA; 2022.
Zhang Y, Hao W, Niu B, Liu K, Wang S, Liu N, et al. Multi-view feature-based SSD failure prediction: what, when, and why. In: 21st USENIX Conference on File and Storage Technologies (FAST 23). Berkeley (CA): USENIX Association; 2023. p. 409-424.
Liu Y, Guan Y, Jiang T, Zhou K, Wang H, Hu G, et al. SPAE: lifelong disk failure prediction via end-to-end GAN-based anomaly detection with ensemble update. Future Gener Comput Syst. 2023;148:460-471. doi:10.1016/j.future.2023.05.020.
Bai X, Pan Z, Meng G, Wang S, Fu Y. Disk failure prediction based on association analysis and SSA-LSTM. J Intell Fuzzy Syst. 2023;45(4):5633-5645. doi:10.3233/JIFS-231268.
Ahmed J, Green RC. Cost aware LSTM model for predicting hard disk drive failures based on extremely imbalanced S.M.A.R.T. sensors data. Eng Appl Artif Intell. 2024;127:107339. doi:10.1016/j.engappai.2023.107339.
Gu Y, Wu C, He X. Exploit both SMART attributes and NAND flash wear characteristics to effectively forecast SSD-based storage failures in clusters. In: 2024 USENIX Annual Technical Conference (USENIX ATC 24). Berkeley (CA): USENIX Association; 2024. p. 1101-1117.
Koh C, Kang J, Kim T, Han SW. Temporal-contextual attention network for solid-state drive failure prediction in data centers. IEEE Access. 2024;12:154455-154466. doi:10.1109/ACCESS.2024.3482368.
Fang X, Guan W, Li J, Cao C, Xia B. SiaDFP: a disk failure prediction framework based on Siamese neural network in large-scale data center. IEEE Trans Serv Comput. 2024;17(5):2890-2903. doi:10.1109/TSC.2024.3394692.
Downloads
Published
Issue
Section
License
Authors retain copyright in their published work. Unless an individual Version of Record identifies different terms, JMRIS articles are distributed under the Creative Commons Attribution 4.0 International License (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/. This permits sharing and adaptation, including commercial use, with appropriate attribution, a link to the license, and indication of changes. Third-party material may be subject to separate terms. Existing articles retain the license originally granted to their Version of Record; a policy update does not retrospectively replace that license.