Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Disk failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Collection of Disk Failure Events from Alpine, the Parallel File System for Summit Supercomputer

This dataset contains disk (HDD) failure events collected from the Alpine storage system of the Summit supercomputer, hosted at OLCF, spanning from January 4, 2019, to December 21, 2023 (a total of 4 years, 11 months, and 18 days), covering 89% of its operational lifetime. It includes 3,766 disk failure events, each recorded with its detection timestamp (in ISO 8601 format) and detailed by its location within the storage system - rack, enclosure, and drive slot number.

97 MATHEMATICS AND COMPUTING

Disk Failure Dataset from the Campaign Storage System

This dataset consists of 1,389 disk (HDD) failure events collected from the Campaign storage system at LANL. The Campaign system supported various compute platforms throughout its lifespan, including Cielo, Fire, Ice, and notably, the Trinity supercomputer. Each recorded event includes its detection timestamp (in ISO 8601 format) and details such as its location within the storage system—rack, enclosure, and drive slot number. The data, spanning from May 4, 2021, to July 25, 2023 (2 years, 2 months, and 22 days), represents failure events from the terminal years of Campaign's operational period, accounting for 26% of its total operational time.

97 MATHEMATICS AND COMPUTING

From Failure to Insight: Analyzing Disk Breakdowns in Large-Scale HPC Environments

Disk failure data provides valuable insights for preventing failures, enhancing storage robustness, guiding system design and deployment, and ensuring reliable operations at data centers. This paper introduces two disk failure datasets collected from large-scale HPC production environments over the past five years, comprising over 5,000 failure records from more than 40,000 disks. We analyzed these datasets across multiple dimensions, including temporal, spatial, and relational trends, and performed a comprehensive reliability assessment. Our analysis yielded numerous observations and insights that influence various operational aspects of HPC storage systems. We believe this study offers a holistic understanding of disk failure trends likely to interest the HPC storage community.

George, Anjus

Rotor fragment protection program: Statistics on aircraft gas turbine engine rotor failures that occurred in US commercial aviation during 1979

Statistical information relating to the number of gas turbine engine rotor failures which occurred during 1979 in commercial aviation service use is provided. The predominant failure mode involved blade fragments, 84 percent of which were contained. No uncontained disk failures occurred and although fewer rotor rim and seal failures occurred, 100 percent and 50 percent, respectively, were uncontained. Sixty-eight percent of the 157 rotor failures occurred during the take-off and climb stages of flight.

Delucia, R. A.

Fatigue Crack Growth Behavior Evaluation of Grainex Mar-M 247 for NASA's High Temperature, High Speed Turbine Seal Test Rig

The fatigue crack growth behavior of Grainex Mar-M 247 is evaluated for NASA s Turbine Seal Test Facility. The facility is used to test air-to-air seals primarily for use in advanced jet engine applications. Because of extreme seal test conditions of temperature, pressure, and surface speeds, surface cracks may develop over time in the disk bolt holes. An inspection interval is developed to preclude catastrophic disk failure by using experimental fatigue crack growth data. By combining current fatigue crack growth results with previous fatigue strain-life experimental work, an inspection interval is determined for the test disk. The fatigue crack growth life of the NASA disk bolt holes is found to be 367 cycles at a crack depth of 0.501 mm using a factor of 2 on life at maximum operating conditions. Combining this result with previous fatigue strain-life experimental work gives a total fatigue life of 1032 cycles at a crack depth of 0.501 mm. Eddy-current inspections are suggested starting at 665 cycles since eddy current detection thresholds are currently at 0.381 mm. Inspection intervals are recommended every 50 cycles when operated at maximum operating conditions.

Delgado, Irebert R.

Tutorial: Performance and reliability in redundant disk arrays

A disk array is a collection of physically small magnetic disks that is packaged as a single unit but operates in parallel. Disk arrays capitalize on the availability of small-diameter disks from a price-competitive market to provide the cost, volume, and capacity of current disk systems but many times their performance. Unfortunately, relative to current disk systems, the larger number of components in disk arrays leads to higher rates of failure. To tolerate failures, redundant disk arrays devote a fraction of their capacity to an encoding of their information. This redundant information enables the contents of a failed disk to be recovered from the contents of non-failed disks. The simplest and least expensive encoding for this redundancy, known as N+1 parity is highlighted. In addition to compensating for the higher failure rates of disk arrays, redundancy allows highly reliable secondary storage systems to be built much more cost-effectively than is now achieved in conventional duplicated disks. Disk arrays that combine redundancy with the parallelism of many small-diameter disks are often called Redundant Arrays of Inexpensive Disks (RAID). This combination promises improvements to both the performance and the reliability of secondary storage. For example, IBM's premier disk product, the IBM 3390, is compared to a redundant disk array constructed of 84 IBM 0661 3 1/2-inch disks. The redundant disk array has comparable or superior values for each of the metrics given and appears likely to cost less. In the first section of this tutorial, I explain how disk arrays exploit the emergence of high performance, small magnetic disks to provide cost-effective disk parallelism that combats the access and transfer gap problems. The flexibility of disk-array configurations benefits manufacturer and consumer alike. In contrast, I describe in this tutorial's second half how parallelism, achieved through increasing numbers of components, causes overall failure rates to rise. Redundant disk arrays overcome this threat to data reliability by ensuring that data remains available during and after component failures.

Gibson, Garth A.

Redundant disk arrays: Reliable, parallel secondary storage

During the past decade, advances in processor and memory technology have given rise to increases in computational performance that far outstrip increases in the performance of secondary storage technology. Coupled with emerging small-disk technology, disk arrays provide the cost, volume, and capacity of current disk subsystems, by leveraging parallelism, many times their performance. Unfortunately, arrays of small disks may have much higher failure rates than the single large disks they replace. Redundant arrays of inexpensive disks (RAID) use simple redundancy schemes to provide high data reliability. The data encoding, performance, and reliability of redundant disk arrays are investigated. Organizing redundant data into a disk array is treated as a coding problem. Among alternatives examined, codes as simple as parity are shown to effectively correct single, self-identifying disk failures.

Gibson, Garth Alan

Redundant Disk Arrays in Transaction Processing Systems

We address various issues dealing with the use of disk arrays in transaction processing environments. We look at the problem of transaction undo recovery and propose a scheme for using the redundancy in disk arrays to support undo recovery. The scheme uses twin page storage for the parity information in the array. It speeds up transaction processing by eliminating the need for undo logging for most transactions. The use of redundant arrays of distributed disks to provide recovery from disasters as well as temporary site failures and disk crashes is also studied. We investigate the problem of assigning the sites of a distributed storage system to redundant arrays in such a way that a cost of maintaining the redundant parity information is minimized. Heuristic algorithms for solving the site partitioning problem are proposed and their performance is evaluated using simulation. We also develop a heuristic for which an upper bound on the deviation from the optimal solution can be established.

Mourad, Antoine Nagib

Braking System for Wind Turbines

Operating turbine stopped smoothly by fail-safe mechanism. Windturbine braking systems improved by system consisting of two large steel-alloy disks mounted on high-speed shaft of gear box, and brakepad assembly mounted on bracket fastened to top of gear box. Lever arms (with brake pads) actuated by spring-powered, pneumatic cylinders connected to these arms. Springs give specific spring-loading constant and exert predetermined load onto brake pads through lever arms. Pneumatic cylinders actuated positively to compress springs and disengage brake pads from disks. During power failure, brakes automatically lock onto disks, producing highly reliable, fail-safe stops. System doubles as stopping brake and "parking" brake.

Krysiak, J. E.

The Competition of Failure Modes in an Additively Manufactured Disk Superalloy

Additive manufacturing of powder metallurgy disk superalloys can produce unique microstructures that are different from those usually encountered in traditional processing by consolidation, forging, and heat treatments. Unusual variations in grain size, major and minor phase precipitate sizes, and defects can occur. The associated failure modes of these unique microstructures are of high interest. The objective of this study was to compare the failure modes for a powder metallurgy disk superalloy LSHR produced by electron beam melting additive manufacturing. Specimens were subsequently given different solution heat treatments and a fixed aging heat treatment. Tensile, creep, and fatigue failure modes were screened in tests at elevated temperatures. Failure modes were considered with respect to these unique microstructures.

additive manufacturing

Competition of Failure Modes in an Additively Manufactured Disk Superalloy

Additive manufacturing of powder metallurgy (PM) disk superalloys can produce unique microstructures that differ from those usually encountered in traditional processing as the result of consolidation, forging, and heat treatments. Unusual variations in grain size, major and minor phase precipitate sizes, and defects can occur. The failure modes associated with these unique microstructures are of high interest. The objective of this study was to compare the failure modes for a low solvus, high refractory (LSHR) PM disk superalloy produced by electron-beam-melting additive manufacturing. Specimens were subsequently given different solution heat treatments and a fixed aging heat treatment. Tensile, creep, and fatigue failure modes were screened in tests at elevated temperatures. Failure modes were considered with respect to these unique microstructures.

additive manufacturing

Advanced turbine disk designs to increase reliability of aircraft engines

Results of analytical studies to improve the low cycle fatigue lives and reliability of turbine disks in high performance gas turbine engines are presented. Advanced disk concepts were evaluated for the first stage high pressure turbines of the CF6-50 and JT8D-17 engines. The advanced disk designs are compared to the existing disks on the bases of cycles to crack initiation and overspeed capability for initially unflawed disks, crack propagation cycles to failure for initially flawed disks, and available kinetic energy of disk burst fragments.

Kaufman, A.

Advanced turbine disk designs to increase reliability of aircraft engines

Results of analytical studies to improve the low cycle fatigue lives and reliability of turbine disks in high-performance gas turbine engines are presented. Advanced disk concepts were evaluated for the first-stage high pressure turbines of the CF6-50 and JT8D-17 engines. The advanced disk designs are compared to the existing disks on the bases of cycles to crack initiation and overspeed capability for initially unflawed disks, crack propagation cycles to failure for initially flawed disks, and available kinetic energy of disk burst fragments.

Kaufman, A.

Fatigue Life of a NiCr-Coated Powder Metallurgy Disk Superalloy After Varied Processing and Exposures

A protective ductile NiCr coating has shown promise to mitigate oxidation and corrosion attack on superalloy disk alloys. The effects of this coating on fatigue life and failure modes of the disk superalloy are an important concern. The objective of this study was to investigate the fatigue life and failure modes of disk superalloy specimens protected by this coating, using varied pre-coating and post-coating processes. Cylindrical gage fatigue specimens of a powder metallurgy-processed disk superalloy were grit blast or wet blast before being coated with a ductile NiCrY coating, then shot peened at low or medium levels after coating. All were then heat treated, some exposed, and finally all were subjected to fatigue at high temperature. The effects of varied pre-coating treatment, post-coating shot peening, and oxidation plus hot corrosion exposures on fatigue life with the coating were compared.

coating

Rotor fragment protection program: Statistics on aircraft gas turbine ngine rotor failures that occurred in U.S. commercial aviation during 1978

This report presents statistical information relating to the number of gas turbine engine rotor failures which occurred in commercial aviation service use. The predominant failure involved blade fragments, 82.4 percent of which were contained. Although fewer rotor rim, disk, and seal failures occurred, 33.3%, 100% and 50% respectively were uncontained. Sixty-five percent of the 166 rotor failures occurred during the takeoff and climb stages of flight.

Delucia, R. A.

Implementing Journaling in a Linux Shared Disk File System

In computer systems today, speed and responsiveness is often determined by network and storage subsystem performance. Faster, more scalable networking interfaces like Fibre Channel and Gigabit Ethernet provide the scaffolding from which higher performance computer systems implementations may be constructed, but new thinking is required about how machines interact with network-enabled storage devices. In this paper we describe how we implemented journaling in the Global File System (GFS), a shared-disk, cluster file system for Linux. Our previous three papers on GFS at the Mass Storage Symposium discussed our first three GFS implementations, their performance, and the lessons learned. Our fourth paper describes, appropriately enough, the evolution of GFS version 3 to version 4, which supports journaling and recovery from client failures. In addition, GFS scalability tests extending to 8 machines accessing 8 4-disk enclosures were conducted: these tests showed good scaling. We describe the GFS cluster infrastructure, which is necessary for proper recovery from machine and disk failures in a collection of machines sharing disks using GFS. Finally, we discuss the suitability of Linux for handling the big data requirements of supercomputing centers.

Preslan, Kenneth W.

The Effects of Hot Corrosion Pits on the Fatigue Resistance of a Disk Superalloy

The effects of hot corrosion pits on low cycle fatigue life and failure modes of the disk superalloy ME3 were investigated. Low cycle fatigue specimens were subjected to hot corrosion exposures producing pits, then tested at low and high temperatures. Fatigue lives and failure initiation points were compared to those of specimens without corrosion pits. Several tests were interrupted to estimate the fraction of fatigue life that fatigue cracks initiated at pits. Corrosion pits significantly reduced fatigue life by 60 to 98 percent. Fatigue cracks initiated at a very small fraction of life for high temperature tests, but initiated at higher fractions in tests at low temperature. Critical pit sizes required to promote fatigue cracking were estimated, based on measurements of pits initiating cracks on fracture surfaces.

Gabb, Timothy P.

Superalloy Disk With Dual-Grain Structure Spin Tested

Advanced nickel-base disk alloys for future gas turbine engines will require greater temperature capability than current alloys, but they must also continue to deliver safe, reliable operation. An advanced, nickel-base disk alloy, designated Alloy 10, was selected for evaluation in NASA s Ultra Safe Propulsion Project. Early studies on small test specimens showed that heat treatments that produced a fine grain microstructure promoted high strength and long fatigue life in the bore of a disk, whereas heat treatments that produced a coarse grain microstructure promoted optimal creep and crack growth resistance in the rim of a disk. On the basis of these results, the optimal combination of performance and safety might be achieved by utilizing a heat-treatment technology that could produce a fine grain bore and coarse grain rim in a nickel-base disk. Alloy 10 disks that were given a dual microstructure heat treatment (DMHT) were obtained from NASA s Ultra-Efficient Engine Technology (UEET) Program for preliminary evaluation. Data on small test specimens machined from a DMHT disk were encouraging. However, the benefit of the dual grain structure on the performance and reliability of the entire disk still needed to be demonstrated. For this reason, a high temperature spin test of a DMHT disk was run at 20 000 rpm and 1500 F at the Balancing Company of Dayton, Ohio, under the direction of NASA Glenn Research Center personnel. The results of that test showed that the DMHT disk exhibited significantly lower crack growth than a disk with a fine grain microstructure. In addition, the results of these tests could be accurately predicted using a two-dimensional, axisymmetric finite element analysis of the DMHT disk. Although the first spin test demonstrated a significant performance advantage associated with the DMHT technology, a second spin test on the DMHT disk was run to determine burst margin. The disk burst in the web at a very high speed, over 39 000 rpm, in line with the predicted location and speed. Furthermore, significant growth of the disk was observed before failure, in line with predictions, clearly demonstrating the reliability and safety of the DMHT technology. Although successful spin testing in Ultra Safe's Nickel Disk Program represents a significant milestone for DMHT technology, realistic engine operation will require repeated loading of a DMHT nickel disk. For this reason, a cyclic spin test study of DMHT nickel disk technology has been proposed to start in fiscal year 2003 under NASA's Aviation Safety Program. The goal of this program will be to determine the fatigue failure mechanism in DMHT nickel disks, thereby demonstrating the reliability and safety of DMHT technology under repetitive loading conditions encountered in realistic engine operation.

Gayda, John