Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “High performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Predictive Models and High-Performance Computing as Tools to Accelerate the Scaling-Up of New Bio-Based Fuels Workshop: Summary Report

This report summarizes the results of a virtual workshop sponsored by the Bioenergy Technologies Office held on June 9–11, 2020. The workshop discussed best practices for utilizing mathematical modeling tools across multiple scales to reduce technology uncertainty and accelerate scaling-up of biorefinery/chemical production equipment and optimize operations.

09 BIOMASS FUELS↗

High Performance Computing to Quantify the Evolution of Microscopic Concentration Gradients During Flash Processing

During the Flash process, the cross section of a plain-carbon or a low-alloy steel is austenitized through rapid heating and transformed on rapid cooling to a predominantly martensite + bainite structure with small amounts of retained austenite. Unlike conventional heat treating, homogeneity is intentionally avoided during Flash processing of steels. The Flash process assembly consists of a pair of rolls that transfer the steel sheets through the heating and cooling stage of the thermal cycle. The initial microstructure of the steel consists of ferrite (body-centered cubic iron) + carbide ((Fe,X)mCn) mixture. The heating rate through the peak temperature is a function of temperature and reaches a peak of about 300-400°C/s and the cooling rate has a maximum value of 3,000-4,000°C/s. The on-heating phase transformations include carbide dissolution, austenite (face-centered cubic iron) nucleation and growth, and diffusion of carbon and other substitutional elements in the steel. The on-cooling phase transformations include formation of martensite (body-centered tetragonal phase containing supersaturated solute) and bainite (ferrite plates with or without fine carbides). In this project, the focus is on Fe-C-Cr steels that are currently Flash processed for armor applications. The modeling effort proposed here will help optimize the Flash thermal cycle for these low alloy steels to achieve the target performance, which is an ongoing effort at SFP Works. A significant feature of Flash processed Fe-C-Cr steels is the presence of scatter in the through-thickness in the sheet. The variability in hardness results from a variability in the bainite + martensite microstructure that is sensitive to the local chemical concentration of C and Cr. Such a chemical inhomogeneity is intentionally obtained in the Flash process. Although such a microstructural gradient is presumably responsible for the exceptional properties of the Flash processed steel, it is very important to quantify the gradients as a function of Flash variabilities in processing parameters and the input microstructure. Understanding the mechanistic pathway that leads to microstructural gradients could be ground-breaking and instrumental for achieving better process control and optimized microstructural state to meet application-specific strength-ductility requirements. Since the final microstructure depends on setting up precise solute concentration gradients through a rapid heating process, and transforming these regions into various phases, it is important to understand how small changes in steel chemistry, input microstructure (carbide size and distribution), and process variables (Flash thermal cycle) will impact the solute concentration gradients.

97 MATHEMATICS AND COMPUTING↗

Los Alamos National Laboratory High Performance Computing Overview [Slides]

We need such big computers because without testing, we don't understand the health of the AGING stockpile. Our big computers store a wealth of test data. We model weapons as they were tested and see if we can match the results to validate models, and we then use validated models along with dismantlement information on aging to certify the stockpile. Sometimes the results of tests can’t be explained well, so even that part requires massive computations, but applying the model to future use is an astonishing amount of computing.

97 MATHEMATICS AND COMPUTING↗

Inference-Optimized AI and High Performance Computing for Gravitational Wave Detection at Scale

We introduce an ensemble of artificial intelligence models for gravitational wave detection that we trained in the Summit supercomputer using 32 nodes, equivalent to 192 NVIDIA V100 GPUs, within 2 h. Once fully trained, we optimized these models for accelerated inference using NVIDIA TensorRT. We deployed our inference-optimized AI ensemble in the ThetaGPU supercomputer at Argonne Leadership Computer Facility to conduct distributed inference. Using the entire ThetaGPU supercomputer, consisting of 20 nodes each of which has 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, our NVIDIA TensorRT-optimized AI ensemble processed an entire month of advanced LIGO data (including Hanford and Livingston data streams) within 50 s. Our inference-optimized AI ensemble retains the same sensitivity of traditional AI models, namely, it identifies all known binary black hole mergers previously identified in this advanced LIGO dataset and reports no misclassifications, while also providing a 3X inference speedup compared to traditional artificial intelligence models. We used time slides to quantify the performance of our AI ensemble to process up to 5 years worth of advanced LIGO data. In this synthetically enhanced dataset, our AI ensemble reports an average of one misclassification for every month of searched advanced LIGO data. We also present the receiver operating characteristic curve of our AI ensemble using this 5 year long advanced LIGO dataset. This approach provides the required tools to conduct accelerated, AI-driven gravitational wave detection at scale.

97 MATHEMATICS AND COMPUTING↗

A high level language for a high performance computer

The proposed computational aerodynamic facility will join the ranks of the supercomputers due to its architecture and increased execution speed. At present, the languages used to program these supercomputers have been modifications of programming languages which were designed many years ago for sequential machines. A new programming language should be developed based on the techniques which have proved valuable for sequential programming languages and incorporating the algorithmic techniques required for these supercomputers. The design objectives for such a language are outlined.

Perrott, R. H.↗

A vectorized Lanczos eigensolver for high-performance computers

The computational strategies used to implement a Lanczos-based-method eigensolver on the latest generation of supercomputers are described. Several examples of structural vibration and buckling problems are presented that show the effects of using optimization techniques to increase the vectorization of the computational steps. The data storage and access schemes and the tools and strategies that best exploit the computer resources are presented. The method is implemented on the Convex C220, the Cray 2, and the Cray Y-MP computers. Results show that very good computation rates are achieved for the most computationally intensive steps of the Lanczos algorithm and that the Lanczos algorithm is many times faster than other methods extensively used in the past.

Bostic, Susan W.↗

A parallel-vector algorithm for rapid structural analysis on high-performance computers

A fast, accurate Choleski method for the solution of symmetric systems of linear equations is presented. This direct method is based on a variable-band storage scheme and takes advantage of column heights to reduce the number of operations in the Choleski factorization. The method employs parallel computation in the outermost DO-loop and vector computation via the 'loop unrolling' technique in the innermost DO-loop. The method avoids computations with zeros outside the column heights, and as an option, zeros inside the band. The close relationship between Choleski and Gauss elimination methods is examined. The minor changes required to convert the Choleski code to a Gauss code to solve non-positive-definite symmetric systems of equations are identified. The results for two large-scale structural analyses performed on supercomputers, demonstrate the accuracy and speed of the method.

Storaasli, Olaf O.↗

High performance computing system for flight simulation at NASA Langley

The computer architecture and components used in the NASA Langley Advanced Real-Time Simulation System (ARTSS) are briefly described and illustrated with diagrams and graphs. Particular attention is given to the advanced Convex C220 processing units, the UNIX-based operating system, the software interface to the fiber-optic-linked Computer Automated Measurement and Control system, configuration-management and real-time supervisor software, ARTSS hardware modifications, and the current implementation status. Simulation applications considered include the Transport Systems Research Vehicle, the Differential Maneuvering Simulator, the General Aviation Simulator, and the Visual Motion Simulator.

Cleveland, Jeff I., II↗