Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “accelerated computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Leveraging Afterglow in Scintillation-based x-ray detectors for spacetime-resolved computed tomography for accelerated acquisition and high-speed event capture

Afterglow in x-ray imaging for high-speed radiography is a constraint that limits imaging systems to low-light/fast decay screens which create poor data. Current approaches focus purely on using low-light yield screens with fast decay to avoid multiple exposure pileup due to afterglow. The goal of this work is to develop a statistical estimation approach to leverage afterglow to improve image quality thus allowing for higher quality imaging components to be used. This will allow for bright screens will slow decay to be used, and then a post-processing step applies the statistical estimation to separate each frame with superior signal compared to low-light/fast decay screens.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Randomized algorithms for accelerating linear algebraic computations

The project supported the development of new methodologies for performing matrix computations that form key building blocks in modern scientific computing, such as low rank approximation of matrices, and efficient representations of global operators that arise in simulations of physical phenomena.

97 MATHEMATICS AND COMPUTING↗

Computed lateral rate and acceleration power spectral response of conventional and STOL airplanes to atmospheric turbulence

Power-spectral-density calculations were made of the lateral responses to atmospheric turbulence for several conventional and short take-off and landing (STOL) airplanes. The turbulence was modeled as three orthogonal velocity components, which were uncorrelated, and each was represented with a one-dimensional power spectrum. Power spectral densities were computed for displacements, rates, and accelerations in roll, yaw, and sideslip. In addition, the power spectral density of the transverse acceleration was computed. Evaluation of ride quality based on a specific ride quality criterion was also made. The results show that the STOL airplanes generally had larger values for the rate and acceleration power spectra (and, consequently, larger corresponding root-mean-square values) than the conventional airplanes. The ride quality criterion gave poorer ratings to the STOL airplanes than to the conventional airplanes.

Lichtenstein, J. H.↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Computational diagnostics for flame acceleration and transition to detonation in a hydrogen/air mixture

A new computational diagnostic method for pressure-induced compressibility is proposed by projecting its local contribution to the chemical explosive mode (CEM) in the chemical explosive mode analysis (CEMA) framework. The new method is validated for the study of detonation development during the deflagration-to-detonation transition (DDT) process. The flame characteristics are identified through the quantification of individual CEM contributions of chemical reaction, diffusion, and pressure-induced compressibility. Numerical simulations are performed to investigate the DDT processes in a stoichiometric hydrogen-air mixture. A Godunov algorithm, fifth-order in space, and third-order in time are used to solve the fully compressible Navier-Stokes equations on a dynamically adapting mesh. A single-step, calibrated chemical diffusive model (CDM) described by Arrhenius kinetics is used for energy release and conservation between the fuel and the product. The new diagnostic method is first applied to onedimensional (1D) canonical flame configurations followed by two-dimensional (2D) simulations of DDT in an obstructed channel where different detonation initiation scenarios are examined using the new CEMA projection formulation. Detailed examinations of the idealized configuration of detonation initiation through shock focusing mechanism at a flame front are also studied using the new formulation. A comparison of the currently proposed CEMA projection and the original formulation by the authors suggests that including the pressure-induced compressibility is essential for the use of CEMA in DDT process. The results also show that the new formulation of CEMA projection can successively capture the detonation initiation through either a gradient mechanism or a direct initiation mechanism, and therefore can be used as an effective local analytical tool for the computational diagnostics of detonation initiation in a DDT process. It was found that detonation development is characterized by a strong contribution of chemistry role to the CEM which is pivotal to the initiation of detonation. The role of compressibility is found enhanced at the edge of the detonation front where diffusion was found to have minimal effects on detonation development.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Formation of Zeolites Responsible for Waste Glass Rate Acceleration: An Experimental and Computational Study for Understanding Thermodynamic and Kinetic Processes

This report summarizes work on the NEUP Project 18-15496, DE-NE0008774 (Formation of Zeolites Responsible for Waste Glass Rate Acceleration: An Experimental and Computational Study for Understanding Thermodynamic and Kinetic Processes). The content of the report is from four journal articles that resulted from work on the project.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Convergence acceleration of the Proteus computer code with multigrid methods

Presented here is the first part of a study to implement convergence acceleration techniques based on the multigrid concept in the Proteus computer code. A review is given of previous studies on the implementation of multigrid methods in computer codes for compressible flow analysis. Also presented is a detailed stability analysis of upwind and central-difference based numerical schemes for solving the Euler and Navier-Stokes equations. Results are given of a convergence study of the Proteus code on computational grids of different sizes. The results presented here form the foundation for the implementation of multigrid methods in the Proteus code.

Demuren, A. O.↗

Convergence acceleration of the Proteus computer code with multigrid methods

This report presents the results of a study to implement convergence acceleration techniques based on the multigrid concept in the two-dimensional and three-dimensional versions of the Proteus computer code. The first section presents a review of the relevant literature on the implementation of the multigrid methods in computer codes for compressible flow analysis. The next two sections present detailed stability analysis of numerical schemes for solving the Euler and Navier-Stokes equations, based on conventional von Neumann analysis and the bi-grid analysis, respectively. The next section presents details of the computational method used in the Proteus computer code. Finally, the multigrid implementation and applications to several two-dimensional and three-dimensional test problems are presented. The results of the present study show that the multigrid method always leads to a reduction in the number of iterations (or time steps) required for convergence. However, there is an overhead associated with the use of multigrid acceleration. The overhead is higher in 2-D problems than in 3-D problems, thus overall multigrid savings in CPU time are in general better in the latter. Savings of about 40-50 percent are typical in 3-D problems, but they are about 20-30 percent in large 2-D problems. The present multigrid method is applicable to steady-state problems and is therefore ineffective in problems with inherently unstable solutions.

Demuren, A. O.↗

Machine-Learning Accelerated Studies of Materials with High Performance and Edge Computing

In the studies of materials, experimental measurements often serve as the reference to verify physics theory and modeling; while theory and modeling provide a fundamental understanding of the physics and principles behind. However, the interactions and cross validation between them have long been a challenge even to-date. Not only that inferring a physics model from experimental data is itself a difficult inverse problem, another major challenge is the orders-of-magnitude longer wall-clock time required to carry out high-fidelity computer modeling to match the timescale of experiments. We envisage that by combining high performance computing, data science, and edge computing technology, the current predicament can be alleviated, and a new paradigm of data-driven physics research will open up. For example, we can accelerate computer simulations by first performing the large-scale modeling on high performance computers and train a machine-learned surrogate model. This computationally inexpensive surrogate model can then be transferred to the computing units residing closely to the experimental facilities to perform high-fidelity simulations at a much higher throughout. The model will also be more amenable to analyzing and validating experimental observations in comparable time scales at a much lower computational cost. Further integration of these accelerated computer simulations with an outer machine learning loop can also inform and direct future experiments, while making the inverse problem of physics model inference more tractable. We will demonstrate a proof-of-concept by using a quantum Monte Carlo application, Dynamical Cluster Approximation (DCA++), to machine-learn a surrogate model and accelerate the study of quantum correlated materials.

Li, Ying Wai↗

An extended BET format for La RC shuttle experiments: Definition and development

A program for shuttle post-flight data reduction is discussed. An extended Best Estimate Trajectory (BET) file was developed. The extended format results in some subtle changes to the header record. The major change is the addition of twenty-six words to each data record. These words include atmospheric related parameters, body axis rate and acceleration data, computed aerodynamic coefficients, and angular accelerations. These parameters were added to facilitate post-flight aerodynamic coefficient determinations as well as shuttle entry air data sensor analyses. Software (NEWBET) was developed to generate the extended BET file utilizing the previously defined ENTREE BET, a dynamic data file which may be either derived inertial measurement unit data or aerodynamic coefficient instrument package data, and some atmospheric information.

Findlay, J. T.↗

Simulations of future particle accelerators: issues and mitigations

The ever increasing demands placed upon machine performance have resulted in the need for more comprehensive particle accelerator modeling. Computer simulations are key to the success of particle accelerators. Many aspects of particle accelerators rely on computer modeling at some point, sometimes requiring complex simulation tools and massively parallel supercomputing. Examples include the modeling of beams at extreme intensities and densities (toward the quantum degeneracy limit), and with ultra-fine control (down to the level of individual particles). In the future, adaptively tuned models might also be relied upon to provide beam measurements beyond the resolution of existing diagnostics. Much time and effort has been put into creating accelerator software tools, some of which are highly successful. However, there are also shortcomings such as the general inability of existing software to be easily modified to meet changing simulation needs. In this paper possible mitigating strategies are discussed for issues faced by the accelerator community as it endeavors to produce better and more comprehensive modeling tools. This includes lack of coordination between code developers, lack of standards to make codes portable and/or reusable, lack of documentation, among others.

43 PARTICLE ACCELERATORS↗

Andes Data Analysis System at the Oak Ridge Leadership Computing Facility

Andes is a (704)-node commodity-type Linux® cluster. Each of Andes’s 704 nodes contain two 16-core 3.0 GHz AMD EPYC 7302 processors with AMD’s Simultaneous Multithreading (SMT) Technology and 256GB of main memory. Andes also has nine large memory GPU nodes. These nodes each have 1TB of main memory and two NVIDIA K80 GPUs with two 14-core 2.30 GHz Intel Xeon processors with HT Technology.

AMD EPYC↗

Hydrodynamic modeling of plasma channel systems for laser plasma accelerators

Structured plasma channels are an essential technology for driving high-gradient, plasma-based acceleration and control of electron and positron beams for advanced concepts accelerators. Laser and gas technologies can permit the generation of long plasma columns known as hydrodynamic, optically-field-ionized (HOFI) channels, which feature low on-axis densities and steep walls. By carefully selecting the background gas and laser properties, one can generate narrow, tunable plasma channels for guiding high intensity laser pulses. Here, we present on the development of simulations of HOFI channels using the FLASH code, a publicly available radiation hydrodynamics code. We explore sensitivities of the channel evolution to laser profile, intensity, and background gas conditions, and identify relevant scalings with laser intensity through a range of practical channel delays.

Accelerators↗

Accelerating Climate Simulations Through Hybrid Computing

Unconventional multi-core processors (e.g., IBM Cell B/E and NYIDIDA GPU) have emerged as accelerators in climate simulation. However, climate models typically run on parallel computers with conventional processors (e.g., Intel and AMD) using MPI. Connecting accelerators to this architecture efficiently and easily becomes a critical issue. When using MPI for connection, we identified two challenges: (1) identical MPI implementation is required in both systems, and; (2) existing MPI code must be modified to accommodate the accelerators. In response, we have extended and deployed IBM Dynamic Application Virtualization (DAV) in a hybrid computing prototype system (one blade with two Intel quad-core processors, two IBM QS22 Cell blades, connected with Infiniband), allowing for seamlessly offloading compute-intensive functions to remote, heterogeneous accelerators in a scalable, load-balanced manner. Currently, a climate solar radiation model running with multiple MPI processes has been offloaded to multiple Cell blades with approx.10% network overhead.

Zhou, Shujia↗

Accelerating Floating-Point Computations with Intel AMX

Intel AMX is a built-in component of recent Intel CPU architectures, first supported by the Intel Sapphire Rapids in 2023, that enables efficient dense matrix multiplications using mixed precision with low-precision data types. The popularity of mixed-precision algorithms has grown recently, primarily due to their use on GPUs to enhance the efficiency of HPC applications, particularly for the training of large language models. The availability of mixed precision on CPUs represents a cost-effective solution for applications where high speed is not critical. This report shows how to use the Intel AMX accelerator through examples in C++ and Python. The examples will focus on mixed-precision floating-point operations obtained by the use of bfloat16 (or BF16) to accelerate code in single precision. We employ a bottom-up methodology, starting from specific register instructions (TMUL operation) to higher-level applications in libraries such as Intel MKL, PyTorch, and TensorFlow, ensuring a comprehensive understanding of the accelerator's potential. Additionally, we provide insights into the expected performance gains when leveraging the accelerator on the Kestrel HPC machine at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗