Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “accelerated computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

AIRNOISE: A Tool for Preliminary Noise-Abatement Terminal Approach Route Design

Noise from aircraft in the airport vicinity is one of the leading aviation-induced environmental issues. The FAA developed the Integrated Noise Model (INM) and its replacement Aviation Environmental Design Tool (AEDT) software to assess noise impact resulting from all aviation activities. However, a software tool is needed that is simple to use for terminal route modification, quick and reasonably accurate for preliminary noise impact evaluation and flexible to be used for iterative design of optimal noise-abatement terminal routes. In this paper, we extend our previous work on developing a noise-abatement terminal approach route design tool, named AIRNOISE, to satisfy this criterion. First, software efficiency has been significantly increased by over tenfold using the C programming language instead of MATLAB. Moreover, a state-of-the-art high performance GPU-accelerated computing module is implemented that was tested to be hundreds time faster than the C implementation. Secondly, a Graphical User Interface (GUI) was developed allowing users to import current terminal approach routes and modify the routes interactively to design new terminal approach routes. The corresponding noise impacts are then calculated and displayed in the GUI in seconds. Finally, AIRNOISE was applied to Baltimore-Washington International Airport terminal approach route to demonstrate its usage.

aircraft noise↗

Constraining the Structure under Lunar Impact Basins with Gravity

The lunar gravity field is used to estimate and constrain the depth of mass anomalies under 19 major lunar impact basins. We use radial gravitational spectra, consisting of accelerations computed either per spherical harmonic degree or cumulatively, at surface locations to obtain the distribution of the gravity signal with spherical harmonic degree and, by implication, to the likely depth below the surface. The results provide estimates for the maximum likely depths of the primary component to the mass anomalies under 19 basins. We find that the maximum depths of the primary source of mascon gravity on the lunar nearside are deeper than the depths for those on the farside when South Pole–Aitken (SPA) is excluded. All basin mass anomalies on the lunar nearside are in the mantle. The maximum depth of the primary source of the mass anomalies is 200 km beneath the surface. The upper 20 km under all basins is largely devoid of anomalies, reflecting predominantly mixing and relaxation associated with impact melt combined with ejecta fallback, as well as homogenization associated with post-basin formation impact bombardment. Except for SPA, all basin anomalies merge with the deep interior at ∼150 km or below, indicating the depth penetration of disruption of the density structure of the lunar interior associated with impact bombardment.

Lunar gravitational field↗

Porting CMS Heterogeneous Pixel Reconstruction to Kokkos

Programming for a diverse set of compute accelerators in addition to the CPU is a challenge. Maintaining separate source code for each architecture would require lots of effort, and development of new algorithms would be daunting if it had to be repeated many times. Fortunately there are several portability technologies on the market such as Alpaka, Kokkos, and SYCL. These technologies aim to improve the developer productivity by making it possible to use the same source code for many different architectures. In this paper we use heterogeneous pixel reconstruction code from the CMS experiment at the CERNL LHC as a realistic use case of a GPU-targeting HEP reconstruction software, and report experience from prototyping a portable version of it using Kokkos. The development was done in a standalone program that attempts to model many of the complexities of a HEP data processing framework such as CMSSW. We also compare the achieved event processing throughput to the original CUDA code and a CPU version of it.

Childers, Taylor↗

CMSSW Scaling Limits on Many-Core Machines

Today the LHC offline computing relies heavily on CPU resources, despite the interest in compute accelerators, such as GPUs, for the longer term future. The number of cores per CPU socket has continued to increase steadily, reaching the levels of 64 cores (128 threads) with recent AMD EPYC processors, and 128 cores on Ampere Altra Max ARM processors. Over the course of the past decade, the CMS data processing framework, CMSSW, has been transformed from a single-threaded framework into a highly concurrent one. The first multithreaded version was brought into production by the start of the LHC Run 2 in 2015. Since then, the framework's threading efficiency has gradually been improved by adding more levels of concurrency and reducing the amount of serial code paths. The latest addition was support for concurrent Runs. In this work we review the concurrency model of the CMSSW, and measure its scalability with real CMS applications, such as simulation and reconstruction, on mode rn many-core machines. We show metrics such as event processing throughput and application memory usage with and without the contribution of I/O, as I/O has been the major scaling limitation for the CMS applications.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Controlling Flexible Robot Arms Using High Speed Dynamics Process

A robot manipulator controller for a flexible manipulator arm having plural bodies connected at respective movable hinges and flexible in plural deformation modes corresponding to respective modal spatial influence vectors relating deformations of plural spaced nodes of respective bodies to the plural deformation modes, operates by computing articulated body quantities for each of the bodies from respective modal spatial influence vectors, obtaining specified body forces for each of the bodies, and computing modal deformation accelerations of the nodes and hinge accelerations of the hinges from the specified body forces, from the articulated body quantities and from the modal spatial influence vectors. In one embodiment of the invention, the controller further operates by comparing the accelerations thus computed to desired manipulator motion to determine a motion discrepancy, and correcting the specified body forces so as to reduce the motion discrepancy. The manipulator bodies and hinges are characterized by respective vectors of deformation and hinge configuration variables, and computing modal deformation accelerations and hinge accelerations is carried out for each one of the bodies beginning with the outermost body by computing a residual body force from a residual body force of a previous body and from the vector of deformation and hinge configuration variables, computing a resultant hinge acceleration from the body force, the residual body force and the articulated hinge inertia, and revising the residual body force modal body acceleration.

Jain, Abhinandan↗

Computational Methods to Accelerate Development of Corrosion Resistant Coatings for Industrial Gas Turbines

Oxidation resistant overlay coatings protect the underlying superalloy component in industrial gas turbines from oxidation attack. Rate of depletion of the Al-rich β-phase in the bond coat governs the lifetime of these coatings. The applicability of a computational method in accelerating the development of corrosion resistant coatings and significantly reducing the extensive experimental effort to predict coating lifetimes and microstructural changes in three-coated Ni-based superalloys for real operational durations (20–40 kh) was undertaken in the present study. Scanning electron microscopy (SEM), energy dispersive X-ray spectroscopy (EDX), and electron microprobe analysis (EPMA) were employed to characterize MCrAlY-coated superalloy substrates (1483, 247 and X4) after exposure at 900 °C in air + 10% H 2 O for up to 20,000 h. The model predicted the longest coating lifetime for the coating on X4 substrate. Precipitation of γ' in the coatings was correctly predicted for all three coating systems. Additionally, the model was able to predict the formation of topologically close packed (TCP)-phases in the investigated coating systems.

Pillai, Rishi R.↗

Portable Acceleration of CMS Computing Workflows with Coprocessors as a Service

Computing demands for large scientific experiments, such as the CMS experiment at the CERN LHC, will increase dramatically in the next decades. To complement the future performance increases of software running on central processing units (CPUs), explorations of coprocessor usage in data processing hold great potential and interest. Coprocessors are a class of computer processors that supplement CPUs, often improving the execution of certain functions due to architectural design choices. We explore the approach of Services for Optimized Network Inference on Coprocessors (SONIC) and study the deployment of this as-a-service approach in large-scale data processing. In the studies, we take a data processing workflow of the CMS experiment and run the main workflow on CPUs, while offloading several machine learning (ML) inference tasks onto either remote or local coprocessors, specifically graphics processing units (GPUs). With experiments performed at Google Cloud, the Purdue Tier-2 computing center, and combinations of the two, we demonstrate the acceleration of these ML algorithms individually on coprocessors and the corresponding throughput improvement for the entire workflow. This approach can be easily generalized to different types of coprocessors and deployed on local CPUs without decreasing the throughput performance. We emphasize that the SONIC approach enables high coprocessor usage and enables the portability to run workflows on different types of coprocessors.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Computational Framework to Accelerate the Discovery of Perovskites for Solar Thermochemical Hydrogen Production: Identification of Gd Perovskite Oxide Redox Mediators

A high-throughput computational framework to identify novel multinary perovskite redox mediators is presented, and this framework is applied to discover the Gd-containing perovskite oxide compositions Gd 2 BB'O 6 , GdA'B 2 O 6 , and GdA'BB'O 6 that split water. The computational scheme uses a sequence of empirical approaches to evaluate the stabilities, electronic properties, and oxygen vacancy thermodynamics of these materials, including contributions to the enthalpies and entropies of reduction, ΔH TR and ΔS TR . This scheme uses the machine-learned descriptor τ to identify compositions that are likely stable as perovskites, the bond valence method to estimate the magnitude and phase of BO 6 octahedral tilting and provide accurate initial estimates of perovskite geometries, and density functional theory including magnetic- and defect-sampling to predict STCH-relevant properties. Eighty-three promising STCH candidate perovskite oxides down-selected from 4392 Gd-containing compositions are reported, three of which are referred to experimental collaborators for characterization and exhibit STCH activity. Our results demonstrate that the high-throughput computational scheme described herein—which is used to evaluate Gd-containing compositions but can be applied to any multinary perovskite oxide compositional space(s) of interest—accelerates the discovery of novel STCH active redox mediators with reasonable computational expense.

36 MATERIALS SCIENCE↗

H-AMR: A New GPU-accelerated GRMHD Code for Exascale Computing with 3D Adaptive Mesh Refinement and Local Adaptive Time Stepping

General relativistic magnetohydrodynamic (GRMHD) simulations have revolutionized our understanding of black hole accretion. Here, we present a GPU-accelerated GRMHD code H-AMR with multifaceted optimizations that, collectively, accelerate computation by 2–5 orders of magnitude for a wide range of applications. First, it introduces a spherical grid with 3D adaptive mesh refinement that operates in each of the three dimensions independently. This allows us to circumvent the Courant condition near the polar singularity, which otherwise cripples high-resolution computational performance. Second, we demonstrate that local adaptive time stepping on a logarithmic spherical-polar grid accelerates computation by a factor of ≲10 compared to traditional hierarchical time-stepping approaches. Jointly, these unique features lead to an effective speed of ~10 9 zone cycles per second per node on 5400 NVIDIA V100 GPUs (i.e., 900 nodes of the OLCF Summit supercomputer). We illustrate H-AMR's computational performance by presenting the first GRMHD simulation of a tilted thin accretion disk threaded by a toroidal magnetic field around a rapidly spinning black hole. With an effective resolution of 13,440 × 4608 × 8092 cells and a total of ≲22 billion cells and ~0.65 × 10 8 time steps, it is among the largest astrophysical simulations ever performed. We find that frame dragging by the black hole tears up the disk into two independently precessing subdisks. The innermost subdisk rotation axis intermittently aligns with the black hole spin, demonstrating for the first time that such long-sought alignment is possible in the absence of large-scale poloidal magnetic fields.

79 ASTRONOMY AND ASTROPHYSICS↗

Navier-Stokes simulation of the supersonic combustion flowfield in a ram accelerator

A computational study of the ram accelerator, a ramjet-in-tube device for accelerating projectiles to ultrahigh velocities, is presented. The analysis is performed using a fully implicit TVD scheme that efficiently solves the Reynolds-averaged Navier-Stokes equations and the species continuity equations associated with a finite rate combustion model. Previous analyses of this concept were based on inviscid assumptions. The present results indicate that viscous effects are of primary importance; in all the cases studied, shock-induced combustion always started in the boundary layer. The effects of Mach number, mixture composition, pressure, and turbulence are investigated for various configurations. Two types of combustion processes, one stable and the other unstable, were observed depending on the inflow conditions. In the unstable case, a detonation wave is formed, which propagates upstream and unstarts the ram accelerator. In the stable case, a solution that converges to steady-state is obtained, in which the combustion wave remains stationary with respect to the ram accelerator projectile. The possibility of stabilizing the detonation wave by means of a backward facing step is also investigated. In addition to these studies, two numerical techniques were tested. These two techniques are vector extrapolation to accelerate convergence, and a diagonal formulation that eliminates the expense of inverting large block matrices that arise in chemically reacting flows.

Yungster, Shaye↗

Navier-Stokes simulation of the supersonic combustion flowfield in a ram accelerator

A computational study of the ram accelerator, a ramjet-in-tube device for accelerating projectiles to ultrahigh velocities, is presented. The analysis is performed using a fully implicit TVD scheme that efficiently solves the Reynolds-averaged Navier-Stokes equations and the species continuity equations associated with a finite rate combustion model. The present results indicate that viscous effects are of primary importance in all the cases studied, shock-induced combustion always started in the boundary layer. The effects of Mach number, mixture composition, pressure and tubulence are investigated for various configurations. Two types of combustion processes, one stable and the other unstable, were observed depending on the inflow conditions. The possibility of stabilizing the detonation wave by means of a backward facing step is also investigated. Two numerical techniques were tested: vector extrapolation, to accelerate convergence, and a diagonal formulation that eliminates the expense of inverting large block matrices which arise in chemically reacting flows.

Yungster, Shaye↗

Computational Workflow for Accelerated Molecular Design Using Quantum Chemical Simulations and Deep Learning Models

Efficient methods for searching the chemical space of molecular compounds are needed to automate and accelerate the design of new functional molecules such as pharmaceuticals. Given the high cost in both resources and time for experimental efforts, computational approaches play a key role in guiding the selection of promising molecules for further investigation. Here, we construct a workflow to accelerate design by combining approximate quantum chemical methods [i.e. density-functional tight-binding (DFTB)], a graph convolutional neural network (GCNN) surrogate model for chemical property prediction, and a masked language model (MLM) for molecule generation. Property data from the DFTB calculations are used to train the surrogate model; the surrogate model is used to score candidates generated by the MLM. The surrogate reduces computation time by orders of magnitude compared to the DFTB calculations, enabling an increased search of chemical space. Furthermore, the MLM generates a diverse set of chemical modifications based on pre-training from a large compound library. We utilize the workflow to search for near-infrared photoactive molecules by minimizing the predicted HOMO-LUMO gap as the target property. Our results show that the workflow can generate optimized molecules outside of the original training set, which suggests that iterations of the workflow could be useful for searching vast chemical spaces in a wide range of design problems.

Blanchard, Andrew↗

An Accelerated Approach for Computationally Efficient Evaluation of Deflagration Time During Abnormal Thermal Events

The challenges of modeling abnormal thermal events for explosives of interest to LLNL can be simplified by an accelerated approach to calculate Prout-Tompkins (P-T) parameters describing autocatalytic deflagration. Rather than depending on detailed modeling in multiphysics codes that include chemical reactivity (e.g. ALE3D), the accelerated approach can calculate P-T parameters based solely on experimental time of deflagration measurements for the explosive of interest. We compare deflagration time for five explosives with previously known P-T parameters at conditions typical of abnormal thermal events and demonstrate that the accelerated approach produces deflagration predictions with a maximum error of 3% and an average error of 1.7%. While the accelerated approach may not be applicable to all experimental conditions, it holds promise of rapid derivation of P-T parameters for accurate modeling of abnormal thermal events of interest to LLNL.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Convergence acceleration of viscous flow computations

A multiple-grid convergence acceleration technique introduced for application to the solution of the Euler equations by means of Lax-Wendroff algorithms is extended to treat compressible viscous flow. Computational results are presented for the solution of the thin-layer version of the Navier-Stokes equations using the explicit MacCormack algorithm, accelerated by a convective coarse-grid scheme. Extensions and generalizations are mentioned.

Johnson, G. M.↗

Harnessing High‐Throughput Computational Methods to Accelerate the Discovery of Optimal Proton Conductors for High‐Performance and Durable Protonic Ceramic Electrochemical Cells

Abstract The pursuit of high‐performance and long‐lasting protonic ceramic electrochemical cells (PCECs) is impeded by the lack of efficient and enduring proton conductors. Conventional research approaches, predominantly based on a trial‐and‐error methodology, have proven to be demanding of resources and time‐consuming. Here, this work reports the findings in harnessing high‐throughput computational methods to expedite the discovery of optimal electrolytes for PCECs. This work methodically computes the oxygen vacancy formation energy (E V ), hydration energy (E H ), and the adsorption energies of H 2 O and CO 2 for a set of 932 oxide candidates. Notably, these findings highlight BaSn x Ce 0.8‐x Yb 0.2 O 3‐δ (BSCYb) as a prospective game‐changing contender, displaying superior proton conductivity and chemical resilience when compared to the well‐regarded BaZr x Ce 0.8‐x Y 0.1 Yb 0.1 O 3‐δ (BZCYYb) series. Experimental validations substantiate the computational predictions; PCECs incorporating BSCYb as the electrolyte achieved extraordinary peak power densities in the fuel cell mode (0.52 and 1.57 W cm −2 at 450 and 600 °C, respectively), a current density of 2.62 A cm −2 at 1.3 V and 600 °C in the electrolysis mode while demonstrating exceptional durability for over 1000‐h when exposed to 50% H 2 O. This research underscores the transformative potential of high‐throughput computational techniques in advancing the field of proton‐conducting oxides for sustainable power generation and hydrogen production.

08 HYDROGEN↗

Modeling of advanced accelerator concepts

Computer modeling is essential to research on Advanced Accelerator Concepts (AAC), as well as to their design and operation. This paper summarizes the current status and future needs of AAC systems and reports on several key aspects of (i) high-performance computing (including performance, portability, scalability, advanced algorithms, scalable I/Os and In-Situ analysis), (ii) the benefits of ecosystems with integrated workflows based on standardized input and output and with integrated frameworks developed as a community, and (iii) sustainability and reliability (including code robustness and usability).

47 OTHER INSTRUMENTATION↗