Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Position Papers for the 2024 ASCR Workshop on Neuromorphic Computing for Science

Engineering novel neuromorphic computing systems with functionalities, capabilities, and energy efficiency similar to biological brains is one of the most exciting and challenging scientific endeavors of our time. This workshop aims to identify key research needs, challenges, and next steps necessary to develop biologically realistic neuromorphic circuits primitives that capture the functionality of neural systems found in nature. Moreover, simulating neuromorphic computing primitives integrated into networks will be key to under standing their behavior at scale, particularly for those computing architectures where full-scale commercial fabrication is not yet readily accessible. Appropriate neuroscience datasets and metrics will have to be established to vet proposed neuromorphic circuits.

97 MATHEMATICS AND COMPUTING↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

AMR-Wind: A Performance-Portable, High-Fidelity Flow Solver for Wind Farm Simulations

We present AMR-Wind, a verified and validated high-fidelity computational-fluid-dynamics code for wind farm flows. AMR-Wind is a block-structured, adaptive-mesh, incompressible-flow solver that enables predictive simulations of the atmospheric boundary layer and wind plants. It is a highly scalable code designed for parallel high-performance computing with a specific focus on performance portability for current and future computing architectures, including graphical processing units (GPUs). In this paper, we detail the governing equations, the numerical methods, and the turbine models. Establishing a foundation for the correctness of the code, we present the results of formal verification and validation. The verification studies, which include a novel actuator line test case, indicate that AMR-Wind is spatially and temporally second-order accurate. The validation studies demonstrate that the key physics capabilities implemented in the code, including actuator disk models, actuator line models, turbulence models, and large eddy simulation (LES) models for atmospheric boundary layers, perform well in comparison to reference data from established computational tools and theory. We conclude with a demonstration simulation of a 12-turbine wind farm operating in a turbulent atmospheric boundary layer, detailing computational performance and realistic wake interactions.

17 WIND ENERGY↗

Mu2e - Extinction Monitor Research & Development

Current efforts are being conducted at Fermi National Laboratory to study potential violations in accepted theory that would otherwise suggest a restructuring of our fundamental understanding of the universe. Mu2e is one of these frontier projects that studies charged lepton flavor violation (CLFV) which if observed, would suggest physics beyond the Standard Model. Therefore, this note encompasses several projects that contribute to the fruition of Mu2e investigations. Due to the broad range of disciplinary inconsistencies that each project requires, all the work is being presented as a means of justifying contribution to Mu2e. The projects are comprised of a G4Beamline simulation analyzing 8GeV proton beam interaction with a titanium window of Recycler ring extinction rates using three Cherenkov radiation-based detectors and supplemental work for the implementation of a micro–Telecommunications Computing Architecture (MicroTCA) crate to establish a peak finding algorithm to ensure that the out-of-time beam is less than 10^-10 fractional level along with inefficiency analysis on scintillation counters for the Cosmic-Ray Veto (CRV) analysis. Preliminary results have been achieved for extinction rate simulation by achieving coincidence rates for 2/3-fold and 3/3-fold on the detectors in the order of 10^-9 and 10^-10, respectively. Only preliminary results of a triangular counter and four rectangular di-counters for the CRV have been realized but other non-experimental contributions were made to the development of the MicroTCA crate.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Cyber-Power Co-Simulation for End-to-End Synchrophasor Network Analysis and Applications

The resiliency, reliability and security of the next generation cyber-power smart grid depend upon efficiently leveraging advanced communication and computing technologies. Also, developing real-time data-driven applications is critical to enable wide-area monitoring and control of the cyber-power grid given high-resolution data from Phasor Measurement Units (PMUs). North American Synchrophasor Initiative Network (NASPlnet) provides guidance for PMU data exchanges. With the advancement in networking and grid operation, it is necessary to evaluate the performance of different data flow architectures suggested by NASPInet and analyze the impact on applications. Therefore, we need a cyber-power co-simulation framework that supports very large-scale co-simulation capable of running in parallel, high-performance computing platforms and capturing real-life network behavior. This work presents an end-to-end automated and user-driven cyber-power co-simulation using NS3 to model communication networks, GridPACK to model the power grid, and HELICS as a co-simulation engine. Comparative analysis of latency in synchrophasor networks and a performance evaluation of a power system stabilizer application utilizing PMU data in an IEEE 39 bus test system is presented using this cosimulation testbed.

Mustafa, Hussain M.↗

Computationally Efficient Multiscale Neural Networks Applied to Fluid Flow in Complex 3D Porous Media

Abstract The permeability of complex porous materials is of interest to many engineering disciplines. This quantity can be obtained via direct flow simulation, which provides the most accurate results, but is very computationally expensive. In particular, the simulation convergence time scales poorly as the simulation domains become less porous or more heterogeneous. Semi-analytical models that rely on averaged structural properties (i.e., porosity and tortuosity) have been proposed, but these features only partly summarize the domain, resulting in limited applicability. On the other hand, data-driven machine learning approaches have shown great promise for building more general models by virtue of accounting for the spatial arrangement of the domains’ solid boundaries. However, prior approaches building on the convolutional neural network (ConvNet) literature concerning 2D image recognition problems do not scale well to the large 3D domains required to obtain a representative elementary volume (REV). As such, most prior work focused on homogeneous samples, where a small REV entails that the global nature of fluid flow could be mostly neglected, and accordingly, the memory bottleneck of addressing 3D domains with ConvNets was side-stepped. Therefore, important geometries such as fractures and vuggy domains could not be modeled properly. In this work, we address this limitation with a general multiscale deep learning model that is able to learn from porous media simulation data. By using a coupled set of neural networks that view the domain on different scales, we enable the evaluation of large ( $$>512^3$$ > 512 3 ) images in approximately one second on a single graphics processing unit. This model architecture opens up the possibility of modeling domain sizes that would not be feasible using traditional direct simulation tools on a desktop computer. We validate our method with a laminar fluid flow case using vuggy samples and fractures. As a result of viewing the entire domain at once, our model is able to perform accurate prediction on domains exhibiting a large degree of heterogeneity. We expect the methodology to be applicable to many other transport problems where complex geometries play a central role.

36 MATERIALS SCIENCE↗

All-Atom Biomolecular Simulation in the Exascale Era

Exascale supercomputers have opened the door to dynamic simulations, facilitated by AI/ML techniques, that model biomolecular motions over unprecedented length and time scales. This new capability holds the potential to revolutionize our understanding of fundamental biological processes. Herein we report on some of the major advances that were discussed at a recent CECAM workshop in Pisa, Italy, on the topic with a primary focus on atomic-level simulations. First, we highlight examples of current large-scale biomolecular simulations and the future possibilities enabled by crossing the exascale threshold. Next, we discuss challenges to be overcome in optimizing the usage of these powerful resources. Finally, we close by listing several grand challenge problems that could be investigated with this new computer architecture.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep learning for in situ data compression of large turbulent flow simulations

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

97 MATHEMATICS AND COMPUTING↗

AEflow (Autoencoder fluid flow compression network) [SWR-22-29]

As the size of turbulent flow simulations continues to grow, in situ data compression is becoming increasingly important for visualization, analysis, and restart checkpointing. For these applications, single-pass compression techniques with low computational and communication overhead are crucial. In this paper we present a deep-learning approach to in situ compression using an autoencoder architecture that is customized for three-dimensional turbulent flows and is well suited for contemporary heterogeneous computing resources. The autoencoder is compared against a recently introduced randomized single-pass singular value decomposition (SVD) for three different canonical turbulent flows: decaying homogeneous isotropic turbulence, a Taylor-Green vortex, and turbulent channel flow. Our proposed fully convolutional autoencoder architecture compresses turbulent flow snapshots by a factor of 64 with a single pass, allows for arbitrarily sized input fields, is cheaper to compute than the randomized single-pass SVD for typical simulation sizes, performs well on unseen flow configurations, and has been made publicly available. The results reported here show that the autoencoder dramatically outperforms a randomized single-pass SVD with similar compression ratio and yields comparable performance to a higher-rank decomposition with an order of magnitude less compression in regard to preserving a number of important statistical quantities such as turbulent kinetic energy, enstrophy, and Reynolds stresses.

King, Ryan↗

The Pele Simulation Suite for Reacting Flows at Exascale

In this work, we present the Pele suite of software tools for compressible and incompressible reacting flows. The Pele suite leverages several different libraries, notably AMReX and SUNDIALS, to achieve performance portability on heterogeneous computing architectures across the supercomputing landscape. The Pele suite is comprised of PeleC, a compressible reacting flow block-structured adaptive mesh refinement solver, PeleLMeX, a low-Mach number reacting flow block-structured adaptive mesh refinement solver, Pele-Physics, a library for transport, thermodynamics, finite rate chemistry, soot, spray and radiation physics. The objective of this paper is (i) to present the code development efforts necessary to achieve highly effective and scalable applications for exascale machines and (ii) to detail the performance results of the Combustion-Pele project applications on Oak Ridge National Laboratory's Frontier. We show good weak and strong scaling results for both PeleC and PeleLMeX up to more than 50 billion cells on more than 4096 Frontier graphics processing unit nodes. We also present a capability demonstration simulation of a dual-fuel pulse compression ignition engine (six adaptive mesh refinement levels, and 60 billion cells or 2.1 trillion degrees of freedom) on Frontier, to date one of the largest simulations performed on the first exascale-class supercomputer.

adaptive mesh refinement↗

Metal additive manufacturing simulation across length, time, and computing scales

Metal additive manufacturing (AM) offers a unique opportunity for production of advanced materials and complex geometries. However, variability in microstructure and properties challenges conventional approaches to design, process optimization, qualification, and materials selection. Modeling and simulation can improve understanding of AM processing and materials, but also poses major challenges for existing computational methods. Simultaneously, modern scientific computing hardware has become increasingly complex, most notably with the adoption of hybrid architectures such as Graphical Processing Units (GPUs). If appropriately utilized, emerging computational capabilities provide an opportunity to reveal new insight into AM processing and the resulting material structure and properties. In this review we describe the computational AM landscape, identify critical gaps, and highlight opportunities to impact the development and application of AM. First, the requirements and challenges of representative AM problem statements will be defined. Here, these problems range from scientific studies to industrial applications and are designed to capture the breadth of challenges facing the AM community. Next, the current state of AM modeling and simulation is evaluated, broken down by enabling hardware and software, process simulation, microstructure simulation, and property simulation. Each section describes the diversity of simulation approaches and associated trade-offs in physical fidelity and computational expense. Each area is then assessed based on their suitability and readiness for current and developing computational architectures. Lastly, the greatest opportunities for future research and application are highlighted, including gaps in modeling capabilities, opportunities for near-term application, and key scientific challenges.

additive manufacturing↗

Engineering Dynamically Decoupled Quantum Simulations with Trapped Ions

An external drive can improve the coherence of a quantum many-body system by averaging out noise sources. It can also be used to realize models that are inaccessible in the static limit, through Floquet Hamiltonian engineering. The full possibilities for combining these tools remain unexplored. We develop the requirements needed for a pulse sequence to decouple a quantum many-body system from an external field without altering the intended dynamics. Demonstrating this technique experimentally in an ion-trap platform, we show that it can provide a large improvement to coherence in real-world applications. Finally, we engineer an approximate quantum simulation of the Haldane-Shastry model, an exactly solvable paradigm for long-range interacting spins. Our results expand and unify the quantum simulation toolbox.

97 MATHEMATICS AND COMPUTING↗

High Flux Isotope Reactor Low-Enriched Uranium Low Density Silicide Fuel Design Parameters

High Flux Isotope Reactor (HFIR) highly enriched uranium (HEU) to low-enriched uranium (LEU) conversion activities are ongoing as part of the Department of Energy (DOE) National Nuclear Security Administration (NNSA)’s nuclear nonproliferation mission. Design activities studying the conversion of HFIR from HEU to LEU fuel explored different fuel design features and shapes with a low density uranium-silicide dispersion (U 3 Si 2 -Al) fuel, which has a uranium density of 4.8 gU/cm 3 . The goal of these studies is to generate several HFIR LEU fuel designs of varying fuel fabrication complexity that meet the current HEU performance metrics and safety requirements. The documented designs will serve as references for fuel fabrication and qualification activities. Recent advancements in modeling and simulation tools enable quick prototyping of fuel designs. Shift, a Monte Carlo neutron transport and depletion tool optimized for high-performance computing (HPC) architectures, is used for efficient fuel cycle and performance metrics calculations. The HFIR Steady State Heat Transfer Code (HSSHTC) is used to vet the thermal safety margin. Also, a new automation tool that connects all fuel design analysis steps, named Python HFIR Analysis and Measurement Engine (PHAME), has been developed to expedite the design study in an efficient and reproducible manner. Leveraging these tools, several candidate fuel designs were selected for varying fabrication complexity. This report provides design feature details for four selected HFIR LEU low density U 3 Si 2 -Al fuel designs and their corresponding performance and safety metrics. Nominal, best-estimate design parameters and irradiation conditions, including fission rate densities, power densities, heat fluxes, and cumulative fission densities are provided for candidate fuel designs relevant to framing irradiation experiments to support fuel qualification efforts. Simulations show that the low density U 3 Si 2 -Al, with design features to enhance safety, can meet HEU core performance metrics and safety requirements if the reactor power is increased from 85 MW (HEU) to 95 MW (LEU) and if the active fuel length is increased from 50.80 cm (HEU) to 55.88 cm (LEU).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Dynamical Complexity of Non-Gaussian Many-Body Systems with Dissipation

We characterize the dynamical state of many-body bosonic and fermionic many-body models with intersite Gaussian couplings, on-site non-Gaussian interactions, and local dissipation comprising incoherent particle loss, particle gain, and dephasing. We first establish that, for fermionic systems, if the dephasing noise is larger than the non-Gaussian interactions, irrespective of the Gaussian coupling strength, the system state is a convex combination of Gaussian states at all times. Furthermore, for bosonic systems, we show that if the particle loss and particle gain rates are larger than the Gaussian intersite couplings, the system remains in a separable state at all times. Building on this characterization, we establish that at noise rates above a threshold, there exists a classical algorithm that can efficiently sample from the system state of both the fermionic and bosonic models. Finally, we show that, unlike fermionic systems, bosonic systems can evolve into states that are not convex Gaussian even when the dissipation is much higher than the on-site non-Gaussianity. Similarly, unlike bosonic systems, fermionic systems can generate entanglement even with noise rates much larger than the intersite couplings.

Computational complexity↗

Special Topic on High Performance Computing in Chemical Physics

Computational modeling and simulation have become indispensable scientific tools in virtually all areas of chemical, biomolecular, and materials systems research. Computation can provide unique and detailed atomic level information that is difficult or impossible to obtain through analytical theories and experimental investigations. In addition, recent advances in micro-electronics have resulted in computer architectures with unprecedented computational capabilities, from the largest supercomputers to common desktop computers. In conclusion, combined with the development of new computational domain science methodologies and novel programming models and techniques, this has resulted in modeling and simulation resources capable of providing results at or better than experimental chemical accuracy and for systems in increasingly realistic chemical environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Framework for Integrating Quantum Simulation and High Performance Computing

Scientific applications are starting to explore the viability of quantum computing. This exploration typically begins with quantum simulations that can run on existing classical platforms, albeit without the performance advantages of real quantum resources. In the context of high-performance computing (HPC), the incorporation of simulation software can often take advantage of the powerful resources to help scale-up the simulation size. The configuration, installation and operation of these quantum simulation packages on HPC resources can often be rather daunting and increases friction for experimentation by scientific application developers. We describe a framework to help streamline access to quantum simulation software running on HPC resources. This includes an interface for circuit-based quantum computing tasks, as well as the necessary resource management infrastructure to make effective use of the underlying HPC resources. The primary contributions of this work include a classification of different usage models for quantum simulation in an HPC context, a review of the software architecture for our approach and a detailed description of the prototype implementation to experiment with these ideas using two different simulators (TNQVM & NWQ-Sim). We include initial experimental results running on the Frontier supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) using a synthetic workload generated via the SupermarQ quantum benchmarking framework.

Shehata, Amir [ORNL] (ORCID:0000000224531426)↗

MFIX-Exa: A path toward exascale CFD-DEM simulations

MFIX-Exa is a computational fluid dynamics–discrete element model (CFD-DEM) code designed to run efficiently on current and next-generation supercomputing architectures. MFIX-Exa combines the CFD-DEM expertise embodied in the MFIX code—which was developed at NETL and is used widely in academia and industry—with the modern software framework, AMReX, developed at LBNL. The fundamental physics models follow those of the original MFIX, but the combination of new algorithmic approaches and a new software infrastructure will enable MFIX-Exa to leverage future exascale machines to optimize the modeling and design of multiphase chemical reactors.

97 MATHEMATICS AND COMPUTING↗

Application-specific machine-learned interatomic potentials: exploring the trade-off between DFT convergence, MLIP expressivity, and computational cost

Machine-learned interatomic potentials (MLIPs) are revolutionizing computational materials science and chemistry by offering an efficient alternative to ab initio molecular dynamics (MD) simulations. However, fitting high-quality MLIPs remains a challenging, time-consuming, and computationally intensive task where numerous trade-offs have to be considered, e.g., How much and what kind of atomic configurations should be included in the training set? Which level of ab initio convergence should be used to generate the training set? Which loss function should be used for fitting the MLIP? Which machine learning architecture should be used to train the MLIP? The answers to these questions significantly impact both the computational cost of MLIP training and the accuracy and computational cost of subsequent MLIP MD simulations. In this study, we use a configurationally diverse beryllium dataset and quadratic spectral neighbor analysis potential. We demonstrate that joint optimization of energy versus force weights, training set selection strategies, and convergence settings of the ab initio reference simulations, as well as model complexity can lead to a significant reduction in the overall computational cost associated with training and evaluating MLIPs. This opens the door to computationally efficient generation of high-quality MLIPs for a range of applications which demand different accuracy versus training and evaluation cost trade-offs.

36 MATERIALS SCIENCE↗