Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Application-specific machine-learned interatomic potentials: exploring the trade-off between DFT convergence, MLIP expressivity, and computational cost

Machine-learned interatomic potentials (MLIPs) are revolutionizing computational materials science and chemistry by offering an efficient alternative to ab initio molecular dynamics (MD) simulations. However, fitting high-quality MLIPs remains a challenging, time-consuming, and computationally intensive task where numerous trade-offs have to be considered, e.g., How much and what kind of atomic configurations should be included in the training set? Which level of ab initio convergence should be used to generate the training set? Which loss function should be used for fitting the MLIP? Which machine learning architecture should be used to train the MLIP? The answers to these questions significantly impact both the computational cost of MLIP training and the accuracy and computational cost of subsequent MLIP MD simulations. In this study, we use a configurationally diverse beryllium dataset and quadratic spectral neighbor analysis potential. We demonstrate that joint optimization of energy versus force weights, training set selection strategies, and convergence settings of the ab initio reference simulations, as well as model complexity can lead to a significant reduction in the overall computational cost associated with training and evaluating MLIPs. This opens the door to computationally efficient generation of high-quality MLIPs for a range of applications which demand different accuracy versus training and evaluation cost trade-offs.

36 MATERIALS SCIENCE↗

Micropolar deep material network

This study extends the Deep Material Network (DMN), a physics-informed machine learning framework, to predict the homogenized mechanical response of composite materials with micropolar (Cosserat-type) constitutive behavior. This extension incorporates microstructure-dependent size effects, enabling accurate, efficient, and size-aware predictions for composites with complex internal architectures. While traditional, direct numerical simulation micropolar models effectively capture size effects by introducing extra local degrees of freedom, they bring significant computational challenges, particularly for multiscale analyses relevant to engineering applications. The micropolar DMN developed in this paper achieves high accuracy while significantly reducing computation time compared to micropolar direct numerical simulations. This advancement enables multiscale analyses and parameter studies that were previously impractical, such as high-cycle fatigue simulations and comprehensive investigations of internal length scale effects notably in size-dependent plastic response and the optimization of lattice structures. By uniting microstructure-sensitive modeling, physics-driven learning, and scalable surrogate modeling, the micropolar DMN paves the way for accelerated material design, large-scale parametric studies, and the reliable incorporation of size-dependent effects across a wide range of engineering applications, including optimization and next-generation composite design.

36 MATERIALS SCIENCE↗

Effects of Photovoltaic Module Materials and Design on Module Deformation Under Load

Static structural finite element models of an aluminum-framed crystalline silicon (c-Si) photovoltaic (PV) module and a glass-glass thin film PV module were constructed and validated against experimental measurements of deflection under uniform pressure loading. Parametric analyses using Latin Hypercube Sampling (LHS) were performed to propagate simulation input uncertainties related to module material properties, dimensions, and manufacturing tolerances into expected uncertainties in simulated deflection predictions. This exercise verifies the applicability and validity of finite element modeling for predicting mechanical behavior of solar modules across architectures and enables computational models to be used with greater confidence in assessment of module mechanical stressors and design for reliability. Sensitivity analyses were also performed on the uncertainty quantification data sets using linear correlation coefficients to elucidate the key parameters influencing module deformation. This information has implications on which materials or parameters may be optimized to best increase module stiffness and reliability, whether the key optimization parameters change with module architecture or loading magnitudes, and whether parameters such as frame design and racking must be replicated in reduced-scale reliability studies to adequately capture full module mechanical behavior.

14 SOLAR ENERGY↗

Applying corrective machine learning in the E3SM atmosphere model in C++ (EAMxx)

The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM) is the newest addition to the family of earth system models capable of explicitly resolving convective systems. SCREAM is a kilometer-scale configuration of the advanced E3SM Atmosphere Model (EAMxx), designed for heterogeneous computing architectures. While the enhanced accuracy of kilometer-scale modeling offers significant benefits, it comes with a substantial computational cost, limiting feasible simulation durations to only a few years to a few decades, even on the fastest supercomputers. Machine learning presents an opportunity for scientists to achieve the high accuracy of storm-resolving models at a significantly reduced cost. Building on the previous success of applying corrective machine learning (ML) to the FV3GFS earth system model, this study explores the effects of implementing corrective-ML in EAMxx-SCREAM. We also address the computational challenges of integrating our implementation of corrective-ML, which is written in Python, with the C++/Kokkos EAMxx driver, as well as potential reasons why this approach has not proved as effective for EAMxx-SCREAM as for FV3GFS.

Environmental sciences↗

Engagement: Hyperparameter Optimization of Generative Adversarial Network Models for High-Energy Physics Simulations

We present our SciDAC FASTMath-HEP partnership results for tuning generative adversarial models (GANs) for high energy physics applications. The GANs are used in hybrid simulations to accelerate otherwise time-consuming computations. We optimize for both, prediction accuracy and variability with the goal to find GAN architectures that are reliable and robust.

high energy physics↗

Towards Ultra-high-resolution E3SM Land Modeling on Exascale Computers

Here we present an ultra-high-resolution E3SM land model (uELM) for high-fidelity land simulations targeting new Exascale computers. After considering modeling infrastructure compatibility and ELM software features, we designed a parallel model for the uELM development targeting hybrid architectures of new US Exascale computers. We also described a function unit test framework to expedite the piece-wise code porting (with compiler directives), verification, and global variable management. Furthermore, in this study, we report an early uELM model development using OpenACC within a function unit test framework on a pre-Exascale computer, demonstrate the performance of a uLEM submodel with a 3.0-time speedup, and summarize the code porting experience regarding global variable handling, deepcopy, memory reduction, and parallel loop reconstruction.

97 MATHEMATICS AND COMPUTING↗

Privacy Preserving Federated Learning for Advanced Scientific Ecosystems

We present a framework to provide privacy preserving (PP) federating learning (FL) across multiple computational and experimental facilities. This work joins the compute capabilities of National Energy Research Scientific Computing Center (NERSC) and Oak Ridge National Laboratory Research Cloud (ORC) with simulated experimental data, such as those produced at the SLAC National Accelerator Laboratory and Spallation Neutron Source (SNS). We describe the software infrastructure developed to provide privacy for computational and experimental networks. We developed algorithmic privacy across the federated system by embedding database security, computation, and communication into the federation architecture, utilizing scientific tools developed by the experimental community.

Archibald, Rick [ORNL] (ORCID:0000000245389780)↗

Mu2e - Research & Development

Current efforts are being conducted at Fermi National Laboratory to study potential violations in accepted theory that would otherwise suggest a restructuring of our fundamental understanding of the universe. Mu2e is one of these frontier projects that studies Charged Lepton Flavor Violation (CLFV) which if observed, would suggest physics beyond the Standard Model. Therefore, this note encompasses several projects that contribute to the fruition of Mu2e investigations. Due to the broad range of disciplinary inconsistencies that each project requires, all the work is being presented as a means of justifying contribution to Mu2e. The projects are comprised of a simulation exploring the extinction level of proton pulses after Recycler ring re-bunching by using G4beamline to simulate an 8GeV proton beam interaction with a titanium target, three Cherenkov radiation-based detectors and 2/3-fold and 3/3-fold coincidence rate analysis. Additionally, supplemental work for the implementation of a Micro Telecommunications Computing Architecture (TCA) crate to establish a peak finding algorithm to ensure that the out-of-time beam is less than 10$^{−10}$ fractional level along with single-layer inefficiency analysis on scintillation counters for the Cosmic-Ray Veto (CRV) analysis to ensure the overall inefficiency is 10$^{−4}$. Preliminary results have been achieved for the G4beamline simulation 2/3-fold and 3/3-fold coincidences which are in the order of 10$^{−9}$ and 10$^{−10}$, respectively. Only preliminary results of a triangular counter and four rectangular di-counters for the CRV have been realized. The microTCA crate development is still ongoing.

43 PARTICLE ACCELERATORS↗

rBahadur: efficient simulation of structured high-dimensional genotype data with applications to assortative mating

Existing methods for generating synthetic genotype data are ill-suited for replicating the effects of assortative mating (AM). We propose rb_dplr, a novel and computationally efficient algorithm for generating high-dimensional binary random variates that effectively recapitulates AM-induced genetic architectures using the Bahadur order-2 approximation of the multivariate Bernoulli distribution. The rBahadur R library is available through the Comprehensive R Archive Network at https://CRAN.R-project.org/package=rBahadur.

59 BASIC BIOLOGICAL SCIENCES↗

Unorthodox parallelization for Bayesian quantum state estimation

Quantum state tomography (QST) allows for the reconstruction of quantum states through measurements and some inference technique under the assumption of repeated state preparations. Bayesian inference provides a promising platform to achieve both efficient QST and accurate uncertainty quantification, yet is generally plagued by the computational limitations associated with long Markov chains. In this work, we present a novel Bayesian QST approach that leverages modern distributed parallel computer architectures to efficiently sample a D-dimensional Hilbert space. Using a parallelized preconditioned Crank–Nicholson Metropolis–Hastings algorithm, we demonstrate our approach on simulated data and experimental results from IBM Quantum systems up to four qubits, showing significant speedups through parallelization. Although highly unorthodox in pooling independent Markov chains, our method proves remarkably practical, with validation ex post facto via diagnostics like the intrachain autocorrelation time. We conclude by discussing scalability to higher-dimensional systems, offering a path toward efficient and accurate Bayesian characterization of large quantum systems.

Bayesian inference↗

Analog-to-digital converter based on voltage-controlled superconducting devices

The increasing demand for cryogenic electronics in superconducting and quantum computing systems calls for ultra-energy-efficient data conversion architectures that remain functional at deep cryogenic temperatures. Here, in this work, we present the first design of a voltage-controlled superconducting flash analog-to-digital converter (ADC) based on a voltage-controlled quantum-enhanced Josephson junction field-effect transistor (JJFET). Exploiting its strong gate tunability and transistor-like behavior, the JJFET offers a scalable alternative to conventional current-controlled superconducting devices while aligning naturally with CMOS-style design methodologies. Building on our previously developed Verilog-A compact model calibrated to experimental data, we design and simulate a three-bit JJFET-based flash ADC targeted for integration within cryogenic control and readout circuitry in quantum computing. The core comparator block is realized through careful bias current selection and augmented with a three-terminal nanocryotron to precisely define reference voltages. Cascaded JJFET comparators ensure robust voltage gain, cascadability, and logic-level restoration across stages. Simulation results demonstrate accurate quantization behavior with ultra-low power dissipation, underscoring the feasibility of voltage-driven superconducting mixed-signal circuits. This work establishes a critical step toward unifying superconducting logic and data conversion, paving the way for scalable cryogenic architectures in quantum–classical co-processors, low-power artificial intelligence accelerators, and next-generation energy-constrained computing platforms.

Analog-to-digital converter↗

Hybrid Quantum-Classical Neural Networks

Deep learning is one of the most successful and far-reaching strategies used in machine learning today. However, the scale and utility of neural networks is still greatly limited by the current hardware used to train them. These concerns have become increasingly pressing as conventional computers are soon expected to approach the physical limitations that will slow their performance improvements in the near future. For these reasons, scientists have begun to explore alternative computing platforms, like quantum computers, for training neural networks. In recent years, variational quantum circuits have emerged as one of the most successful approaches to quantum deep learning on noisy intermediate scale quantum devices. We propose a hybrid quantum-classical neural network architecture where each neuron is a variational quantum circuit. We empirically analyze the performance of this hybrid neural network on a series of binary classification data sets using a simulated IBM universal quantum computer and a state-of-the-art IBM universal quantum computer. On the simulated hardware, we observe that the hybrid neural network achieves around 10% higher classification accuracy and 20% better minimization of the cost function than an individual variational quantum circuit. On the quantum hardware, we observe that each model only performs well when the qubit and gate count is sufficiently small.

Arthur, Davis↗

A Hybrid Climate Modeling System Using AI-assisted Process Emulators

This white paper addresses Focus Area II. We advocate developing a hybrid modeling system to improve the understanding of decadal- and longer-scale predictability of high impact water cycle components. This hybrid model combines a partial differential equation (PDE)-based dynamic core with AI/ML based emulators to represent many of the computationally expensive processes in Earth’s climate models. The hybrid modeling system has the potential to exploit emerging graphics processing unit (GPU)-accelerated architectures and allows for the generation of large ensemble (~1000’s) simulations to better characterize the model uncertainty and understand predictability.

58 GEOSCIENCES↗

A compute-bound formulation of Galerkin model reduction for linear time-invariant dynamical systems

This work aims to advance computational methods for projection-based reduced-order models (ROMs) of linear time-invariant (LTI) dynamical systems. For such systems, current practice relies on ROM formulations expressing the state as a rank-1 tensor (i.e., a vector), leading to computational kernels that are memory bandwidth bound and, therefore, ill-suited for scalable performance on modern architectures. This weakness can be particularly limiting when tackling many-query studies, where one needs to run a large number of simulations. This work introduces a reformulation, called rank-2 Galerkin, of the Galerkin ROM for LTI dynamical systems which converts the nature of the ROM problem from memory bandwidth to compute bound. We present the details of the formulation and its implementation, and demonstrate its utility through numerical experiments using, as a test case, the simulation of elastic seismic shear waves in an axisymmetric domain. We quantify and analyze performance and scaling results for varying numbers of threads and problem sizes. In conclusion, we present an end-to-end demonstration of using the rank-2 Galerkin ROM for a Monte Carlo sampling study. We show that the rank-2 Galerkin ROM is one order of magnitude more efficient than the rank-1 Galerkin ROM (the current practice) and about 970 times more efficient than the full-order model, while maintaining accuracy in both the mean and statistics of the field.

97 MATHEMATICS AND COMPUTING↗

Toward Exascale: Overview of Large Eddy Simulations and Direct Numerical Simulations of Nuclear Reactor Flows with the Spectral Element Method in Nek5000

At the beginning of the last decade, Petascale supercomputers (i.e., computers capable of more than 1 petaFLOP) emerged. Now, at the dawn of exascale supercomputing, we provide a review of recent landmark simulations of portions of reactor components with turbulence-resolving techniques that this computational power has made possible. In fact, these simulations have provided invaluable insight into flow dynamics, which is difficult or often impossible to obtain with experiments alone. We focus on simulations performed with the spectral element method, as this method has emerged as a powerful tool to deliver massively parallel calculations at high fidelity by using large eddy simulation or direct numerical simulation. We also limit this paper to constant-property incompressible flow of a Newtonian fluid in the absence of other body or external forces, although the method is by no means limited to this class of flows. We briefly review the fundamentals of the method and the reasons it is compelling for the simulation of nuclear engineering flows. We review in detail a series of Petascale simulations, including the simulations of helical coil steam generators, fuel assemblies, and pebble beds. Even with Petascale computing, however, limitations for nuclear modeling and simulation tools remain. In particular, the size and scope of turbulence-resolving simulations are still limited by computing power and resolution requirements, which scale with the Reynolds number. In the final part of this paper, we discuss the future of the field, including recent advancements in emerging architectures such as GPUbased supercomputers, which are expected to power the next generation of high-performance computers.

computational fluid dynamics↗

PythonFOAM: In-situ data analyses with OpenFOAM and Python

Here, we outline the development of a general-purpose Python-based data analysis tool for OpenFOAM. Our implementation relies on the construction of OpenFOAM applications that have bindings to data analysis libraries in Python. Double precision data in OpenFOAM is cast to a NumPy array using the NumPy C-API and Python modules may then be used for arbitrary data analysis and manipulation on flow-field information. We highlight how the proposed wrapper may be used for an in-situ online singular value decomposition (SVD) implemented in Python and accessed from the OpenFOAM solver PimpleFOAM. Here, 'in-situ' refers to a programming paradigm that allows for a concurrent computation of the data analysis on the same computational resources utilized for the partial differential equation solver. In addition, to demonstrate parallel deployments, we deploy a distributed SVD, which collects snapshot data across the ranks of a distributed simulation to compute the global left singular vectors. Crucially, both OpenFOAM and Python share the same message passing interface (MPI) communicator for this deployment which allows Python objects and functions to exchange NumPy arrays across ranks. Subsequently, we provide scaling assessments of this distributed SVD on multiple nodes of Intel Broadwell and KNL architectures for canonical test cases such as the large eddy simulations of a backward facing step and a channel flow at friction Reynolds number of 395. Finally, we demonstrate the deployment of a deep neural network for compressing the flow-field information using an autoencoder to demonstrate an ability to use state-of-the-art machine learning tools in the Python ecosystem.

97 MATHEMATICS AND COMPUTING↗

Developments in Performance and Portability for MadGraph5_aMC@NLO

Event generators simulate particle interactions using Monte Carlo techniques, providing the primary connection between experiment and theory in experimental high energy physics. These software packages, which are the first step in the simulation worflow of collider experiments, represent approximately 5 to 20% of the annual WLCG usage for the ATLAS and CMS experiments. With computing architectures becoming more heterogeneous, it is important to ensure that these key software frameworks can be run on future systems, large and small. In this contribution, recent progress on porting and speeding up the Madgraph5_aMC@NLO event generator on hybrid architectures, i.e. CPU with GPU accelerators, is discussed. The main focus of this work has been in the calculation of scattering amplitudes and "matrix elements", which is the computational bottleneck of an event generation application. For physics processes limited to QCD leading order, the code generation toolkit has been expanded to produce matrix element calculations using C++ vector instructions on CPUs and using CUDA for NVidia GPUs, as well as using Alpaka, Kokkos and SYCL for multiple CPU and GPU architectures. Performance is reported in terms of matrix element calculations per time on NVidia, Intel, and AMD devices. The status and outlook for the integration of this work into a production release usable by the LHC experiments, with the same functionalities and very similar user interfaces as the current Fortran version, is also described.

Valassi, Andrea↗

Intracardiac Electrical Imaging using the 12-lead ECG: A Machine Learning Approach using Synthetic Data

Current state-of-the-art techniques for non-invasive imaging of cardiac electrical phenomena require voltage recordings from dozens of different torso locations and anatomical models built from expensive medical diagnostic imaging procedures. Here this study aimed to assess if recent machine learning advances could alternatively reconstruct electroanatomical maps at clinically relevant resolutions using only the standard 12-lead electrocardiogram (ECG) as input. To that end, a computational study was conducted to generate a dataset of over 16000 detailed cardiac simulations, which was then used to train neural network (NN) architectures designed to exploit both spatial and temporal correlations in the ECG signal. Analysis over a validation set showed average errors in activation map reconstruction below 1.7 msec over 75 intracardiac locations. Furthermore, phenotypical patterns of activation and the morphology of the activation potential were correctly reconstructed. The approach offers opportunities to stratify patients non-invasively, both retrospectively and prospectively, using metrics otherwise only available through invasive clinical procedures.

59 BASIC BIOLOGICAL SCIENCES↗