Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer architecture simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

High Fidelity CFD Simulations Supporting the KP-FHR

Kairos Power, LLC, is developing its version of the Fluoride-cooled High-temperature Reactor, the KP-FHR. The design uses a pebble bed core with fluoride salt as a coolant. The pebbles used in the KP-FHR have a diameter of 4 cm, with a shell fuel region where TRISO particles are embedded. A Pebble bed core design is adopted by several Gen IV reactors, They boast many benefits, such as fuel integrity, highly efficient heat transfer, and passive safety. However, it is challenging to accurately predict temperature and flow inside a pebble bed. Traditional approaches use the porous media model, which regards the pebble bed as a continuous medium, but with different temperature fields representing different levels, such as the fluid temperature, pebble surface temperature, and pebble center temperature. Empirical heat transfer correlations are adopted to calculate the heat transfer coefficient between different phases. However, empirical correlations are usually validated with experimental data, which usually lacks detail inside the pebble bed. The available experimental data is also generally at a high Reynolds number, which falls outside of the conditions of KP-FHR. Explicit computational fluid dynamics (CFD) simulations of randomly packed pebble beds have only become feasible recently. This is thanks to the rapid development of computational power and scalable algorithms. In this work, we used the Spectral Element Method (SEM) CFD code NekRS to simulate the randomly packed pebble bed in a cylindrical container. NekRS, which is the GPU variant of Nek5000, but refactored to utilize the computational power of GPUs using the OCCA library to run on hybrid architecture high performance computing systems. It was initially developed with the libParamunal library, but truncated and tuned for large-scale turbulence simulation. As a result, the SEM reaches higher precision with the same degrees of freedom by using a high-order Lagrange polynomial basis distributed on Gauss-Lobatto-Legendre quadrature inside each element, compared to lower-order methods, such the Finite Volume Method and Finite Element Method. The report is divided into five parts. We start with a general discussion of the pebble bed reactor, along with a specific investigation into the KP-FHR. The second part presents the numerical methodology. In the third part, we study a modular pebble bed with 1741 pebbles in a container of 7 pebble-diameter radius. Beyond LES simulations done by NekRS, we also leveraged the thermal radiation model in OpenFOAM to study heat transfer under no-forced-flow scenarios. Then, in the fourth part we simulated a pebble bed similar to the size of the Hermes Test Reactor. The total number of pebbles is in these simulations is 34,374. The container radius is 14 pebble-diameters. Finally, the report concludes in part five, with a discussion of future work.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Development of a machine learning model for polyethylene pyrolysis using a detailed reaction mechanism

Waste plastics have recently received significant attention as the issue of waste generation continues to increase. Thermal conversion processes, such as pyrolysis and gasification, are attractive potential technologies for utilizing waste plastics and reducing overall waste generation. Efficient utilization of plastics requires a detailed understanding of the conversion process such as pyrolysis and gasification. However, a mechanistic understanding of these processes lead to large and complex kinetic schemes that are not suited for large-scale and long-time simulation methods. Currently, most modeling approaches for pyrolysis and gasification rely on globally lumped, simplified kinetic schemes that provide results that are classified by their product type and not individual species, which limit the level of fidelity achieved via modeling. A machine learning (ML) model has been developed for the primary reactions of high-density polyethylene (HDPE) in an attempt to increase computational efficiency while still maintaining a high level of detail and accuracy. The ML model is trained on a detailed reaction mechanism containing 42 total species and 737 chemical reactions. A DeepONet branch and trunk architecture was adopted to train the model using time-steps relevant to computational fluid dynamics simulations. The ML used physics-informed loss functions to ensure mass conservation. The surrogate model has been deployed in simple MFiX CFD simulations, single particle and an experimental drop tube reactor, and has shown promising performance compared to the original scheme.

Houston, Ross↗

Probabilistic Neural Computing with Stochastic Devices

Abstract The brain has effectively proven a powerful inspiration for the development of computing architectures in which processing is tightly integrated with memory, communication is event‐driven, and analog computation can be performed at scale. These neuromorphic systems increasingly show an ability to improve the efficiency and speed of scientific computing and artificial intelligence applications. Herein, it is proposed that the brain's ubiquitous stochasticity represents an additional source of inspiration for expanding the reach of neuromorphic computing to probabilistic applications. To date, many efforts exploring probabilistic computing have focused primarily on one scale of the microelectronics stack, such as implementing probabilistic algorithms on deterministic hardware or developing probabilistic devices and circuits with the expectation that they will be leveraged by eventual probabilistic architectures. A co‐design vision is described by which large numbers of devices, such as magnetic tunnel junctions and tunnel diodes, can be operated in a stochastic regime and incorporated into a scalable neuromorphic architecture that can impact a number of probabilistic computing applications, such as Monte Carlo simulations and Bayesian neural networks. Finally, a framework is presented to categorize increasingly advanced hardware‐based probabilistic computing technologies.

Misra, Shashank↗

Scalable Deep Learning-Based Microarchitecture Simulation on GPUs

Cycle-accurate microarchitecture simulators are essential tools for designers to architect, estimate, optimize, and manufacture new processors that meet specific design expectations. However, conventional simulators based on discrete-event methods often require an exceedingly long time-to-solution for the simulation of applications and architectures at full complexity and scale. Given the excitement around wielding the machine learning (ML) hammer to tackle various architecture problems, there have been attempts to employ ML to perform architecture simulations, such as Ithemal and SimNet. However, the direct application of existing ML approaches to architecture simulation may be even slower due to overwhelming memory traffic and stringent sequential computation logic. This work proposes the first graphics processing unit (GPU)-based microarchitecture simulator that fully unleashes the potential of GPUs to accelerate state-of-the-art ML-based simulators. First, considering the application traces are loaded from central processing unit (CPU) to GPU for simulation, we introduce various designs to reduce the data movement cost between CPUs and GPUs. Second, we propose a parallel simulation paradigm that partitions the application trace into sub-traces to simulate them in parallel with rigorous error analysis and effective error correction mechanisms. Combined, this scalable GPU-based simulator outperforms by orders of magnitude the traditional CPU-based simulators and the state-of-the-art ML-based simulators, i.e., SimNet and Ithemal.

97 MATHEMATICS AND COMPUTING↗

5G Enabled Transformative Co-design and Co-simulation Framework for Grid Decarbonization and Modernization

5G is a breakthrough technology to enable a fully mobile and connected society, and a 5G-enabled digital continuum will be one of the critical foundations for clean energy economy and grid modernization. The clean energy transformation not only requires innovative manufacturing, algorithm, application, and cross-domain co-simulation, but also a systematic evaluation and comprehensive support on the microelectronic and architectural computing customization to enable and maximize the high renewable penetration in the future power grid. The co-design of decarbonized grid, reliable yet efficient communication, and affordable yet ubiquitous computing will unlock the potential of almost any available tools for the undergoing clean energy transformation. 5G is also an opportunity to rethink the paradigm of infrastructure planning, as computing will be one of the core services for society alongside electricity and communication services.

5G↗

Scalable deep learning for watershed model calibration

Watershed models such as the Soil and Water Assessment Tool (SWAT) consist of high-dimensional physical and empirical parameters. These parameters often need to be estimated/calibrated through inverse modeling to produce reliable predictions on hydrological fluxes and states. Existing parameter estimation methods can be time consuming, inefficient, and computationally expensive for high-dimensional problems. In this paper, we present an accurate and robust method to calibrate the SWAT model (i.e., 20 parameters) using scalable deep learning (DL). We developed inverse models based on convolutional neural networks (CNN) to assimilate observed streamflow data and estimate the SWAT model parameters. Scalable hyperparameter tuning is performed using high-performance computing resources to identify the top 50 optimal neural network architectures. We used ensemble SWAT simulations to train, validate, and test the CNN models. We estimated the parameters of the SWAT model using observed streamflow data and assessed the impact of measurement errors on SWAT model calibration. We tested and validated the proposed scalable DL methodology on the American River Watershed, located in the Pacific Northwest-based Yakima River basin. Our results show that the CNN-based calibration is better than two popular parameter estimation methods (i.e., the generalized likelihood uncertainty estimation [GLUE] and the dynamically dimensioned search [DDS], which is a global optimization algorithm). For the set of parameters that are sensitive to the observations, our proposed method yields narrower ranges than the GLUE method but broader ranges than values produced using the DDS method within the sampling range even under high relative observational errors. The SWAT model calibration performance using the CNNs, GLUE, and DDS methods are compared using R 2 and a set of efficiency metrics, including Nash-Sutcliffe, logarithmic Nash-Sutcliffe, Kling-Gupta, modified Kling-Gupta, and non-parametric Kling-Gupta scores, computed on the observed and simulated watershed responses. The best CNN-based calibrated set has scores of 0.71, 0.75, 0.85, 0.85, 0.86, and 0.91. The best DDS-based calibrated set has scores of 0.62, 0.69, 0.8, 0.77, 0.79, and 0.82. The best GLUE-based calibrated set has scores of 0.56, 0.58, 0.71, 0.7, 0.71, and 0.8. The scores above show that the CNN-based calibration leads to more accurate low and high streamflow predictions than the GLUE and DDS sets. Our research demonstrates that the proposed method has high potential to improve our current practice in calibrating large-scale integrated hydrologic models.

54 ENVIRONMENTAL SCIENCES↗

Large-scale physically accurate modelling of real proton exchange membrane fuel cell with deep learning

Proton exchange membrane fuel cells, consuming hydrogen and oxygen to generate clean electricity and water, suffer acute liquid water challenges. Accurate liquid water modelling is inherently challenging due to the multi-phase, multi-component, reactive dynamics within multi-scale, multi-layered porous media. In addition, currently inadequate imaging and modelling capabilities are limiting simulations to small areas (<1 mm 2 ) or simplified architectures. Herein, an advancement in water modelling is achieved using X-ray micro-computed tomography, deep learned super-resolution, multi-label segmentation, and direct multi-phase simulation. The resulting image is the most resolved domain (16 mm 2 with 700 nm voxel resolution) and the largest direct multi-phase flow simulation of a fuel cell. This generalisable approach unveils multi-scale water clustering and transport mechanisms over large dry and flooded areas in the gas diffusion layer and flow fields, paving the way for next generation proton exchange membrane fuel cells with optimised structures and wettabilities.

25 ENERGY STORAGE↗

Mesoscopic Modeling and Rapid Simulation of Incremental Changes in Epidemic Scenarios on GPUs

In simulation-based studies and analyses of epidemics, a major challenge lies in resolving the conflict between fidelity of models and the speed of their simulation. Another related challenge arises in dealing with the large number of what–if scenarios that need to be explored. Here, we describe new computational methods that together provide an approach to dealing with both challenges. A mesoscopic modeling approach is described that strikes a middle ground between macroscopic models based on coupled differential equations and microscopic models built on fine-grained behaviors at the individual entity level. The mesoscopic approach offers the ability to incorporate complex compositions of multiple layers of dynamics even while retaining the potential for aggregate behaviors at varying levels. It also is an excellent match to the accelerator-based architectures of modern computing platforms in which graphical processing units (GPUs) can be exploited for fast simulation via the parallel execution mode of single instruction multiple thread (SIMT). The challenge of simulating a large number of scenarios is addressed via a method of sharing model state and computation across a tree of what–if scenarios that are localized, incremental changes to a large base simulation. A combination of the mesoscopic modeling approach and the incremental what–if scenario tree evaluation has been implemented in the software on modern GPUs. Synthetic simulation scenarios are presented to demonstrate the computational characteristics of our approach. Results from the experiments with large population data, including USA, UK, and India, illustrate the modeling methodology and computational performance on thousands of synthetically generated what–if scenarios. Execution of our implementation scaled to 8192 GPUs of supercomputing platforms demonstrates the ability to rapidly evaluate what–if scenarios several orders of magnitude faster than the conventional methods.

97 MATHEMATICS AND COMPUTING↗

Celeritas: Accelerating Geant4 with GPUs

Celeritas [1] is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications (specifically, High Lumi- nosity Large Hadron Collider (HL-LHC) simulation) on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a GPU-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.3, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling integration of Celeritas into high energy physics (HEP) experimental frameworks such as CMSSW. On the Summit supercomputer, Celeritas performs EM physics between 6 and 32 faster using the machine’s Nvidia GPUs compared to using only CPUs. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.8–2.3× on GPU and 1.2× on CPU.

Johnson, Seth R.↗

FPGA-Accelerated Range-Limited Molecular Dynamics

Long timescale Molecular Dynamics (MD) simulation of small molecules is crucial in drug design and basic science. To accelerate a small data set that is executed for a large number of iterations, high-efficiency is required. Recent work in this domain has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors. The problem addressed here is that, as the number of on-chip processors has increased from fewer than 10 into the hundreds, previous intra-chip routing solutions are no longer viable. We find, however, that through various design innovations, high efficiency can be maintained. These include replacing the previous broadcast networks with ring-routing and then augmenting the rings with out-of-order and caching mechanisms. Others are adding a level of hierarchical filtering and memory recycling. Two novel optimized architectures emerge, together with a number of variations. These are validated, analyzed, and evaluated. We find that in the domain of interest speed-ups over GPUs are achieved. Finally, the potential impact is that this system promises to be the basis for scalable long timescale MD with commodity clusters.

97 MATHEMATICS AND COMPUTING↗

Technical Characterization and Benefit Evaluation of 5G-Enabled Grid Data Transport and Applications

This report summarizes the Year 1 work of Pacific Northwest National Laboratory’s (PNNL’s) 5G Fabricated Resource and Asset Management Encompassment for energy infrastructure (Energy FRAME) project funded by the Department of Energy Office of Science’s Advanced Scientific Computing Research Program. 5G is a breakthrough technology that enables a fully mobile and connected society, and a 5G-enabled digital continuum will be one of the critical foundations for a clean energy economy and grid modernization. In collaboration with PNNL’s Advanced Wireless Communication team and Center for Advanced Technology Evaluation team, the project team has been evaluating the system performance of 5G testbeds in the PNNL 5G Innovation Studio, and has formulated a co-simulation test case of power system transmission, distribution, and communication (T&D&C) networks considering 5G technology and high penetration of distributed energy resources. The methodology developed in the 5G Energy FRAME project can be customized to fit different future grid scenarios to evaluate multiple (dynamic) configurations (computing, sensing, communication, environment) for different stakeholders. In summary, our main technical highlights in project Year 1 are as follows: 1) Technical characterization of 5G standalone architectures, 2) Formulation of co-simulation test case of T&D&C networks embedded with 5G, 3) Initial benefit evaluation of 5G communication platform for grid use cases, and 4) Additional extended discussions on edge computing, artificial intelligence and machine learning, and high-performance computing and cloud computing adoptions. In addition, a collection of system performance data is shared through the publicly available weblink, https://www.pnnl.gov/projects/5g-energy-frame/publications

24 POWER TRANSMISSION AND DISTRIBUTION↗

Parallel Multigrid in Time and Space for Extreme-Scale Computational Science

The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations, which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of programmatic interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures.

97 MATHEMATICS AND COMPUTING↗

Automation for Grid Interconnected Laboratory Emulation

As computational capabilities improve, digital twins are becoming vital for evaluating equipment realistically in laboratories. This paper outlines a digital twin architecture for the power grid, employing electromagnetic transient (EMT) simulation alongside real-time simulation of power hardware and hierarchical control systems. EMT simulation occurs on a high-performance computing server for scalability. Additionally, the paper describes a workflow and real-time data streaming software facilitating connectivity among EMT simulation, hierarchical control systems, and power hardware. This software enables automated equipment connectivity in the laboratory for realistic evaluations, aiding in identifying necessary upgrades for both equipment control systems and the power grid.

Marthi, Phani Ratna Vanamali [ORNL] (ORCID:0000000↗

Evaluation of data driven low-rank matrix factorization for accelerated solutions of the Vlasov equation

Low-rank methods have shown success in accelerating simulations of a collisionless plasma described by the Vlasov equation, but still rely on computationally costly linear algebra every time step. We propose a data-driven factorization method using artificial neural networks, specifically with convolutional layer architecture, that trains on existing simulation data. At inference time, the model outputs a low-rank decomposition of the distribution field of the charged particles, and we demonstrate that this step is faster than the standard linear algebra technique. Numerical experiments show that the method achieves comparable reconstruction accuracy for interpolation tasks, generalizing to unseen test data in a manner beyond just memorizing training data; patterns in factorization also inherently followed the same numerical trend as those within algebraic methods (e.g., truncated singular-value decomposition). However, when training on the first 70% of a time-series data and testing on the remaining 30%, the method fails to meaningfully extrapolate. Despite this limiting result, the technique may have benefits for simulations in a statistical steady-state or otherwise showing temporal stability. These results suggest that while the model offers a computationally efficient alternative for datasets with temporal stability, its current formulation is best suited for interpolation rather than for predicting future states in time-evolving systems. This study thus lays the groundwork for further refinement of neural network-based approaches to low-rank matrix factorization in high-dimensional plasma simulations.

97 MATHEMATICS AND COMPUTING↗

On the use of a multigrid-reduction-in-time algorithm for multiscale convergence of turbulence simulations

Simulations of turbulent flow present challenges in terms of accuracy and affordability on modern highly-parallel computer architectures. A multigrid-reduction-in-time algorithm is used to provide a framework for separately evolving different scales of turbulence and for parallelizing the temporal domain, thereby increasing the concurrency. It is hypothesized that the space–time locality of the small scales of turbulence can be used to circumvent difficulties in applying temporal multigrid to flows dominated by inertial physics. For algorithms that fall well short of spectral accuracy (fourth-order is used in this work) attention must be paid to the accuracy of features on scales transferred between multigrid levels. Numerical experiments were performed using implicit large-eddy simulation. Results from applying the approach to an infinite-Reynolds number Taylor–Green flow and a double-shear flow at a Reynolds number of 11650 provide strong evidence that the approach has merit. The multigrid-reduction-in-time framework can be used to parallelize the temporal domain of a high-Reynolds-number turbulent flow and permit independent convergence of different scales. Establishing this foundation allows for future research in reducing the wall-clock time to solve turbulent flows while retaining the same accuracy as sequential solvers. In conclusion, current performance results from parallelizing the temporal domain are not competitive with those from sequential-in-time methods.

97 MATHEMATICS AND COMPUTING↗

Real-space Kohn–Sham density functional theory for complex energy applications

Real-space Kohn-Sham density functional theory (real-space KS-DFT) enables large-scale electronic structure simulations that is particularly well-suited for the modern high-performance computing (HPC) architectures. This feature article reviews its theoretical foundations, highlights the algorithmic advances and recent developments, and showcases applications in complex nano systems. We aim to provide a perspective on the trajectory of real-space KS-DFT as an emerging tool for computational chemistry and materials science in the exascale era.

Zhang, Zeyi↗

Multi-Core Microcontroller Hardware In the Loop System for Electric Machine Control

Hardware in the Loop (HIL) is a simulation technique used to reduce the software development cycle and test control systems in a non-destructive environment. This work describes a cost effective HIL simulator on a dual core microcontroller in which one core acts as a controller and the other emulates the system under control. The emulator runs one step per Pulse Width Modulation (PWM) period in real time. To handle the computational burden and prioritize execution of simulation and control tasks, an interrupt-based software architecture with task prioritization has been developed. As a demonstration, the HIL has been implemented on a Texas Instruments TMS320F28379D dual core microcontroller, which emulates a Permanent Magnet Synchronous Machine (PMSM) with resolver feedback. Hardware peripherals are developed and tested concurrently with the control system, providing higher confidence in the software. By using the peripherals in the HIL development, the controller exercises either the HIL emulation or a pin compatible PMSM testbench. To quantify performance and validate the processor based emulator, the HIL results are compared to the preexisting testbench for accuracy benchmarking at no-load and under load for a range of operating points.

33 ADVANCED PROPULSION SYSTEMS↗

Implementation of Detailed Polyethylene Pyrolysis Kinetics into CFD Simulations using Machine Learning

Municipal solid waste (MSW) and waste plastics have received significant attention due to the issues of waste generation and storage, as well as their potential as an energy resource. High-density polyethylene (HDPE) makes up a large portion of plastic waste and has been the subject of several conversion studies. However, the mechanisms associated with converting HDPE through pyrolysis and gasification are extensive and complex making them difficult to implement into high-fidelity computational fluid dynamic (CFD) simulations. For this project, a primary pyrolysis mechanism containing 42 unique species and 737 heterogeneous reactions was used to generate kinetic data over a range of operating conditions. A machine learning (ML) model was developed to replicate the results of the detailed pyrolysis mechanism while significantly increasing the computational efficiency. A deep operator network (DeepONet) architecture was adopted to train the model using time steps relevant to CFD simulations. The ML used physics-based loss functions to ensure mass conservation. The ML model has been deployed in simple MFiX CFD simulations, single particle, and an experimental drop tube reactor, and has shown promising performance compared to the original scheme.

Houston, Ross↗