Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Structural Simluation Toolkit (SST) v.11.0

The SST provides a parallel framework to perform system simulation of computer architectures to determine their performance and power consumption. Additionally, the SST contains basic models of a computer processor, and interconnect and can connect to an external memory simulator (DRAMSim II). The SST framework provides a simple interface by which other computer simulation models can be combined under a common parallel discrete event-based simulation environment. This allows design exploration of future architectures, analysis of how current computer programs will function on future architectures. The SST provides a parallel discrete event simulation framework, including partitioning and object distribution over MPI. It also provides a mechanism by which components can report their power consumption for analysis.

Rodrigues, ArunF.↗

Development of Steady-State and Dynamic Mass and Energy Constrained Neural Networks for Distributed Chemical Systems Using Noisy Transient Data

The paper presents the development of algorithms for mass and energy constrained neural network models that can exactly conserve the overall mass and energy of distributed chemical process systems, even though the noisy transient data used for optimal model training violate the same. In contrast to approximately satisfying mass and energy balance constraints of a system by soft penalization of objective function, algorithms have been developed for solving equality-constrained nonlinear optimization problems, thus providing the guarantee of exactly satisfying the system mass and energy conservation laws. For developing dynamic mass-energy constrained network models for distributed systems, hybrid series and parallel dynamic-static neural networks have been leveraged. The developed algorithms for solving both the training and forward problems are validated using both steady-state and dynamic data in the presence of various noise characteristics. The developed data-driven algorithms are flexible to exactly satisfy mass and energy balance constraints for dynamic chemical processes if the system holdup information is available. The proposed network structures and algorithms are applied to the development of data-driven lumped and distributed models of an adiabatic superheater/reheater system, a nonisothermal continuous stirred tank reactor, as well as an electrically heated plug-flow reactor system where one form of energy gets transformed to another. It has been observed that the mass-energy constrained neural networks yield a root mean squared error of <1% with respect to the system truth for the case studies evaluated in this work.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multigrid Reduction in Time for Chaotic and Hyperbolic Problems (Final Report)

The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations (PDEs), which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures. The chosen methodology for solving these problems parallel-in-time is multigrid, because multigrid (when it works) is a powerful, optimal, and scalable solver for discretized PDEs. Multigrid is already commonly used in many DOE simulations for scalably and optimally solving space-only PDE problems. The areas of hyperbolic and chaotic problems are chosen because of their relevance to problems of programmatic interest to DOE. However, these problems are also well-known to be difficult for parallelin-time methods, with the most common method, parareal, diverging in many cases. The current state-of-the-art for parallel-in-time at LLNL is the multigrid reduction in time (MGRIT) XBraid package, which also struggles for such problems, while still showing some improvement over parareal. In summary, new methods are needed for an efficient parallel-in-time scheme for hyperbolic and chaotic problems, and this work shall research promising new multigrid methods in this area. In particular, this work shall continue researching the directions from the current collaboration with Dr. Falgout, which are laid out in the work Toward Parallel in Time for Chaotic Dynamical Systems and showed the first known results of a parallel-in-time speedup for a chaotic problem. This work outlines two key improvements to XBraid for chaotic problems, the so-called “theta” and “delta-correction” methods. Here, these two improvements will be implemented in a high-performance but general way in XBraid and explored for more complicated problems. We will additionally research, as time allows, improvements to these techniques, as well as multigrid relaxation techniques based on Least Squares Shadowing (LSS by Wang) and a nonintrusive block tridiagonal solver based on MGRIT, called TriMGRIT.

97 MATHEMATICS AND COMPUTING↗

Multigrid Reduction in Time for Chaotic and Hyperbolic Problems (Final Report)

The coming massive parallelism of exascale computing presents a pressing challenge for the many DOE simulations of time-dependent partial differential equations (PDEs), which typically use traditional sequential time stepping methods. Since this traditional approach is inherently serial, it presents a sequential bottleneck when moving to exascale computing, because future performance gains will come through greater concurrency, not faster clock speeds. Thus, the goal of this work is to research parallelism in time, i.e., methods that compute multiple time values simultaneously, not sequentially. The focus will be on hyperbolic and chaotic problems of interest to DOE, with the goal of enabling scalable simulations of time-dependent hyperbolic and chaotic problems on future architectures. The chosen methodology for solving these problems parallel-in-time is multigrid, because multigrid (when it works) is a powerful, optimal, and scalable solver for discretized PDEs. Multigrid is already commonly used in many DOE simulations for scalably and optimally solving space-only PDE problems. The areas of hyperbolic and chaotic problems are chosen because of their relevance to problems of programmatic interest to DOE. However, these problems are also well-known to be difficult for parallel-in-time methods, with the most common method, parareal, diverging in many cases. The current state of-the-art for parallel-in-time at LLNL is the multigrid reduction in time (MGRIT) XBraid package, which also struggles for such problems, while still showing some improvement over parareal. In summary, new methods are needed for an efficient parallel-in-time scheme for hyperbolic and chaotic problems, and this work shall research promising new multigrid methods in this area. In particular, this work shall continue researching the directions from the current collaboration with Dr. Falgout, which are laid out in the work Toward Parallel in Time for Chaotic Dynamical Systems and showed the first known results of a parallel-in-time speedup for a chaotic problem. This work outlines two key improvements to XBraid for chaotic problems, the so-called “theta” and “delta-correction” methods. Here, these two improvements will be further researched and improved (including with a new relaxation method inspired by on Least Squares Shadowing (LSS)) and explored for more complicated problems.

97 MATHEMATICS AND COMPUTING↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Methods for Method-of-Lines Discretizations of Nonlinear Hyperbolic PDEs and Systems (Final Report)

The work for the subcontract is situated in the area of parallel-in-time integration for hyperbolic partial differential equations (PDEs). Parallel-in-time integration is an active area of research due to its ability to enable faster numerical simulations for applications throughout many areas of science. The work in this subcontract builds on a variety of results that were obtained, as part of the work performed for Subcontract No. B648355, for the Multigrid Reduction-in-Time (MGRIT) method from [1] applied to hyperbolic PDEs. This subcontract extends these results further to more efficient methods and to the case of method-of-lines discretizations for nonlinear hyperbolic PDES and systems of PDEs. The following is a summary of the research performed and results achieved during milestone periods 1, 2 and 3 by the PI (Hans De Sterck) and Postdoctoral Research Associate (Oliver Krzysik), for required tasks 1-4 (as listed in the Statement of Work): Research over the previous year has been split into three main projects: (i) solution of acoustic equation system; (ii) solution of nonlinear scalar hyperbolic PDEs; (iii) solution of nonlinear hyperbolic systems of PDEs.

97 MATHEMATICS AND COMPUTING↗

Ensemble models for circuit topology estimation, fault detection and classification in distribution systems

This paper presents a methodology for simultaneous fault detection, classification, and topology estimation for adaptive protection of distribution systems. The methodology estimates the probability of the occurrence of each one of these events by using a hybrid structure that combines three sub-systems, a convolutional neural network for topology estimation, a fault detection based on predictive residual analysis, and a standard support vector machine with probabilistic output for fault classification. The input to all these sub-systems is the local voltage and current measurements. A convolutional neural network uses these local measurements in the form of sequential data to extract features and estimate the topology conditions. The fault detector is constructed with a Bayesian stage (a multitask Gaussian process) that computes a predictive distribution (assumed to be Gaussian) of the residuals using the input. Since the distribution is known, these residuals can be transformed into a Standard distribution, whose values are then introduced into a one-class support vector machine. The structure allows using a one-class support vector machine without parameter cross-validation, so the fault detector is fully unsupervised. Finally, a support vector machine uses the input to perform the classification of the fault types. All three sub-systems can work in a parallel setup for both performance and computation efficiency. In conclusion, we test all three sub-systems included in the structure on a modified IEEE123 bus system, and we compare and evaluate the results with standard approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

GASNet-EX RMA Communication Performance on Recent Supercomputing Systems

Partitioned Global Address Space (PGAS) programming models, typified by systems such as Unified Parallel C (UPC) and Fortran coarrays, expose one-sided Remote Memory Access (RMA) communication as a key building block for High Performance Computing (HPC) applications. Architectural trends in supercomputing make such programming models increasingly attractive, and newer, more sophisticated models such as UPC++, Legion and Chapel that rely upon similar communication paradigms are gaining popularity. GASNet-EX is a portable, open-source, high-performance communication library designed to efficiently support the networking requirements of PGAS runtime systems and other alternative models in emerging exascale machines. The library is an evolution of the popular GASNet communication system, building upon 20 years of lessons learned. We present microbenchmark results which demonstrate the RMA performance of GASNet-EX is competitive with MPI implementations on four recent, high-impact, production HPC systems. These results are an update relative to previously published results on older systems. The networks measured here are representative of hardware currently used in six of the top ten fastest supercomputers in the world, and all of the exascale systems on the U.S. DOE road map.

Hargrove, Paul H↗

Uncovering I/O demands on HPC platforms: Peeking under the hood of Santos Dumont

High-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues.

97 MATHEMATICS AND COMPUTING↗

Disk Failure Dataset from the Campaign Storage System

This dataset consists of 1,389 disk (HDD) failure events collected from the Campaign storage system at LANL. The Campaign system supported various compute platforms throughout its lifespan, including Cielo, Fire, Ice, and notably, the Trinity supercomputer. Each recorded event includes its detection timestamp (in ISO 8601 format) and details such as its location within the storage system—rack, enclosure, and drive slot number. The data, spanning from May 4, 2021, to July 25, 2023 (2 years, 2 months, and 22 days), represents failure events from the terminal years of Campaign's operational period, accounting for 26% of its total operational time.

97 MATHEMATICS AND COMPUTING↗

Two-Phase Turbulence Statistics from High Fidelity Dispersed Droplet Flow Simulations in a Pressurized Water Reactor (PWR) Sub-Channel with Mixing Vanes

In the dispersed flow film boiling regime (DFFB), which exists under post-LOCA (loss-of-coolant accident) conditions in pressurized water reactors (PWRs), there is a complex interplay between droplet dynamics and turbulence in the surrounding steam. Experiments have accredited particular significance to droplet collision with the spacer-grids and mixing vane structures and their consequent positive feedback to the heat transfer recorded in the immediate downstream vicinity. Enabled by high-performance computing (HPC) systems and a massively parallel finite element-based flow solver—PHASTA (Parallel Hierarchic Adaptive Stabilized Transient Analysis)—this work presents high fidelity interface capturing, two-phase, adiabatic simulations in a PWR sub-channel with spacer grids and mixing vanes. Selected flow conditions for the simulations are informed by the experimental data found in the literature, including the steam Reynolds number and collision Weber number (Wec={40,80}), and are characteristic of the DFFB regime. Data were collected from the simulations at an unprecedented resolution, which provides detailed insights into the continuous phase turbulence statistics, highlighting the effects of the presence of droplets and the comparative effect of different Weber numbers on turbulence in the surrounding steam. Further, axial evolution of droplet dynamics was analyzed through cross-sectionally averaged quantities, including droplet volume, surface area and Sauter mean diameter (SMD). The downstream SMD values agree well with the existing empirical correlations for the selected range of Wec. The high-resolution data repository from the simulations herein is expected to be of significance to guide model development for system-level thermal hydraulic codes.

Saini, Nadish↗

Overcoming Shockley-Queisser limit using halide perovskite platform?

The Intergovernmental Panel on Climate Change (IPCC) reveals that the global temperature has reached its highest level in the last 2,000 years. Development of emissions-free electrification technologies such as photovoltaics (PV) can be a great paramountcy to balance the climate pressure and the growing demand on energy. Although many PV technologies have been demonstrated so far, all the single-junction PVs are still subjected to the well-known efficiency cap of the Shockley-Queisser limit (33.7%), with most of the absorbed solar energy lost into heat. In parallel to delicate system-level designs, such as tandem, multi-junction, photothermal, or up-/down-conversion, we ask the question about the feasibility of overcoming the SQ limit at material level. In this study, we first dissect the origin of the limit, and then, by using the emerging perovskite as the platform, we list several potential pathways (i.e., hot carrier, multi-exciton generation, intermediate band gap, and ferroelectricity that have been newly discovered or can be introduced in the perovskite) to the roadmap of exceeding the SQ limit.

14 SOLAR ENERGY↗

Simulating MADMAX in 3D: Requirements for Dielectric Axion Haloscopes

We present 3D calculations for dielectric haloscopes such as the currently envisioned MADMAX experiment. For ideal systems with perfectly flat, parallel and isotropic dielectric disks of finite diameter, we find that a geometrical form factor reduces the emitted power by up to 30% compared to earlier 1D calculations. We derive the emitted beam shape, which is important for antenna design. We show that realistic dark matter axion velocities of 10 -3 c and inhomogeneities of the external magnetic field at the scale of 10% have negligible impact on the sensitivity of MADMAX. We investigate design requirements for which the emitted power changes by less than 20% for a benchmark boost factor with a bandwidth of 50 MHz at 22 GHz, corresponding to an axion mass of 90 μeV. We find that the maximum allowed disk tilt is 100 μm divided by the disk diameter, the required disk planarity is 20 μm (min-to-max) or better, and the maximum allowed surface roughness is 100 μm (min-to-max). We show how using tiled dielectric disks glued together from multiple smaller patches can affect the beam shape and antenna coupling.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗

Surface electron modulation of metal oxide‐based electrochemical devices by surface additives—linking sensors and fuel cells

Abstract The interaction between ambient oxygen and the metal oxide surface is central to various electrochemical devices. Solid oxide fuel cell cathodes and semiconducting metal oxide gas sensors are two prominent examples. Parallels between these two systems are highlighted in this perspective. In both cases, the presence of foreign surface species has been found to significantly alter the interactions of the metal oxides’ surfaces with oxygen. On the one hand, interactions are hindered due to the presence of certain impurities, such as the degradation of fuel cell cathode by Si‐species that originate from processing or from operating environments. On the other hand, electrochemically active noble metal additives have been intentionally used to tune the sensor response of metal oxides. In the case of electrochemically active additives, the need for operando spectroscopies to elucidate the mechanism is discussed. Electronic coupling of the surface additive with the metal oxide is found to play a central role in both cases. We use these results to demonstrate that insights gained from either field can be effectively applied to the other.

Materials Science↗

Microcontroller-based aerosol jet printer control software

This software is used with a microcontroller to control an aerosol jet printing system. It was written for a 32-bit ARM microcontroller, providing greater processing speed and complexity relative to more conventional low-cost printer controllers. The microcontroller software was written using a real-time operating system (FreeRTOS), which simplifies parallel operation of multiple functions and further customization. It supports simultaneous motion planning, pulse generation for stepper motors, reading encoder feedback, communicating with multiple mass flow controllers, controlling solid state relays, handling a software-based emergency stop, and sending data to the client PC. The program is written in C and communicates with the client PC via USB connection. Sandia National Laboratories is a multi-mission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-0436 O

Secor, EthanBenjamin↗

Hydration Mechanisms in Nanoparticle Interaction and Surface Energetics

Work on the project advanced molecular level understanding, prediction and control of nanoscale hydration in salt solutions under electric stimuli or ionic patterning. Permeation of nanoporous electrodes at preset voltage underlies the function of ultracapacitors. Transitory regulation of wetting in nanoporous media by electric field spans an array of applications in materials, energy storage, and separation sciences. The physically related modulation of nanoparticle solubility by surface charges can significantly extend the range of the nanomaterial applications, improve processing techniques, and potentially alleviate environmental concerns. Pore/solution equilibria, and activated kinetics of liquid gating by nanoconfined electrolytes are challenging problems at the forefront of experimental and theoretical research. Addressing these problems from a molecular perspective required the development of state-of-the-art simulation algorithms in statistical mechanics to capture complex processes in open systems under electric control. Parallel studies of wetting and dispersibility of polar and ionizing particles aim to uncover predictive relations between electrowetting, chemical functionalization, and geometry of nanomaterial particles. By nonequilibrium dynamic modeling, we elucidated electrolyte flow in nanochanels and associated electrokinetic energy conversion. Research on the program provided training opportunities for the next generation of scientists in computational chemistry. Insights, and methods from the project implicate broad segment of researchers in energy and nanosciences, materials and surface chemistry.

74 ATOMIC AND MOLECULAR PHYSICS↗