Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Kinetic Plasma Simulation Capabilities in the MOOSE Framework: Verification of Particle-Particle Collisions

High-fidelity simulations of complex plasma systems allow researchers to gain key insights into and understanding of these systems. To facilitate massively parallel high-fidelity plasma simulations, finite-element-based particle-in-cell capabilities are being developed within the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) based framework called Software for Advanced Large-scale Analysis of MAgnetic confinement for Numerical Design, Engineering & Research (SALAMANDER). While SALAMANDER’s primary objective is modeling edge plasmas and plasma-facing components in fusion devices, the particle-in-cell capabilities being developed are general and will support modeling low-temperature plasmas as well. Previously, collisionless magnetostatic simulation capabilities have been verified with the two-stream and Dorey-Guest-Harris instabilities, and single particle motion. Collisions were implemented using the direct simulation Monte Carlo method, and verification of this capability will be presented here several verification problems: relaxation of a randomly initialized gas to a Maxwellian distribution, Fourier heat flow, and comparison of reaction rates to both analytic calculations and those calculated using a multi-term Boltzmann solver.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Cyber Security Analysis for Nuclear Reactor Control Systems (Final Technical Report)

This project investigated the cyber-security impacts of moving from an all analog, point-to-point, instrumentation and control (I&C) system to a digital I&C system based on Modbus and a shared communication medium. A formalism called a hybrid attack graph was expanded to support the nuclear research reactor system. The hybrid attack graph allows one to check a system for vulnerabilities, in this case cyber-security vulnerabilities, and to document the attack vectors (scenarios) causing those vulnerabilities. In parallel, a simulation of the system was developed to model both the physical reactor parameters and operations, as well as the network interconnects and communications. This simulation platform was modeled on the nuclear research reactor located at Washington State University. The simulation platform provided a sandbox to evaluate and quantify the impact of identified and proposed vulnerabilities in the system and to determine the effectiveness of countermeasures at stopping these attacks. The simulation and hybrid attack graph tools were integrated to provide a streamlined process of generating attack scenarios, playing those scenarios out in the simulation, and then analyzing the results to correlate system state to states in the hybrid attack graph. This process was used to (1) quantify the impact of attack scenarios and (2) to determine if the system moved through the hybrid attack graph as anticipated. The hybrid attack graph tool was extended and customized to produce a tool to automatically identify critical assets (CAs) and critical digital assets (CDAs) as defined by NRC Regulatory Guide 5.71. This tool was verified using the nuclear research reactor at Washington State University. Finally, a series of educational modules covering the findings of the different aspects of this research have been created.

97 MATHEMATICS AND COMPUTING↗

Weak scaling of the contact distance between two fluctuating interfaces with system size

A pair of flat parallel surfaces, each freely diffusing along the direction of their separation, will eventually come into contact. If the shapes of these surfaces also fluctuate, then contact will occur when their centers-of-mass remain separated by a nonzero distance ℓ. An example of such a situation is the motion of interfaces between two phases at conditions of thermodynamic coexistence, and in particular the annihilation of domain wall pairs under periodic boundary conditions. Here we present a general approach to calculate the probability distribution of the contact distance ℓ and determine how its most likely value ℓ* depends on the surfaces' lateral size L. Using the Edward-Wilkinson equation as a model for interfaces, we demonstrate that ℓ* scales weakly with system size, i.e., the dependence of ℓ* on L for both (1+1)- and (2+1)-dimensional interfaces is such that lim L→∞ (ℓ*/L) = 0. In particular, for (2+1)-dimensional interfaces ℓ* is an algebraic function of logL, a result that is confirmed by computer simulations of slab-shaped domains formed under periodic boundary conditions. Overall, this weak scaling implies that such domains remain topologically intact until ℓ becomes very small compared to the lateral size of the interface, contradicting expectations from equilibrium thermodynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Intelligent Partitioning based Fully Parallel AC Security-Constrained Optimal Power Flow

Today’s power grid is becoming more diverse and integrated with high-level distributed energy resources and smart control technologies that is creating a new set of grid management challenges in terms of large-scale, nonlinear, and non-convex problem modeling, complex and time-consuming computation, as well as difficult uncertainty handling. This project focused on solving a challenging multi-period security-constrained generation scheduling problem, which is of great importance for maximizing the social welfare of real-time dispatch, day-ahead market, as well as weekly planning of power systems. Our developed software explored parallel optimization algorithms for complex and realistic power system models, and develop fast, efficient, and robust grid optimization solutions on the high-performance computing platform that will enable increased grid economics, flexibility, resilience, as well as energy security in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

3. Motion Platforms and Kinematic Arrangements

Within a machine, mechanisms and motion are organized in what is known as a “kinematic arrangement,” which helps classify machines based on how they move. The most common kinematic arrangements for additive manufacturing systems are Cartesian, followed by delta, and then six-degrees-of-freedom robotic arms. However, there are a multitude of less common systems, such as the SCARA, polar robots, cable driven parallel robots, mobile platforms, and multi-agent systems. This chapter surveys these various kinematic arrangements to give a broad understanding of the mechanisms underlying motion within additive manufacturing systems. Understanding these mechanisms and their resulting motion provides a framework for discussing path planning for all scales and families of additive manufacturing.

Wang, Peter↗

Simulating the Impact of Dynamic Rerouting on Metropolitan-scale Traffic Systems

The rapid introduction of mobile navigation aides that use real-time road network information to suggest alternate routes to drivers is making it more difficult for researchers and government transportation agencies to understand and predict the dynamics of congested transportation systems. Computer simulation is a key capability for these organizations to analyze hypothetical scenarios; however, the complexity of transportation systems makes it challenging for them to simulate very large geographical regions, such as multi-city metropolitan areas. In this article, we describe enhancements to the Mobiliti parallel traffic simulator to model dynamic rerouting behavior with the addition of vehicle controller actors and vehicle-to-controller reroute requests. The simulator is designed to support distributed-memory parallel execution using discrete event simulation and be scalable on high-performance computing platforms. We demonstrate the potential of the simulator by analyzing the impact of varying the population penetration rate of dynamic rerouting on the San Francisco Bay Area road network. Using high-performance parallel computing, we can simulate a day in the San Francisco Bay Area with 19 million vehicle trips with 50 percent dynamic rerouting penetration over a road network with 0.5 million nodes and 1 million links in less than three minutes. We present a sensitivity study on the dynamic rerouting parameters, discuss the simulator’s parallel scalability, and analyze system-level impacts of changing the dynamic rerouting penetration. Furthermore, we examine the varying effects on different functional classes and geographical regions and present a validation of the simulation results compared to real-world data.

97 MATHEMATICS AND COMPUTING↗

Design and Analysis of 10” Parallel Plate Relief Device

All Cryomodule (CM) and Cryogenic Distribution System (CDS) relieving into Helium Low Pressure (LP) return header, which is connected to compressor suction so, helium can be preserved during small flow relieving event and recirculated to system. However, during worst case scenario, Helium LP header requires a parallel plate relief device to relieve excess pressure from header. To complete the CDS Warm piping header, a new design for a 10 parallel plate relief device is necessary to relieve outside of the tunnel into atmosphere.

Chicas, Kelly↗

AENET–LAMMPS and AENET–TINKER : Interfaces for accurate and efficient molecular dynamics simulations with machine learning potentials

Machine-learning potentials (MLPs) trained on data from quantum-mechanics based first-principles methods can approach the accuracy of the reference method at a fraction of the computational cost. To facilitate efficient MLP-based molecular dynamics and Monte Carlo simulations, an integration of the MLPs with sampling software is needed. Here, we develop two interfaces that link the atomic energy network (ænet) MLP package with the popular sampling packages TINKER and LAMMPS. The three packages, ænet, TINKER, and LAMMPS, are free and open-source software that enable, in combination, accurate simulations of large and complex systems with low computational cost that scales linearly with the number of atoms. Scaling tests show that the parallel efficiency of the ænet–TINKER interface is nearly optimal but is limited to shared-memory systems. The ænet–LAMMPS interface achieves excellent parallel efficiency on highly parallel distributed memory systems and benefits from the highly optimized neighbor list implemented in LAMMPS. We demonstrate the utility of the two MLP interfaces for two relevant example applications: the investigation of diffusion phenomena in liquid water and the equilibration of nanostructured amorphous battery materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

High-Pressure Electrides: A Quantum Chemical Perspective

It has long been assumed that all matter will adopt simple close-packed lattices and become metallic under pressure, in accordance with the Thomas–Fermi–Dirac (TFD) model. However, this model struggles to explain pressure-driven complex structural transitions that have been observed in elements, including sodium, challenging our conventional understanding of compressed matter. Moreover, in stark contrast to the TFD model, first-principles calculations suggest that various elements and compounds become electrides under pressure. Electrides, characterized by concentrations of charge density at interstitial regions, can be thought of as ionic compounds where electrons behave as the anions. Though ambient-pressure molecular electrides have been extensively studied via experiments and computations, high-pressure electrides (HPEs) are not well-understood. The identification and characterization of HPEs have been, to date, based purely on theory, including topological analysis of the electron density and the electron localization function. Here, we review these theoretical analysis tools and suggest guidelines that can be used to classify systems as electrides. Moreover, we describe models used to rationalize the electronic structure of HPEs, drawing parallels with ambient-pressure molecular systems, and encourage the development of experimental techniques that provide evidence for the theoretically calculated charge localization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Geometry, Disorder and Phase Transitions in Topological States of Matter

The quantum Hall effect is the birthplace of topological states of matter, a major theme at the forefront of condensed matter physics in the past two decades. The fractional quantum Hall (FQH) effect revolutionized our understanding of phases of electronic matter. FQH states support exotic fractionally charged excitations that obey Abelian or non-Abelian fractional statistics, which are topological excitations that result from the underlying topological order. During this project, our group discovered a previously unrecognized geometric degree of freedom of incompressible FQH states and studied that for a variety of gapped FQH states. We brought this new concept into direct contact with experiments for the first time by generalizing it to Fermi-liquid states of composite fermions. Using the newly formulated powerful infinite Density Matrix Renormalization Group method, our numerical calculations yielded a parameter free prediction that was found to be in excellent agreement with experimental findings on electron systems in semiconductor heterostructures. In parallel, we performed extensive numerical studies on different, competing phases at various Landau level filling factors, and quantum phase transitions that result from such a competition, e.g. Abelian-non-Abelian phase transitions in bilayer systems. We studied geometrical excitations dubbed “gravitons” (because of their analogy with excitations in the theory of gravitation) and ways to excite and detect them, and explored how they couple with topological excitations. In graphene-based chiral materials, we realized the ability to tune through different incompressible and compressible states in a single Landau level, and found appropriate experimental parameters for the exploration of universal Luttinger liquid behavior not obtained in semiconductor-based electron systems. We showed that topological systems had a very different response from nontopological systems to strong disorder (many-body localization) as well as periodic drives.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing

The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this work, we propose ARENA – an asynchronous reconfigurable accelerator ring architecture as a potential scenario on how the future HPC and data centers will be like. Despite using the coarse-grained reconfigurable arrays (CGRAs) as the substrate platform, our key contribution is not only the CGRA-cluster design itself, but also the ensemble of a new architecture and programming model that enables asynchronous tasking across a cluster of reconfigurable nodes, so as to bring specialized computation to the data rather than the reverse. We presume distributed data storage without asserting any prior knowledge on the data distribution. Hardware specialization occurs at runtime when a task finds the majority of data it requires are available at the present node. In other words, we dynamically generate specialized CGRA accelerators where the data reside. The asynchronous tasking for bringing computation to data is achieved by circulating the task token, which describes the dataflow graphs to be executed for a task, among the CGRA cluster connected by a fast ring network. Evaluations on a set of HPC and data-driven applications across different domains show that ARENA can provide better parallel scalability with reduced data movement (53.9 percent). Compared with contemporary compute-centric parallel models, ARENA can bring on average 4.37× speedup. The synthesized CGRAs and their task-dispatchers only occupy 2.93mm 2 chip area under 45nm process technology and can run at 800MHz with on average 759.8mW power consumption. ARENA also supports the concurrent execution of multi-applications, offering ideal architectural support for future high-performance parallel computing and data analytics systems.

97 MATHEMATICS AND COMPUTING↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

Milestone report for MRT 8479: High Yield Xray Imager Preliminary Design Review

The High Yield Xray Imager (HYXI) is a new NIF target diagnostic system currently under development. The goal of HYXI is to provide high-fidelity, high temporal resolution x-ray imaging capability on high yield NIF implosions at 10MJ and above. The HYXI instrument design concept is based on the combination of two technologies that have been successfully utilized at the NIF on previous instruments, electron pulse-dilation and hybrid-CMOS sensor imaging. The combination of these two techniques will give HYXI sufficient data quality to ascertain differences in hot spot formation dynamics between high and low yield implosions. This information will highlight the critical hot spot conditions needed for ignition and burn. The HYXI design leverages the successful operation of the PDIXI x-ray imager at the NIF on multi-MJ yield shots. Also, a new radiation tolerance CMOS imaging array is being developed to eliminate the significant background noise which limits the data quality of PDIXI. The HYXI Preliminary Design Review was completed at the end of Q4 FY23. The HYXI project is a multi-year effort with a phased approach to be bring up system functionality over time in parallel with the development and fabrication effort of the new CMOS imaging array. Initial time integrated NIF data will be collected in Q2 FY25, first time resolved imaging with Daedalus starting in Q2 FY26 and the final performance qualification of the complete HYXI system is scheduled for Q2 FY27.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Advanced architectures for high-performance quantum networking

As practical quantum networks prepare to serve an ever-expanding number of nodes, there has grown a need for advanced auxiliary classical systems that support the quantum protocols and maintain compatibility with the existing fiber-optic infrastructure. We propose and demonstrate a quantum local area network design that addresses current deployment limitations in timing and security in a scalable fashion using commercial off-the-shelf components. First, we employ White Rabbit switches to synchronize three remote nodes with ultra-low timing jitter, significantly increasing the fidelities of the distributed entangled states over previous work with Global Positioning System clocks. Second, using a parallel quantum key distribution channel, we secure the classical communications needed for instrument control and data management. Therefore, the conventional network that manages our entanglement network is secured using keys generated via an underlying quantum key distribution layer, preserving the integrity of the supporting systems and the relevant data in a future-proof fashion.

97 MATHEMATICS AND COMPUTING↗

Utilizing waste heat in wastewater treatment plants for water desalination: Modeling and Multi-Objective optimization of a Multi-Effect desalination system using Decision Tree Regression and Pelican optimization algorithm

This paper examines the feasibility of using waste heat from wastewater treatment plants (WWTPs) for water desalination. A model was developed to utilize waste heat from the gensets at As Samra WWTP in Jordan, using real data and TRNSYS® software to calculate available waste heat. The desalination process was then modeled with ASPEN PLUS® software, focusing on multi-effect desalination (MED). Both series and parallel configurations for the MED system were compared. The study investigated the effects of system feeding flow rate, feeding pressure, and heat input on productivity, performance ratio, and recovery ratio. The study also introduces a novel optimization technique combining machine learning and modern optimization algorithms to maximize system productivity and performance. Initially, a decision tree regression (DTR) model is developed to establish relationships between key independent variables (flow rate, feed pressure, and heat input) and dependent variables (productivity, performance ratio, and recovery ratio). The Pelican Optimization Algorithm (POA) is then used to identify the optimal values of the independent variables for maximum productivity and performance. The results show that using a series configuration yields a system productivity of 3984.2 kg/hr, a performance ratio of 3.78, and a recovery ratio of 0.991 at a feed flow rate of 4000 kg/hr, feed pressure of 3 bars, and heat input of 719 kW. Optimal productivity (4421 kg/hr), performance ratio (3.81), and recovery ratio (0.851) are achieved at a feed flow rate of 5166 kg/hr, feed pressure of 3.2 bars, and heat input of 794 kW. In conclusion, the techno-economic assessment indicates a levelized cost of water of 1.63 USD/m 3 for parallel configurations and 1.65 USD/m 3 for series configurations, with a payback period of less than two years.

42 ENGINEERING↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Julia as a unifying end-to-end workflow language on the Frontier exascale system

We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy’s first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational kernel on AMD’s MI250x GPUs, (ii) weak scaling up to 4,096 MPI processes/GPUs or 512 nodes, (iii) parallel I/O writes using the ADIOS2 library bindings, and (iv) Jupyter Notebooks for interactive analysis. Results suggest that although Julia generates a reasonable LLVM-IR, a nearly 50% performance difference exists vs. native AMD HIP stencil codes when running on the GPUs. As expected, we observed near-zero overhead when using MPI and parallel I/O bindings for system-wide installed implementations. Consequently, Julia emerges as a compelling high-performance and high-productivity workflow composition language, as measured on the fastest supercomputer in the world.

Godoy, William↗