Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Neptunium mononitride as a target material for Pu-238 production

Deep space exploration requires specialized sources for both thermal and power applications. Radioactive decay heat of plutonium-238 (238Pu) provides these sources in the form of radioisotope thermoelectric generators (RTGs). The 238 Pu is produced via neutron capture reaction involving neptunium-237 ( 237 Np) target material. Continual optimization of 237 Np target materials and evaluation of potential alternative targets for production of 238 Pu RTGs are advantageous for meeting ongoing space power system resource requirements. Current production of 238 Pu for RTGs for the United States space program utilizes neptunium dioxide ( 237 NpO 2 ) targets; however, the use of neptunium mononitride ( 237 NpN) presents an opportunity to increase the mass of 237 Np per target compared to the dioxide form, as well as increase the thermal conductivity of the target. To assess the viability of a 237 NpN target material, the material chemistry must be thoroughly evaluated, including synthesis methods and dissolution and reprocessing schemes. This review presents a summary of synthesis pathways for 237 NpN based on published literature on actinide mononitrides. Specific literature on 237 NpN is limited, necessitating evaluation of other actinide systems to gather parallels. This suggests a need for additional experimental studies on 237 NpN. A particular limitation in the existing literature is a lack of information on the differences in material characteristics, such as morphology, particle size, and trace chemical impurities, as a function of synthesis method. These parameters may affect subsequent reactor performance or dissolution of irradiated targets. The evaluation of existing literature is presented with a focus on the efficacy of 237 NpN targets for 238 Pu production.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Performance and Feature Improvements in Parareal-based Power System Dynamic Simulation

In recent years, a novel Parareal-based approach has been developed for fast transient simulations of large power system interconnections. Parareal belongs to the class of Parallel-in-time algorithms for solution of systems of differential-algebraic equations in parallel over an interval of time. The selection of a reasonably fast and accurate coarse solution is crucial to improve the performance of Parareal algorithm. Semi-analytical solution methods are one promising approach to achieve this goal. They have been investigated, and some preliminary results are presented here. In addition, Parareal-based simulator has been expanded to enable co-simulation with OpenDSS, a widely used open-source distribution system simulator. Preserving the parallel nature of the Parareal approach and taking advantage of the parallel capabilities of the latest versions of OpenDSS, each distribution system can be solved in their entirety on different processors in parallel within the main Parareal simulator. This paper also presents the structure of the transmission and distribution co-simulation and some results with different dynamic models of inverter-based resources in the distribution systems.

Park, Byungkwon↗

Grafted nickel-promoter catalysts for dry reforming of methane identified through high-throughput experimentation

High-throughput synthesis of a series of monometallic and bimetallic catalysts (45 bimetallic and 50 monometallic samples) consisting of nickel and one of nine different metal promoters (B, Co, Cu, Fe, Mg, Mn, Sn, V and Zn) supported on one of six different metal oxides alumina, ceria, magnesia, silica and titania) is carried out via organometallic grafting using a robotic platform. The catalysts are evaluated for their activity and selectivity for the dry reforming of methane at a feed ratio of CH 4 :CO 2 of 1 at 650–800 °C in a parallel flow reactor system. The type of oxide support prevails over the type of additive for both catalyst activity and stability. On Al 2 O 3 and MgO, Fe was found to be the best promoter; on SiO 2 , Cu is the best promoter at 700 °C and higher, while on TiO 2 , Mn is found to enhance the conversion at 800 °C. On CeO 2 , all additives except Fe have beneficial effects. Twenty-five catalysts show > 90% methane conversion with ten catalysts showing > 95% conversion at 800 °C with the H 2 :CO ratios ranging from 0.8 to 1.2. Amongst the ten highest performers, NiFe/Al 2 O 3 and NiFe/MgO are more active than Ni/Al 2 O 3 and Ni/MgO, respectively and were stable over a period of 25 h at 800 °C. Characterization on the as-prepared samples reveals highly dispersed phase, while after reduction in H 2 , highly dispersed and reduced nickel particles up to 10 nm are formed. The particles do not increase in size under dry reforming reaction conditions at 800 °C. An increased hydrogen consumption observed during H 2 -TPR of the nickel particles is positively correlated with methane conversion for Al 2 O 3 -based catalysts. The resistance to deactivation by coking and variation in coke structure are investigated by spectroscopic and microscopic methods to identify the relationship between metal promoters, alloy formation, and type of surface carbon deposits. Carbon whiskers were observed on the ten selected spent samples and are preferentially deposited on Ni rather than on the promoters. Carbon nanotube formation and metal particle removal from support were not observed to cause deactivation while amorphous carbon formation was clearly linked to catalyst deactivation, as amorphous carbon could encapsulate nickel, either on the support or at the end of the carbon nanotube. Furthermore, the organometallic grafting technique is an efficient and suitable technique for synthesizing highly dispersed and homogeneous phases which lead to high conversion and high durability for dry reforming of methane.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electromechanical memcapacitor model offering biologically plausible spiking.

In this article, we introduce a new nanoscale electromechanical device - a leaky memcapacitor - and show that it may be useful for the hardware implementation of spiking neurons. The leaky memcapacitor is a movableplate capacitor that becomes quite conductive when the plates come close to each other. The equivalent circuit of the leaky memcapacitor involves a memcapacitive and memristive system connected in parallel. In the leaky memcapacitor, resistance and capacitance depend on the same internal state variable, which is the displacement of the movable plate. We have performed a comprehensive analysis showing that several types of spiking observed in biological neurons can be implemented with the leaky memcapacitor. Significant attention is paid to the dynamic properties of the model. As in leaky memcapacitors the capacitive, leaking resistive, and reset functionalities are implemented naturally within the same device structure, their use will simplify the creation of spiking neural networks.

Zhang, Zixi↗

Translational research in the MPICH project

The MPICH project is an example of translational research in computer science before that term was well known or even coined. The project began in 1992 as an effort to develop a portable, high-performance implementation of the emerging Message-Passing Interface (MPI) Standard. It has enabled the widespread adoption of MPI as a way to write scalable parallel applications on systems of all sizes including upcoming exascale supercomputers. In this paper, we describe how the translational research process was used in MPICH, how that led to its success, the challenges encountered and lessons learned, and how the process could be applied to other similar projects.

97 MATHEMATICS AND COMPUTING↗

BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC Applications

Testing is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Finally, our results show that BeeSwarm can be used for scalability and performance testing of HPC applications.

97 MATHEMATICS AND COMPUTING↗

Secondary Use-Plug-and-Play Energy Storage System Composed of Multiple Energy Storage Technologies

Low-cost, grid-connectable energy storage technologies represent a significant challenge for the electric grid of the future. Energy storage technologies are in rapid development with targets to reduce the storage medium cost. However, a significant cost to deployment also comes in the integration. This paper presents the development of a plug-and-play system for supporting secondary use multiple battery systems into a single grid connectable unit. Results of the system design are demonstrated in a controller hardware in the loop (CHIL) platform. Simulations of two energy storage systems operating in parallel and dispatched optimally are presented.

Starke, Michael↗

PLEXUS: A Pattern-Oriented Runtime System Architecture for Resilient Extreme-Scale High-Performance Computing Systems

For high-performance computing (HPC) system designers and users, meeting the myriad challenges of next-generation exascale supercomputing systems requires rethinking their approach to application and system software design. Among these challenges, providing resiliency and stability to the scientific applications in the presence of high fault rates requires new approaches to software architecture and design. As HPC systems become increasingly complex, they require intricate solutions for detection and mitigation for various modes of faults and errors that occur in these large-scale systems, as well as solutions for failure recovery. These resiliency solutions often interact with and affect other system properties, including application scalability, power and energy efficiency. Therefore, resilience solutions for HPC systems must be thoughtfully engineered and deployed.In previous work, we developed the concept of resilience design patterns, which consist of templated solutions based on well-established techniques for detection, mitigation and recovery. In this paper, we use these patterns as the foundation to propose new approaches to designing runtime systems for HPC systems. The instantiation of these patterns within a runtime system enables flexible and adaptable end-to-end resiliency solutions for HPC environments. The paper describes the architecture of the runtime system, named Plexus, and the strategies for dynamically composing and adapting pattern instances under runtime control. This runtime-based approach enables actively balancing the cost-benefit trade-off between performance overhead and protection coverage of the resilience solutions. Based on a prototype implementation of PLEXUS, we demonstrate the resiliency and performance gains achieved by the pattern-based runtime system for a parallel linear solver application.

Hukerikar, Saurabh↗

libEnsemble: A complete Python toolkit for dynamic ensembles of calculations

Almost all science and engineering applications eventually stop scaling: their runtime no longer decreases as available computational resources increase. Therefore, many applications will struggle to efficiently use emerging extreme-scale high-performance, parallel, and distributed systems. libEnsemble is a complete Python toolkit and workflow system for intelligently driving ensembles of experiments or simulations at massive scales. It enables and encourages multidisciplinary design, decision, and inference studies portably running on laptops, clusters, and supercomputers.

97 MATHEMATICS AND COMPUTING↗

In Situ High-Temperature Ultrafast Electron Diffraction through Integrated Furnace and MEMS Platforms

Temperature fundamentally governs phase stability, defect evolution, and transport behavior in materials. Despite its central role, direct measurements of structural evolution at elevated temperatures on ultrafast timescales have remained limited. Here, we report the design, integration, and validation of 2 complementary in situ heating platforms that substantially extend the thermal operating range of ultrafast electron diffraction (UED). A compact furnace-type heating stage enables stable diffraction measurements from room temperature to 800 K with ±0.1 K stability under ultrahigh vacuum, achieved through multi-sensor feedback control, dual air-cooling channels, and a thermally isolated motion stage. In parallel, a microelectromechanical system (MEMS)-based heating platform provides rapid thermal response and access to extreme temperatures ≥1,373 K with ±0.1 K stability over hundreds-micrometer regions while supporting simultaneous electrical biasing for electrothermal coupling studies. Absolute temperature calibration is established using diffraction-based thermometry via aluminum lattice expansion and independently validated through in situ melting of bismuth thin films. UED measurements further reveal pronounced temperature-dependent nonequilibrium lattice dynamics in bismuth, including modifications to electron–phonon coupling and Debye–Waller behavior, as well as enhanced ultrafast diffuse scattering in aluminum at elevated temperatures. Together, these developments establish a practical framework for quantitative, time-resolved studies of temperature-driven kinetics and nonequilibrium structural dynamics under extreme thermal environments.

Bai, Qianqian [Chinese Academy of Sciences (CAS), ↗

Error containment for enabling local checkpoint and recovery

Various embodiments include a parallel processing computer system that detects memory errors as a memory client loads data from memory and disables the memory client from storing data to memory, thereby reducing the likelihood that the memory error propagates to other memory clients. The memory client initiates a stall sequence, while other memory clients continue to execute instructions and the memory continues to service memory load and store operations. When a memory error is detected, a specific bit pattern is stored in conjunction with the data associated with the memory error. When the data is copied from one memory to another memory, the specific bit pattern is also copied, in order to identify the data as having a memory error.

Cherukuri, Naveen↗

Leveraging the Run 3 experience for the evolution of the ATLAS software-based readout towards HL-LHC

The High-Luminosity Large Hadron Collider (HL-LHC), scheduled to start operating in 2030, aims to increase the instantaneous luminosity by a factor of 10 compared to the LHC. To match this increase, the ATLAS experiment has been implementing a major upgrade program divided into two phases. The first phase (Phase-I), completed in 2022, introduced new trigger and detector systems that have been used during the Run 3 data taking period which began in July 2022. These systems have been used in conjunction with the new Data Acquisition (DAQ) Readout system, based on a software application called Software Readout Driver (SW ROD). SW ROD receives and aggregates data from the front-end electronics via the Front-End Link eXchange (FELIX) system and passes aggregated data fragments to the High-Level Trigger (HLT) system. During Run 3, SW ROD operates in parallel with the legacy Readout System (ROS) at an input rate of 100 kHz. For the Phase-II, the legacy ROS will be completely replaced with a new system based on the next generation of FELIX and an evolution of the SW ROD application called Data Handler. Data Handler has the same functional requirements as SW ROD but must be able to operate at an input rate of 1 MHz. To facilitate this evolution the SW ROD has been implemented using plugin architecture. This contribution presents the design and implementation of the SW ROD application for Run 3, along with the strategy for its evolution to the Phase-II Readout system. It discusses the lessons learned during Run 3 and describes the challenges that have been addressed to accomplish the demanding performance requirements of HL-LHC.

Kolos, Serguei [Univ. of California, Irvine, CA (U↗

GPU-Accelerated Solution of the Bethe–Salpeter Equation for Large and Heterogeneous Systems

We present a massively parallel GPU-accelerated implementation of the Bethe–Salpeter equation (BSE) for the calculation of the vertical excitation energies (VEEs) and optical absorption spectra of condensed and molecular systems, starting from single-particle eigenvalues and eigenvectors obtained with density functional theory. The algorithms adopted here circumvent the slowly converging sums over empty and occupied states and the inversion of large dielectric matrices through a density matrix perturbation theory approach and a low-rank decomposition of the screened Coulomb interaction, respectively. Further computational savings are achieved by exploiting the nearsightedness of the density matrix of semiconductors and insulators to reduce the number of screened Coulomb integrals. We scale our calculations to thousands of GPUs with a hierarchical loop and data distribution strategy. The efficacy of our method is demonstrated by computing the VEEs of several spin defects in wide-band-gap materials, showing that supercells with up to 1000 atoms are necessary to obtain converged results. We discuss the validity of the common approximation that solves the BSE with truncated sums over empty and occupied states. In conclusion, we then apply our GW-BSE implementation to a diamond lattice with 1727 atoms to study the symmetry breaking of triplet states caused by the interaction of a point defect with an extended line defect.

Absorption spectra↗

Reducing communication in algebraic multigrid with multi-step node aware communication

Algebraic multigrid (AMG) is often viewed as a scalable [Formula: see text] solver for sparse linear systems. Yet, AMG lacks parallel scalability due to increasingly large costs associated with communication, both in the initial construction of a multigrid hierarchy and in the iterative solve phase. This work introduces a parallel implementation of AMG that reduces the cost of communication, yielding improved parallel scalability. It is common in Message Passing Interface (MPI), particularly in the MPI-everywhere approach, to arrange inter-process communication, so that communication is transported regardless of the location of the send and receive processes. Performance tests show notable differences in the cost of intra- and internode communication, motivating a restructuring of communication. In this case, the communication schedule takes advantage of the less costly intra-node communication, reducing both the number and the size of internode messages. Node-centric communication extends to the range of components in both the setup and solve phase of AMG, yielding an increase in the weak and strong scaling of the entire method.

Computer Science↗

Initial Measurements with the Prototype Parallel-Slit Ring Collimator Fast Neutron Emission Tomography System

Since 2017, Oak Ridge National Laboratory (ORNL) has been developing a passive fast-neutron emission tomography capability. The goal of this development is the ability to quantify the neutron source strength of individual fuel pins (rods) in spent nuclear fuel assemblies. Such a system could be used to measure the burnup of each fuel pin in a spent fuel assembly to take burnup credit when loading dry storage casks or to count individual fuel pins in spent fuel assemblies for safeguards purposes. At present, a laboratory prototype imager has been built and initial imaging measurements performed. The purpose of this prototype is to demonstrate imaging capability sufficient to resolve individual fuel pins in spent fuel assemblies, and in initial measurements, neutron sources separated by a spacing of 1.27 cm (similar to the spacing between fuel pins in commercial pressurized water reactor 17×17 nuclear fuel assemblies) have been resolved. This report documents the as-built imager, first measurements performed with it, and tomographic reconstructions performed using the measured data.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗