Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Dynamic Response of a Semiactive Suspension System with Hysteretic Nonlinear Energy Sink Based on Random Excitation by means of Computer Simulation

This paper aims to investigate the property and behavior of the hysteretic nonlinear energy sink (HNES) coupled to a half vehicle system which is a nine-degree-of-freedom, nonlinear, and semiactive suspension system in order to improve the ride comfort and increase the stability in shock mitigation by using the computer simulation method. The HNES model is a semiactive suspension device, which comprises the famous Bouc–Wen (B-W) model employed to describe the force produced by both the purely hysteretic spring and linear elastic spring of potentially negative stiffness connected in parallel, for the half vehicle system. Nine nonlinear motion equations of the half vehicle system are derived in terms of the seven displacements and the two dimensionless hysteretic variables, which are integrated numerically by employing the direct time integration method for studying both the variables of vertical displacements, velocities, accelerations, chassis pitch angle, and the ride comfort and driver safety, respectively, based on the bump and random road inputs of the pseudoexcitation method as excitation signal. Simulation results show that, compared with the HNES model and the magnetorheological (MR) model coupled to the half vehicle system, the ride comfort and stability have been evidently improved. A successful validation process has been performed, which indicated that both the ride comfort and driver safety properties of the HNES model coupled to half vehicle significantly improved.

Chen, Hui↗

Century: Zap Energy’s 100-kW-Scale Repetitive Sheared-Flow-Stabilized Z -Pinch System with Liquid Metal Cooling

Zap Energy is developing the sheared-flow-stabilized (SFS) Z-pinch concept for commercial applications. The SFS Z pinch relies on plasma self-organization, in the sense that plasma dynamics play a critical role in confinement. Using plasma axial current for confinement and compression eliminates the need for external confinement or heating technologies. This compact magnetic confinement technology could, in turn, provide the basis for a cost-effective deuterium-tritium fusion power plant. In addition to a robust experimental program pushing plasma performance towards breakeven conditions, Zap Energy has parallel programs developing power handling systems suitable for future power plants. Technologies under development include high average-power repetitive pulsed power, high duty-cycle cathodes, and liquid metal wall systems. Century is the name of Zap Energy’s first effort to integrate these three components into an operational system capable of firing non-reacting hydrogen SFS Z-pinch plasmas into a liquid-metal-lined container at sustained repetition rates on the order of 0.1 Hz. Here, the pulsed power driver and liquid metal heat exchanger are both designed to sustain input powers of 100 kW. Construction and initial operations with an interim ~10 kW liquid metal heat exchanger are described.

Century↗

Advanced Computing is at the Forefront of a New “Moonshot” Revolutionizing the North American Power Grid

In the 50+ years since the first humans landed on the moon, computing has grown at breakneck speed. We are faced with another challenge that is just as daunting, and just as important to overcome-modernizing the North American electric power grid-and high-performance computing (HPC) systems with specialized software will be an important element in rising to this challenge. We describe at a high level how software developed in the ExaSGD project addresses this "moonshot" goal by utilizing exascale computing and a novel high performance solver software stack to support the mission of decarbonizing power grid operations in an environment of uncertain weather and climate. To reach the exascale benchmark the team has made a number of first-of-their-kind innovations, including novel method for stochastic optimization, fine grained parallel methods for modeling power systems, and GPU resident sparse numerical linear solvers.

17 WIND ENERGY↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

HDF5 in the exascale era: Delivering efficient and scalable parallel I/O for exascale applications

Accurately modeling real-world systems requires scientific applications at exascale to generate massive amounts of data and manage data storage efficiently. However, parallel input and output (I/O) faces challenges due to new application workflows and the state-of-the-art memory, interconnect, and storage architectures considered in exascale designs. The storage hierarchy has expanded with node-local persistent memory, solid-state storage, and traditional disk and tape-based storage, thus requiring efficiency at each layer and much more efficient data movement among these layers. This paper discusses how the ExaHDF5 project improved the I/O performance and data management for exascale architectures by enhancing HDF5, a widely used parallel I/O library. The team developed an Asynchronous I/O Virtual Object Layer (VOL) connector that allowed overlapping I/O with computation. They also created a Cache VOL to complement asynchronous I/O by incorporating fast storage layers, such as burst buffer and node-local storage, into the parallel I/O workflow through caching and staging data. Additionally, the team enabled data aggregation and I/O at the node level by using a Subfiling Virtual File Driver (VFD). To demonstrate superior I/O performance with HDF5 at exascale, the ExaHDF5 team collaborated with several exascale applications. In this paper, we show I/O performance improvements for three applications: Cabana (a particle-based simulation library), EQSIM (a regional earthquake simulation software), and E3SM (a climate system modeling library).

Asynchronous I/Ol↗

Structural Simluation Toolkit (SST) v.12.0

The Structural Simulation Toolkit (SST) was developed to explore innovations in highly concurrent computing systems where the instruction set architecture (ISA), micro-architecture, and memory interact with the programming model and communications system. The package provides a fully modular design for extensive exploration of an individual system parameter as well as a parallel simulation environment based on message passing interface (MPI) which enable a high level of performance as well as the ability to look at large systems. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodrigues, ArunF.↗

Water–Hydrocarbon Interactions in Anionic Pyrene Monohydrate

Interactions between water and polycyclic aromatic hydrocarbons are essential in many aspects of chemistry, from interstellar and atmospheric processes to interfacial hydrophobicity and wetting phenomena. Despite their growing importance, the intermolecular potentials of the water-hydrocarbon interactions are underdeveloped compared to water-water potentials, and there are similarly few experimental probes that are sensitive to the details of the water-hydrocarbon potential. We present a combined experimental and computational study of anionic pyrene monohydrate, one of the simplest water/hydrocarbon clusters. The action spectrum in the OH region of the mass-selected cluster ion provides a rigorous benchmark for intermolecular potentials and computational methodologies. We identify missing intermolecular interactions and shortcomings in conventional dynamics calculations by comparing experimental data to density functional theory and classical molecular dynamics calculations. Kinetic trapping is prevalent, even for one water molecule and one pyrene molecule, leading to slow equilibration in conventional molecular dynamics calculations, even on nanosecond timescales and at low temperatures (50 K). At constant energy, temperature fluctuations for the pair of molecules are substantial. Immersing the system in a bath of soft spheres and employing parallel tempering alleviates kinetic trapping and dampens temperature fluctuations, bringing the system closer to the thermodynamic limit. With such augmented sampling, a simple, flexible water model reproduces the linewidth and the asymmetric broadening of the symmetric OH stretching mode, which we assign to spectral diffusion. In the OH stretching region, dynamics calculations predict a more intense antisymmetric peak than experiments observe but do not predict the bimodal split symmetric peak that the experiments show. Furthermore, our work suggests that electronic polarization, missing in the empirical force field, is responsible for the first discrepancy and that quantum nuclear effects, captured neither in density functional theory nor in classical dynamics, may be responsible for the second.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

Development of algorithms for augmenting and replacing conventional process control using reinforcement learning

Here, this work seeks to allow for the online operation and training of model-free reinforcement learning (RL) agents but limit the risk to system equipment and personnel. The parallel implementation of RL alongside more conventional process control (CPC) allows for the RL algorithm to learn from CPC. The past performance of both methods are assessed on a continuous basis allowing for a transition from CPC to RL and, if needed, transitioning back to CPC from RL. This allows for the RL algorithm to slowly and safely assume control of the process without significant degradation in control performance. It is shown that the RL can derive a near optimal policy even when coupled with a suboptimal CPC. It is also demonstrated that the coupled RL-CPC algorithm learns at a faster rate than traditional RL methods of exploration while the algorithm’s performance does not deteriorate below CPC, even when exposed to an unknown operating condition.

30 DIRECT ENERGY CONVERSION↗

Pass-efficient methods for compression of high-dimensional turbulent flow data

The future of high-performance computing, specifically on future Exascale computers, will presumably see memory capacity and bandwidth fail to keep pace with data generated, for instance, from massively parallel partial differential equation (PDE) systems. Current strategies proposed to address this bottleneck entail the omission of large fractions of data, as well as the incorporation of in situ compression algorithms to avoid overuse of memory. To ensure that post-processing operations are successful, this must be done in a way that a sufficiently accurate representation of the solution is stored. Moreover, in situations where the input/output system becomes a bottleneck in analysis, visualization, etc., or the execution of the PDE solver is expensive, the number of passes made over the data must be minimized. In the interest of addressing this problem, this work focuses on the utility of pass-efficient, parallelizable, low-rank, matrix decomposition methods in compressing high-dimensional simulation data from turbulent flows. Additionally, a particular emphasis is placed on using coarse representation of the data – compatible with the PDE discretization grid – to accelerate the construction of the low-rank factorization. This includes the presentation of a novel single-pass matrix decomposition algorithm for computing the so-called interpolative decomposition. The methods are described extensively and numerical experiments on two turbulent channel flow data are performed. In the first (unladen) channel flow case, compression factors exceeding 400 are achieved while maintaining accuracy with respect to first- and second-order flow statistics. In the particle-laden case, compression factors of 100 are achieved and the compressed data is used to recover particle velocities. These results show that these compression methods can enable efficient computation of various quantities of interest in both the carrier and disperse phases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Xcompact3D: An open-source framework for solving turbulence problems on a Cartesian mesh

Xcompact3D is a Fortran 90–95 open-source framework designed for fast and accurate simulations of turbulent flows, targeting CPU-based supercomputers. It is an evolution of the flow solver Incompact3D which was initially designed in France in the mid-90’s for serial processors to solve the incompressible Navier–Stokes equations. Incompact3D was then ported to parallel High Performance Computing (HPC) systems in the early 2010’s. Very recently the capabilities of Incompact3D have been extended so that it can now tackle more flow regimes (from incompressible flows to compressible flows at low Mach numbers), resulting in the design of a new user-friendly framework called Xcompact3D. The present manuscript presents an overview of Xcompact3D with a particular focus on its functionalities, its ready-to-run simulations and a few case studies to demonstrate its impact.

17 WIND ENERGY↗

Low-abundance populations distinguish microbiome performance in plant cell wall deconstruction

Abstract Background Plant cell walls are interwoven structures recalcitrant to degradation. Native and adapted microbiomes can be particularly effective at plant cell wall deconstruction. Although most understanding of biological cell wall deconstruction has been obtained from isolates, cultivated microbiomes that break down cell walls have emerged as new sources for biotechnologically relevant microbes and enzymes. These microbiomes provide a unique resource to identify key interacting functional microbial groups and to guide the design of specialized synthetic microbial communities. Results To establish a system assessing comparative microbiome performance, parallel microbiomes were cultivated on sorghum ( Sorghum bicolor L. Moench) from compost inocula. Biomass loss and biochemical assays indicated that these microbiomes diverged in their ability to deconstruct biomass. Network reconstructions from gene expression dynamics identified key groups and potential interactions within the adapted sorghum-degrading communities, including Actinotalea , Filomicrobium , and Gemmatimonadetes populations. Functional analysis demonstrated that the microbiomes proceeded through successive stages that are linked to enzymes that deconstruct plant cell wall polymers. The combination of network and functional analysis highlighted the importance of cellulose-degrading Actinobacteria in differentiating the performance of these microbiomes. Conclusions The two-tier cultivation of compost-derived microbiomes on sorghum led to the establishment of microbiomes for which community structure and performance could be assessed. The work reinforces the observation that subtle differences in community composition and the genomic content of strains may lead to significant differences in community performance.

59 BASIC BIOLOGICAL SCIENCES↗

GLUE Code: A framework handling communication and interfaces between scales

Many scientific applications are inherently multiscale in nature. Such complex physical phenomena often require simultaneous execution and coordination of simulations spanning multiple time and length scales. This is possible by combining expensive small-scale simulations (such as molecular dynamics simulations) with larger scale simulations (such continuum limit/hydro solvers) to allow for considerably larger systems using task and data parallelism. However, the granularity of the tasks can be very large and often leads to load imbalance. Traditionally, we use approximations to streamline the computation of the more costly interactions and this introduces trade-offs between simulation cost and accuracy. In recent years, the available computational power and the advances in machine learning have made computing these scale-bridging interactions and multiscale simulations more feasible. One driving application has been plasma modeling in inertial confinement fusion (ICF), which is fundamentally multiscale in nature. This requires deep understanding of how to extrapolate microscopic information into macroscopically relevant scales. For example, in ICF one needs an accurate understanding of the connection between experimental observables and the underlying microphysics. The properties of the larger scales are often affected by the microscale behavior incorporated usually into the equations of state and ionic and electronic transport coefficients (Liboff, 1959; Rinderknecht et al., 2014; Rosenberg et al., 2015; Ross et al., 2017). Instead of incorporating this information using reliable molecular dynamics (MD) simulations, one often needs to use theoretical models, due to the inability of MD to reach engineering scales (Glosli et al., 2007; Marinak et al., 1998). One approach to resolve this issue is by coupling two MD simulations of different scales via force interpolation, e.g., the AdResS method (Krekeler et al., 2018; Nagarajan et al., 2013). Another approach, which we will pursue in the scope of this work, is by enabling scale bridging between MD simulations and meso/macro-scale models through the development and support of application programming interfaces that these different applications can interact through.

54 ENVIRONMENTAL SCIENCES↗

High Yield Xray Imager Final Design Review

The High Yield Xray Imager (HYXI) is a new NIF target diagnostic system currently under development. The goal of HYXI is to provide high-fidelity, high temporal resolution x-ray imaging capability on high yield NIF implosions at 10MJ and above. The HYXI instrument design concept is based on the combination of two technologies that have been successfully utilized at the NIF on previous instruments, electron pulse-dilation and hybrid-CMOS sensor imaging. The combination of these two techniques will give HYXI sufficient data quality to ascertain differences in hot spot formation dynamics between high and low yield implosions. This information will highlight the critical hot spot conditions needed for ignition and burn. The HYXI design leverages the successful operation of the PDIXI x-ray imager at the NIF on multi MJ yield shots. A new radiation tolerant CMOS imaging array (HYPERION) is being developed to eliminate the significant background noise which limits the data quality of PDIXI. We successfully placed the contract with Advanced hCMOS Systems (AHS) to develop the HYPERION sensor, which fulfils our criteria to place long lead time item procurements by end of FY24. The HYXI Final Design Review was completed at the end of Q4 FY24 (Sep 24 th and Sep 30 th ). The HYXI project is a multi-year effort with a phased approach to be bring up system functionality over time in parallel with the development and fabrication effort of the HYPERION CMOS imaging array. In Phase 1, time-integrated x-ray images on NIF DT experiments will be collected starting in Q3 FY25. In Phase 2 of the project, time-resolved imaging with HYXI utilizing a spare microchannel plate detector back-end will begin in Q3 FY26. Phase 3 concludes the project with the installation of the HYPERION sensor array and the final performance qualification of the HYXI instrument which is scheduled for Q3 FY27 as discussed in the PDR and MRT report on this project in FY23.

42 ENGINEERING↗

Using Hardware-In-The-Loop Methodology to Develop Test Systems

Hardware in the Loop (HIL) testing methodologies have become widespread in industry. Typically, they focus on developing control algorithms for systems such as autonomous vehicles or aircraft. An oft overlooked aspect of product development is the design and fabrication of a test system for validating that the product meets requirements. Abstractly, a test system differs little from a control system—testers provide signals to the unit, monitor feedback, and base decisions on the results. While the time scales may differ, the functionalities are conceptually similar. Viewed in this light, it becomes natural to extend HIL approaches to tester development. By replacing a physical unit with a proxy model deployed to a real-time or pseudo real-time target, test systems can be developed in parallel with the design and fabrication of a first production unit. This saves considerable time in the life cycle from conceptual design to realized product. This manuscript demonstrates the process flow using a capacitive discharge unit as an exemplar.

42 ENGINEERING↗

Development, construction and qualification tests of the mechanical structures of the electromagnetic calorimeter of the Mu2e experiment at Fermilab

The “muon-to-electron conversion” (Mu2e) experiment at Fermilab will search for the Charged Lepton Flavour Violating neutrino-less coherent conversion of a muon into an electron in the field of an aluminum nucleus. The observation of this process would be the unambiguous evidence of physics beyond the Standard Model. Mu2e detectors comprise a straw-tracker, an electromagnetic calorimeter and an external veto for cosmic rays. The calorimeter provides excellent electron identification, complementary information to aid pattern recognition and track reconstruction, and a fast calorimetric online trigger. The detector has been designed as a state-of-the-art crystal calorimeter and employs 1340 pure Cesium Iodide (CsI) crystals readout by UV-extended silicon photosensors and fast front-end and digitization electronics. A design consisting of two identical annular matrices (named “disks”) positioned at the relative distance of 70 cm downstream the aluminum target along the muon beamline satisfies the Mu2e physics requirements.The hostile Mu2e operational conditions, in terms of radiation levels (total ionizing dose of 12 krad and a neutron fluence of 5x1010 n/cm2 @ 1 MeVeq (Si)/y), magnetic field intensity (1 T) and vacuum level (10$^{-4}$ Torr) have posed tight constraints on the design of the detector mechanical structures and materials choice. The support structure of the two 670 crystal matrices employs two aluminum hollow rings and parts made of open-cell vacuum-compatible carbon fiber. The photosensors and service front-end electronics for each crystal are assembled in a unique mechanical unit inserted in a machined copper holder. The 670 units are supported by a machined plate made of vacuum-compatible plastic material. The plate also integrates the cooling system made of a network of copper lines flowing a low temperature radiation-hard fluid and placed in thermal contact with the copper holders to constitute a low resistance thermal bridge. The data acquisition electronics is hosted in aluminum custom crates positioned on the external lateral surface of the two disks. The crates also integrate the electronics cooling system as lines running in parallel to the front-end system.The constraints on the calorimeter mechanical structures design, the development from the conceptual design to the specifications of all the structural components, the status of components production and the components quality assurance tests are presented.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Symmetric Random Butterfly Transform (SRBT) Based Preconditioner

Summary of work using Symmetric Random Butterfly Transformation (SRBT) in conjunction with Incomplete LDL T factorization as a preconditioner for FGMRES solver as a way of solving linear systems arising from interior point methods applied to power system problems. These linear systems have proven difficult to parallelize and this represents a possible route forward.

interior point optimization↗