Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Development, construction and qualification tests of the mechanical structures of the electromagnetic calorimeter of the Mu2e experiment at Fermilab

The “muon-to-electron conversion” (Mu2e) experiment at Fermilab will search for the Charged Lepton Flavour Violating neutrino-less coherent conversion of a muon into an electron in the field of an aluminum nucleus. The observation of this process would be the unambiguous evidence of physics beyond the Standard Model. Mu2e detectors comprise a straw-tracker, an electromagnetic calorimeter and an external veto for cosmic rays. The calorimeter provides excellent electron identification, complementary information to aid pattern recognition and track reconstruction, and a fast calorimetric online trigger. The detector has been designed as a state-of-the-art crystal calorimeter and employs 1340 pure Cesium Iodide (CsI) crystals readout by UV-extended silicon photosensors and fast front-end and digitization electronics. A design consisting of two identical annular matrices (named “disks”) positioned at the relative distance of 70 cm downstream the aluminum target along the muon beamline satisfies the Mu2e physics requirements.The hostile Mu2e operational conditions, in terms of radiation levels (total ionizing dose of 12 krad and a neutron fluence of 5x1010 n/cm2 @ 1 MeVeq (Si)/y), magnetic field intensity (1 T) and vacuum level (10$^{-4}$ Torr) have posed tight constraints on the design of the detector mechanical structures and materials choice. The support structure of the two 670 crystal matrices employs two aluminum hollow rings and parts made of open-cell vacuum-compatible carbon fiber. The photosensors and service front-end electronics for each crystal are assembled in a unique mechanical unit inserted in a machined copper holder. The 670 units are supported by a machined plate made of vacuum-compatible plastic material. The plate also integrates the cooling system made of a network of copper lines flowing a low temperature radiation-hard fluid and placed in thermal contact with the copper holders to constitute a low resistance thermal bridge. The data acquisition electronics is hosted in aluminum custom crates positioned on the external lateral surface of the two disks. The crates also integrate the electronics cooling system as lines running in parallel to the front-end system.The constraints on the calorimeter mechanical structures design, the development from the conceptual design to the specifications of all the structural components, the status of components production and the components quality assurance tests are presented.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Symmetric Random Butterfly Transform (SRBT) Based Preconditioner

Summary of work using Symmetric Random Butterfly Transformation (SRBT) in conjunction with Incomplete LDL T factorization as a preconditioner for FGMRES solver as a way of solving linear systems arising from interior point methods applied to power system problems. These linear systems have proven difficult to parallelize and this represents a possible route forward.

interior point optimization↗

Modernizing the Legacy Fission Wire Measurement System for the Advanced Test Reactor-Critical Facility

Operational lifetime extensions of existing research reactors have emphasized the need for refurbishment, replacements, and upgrades to supporting equipment and instrumentation. The Advanced Test Reactor (ATR) at Idaho National Laboratory (INL), which entered service in 1967, has recently completed the sixth core internals change-out and has scheduled operations until at least 2040. Reactor maintenance and operational risk management is critically important in the research reactor community, however supporting measurement systems sometimes get overlooked when maintenance is planned. The Fission Wire Measurement System (FWMS) is a custom measurement system designed in the 1960s to measure the beta-particle activity of irradiated uranium-aluminum fission wires. This measurement is conducted to determine the fission rate profile of the Advanced Reactor Test Critical (ATR-C) facility. The ATR-C is an open-pool, low-power test reactor that was purpose driven to resemble ATR and is used to qualify experiment configurations and verify core models prior to full-power experiment irradiations in ATR. A power distribution measurement in ATR-C uses uranium-aluminum wires that are distributed throughout the ATR-C core to validate simulation and modeling results. These measurements require 340 to 1500 wires to be irradiated and measured within a 12-hour window. The activity of the wires is measured in the required time with the FWMS, which was put into service in 1965 at the Radiation Measurements Laboratory (RML). The system consists of 4 measurement channels and one reference channel, each with a 2-pi proportional gas flow detector and the measurement channels each have an automated sample changer. This legacy system is crucial to the continued operations of ATR and has undergone some minor hardware upgrades since 1965, however the system presently relies on custom control boards, custom gas ion chambers, analog amplifiers/discriminators, and a user interface (UI) for the system written in outdated code. Much of the equipment and software is custom with no commercial replacements or support and limited documentation. The existing control software requires an operating system that is no longer supported, creating more vulnerabilities to continued operations. A project is underway with a third-party vendor to design, build, and document a new control and data acquisition system (CDAS) for the FWMS. The new upgrade will replace the control system, computer, UI, sample changer motors, and main power supply while maintaining the interface with existing detector hardware. The upgraded system will be operated in parallel with the current hardware and software to conduct validation testing. This equipment upgrade demonstrates the commitment at ATR to ensuring successful operations and potential future research reactors at INL.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Xyce Case Study

With the elimination of underground nuclear testing and declining defense budgets, science-based stockpile stewardship requires increased reliance on high performance modeling and simulation of weapon systems. Today's weapon systems are comprised of various electrical components and systems. As a result, there is a need for tools that will allow the use of massively parallel modeling and simulation techniques on high performance computers in existing and future weapons' electrical systems models. The Xyce Parallel Electronic Simulator is a SPICE (Simulation Program with Integrated Circuit Emphasis)- compatible circuit simulator designed to run on large-scale parallel computing platforms, though it can also execute efficiently on a variety of architectures including single processor workstations. As a mature platform for large-scale parallel circuit simulation, Xyce supports standard capabilities available in commercial simulators, in addition to various devices and models specific to Sandia's needs. Specifically, Xyce aids in the design and verification of electrical and electronic circuits and systems prior to weapons' manufacturing and deployment.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Xyce Case Study

With the elimination of underground nuclear testing and declining defense budgets, science-based stockpile stewardship requires increased reliance on high performance modeling and simulation of weapon systems. Today's weapon systems are comprised of various electrical components and systems. As a result, there is a need for tools that will allow the use of massively parallel modeling and simulation techniques on high performance computers in existing and future weapons' electrical systems models. The Xyce Parallel Electronic Simulator is a SPICE (Simulation Program with Integrated Circuit Emphasis)- compatible circuit simulator designed to run on large-scale parallel computing platforms, though it can also execute efficiently on a variety of architectures including single processor workstations. As a mature platform for large-scale parallel circuit simulation, Xyce supports standard capabilities available in commercial simulators, in addition to various devices and models specific to Sandia's needs. Specifically, Xyce aids in the design and verification of electrical and electronic circuits and systems prior to weapons' manufacturing and deployment.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Parallel derivative-free optimization for simulation-based design of behind-the-meter energy systems

In this work, the integrated design and dispatch of behind-the-meter or distributed resources (e.g. stationary battery storage and solar PV generation) is considered. A simulation-based framework is employed, generating high-fidelity results with closed-loop predictive control at a fine resolution, at the expense of high computational cost (several minutes to a few hours per design point). To address this challenge, parallel derivative-free design methods are considered. Four methods are compared, including state-of-the-art surrogate-based methods (Radial-Basis Functions and Gaussian processes) and sampling strategies, an evolutionary-based method, and a simple sequential grid refinement method. As a case study, two types of design problem with increasing complexity are considered, namely, the design of behind-the-meter resources (three design variables) and the inclusion of grid capacity (four design variables). The second yields a constrained design problem for which violations can only be determined after solving the computationally expensive simulation. For the three-dimensional case, all methods present a good performance, achieving a solution within 1% of the optimum after the first iteration, with the sequential grid refinement exhibiting the fastest convergence and achieving the best final objective value. This indicates that the parallel evaluation of multiple sampling points may be more important than the choice of method for small decision spaces. For the four-dimensional constrained case, the Genetic Algorithm presents the best tradeoff between performance and computational effort, while the rough objective function terrain generated by constraint violation penalties reduces the performance of surrogate-based methods. Contour plots with flat regions indicate flexibility in the optimal design and highlight the importance of characterizing the solution space.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Function, Structure, and Regulation of Nitrogen Fixation-like Metalloproteins for Nitrogen, Energy, Carbon, and Sulfur Metabolism

Nitrogenases (N 2 ases) and nitrogen fixation-like (NFL) systems play distinct roles in nitrogen, carbon, sulfur, and energy metabolism based on their fundamental differences in structure and metallocofactor identity. As new NFL systems have recently been identified and characterized, striking parallels and differences compared to N 2 ase structure, catalysis, and regulation have emerged. NFL systems use metallocofactors that span from simple [4Fe-4S] clusters to complex clusters akin to FeMo-co, previously only thought to occur in N 2 ase. This review describes the present state of knowledge on the function, structure, catalytic mechanisms, and regulation of NFL systems that perform distinct biological roles across all three domains of life. Recent advancements in N 2 ase spectroscopic techniques for probing metallocofactor structure and electronic states guide current and future work on how each NFL system catalyzes its specific biological reaction(s). Key knowledge gaps and needed areas of research for uncovering the specific metallocofactors and structural motifs that are at the heart of NFL system reaction specificity, along with how these systems are regulated, are discussed.

Bacteria↗

Coupling of CTF and TRACE for Modeling of Transients

This report documents the improvements that have been made to the capabilities for coupling CTF to systems codes-specifically, the US Nuclear Regulatory Commission (NRC) TRACE code. An initial systems coupling capability had been set up previously using a nonoverlapping domain approach with the codes exchanging data at the core boundaries. The present work adds a new approach using overlapping domains, in which the system code models the core as well. A new input format has been added to allow the user to specify the physical quantities to be exchanged and their location in the system model, which gives the flexibility of applying one-way or two-way coupling between the codes using the desired data exchanges. In addition to applying thermal-hydraulics (T/H) boundary condition (BC) values obtained from TRACE, a capability was added to allow CTF to apply flow resistance feedback to TRACE to match the CTF core pressure drop. Support was added for executing parallel CTF models within the CTF-systems coupling. The system coupling capability was successfully applied to a parallel MSLB transient, demonstrating that both the one-way and two-way coupling behaved as expected and provided substantial improvements to numerical stability and routine compared to the previous nonoverlapping domain coupling. An initial capability was also developed for performing restart calculations in CTF which will be used in the future for restarting CTF-TRACE simulations at specific points in the transient simulation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Isochronous Architecture-Based Voltage-Active Power Droop for Multi-Inverter Systems

This article proposes an isochronous architecture for parallel inverters with only voltage-active power droop (VP-D) control for improving active power sharing as well as plug-and-play of multi-inverter-based distributed energy resources (DERs). The isochronous framework obviates the need for an explicit regulation of the frequency while allowing for sharing of reactive power. The article shows that the detrimental effects of circulating currents between inverters can be addressed in the framework developed. The isochronous architecture is implemented by employing a GPS to disseminate clock timing signals that enable the microgrid to maintain the nominal system frequency. Small signal eigenvalue analysis of a Multi-inverter Microgrid system near the steady-state operating point is presented to evaluate the system stability. Moreover, unlike traditional droop-based methods, the isochronous strategy lends itself to analytical guarantees which are developed in the article. Furthermore, the effect of delays in receiving GPS clock signals by an inverter is experimentally characterized and an inverter isolation mechanism is proposed that disconnects inverters facing large GPS communication delays. Hardware experiments on an 1.2 kVA-prototype and controller-hardware-in-the-loop validation on the CIGRE distribution network are conducted to demonstrate the effectiveness of the proposed architecture towards active and reactive power sharing between inverters with load scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗

All-electron APW+ lo calculation of magnetic molecules with the SIRIUS domain-specific package

We report APW+lo (augmented plane wave plus local orbital) density functional theory (DFT) calculations of large molecular systems using the domain specific SIRIUS multi-functional DFT package. The APW and FLAPW (full potential linearized APW) task and data parallelism options and the advanced eigen-system solver provided by SIRIUS can be exploited for performance gains in ground state Kohn–Sham calculations on large systems. This approach is distinct from our prior use of SIRIUS as a library backend to another APW+lo or FLAPW code. We benchmark the code and demonstrate performance on several magnetic molecule and metal organic framework systems. We show that the SIRIUS package in itself is capable of handling systems as large as a several hundred atoms in the unit cell without having to make technical choices that result in the loss of accuracy with respect to that needed for the study of magnetic systems.

Chemistry↗

Multigrid Reduction in Time for Chaotic Dynamical Systems

As CPU clock speeds have stagnated and high performance computers continue to have ever higher core counts, increased parallelism is needed to take advantage of these new architectures. Traditional serial time-marching schemes can be a significant bottleneck, as many types of simulations require large numbers of time-steps which must be computed sequentially. Parallel-in-time schemes, such as the Multigrid Reduction in Time (MGRIT) method, remedy this by parallelizing across time-steps and have shown promising results for parabolic problems. However, chaotic problems have proved more difficult, since chaotic initial value problems (IVPs) are inherently ill-conditioned. MGRIT relies on a hierarchy of successively coarser time-grids to iteratively correct the solution on the finest time-grid, but due to the nature of chaotic systems, small inaccuracies on the coarser levels can be greatly magnified and lead to poor coarse-grid corrections. Here we introduce a modified MGRIT algorithm based on an existing quadratically converging nonlinear extension to the multigrid Full Approximation Scheme (FAS), as well as a novel time-coarsening scheme. Together, these approaches better capture long-term chaotic behavior on coarse-grids and greatly improve convergence of MGRIT for chaotic IVPs. Further, we introduce a novel low-memory variant of the algorithm for solving chaotic PDEs with MGRIT which not only solves the IVP, but also provides estimates for the unstable Lyapunov vectors of the system. Finally, we provide supporting numerical results for the Lorenz system and demonstrate parallel speedup for the chaotic Kuramoto–Sivashinsky PDE over a significantly longer time-domain than in previous works.

97 MATHEMATICS AND COMPUTING↗

OpenMP Target Task: Tasking and Target Offloading on Heterogeneous Systems

This work evaluated the use of OpenMP tasking with target GPU offloading as a potential solution for programming productivity and performance on heterogeneous systems. Also, it is proposed a new OpenMP specification to make the implementation of heterogeneous codes simpler by using OpenMP target task, which integrates both OpenMP tasking and target GPU offloading in a single OpenMP pragma. As a test case, the authors used one of the most popular and widely used Basic Linear Algebra Subprogram Level-3 routines: triangular solver (TRSM). To benefit from the heterogeneity of the current high-performance computing systems, the authors propose a different parallelization of the algorithm by using a nonuniform decomposition of the problem. This work used target GPU offloading inside OpenMP tasks to address the heterogeneity found in the hardware. This new approach can outperform the state-of-the-art algorithms, which use a uniform decomposition of the data, on both the CPU-only and hybrid CPU-GPU systems, reaching speedups of up to one order of magnitude. The performance that this approach achieves is faster than the IBM ESSL math library on CPU and competitive relative to a highly optimized heterogeneous CUDA version. One node of Oak Ridge National Laboratory’s supercomputer, Summit, was used for performance analysis.

Valero Lara, Pedro↗

Real-Time Bayesian Inference at Extreme Scale: A Digital Twin for Tsunami Early Warning Applied to the Cascadia Subduction Zone

We present a Bayesian inversion-based digital twin that employs acoustic pressure data from seafloor sensors, along with 3D coupled acoustic–gravity wave equations, to infer earthquake-induced spatiotemporal seafloor motion in real time and forecast tsunami propagation toward coastlines for early warning with quantified uncertainties. Our target is the Cascadia subduction zone, with one billion parameters. Computing the posterior mean alone would require 50 years on a 512 GPU machine. Instead, exploiting the shift invariance of the parameter-to-observable map and devising novel parallel algorithms, we induce a fast offline–online decomposition. The offline component requires just one adjoint wave propagation per sensor; using MFEM, we scale this part of the computation to the full El Capitan system (43,520 GPUs) with 92% weak parallel efficiency. Moreover, given real-time data, the online component exactly solves the Bayesian inverse and forecasting problems in 0.2 seconds on a modest GPU system, a ten-billion-fold speedup.

97 MATHEMATICS AND COMPUTING↗

Design and implementation of dynamic I/O control scheme for large scale distributed file systems

In this paper, we have analyzed the input/output (I/O) activities of Cori, which is a high-performance computing system at the National Energy Research Scientific Computing Center at Lawrence Berkeley National Laboratory. Our analysis results indicate that most users do not adjust storage configurations but rather use the default settings. In addition, owing to the interference from many applications running simultaneously, the performance varies based on the system status. To configure file systems autonomously in complex environments, we developed DCA-IO, a dynamic distributed file system configuration adjustment algorithm that utilizes the system log information to adjust storage configurations automatically. Our scheme aims to improve the application performance and avoid interference from other applications without user intervention. Moreover, DCA-IO uses the existing system logs and does not require code modifications, an additional library, or user intervention. To demonstrate the effectiveness of DCA-IO, we performed experiments using I/O kernels of real applications in both an isolated small-sized Lustre environment and Cori. Our experimental results shows that our scheme can improve the performance of HPC applications by up to 263% with the default Lustre configuration.

97 MATHEMATICS AND COMPUTING↗

Frequency Recovery in Power Grids Using High-Performance Computing

Maintaining electric power system stability is paramount, especially in extreme contingencies involving unexpected outages of multiple generators or transmission lines that are typical during severe weather events. Such outages often lead to large supply-demand mismatches followed by subsequent system frequency deviations from their nominal value. The extent of frequency deviations is an important metric of system resilience, and its timely mitigation is a central goal of power system operation and control. This paper develops a novel nonlinear model predictive control (NMPC) method to minimize frequency deviations when the grid is affected by an unforeseen loss of multiple components. Our method is based on a novel multi-period alternating current optimal power flow (ACOPF) formulation that accurately models both nonlinear electric power flow physics and the primary and secondary frequency response of generator control mechanisms. We develop a distributed parallel Julia package for solving the large-scale nonlinear optimization problems that result from our NMPC method and thereby address realistic test instances on existing high-performance computing architectures. Our method demonstrates superior performance in terms of frequency recovery over existing industry practices, where generator levels are set based on the solution of single-period classical ACOPF models.

nonlinear model predictive control↗

A Multiport DC Transformer to Enable Flexible Scalable DC as a Service

The rapid adoption of new DC loads and sources, including photovoltaic arrays, DC fast charging of electric vehicle, battery energy storage and data centers, requires a large amount of DC power conversion. A majority of new deployments also necessitate the integration of multiple DC loads and sources at one site. The traditional approach to serve these new applications relies on multiple standard power converters, each dedicated to a source or load, and results in highly customized systems, with challenging control coordination, complex protection strategies, and poor scalability. Instead, this paper proposes the concept of a multiport DC transformer (MDCT) as a modular building block for realizing a flexible, scalable DC as a Service system and address this rapidly growing need. The MDCT uses the S4T to achieve very tight control of cycle-by-cycle energy exchange between multiple ports, with high efficiency. A single multiport converter replaces 4-6 distinct converters, integrates all energy flows and manages protection. As a result, a new layered control architecture is introduced to ensure stability and scalability of the MDCT, with multiple S4T power converter building blocks connected in parallel to realize a fully modular system and reach the target power levels.

30 DIRECT ENERGY CONVERSION↗

Resilient Inverter-Driven Black Start with Collective Parallel Grid-Forming Operation

As modern power systems are experiencing exceptional changes with increasing penetrations of inverter-based resources (IBRs), system restoration using IBRs has received attention. Using local grid-forming (GFM) assets near consumers, engineered to establish grid voltages in the absence of a stiff grid, i.e., bottom-up restoration, a distribution system could obtain high system resilience by not relying on the bulk power system restoration, which requires significant human intervention and procedure. This paper studies the technical feasibility of the novel approach with detailed electromagnetic transient (EMT) simulations. To thoroughly evaluate the potential of GFM inverters and the technical challenges in IBR-driven black start, a detailed three-phase inverter model is developed, including negative-sequence control for voltage balance and a phase-by- phase current limiter to sustain momentary overloading during the black start. To examine dynamic aspects of the black-start process, the EMT simulation also models transformer and motor dynamics to emulate their inrush and startup behaviors as well as network dynamics. In addition, active involvement of grid- following distributed energy resources is also studied to facilitate the black-start process. By allowing multiple GFM inverters to collectively black start without leader-follower coordination, we demonstrate that a system can achieve high resilience even with a fraction of assets lost. Two test cases of inverter-driven black start, using two and one GFM inverters, respectively, for a heavily unbalanced 2-MVA distribution feeder are demonstrated. Takeaways for further study and field deployment are provided.

grid-forming inverter↗