Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

AMG Preconditioners based on parallel hybrid coarsening and multi-objective graph matching

We describe preliminary results from a multi-objective graph matching algorithm, in the coarsening step of an aggregation-based Algebraic MultiGrid (AMG) preconditioner, for solving large and sparse linear systems of equations on high-end parallel computers. We have two objectives. First, we wish to improve the convergence behavior of the AMG method when applied to highly anisotropic problems. Second, we wish to extend the parallel package \texttt{PSCToolkit} to exploit multi-threaded parallelism at the node level on multi-core processors. Our matching proposal balances the need to simultaneously compute high weights and large cardinalities by a new formulation of the weighted matching problem combining both these objectives using a parameter $\lambda$. We compute the matching by a parallel $2/3-\varepsilon$-approximation algorithm for maximum weight matchings. Results with the new matching algorithm show that for a suitable choice of the parameter $\lambda$ we compute effective preconditioners in the presence of anisotropy, i.e., smaller solve times, setup times, iterations counts, and operator complexity.

D'Ambra, Pasqua↗

Fast and robust all-electron density functional theory calculations in solids using orthogonalized enriched finite elements

Here, we present a computationally efficient approach to perform systematically convergent real-space all-electron Kohn-Sham density functional theory calculations for solids using an enriched finite element (FE) basis. The enriched FE basis is constructed by augmenting the classical FE basis with atom-centered numerical basis functions, comprising of atomic solutions to the Kohn-Sham problem. Notably, to improve the conditioning, we orthogonalize the enrichment functions with respect to the classical FE basis, without sacrificing the locality of the resultant basis. In addition to improved conditioning, this orthogonalization procedure also renders the overlap matrix block diagonal, greatly simplifying its inversion. Subsequently, we use a Chebyshev polynomial based filtering technique to efficiently compute the occupied eigenspace in each self-consistent field iteration. We demonstrate the accuracy and efficiency of the proposed approach on periodic unit cells and supercells. The benchmark studies show a staggering 130× speedup of the orthogonalized enriched FE basis over the classical FE basis. We also present a comparison of the orthogonalized enriched FE basis with the linearized augmented plane-wave + local orbitals basis, both in terms of accuracy and efficiency. Notably, we demonstrate that the orthogonalized enriched FE basis can handle large system sizes of ~10 000 electrons. Finally, we observe good parallel scalability of our implementation with 92% efficiency at 22× speedup for a system with 620 electrons

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Design of a Multi-Chemistry Battery Pack System for Behind-the-Meter Storage Applications

Battery management systems (BMS) are essential for a battery pack's safe operation and longevity. This paper presents an active balancing method-based BMS for different cell chemistry structures to be used in behind-the-meter storage (BTMS) applications. The proposed system utilizes modular isolated dual active bridge (DAB) DC/DC converters to actively balance the battery pack through a low voltage (LV) bus. A supervisory controller monitors all the cell voltage, current, and state of charge (SOC) values. Based on the estimation of the SOCs, reference currents for the DAB converters are generated by the supervisory controller. Detailed modeling and the control approach of the modular DAB converters are presented in the paper. Moreover, the control strategy of the supervisory control is also analyzed. The proposed method and structure can be extended to any combination of the number of cells to design the battery pack. Simulation results are provided for a system consisting of three cells in parallel to form a cell block and three cell blocks in series to form the battery module. Experimental results are provided for three modular DAB converters operating with a LiFeMnPO4 prismatic cell with 3.2V, 20Ah rated values.

battery management systems↗

A memory-driven mapping algorithm for heterogeneous systems

mpibind is a memory-driven algorithm to map parallel hybrid applications to the underlying hardware resources transparently, efficiently, and portably. There are two fundamental aspects of this algorithm. First, unlike existing mappings, its primary design point is the memory system. Compute elements are selected based on the identified memory components and not vice versa. Second, it embodies a global awareness of hybrid programming abstractions as well as heterogeneous devices.

Leon Borja, EdgarA↗

Development of a Test-Bed for Testing and Refining EarthEn’s Supercritical CO 2 Based Energy Storage System

EarthEn’s energy storage concept leverages supercritical carbon dioxide (sCO 2 ) as a working fluid and relies on compact, high-performance components operating at elevated pressures and temperatures. To accelerate component development and reduce technical risk prior to larger-scale demonstrations, Oak Ridge National Laboratory (ORNL) developed a 100 kW-scale sCO 2 test-bed under a Cooperative Research and Development Agreement with EarthEn (CRADA NO. NFE-24-10050). The objective of the work was to design and construct a flexible experimental facility capable of reproducing key thermodynamic state points and heat-transfer conditions relevant to EarthEn’s thermal energy storage (TES) cycle, with particular emphasis on enabling development and evaluation of next-generation heat exchangers and TES concepts. The test-bed consists of a closed-loop sCO 2 circulation system housed within an open-topped enclosure. In its as-installed configuration, dense-phase sCO 2 is recirculated through a printed circuit recuperator, an electrically heated section, a throttling device used to simulate turbine expansion, and a water-cooled printed circuit heat exchanger that rejects heat to the building chilled-water system before returning to the pump. The pump is driven by a variable frequency drive, enabling controlled adjustment of flow and operating point. A comprehensive instrumentation suite was integrated to support both safe operation and high-quality data collection. Installed sensors include Coriolis flow meters for sCO 2 flow rate and density, resistance temperature detectors and thermocouples distributed throughout the loop (including the heated section and key heat exchanger ports), and pressure transducers for absolute and differential pressure measurements. The facility was designed to support high-pressure (19 MPa nominal) and high-temperature (575°C nominal) operation with credited overpressure protection provided by a rupture disk. Nominal operating conditions were selected to support 100 kW-class testing while maintaining flexibility for non-heated and heated shakedown, control development, and future integration of advanced TES test sections. In parallel with facility development, a system-level thermal-hydraulic model was created using Modelica-based tools to support component sizing, anticipate performance over targeted test conditions, and establish a framework for future model calibration against experimental data. At the conclusion of the project performance period, the facility was in final assembly, and the pressure boundary was nearly completed. However, several practical challenges associated with high-pressure/high-temperature systems and specialized component procurement impacted schedule and prevented initial pump-driven operation and full commissioning within the available resources. This report documents the as-built design, operating capabilities, and instrumentation, and it summarizes key lessons learned related to heater fabrication and testing, first-of-a-kind assembly factors, specialty flange supply constraints, and fill pump corrective actions. Finally, it outlines a phased plan for future commissioning and experimental campaigns, including control and instrumentation shakedown, heater characterization, model calibration, and testing at state points representative of EarthEn’s TES cycle.

25 ENERGY STORAGE↗

A Sheaf Theoretical Approach to Uncertainty Quantification of Heterogeneous Geolocation Information

Integration of multiple, heterogeneous sensors is a challenging problem across a range of applications. Prominent among these are multi-target tracking, where one must combine observations from different sensor types in a meaningful and efficient way to track multiple targets. Because different sensors have differing error models, we seek a theoretically justified quantification of the agreement among ensembles of sensors, both overall for a sensor collection, and also at a fine-grained level specifying pairwise and multi-way interactions among sensors. We demonstrate that the theory of mathematical sheaves provides a unified answer to this need, supporting both quantitative and qualitative data. Furthermore, the theory provides algorithms to globalize data across the network of deployed sensors, and to diagnose issues when the data do not globalize cleanly. We demonstrate and illustrate the utility of sheaf-based tracking models based on experimental data of a wild population of black bears in Asheville, North Carolina. A measurement model involving four sensors deployed among the bears and the team of scientists charged with tracking their location is deployed. This provides a sheaf-based integration model which is small enough to fully interpret, but of sufficient complexity to demonstrate the sheaf’s ability to recover a holistic picture of the locations and behaviors of both individual bears and the bear-human tracking system. A statistical approach was developed in parallel for comparison, a dynamic linear model which was estimated using a Kalman filter. This approach also recovered bear and human locations and sensor accuracies. When the observations are normalized into a common coordinate system, the structure of the dynamic linear observation model recapitulates the structure of the sheaf model, demonstrating the canonicity of the sheaf-based approach. However, when the observations are not so normalized, the sheaf model still remains valid.

97 MATHEMATICS AND COMPUTING↗

Graph-Learning-Assisted State and Event Tracking for Solar-Penetrated Power Grids with Heterogeneous Data Sources

Unlike transmission systems, distribution systems do not typically contain sufficient metering to enable real-time state estimation. The lack of sufficient real-time measurements prohibits accurate and timely monitoring of the state of distribution systems. As a result, control and optimal operation of distribution systems, especially those containing large numbers of renewable generation units are not possible without proper data and information about the current state of the system. The main motivation of this project is to address this shortcoming by developing an approach which provides “predicted” real-time measurements so that they can be used to execute a distribution system state estimator. Thus, the objective of the project is to make the distribution systems fully observable, such that the hosting capacity for solar generation can be accurately estimated, and unnecessary solar curtailments can be avoided. In order to accomplish this goal, the project investigated the use of a grid-model-informed machine learning (ML) tool which integrates heterogeneous data streams obtained from AMI meters, SCADA as well as PMU measurements and created synchronous measurement snapshots for the state estimator (SE); and developed a hybrid robust SE which provides not only accurate state estimates but also real-time feedback for the ML model refinement.

14 SOLAR ENERGY↗

Impact of organic acids and sulfate on the biogeochemical properties of soil from urban subsurface environments

Urban subsurface environments are often different from undisturbed subsurface environments due to the impacts of human activities. For example, deterioration of underground infrastructure can introduce elevated levels of Ca, Fe, and heavy metals into subsurface soils and groundwater. Likewise, leakage from sewer systems can lead to contamination by organic C, N, S, and P. However, the impact of these organic and inorganic compounds on biogeochemical processes including microbial redox reactions, mineral transformations, and microbial community transitions in urban subsurface environments is poorly understood. Here we conducted a microcosm experiment with soil samples from an urban construction site to investigate the possible biotic and abiotic processes impacted when sulfate and acetate or lactate were introduced into an urban subsurface environment. In the top-layer soil (0-0.3 m) microcosms, which were highly alkaline (pH > 10), the major impact was on abiotic processes such as secondary mineral precipitation. In the mid-layer (2-3 m) soil microcosms, the rate of Fe(III)-reduction and the amount of Fe(II) produced were greatly impacted by the specific organic acid added, and sulfate-reduction was not observed until after Fe(III)-reduction was complete. Near the end of the incubation, some genera related to syntrophic acetate oxidation and methanogenesis were observed in the lactateamended microcosms. In the bottom-layer (7-8 m) soil microcosms, the rate of Fe(III)-reduction and the amount of Fe(II) produced were affected by the concentration of amended sulfate. Sulfate-reduction was concurrent with Fe(III)-reduction, suggesting that Fe(II) production was likely due to abiotic reduction of Fe(III) by sulfide produced by microbial sulfate reduction. The slightly acidic initial pH (~5.8) of the mid-soil system was a major factor controlling sequential microbial Fe(III) and sulfate reduction versus parallel Fe(III) and sulfate reduction in the bottom soil system, which had a neutral initial pH (~7.2). Finally, 16S rRNA gene-based community analysis revealed a variety of indigenous microbial groups including alkaliphiles, dissimilatory iron and sulfate reducers, syntrophes, and methanogens tightly coupled with, and impacted by, these complex abiotic and biogeochemical processes occurring in urban subsurface environments.

54 ENVIRONMENTAL SCIENCES↗

Unorthodox parallelization for Bayesian quantum state estimation

Quantum state tomography (QST) allows for the reconstruction of quantum states through measurements and some inference technique under the assumption of repeated state preparations. Bayesian inference provides a promising platform to achieve both efficient QST and accurate uncertainty quantification, yet is generally plagued by the computational limitations associated with long Markov chains. In this work, we present a novel Bayesian QST approach that leverages modern distributed parallel computer architectures to efficiently sample a D-dimensional Hilbert space. Using a parallelized preconditioned Crank–Nicholson Metropolis–Hastings algorithm, we demonstrate our approach on simulated data and experimental results from IBM Quantum systems up to four qubits, showing significant speedups through parallelization. Although highly unorthodox in pooling independent Markov chains, our method proves remarkably practical, with validation ex post facto via diagnostics like the intrachain autocorrelation time. We conclude by discussing scalability to higher-dimensional systems, offering a path toward efficient and accurate Bayesian characterization of large quantum systems.

Bayesian inference↗

RAPID

Parallel computer code for the simulator for dynamics of power systems which has the capability to initiate the system and create different faults for the dynamic analysis. The code is based on time-parallel method (Parareal) with Adaptive Method Reduction (AMR). The coarse solvers for the Parareal algorithm include several Semi Analytical Solution methods. Also, Integrated simulation of coupled transmission and distribution systems can be studied.

Simunovic, Srdjan [Oak Ridge National Lab. (ORNL),↗

Parallel quantum annealing

Quantum annealers of D-Wave Systems, Inc., offer an efficient way to compute high quality solutions of NP-hard problems. This is done by mapping a problem onto the physical qubits of the quantum chip, from which a solution is obtained after quantum annealing. However, since the connectivity of the physical qubits on the chip is limited, a minor embedding of the problem structure onto the chip is required. In this process, and especially for smaller problems, many qubits will stay unused. We propose a novel method, called parallel quantum annealing, to make better use of available qubits, wherein either the same or several independent problems are solved in the same annealing cycle of a quantum annealer, assuming enough physical qubits are available to embed more than one problem. Although the individual solution quality may be slightly decreased when solving several problems in parallel (as opposed to solving each problem separately), we demonstrate that our method may give dramatic speed-ups in terms of the Time-To-Solution (TTS) metric for solving instances of the Maximum Clique problem when compared to solving each problem sequentially on the quantum annealer. Additionally, we show that solving a single Maximum Clique problem using parallel quantum annealing reduces the TTS significantly.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A VSC-HVDC-Assisted Black-Start Strategy in Bulk Power Systems a Case Study in San Diego

With the worldwide growth in deploying high-voltage direct current (HVDC) transmission systems, their ability to facilitate black-start (BS) restoration has been a research topic of interest. In this context, voltage source converter (VSC)-HVDC is regarded as a BS resource, and this paper proposes a VSC-HVDC-assisted parallel BS restoration strategy in bulk power systems. The proposed strategy consists of two stages: 1) determination of the VSC and generator startup sequence and 2) load restoration simulation. In the first stage, the entire blackout system is sectionalized into multiple subsystems. Each subsystem includes a VSC-HVDC station or traditional BS unit, it independently determines its generator startup timeline and the energization timelines for buses and lines. The second stage involves load restoration, conceptualized as a modified unit commitment problem, with the timelines established in the first stage work as critical inputs. The proposed BS restoration strategy is tested on the San Diego power system to simulate the 2011 Southwest blackout. The simulation results validate the effectiveness of using VSC-HVDC links as a BS resource which not only speeds up the restoration process but also reduces both energy and economic losses.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,↗

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING↗

A Task Based Approach for Co-Scheduling Ensemble Workloads on Heterogeneous Nodes

Scientific workflows consist of multiple, connected applications, with data and results flowing from one to another in a pipeline. Traditionally, such workflows are executed in sequential order, storing intermediate data in storage disks. Co-scheduling application workflows concurrently on the same compute nodes would greatly reduce the cost of moving data to/from storage and allow real-time analysis of intermediate results. Nevertheless, most parallel programming runtimes do not allow seamless integration of various applications in a scientific workflow, in part due to the complexity of managing data and resources. The situation is even more complicated for heterogeneous systems. In this work we extend the Minos Computing Library (MCL) runtime to accelerate pipe-lined and parallel workloads where multiple applications are running in the same system. MCL’s asynchronous task library and runtime dynamically manages resources to allow co-scheduling of multiple processes sharing heterogeneous resources. In addition, we design a custom ex- tension of the Open Compute Language (OpenCL) to enable multiple processes to share device memory. We enable MCL to coordinate these shared buffers to allow for easy, fast data sharing between applications. Using malleable micro-benchmarks and two application workflows that combine scientific simulation and AI-based analysis, we show that our method outperforms traditional approaches.

Index Terms—Parallel systems, Scheduling and Task ↗

An optical-input Maximum Likelihood Estimation feedback system demonstrated on tokamak horizontal equilibrium control

A readily parallelized Maximum Likelihood Estimation (MLE) algorithm with linear computational complexity is demonstrated in real time using only measurements from an extreme ultraviolet (EUV) diagnostic to control the horizontal position of a tokamak plasma. A set of trial emissivity profiles are parameterized by the control quantity of interest (R m ), and the MLE is identified from the profile which minimizes the signal reconstruction residual. The algorithm depends on an empirically determined likelihood function with exponential form. EUV emission (λ ≈ 15eV-1keV) is captured in a poloidal plane by four 16-channel AXUV diodes mounted at different poloidal angles with radial and angular resolution sufficient to discern plasma equilibrium evolution in HBT-EP. Calculations of the plasma major radius by the system are consistent within diagnostic uncertainty for the majority of the discharge with those of: a weighted average of vertical soft X-ray or EUV chords, magnetic sensors, and an equilibrium reconstruction. The feedback system corrects for a horizontal displacement of the major radius equal to 20% of the plasma minor radius by adjusting the vertical field produced from 40 in-vessel control coils in real time. The MLE calculation is performed on a GPU in a 15 μs cycle, with similar performance in this application to a simple weighted average of vertical chords. Finally, results demonstrate horizontal position control using magnetic actuators and an optical observer.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

CI/CD Efforts for Validation, Verification and Benchmarking OpenMP Implementations

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.

Jarmusch, Aaron↗