Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

An agent-based deployment decision-support system for electric vehicle services

METS-R ADDSEVS simulator is a high fidelity, parallel, agent-based evacuation simulator for multi-modal energy-optimal trip scheduling in real-time (METS-R) at transportation hubs. It consists of two modules. The first one is the traffic simulator module; the second one is the high-performance computing (HPC) module. More details can be found at https://umnilab.github.io/METS-R_doc/.

Lei, Zengxiang↗

PPO And Friends

PPO and Friends (PPOAF) is a pytorch implementation of proximal policy optimization for single- and multi-agent reinforcement learning (the PPO), along with several optimizations and add-ons (the Friends) to enable efficient MPI-parallelized model training on HPC clusters.

Maguire, AlisterO↗

Design Status of the Electron-Ion Collider

The Electron-Ion Collider is gearing up for "Critical Decision 2", theproject baseline with defined scope, cost and schedule.Lattice designs are beingfinalized, and preliminary component design is being carried out. Beam dynamicsstudies such as dynamic aperture optimization, instability and polarizationstudies, and beam-beam simulations are continuing in parallel. We report onthe latest developments and the overall status of the project, and presentthe plans for future activities.

43 PARTICLE ACCELERATORS↗

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

Systems and methods for tensor scheduling

A technique for efficient scheduling of operations in a program for parallelized execution thereof using a multi-processor runtime environment having two or more processors includes constraining the type or number of loop optimization transforms that may be explored such that memory and processing capacity available for the scheduling task are not exceeded, while facilitating a tradeoff between memory locality, parallelization, and/or data communication between memory modules of the multi-processor runtime environment.

Meister, Benoit J.↗

GronOR: Massively Parallel and GPU-Accelerated Non-Orthogonal Configuration Interaction for Large Molecular Systems

GronOR is a program package for non-orthogonal configuration interaction calculations for an electronic wave function built in terms of anti-symmetrized products of multi-configuration molecular fragment wave functions. The two-electron integrals that have to be processed may be expressed in terms of atomic orbitals or in terms of an orbital basis determined from the molecular orbitals of the fragments. The code has been specifically designed for execution on distributed memory massively parallel and Graphics Processing Unit (GPU)-accelerated computer architectures, using an MPI+OpenACC/OpenMP programming approach. The task-based execution model used in the implementation allows for linear scaling with the number of nodes on the largest pre-exascale architectures available, provides hardware fault resiliency, and enables effective execution on systems with distinct central processing unit-only and GPU-accelerated partitions. The code interfaces with existing multi-configuration electronic structure codes that provide optimized molecular fragment orbitals, configuration interaction coefficients, and the required integrals. Algorithm and implementation details, parallel and accelerated performance benchmarks, and an analysis of the sensitivity of the accuracy of results and computational performance to thresholds used in the calculations are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Innovative rail transport of a supersized land-based wind turbine blade

Wind turbine blade logistic providers are being challenged with escalating costs and routing complexities as one-piece blade approach lengths of 75 m in various regions of the U.S. land-based market. New lower cost solutions are needed to enable further reductions in the levelized cost of energy (LCOE) and continued market expansion. In this paper, a novel method of using existing U.S. rail infrastructure to deploy 100-m, one-piece blades to U.S. land-based wind sites is numerically investigated. The study removes the constraint that blades must be kept rigid during transport, and it allows bending to keep blades within a clearance profile while navigating horizontal and vertical curvatures. Novel system optimization and blade design processes consider blade structural constraints and rail logistic constraints in parallel to develop a highly flexible, rail-transportable blade. Results indicate maximum deployment potential in the Interior region of the United States and limited deployment potential in other regions. The study concludes that innovative rail transportation solutions combined with advanced rotor technologies can provide a feasible alternative to segmentation and support continued LCOE reductions in the U.S. land-based wind energy market.

17 WIND ENERGY↗

Parallelized POD-based suboptimal economic model predictive control of a state-constrained Boussinesq approximation

Motivated by an energy efficient building application, we want to optimize a quadratic cost functional subject to the Boussinesq approximation of the Navier-Stokes equations and to bilateral state and control constraints. Since the computation of such an optimal solution is numerically costly, we design an efficient strategy to compute a sub-optimal (but applicationally acceptable) solution with significantly reduced computational effort. We employ an economic Model Predictive Control (MPC) strategy to obtain a feedback control. The MPC sub-problems are based on a linear-quadratic optimal control problem subjected to mixed control and state constraints and a convection-diffusion equation, reduced with proper orthogonal decomposition. Finally, to solve each sub-problem, we apply a primal-dual active set strategy. The method can be fully parallelized, which enables the solution of large problems with real-world parameters.

97 MATHEMATICS AND COMPUTING↗

Moments-based interface reconstruction, remap and advection

Here, we present a new moment-of-fluid (MOF 2 ) interface reconstruction method. It uses the zeroth, first, and second moments of the fragment of material inside a cell of the mesh to reconstruct a convex material polygon or a union of convex polygons that approximate the respective material fragment. The new method requires information about the material moments only for the cell under consideration. The MOF 2 method allows to exactly reproduce several convex shapes: corners, filaments, and some concave shapes: cell-complements to corners and filaments. Interface reconstruction is formulated as a local (for each cell), non-linear, equality constrained optimization problem, which does not require additional communication and allows for an efficient parallel implementation. We present an extensive set of test problems, both for interface reconstruction on a single cell, and for reconstruction of a variety of shapes on a variety of meshes. We describe how to perform two-material advection using the MOF 2 method and present the results for the classical advection tests. We also show the examples of material interface remapping needed in the framework of multi-material arbitrary Lagrangian-Eulerian methods, and give a brief description of a procedure that can be used to update the material moments on the Lagrangian stage of those methods.

97 MATHEMATICS AND COMPUTING↗

An adaptive moments-based interface reconstruction using intersection of the cell with one half-plane, two half-planes and a circle

We present a new adaptive moment-of-fluid (A-MOF) interface reconstruction method. It uses the zeroth, first, and second moments of the fragment of material inside a cell of the mesh to construct a shape that approximates the respective material fragment. The new method requires information about the material moments only for the cell under consideration. The adaptive method chooses between shapes obtained by the intersection of the cell with one half-plane, two half-planes, or a circle. The A-MOF method allows to exactly reproduce several convex shapes: corners, filaments, and their concave cell-complements; as well as pieces of the circles and its cell-compliments. Interface reconstruction is formulated as a local (for each cell), non-linear, equality constrained optimization problem, which does not require additional communication and allows for an efficient parallel implementation. In conclusion, we present an extensive set of test problems, both for interface reconstruction on a single cell, and for reconstruction of a variety of shapes on the entire mesh.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Toward fully automated UED operation using two-stage machine learning model

To demonstrate the feasibility of automating UED operation and diagnosing the machine performance in real time, a two-stage machine learning (ML) model based on self-consistent start-to-end simulations has been implemented. This model will not only provide the machine parameters with adequate precision, toward the full automation of the UED instrument, but also make real-time electron beam information available as single-shot nondestructive diagnostics. Furthermore, based on a deep understanding of the root connection between the electron beam properties and the features of Bragg-diffraction patterns, we have applied the hidden symmetry as model constraints, successfully improving the accuracy of energy spread prediction by a factor of five and making the beam divergence prediction two times faster. The capability enabled by the global optimization via ML provides us with better opportunities for discoveries using near-parallel, bright, and ultrafast electron beams for single-shot imaging. It also enables directly visualizing the dynamics of defects and nanostructured materials, which is impossible using present electron-beam technologies.

36 MATERIALS SCIENCE↗

Optimizing Error-Bounded Lossy Compression for Scientific Data on GPUs

Error-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. With ever-emerging heterogeneous high-performance computing (HPC) architecture, GPU-accelerated error-bounded compressors (such as CUSZ and cuZFP) have been developed. However, they suffer from either low performance or low compression ratios. To this end, we propose CUSZ+ to target both high compression ratios and throughputs. We identify that data sparsity and data smoothness are key factors for high compression throughputs. Our key contributions in this work are fourfold: (1) We propose an efficient compression workflow to adaptively perform run-length encoding and/or variable-length encoding. (2) We derive Lorenzo reconstruction in decompression as multidimensional partial-sum computation and propose a fine-grained Lorenzo reconstruction algorithm for GPU architectures. (3) We carefully optimize each of CUSZ kernels by leveraging state-of-the-art CUDA parallel primitives. (4) We evaluate CUSZ+ using seven real-world HPC application datasets on V100 and A100 GPUs. Experiments show CUSZ+ improves the compression throughputs and ratios by up to 18.4x and 5.3x, respectively, over CUSZ on the tested datasets.

Tian, Jiannan↗

A Faster-Than-Real-Time Framework for Reliability-Oriented Simulation of PV Inverters

Physics-of-Failure (PoF) based reliability assessment for photovoltaic (PV) inverters requires long-duration electrical and electrothermal stress histories, yet generating such stress histories with high-fidelity switching models over year long mission profiles is computationally prohibitive. Conventional methods either sacrifice modeling fidelity for speed or require runtimes that are impractical for design iteration and uncertainty studies. To address this bottleneck, this paper presents a High-Performance Computing (HPC) based simulation frame work for faster-than-real-time reliability-oriented simulation. The proposed framework integrates the Average-to-Switching (A2S) method with parallel computing techniques to accelerate switching-level waveform reconstruction. We further introduce optimization strategies, including cluster merging and sensitivity based mission profile screening, to reduce the computational burden. Evaluated using real-world mission profile inputs and a MATLAB/Simulink switching-model reference, the framework reduces the simulation time for a one-year mission from an intractable multi-year duration to approximately 7.3 minutes while maintaining low waveform error. This acceleration provides a practical reliability-oriented simulation engine that can be coupled with component-specific aging models for subsequent PV inverter PoF assessment.

High-performance Computing↗

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

TCF High Efficiency Anaerobic Electroporation (CRADA Final Report)

The Joint BioEnergy Institute (JBEI) researchers were co-inventors of the technology that will be used on this project and have developed a more current version of the chip and controller. JBEI will also assist with the design of the pathways and implementation of pathways on the chip. LanzaTech has developed novel gas fermentation technology that captures and utilizes greenhouse gases for production of fuels and chemicals. In contrast to traditional fermentation that uses sugars as substrate (and releases CO2 as a byproduct), gas fermentation utilizes C1 substrates carbon monoxide (CO) or CO2. This enables a diverse range of feedstock options including waste gases from industrial sources (e.g., steel mills and processing plants) or syngas generated from any biomass resource (e.g., agricultural waste, municipal solid waste, or organic industrial waste). Biomass is then gasified, allowing for maximum yields and complete carbon utilization including the recalcitrant lignocellulosic fraction that cannot be utilized in traditional sugar fermentation. To maximize the value that can be added to the array of gas resources that the LanzaTech process can use as an input, LanzaTech has pioneered genetic modification of acetogens and developed a comprehensive set of genetic tools to perform routine strain modification, including genome editing tools as CRISPR/Cas9 and libraries of validated genetic parts as promoters and terminators. Using this platform, production of over 50 new molecules have been demonstrated directly from gas. For a few selected molecules production rates and yields have been optimized and surpass production of native producers and engineered E. coli or yeast strains, but a higher throughput approach for combinatorial optimization of pathways is required to further advance synthesis of additional products in parallel.

09 BIOMASS FUELS↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

Microgrid energy scheduling under uncertain extreme weather: Adaptation from parallelized reinforcement learning agents

Microgrids are useful solutions for integrating renewable energy resources and providing seamless green electricity to minimize carbon footprint. In recent years, extreme weather events happened often worldwide and caused significant economic and societal losses. Such events bring uncertainties to the microgrid energy scheduling problems and increase the challenges of microgrid operation. Traditional optimization approaches suffer from the inaccuracy of the uncertain microgrid model and the unseen events. Existing reinforcement learning (RL) - based approaches are also hampered by the limited generalization and the increasing computational burden when stochastic formulations are required to accommodate the uncertainties. This paper proposes a new parallelized reinforcement learning (PRL) method based on the probabilistic events to handle the microgrid energy uncertainties. Specifically, several local learning agents are employed to interact with pertinent microgrid environments in a distributed manner and report outcomes to the global agent, which will optimize microgrid energy resources online during extreme events. The stochastic microgrid energy optimization problem is reformulated to include all possible scenarios with probabilities. The advantage estimate functions of learning agents are designed with a backward sweep to transfer the outcomes to the value function updating process. Two simulation studies, stochastic optimization and online testing, are performed to compare with several existing RL approaches. Results substantiate that the proposed PRL method can achieve up to 20% optimization performance improvement with 4 and 28 times less computation cost than Q-learning with experience replay and multi-agent Q-learning approaches, respectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗