Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Secondary Use-Plug-and-Play Energy Storage System Composed of Multiple Energy Storage Technologies

Low-cost, grid-connectable energy storage technologies represent a significant challenge for the electric grid of the future. Energy storage technologies are in rapid development with targets to reduce the storage medium cost. However, a significant cost to deployment also comes in the integration. This paper presents the development of a plug-and-play system for supporting secondary use multiple battery systems into a single grid connectable unit. Results of the system design are demonstrated in a controller hardware in the loop (CHIL) platform. Simulations of two energy storage systems operating in parallel and dispatched optimally are presented.

Starke, Michael↗

Advanced Computing is at the Forefront of a New “Moonshot” Revolutionizing the North American Power Grid

In the 50+ years since the first humans landed on the moon, computing has grown at breakneck speed. We are faced with another challenge that is just as daunting, and just as important to overcome-modernizing the North American electric power grid-and high-performance computing (HPC) systems with specialized software will be an important element in rising to this challenge. We describe at a high level how software developed in the ExaSGD project addresses this "moonshot" goal by utilizing exascale computing and a novel high performance solver software stack to support the mission of decarbonizing power grid operations in an environment of uncertain weather and climate. To reach the exascale benchmark the team has made a number of first-of-their-kind innovations, including novel method for stochastic optimization, fine grained parallel methods for modeling power systems, and GPU resident sparse numerical linear solvers.

17 WIND ENERGY↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

RX-PSA

This code is built on top of the ML-PSA utility and implements the ability to implement a multi-fidelity parallel simulated annealing optimization for a range of engineering problems. RX-PSA provides specializations to interact with modern nuclear reactor codes such as VERA to perform assembly and core optimization in a robust way. RX-PSA also provides LWR specific objective functions and constraints to provide a straightforward user interface that reactor designers are familiar with.

Collins, BenjaminS.↗

Wavefront shaping with a Hadamard basis for scattering soil imaging

Here, soil is a scattering medium that inhibits imaging of plant-microbial-mineral interactions that are essential to plant health and soil carbon sequestration. However, optical imaging in the complex medium of soil has been stymied by the seemingly intractable problems of scattering and contrast. Here, we develop a wavefront shaping method based on adaptive stochastic parallel gradient descent optimization with a Hadamard basis to focus light through soil mineral samples. Our approach allows a sparse representation of the wavefront with reduced dimensionality for the optimization. We further divide the used Hadamard basis set into subsets and optimize a certain subset at once. Simulation and experimental optimization results demonstrate our method has an approximately seven times higher convergence rate and overall better performance compared to that with optimizing all pixels at once. The proposed method can benefit other high-dimensional optimization problems in adaptive optics and wavefront shaping.

47 OTHER INSTRUMENTATION↗

Exago TM Users Manual: Version 1.0

The Exascale Grid Optimization (ExaGOTM) toolkit is an open source package for solving large-scale power grid optimization problems on parallel and distributed architectures, particularly targeted for exascale machines with heteregenous architectures (GPU). This manual is a guide to ExaGO's working including installation, formulation, and usage.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dynamic Modeling, Trajectory Optimization, and Linear Control of Cable-Driven Parallel Robots for Automated Panelized Building Retrofits

The construction industry faces a growing need for automation to reduce costs, improve accuracy and productivity, and address labor shortages. One area that stands to benefit significantly from automation is panelized prefabricated building envelope retrofits, which can improve a building’s energy efficiency in heating and cooling interior spaces. In this paper, we propose using cable-driven parallel robots (CDPRs), which can effectively lift and handle large objects, to install these panels. However, implementing CDPRs presents significant challenges because of their nonlinear dynamics, complex trajectory planning, and precise control requirements. To tackle these challenges, this work focuses on a new application of established control and trajectory optimization theories in a CDPR simulation of a building envelope retrofit under real-world conditions. We first model the dynamics of CDPRs, highlighting the critical role of damping in system behavior. Building on this dynamic model, we formulate a trajectory optimization problem to generate feasible and efficient motion plans for the robot under operational and environmental constraints. Given the high precision required in the construction industry, accurately tracking the optimized trajectory is essential. However, challenges such as partial observability and external vibrations complicate this task. To address these issues, a Linear Quadratic Gaussian control framework is applied, enabling the robot to track the optimized trajectories with precision. Simulation results show that the proposed controller enables precise end effector positioning with errors under 4 mm, even in the presence of external wind disturbances. Through comprehensive simulations, our approach allows for an in-depth exploration of the system’s nonlinear dynamics, trajectory optimization, and control strategies under controlled yet highly realistic conditions. The results demonstrate the feasibility of CDPRs for automating panel installation and provide insights into their practical deployment.

CDPR↗

Evaluation and Optimization of Well Completion Options for the Utah FORGE Site

Orientation and completion for well pairs that have been subjected to multi-zonal stimulation play a critical role in the long-term performance of an Enhanced Geothermal Reservoir. Enhanced geothermal systems often rely on preferential flow along fractures between well injection and production locations. Modeling this preferential flow using discrete fracture networks (DNF) relies on stochastic realizations of the DFN based on geological sampling. Here we present the development of a stochastic optimization methodology to determine well completion options in a discrete fracture network based on using parallel subset simulation. Stochastic optimization will provide insight into regions where placements of the injection and production wells are optimal. An example optimization of well-pair location optimization based on a deterministic-stochastic DFN model representing FORGE follows a discussion of the theory.

15 GEOTHERMAL ENERGY↗

Optimizing temperature distributions for training neural quantum states using parallel tempering

Parametrized artificial neural networks (ANNs) can be very expressive ansatzes for variational algorithms, reaching state-of-the-art energies on many quantum many-body Hamiltonians. Nevertheless, the training of the ANN can be slow and stymied by the presence of local minima in the parameter landscape. One approach to mitigate this issue is to use parallel tempering methods, and in this work, we focus on the role played by the temperature distribution of the parallel tempering replicas. Using an adaptive method that adjusts the temperatures in order to equate the exchange probability between neighboring replicas, we show that this temperature optimization can significantly increase the success rate of the variational algorithm with negligible computational cost by eliminating bottlenecks in the replicas' random walk. Furthermore, we demonstrate this using two different neural networks, a restricted Boltzmann machine and a feedforward network, which we use to study a toy problem based on a permutation invariant Hamiltonian with a pernicious local minimum and the 𝐽 1 −𝐽 2 model on a rectangular lattice.

Neural network simulations↗

Optimization Plugin Library

The Optimization Plugin library ("op") is a lightweight general optimization solver interface. The primary purpose of op is to simplify the process of integrating different optimization solvers (serial or parallel) with scalable parallel physics engines. By design it has several features that help make this a reality. The core abstraction interface was developed to encompass a large class of optimization problems in an optimizer-agnostic way. This enables us to describe the optimization problem once and then use a variety of supported "op" optimizers with ideally no code-changes. The abstraction interface is made up of lightweight wrappers that make it easy to integrate with existing simulation codes. This makes integration less intrusive and should minimize changes to existing physics codes. The "op" interface includes an assortment of utility methods that help specify parallel communication patterns as well as methods to convert from optimization-specific interfaces to the more general "op" interface. Lastly a dynamic library linking interface is provided to allow for use of proprietary optimization engines without explicit reference in the source code, along with standard shared library interfaces for opensource engines.

Jekel, CharlesF↗

GOAT. jl

SF-23-008 This project is a Julia implementation of the Gradient Optimization of Analytic conTrols (GOAT) optimal control methodology. It integrates with other packages in Julia's ecosystem to provide memory-efficient, parallelized solutions to quantum optimal control tasks. A prototype implementation of some of these algorithms was initially developed and funded by the ASCR Early Career Research Award program under PI Travis Humble at Oak Ridge National Laboratory. The current version was funded under the ASCR AIDE-QC Program under PI Paul Hovland. The current package to be released has a novel implementation, syntax, and structure making it substantially different than the original prototype (which was not released under copyright to the best of my knowledge).

KAIRYS, PAUL↗

A parallel hub-and-spoke system for large-scale scenario-based optimization under uncertainty

Practical solution of stochastic programming problems generally requires the use of parallel computing resources. Here, we describe the open source package mpi-sppy, in which efficient and scalable parallelization is a central feature. We report computational experiments that demonstrate the ability to solve very large stochastic programming problems - including mixed-integer variants - in minutes of wall clock time, efficiently leveraging significant parallel computing resources. We report results for the largest publicly available instances of stochastic mixed-integer unit commitment problems, solving to provably tight optimality gaps. In addition, we introduce a novel software architecture that facilitates combinations of methods for accelerating convergence that can be combined in plug-and-play manner. Finally, the mpi-sppy package is written in Python, leverages the widely used Pyomo (http://www.pyomo.org) library for modeling mathematical programs, builds on existing MPI implementations to ensure efficiency and scalability, and is available via http://github.com/Pyomo/mpi-sppy.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters

We present an optimized Floyd-Warshall (Floyd-Warshall) algorithm that computes the All-pairs shortest path (APSP) for GPU accelerated clusters. The Floyd-Warshall algorithm due to its structural similarities to matrix-multiplication is well suited for highly parallel GPU architectures. To achieve high parallel efficiency, we address two key algorithmic challenges: reducing high communication overhead and addressing limited GPU memory. To reduce high communication costs, we redesign the parallel (a) to expose more parallelism, (b) aggressively overlap communication and computation with pipelined and asynchronous scheduling of operations, and (c) tailored MPI-collective. To cope with limited GPU memory, we employ an offload model, where the data resides on the host and is transferred to GPU on-demand. The proposed optimizations are supported with detailed performance models for tuning. Our optimized parallel Floyd-Warshall implementation is up to 5x faster than a strong baseline and achieves 8.1 PetaFLOPS/sec on 256~nodes of the Summit supercomputer at Oak Ridge National Laboratory. This performance represents 70% of the theoretical peak and 80% parallel efficiency. The offload algorithm can handle 2.5x larger graphs with a 20% increase in overall running time.

Sao, Piyush↗

Differences In High Burnup Fuel Management Strategies to Minimize FFRD and Increase Economic Viability

The nuclear industry is pursuing approval of an increase in the length of the pressurized water reactor (PWR) cycle from 18 months to 24 months to reduce reactor downtime and enhance the economic competitiveness of nuclear energy. Such an increase in reactor cycle length will require that the maximum rod average burnup exceeds the current regulatory limit of 62 GWd/MTU, and it could peak at approximately 75 GWd/MTU, posing potential reactor safety and performance concerns. One such concern is that fuel fragmentation, relocation, and dispersal (FFRD) could occur during a severe loss-of coolant accident (LOCA) in which a fuel rod balloons and bursts, and pulverized fuel fragments are dispersed throughout the reactor’s primary coolant system. Previous analyses have identified which reactor operating conditions leave the core more susceptible to FFRD and have shown that FFRD susceptibility is strongly linked to fuel rod burnup and linear heat rate (LHR) history. The work described in this report uses an optimization strategy known as parallel simulated annealing (PSA) and a coarse mesh Purdue Advanced Reactor Core Simulator (PARCS) reactor physics model to develop two core fuel loading patterns, each with a different optimization objective. One core optimization maximized the core’s cycle length while still respecting regulatory limits on the radial peaking factor and soluble boron concentration with a peak rod average burnup of 75 GWd/MTU. The second optimization was aimed at minimizing FFRD susceptibility while still targeting a 24-month cycle length and respecting regulatory limits. PARCS model predictions were verified using the high-fidelity Virtual Environment for Reactor Applications (VERA). The two core designs were compared to highlight core design strategies to minimize FFRD susceptibility and to maximize economic viability.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗