Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and hermitian eigenproblems

The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. Furthermore, we benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU–GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.

97 MATHEMATICS AND COMPUTING↗

Solving larger maximum clique problems using parallel quantum annealing

Quantum annealing has the potential to find low energy solutions of NP-hard problems that can be expressed as quadratic unconstrained binary optimization problems. However, the hardware of the quantum annealer manufactured by D-Wave Systems, which we consider in this work, is sparsely connected and moderately sized (on the order of thousands of qubits), thus necessitating a minor-embedding of a logical problem onto the physical qubit hardware. The combination of relatively small hardware sizes and the necessity of a minor-embedding can mean that solving large optimization problems is not possible on current quantum annealers. In this research, we show that a hybrid approach combining parallel quantum annealing with graph decomposition allows one to solve larger optimization problem accurately. We apply the approach to the Maximum Clique problem on graphs with up to 120 nodes and 6395 edges.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Exploiting Power Flow Manifold to Solve AC Optimal Power Flow

AC optimal power flow has proven difficult to solve with interior point methods on GPUs. This is largely due to challenging linear algebra problems that current state of the art massively parallel linear solvers struggle with. However, the advent of Riemannian optimization techniques and the fact that the power flow equations form a smooth manifold present an alternative approach. In this talk, we present the basics of Riemannian optimization techniques in which optimization is done directly on a manifold. Then we present computational results showing that Riemannian techniques are capable of producing solutions of comparable quality as interior point methods.

AC optimal power flow↗

Adaptive, Active Learning, and Multifidelity Monte Carlo Methods in the MOOSE Stochastic Tools Module

MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.

97 MATHEMATICS AND COMPUTING↗

Optimization of the moderators in the STS preliminary design

This report details the results for an optimization of the dimensions of the moderators in the preliminary design of the Spallation Neutron Source Second Target Station (STS). This study uses the optimization algorithms of Dakota and an unstructured mesh model for the moderators in MCNP. More details on the unstructured mesh model and the automated mesh generation can be found in [3]. Parallel to this effort, the same moderator geometries have been optimized using a constructive solid geometry (CSG) MCNP model. More details on this model and its results can be found in [4]. Three optimal designs are selected for each moderator: one that is optimized for maximum peak brightness, one for maximum time-integrated brightness, and one for a combination of peak and time-integrated brightness. The backbone of the optimization work flow is provided by Dakota. For each set of design parameters requested by Dakota, a new solid geometry is automatically built in Creo and SpaceClaim, and subsequently exported to Attila4MC to generate an unstructured mesh geometry for MCNP. After the MCNP calculation is finished, the objective function (e.g., brightness metric) is returned to Dakota. After the new design has been evaluated, a result-file is written, and Dakota proposes the next set of design parameters to be evaluated. The loop continues until a specified convergence criterion has been met. The design parameters of the cylindrical (upper) moderator include the hydrogen radius, the premoderator thickness (top, bottom, radial), the beryllium radius and the horizontal position of the moderator. The crucial design choice is the hydrogen radius. A radius of 62 mm is shown to provide the maximum time-integrated brightness. The maximum peak brightness occurs with a radius of 40 mm. A combined (middle) design, which balances peak and time-integrated brightnesses, is obtained with a hydrogen radius of 50 mm. The premoderator thicknesses and the beryllium radius are slightly larger in the design optimized for time-integrated brightness than in the design optimized for peak brightness. The sensitivity to these two parameters is relatively small close to the optimal configurations. The hydrogen vessel and vacuum vessel wall thicknesses are dependent on the radius of the liquid hydrogen due to structural integrity requirements. The increased wall thicknesses for larger vessels significantly penalize the time-integrated brightness, with the maximum obtainable value reduced by more than 10% relative to earlier studies which used fixed vessel wall thicknesses. The impact of the variable wall thicknesses is much less for the peak brightness and combined brightness designs. The design parameters of the tube (lower) moderator selected for the optimization are the tube length, the annular premoderator thickness, the beryllium radius and the horizontal position of the moderator. The tube length is the crucial parameter and is chosen large (210 mm) and small (125 mm) in the designs optimized for time-integrated and peak brightness respectively. A combined optimal design has a tube length of 170 mm. The premoderator thickness and the beryllium radius are chosen larger in the design optimized for time-integrated brightness.

42 ENGINEERING↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Avatar Tools

Supervised machine learning is the process of using past experience to predict the future. "Ensembles" are a machine-learning meta-method that can be applied to most machine learning algorithms. Ensembles generally greatly improve accuracy, reduce or remove most of the design issues presented by machine learning, and are admirably suited to parallel and distributed computation. The Avatar Tools codes are an implementation of ensembles specifically for decision trees. Some features that distinguish Avatar Tools from other "ensembles for decision trees" codes are: (1) Does the bookkeeping necessary for out of bag (OOB) validation. (2) Can use OOB validation to automatically determine optimal ensemble size. (3) Provides an MPI-based parallel implementation, for distributed operation. (4) Provides convenient tools for cross-validation, to assess the accuracy provided by a training set. SAND2020-3858 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Siefert, Christopher↗

Invited Paper: Benchmarking and Optimizing Data Movement on Emerging Heterogeneous Architectures

As supercomputers evolve, nodes are continually increasing in complexity. As a result, each generation of parallel systems brings new performance challenges. For instance, on recent systems inter-node communication has outperformed inter-socket, resulting in poor performance of many node-aware communication optimizations. Communication optimizations are critical for the performance and scalability of parallel applications, but are dependent on the parallel architecture, which varies significantly among recent generations of supercomputers. Furthermore, this paper investigates the performance of various paths of data movement on recent generations of systems, and analyzes the increased complexity of communication, particularly on recent heterogeneous systems. The paper also introduces MPI Advance, a communication library that enables optimizations to be created based on benchmark analysis of each emerging system.

benchmarking↗

Optimized structure and electronic band gap of monolayer GeSe from quantum Monte Carlo methods

Here, we have used highly accurate quantum Monte Carlo methods to determine the chemical structure and electronic band gaps of monolayer GeSe. Two-dimensional (2D) monolayer GeSe has received a great deal of attention due to its unique thermoelectric, electronic, and optoelectronic properties with a wide range of potential applications. Density functional theory (DFT) methods have usually been applied to obtain optical and structural properties of bulk and 2D GeSe. For the monolayer, DFT typically yields a larger band-gap energy than for bulk GeSe but cannot conclusively determine if the monolayer has a direct or indirect gap. Moreover, the DFT-optimized lattice parameters and atomic coordinates for monolayer GeSe depend strongly on the choice of approximation for the exchange-correlation functional, which makes the ideal structure-and its electronic properties-unclear. In order to obtain accurate lattice parameters and atomic coordinates for the monolayer, we use a surrogate Hessian-based parallel line search within diffusion Monte Carlo to fully optimize the GeSe monolayer structure. The DMC-optimized structure is different from those obtained using DFT, as are calculated band gaps. The potential energy surface has a shallow minimum at the optimal structure. This, combined with the sensitivity of the electronic structure to strain, suggests that the optical properties of monolayer GeSe are highly tunable by strain.

36 MATERIALS SCIENCE↗

Optimizing Desalination Operations for Energy Flexibility

Despite the value of energy optimization in desalination processes, modeling dynamic operations for monthly billing periods has remained a computational challenge. This work proposes a framework for energy flexibility optimization, which includes new modeling features for independent operation of parallel skids, start-up delays associated with chemical stabilization, the consideration of industrial energy tariff structures, and inclusion of hourly electrical carbon intensities. This is done using a modular and computationally efficient formulation that guarantees a globally optimal solution with standard optimization solvers. In this study, the approach is demonstrated in two distinct case studies: a seawater desalination plant in Santa Barbara, CA, and an indirect potable reuse facility in San Jose, CA. Trends predicted from the model are validated against operational facility measurements from a demand response shutdown event. Preliminary results show that optimizing energy flexibility can result in 18.51% monthly cost savings over energy efficiency-optimized operation. The value extracted from a facility-wide shutdown during peak electricity price hours is hampered by start-up delays in post-treatment chemical stabilization. In cases in which a facility does not have much excess capacity, using a flow equalization tank or operating over a wide recovery range may be cost-effective.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel algorithms for finding connected components using linear algebra

Finding connected components is one of the most widely used operations on a graph. Optimal serial algorithms for the problem have been known for half a century, and many competing parallel algorithms have been proposed over the last several decades under various different models of parallel computation. This paper presents a class of parallel connected-component algorithms designed using linear-algebraic primitives. These algorithms are based on a PRAM algorithm by Shiloach and Vishkin and can be designed using standard GraphBLAS operations. Here, we demonstrate two algorithms of this class, one named LACC for Linear Algebraic Connected Components, and the other named FastSV which can be regarded as LACC’s simplification. With the support of the highly-scalable Combinatorial BLAS library, LACC and FastSV outperform the previous state-of-the-art algorithm by a factor of up to 12x for small to medium scale graphs. For large graphs with more than 50B edges, LACC and FastSV scale to 4K nodes (262K cores) of a Cray XC40 supercomputer and outperform previous algorithms by a significant margin. This remarkable performance is accomplished by (1) exploiting sparsity that was not present in the original PRAM algorithm formulation, (2) using high-performance primitives of Combinatorial BLAS, and (3) identifying hot spots and optimizing them away by exploiting algorithmic insights.

97 MATHEMATICS AND COMPUTING↗

A scoping study of far-SOL main-wall protection limiters for steady-state operation of compact pilot plant tokamaks

We present a novel method for handling steady-state heat fluxes incident on the main wall of pilot plant-scale magnetic fusion devices, based on the utilization of protection limiters in the far scrape-off layer (SOL). This method helps avoid large plasma-wall gaps, without excessively compromising blanket performance. We present an optimization algorithm for determining the appropriate size and scale of these protection limiters given (1) probability distributions of SOL plasma parameters and (2) assumed risk tolerance. As part of this optimization, we have developed an analytic description of parallel heat fluxes across limiter shadows, and an objective cost function (the ‘Far-SOL Marginal Cost’) to quantify the impact that different main-wall thermal management design choices have on reactor capital cost. Applying the model to a midscale fusion pilot plant concept shows that making use of far-SOL protection limiters can reduce capital costs on the order of $500 M, relative to naively increasing the plasma-wall gap. Our analysis demonstrates that the far-SOL power decay length is the highest-leverage plasma assumption for thermal loading of the first wall, and the primary cost driver for main wall thermal management. The relative cost efficiency of protection limiters increases as assumptions on the far-SOL heat flux become more pessimistic. The concepts described in this paper motivate the further development of far-SOL protection limiters as part of larger efforts to design economical core-edge-wall compatible solutions for a fusion pilot plant.

Design under uncertainty↗

Efficient smoothed particle radiation hydrodynamics I: Thermal radiative transfer

This work presents efficient solution techniques for radiative transfer in the smoothed particle hydrodynamics discretization. Two choices that impact efficiency are how the material and radiation energy are coupled, which determines the number of iterations needed to converge the emission source, and how the radiation diffusion equation is solved, which must be done in each iteration. The coupled material and radiation energy equations are solved using an inexact Newton iteration scheme based on nonlinear elimination, which reduces the number of Newton iterations needed to converge within each time step. During each Newton iteration, the radiation diffusion equation is solved using Krylov iterative methods with a multigrid preconditioner, which abstracts and optimizes much of the communication when running in parallel. The code is verified for an infinite medium problem, a one-dimensional Marshak wave, and a two and three-dimensional manufactured problem, and exhibits first-order convergence in time and second-order convergence in space. For these problems, the number of iterations needed to converge the inexact Newton scheme and the diffusion equation is independent of the number of spatial points and the number of processors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Successes and Opportunities for Discovery of Metal Oxide Photoanodes for Solar Fuels Generators

We report that the importance of metal oxide photoanodes in solar fuels technology has garnered concerted efforts in photoanode discovery in recent decades, which complement parallel efforts in development of analytical techniques and optimization strategies using standard photoanodes such as TiO 2 , Fe 2 O 3 and BiVO 4 . Theoretical guidance of high-throughput experiments has been particularly effective in dramatically increasing the portfolio of metal oxide photoanodes, motivating a new era of photoanode development where the characterization and optimization techniques developed on traditional materials are applied to nascent photoanodes that exhibit visible light photoresponse. The compendium of metal oxide photoanodes presented in the present work can also serve as the basis for further technique development, with a primary goal to establish workflows for discovery of materials that perform better against the critical criteria of operational stability, visible light photoresponse, and photovoltage suitable for tandem absorber architectures.

14 SOLAR ENERGY↗

Accelerating computational modeling and design of high-entropy alloys

High-entropy alloys, with N elements and compositions {$c_{ν = 1,N}$} in competing crystal structures, have large design spaces for unique chemical and mechanical properties. In this work, to enable computational design, we use a metaheuristic hybrid Cuckoo search (CS) to construct alloy configurational models on the fly that have targeted atomic site and pair probabilities on arbitrary crystal lattices, given by supercell random approximates (SCRAPs) with S sites. Our Hybrid CS permits efficient global solutions for large, discrete combinatorial optimization that scale linearly in a number of parallel processors, and linearly in sites S for SCRAPs. For example, a four-element, 128-site SCRAP is found in seconds—a more than 13,000-fold reduction over current strategies. Our method thus enables computational alloy design that is currently impractical. We qualify the models and showcase application to real alloys with targeted atomic short-range order. Being problem-agnostic, our Hybrid CS offers potential applications in diverse fields.

36 MATERIALS SCIENCE↗

Hvac: Removing I/O Bottleneck for Large-Scale Deep Learning Applications

Scientific communities are increasingly adopting deep learning (DL) models in their applications to accelerate scientific discovery processes. However, with rapid growth in the computing capabilities of HPC supercomputers, large-scale DL applications have to spend a significant portion of training time performing I/O to a parallel storage system. Previous research works have investigated optimization techniques such as prefetching and caching. Unfortunately, there exist non-trivial challenges to adopting the existing solutions on HPC supercomputers for large-scale DL training applications, which include non-performance and/or failures at extreme scale, lack of portability and generality in design, complex deployment methodology, and being limited to a specific application or dataset. To address these challenges, we propose High-Velocity AI Cache (HVAC), a distributed read-cache layer that targets and fully exploits the node-local storage or near node-local storage technology. HVAC seamlessly accelerates read I/O by aggregating node-local or near node-local storage, avoiding metadata lookups and file locking while preserving portability in the application code. We deploy and evaluate HVAC on 1,024 nodes (with over 6000 NVIDIA V100 GPUS) of the Summit supercomputer. In particular, we evaluate the scalability, efficiency, accuracy, and load distribution of HVAC compared to GPFS and XFS-on-NVMe. With four different DL applications, we observe an average 25 % performance improvement atop GPFS and 9% drop against XFS-on-NVMe, which scale linearly and are considered the performance upper bound. We envision HVAC as an important caching library for upcoming HPC supercomputers such as Frontier.

Khan, Awais↗

Solving Unit Commitment Problems with Demand Responsive Loads: Preprint

This paper focuses on using variations of the Frank-Wolfe algorithm for solving unit commitment problems with high volumes of demand responsive loads on the power grid. We present a formulation of the unit commitment problem with demand responsive loads. We then show through reformulation and relaxations of the problem that variations of the Frank-Wolfe algorithm can be used to determine the time series decisions for the demand responsive loads. We show through computational experiments on the RTS-GMLC test system that the time series of demand responsive load decisions obtained through our approach are near optimal and provide details regarding how large scale parallel implementations of our approach can be highly computationally efficient.

demand response↗

Efficient Ensemble-Based Stochastic Gradient Methods for Optimization Under Geological Uncertainty

Ensemble-based stochastic gradient methods, such as the ensemble optimization (EnOpt) method, the simplex gradient (SG) method, and the stochastic simplex approximate gradient (StoSAG) method, approximate the gradient of an objective function using an ensemble of perturbed control vectors. These methods are increasingly used in solving reservoir optimization problems because they are not only easy to parallelize and couple with any simulator but also computationally more efficient than the conventional finite-difference method for gradient calculations. In this work, we show that EnOpt may fail to achieve sufficient improvement of the objective function when the differences between the objective function values of perturbed control variables and their ensemble mean are large. On the basis of the comparison of EnOpt and SG, we propose a hybrid gradient of EnOpt and SG to save on the computational cost of SG. We also suggest practical ways to reduce the computational cost of EnOpt and StoSAG by approximating the objective function values of unperturbed control variables using the values of perturbed ones. We first demonstrate the performance of our improved ensemble schemes using a benchmark problem. Results show that the proposed gradients saved about 30–50% of the computational cost of the same optimization by using EnOpt, SG, and StoSAG. As a real application, we consider pressure management in carbon storage reservoirs, for which brine extraction wells need to be optimally placed to reduce reservoir pressure buildup while maximizing the net present value. Results show that our improved schemes reduce the computational cost significantly.

58 GEOSCIENCES↗