Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

m-CUBES An efficient and portable implementation of multi-dimensional integration for gpus

The task of multi-dimensional numerical integration is frequently encountered in physics and other scientific fields, e.g., in modeling the effects of systematic uncertainties in physical systems and in Bayesian parameter estimation. Multi-dimensional integration is often time-prohibitive on CPUs. Efficient implementation on many-core architectures is challenging as the workload across the integration space cannot be predicted a priori. We propose m-Cubes, a novel implementation of the well-known Vegas algorithm for execution on GPUs. Vegas transforms integration variables followed by calculation of a Monte Carlo integral estimate using adaptive partitioning of the resulting space. m-Cubes improves performance on GPUs by maintaining relatively uniform workload across the processors. As a result, our optimized Cuda implementation for Nvidia GPUs outperforms parallelization approaches proposed in past literature. We further demonstrate the efficiency of m-Cubes by evaluating a six-dimensional integral from a cosmology application, achieving significant speedup and greater precision than the CUBA library's CPU implementation of VEGAS. We also evaluate m-Cubes on a standard integrand test suite. m-Cubes outperforms the serial implementations of the Cuba and GSL libraries by orders of magnitude speedup while maintaining comparable accuracy. Our approach yields a speedup of at least 10 when compared against publicly available Monte Carlo based GPU implementations. In summary, m-Cubes can solve integrals that are prohibitively expensive using standard libraries and custom implementations. A modern C++ interface header-only implementation makes m-Cubes portable, allowing its utilization in complicated pipelines with easy to define stateful integrals. Compatibility with non-Nvidia GPUs is achieved with our initial implementation of m-Cubes using the Kokkos framework.

Sakiotis, Ioannis↗

Case Study of Using Kokkos and SYCLs Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

Six of the top ten supercomputers in the TOP500 list from June 2021 rely on NVIDIA GPUs to achieve their peak compute bandwidth. With the announcement of Aurora, Frontier, and El Capitan, Intel and AMD have also entered the domain of providing GPUs for scientific computing. A consequence of the increased diversity in the GPU landscape is the emergence of portable programming models such as Kokkos, SYCL, OpenCL, and OpenMP, which allow application developers to maintain a single-source code across a diverse range of hardware architectures. While the portable frameworks try to optimize the compute resource usage on a given architecture, it is the programmers responsibility to expose parallelism in an application that can take advantage of thousands of processing elements available on GPUs. In this paper, we introduce a GPU-friendly parallel implementation of Milc-Dslash that exposes multiple hierarchies of parallelism in the algorithm. Milc-Dslash was designed to serve as a benchmark with highly optimized matrix-vector multiplications to measure the resource utilization on the GPU systems. The parallel hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware using Kokkos and SYCL programming models. We present the performance achieved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU, AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos and SYCL performances with those obtained from the versions written in CUDA and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectively.

Dufek, Amanda S↗

Parallel computation of manipulator inverse dynamics

In this article, parallel computation of manipulator inverse dynamics is investigated. A hierarchical graph-based mapping approach is devised to analyze the inherent parallelism in the Newton-Euler formulation at several computational levels, and to derive the features of an abstract architecture for exploitation of parallelism. At each level, a parallel algorithm represents the application of a parallel model of computation that transforms the computation into a graph whose structure defines the features of an abstract architecture, i.e., number of processors, communication structure, etc. Data-flow analysis is employed to derive the time lower bound in the computation as well as the sequencing of the abstract architecture. The features of the target architecture are defined by optimization of the abstract architecture to exploit maximum parallelism while minimizing architectural complexity. An architecture is designed and implemented that is capable of efficient exploitation of parallelism at several computational levels. The computation time of the Newton-Euler formulation for a 6-degree-of-freedom (dof) general manipulator is measured as 187 microsec. The increase in computation time for each additional dof is 23 microsec, which leads to a computation time of less than 500 microsec, even for a 12-dof redundant arm.

Fijany, Amir↗

Design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a three-dimensional unstructured-grid Euler solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data structure has been adopted, and a number of optimizations have been devised and implemented to accelerate the parallel computational rates. The implementation is carried out by creating a set of software tools, which provide an interface between the parallelization issues and the sequential code, while providing a basis for future automatic run-time compilation support. Large practical unstructured grid problems are solved on the Intel iPSC/860 hypercube and Intel Touchstone Delta machine. The quantitative effects of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of 3 performance improvement. The overall solution efficiency is compared with that obtained on the Cray Y-MP vector supercomputer.

Das, R.↗

Enabling Ultra-Compact, Lightweight, Efficient, and Reliable 6.6 kW On-Board Bi-Directional Electric Vehicle Charger with Advanced Topology and Control

The research explored new topologies, control methods, mechanical integration, and thermal management methods for electric vehicle (EV) on-board chargers. The team investigated capacitor-based power conversion, leveraging the high energy densities inherent to capacitive energy storage compared to inductive methods. The proposed topologies simultaneously enabled high power density and high efficiency of the design. The proposed architecture was demonstrated in a 6.6 kW bi-directional charger prototype. The research pursued several directions to improve system performance. Innovative topologies were studied for both the main power conversion stage as well as the single-phase twice-line-frequency energy buffer. To ensure robust and efficient operation, new control methods were developed to integrate these two subsystems. To achieve high power density in the full system solution, the mechanical structure of the charger is highly optimized to maximally fill the converter box volume. In parallel with the mechanical design effort, the converter was packaged with high-performance cooling methods which removed heat from key areas of power dissipation in the converter. The thermal management system was optimized to minimize its weight and volume, ultimately motivating the design of a custom additively manufactured cold-plate. The full system achieves a peak power of 7 kW with less than 0.3% total harmonic distortion (THD) and greater than 0.994 power factor in power factor correction (PFC) operation, corresponding to a total box-volume power density of 47.9 kW/L and gravimetric power density of 24.6 W/g. The system achieves a peak efficiency of 98.9%, with 97.9% efficiency at maximum power.

33 ADVANCED PROPULSION SYSTEMS↗

On the accuracy of solving triangular systems in parallel

An error complexity analysis of two algorithms for solving a unit-diagonal triangular system is given. The results show that the unusual sequential algorithm is optimal in terms of having the minimal maximum and cumulative error complexity measures. The parallel algorithm described by Sameh and Brent is shown to be essentially equivalent to the optimal sequential one. Some numerical experiments are also taught.

Tsao, Nai-Kuan↗

New Results on Communication- and Memory-Aware Load Balancing Model and Algorithms

While load balancing in distributed-memory computing has been well-studied, we present an innovative approach to this problem: a unified, reduced-order model that combines three key components to describe “work” in a distributed system: computation, communication, and memory. Our model enables an optimizer to explore complex tradeoffs in task placement, such as augmented parallelism, at the expense of data replication increasing memory usage. We propose a fully distributed, heuristic-based load balancing optimization algorithm, and demonstrate that it quickly finds close-to-optimal solutions. We formalize the complex optimization problem as a mixed-integer linear program, and compare it to our strategy. Finally, we show that when applied to an electromagnetics code, our approach obtains up to 2.3x speedups for the imbalanced execution.

97 MATHEMATICS AND COMPUTING↗

High speed civil transport: Sonic boom softening and aerodynamic optimization

An improvement in sonic boom extrapolation techniques has been the desire of aerospace designers for years. This is because the linear acoustic theory developed in the 60's is incapable of predicting the nonlinear phenomenon of shock wave propagation. On the other hand, CFD techniques are too computationally expensive to employ on sonic boom problems. Therefore, this research focused on the development of a fast and accurate sonic boom extrapolation method that solves the Euler equations for axisymmetric flow. This new technique has brought the sonic boom extrapolation techniques up to the standards of the 90's. Parallel computing is a fast growing subject in the field of computer science because of its promising speed. A new optimizer (IIOWA) for the parallel computing environment has been developed and tested for aerodynamic drag minimization. This is a promising method for CFD optimization making use of the computational resources of workstations, which unlike supercomputers can spend most of their time idle. Finally, the OAW concept is attractive because of its overall theoretical performance. In order to fully understand the concept, a wind-tunnel model was built and is currently being tested at NASA Ames Research Center. The CFD calculations performed under this cooperative agreement helped to identify the problem of the flow separation, and also aided the design by optimizing the wing deflection for roll trim.

Cheung, Samson↗

Optimal rocket thrust profile shaping using third degree spline function interpolation

Optimal solid-rocket thrust profiles for the parallel-burn, solid-rocket-assisted space shuttle are investigated. Solid-rocket thrust profiles are simulated by using third-degree spline functions, with the values of the thrust ordinates defined as parameters. The profiles are optimized parametrically, using the Davidon-Fletcher-Powell penalty function method, by minimizing propellant weight subject to state and control inequality constraints and to terminal boundary conditions. This study shows that optimizing a control variable parametrically by using third-degree spline function interpolation allows the control to be shaped so that inequality constraints are strictly adhered to and all corners are eliminated. The absence of corners, which is realistic in nature, makes this method attractive from the viewpoint of solid rocket grain design.

Johnson, I. L.↗

Multimodal Bayesian registration of noisy functions using Hamiltonian Monte Carlo

Functional data registration is a necessary processing step for many applications. The observed data can be inherently noisy, often due to measurement error or natural process uncertainty; which most functional alignment methods cannot handle. A pair of functions can also have multiple optimal alignment solutions, which is not addressed in current literature. In this paper, a flexible Bayesian approach to functional alignment is presented, which appropriately accounts for noise in the data without any pre-smoothing required. Additionally, by running parallel MCMC chains, the method can account for multiple optimal alignments via the multi-modal posterior distribution of the warping functions. To most efficiently sample the warping functions, the approach relies on a modification of the standard Hamiltonian Monte Carlo to be well-defined on the infinite-dimensional Hilbert space. In this work, this flexible Bayesian alignment method is applied to both simulated data and real data sets to show its efficiency in handling noisy functions and successfully accounting for multiple optimal alignments in the posterior; characterizing the uncertainty surrounding the warping functions.

97 MATHEMATICS AND COMPUTING↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and hermitian eigenproblems

The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. Furthermore, we benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU–GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.

97 MATHEMATICS AND COMPUTING↗

Solving larger maximum clique problems using parallel quantum annealing

Quantum annealing has the potential to find low energy solutions of NP-hard problems that can be expressed as quadratic unconstrained binary optimization problems. However, the hardware of the quantum annealer manufactured by D-Wave Systems, which we consider in this work, is sparsely connected and moderately sized (on the order of thousands of qubits), thus necessitating a minor-embedding of a logical problem onto the physical qubit hardware. The combination of relatively small hardware sizes and the necessity of a minor-embedding can mean that solving large optimization problems is not possible on current quantum annealers. In this research, we show that a hybrid approach combining parallel quantum annealing with graph decomposition allows one to solve larger optimization problem accurately. We apply the approach to the Maximum Clique problem on graphs with up to 120 nodes and 6395 edges.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Exploiting Power Flow Manifold to Solve AC Optimal Power Flow

AC optimal power flow has proven difficult to solve with interior point methods on GPUs. This is largely due to challenging linear algebra problems that current state of the art massively parallel linear solvers struggle with. However, the advent of Riemannian optimization techniques and the fact that the power flow equations form a smooth manifold present an alternative approach. In this talk, we present the basics of Riemannian optimization techniques in which optimization is done directly on a manifold. Then we present computational results showing that Riemannian techniques are capable of producing solutions of comparable quality as interior point methods.

AC optimal power flow↗

Adaptive, Active Learning, and Multifidelity Monte Carlo Methods in the MOOSE Stochastic Tools Module

MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.

97 MATHEMATICS AND COMPUTING↗

Optimization of the moderators in the STS preliminary design

This report details the results for an optimization of the dimensions of the moderators in the preliminary design of the Spallation Neutron Source Second Target Station (STS). This study uses the optimization algorithms of Dakota and an unstructured mesh model for the moderators in MCNP. More details on the unstructured mesh model and the automated mesh generation can be found in [3]. Parallel to this effort, the same moderator geometries have been optimized using a constructive solid geometry (CSG) MCNP model. More details on this model and its results can be found in [4]. Three optimal designs are selected for each moderator: one that is optimized for maximum peak brightness, one for maximum time-integrated brightness, and one for a combination of peak and time-integrated brightness. The backbone of the optimization work flow is provided by Dakota. For each set of design parameters requested by Dakota, a new solid geometry is automatically built in Creo and SpaceClaim, and subsequently exported to Attila4MC to generate an unstructured mesh geometry for MCNP. After the MCNP calculation is finished, the objective function (e.g., brightness metric) is returned to Dakota. After the new design has been evaluated, a result-file is written, and Dakota proposes the next set of design parameters to be evaluated. The loop continues until a specified convergence criterion has been met. The design parameters of the cylindrical (upper) moderator include the hydrogen radius, the premoderator thickness (top, bottom, radial), the beryllium radius and the horizontal position of the moderator. The crucial design choice is the hydrogen radius. A radius of 62 mm is shown to provide the maximum time-integrated brightness. The maximum peak brightness occurs with a radius of 40 mm. A combined (middle) design, which balances peak and time-integrated brightnesses, is obtained with a hydrogen radius of 50 mm. The premoderator thicknesses and the beryllium radius are slightly larger in the design optimized for time-integrated brightness than in the design optimized for peak brightness. The sensitivity to these two parameters is relatively small close to the optimal configurations. The hydrogen vessel and vacuum vessel wall thicknesses are dependent on the radius of the liquid hydrogen due to structural integrity requirements. The increased wall thicknesses for larger vessels significantly penalize the time-integrated brightness, with the maximum obtainable value reduced by more than 10% relative to earlier studies which used fixed vessel wall thicknesses. The impact of the variable wall thicknesses is much less for the peak brightness and combined brightness designs. The design parameters of the tube (lower) moderator selected for the optimization are the tube length, the annular premoderator thickness, the beryllium radius and the horizontal position of the moderator. The tube length is the crucial parameter and is chosen large (210 mm) and small (125 mm) in the designs optimized for time-integrated and peak brightness respectively. A combined optimal design has a tube length of 170 mm. The premoderator thickness and the beryllium radius are chosen larger in the design optimized for time-integrated brightness.

42 ENGINEERING↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Avatar Tools

Supervised machine learning is the process of using past experience to predict the future. "Ensembles" are a machine-learning meta-method that can be applied to most machine learning algorithms. Ensembles generally greatly improve accuracy, reduce or remove most of the design issues presented by machine learning, and are admirably suited to parallel and distributed computation. The Avatar Tools codes are an implementation of ensembles specifically for decision trees. Some features that distinguish Avatar Tools from other "ensembles for decision trees" codes are: (1) Does the bookkeeping necessary for out of bag (OOB) validation. (2) Can use OOB validation to automatically determine optimal ensemble size. (3) Provides an MPI-based parallel implementation, for distributed operation. (4) Provides convenient tools for cross-validation, to assess the accuracy provided by a training set. SAND2020-3858 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Siefert, Christopher↗