Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters

We present an optimized Floyd-Warshall (Floyd-Warshall) algorithm that computes the All-pairs shortest path (APSP) for GPU accelerated clusters. The Floyd-Warshall algorithm due to its structural similarities to matrix-multiplication is well suited for highly parallel GPU architectures. To achieve high parallel efficiency, we address two key algorithmic challenges: reducing high communication overhead and addressing limited GPU memory. To reduce high communication costs, we redesign the parallel (a) to expose more parallelism, (b) aggressively overlap communication and computation with pipelined and asynchronous scheduling of operations, and (c) tailored MPI-collective. To cope with limited GPU memory, we employ an offload model, where the data resides on the host and is transferred to GPU on-demand. The proposed optimizations are supported with detailed performance models for tuning. Our optimized parallel Floyd-Warshall implementation is up to 5x faster than a strong baseline and achieves 8.1 PetaFLOPS/sec on 256~nodes of the Summit supercomputer at Oak Ridge National Laboratory. This performance represents 70% of the theoretical peak and 80% parallel efficiency. The offload algorithm can handle 2.5x larger graphs with a 20% increase in overall running time.

Sao, Piyush↗

A History of Orion Mission Design, Copernicus Software Development, and the Artemis I Trajectory

This paper describes the history of the on-orbit trajectory design and optimization for the Orion spacecraft at NASA JSC, from the initial design through the execution of the Artemis I test flight. In parallel, the Copernicus trajectory optimization tool was also being developed and was the main tool used for Orion trajectory design during this period. Finally, the paper gives an overview of the Artemis I trajectory that was flown during the Artemis I mission from November 16 - December 11, 2022.

Orion↗

Optimal Control within the Context of Multidisciplinary Design, Analysis, and Optimization

Multidisciplinary design, analysis and optimization involves modeling the interactions of complex systems across a variety of disciplines. The optimization of such systems can be a computationally expensive exercise with multiple levels of nested nonlinear solvers running under an optimizer.The application of optimal control in project development often involves performing trajectory optimization for fixed vehicle designs or parametric sweeps across some key vehicle properties.This information is then relayed to the subsystem design teams who update their designs and relay some bulk characteristics back to the trajectory optimization procedure.This iteration is then repeated until the design closes.However, with increasing interest in more tightly coupled systems, such as electric and hybrid-electric aircraft propulsion and boundary layer ingestion, this process is prone to ignore subtle coupling between vehicle subsystem designs and vehicle operation on a given mission.Integrating trajectory optimization into a tightly coupled multidisciplinary design procedure can be computationally prohibitive, depending on the complexity of the subsystem analyses and the optimal control technique applied.To address these issues a new optimal control software tool, Dymos, has been developed.Dymos is built upon NASA's OpenMDAO software and can leverage its capabilities to efficiently compute gradients for the optimization and optimize complex models in parallel on distributed memory systems.This report provides some explanation into the numerical methods employed in Dymos and provides several use cases that demonstrate its performance on traditional optimal control problems and improvements ino techniques have been used extensively in recent decades to solve a variety of optimal control problems, typically in the form of aerospace vehicle trajectory optimization.

pseudospectral↗

Differences In High Burnup Fuel Management Strategies to Minimize FFRD and Increase Economic Viability

The nuclear industry is pursuing approval of an increase in the length of the pressurized water reactor (PWR) cycle from 18 months to 24 months to reduce reactor downtime and enhance the economic competitiveness of nuclear energy. Such an increase in reactor cycle length will require that the maximum rod average burnup exceeds the current regulatory limit of 62 GWd/MTU, and it could peak at approximately 75 GWd/MTU, posing potential reactor safety and performance concerns. One such concern is that fuel fragmentation, relocation, and dispersal (FFRD) could occur during a severe loss-of coolant accident (LOCA) in which a fuel rod balloons and bursts, and pulverized fuel fragments are dispersed throughout the reactor’s primary coolant system. Previous analyses have identified which reactor operating conditions leave the core more susceptible to FFRD and have shown that FFRD susceptibility is strongly linked to fuel rod burnup and linear heat rate (LHR) history. The work described in this report uses an optimization strategy known as parallel simulated annealing (PSA) and a coarse mesh Purdue Advanced Reactor Core Simulator (PARCS) reactor physics model to develop two core fuel loading patterns, each with a different optimization objective. One core optimization maximized the core’s cycle length while still respecting regulatory limits on the radial peaking factor and soluble boron concentration with a peak rod average burnup of 75 GWd/MTU. The second optimization was aimed at minimizing FFRD susceptibility while still targeting a 24-month cycle length and respecting regulatory limits. PARCS model predictions were verified using the high-fidelity Virtual Environment for Reactor Applications (VERA). The two core designs were compared to highlight core design strategies to minimize FFRD susceptibility and to maximize economic viability.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

PRO-X Parallelization Study

The proliferation resistance optimization (PRO-X) program is actively supporting the design of nuclear systems by developing a framework to both optimize the fuel cycle infrastructure for nuclear reactor (including both advanced reactors (ARs) and research reactors (RRs)) and minimize the potential for production of weapons-usable nuclear material (Figure 1). One area of interest is in the impact a modular approach to bulk handling fuel cycle facilities could have on meeting safeguards requirements to identify future areas of growth within the proliferation resistance space. This study evaluates how changing the number of streams within a fuel cycle facility could impact a facilities ability to meet both domestic and international safeguards requirements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Importance of Higher Fidelity Model Geometries during Optimization of Critical Experiments

PARADIGM, PARallel Approach of Differential and InteGral Measurements, is a cross-collaborative effort at Los Alamos National Laboratory between nuclear data theorists, differential and integral experimenters, as well as machine learning statisticians to tackle uncertainties in the intermediate region of 239 Pu. In essence, the idea behind PARADIGM is to remove the linear conceptualization of the nuclear data pipeline, shown in Figure 1, and replace it with a far more parallelized approach. The novel approach leverages machine learning to guide which differential measurements and integral experiments will result in the largest decrease in uncertain ties for a nuclide reaction pair in a given energy range. The concept builds off earlier work, EUCLID, which focused on the fast region of 239 Pu. The practical benefit of having evaluation, differential measurement, and integral experiment personnel in collaboration with machine learning is to represent the entire nuclear data in one snapshot. This enable large reduction in the time to deliver improved nuclear data, which using the PARADIGM approach could be done in 3 years. A general outline of PARADIGM and specific topics are available in other papers. The discussion here will pertain directly to the integral experiment design. More specifically, the process of taking a rough design and transforming it into a finalized neutronic model will be discussed.

97 MATHEMATICS AND COMPUTING↗

TEAM Project Review, Year 2

This report summarizes our research activities within the TEAM project between December 2020 and December 2021, funded by the ASCR Advanced Research in Quantum Computing program. During the reporting period the LLNL-MSU team has made progress on several fronts. An overarching goal of the team is to provide a comprehensive suite of software tools that can be used for the Characterize-Optimize-Compute loop needed to implement and execute algorithms on quantum devices. We are concurrently developing lightweight solvers that can be used on desktop computers to find optimal control pulses and to characterize small quantum systems (consisting of a few transmons and cavities). However, desktop computers are insufficient for simulating and characterizing larger quantum systems. We have therefore also developed parallel, distributed memory, simulators and optimization solvers, both for open and closed quantum systems. These parallel solvers have, for example, been used to study quantum optimal control for pure-state preparation, utilizing 1000’s of cores on a modern high-performance computing (HPC) platform.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Optimizing the hypre solver for manycore and GPU architectures

The solution of large-scale combustion problems with codes such as Uintah on modern computer architectures requires the use of multithreading and GPUs to achieve performance. Uintah uses a low-Mach number approximation that requires iteratively solving a large system of linear equations. The Hypre iterative solver has solved such systems in a scalable way for Uintah, but the use of OpenMP with Hypre leads to at least slowdown due to OpenMP overheads. The proposed solution uses the MPI Endpoints within Hypre, where each team of threads acts as a different MPI rank. This approach minimizes OpenMP synchronization overhead and performs as fast or (up to 1.44) faster than Hypre's MPI-only version, and allows the rest of Uintah to be optimized using OpenMP. The profiling of the GPU version of Hypre shows the bottleneck to be the launch overhead of thousands of micro-kernels. The GPU performance was improved by fusing these micro-kernels and was further optimized by using Cuda-aware MPI, resulting in an overall speedup of 1.16—1.44 compared to the baseline GPU implementation. The above optimization strategies were published in the International Conference on Computational Science 2020 [1]. This work extends the previously published research by carrying out the second phase of communication-centered optimizations in Hypre to improve its scalability on large-scale supercomputers. Additionally, this includes an efficient non-blocking inter-thread communication scheme, communication-reducing patch assignment, and expression of logical communication parallelism to a new version of the MPICH library that utilizes the underlying network parallelism [2]. The above optimizations avoid communication bottlenecks previously observed during strong scaling and improve performance by up to 2 on 256 nodes of Intel Knight's Landing processor.

97 MATHEMATICS AND COMPUTING↗

Developing Information Power Grid Based Algorithms and Software

This exploratory study initiated our effort to understand performance modeling on parallel systems. The basic goal of performance modeling is to understand and predict the performance of a computer program or set of programs on a computer system. Performance modeling has numerous applications, including evaluation of algorithms, optimization of code implementations, parallel library development, comparison of system architectures, parallel system design, and procurement of new systems. Our work lays the basis for the construction of parallel libraries that allow for the reconstruction of application codes on several distinct architectures so as to assure performance portability. Following our strategy, once the requirements of applications are well understood, one can then construct a library in a layered fashion. The top level of this library will consist of architecture-independent geometric, numerical, and symbolic algorithms that are needed by the sample of applications. These routines should be written in a language that is portable across the targeted architectures.

Dongarra, Jack↗

Machine learning-based optimization of air-cooled heat sinks

Machine learning-based models using Artificial Neural Network (ANN) and greedy search algorithm are used to optimize air-cooled parallel plate-finned heat sinks (PPFHSs) subjected to laminar flow over an extensive range of design parameters. Here, the thermal and hydraulic performances of PPFHSs are represented by heat transfer coefficient (h) and pressure drop (ΔP), respectively. Optimization objectives for PPFHS designs can vary from industry to industry depending on their design priorities. The present study proposes a novel and generalized optimization method that defines practical optimization objectives and provides an accurate optimization process to design effective PPFHSs for a wide range of industrial applications with different design requirements. Three optimization objectives are presented in this study: (i) the largest h PΔ, (ii) the largest h within a specified maximum allowed flow rate, and (iii) the lowest weight that maximizes h for operation within the maximum allowed flow rate. While the shortcoming of the first objective is demonstrated, the other two objectives are found to be suitable for designing effective heat sinks (HSs) across different applications. Results suggest a promising trend from the third objective to develop HSs with ~ 37-68% lower weight, 80-85% reduced ΔP, and negligible penalty in h compared with optimized HSs obtained from the second objective. However, since the third objective leads to HSs with thinner fins, structural analysis should be performed to ensure reliable operation of the HSs.

42 ENGINEERING↗

The Sizing and Optimization Language, (SOL): Computer language for design problems

The Sizing and Optimization Language, (SOL), a new high level, special purpose computer language was developed to expedite application of numerical optimization to design problems and to make the process less error prone. SOL utilizes the ADS optimization software and provides a clear, concise syntax for describing an optimization problem, the OPTIMIZE description, which closely parallels the mathematical description of the problem. SOL offers language statements which can be used to model a design mathematically, with subroutines or code logic, and with existing FORTRAN routines. In addition, SOL provides error checking and clear output of the optimization results. Because of these language features, SOL is best suited to model and optimize a design concept when the model consits of mathematical expressions written in SOL. For such cases, SOL's unique syntax and error checking can be fully utilized. SOL is presently available for DEC VAX/VMS systems. A SOL package is available which includes the SOL compiler, runtime library routines, and a SOL reference manual.

Lucas, Stephen H.↗

Three-dimensional Finite Element Formulation and Scalable Domain Decomposition for High Fidelity Rotor Dynamic Analysis

This paper has two objectives. The first objective is to formulate a 3-dimensional Finite Element Model for the dynamic analysis of helicopter rotor blades. The second objective is to implement and analyze a dual-primal iterative substructuring based Krylov solver, that is parallel and scalable, for the solution of the 3-D FEM analysis. The numerical and parallel scalability of the solver is studied using two prototype problems - one for ideal hover (symmetric) and one for a transient forward flight (non-symmetric) - both carried out on up to 48 processors. In both hover and forward flight conditions, a perfect linear speed-up is observed, for a given problem size, up to the point of substructure optimality. Substructure optimality and the linear parallel speed-up range are both shown to depend on the problem size as well as on the selection of the coarse problem. With a larger problem size, linear speed-up is restored up to the new substructure optimality. The solver also scales with problem size - even though this conclusion is premature given the small prototype grids considered in this study.

Datta, Anubhav↗

Multi-objective optimization with an integrated electromagnetics and beam dynamics workflow

In particle accelerators, RF cavities are used to accelerate charged particle beams to designed high energy for physical applications. In a typical accelerator design, the optimization of RF cavities and the optimization of beam dynamics are carried out in separate studies. For a more general and unrestricted accelerator design, a coupled optimization of the RF cavities and the beam parameters is required. For this coupled optimization problem, we have developed an integrated electromagnetics and beam dynamics workflow management system. Within this system, the geometries for a set of cavity components are first adjusted; the field modes are then computed with an electromagnetics program, and imported into a beam dynamics program for beam dynamics simulation. This workflow is encapsulated into a parallel multi-objective optimizer to achieve the integrated accelerator design optimization. A multi fidelity strategy is developed to improve the speed of the optimizer. Furthermore, this integrated global optimization capability is illustrated using a photoinjector design example and yields an improved design.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Homotopy Solver

This software implements parallel versions of an interior-point solver, based on the publicly available ipopt solver. Here we have full control over the linear solver and our algorithm is fully parallel thus enabling scalability to large-scale optimization problems. This package also has a parallel implementation of a homotopy solver developed under the scalable methods for contact LDRD project 23-ERD-017. This solver is an mfem-based implementation of algorithm described in ``A filter trust-region Newton continuation method for nonlinear complementarity problems''. Cosmin G. Petra, Nai-Yuan Chiang, Jingyi Wang, Tucker Hartland, and Michael Puso (submitted), LLNL-JRNL-869761.

Hartland, Tucker [Lawrence Livermore National Labo↗

Parallel IO Libraries for Managing HEP Experimental Data

The computing and storage requirements of the energy and intensity frontiers will grow significantly during the Run 4 & 5 and the HL-LHC era. Similarly, in the intensity frontier, with larger trig ger readouts during supernovae explosions, the Deep Underground Neutrino Experiment (DUNE) will have unique computing challenges that could be addressed by the use of parallel and accelerated dataprocessing capabilities. Most of the requirements of the energy and intensity frontier experiments rely on increasing the role of high performance computing (HPC) in the HEP community. In this presentation, we will describe our ongoing efforts that are focused on using HPC resources for the next generation HEP experiments. The HEPCCE (High Energy Physics-Center for Computational Excellence) IOS (Input/Output and Storage) group has been developing approaches to map HEP data to the HDF5 , an IO library optimized for the HPC platforms to store the intermediate HEP data. The complex HEP data products are serialized using ROOT to allow for experiment independent general mapping approaches of the HEP data to the HDF5 format. The mapping approaches can be optimized for high performance parallel IO. Similarly, simpler data can be directly mapped into the HDF5, which can also be suitable for offloading into the GPUs directly. We will present our works on both complex and simple data model models.

Bashyal, Amit↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗