Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

PRO-X Parallelization Study

The proliferation resistance optimization (PRO-X) program is actively supporting the design of nuclear systems by developing a framework to both optimize the fuel cycle infrastructure for nuclear reactor (including both advanced reactors (ARs) and research reactors (RRs)) and minimize the potential for production of weapons-usable nuclear material (Figure 1). One area of interest is in the impact a modular approach to bulk handling fuel cycle facilities could have on meeting safeguards requirements to identify future areas of growth within the proliferation resistance space. This study evaluates how changing the number of streams within a fuel cycle facility could impact a facilities ability to meet both domestic and international safeguards requirements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Importance of Higher Fidelity Model Geometries during Optimization of Critical Experiments

PARADIGM, PARallel Approach of Differential and InteGral Measurements, is a cross-collaborative effort at Los Alamos National Laboratory between nuclear data theorists, differential and integral experimenters, as well as machine learning statisticians to tackle uncertainties in the intermediate region of 239 Pu. In essence, the idea behind PARADIGM is to remove the linear conceptualization of the nuclear data pipeline, shown in Figure 1, and replace it with a far more parallelized approach. The novel approach leverages machine learning to guide which differential measurements and integral experiments will result in the largest decrease in uncertain ties for a nuclide reaction pair in a given energy range. The concept builds off earlier work, EUCLID, which focused on the fast region of 239 Pu. The practical benefit of having evaluation, differential measurement, and integral experiment personnel in collaboration with machine learning is to represent the entire nuclear data in one snapshot. This enable large reduction in the time to deliver improved nuclear data, which using the PARADIGM approach could be done in 3 years. A general outline of PARADIGM and specific topics are available in other papers. The discussion here will pertain directly to the integral experiment design. More specifically, the process of taking a rough design and transforming it into a finalized neutronic model will be discussed.

97 MATHEMATICS AND COMPUTING↗

TEAM Project Review, Year 2

This report summarizes our research activities within the TEAM project between December 2020 and December 2021, funded by the ASCR Advanced Research in Quantum Computing program. During the reporting period the LLNL-MSU team has made progress on several fronts. An overarching goal of the team is to provide a comprehensive suite of software tools that can be used for the Characterize-Optimize-Compute loop needed to implement and execute algorithms on quantum devices. We are concurrently developing lightweight solvers that can be used on desktop computers to find optimal control pulses and to characterize small quantum systems (consisting of a few transmons and cavities). However, desktop computers are insufficient for simulating and characterizing larger quantum systems. We have therefore also developed parallel, distributed memory, simulators and optimization solvers, both for open and closed quantum systems. These parallel solvers have, for example, been used to study quantum optimal control for pure-state preparation, utilizing 1000’s of cores on a modern high-performance computing (HPC) platform.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Optimizing the hypre solver for manycore and GPU architectures

The solution of large-scale combustion problems with codes such as Uintah on modern computer architectures requires the use of multithreading and GPUs to achieve performance. Uintah uses a low-Mach number approximation that requires iteratively solving a large system of linear equations. The Hypre iterative solver has solved such systems in a scalable way for Uintah, but the use of OpenMP with Hypre leads to at least slowdown due to OpenMP overheads. The proposed solution uses the MPI Endpoints within Hypre, where each team of threads acts as a different MPI rank. This approach minimizes OpenMP synchronization overhead and performs as fast or (up to 1.44) faster than Hypre's MPI-only version, and allows the rest of Uintah to be optimized using OpenMP. The profiling of the GPU version of Hypre shows the bottleneck to be the launch overhead of thousands of micro-kernels. The GPU performance was improved by fusing these micro-kernels and was further optimized by using Cuda-aware MPI, resulting in an overall speedup of 1.16—1.44 compared to the baseline GPU implementation. The above optimization strategies were published in the International Conference on Computational Science 2020 [1]. This work extends the previously published research by carrying out the second phase of communication-centered optimizations in Hypre to improve its scalability on large-scale supercomputers. Additionally, this includes an efficient non-blocking inter-thread communication scheme, communication-reducing patch assignment, and expression of logical communication parallelism to a new version of the MPICH library that utilizes the underlying network parallelism [2]. The above optimizations avoid communication bottlenecks previously observed during strong scaling and improve performance by up to 2 on 256 nodes of Intel Knight's Landing processor.

97 MATHEMATICS AND COMPUTING↗

Machine learning-based optimization of air-cooled heat sinks

Machine learning-based models using Artificial Neural Network (ANN) and greedy search algorithm are used to optimize air-cooled parallel plate-finned heat sinks (PPFHSs) subjected to laminar flow over an extensive range of design parameters. Here, the thermal and hydraulic performances of PPFHSs are represented by heat transfer coefficient (h) and pressure drop (ΔP), respectively. Optimization objectives for PPFHS designs can vary from industry to industry depending on their design priorities. The present study proposes a novel and generalized optimization method that defines practical optimization objectives and provides an accurate optimization process to design effective PPFHSs for a wide range of industrial applications with different design requirements. Three optimization objectives are presented in this study: (i) the largest h PΔ, (ii) the largest h within a specified maximum allowed flow rate, and (iii) the lowest weight that maximizes h for operation within the maximum allowed flow rate. While the shortcoming of the first objective is demonstrated, the other two objectives are found to be suitable for designing effective heat sinks (HSs) across different applications. Results suggest a promising trend from the third objective to develop HSs with ~ 37-68% lower weight, 80-85% reduced ΔP, and negligible penalty in h compared with optimized HSs obtained from the second objective. However, since the third objective leads to HSs with thinner fins, structural analysis should be performed to ensure reliable operation of the HSs.

42 ENGINEERING↗

Multi-objective optimization with an integrated electromagnetics and beam dynamics workflow

In particle accelerators, RF cavities are used to accelerate charged particle beams to designed high energy for physical applications. In a typical accelerator design, the optimization of RF cavities and the optimization of beam dynamics are carried out in separate studies. For a more general and unrestricted accelerator design, a coupled optimization of the RF cavities and the beam parameters is required. For this coupled optimization problem, we have developed an integrated electromagnetics and beam dynamics workflow management system. Within this system, the geometries for a set of cavity components are first adjusted; the field modes are then computed with an electromagnetics program, and imported into a beam dynamics program for beam dynamics simulation. This workflow is encapsulated into a parallel multi-objective optimizer to achieve the integrated accelerator design optimization. A multi fidelity strategy is developed to improve the speed of the optimizer. Furthermore, this integrated global optimization capability is illustrated using a photoinjector design example and yields an improved design.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Homotopy Solver

This software implements parallel versions of an interior-point solver, based on the publicly available ipopt solver. Here we have full control over the linear solver and our algorithm is fully parallel thus enabling scalability to large-scale optimization problems. This package also has a parallel implementation of a homotopy solver developed under the scalable methods for contact LDRD project 23-ERD-017. This solver is an mfem-based implementation of algorithm described in ``A filter trust-region Newton continuation method for nonlinear complementarity problems''. Cosmin G. Petra, Nai-Yuan Chiang, Jingyi Wang, Tucker Hartland, and Michael Puso (submitted), LLNL-JRNL-869761.

Hartland, Tucker [Lawrence Livermore National Labo↗

Parallel IO Libraries for Managing HEP Experimental Data

The computing and storage requirements of the energy and intensity frontiers will grow significantly during the Run 4 & 5 and the HL-LHC era. Similarly, in the intensity frontier, with larger trig ger readouts during supernovae explosions, the Deep Underground Neutrino Experiment (DUNE) will have unique computing challenges that could be addressed by the use of parallel and accelerated dataprocessing capabilities. Most of the requirements of the energy and intensity frontier experiments rely on increasing the role of high performance computing (HPC) in the HEP community. In this presentation, we will describe our ongoing efforts that are focused on using HPC resources for the next generation HEP experiments. The HEPCCE (High Energy Physics-Center for Computational Excellence) IOS (Input/Output and Storage) group has been developing approaches to map HEP data to the HDF5 , an IO library optimized for the HPC platforms to store the intermediate HEP data. The complex HEP data products are serialized using ROOT to allow for experiment independent general mapping approaches of the HEP data to the HDF5 format. The mapping approaches can be optimized for high performance parallel IO. Similarly, simpler data can be directly mapped into the HDF5, which can also be suitable for offloading into the GPUs directly. We will present our works on both complex and simple data model models.

Bashyal, Amit↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

Privateer

Privateer is a general-purpose data store that optimizes the tradeoff between storage space utilization and I/O performance. Privateer uses memory-mapped I/O with private mapping and an optimized writeback mechanism to maximize write parallelism and eliminate redundant writes; it also uses contentaddressable storage to optimize storage space via de-duplication.

Iwabuchi, Keita↗

Layer-Parallel Training of Deep Residual Neural Networks

Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Finally, using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.

97 MATHEMATICS AND COMPUTING↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TeraChem: Accelerating electronic structure and ab initio molecular dynamics with graphical processing units

Developed over the past decade, TeraChem is an electronic structure and ab initio molecular dynamics software package designed from the ground up to leverage graphics processing units (GPUs) to perform large-scale ground and excited state quantum chemistry calculations in the gas and the condensed phase. TeraChem’s speed stems from the reformulation of conventional electronic structure theories in terms of a set of individually optimized high-performance electronic structure operations (e.g., Coulomb and exchange matrix builds, one- and two-particle density matrix builds) and rank-reduction techniques (e.g., tensor hypercontraction). Recent efforts have encapsulated these core operations and provided language-agnostic interfaces. Finally, this greatly increases the accessibility and flexibility of TeraChem as a platform to develop new electronic structure methods on GPUs and provides clear optimization targets for emerging parallel computing architectures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TEAL

TEAL is a financial performance calculator plugin for the RAVEN code, framework, resolving around the computation of Net Present Value and associated financial metrics. TEAL can make use of inflation rates, taxation, escalation factors, capital expenditure economy of scale scaling factors. The unique feature of TEAL is the capability to be linked with RAVEN external models and build corresponding cash flows using the variables computed by those external models. In addition to be able to use the capability to generate cash flows derived from complex physical models generated by RAVEN, another distinctive feature of TEAL is the capability to provide financial risk/probabilistic metrics that can empower RAVEN to perform optimization/analysis driven by financial risk augmentations. Optimization, robust optimization, parametric studies, large parallel simulations, sensitivity analysis, data mining, etc. are just some of the capabilities that can be leveraged.

Alfonsi, Andrea↗

Parapint

Parapint is a Python package for parallel solution of dynamic optimization problems. SAND2020-12446 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Bynum, Michael↗

Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures

This study presents the first constrained sparse tensor factorization (cSTF) framework that optimizes and fully offloads computation to massively parallel GPU architectures, and the first performance characterization of cSTF on GPU architectures. In contrast to prior work on tensor factorization, where the matricized tensor times Khatri-Rao product (MTTKRP) is the primary performance bottleneck, our systematic analysis of the cSTF algorithm on GPUs reveals that adding constraints creates an additional bottleneck in the update operation for many real-world sparse tensors. While executing the update operation on the GPU brings significant speedup over its CPU counterpart, it remains a significant bottleneck. To further accelerate the update operation, we propose cuADMM, a new update algorithm that leverages algorithmic and code optimization strategies to minimize both computation and data movement on GPUs. As a result, our framework delivers significantly improved performance compared to prior state-of-the-art. On 10 real-world sparse tensors, our framework achieves geometric mean speedup of 5.1 × (max 41.59 ×) and 7.01 × (max 58.05 ×) on the NIVIDA A100 and H100 GPUs, respectively, over the state-of-the-art SPLATT library running on a 26-core Intel Ice Lake Xeon CPU.

Soh, Yongseok↗