Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Invited Paper: Benchmarking and Optimizing Data Movement on Emerging Heterogeneous Architectures

As supercomputers evolve, nodes are continually increasing in complexity. As a result, each generation of parallel systems brings new performance challenges. For instance, on recent systems inter-node communication has outperformed inter-socket, resulting in poor performance of many node-aware communication optimizations. Communication optimizations are critical for the performance and scalability of parallel applications, but are dependent on the parallel architecture, which varies significantly among recent generations of supercomputers. Furthermore, this paper investigates the performance of various paths of data movement on recent generations of systems, and analyzes the increased complexity of communication, particularly on recent heterogeneous systems. The paper also introduces MPI Advance, a communication library that enables optimizations to be created based on benchmark analysis of each emerging system.

benchmarking↗

The Galley Parallel File System

Most current multiprocessor file systems are designed to use multiple disks in parallel, using the high aggregate bandwidth to meet the growing I/0 requirements of parallel scientific applications. Many multiprocessor file systems provide applications with a conventional Unix-like interface, allowing the application to access multiple disks transparently. This interface conceals the parallelism within the file system, increasing the ease of programmability, but making it difficult or impossible for sophisticated programmers and libraries to use knowledge about their I/O needs to exploit that parallelism. In addition to providing an insufficient interface, most current multiprocessor file systems are optimized for a different workload than they are being asked to support. We introduce Galley, a new parallel file system that is intended to efficiently support realistic scientific multiprocessor workloads. We discuss Galley's file structure and application interface, as well as the performance advantages offered by that interface.

Nieuwejaar, Nils↗

On Improving Efficiency of Differential Evolution for Aerodynamic Shape Optimization Applications

Differential Evolution (DE) is a simple and robust evolutionary strategy that has been provEn effective in determining the global optimum for several difficult optimization problems. Although DE offers several advantages over traditional optimization approaches, its use in applications such as aerodynamic shape optimization where the objective function evaluations are computationally expensive is limited by the large number of function evaluations often required. In this paper various approaches for improving the efficiency of DE are reviewed and discussed. Several approaches that have proven effective for other evolutionary algorithms are modified and implemented in a DE-based aerodynamic shape optimization method that uses a Navier-Stokes solver for the objective function evaluations. Parallelization techniques on distributed computers are used to reduce turnaround times. Results are presented for standard test optimization problems and for the inverse design of a turbine airfoil. The efficiency improvements achieved by the different approaches are evaluated and compared.

Madavan, Nateri K.↗

Optimized structure and electronic band gap of monolayer GeSe from quantum Monte Carlo methods

Here, we have used highly accurate quantum Monte Carlo methods to determine the chemical structure and electronic band gaps of monolayer GeSe. Two-dimensional (2D) monolayer GeSe has received a great deal of attention due to its unique thermoelectric, electronic, and optoelectronic properties with a wide range of potential applications. Density functional theory (DFT) methods have usually been applied to obtain optical and structural properties of bulk and 2D GeSe. For the monolayer, DFT typically yields a larger band-gap energy than for bulk GeSe but cannot conclusively determine if the monolayer has a direct or indirect gap. Moreover, the DFT-optimized lattice parameters and atomic coordinates for monolayer GeSe depend strongly on the choice of approximation for the exchange-correlation functional, which makes the ideal structure-and its electronic properties-unclear. In order to obtain accurate lattice parameters and atomic coordinates for the monolayer, we use a surrogate Hessian-based parallel line search within diffusion Monte Carlo to fully optimize the GeSe monolayer structure. The DMC-optimized structure is different from those obtained using DFT, as are calculated band gaps. The potential energy surface has a shallow minimum at the optimal structure. This, combined with the sensitivity of the electronic structure to strain, suggests that the optical properties of monolayer GeSe are highly tunable by strain.

36 MATERIALS SCIENCE↗

On the minimax feedback control of uncertain dynamic systems.

In this paper the problem of optimal feedback control of uncertain discrete-time dynamic systems is considered where the uncertain quantities do not have a stochastic description but instead are known to belong to given sets. The problem is converted to a sequential minimax problem and dynamic programming is suggested as a general method for its solution. The notion of a sufficiently informative function, which parallels the notion of a sufficient statistic of stochastic optimal control, is introduced, and conditions under which the optimal controller decomposes into an estimator and an actuator are identified.

Bertsekas, D. P.↗

Sufficiently informative functions and the minimax feedback control of uncertain dynamic systems.

The problem of optimal feedback control of uncertain discrete-time dynamic systems is considered where the uncertain quantities do not have a stochastic description but instead are known to belong to given sets. The problem is converted to a sequential minimax problem and dynamic programming is suggested as a general method for its solution. The notion of a sufficiently informative function, which parallels the notion of a sufficient statistic of stochastic optimal control, is introduced, and conditions under which the optimal controller decomposes into an estimator and an actuator are identified.

Bertsekas, D. P.↗

An Efficient Multiblock Method for Aerodynamic Analysis and Design on Distributed Memory Systems

The work presented in this paper describes the application of a multiblock gridding strategy to the solution of aerodynamic design optimization problems involving complex configurations. The design process is parallelized using the MPI (Message Passing Interface) Standard such that it can be efficiently run on a variety of distributed memory systems ranging from traditional parallel computers to networks of workstations. Substantial improvements to the parallel performance of the baseline method are presented, with particular attention to their impact on the scalability of the program as a function of the mesh size. Drag minimization calculations at a fixed coefficient of lift are presented for a business jet configuration that includes the wing, body, pylon, aft-mounted nacelle, and vertical and horizontal tails. An aerodynamic design optimization is performed with both the Euler and Reynolds Averaged Navier-Stokes (RANS) equations governing the flow solution and the results are compared. These sample calculations establish the feasibility of efficient aerodynamic optimization of complete aircraft configurations using the RANS equations as the flow model. There still exists, however, the need for detailed studies of the importance of a true viscous adjoint method which holds the promise of tackling the minimization of not only the wave and induced components of drag, but also the viscous drag.

Reuther, James↗

The design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a 3D unstructured-grid Euler-solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data-structure has been adopted, and a number of optimizations have been devised and implemented in order to accelerate the parallel computational rates. The implementation is carried out by creating a set of software tools, which ease the implementation of computational problems on parallel architecture machines by relieving the user of the low-level machine specific issues. The quantitative effect of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of three performance improvement. The overall solution efficiency is compared with that obtained on the CRAY-YMP vector supercomputer.

Das, R.↗

Optimizing Desalination Operations for Energy Flexibility

Despite the value of energy optimization in desalination processes, modeling dynamic operations for monthly billing periods has remained a computational challenge. This work proposes a framework for energy flexibility optimization, which includes new modeling features for independent operation of parallel skids, start-up delays associated with chemical stabilization, the consideration of industrial energy tariff structures, and inclusion of hourly electrical carbon intensities. This is done using a modular and computationally efficient formulation that guarantees a globally optimal solution with standard optimization solvers. In this study, the approach is demonstrated in two distinct case studies: a seawater desalination plant in Santa Barbara, CA, and an indirect potable reuse facility in San Jose, CA. Trends predicted from the model are validated against operational facility measurements from a demand response shutdown event. Preliminary results show that optimizing energy flexibility can result in 18.51% monthly cost savings over energy efficiency-optimized operation. The value extracted from a facility-wide shutdown during peak electricity price hours is hampered by start-up delays in post-treatment chemical stabilization. In cases in which a facility does not have much excess capacity, using a flow equalization tank or operating over a wide recovery range may be cost-effective.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel algorithms for finding connected components using linear algebra

Finding connected components is one of the most widely used operations on a graph. Optimal serial algorithms for the problem have been known for half a century, and many competing parallel algorithms have been proposed over the last several decades under various different models of parallel computation. This paper presents a class of parallel connected-component algorithms designed using linear-algebraic primitives. These algorithms are based on a PRAM algorithm by Shiloach and Vishkin and can be designed using standard GraphBLAS operations. Here, we demonstrate two algorithms of this class, one named LACC for Linear Algebraic Connected Components, and the other named FastSV which can be regarded as LACC’s simplification. With the support of the highly-scalable Combinatorial BLAS library, LACC and FastSV outperform the previous state-of-the-art algorithm by a factor of up to 12x for small to medium scale graphs. For large graphs with more than 50B edges, LACC and FastSV scale to 4K nodes (262K cores) of a Cray XC40 supercomputer and outperform previous algorithms by a significant margin. This remarkable performance is accomplished by (1) exploiting sparsity that was not present in the original PRAM algorithm formulation, (2) using high-performance primitives of Combinatorial BLAS, and (3) identifying hot spots and optimizing them away by exploiting algorithmic insights.

97 MATHEMATICS AND COMPUTING↗

A scoping study of far-SOL main-wall protection limiters for steady-state operation of compact pilot plant tokamaks

We present a novel method for handling steady-state heat fluxes incident on the main wall of pilot plant-scale magnetic fusion devices, based on the utilization of protection limiters in the far scrape-off layer (SOL). This method helps avoid large plasma-wall gaps, without excessively compromising blanket performance. We present an optimization algorithm for determining the appropriate size and scale of these protection limiters given (1) probability distributions of SOL plasma parameters and (2) assumed risk tolerance. As part of this optimization, we have developed an analytic description of parallel heat fluxes across limiter shadows, and an objective cost function (the ‘Far-SOL Marginal Cost’) to quantify the impact that different main-wall thermal management design choices have on reactor capital cost. Applying the model to a midscale fusion pilot plant concept shows that making use of far-SOL protection limiters can reduce capital costs on the order of $500 M, relative to naively increasing the plasma-wall gap. Our analysis demonstrates that the far-SOL power decay length is the highest-leverage plasma assumption for thermal loading of the first wall, and the primary cost driver for main wall thermal management. The relative cost efficiency of protection limiters increases as assumptions on the far-SOL heat flux become more pessimistic. The concepts described in this paper motivate the further development of far-SOL protection limiters as part of larger efforts to design economical core-edge-wall compatible solutions for a fusion pilot plant.

Design under uncertainty↗

Efficient smoothed particle radiation hydrodynamics I: Thermal radiative transfer

This work presents efficient solution techniques for radiative transfer in the smoothed particle hydrodynamics discretization. Two choices that impact efficiency are how the material and radiation energy are coupled, which determines the number of iterations needed to converge the emission source, and how the radiation diffusion equation is solved, which must be done in each iteration. The coupled material and radiation energy equations are solved using an inexact Newton iteration scheme based on nonlinear elimination, which reduces the number of Newton iterations needed to converge within each time step. During each Newton iteration, the radiation diffusion equation is solved using Krylov iterative methods with a multigrid preconditioner, which abstracts and optimizes much of the communication when running in parallel. The code is verified for an infinite medium problem, a one-dimensional Marshak wave, and a two and three-dimensional manufactured problem, and exhibits first-order convergence in time and second-order convergence in space. For these problems, the number of iterations needed to converge the inexact Newton scheme and the diffusion equation is independent of the number of spatial points and the number of processors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Successes and Opportunities for Discovery of Metal Oxide Photoanodes for Solar Fuels Generators

We report that the importance of metal oxide photoanodes in solar fuels technology has garnered concerted efforts in photoanode discovery in recent decades, which complement parallel efforts in development of analytical techniques and optimization strategies using standard photoanodes such as TiO 2 , Fe 2 O 3 and BiVO 4 . Theoretical guidance of high-throughput experiments has been particularly effective in dramatically increasing the portfolio of metal oxide photoanodes, motivating a new era of photoanode development where the characterization and optimization techniques developed on traditional materials are applied to nascent photoanodes that exhibit visible light photoresponse. The compendium of metal oxide photoanodes presented in the present work can also serve as the basis for further technique development, with a primary goal to establish workflows for discovery of materials that perform better against the critical criteria of operational stability, visible light photoresponse, and photovoltage suitable for tandem absorber architectures.

14 SOLAR ENERGY↗

Accelerating computational modeling and design of high-entropy alloys

High-entropy alloys, with N elements and compositions {$c_{ν = 1,N}$} in competing crystal structures, have large design spaces for unique chemical and mechanical properties. In this work, to enable computational design, we use a metaheuristic hybrid Cuckoo search (CS) to construct alloy configurational models on the fly that have targeted atomic site and pair probabilities on arbitrary crystal lattices, given by supercell random approximates (SCRAPs) with S sites. Our Hybrid CS permits efficient global solutions for large, discrete combinatorial optimization that scale linearly in a number of parallel processors, and linearly in sites S for SCRAPs. For example, a four-element, 128-site SCRAP is found in seconds—a more than 13,000-fold reduction over current strategies. Our method thus enables computational alloy design that is currently impractical. We qualify the models and showcase application to real alloys with targeted atomic short-range order. Being problem-agnostic, our Hybrid CS offers potential applications in diverse fields.

36 MATERIALS SCIENCE↗

Hvac: Removing I/O Bottleneck for Large-Scale Deep Learning Applications

Scientific communities are increasingly adopting deep learning (DL) models in their applications to accelerate scientific discovery processes. However, with rapid growth in the computing capabilities of HPC supercomputers, large-scale DL applications have to spend a significant portion of training time performing I/O to a parallel storage system. Previous research works have investigated optimization techniques such as prefetching and caching. Unfortunately, there exist non-trivial challenges to adopting the existing solutions on HPC supercomputers for large-scale DL training applications, which include non-performance and/or failures at extreme scale, lack of portability and generality in design, complex deployment methodology, and being limited to a specific application or dataset. To address these challenges, we propose High-Velocity AI Cache (HVAC), a distributed read-cache layer that targets and fully exploits the node-local storage or near node-local storage technology. HVAC seamlessly accelerates read I/O by aggregating node-local or near node-local storage, avoiding metadata lookups and file locking while preserving portability in the application code. We deploy and evaluate HVAC on 1,024 nodes (with over 6000 NVIDIA V100 GPUS) of the Summit supercomputer. In particular, we evaluate the scalability, efficiency, accuracy, and load distribution of HVAC compared to GPFS and XFS-on-NVMe. With four different DL applications, we observe an average 25 % performance improvement atop GPFS and 9% drop against XFS-on-NVMe, which scale linearly and are considered the performance upper bound. We envision HVAC as an important caching library for upcoming HPC supercomputers such as Frontier.

Khan, Awais↗

Analytical methods for performance evaluation of nonlinear filters.

In the investigation, the filtering problem is considered in the continuous time domain. The postulated simple suboptimal nonlinear filter structure closely parallels the structure of the Kalman-Bucy optimal linear filter algorithm. Two filter performance evaluation methods are developed based on the Kolmogorov equations for the transition density of Markov processes. The expansions in the approximations for the nonlinear system and observation functions are in effect carried out up to second-order terms in both methods. The description of the filter's performance is sought in terms of second-order statistics in both methods.

Bejczy, A. K.↗

New Approaches to HSCT Multidisciplinary Design and Optimization

New approaches to MDO have been developed and demonstrated during this project on a particularly challenging aeronautics problem- HSCT Aeroelastic Wing Design. To tackle this problem required the integration of resources and collaboration from three Georgia Tech laboratories: ASDL, SDL, and PPRL, along with close coordination and participation from industry. Its success can also be contributed to the close interaction and involvement of fellows from the NASA Multidisciplinary Analysis and Optimization (MAO) program, which was going on in parallel, and provided additional resources to work the very complex, multidisciplinary problem, along with the methods being developed. The development of the Integrated Design Engineering Simulator (IDES) and its initial demonstration is a necessary first step in transitioning the methods and tools developed to larger industrial sized problems of interest. It also provides a framework for the implementation and demonstration of the methodology. Attachment: Appendix A - List of publications. Appendix B - Year 1 report. Appendix C - Year 2 report. Appendix D - Year 3 report. Appendix E - accompanying CDROM.

Schrage, Daniel P.↗

Low Temperature Performance of High-Speed Neural Network Circuits

Artificial neural networks, derived from their biological counterparts, offer a new and enabling computing paradigm specially suitable for such tasks as image and signal processing with feature classification/object recognition, global optimization, and adaptive control. When implemented in fully parallel electronic hardware, it offers orders of magnitude speed advantage. Basic building blocks of the new architecture are the processing elements called neurons implemented as nonlinear operational amplifiers with sigmoidal transfer function, interconnected through weighted connections called synapses implemented using circuitry for weight storage and multiply functions either in an analog, digital, or hybrid scheme.

artificial neural networks image processing signal↗