Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Frequency Recovery in Power Grids Using High-Performance Computing

Maintaining electric power system stability is paramount, especially in extreme contingencies involving unexpected outages of multiple generators or transmission lines that are typical during severe weather events. Such outages often lead to large supply-demand mismatches followed by subsequent system frequency deviations from their nominal value. The extent of frequency deviations is an important metric of system resilience, and its timely mitigation is a central goal of power system operation and control. This paper develops a novel nonlinear model predictive control (NMPC) method to minimize frequency deviations when the grid is affected by an unforeseen loss of multiple components. Our method is based on a novel multi-period alternating current optimal power flow (ACOPF) formulation that accurately models both nonlinear electric power flow physics and the primary and secondary frequency response of generator control mechanisms. We develop a distributed parallel Julia package for solving the large-scale nonlinear optimization problems that result from our NMPC method and thereby address realistic test instances on existing high-performance computing architectures. Our method demonstrates superior performance in terms of frequency recovery over existing industry practices, where generator levels are set based on the solution of single-period classical ACOPF models.

nonlinear model predictive control↗

The Design and Evaluation of "CAPTools"--A Computer Aided Parallelization Toolkit

Writing applications for high performance computers is a challenging task. Although writing code by hand still offers the best performance, it is extremely costly and often not very portable. The Computer Aided Parallelization Tools (CAPTools) are a toolkit designed to help automate the mapping of sequential FORTRAN scientific applications onto multiprocessors. CAPTools consists of the following major components: an inter-procedural dependence analysis module that incorporates user knowledge; a 'self-propagating' data partitioning module driven via user guidance; an execution control mask generation and optimization module for the user to fine tune parallel processing of individual partitions; a program transformation/restructuring facility for source code clean up and optimization; a set of browsers through which the user interacts with CAPTools at each stage of the parallelization process; and a code generator supporting multiple programming paradigms on various multiprocessors. Besides describing the rationale behind the architecture of CAPTools, the parallelization process is illustrated via case studies involving structured and unstructured meshes. The programming process and the performance of the generated parallel programs are compared against other programming alternatives based on the NAS Parallel Benchmarks, ARC3D and other scientific applications. Based on these results, a discussion on the feasibility of constructing architectural independent parallel applications is presented.

Yan, Jerry↗

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher↗

Distributed Parallel Processing and Dynamic Load Balancing Techniques for Multidisciplinary High Speed Aircraft Design

Multidisciplinary design optimization (MDO) for large-scale engineering problems poses many challenges (e.g., the design of an efficient concurrent paradigm for global optimization based on disciplinary analyses, expensive computations over vast data sets, etc.) This work focuses on the application of distributed schemes for massively parallel architectures to MDO problems, as a tool for reducing computation time and solving larger problems. The specific problem considered here is configuration optimization of a high speed civil transport (HSCT), and the efficient parallelization of the embedded paradigm for reasonable design space identification. Two distributed dynamic load balancing techniques (random polling and global round robin with message combining) and two necessary termination detection schemes (global task count and token passing) were implemented and evaluated in terms of effectiveness and scalability to large problem sizes and a thousand processors. The effect of certain parameters on execution time was also inspected. Empirical results demonstrated stable performance and effectiveness for all schemes, and the parametric study showed that the selected algorithmic parameters have a negligible effect on performance.

Krasteva, Denitza T.↗

Parallel Memory-Independent Communication Bounds for SYRK

In this paper, we focus on the parallel communication cost of multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK). SYRK requires half the computation of general matrix multiplication because of the symmetry of the output matrix. Recent work (Beaumont et al., SPAA '22) has demonstrated that the sequential I/O complexity of SYRK is also a constant factor smaller than that of general matrix multiplication. Inspired by this progress, we establish memory-independent parallel communication lower bounds for SYRK with smaller constants than general matrix multiplication, and we show that these constants are tight by presenting communication-optimal algorithms. The crux of the lower bound proof relies on extending a key geometric inequality to symmetric computations and analytically solving a constrained nonlinear optimization problem. Here, the optimal algorithms use a triangular blocking scheme for parallel distribution of the symmetric output matrix and corresponding computation.

Communication costs↗

NEXTorch: A Design and Bayesian Optimization Toolkit for Chemical Sciences and Engineering

Automation and optimization of chemical systems require well-informed decisions on what experiments to run to reduce time, materials, and/or computations. Data-driven active learning algorithms have emerged as valuable tools to solve such tasks. Bayesian optimization, a sequential global optimization approach, is a popular active-learning framework. Past studies have demonstrated its efficiency in solving chemistry and engineering problems. Here we introduce NEXTorch, a library in Python/PyTorch, to facilitate laboratory or computational design using Bayesian optimization. NEXTorch offers fast predictive modeling, flexible optimization loops, visualization capabilities, easy interfacing with legacy software, and multiple types of parameters and data type conversions. It provides GPU acceleration, parallelization, and state-of-the-art Bayesian optimization algorithms and supports both automated an d human-in-the-loop optimization. The comprehensive online documentation introduces Bayesian optimization theory and several examples from catalyst synthesis, reaction condition optimization, parameter estimation, and reactor geometry optimization. NEXTorch is open-source and available on GitHub

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallelization of an Object-Oriented Unstructured Aeroacoustics Solver

A computational aeroacoustics code based on the discontinuous Galerkin method is ported to several parallel platforms using MPI. The discontinuous Galerkin method is a compact high-order method that retains its accuracy and robustness on non-smooth unstructured meshes. In its semi-discrete form, the discontinuous Galerkin method can be combined with explicit time marching methods making it well suited to time accurate computations. The compact nature of the discontinuous Galerkin method also makes it well suited for distributed memory parallel platforms. The original serial code was written using an object-oriented approach and was previously optimized for cache-based machines. The port to parallel platforms was achieved simply by treating partition boundaries as a type of boundary condition. Code modifications were minimal because boundary conditions were abstractions in the original program. Scalability results are presented for the SCI Origin, IBM SP2, and clusters of SGI and Sun workstations. Slightly superlinear speedup is achieved on a fixed-size problem on the Origin, due to cache effects.

Baggag, Abdelkader↗

Solving Unit Commitment Problems with Demand Responsive Loads

This work focuses on using variations of the Frank-Wolfe (FW) algorithm for solving unit commitment problems with high volumes of demand responsive loads on the power grid. We present a formulation of the unit commitment problem with demand responsive loads. We then show through reformulation and relaxations of the problem that variations of the Frank-Wolfe algorithm can be used to determine the time series decisions for the demand responsive loads. We show through computational experiments on the IEEE Reliability Test System that the timeseries of demand responsive load decisions obtained through our approach are near optimal and describe how large-scale parallel implementations of our approach can be highly computationally efficient.

demand response↗

Parallel simulated annealing with embedded machine learning and multifidelity models for reactor core design

This paper presents extensions to a penalty-free, parallel simulated annealing (SA) algorithm for multi-constrained combinatorial optimization with the aim of embedding multi-fidelity physics models into the annealing procedure. The method uses a low-fidelity, quickly executing model for rapid design space exploration and a high-fidelity model for detailed constraint resolution and on-the-fly bias correction. Machine learning models updated within the annealing procedure were used to bridge the gap between the multi-fidelity models, which led to accurate rapid exploration and efficient detailed constraint resolution. A software implementation of the new multi-fidelity optimization methods, called ML-PSA, was demonstrated on a continuous multi-fidelity optimization problem and a constrained combinatorial PWR lattice design problem. These problems demonstrate some of the features, parallel performance characteristics, and extensible nature of the multi-fidelity SA methods. This paper shows that the developed software and procedure are a general optimization tool that can be applied to a wide variety of scientific and engineering design optimization applications. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Selection of Thermal Worst-Case Orbits via Modified Efficient Global Optimization

Efficient Global Optimization (EGO) was used to select orbits with worst-case hot and cold thermal environments for the Stratospheric Aerosol and Gas Experiment (SAGE) III. The SAGE III system thermal model changed substantially since the previous selection of worst-case orbits (which did not use the EGO method), so the selections were revised to ensure the worst cases are being captured. The EGO method consists of first conducting an initial set of parametric runs, generated with a space-filling Design of Experiments (DoE) method, then fitting a surrogate model to the data and searching for points of maximum Expected Improvement (EI) to conduct additional runs. The general EGO method was modified by using a multi-start optimizer to identify multiple new test points at each iteration. This modification facilitates parallel computing and decreases the burden of user interaction when the optimizer code is not integrated with the model. Thermal worst-case orbits for SAGE III were successfully identified and shown by direct comparison to be more severe than those identified in the previous selection. The EGO method is a useful tool for this application and can result in computational savings if the initial Design of Experiments (DoE) is selected appropriately.

Moeller, Timothy M.↗

High-throughput feedback-enabled optogenetic stimulation and spectroscopy in microwell plates

Abstract The ability to perform sophisticated, high-throughput optogenetic experiments has been greatly enhanced by recent open-source illumination devices that allow independent programming of light patterns in single wells of microwell plates. However, there is currently a lack of instrumentation to monitor such experiments in real time, necessitating repeated transfers of the samples to stand-alone analytical instruments, thus limiting the types of experiments that could be performed. Here we address this gap with the development of the optoPlateReader (oPR), an open-source, solid-state, compact device that allows automated optogenetic stimulation and spectroscopy in each well of a 96-well plate. The oPR integrates an optoPlate illumination module with a module called the optoReader, an array of 96 photodiodes and LEDs that allows 96 parallel light measurements. The oPR was optimized for stimulation with blue light and for measurements of optical density and fluorescence. After calibration of all device components, we used the oPR to measure growth and to induce and measure fluorescent protein expression in E. coli . We further demonstrated how the optical read/write capabilities of the oPR permit computer-in-the-loop feedback control, where the current state of the sample can be used to adjust the optical stimulation parameters of the sample according to pre-defined feedback algorithms. The oPR will thus help realize an untapped potential for optogenetic experiments by enabling automated reading, writing, and feedback in microwell plates through open-source hardware that is accessible, customizable, and inexpensive.

59 BASIC BIOLOGICAL SCIENCES↗

LLNL/thicket

Thicket is a python-based toolkit for Exploratory Data Analysis (EDA) of parallel performance data that enables performance optimization and understanding of applications’ performance on supercomputers. It bridges the performance tool gap between being able to consider only a single instance of a simulation run (e.g., single platform, single measurement tool, or single scale) and finding actionable insights in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets.

Brink, Stephanie Labasan↗

Tandem particle-slurry batch reactors for solar water splitting (Final Scientific/Technical Report)

Economically, particle slurry reactors are projected to be one of the most promising technologies for solar photoelectrochemical hydrogen production, according to a 2009 techno-economic analysis commissioned by the US DOE and performed by Directed Technologies, Inc. The Fuel Cell Technologies Office’s Multi-Year Research, Development and Demonstration (MYRD&D) goals and targets are to reduce the cost of H 2 produced from renewable sources at the plant gate (i.e. not including delivery, dispensing, or storage) to < $2.00/gge, equivalent to ~$2.00/kg H 2 . Research results from our techno-economic modeling research suggest that this target could be met using particle slurry reactors assuming STH efficiencies in the range of 5 – 10%, materials lifetimes of < 1 year, and nanoparticles that cost up to 20 times more than projected costs of TiO 2 -coated Fe 2 O 3 nanoparticles. Although most large worldwide research efforts directed at solar photoelectrochemical hydrogen production focus on wafer-based designs, the projected lower cost for a particle slurry reactor at these disparate projected STH efficiencies clearly suggests that particle slurry reactors could be a scalable and deployable technology, assuming several challenges are overcome. These major technological challenges include the demonstration of a vertically-stacked-vessel architecture that is capable of operating sustainably while mostly relying on diffusion and natural convection to mix the redox shuttles between the vessels, and the demonstration that photocatalyst particles can operate at an overall 1% STH efficiency or larger when incorporated into this two-vessel design. Our research adds to the understanding of photocatalytic reactors for solar water splitting through numerical modeling results and experimental results. Numerical models were developed to simulate relevant device physics including particle and reactor dimensions which affect optical, transport, and rheological properties, electrocatalytic and photovoltaic properties of particles at various temperatures, and properties of redox shuttles and separators. Moreover, theoretical maximum solar-to-hydrogen efficiencies for ensembles of particles like in photocatalyst reactors were modeled and simulated and shown to equal or exceed those of photoelectrochemical designs under most scenarios. These results help determine constraints on the reactor that will enable more optimal designs for future prototypes. In parallel, experiments were performed to identify the most effective redox shuttles and to empirically validate the numerical models and simulations. Toward the latter, state-of-the-art light-absorbing particles and electrocatalysts were synthesized and characterized physically and photoelectrochemically for water electrolysis and redox chemistry with redox shuttles in the form factor of mesoporous electrodes and free-floating particles. The most promising materials candidates were used in a suspension reactor to evaluate performance toward photocatalytic H 2 production and results from the two measurements were compared. Predominantly, state-of-the-art cocatalyst-modified Rh-doped SrTiO 3 and BiVO 4 particles were further characterized to assess for their ability to perform visible-light-driven H 2 and O 2 evolution, respectively, and results were similar to those reported for the state-of-the-art in the peer-reviewed literature. Outcomes from this work inform the public of the effectiveness and promise of solar photocatalytic water splitting for clean and renewable hydrogen production. This work may also help increase research interest and funding for photocatalysis projects, which will accelerate development of a technology that will benefit the public by generating fuel while emitting few greenhouse gases and pollutants.

08 HYDROGEN↗

Stencils and problem partitionings: Their influence on the performance of multiple processor systems

Given a discretization stencil, partitioning the problem domain is an important first step for the efficient solution of partial differential equations on multiple processor systems. Partitions are derived that minimize interprocessor communication when the number of processors is known a priori and each domain partition is assigned to a different processor. This partitioning technique uses the stencil structure to select appropriate partition shapes. For square problem domains, it is shown that non-standard partitions (e.g., hexagons) are frequently preferable to the standard square partitions for a variety of commonly used stencils. This investigation is concluded with a formalization of the relationship between partition shape, stencil structure, and architecture, allowing selection of optimal partitions for a variety of parallel systems.

Reed, D. A.↗

Stencils and problem partitionings - Their influence on the performance of multiple processor systems

Given a discretization stencil, partitioning the problem domain is an important first step for the efficient solution of partial differential equations on multiple processor systems. Partitions are derived that minimize interprocessor communication when the number of processors is known a priori and each domain partition is assigned to a different processor. This partitioning technique uses the stencil structure to select appropriate partition shapes. For square problem domains, it is shown that non-standard partitions (e.g., hexagons) are frequently preferable to the standard square partitions for a variety of commonly used stencils. This investigation is concluded with a formalization of the relationship between partition shape, stencil structure, and architecture, allowing selection of optimal partitions for a variety of parallel systems.

Reed, Daniel A.↗

Accelerated Simulation of Air Pollution Using NVIDIA RAPIDS

Atmospheric chemistry models are a central tool to study and forecast the impact of air pollution on the environment, vegetation, and human health. However, the numerical simulation of chemical kinetics is computationally expensive due to the stiffness of the system of ordinary differential equations that describes atmospheric chemistry. Here we present an alternative approach to the computation of atmospheric chemistry based on machine learning. Our training data set is produced using the NASA Goddard Earth Observing System (GEOS) model with GEOS-Chem chemistry, run on the NASA Center for Climate Simulation (NCCS) Discover supercomputing cluster on 384 Intel Xeon Haswell cores. This model spends more than 50% of total run time on solving atmospheric chemistry. The data set contains as input features the air pollution concentrations before solving the differential equations, together with some key physical parameters such as temperature and sun intensity. As target variables we define the air pollution concentrations after solving the differential equations. Using Dask-cuDF and Dask-XGBoost on the NVIDIA RAPIDS platform on 8 Tesla V100 GPUs, we generate from this training set gradient boosted decision tree models that can reproduce the simulation of chemical kinetics. We do this on the NCCS Advanced Data Analytics Platform (ADAPT) science cloud environment. Our application takes full advantage of recent advances in Dask-XGBoost, such as multi-node and multi-GPU scaling for distributed training with large data sets. The increase in training data size enabled by this is critical to capture the full range of chemical environments encountered across the globe and all annual seasons.The boosted tree models offer good predictability and show many of the features of the full chemistry reference simulation. Further improvements can be achieved through mass balance considerations and by accounting for error correlations. We incorporate the boosted tree models into the GEOS reference model using XGBoost's C API. This enables a seamless integration of the GPU trained models into GEOS-Chem, which is written in Fortran and optimized for use in a massively parallel CPU environment. We show the benefits of this approach and discuss the potential speedup of this machine learning accelerated atmospheric chemistry model.

Keller, Christoph A.↗