Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

LLNL/thicket

Thicket is a python-based toolkit for Exploratory Data Analysis (EDA) of parallel performance data that enables performance optimization and understanding of applications’ performance on supercomputers. It bridges the performance tool gap between being able to consider only a single instance of a simulation run (e.g., single platform, single measurement tool, or single scale) and finding actionable insights in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets.

Brink, Stephanie Labasan↗

Tandem particle-slurry batch reactors for solar water splitting (Final Scientific/Technical Report)

Economically, particle slurry reactors are projected to be one of the most promising technologies for solar photoelectrochemical hydrogen production, according to a 2009 techno-economic analysis commissioned by the US DOE and performed by Directed Technologies, Inc. The Fuel Cell Technologies Office’s Multi-Year Research, Development and Demonstration (MYRD&D) goals and targets are to reduce the cost of H 2 produced from renewable sources at the plant gate (i.e. not including delivery, dispensing, or storage) to < $2.00/gge, equivalent to ~$2.00/kg H 2 . Research results from our techno-economic modeling research suggest that this target could be met using particle slurry reactors assuming STH efficiencies in the range of 5 – 10%, materials lifetimes of < 1 year, and nanoparticles that cost up to 20 times more than projected costs of TiO 2 -coated Fe 2 O 3 nanoparticles. Although most large worldwide research efforts directed at solar photoelectrochemical hydrogen production focus on wafer-based designs, the projected lower cost for a particle slurry reactor at these disparate projected STH efficiencies clearly suggests that particle slurry reactors could be a scalable and deployable technology, assuming several challenges are overcome. These major technological challenges include the demonstration of a vertically-stacked-vessel architecture that is capable of operating sustainably while mostly relying on diffusion and natural convection to mix the redox shuttles between the vessels, and the demonstration that photocatalyst particles can operate at an overall 1% STH efficiency or larger when incorporated into this two-vessel design. Our research adds to the understanding of photocatalytic reactors for solar water splitting through numerical modeling results and experimental results. Numerical models were developed to simulate relevant device physics including particle and reactor dimensions which affect optical, transport, and rheological properties, electrocatalytic and photovoltaic properties of particles at various temperatures, and properties of redox shuttles and separators. Moreover, theoretical maximum solar-to-hydrogen efficiencies for ensembles of particles like in photocatalyst reactors were modeled and simulated and shown to equal or exceed those of photoelectrochemical designs under most scenarios. These results help determine constraints on the reactor that will enable more optimal designs for future prototypes. In parallel, experiments were performed to identify the most effective redox shuttles and to empirically validate the numerical models and simulations. Toward the latter, state-of-the-art light-absorbing particles and electrocatalysts were synthesized and characterized physically and photoelectrochemically for water electrolysis and redox chemistry with redox shuttles in the form factor of mesoporous electrodes and free-floating particles. The most promising materials candidates were used in a suspension reactor to evaluate performance toward photocatalytic H 2 production and results from the two measurements were compared. Predominantly, state-of-the-art cocatalyst-modified Rh-doped SrTiO 3 and BiVO 4 particles were further characterized to assess for their ability to perform visible-light-driven H 2 and O 2 evolution, respectively, and results were similar to those reported for the state-of-the-art in the peer-reviewed literature. Outcomes from this work inform the public of the effectiveness and promise of solar photocatalytic water splitting for clean and renewable hydrogen production. This work may also help increase research interest and funding for photocatalysis projects, which will accelerate development of a technology that will benefit the public by generating fuel while emitting few greenhouse gases and pollutants.

08 HYDROGEN↗

User-Oriented Improvements in the MOOSE framework in support of Multiphysics Simulation

The MOOSE Framework is a foundational capability used by NEAMS to create over 15 different simulation tools for advanced nuclear reactors. Due to this ubiquity, improvements to the framework in support of modeling and simulation goals are critical to the program. These improvements can take many forms including optimization, improved user experience, streamlined APIs, parallelism, and new capability. The work transcribed in the report was in direct support of the simulation tools and is already deployed or will be deployed in the coming months. The capabilities implemented were, in the same order as this report, chainable execution objects or executors, support for transfers between applications at the same level in a coupling scheme, support for boundary/subdomain restricted transfers, support for transfers between applications with different coordinate or unit systems, support for MOOSE applications in the NEAMS workbench, deployment of MOOSE application of the INL HPC OnDemand platform, addition of a triangular meshing library in libMesh and increased support of face variables.

97 MATHEMATICS AND COMPUTING↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

Optimization of Thermal Conductance at Interfaces Using Machine Learning Algorithms

We report optimization of thermal transport across the interface of two different materials is critical to micro-/nanoscale electronic, photonic, and phononic devices. Although several examples of compositional intermixing at the interfaces having a positive effect on interfacial thermal conductance (ITC) have been reported, an optimum arrangement has not yet been determined because of the large number of potential atomic configurations and the significant computational cost of evaluation. On the other hand, computation-driven materials design efforts are rising in popularity and importance. Yet, the scalability and transferability of machine learning models remain as challenges in creating a complete pipeline for the simulation and analysis of large molecular systems. In this work we present a scalable Bayesian optimization framework, which leverages dynamic spawning of jobs through the Message Passing Interface (MPI) to run multiple parallel molecular dynamics simulations within a parent MPI job to optimize heat transfer at the silicon and aluminum (Si/Al) interface. We found a maximum of 50% increase in the ITC when introducing a two-layer intermixed region that consists of a higher percentage of Si. Because of the random nature of the intermixing, the magnitude of increase in the ITC varies. We observed that both homogeneity/heterogeneity of the intermixing and the intrinsic stochastic nature of molecular dynamics simulations account for the variance in ITC.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Reassessing the MCNP Random Number Generator

Random number generators are integral components to Monte Carlo codes. They provide the pseudorandom number sequence used to actually sample the distributions of interest. As a result, they are one of the most important components to the software. The current recommended MCNP random number generator is a 63-bit linear congruential generator (LCG). This generator is quite fast, but it has some drawbacks. First, it only has a period of 2 63 . Due to the necessarily non-optimal usage of random numbers to ensure parallel reproducibility, this amount is too few to guarantee random number sequences are not reused in all configurations the code runs under. As simulation size increases, users will need to be aware of the limitations of the generator and tune configuration variables to best suit their simulations, or they will need to assume that reuse is not negatively affecting their answers. Neither of these are optimal. Second, small LCGs are fairly weak in bit generation quality, and this can have an unknown impact on the quality of the simulation. This paper is an investigation into whether or not more modern random number generators can supersede the current ones. The goal is to find a generator that is similar or superior in speed to the LCGs, has a state space large enough to make strong guarantees about random number reuse, and passes all modern random number test suites. If such a generator is found, it would eliminate the need for the user to even be aware of the limitations of the random number generator and would simplify the use of the code. This paper will be broken into several parts. Sec. 2 will discuss the evolution of the random number generator within the MCNP code. Sec. 3 will go over what a Monte Carlo code needs from a generator to be reproducible and portable and how the current generator behaves in that light. Sec. 4 goes through how each generator was tested. Finally, Sec. 5 will discuss improvements that could be made to the code.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Efficient Probabilistic Computing with Stochastic Perovskite Nickelates

Probabilistic computing has emerged as a viable approach to solve hard optimization problems. Devices with inherent stochasticity can greatly simplify their implementation in electronic hardware. In this report we demonstrate intrinsic stochastic resistance switching controlled via electric fields in perovskite nickelates doped with hydrogen. The ability of hydrogen ions to reside in various metastable configurations in the lattice leads to a distribution of transport gaps. With experimentally characterized p-bits, a shared-synapse p-bit architecture demonstrates highly parallelized and energy-efficient solutions to optimization problems such as integer factorization and Boolean satisfiability. The results introduce perovskite nickelates as scalable potential candidates for probabilistic computing and showcase the potential of light-element dopants in next-generation correlated semiconductors.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Parameter Optimization Toolbox for NS-3 network optimization, NS-3 Parameter Optimization Framework [SWR-18-60]

This simulation-based parameter optimization framework is proposed to tune parameters of different types of communication networks using ns-3 to achieve the optimal network performance. It consists of three main components: an ns-3 packet reporting module; a sampler running simulations with all possible parameter sets for the input parameter variables by using a parallel executor at each generation; and a hybrid optimization algorithm for tuning configurable parameters of hybrid designs and application parameter variables. The proposed hybrid metaheuristic optimization algorithm combines an evolutionary algorithm with a gradient descent function to quickly achieve an approximate globally optimum solution. This software is designed to be used in a multi-core processing Linux environment and run over a long duration of time. The execution time varies depending mainly upon the nature of the ns-3 configuration being simulated. This software includes a custom ns-3 QoS measurement application which must be included with the ns-3 source code during installation of the software.

Hasandka, Adarsh↗

Throughput Analytics of Cloud Networks

A network of virtual machines at cloud server sites connected over virtual IO connections is a flexible, easily deployable, and cost-effective alternative to a physical network infrastructure with dedicated servers connected over leased fiber lines. We study the throughput performance of such a cloud network by collecting measurements over Google Cloud infrastructure spanning multiple continents. To study its ideal performance and impact of packet losses, we utilize its emulation using dedicated servers and connection hardware emulation devices. We compare the measurements over the cloud network to those over its emulation on a testbed. We examine the throughput profiles of both networks as a function of the round trip time and their utilization-concavity coefficients, estimated using measurements for common TCP versions. The throughput profile's concave-convex shape and its coefficient are critical indicators of the network performance, qualitatively and quantitatively, respectively. The results indicate their overall agreement between the production cloud network and its emulation using dedicated connections, and a near optimal throughput performance of the former except for a few under-performing connections. Also, the number of parallel flows is found to be a dominant factor in optimizing the throughput across various conditions and TCP versions.

Phanekham, Derek↗

Parallelized domain decomposition for multi-dimensional Lagrangian random walk mass-transfer particle tracking schemes

Lagrangian particle tracking schemes allow a wide range of flow and transport processes to be simulated accurately, but a major challenge is numerically implementing the inter-particle interactions in an efficient manner. This article develops a multi-dimensional, parallelized domain decomposition (DDC) strategy for mass-transfer particle tracking (MTPT) methods in which particles exchange mass dynamically. We show that this can be efficiently parallelized by employing large numbers of CPU cores to accelerate run times. In order to validate the approach and our theoretical predictions we focus our efforts on a well-known benchmark problem with pure diffusion, where analytical solutions in any number of dimensions are well established. In this work, we investigate different procedures for “tiling” the domain in two and three dimensions (2-D and 3-D), as this type of formal DDC construction is currently limited to 1-D. An optimal tiling is prescribed based on physical problem parameters and the number of available CPU cores, as each tiling provides distinct results in both accuracy and run time. We further extend the most efficient technique to 3-D for comparison, leading to an analytical discussion of the effect of dimensionality on strategies for implementing DDC schemes. Increasing computational resources (cores) within the DDC method produces a trade-off between inter-node communication and on-node work. For an optimally subdivided diffusion problem, the 2-D parallelized algorithm achieves nearly perfect linear speedup in comparison with the serial run-up to around 2700 cores, reducing a 5 h simulation to 8 s, while the 3-D algorithm maintains appreciable speedup up to 1700 cores.

97 MATHEMATICS AND COMPUTING↗

Scalable Bayesian optimization with randomized prior networks

Several fundamental problems in science and engineering consist of global optimization tasks involving unknown high-dimensional (black-box) functions that map a set of controllable variables to the outcomes of an expensive experiment. Bayesian Optimization (BO) techniques are known to be effective in tackling global optimization problems using a relatively small number objective function evaluations, but their performance suffers when dealing with high-dimensional outputs. To overcome the major challenge of dimensionality, here we propose a deep learning framework for BO and sequential decision making based on bootstrapped ensembles of neural architectures with randomized priors. Using appropriate architecture choices, we show that the proposed framework can approximate functional relationships between design variables and quantities of interest, even in cases where the latter take values in high-dimensional vector spaces or even infinite-dimensional function spaces. In the context of BO, we augmented the proposed probabilistic surrogates with re-parameterized Monte Carlo approximations of multiple-point (parallel) acquisition functions, as well as methodological extensions for accommodating black-box constraints and multi-fidelity information sources. We test the proposed framework against state-of-the-art methods for BO and demonstrate superior performance across several challenging tasks with high-dimensional outputs, including a constrained multi-fidelity optimization task involving shape optimization of rotor blades in turbo-machinery.

97 MATHEMATICS AND COMPUTING↗

Decomposing Loosely Coupled Mixed-Integer Programs for Optimal Microgrid Design

Microgrids are frequently employed in remote regions, in part because access to a larger electric grid is impossible, difficult, or compromises reliability and independence. Although small microgrids often employ spot generation, in which a diesel generator is attached directly to a load, microgrids that combine these individual loads and augment generators with photovoltaic cells and batteries as a distributed energy system are emerging as a safer, less costly alternative. In this work, we present a model that seeks the minimum-cost microgrid design and ideal dispatched power to support a small remote site for one year with hourly fidelity under a detailed battery model; this mixed-integer nonlinear program (MINLP) is intractable with commercial solvers but loosely coupled with respect to time. A mixed-integer linear program (MIP) approximates the model, and a partitioning scheme linearizes the bilinear terms. We introduce a novel policy for loosely coupled MIPs in which the system reverts to equivalent conditions at regular time intervals; this separates the problem into subproblems that we solve in parallel. We obtain solutions within 5% of optimality in at most six minutes across 14 MIP instances from the literature and solutions within 5% of optimality to the MINLP instances within 20 minutes.

97 MATHEMATICS AND COMPUTING↗

A Framework for the Analysis of Compiler Optimizations

Compilers transform program source code to machine executable code. During this transformation, they perform a number of compiler optimizations to improve the performance of the generated executable code. Importantly, applying those optimizations depends on the source code structure, such as the parallel programming model used to parallelize an algorithm. Often, implementations of the same algorithm with different programming models have vastly different performance because the compiler optimized them differently. We create FAROS, a framework to structure and automate the analysis of compiler optimizations on programs. FAROS automates the building process, execution profiling, and analysis of compiler optimization of programs, through a configuration interface. It outputs compiler optimization reports to show which optimizations applied to which line of source code, leveraging compilation remarks output by the compiler. Also, FAROS supports benchmarking performance of different program versions by collecting execution time results. In this first release of FAROS, we provide a configuration file to analyze compiler optimization differences for sequential vs OpenMP compilation, including 38 programs consisting of HPC proxy/mini/large applications, and NAS and Rodina kernels for analysis.

Georgakoudis, Giorgis↗

PETSc/TAO Users Manual V.3.21

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and optimizers that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started ). The library is organized hierarchically, enabling users to employ the abstraction level most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users.

97 MATHEMATICS AND COMPUTING↗

Brief Announcement: Communication Optimal Sparse LU Factorization for Planar Matrices

We introduce a new parallel algorithm for solving sparse LU factorization of planar matrices, which commonly arise in the finite element method for 2D PDEs. Existing scalable methods, such as the multifrontal approach with subtree-to-subcube mapping by Gupta et al. [1] and right-looking with 3D mapping by Sao et al. [2] fail to achieve optimal communication costs for these matrices. Our new algorithm combines 3D mapping and subtree-to-subcube mapping to minimize communication costs while allowing trade-offs between extra memory and reduced communication. We demonstrate that our proposed algorithm attains the communication lower bound up to a factor of O(log log n) in the memory-optimal case and up to a factor of O(log P) in the memory-independent case for an n-dimensional planar sparse matrix on P processors.

Sao, Piyush↗

Design Considerations for GPU-based Mixed Integer Programming on Parallel Computing Platforms

Mixed Integer Programming (MIP) is a powerful abstraction in combinatorial optimization that finds real-life application across many significant sectors. The recent proliferation of graphical processing unit (GPU)-based accelerated computing architectures in large-scale parallel computing or supercomputing presents new opportunities as well as challenges in the advancement of MIP solver technology to effectively use the new accelerated computing platforms and scale to large parallel systems. Here, we recount the conventional processor-based strategies and focus on configurations where the most promising intersection lies between parallel MIP solver approaches and the specific strengths of accelerated parallel platforms. We note that the best potential lies in solving problems whose individual matrix sizes (of the linear program relaxation) fit entirely within one accelerator's memory and whose branch-and-bound (or branch-and-cut) trees cannot be fully contained within a small number of computational nodes. Additionally, we identify ideal features of computational linear algebra support on GPU accelerators that would help advance this direction of scalable parallel solution of MIP problems on GPU-based accelerated computing architectures.

Perumalla, Kalyan↗

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗