Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98

Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study

Many parallel and distributed computing research results are obtained in simulation, using simulators that mimic real-world executions on some target system. Each such simulator is configured by picking values for parameters that define the behavior of the underlying simulation models it implements. The main concern for a simulator is accuracy: simulated behaviors should be as close as possible to those observed in the real-world target system. This requires that values for each of the simulator's parameters be carefully picked, or “calibrated,” based on ground-truth real-world executions. Examining the current state of the art shows that simulator calibration, at least in the field of parallel and distributed computing, is often undocumented (and thus perhaps often not performed) and, when documented, is described as a labor-intensive, manual process. In this work we evaluate the benefit of automating simulation calibration using simple algorithms. Specifically, we use a real-world case study from the field of High Energy Physics and compare automated calibration to calibration performed by a domain scientist. Our main finding is that automated calibration is on par with or significantly outperforms the calibration performed by the domain scientist. Furthermore, automated calibration makes it straightforward to operate desirable tradeoffs between simulation accuracy and simulation speed.

Mc donald, Jesse↗

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S↗

Distributed fiber sensing systems for 3D combustion temperature field monitoring in coal-fired boilers using optically generated acoustic waves (Final Report)

In this project, we have developed and tested three kinds of fiber optic sensing systems for real time monitoring of temperature variations within an industrial scale boiler furnace. The fiber optic sensing systems target spatial and temporal distributions of high temperature profiles in a boiler furnace in fossil power plants. The reconstructed temperature profile will provide critical input for the control mechanisms to optimize the combustion process. This temperature profile will address the essential problem for fossil power plants in achieving higher efficiency and fewer pollutant emissions. Acoustic pyrometer systems have been used to reconstruct temperature field of power plant boilers based on measuring TOF (times-of-flight) of sound waves along some straight paths in a 2D cross-section of the boiler. In this project, optically generated acoustic signals from a fiber optic sensing system have replaced the acoustic signals generated from an electrical transducer. A 3D reconstruction algorithm replaced the previous 2D model. In this project, three kinds of fiber optic sensing systems have been developed and tested. They are fiber optic sensing system I, fiber optic sensing system II (Distributed Sensing System I) and fiber optic sensing system III (Distributed Sensing System II). For fiber optic sensing system I, the fiber optic ultrasound generator acts as a signal generator. A microphone, hydrophone or other electronic devices serve as a signal receiver. In this system, there are one generator and one receiver. Distance test, water temperature test, air temperature test, air temperature reconstruction, and GE ISBF pilot test were performed by Fiber optic sensing system I. The fiber optic sensing system I successfully detected temperature in all these tests. We got 2D temperature reconstruction results by using the fiber optic sensing system I and it matched the reference data. The fiber optic sensing system I successfully survived in GE ISBF boiler environment (480 °F). For fiber optic sensing system II (Distributed Sensing System I), it is an all optical ultrasound system. The fiber optic ultrasound generator acts as a signal generator. Fiber Bragg Grating (FBG) and Fabry-Perot (FP) sensor act as a signal receiver. In this system, there is one generator and one receiver. Aluminum plate temperature test, furnace high temperature test, and GE ISBF pilot test were performed by the fiber optic sensing system II. Fiber optic sensing system II successfully detected the temperature in all these tests. The fiber optic sensing system II successfully survived at up to 700 °C furnace environment and 320 °C GE ISBF boiler environment. For fiber optic sensing system III (Distributed Sensing System II), it is also an all optical ultrasound system. The fiber optic ultrasound generator acts as a signal generator. Multiple FP fiber sensors act as signal receivers. In this system, there is one generator and three receivers. Three GE ISBF pilot tests were performed by fiber optic sensing system III. The fiber optic sensing system III survived in the cold flow tests in GE’s ISBF pilot test facility. However, we didn’t get high temperature data by using this system since the nanosecond laser issues. During the period of the project, test trials, data simulation and algorithm optimization was performed successfully. For real time temperature field construction, the sampling rate must be fast enough to capture the field variations. The technology of Code-division multiple access (CDMA) is well studied which could allow parallel multiplexing, even if signals overlap in time or frequencies. Moreover, it has been known that extending the length of signal significantly improves SNR. For acoustic signals, these multiplexing techniques have also been widely used, mainly for sonar and acoustic communications. The CDMA modulation technique has been proposed and studied to guarantee high network throughput, low channel access delay and low energy consumption. We have studied the temperature field reconstruction using Gaussian Radial Basis Functions (GRBF)-based approximation approach. Reconstruction of 3D temperature field using Neural Networks with measured TOF and known propagation paths is feasible. 2D and 3D temperature field reconstruction simulation results are achieved. The milestone status is shown in Table 1. We finished milestone 1-8 and milestone 10. For milestone 9, we did three pilot tests by using the fiber optic sensing system III (Distributed Sensing System II) at GE Power. However, due to the failure of the ns laser, we did not get the temperature results. We conducted some additional tasks that were not originally proposed: 1) We fabricated a fiber optic sensing system I and did a pilot test based on this system. 2) In the proposal, we proposed two pilot tests at GE Power. In reality, we finished at least seven pilot tests at GE Power. GE Power has made a lot of efforts for supporting the pilot tests. 3) We got a simulation results based on CDMA. In summary, most of the tasks have been accomplished. The outcome of this project removed a few barriers that hinder the achievement of the final product of the distributed sensing systems. With the successful accomplishment of this project, a prototype of the fiber optic sensing system can be fabricated to attract more interests from companies and other funding agencies.

47 OTHER INSTRUMENTATION↗

Algorithms for on-line parameter and mode shape estimation

Algorithms are presented for on-line parameter and mode-shape estimation. The approach used is based upon a modal decomposition of the dynamic response of the flexible structure and is designed to make use of the parallel processing features of modern minicomputers. Satisfactory performance of the parallel structure identification technique used can be achieved only when the approximation functions noted correspond to the natural modes of the flexible structure. The work summarized here presents a technique for estimating both mode shapes and modal parameters.

Thau, F. E.↗

ICASE semiannual report, April 1 - September 30, 1989

The Institute conducts unclassified basic research in applied mathematics, numerical analysis, and computer science in order to extend and improve problem-solving capabilities in science and engineering, particularly in aeronautics and space. The major categories of the current Institute for Computer Applications in Science and Engineering (ICASE) research program are: (1) numerical methods, with particular emphasis on the development and analysis of basic numerical algorithms; (2) control and parameter identification problems, with emphasis on effective numerical methods; (3) computational problems in engineering and the physical sciences, particularly fluid dynamics, acoustics, and structural analysis; and (4) computer systems and software, especially vector and parallel computers. ICASE reports are considered to be primarily preprints of manuscripts that have been submitted to appropriate research journals or that are to appear in conference proceedings.

Source record↗

Acquisition and tracking performance measurements for a high speed area array detector system

A proof-of-concept (POC) demonstration system has been developed which demonstrates acquisition, tracking and point-ahead angle sensing for a space optical communications terminal utilizing a single high speed area array detector. The detector is the 128 x 128 pixel Kodak HS-40 photodiode array. It has 64 parallel readout channels and can operate at frames rates up to 40,000 frames/sec with rms readout noise of 20 photoelectrons. A windowing scheme and special purpose digital signal processing electronics are employed to implement acquisition and tracking algorithms. The system operates at greater than 1 kHz sample (frame) rates. Acquisition can be performed in as little as 30 milliseconds with less than 1 picowatt of 0.85 micron beacon power on the detector. At the same power level, the rms tracking accuracy is approximately 1/16 pixel. Results of system analysis and measurements using the POC system are presented.

Short, R. C.↗

Three-Dimensional Deformable Grid Electromagnetic Particle-in-cell for Parallel Computers

We describe a new parallel, non-orthogonal grid, three-dimensional electromagnetic particle-in-cell (EMPIC) code based on a finite-volume formulation. This code uses a logically Cartesian grid of deformable hexahedral cells, a discrete surface integral (DSI) algorithm to calculate the electromagnetic field, and a hybrid logical-physical space algorithm to push particles.

Cartesian grid Grid Electromagnetic electromagneti↗

Kepler Science Operations Center Architecture

We give an overview of the operational concepts and architecture of the Kepler Science Data Pipeline. Designed, developed, operated, and maintained by the Science Operations Center (SOC) at NASA Ames Research Center, the Kepler Science Data Pipeline is central element of the Kepler Ground Data System. The SOC charter is to analyze stellar photometric data from the Kepler spacecraft and report results to the Kepler Science Office for further analysis. We describe how this is accomplished via the Kepler Science Data Pipeline, including the hardware infrastructure, scientific algorithms, and operational procedures. The SOC consists of an office at Ames Research Center, software development and operations departments, and a data center that hosts the computers required to perform data analysis. We discuss the high-performance, parallel computing software modules of the Kepler Science Data Pipeline that perform transit photometry, pixel-level calibration, systematic error-correction, attitude determination, stellar target management, and instrument characterization. We explain how data processing environments are divided to support operational processing and test needs. We explain the operational timelines for data processing and the data constructs that flow into the Kepler Science Data Pipeline.

Middour, Christopher↗

An Ensemble Investigation of the Causes for Regional Air-Quality Model Critical Load Exceedances Prediction Variability in European and North American Domains Using Diagnostics From Phase 4 of the Air Quality Model Evaluation International Initiative

We summarize tentative findings from multi air quality model ensembles for the years 2009 and 2010 in Europe (EU), and 2010 and 2016 in North America (NA), under AQMEII-4. The model predictions of sulphur and nitrogen deposition were used to estimate exceedances of critical loads for acidification and eutrophication, to show the extent to which the ensemble members agree in the magnitude and the trend of ecologically meaningful impacts. Model exceedance variability was analyzed using AQMEII-4 diagnostics. Evaluation against concentration and wet deposition observations, coupled with these diagnostics, identified specific process representations as the causes for variability between model predictions and for reduced model performance. All models predicted reductions in ecosystem acidification impacts in North America between the years 2010 and 2016, in accord with SO2 emissions reduction legislation which started in 2010 (SO2 SIP) However, all models in EU and NA domains had net negative biases for wet deposition of sulphur and nitrogen relative to observations. The wet S deposition average mean bias for the NA ensemble was -0.17 eq ha-1 d-1, and for the EU ensemble -1.15 eq ha-1 d-1. The NA daily wet deposition average mean bias for NH4+ was -0.37 eq ha-1d-1; EU -1.19 eq ha-1 d-1. The daily NA wet NO3- deposition average mean bias was -0.24 eq ha-1d-1; EU -0.69 eq ha-1 d-1. The members of the ensemble diverged (factor of 10) in their North American predictions for Ndep and consequently their eutrophication exceedances. The models with the highest eutrophication predictions also predicted the highest levels of gas-phase ammonia dry deposition (standard deviation of ammonia dry deposition flux across ensemble members was larger than the ensemble average). These models also had negative biases of predicted ammonia concentrations; average mean biases of -0.63 (satellite NH3) and -0.85 ppbv (surface NH3) compared to ensemble averages of -0.30 and -0.34 ppbv. Diagnostics showed that these differences resulted from the manner in which bidirectional ammonia fluxes were parameterized within these models. The second largest source of NA eutrophication prediction variability were models with positive biases in particulate ammonium and nitrate concentrations, and higher particle nitrogen deposition levels ( particle ammonium concentration bias +0.35 ug m-3; ensemble bias +0.15 ug m-3). We believe two factors may have led to these latter overestimates: higher levels of fine mode particle nitrate formation compared to other models (due to the use of an inorganic heterogeneous chemistry algorithm which did not take base cation chemistry into account), and updates to particle dry deposition velocities carried out in the absence of concurrent updates to wet scavenging algorithms. The relative importance of dry gas, dry particulate, and wet deposition towards total sulphur and nitrogen deposition totals differed between EU and North American domains, though all models had negative biases in wet deposition as noted above. Parallel and subsequent work suggests that multiphase hydrometeor scavenging may improve model wet deposition performance. An increased research focus is recommended for four model processes: multiphase hydrometeor scavenging, ammonia bidirectional fluxes, base cation chemistry and emissions, and particle dry deposition.

regional air-quality model↗

Scalability of high-performance PDE solvers

Performance tests and analyses are critical to effective high-performance computing software development and are central components in the design and implementation of computational algorithms for achieving faster simulations on existing and future computing architectures for large-scale application problems. In this article, we explore performance and space-time trade-offs for important compute-intensive kernels of large-scale numerical solvers for partial differential equations (PDEs) that govern a wide range of physical applications. We consider a sequence of PDE-motivated bake-off problems designed to establish best practices for efficient high-order simulations across a variety of codes and platforms. We measure peak performance (degrees of freedom per second) on a fixed number of nodes and identify effective code optimization strategies for each architecture. In addition to peak performance, we identify the minimum time to solution at 80% parallel efficiency. The performance analysis is based on spectral and p-type finite elements but is equally applicable to a broad spectrum of numerical PDE discretizations, including finite difference, finite volume, and h-type finite elements.

97 MATHEMATICS AND COMPUTING↗

A massively parallel time-domain coupled electrodynamics–micromagnetics solver

We present a high-performance coupled electrodynamics–micromagnetics solver for full physical modeling of signals in microelectronic circuitry. The overall strategy couples a finite-difference time-domain approach for Maxwell’s equations to a magnetization model described by the Landau–Lifshitz–Gilbert equation. The algorithm is implemented in the Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. Furthermore, the code leverages ongoing developments of the Exascale Application Code, WarpX, which is primarily being developed for plasma wakefield accelerator modeling. Our temporal coupling scheme provides second-order accuracy in space and time by combining the integration steps for the magnetic field and magnetization into an iterative sub-step that includes a trapezoidal temporal discretization for the magnetization. The performance of the algorithm is demonstrated by the excellent scaling results on NERSC multicore and GPU systems, with a significant (59×) speedup on the GPU using a node-by-node comparison. We demonstrate the utility of our code by performing simulations of an electromagnetic waveguide and a magnetically tunable filter.

97 MATHEMATICS AND COMPUTING↗

Oineus v1.0

A library for multi-threaded computation of persistence diagrams. The algorithm for computing persistence diagrams in a lock-free manner was published in the 'Towards Lock-free Persistent Homology' (D. Morozov, A. Nigmetov, Brief Announcement: SPAA 2020); it scales better than the only other shared-memory parallel implementation of persistent homology computation PHAT. Library includes python bindings to compute persistence diagrams and their vectorizations for lower-star filtrations on grid data. It is intended to be used by scientists working on Topological Data Analysis and its applications in different areas.

Nigmetov, Arnur↗

Neural network based decomposition in optimal structural synthesis

The present paper describes potential applications of neural networks in the multilevel decomposition based optimal design of structural systems. The generic structural optimization problem of interest, if handled as a single problem, results in a large dimensionality problem. Decomposition strategies allow for this problem to be represented by a set of smaller, decoupled problems, for which solutions may either be obtained with greater ease or may be obtained in parallel. Neural network models derived through supervised training, are used in two distinct modes in this work. The first uses neural networks to make available efficient analysis models for use in repetitive function evaluations as required by the optimization algorithm. In the second mode, neural networks are used to represent the coupling that exists between the decomposed subproblems. The approach is illustrated by application to the multilevel decomposition-based synthesis of representative truss and frame structures.

Hajela, P.↗

CFD analysis of hypersonic, chemically reacting flow fields

Design studies are underway for a variety of hypersonic flight vehicles. The National Aero-Space Plane will provide a reusable, single-stage-to-orbit capability for routine access to low earth orbit. Flight-capable satellites will dip into the atmosphere to maneuver to new orbits, while planetary probes will decelerate at their destination by atmospheric aerobraking. To supplement limited experimental capabilities in the hypersonic regime, computational fluid dynamics (CFD) is being used to analyze the flow about these configurations. The governing equations include fluid dynamic as well as chemical species equations, which are being solved with new, robust numerical algorithms. Examples of CFD applications to hypersonic vehicles suggest an important role this technology will play in the development of future aerospace systems. The computational resources needed to obtain solutions are large, but solution adaptive grids, convergence acceleration, and parallel processing may make run times manageable.

Edwards, T. A.↗

On the Computational Capabilities of Physical Systems: The Impossibility of Infallible Computation - Part 1

In this first of two papers, strong limits on the accuracy of physical computation are established. First it is proven that there cannot be a physical computer C to which one can pose any and all computational tasks concerning the physical universe. Next it is proven that no physical computer C can correctly carry out any computational task in the subset of such tasks that can be posed to C. This result holds whether the computational tasks concern a system that is physically isolated from C, or instead concern a system that is coupled to C. As a particular example, this result means that there cannot be a physical computer that can, for any physical system external to that computer, take the specification of that external system's state as input and then correctly predict its future state before that future state actually occurs; one cannot build a physical computer that can be assured of correctly 'processing information faster than the universe does'. The results also mean that there cannot exist an infallible, general-purpose observation apparatus, and that there cannot be an infallible, general-purpose control apparatus. These results do not rely on systems that are infinite, and/or non-classical, and/or obey chaotic dynamics. They also hold even if one uses an infinitely fast, infinitely dense computer, with computational powers greater than that of a Turing Machine. This generality is a direct consequence of the fact that a novel definition of computation - a definition of 'physical computation' - is needed to address the issues considered in these papers. While this definition does not fit into the traditional Chomsky hierarchy, the mathematical structure and impossibility results associated with it have parallels in the mathematics of the Chomsky hierarchy. The second in this pair of papers presents a preliminary exploration of some of this mathematical structure, including in particular that of prediction complexity, which is a 'physical computation analogue' of algorithmic information complexity. It is proven in that second paper that either the Hamiltonian of our universe proscribes a certain type of computation, or prediction complexity is unique (unlike algorithmic information complexity), in that there is one and only version of it that can be applicable throughout our universe.

Wolpert, David H.↗

Algorithm 1028: VTMOP: Solver for Blackbox Multiobjective Optimization Problems

VTMOP is a Fortran 2008 software package containing two Fortran modules for solving computationally expensive bound-constrained blackbox multiobjective optimization problems. VTMOP implements the algorithm of [32], which handles two or more objectives, does not require any derivatives, and produces well-distributed points over the Pareto front. The first module contains a general framework for solving multiobjective optimization problems by combining response surface methodology, trust region methodology, and an adaptive weighting scheme. The second module features a driver subroutine that implements this framework when the objective functions can be wrapped as a Fortran subroutine. Lastly, support is provided for both serial and parallel execution paradigms, and VTMOP is demonstrated on several test problems as well as one real-world problem in the area of particle accelerator optimization.

97 MATHEMATICS AND COMPUTING↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

SNoGloDe: A Structured Nonlinear Global Decomposition Solver

Large-scale optimization problems often require decomposition strategies and customized algorithms to achieve optimal solutions within a reasonable time. Building on the work of Cao and Zavala (2019) for solving nonlinear two-stage stochastic programs to global optimality, we implement and extend their approach. We generalize to optimization problems reformulated with a block-angular constraint structure (e.g., temporal decomposition). Our framework, written in Python using Pyomo, is highly customizable and enables parallel execution of the decomposition. SNoGloDe allows tailored branching strategies, lower bounding problems, and candidate generators to leverage problem-specific knowledge. To demonstrate effectiveness, we compare SNoGloDe’s performance with Gurobi on a temporally decomposed produced water case study.

algorithms↗