Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

3D Modeling of Ultrasonic Wave Interaction with Disbonds and Weak Bonds

Ultrasonic techniques, such as the use of guided waves, can be ideal for finding damage in the plate and pipe-like structures used in aerospace applications. However, the interaction of waves with real flaw types and geometries can lead to experimental signals that are difficult to interpret. 3-dimensional (3D) elastic wave simulations can be a powerful tool in understanding the complicated wave scattering involved in flaw detection and for optimizing experimental techniques. We have developed and implemented parallel 3D elastodynamic finite integration technique (3D EFIT) code to investigate Lamb wave scattering from realistic flaws. This paper discusses simulation results for an aluminum-aluminum diffusion disbond and an aluminum-epoxy disbond and compares results from the disbond case to the common artificial flaw type of a flat-bottom hole. The paper also discusses the potential for extending the 3D EFIT equations to incorporate physics-based weak bond models for simulating wave scattering from weak adhesive bonds.

Leckey, C.↗

Aircraft Configuration and Flight Crew Compliance with Procedures While Conducting Flight Deck Based Interval Management (FIM) Operations

Flight deck based Interval Management (FIM) applications using ADS-B are being developed to improve both the safety and capacity of the National Airspace System (NAS). FIM is expected to improve the safety and efficiency of the NAS by giving pilots the technology and procedures to precisely achieve an interval behind the preceding aircraft by a specific point. Concurrently but independently, Optimized Profile Descents (OPD) are being developed to help reduce fuel consumption and noise, however, the range of speeds available when flying an OPD results in a decrease in the delivery precision of aircraft to the runway. This requires the addition of a spacing buffer between aircraft, reducing system throughput. FIM addresses this problem by providing pilots with speed guidance to achieve a precise interval behind another aircraft, even while flying optimized descents. The Interval Management with Spacing to Parallel Dependent Runways (IMSPiDR) human-in-the-loop experiment employed 24 commercial pilots to explore the use of FIM equipment to conduct spacing operations behind two aircraft arriving to parallel runways, while flying an OPD during high-density operations. This paper describes the impact of variations in pilot operations; in particular configuring the aircraft, their compliance with FIM operating procedures, and their response to changes of the FIM speed. An example of the displayed FIM speeds used incorrectly by a pilot is also discussed. Finally, this paper examines the relationship between achieving airline operational goals for individual aircraft and the need for ATC to deliver aircraft to the runway with greater precision. The results show that aircraft can fly an OPD and conduct FIM operations to dependent parallel runways, enabling operational goals to be achieved efficiently while maintaining system throughput.

Shay, Rick↗

Transient Growth Analysis of Compressible Boundary Layers with Parabolized Stability Equations

The linear form of parabolized linear stability equations (PSE) is used in a variational approach to extend the previous body of results for the optimal, non-modal disturbance growth in boundary layer flows. This methodology includes the non-parallel effects associated with the spatial development of boundary layer flows. As noted in literature, the optimal initial disturbances correspond to steady counter-rotating stream-wise vortices, which subsequently lead to the formation of stream-wise-elongated structures, i.e., streaks, via a lift-up effect. The parameter space for optimal growth is extended to the hypersonic Mach number regime without any high enthalpy effects, and the effect of wall cooling is studied with particular emphasis on the role of the initial disturbance location and the value of the span-wise wavenumber that leads to the maximum energy growth up to a specified location. Unlike previous predictions that used a basic state obtained from a self-similar solution to the boundary layer equations, mean flow solutions based on the full Navier-Stokes (NS) equations are used in select cases to help account for the viscous-inviscid interaction near the leading edge of the plate and also for the weak shock wave emanating from that region. These differences in the base flow lead to an increasing reduction with Mach number in the magnitude of optimal growth relative to the predictions based on self-similar mean-flow approximation. Finally, the maximum optimal energy gain for the favorable pressure gradient boundary layer near a planar stagnation point is found to be substantially weaker than that in a zero pressure gradient Blasius boundary layer.

Compressible boundary layer↗

Addressing the Big-Earth-Data Variety Challenge with the Hierarchical Triangular Mesh

We have implemented an updated Hierarchical Triangular Mesh (HTM) as the basis for a unified data model and an indexing scheme for geoscience data to address the variety challenge of Big Earth Data. We observe that, in the absence of variety, the volume challenge of Big Data is relatively easily addressable with parallel processing. The more important challenge in achieving optimal value with a Big Data solution for Earth Science (ES) data analysis, however, is being able to achieve good scalability with variety. With HTM unifying at least the three popular data models, i.e. Grid, Swath, and Point, used by current ES data products, data preparation time for integrative analysis of diverse datasets can be drastically reduced and better variety scaling can be achieved. In addition, since HTM is also an indexing scheme, when it is used to index all ES datasets, data placement alignment (or co-location) on the shared nothing architecture, which most Big Data systems are based on, is guaranteed and better performance is ensured. Moreover, our updated HTM encoding turns most geospatial set operations into integer interval operations, gaining further performance advantages.

SciDB↗

Personal Rotorcraft Design and Performance with Electric Hybridization

Recent and projected improvements for more or all-electric aviation propulsion systems can enable greater personal mobility, while also reducing environmental impact (noise and emissions). However, all-electric energy storage capability is significantly less than present, hydrocarbon-fueled systems. A system study was performed exploring design and performance assuming hybrid propulsion ranging from traditional hydrocarbon-fueled cycles (gasoline Otto and diesel) to all-electric systems using electric motors generators, with batteries for energy storage and load leveling. Study vehicles were a conventional, single-main rotor (SMR) helicopter and an advanced vertical takeoff and landing (VTOL) aircraft. Vehicle capability was limited to two or three people (including pilot or crew); the design range for the VTOL aircraft was set to 150 miles (about one hour total flight). Search and rescue (SAR), loiter, and cruise-dominated missions were chosen to illustrate each vehicle and degree of hybrid propulsion strengths and weaknesses. The traditional, SMR helicopter is a hover-optimized design; electric hybridization was performed assuming a parallel hybrid approach by varying degree of hybridization. Many of the helicopter hybrid propulsion combinations have some mission capabilities that might be effective for short range or on-demand mobility missions. However, even for 30 year technology electrical components, all hybrid propulsion systems studied result in less available fuel, lower maximum range, and reduced hover and loiter duration than the baseline vehicle. Results for the VTOL aircraft were more encouraging. Series hybrid combinations reflective of near-term systems could improve range and loiter duration by 30. Advanced, higher performing series hybrid combinations could double or almost triple the VTOL aircrafts range and loiter duration. Additional details on the study assumptions and work performed are given, as well as suggestions for future study effort.

electric motors↗

Solving Unit Commitment Problems with Demand Responsive Loads: Preprint

This paper focuses on using variations of the Frank-Wolfe algorithm for solving unit commitment problems with high volumes of demand responsive loads on the power grid. We present a formulation of the unit commitment problem with demand responsive loads. We then show through reformulation and relaxations of the problem that variations of the Frank-Wolfe algorithm can be used to determine the time series decisions for the demand responsive loads. We show through computational experiments on the RTS-GMLC test system that the time series of demand responsive load decisions obtained through our approach are near optimal and provide details regarding how large scale parallel implementations of our approach can be highly computationally efficient.

demand response↗

The design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a three-dimensional unstructured grid Euler-solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data structure has been adopted, and a number of optimizations have been devised and implemented in order to accelerate the parallel communication rates. The implementation is carried out by creating a set of software tools, which provide an interface between the parallelization issues and the sequential code, while providing a basis for future automatic run-time compilation support. Large practical unstructured grid problems are solved on the Intel iPSC/860 hypercube and Intel Touchstone Delta machine. The quantitative effect of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of three performance improvement. The overall solution efficiency is compared with that obtained on the CRAY-YMP vector supercomputer.

Das, R.↗

Parallel computation of 3-D Navier-Stokes flowfields for supersonic vehicles

Multidisciplinary design optimization of aircraft will require unprecedented capabilities of both analysis software and computer hardware. The speed and accuracy of the analysis will depend heavily on the computational fluid dynamics (CFD) module which is used. A new CFD module has been developed to combine the robust accuracy of conventional codes with the ability to run on parallel architectures. This is achieved by parallelizing the ARC3D algorithm, a central-differenced Navier-Stokes method, on the Intel iPSC/860. The computed solutions are identical to those from conventional machines. Computational speed on 64 processors is comparable to the rate on one Cray Y-MP processor and will increase as new generations of parallel computers become available.

Ryan, James S.↗

Efficient Ensemble-Based Stochastic Gradient Methods for Optimization Under Geological Uncertainty

Ensemble-based stochastic gradient methods, such as the ensemble optimization (EnOpt) method, the simplex gradient (SG) method, and the stochastic simplex approximate gradient (StoSAG) method, approximate the gradient of an objective function using an ensemble of perturbed control vectors. These methods are increasingly used in solving reservoir optimization problems because they are not only easy to parallelize and couple with any simulator but also computationally more efficient than the conventional finite-difference method for gradient calculations. In this work, we show that EnOpt may fail to achieve sufficient improvement of the objective function when the differences between the objective function values of perturbed control variables and their ensemble mean are large. On the basis of the comparison of EnOpt and SG, we propose a hybrid gradient of EnOpt and SG to save on the computational cost of SG. We also suggest practical ways to reduce the computational cost of EnOpt and StoSAG by approximating the objective function values of unperturbed control variables using the values of perturbed ones. We first demonstrate the performance of our improved ensemble schemes using a benchmark problem. Results show that the proposed gradients saved about 30–50% of the computational cost of the same optimization by using EnOpt, SG, and StoSAG. As a real application, we consider pressure management in carbon storage reservoirs, for which brine extraction wells need to be optimally placed to reduce reservoir pressure buildup while maximizing the net present value. Results show that our improved schemes reduce the computational cost significantly.

58 GEOSCIENCES↗

Genetically Engineered Microelectronic Infrared Filters

A genetic algorithm is used for design of infrared filters and in the understanding of the material structure of a resonant tunneling diode. These two components are examples of microdevices and nanodevices that can be numerically simulated using fundamental mathematical and physical models. Because the number of parameters that can be used in the design of one of these devices is large, and because experimental exploration of the design space is unfeasible, reliable software models integrated with global optimization methods are examined The genetic algorithm and engineering design codes have been implemented on massively parallel computers to exploit their high performance. Design results are presented for the infrared filter showing new and optimized device design. Results for nanodevices are presented in a companion paper at this workshop.

Cwik, Tom↗

A second-order distributed memory parallel fast sweeping method for the Eikonal equation

The Eikonal equation is used to calculate wave propagation and distance fields, and due to its complexity requires numerical treatment for its solution. In this work, we present a second-order distributed memory parallel fast sweeping method. The second-order solution switches on a two-point stencil when two upwind points are available, and reverts to first-order otherwise. In all examples, the second-order method improves the solution over the first-order, allowing for significant savings in memory while achieving the same accuracy. Parallelization over distributed memory saw good weak scaling with optimal convergence. The computational time for second-order was approximately 2.5 times slower than first-order, where the largest amount of mesh points ran on 144 cores (512 GB) was ≈20 billion. The savings in memory from the second-order method combined with the distributed memory algorithm result in the ability to solve problems much larger than are possible with the serial first-order method.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Method for Optimizing Non-Axisymmetric Liners for Multimodal Sound Sources

Central processor unit times and memory requirements for a commonly used solver are compared to that of a state-of-the-art, parallel, sparse solver. The sparse solver is then used in conjunction with three constrained optimization methodologies to assess the relative merits of non-axisymmetric versus axisymmetric liner concepts for improving liner acoustic suppression. This assessment is performed with a multimodal noise source (with equal mode amplitudes and phases) in a finite-length rectangular duct without flow. The sparse solver is found to reduce memory requirements by a factor of five and central processing time by a factor of eleven when compared with the commonly used solver. Results show that the optimum impedance of the uniform liner is dominated by the least attenuated mode, whose attenuation is maximized by the Cremer optimum impedance. An optimized, four-segmented liner with impedance segments in a checkerboard arrangement is found to be inferior to an optimized spanwise segmented liner. This optimized spanwise segmented liner is shown to attenuate substantially more sound than the optimized uniform liner and tends to be more effective at the higher frequencies. The most important result of this study is the discovery that when optimized, a spanwise segmented liner with two segments gives attenuations equal to or substantially greater than an optimized axially segmented liner with the same number of segments.

Watson, W. R.↗

Vertex Reordering for Real-world Graphs and Applications: An Empirical Evaluation

Vertex reordering is a way to improve locality in graph computations. Given an input (or ``natural'') order, reordering aims to compute an alternate permutation of the vertices that is aimed at maximizing a locality-based objective. Given decades of research on this topic, there are tens of graph reordering schemes, and there are also several linear arrangement ``gap'' measures for treatment as objectives. However, a comprehensive empirical analysis of the efficacy of the ordering schemes against the different gap measures, and against real-world applications is currently lacking. In this study, we present an extensive empirical evaluation of up to 11 ordering schemes, taken from different classes of approaches, on a set of 34 real-world graphs emerging from different application domains. Our study is presented in two parts: a) a thorough comparative evaluation of the different ordering schemes on their effectiveness to optimize different linear arrangement gap measures, relevant to preserving locality; and b) extensive evaluation of the impact of the ordering schemes on two real-world, parallel graph applications, namely, community detection and influence maximization. Our studies show a significant divergence among the ordering schemes (up to $40\times$ between the best and the poor) in their effectiveness to reduce the gap measures; and a wide ranging impact of the ordering schemes on various aspects including application runtime (up to $4\times$), memory and cache use, load balancing, and parallel work and efficiency. The comparative study also help reveal the nuances of a parallel environment (compared to serial) on the ordering schemes and their role in optimizing applications.

Barik, Reet↗

An optimized FM-index library for nucleotide and amino acid search

Abstract Background Pattern matching is a key step in a variety of biological sequence analysis pipelines. The FM-index is a compressed data structure for pattern matching, with search run time that is independent of the length of the database text. Implementation of the FM-index is reasonably complicated, so that increased adoption will be aided by the availability of a fast and flexible FM-index library. Results We present AvxWindowedFMindex (AWFM-index), a lightweight, open-source, thread-parallel FM-index library written in C that is optimized for indexing nucleotide and amino acid sequences. AWFM-index introduces a new approach to storing FM-index data in a strided bit-vector format that enables extremely efficient computation of the FM-index occurrence function via AVX2 bitwise instructions, and combines this with optional on-disk storage of the index’s suffix array and a cache-efficient lookup table for partial k-mer searches. The AWFM-index performs exact match count and locate queries faster than SeqAn3’s FM-index implementation across a range of comparable memory footprints. When optimized for speed, AWFM-index is $$\sim $$ ∼ 2–4x faster than SeqAn3 for nucleotide search, and $$\sim $$ ∼ 2–6x faster for amino acid search; it is also $$\sim $$ ∼ 4x faster with similar memory footprint when storing the suffix array in on-disk SSD storage. Conclusions AWFM-index is easy to incorporate into bioinformatics software, offers run-time performance parameterization, and provides clients with FM-index functionality at both a high-level (count or locate all instances of a query string) and low-level (step-wise control of the FM-index backward-search process). The open-source library is available for download at https://github.com/TravisWheelerLab/AvxWindowFmIndex.

59 BASIC BIOLOGICAL SCIENCES↗

The Inspectability Metric: A Formalized System Of Measurement Enabling The Design For Inspection Framework

Nondestructive evaluation (NDE) engineers are often confronted with structural design choices that present challenges to meeting inspection requirements. These challenges, at best, increase the resources needed to design an inspection solution and, at worst, require resource intensive redesign of the structure. If the inspectability of the structure can be determined early in the design cycle, these challenging inspection scenarios can be avoided. The emergence of additive manufacturing has further compounded this problem by enabling the creation of highly optimized structures with no regard to inspection constraints. Design for inspection (DFI) offers a framework to integrate nondestructive evaluation (NDE) into the design process to alleviate the mechanisms that produce uninspectable designs. DFI is the concept of including inspectability in a multi-objective optimization framework so that it can be considered in parallel to other metrics such as mass and manufacturability. This allows rapid evaluation of the trade-off between design metrics to find solutions that meet the inspection needs of a particular material system, structural concept, or vehicle program. To enable DFI, there must be a system by which the inspectability of a structure can be measured. This system must be agile to produce results quickly, it must be versatile to work with the type of incomplete information one would encounter early in the design process (such as lack of inspection requirements), and it must be delivered in a form that is easily understood by designers. To meet this need, this presentation introduces the novel inspectability metric as a system to measure inspectability. The inspectability metric is a standardized, automation friendly procedure that uses simulations to determine inspectability. Along with guidelines to properly process designs and integrate with existing workflows, the inspectability metric provides a suite of simulation tests to interrogate the ability to find defects and the sensitivity to variability. The testing rubric is designed to maximize the coverage of the parameter space while minimizing the number of simulations needed. The inspectability metric has been in development in collaboration with industry partners to ensure compatibility with modern simulation tools and aerospace design workflows. In this study, we will demonstrate how the inspectability metric is able to determine the inspectability of multiple types of structures, including aerospace composites and additively manufactured parts. We will then show how the inspectability score can be plugged into existing design optimization tasks, such as structural sizing algorithms or design for manufacturing (DFM) frameworks.

Design for inspection↗

Reductive Analysis with Compiler-Guided Large Language Models for Input-Centric Code Optimizations

Input-centric program optimization aims to optimize code by considering the relations between program inputs and program behaviors. Despite its promise, a long-standing barrier for its adoption is the difficulty of automatically identifying critical features of complex inputs. This paper introduces a novel technique, reductive analysis through compiler-guided Large Language Models (LLMs), to solve the problem through a synergy between compilers and LLMs. It uses a reductive approach to overcome the scalability and other limitations of LLMs in program code analysis. The solution, for the first time, automates the identification of critical input features without heavy instrumentation or profiling, cutting the time needed for input identification by 44× (or 450× for local LLMs), reduced from 9.6 hours to 13 minutes (with remote LLMs) or 77 seconds (with local LLMs) on average, making input characterization possible to be integrated into the workflow of program compilations. Optimizations on those identified input features show similar or even better results than those identified by previous profiling-based methods, leading to optimizations that yield 92.6% accuracy in selecting the appropriate adaptive OpenMP parallelization decisions, and 20-30% performance improvement of serverless computing while reducing resource usage by 50-60%.

Input-Centric Optimization↗

Hierarchical Parallelism in Finite Difference Analysis of Heat Conduction

Based on the concept of hierarchical parallelism, this research effort resulted in highly efficient parallel solution strategies for very large scale heat conduction problems. Overall, the method of hierarchical parallelism involves the partitioning of thermal models into several substructured levels wherein an optimal balance into various associated bandwidths is achieved. The details are described in this report. Overall, the report is organized into two parts. Part 1 describes the parallel modelling methodology and associated multilevel direct, iterative and mixed solution schemes. Part 2 establishes both the formal and computational properties of the scheme.

Padovan, Joseph↗

Reconfigurable Model Execution in the OpenMDAO Framework

NASA's OpenMDAO framework facilitates constructing complex models and computing their derivatives for multidisciplinary design optimization. Decomposing a model into components that follow a prescribed interface enables OpenMDAO to assemble multidisciplinary derivatives from the component derivatives using what amounts to the adjoint method, direct method, chain rule, global sensitivity equations, or any combination thereof, using the MAUD architecture. OpenMDAO also handles the distribution of processors among the disciplines by hierarchically grouping the components, and it automates the data transfer between components that are on different processors. These features have made OpenMDAO useful for applications in aircraft design, satellite design, wind turbine design, and aircraft engine design, among others. This paper presents new algorithms for OpenMDAO that enable reconfigurable model execution. This concept refers to dynamically changing, during execution, one or more of: the variable sizes, solution algorithm, parallel load balancing, or set of variables-i.e., adding and removing components, perhaps to switch to a higher-fidelity sub-model. Any component can reconfigure at any point, even when running in parallel with other components, and the reconfiguration algorithm presented here performs the synchronized updates to all other components that are affected. A reconfigurable software framework for multidisciplinary design optimization enables new adaptive solvers, adaptive parallelization, and new applications such as gradient-based optimization with overset flow solvers and adaptive mesh refinement. Benchmarking results demonstrate the time savings for reconfiguration compared to setting up the model again from scratch, which can be significant in large-scale problems. Additionally, the new reconfigurability feature is applied to a mission profile optimization problem for commercial aircraft where both the parametrization of the mission profile and the time discretization are adaptively refined, resulting in computational savings of roughly 10% and the elimination of oscillations in the optimized altitude profile.

Hwang, John T.↗