Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Proactive Operations and Investment Planning via Stochastic Optimization to Enhance Power Systems’ Extreme Weather Resilience

We present scalable stochastic optimization approaches for improving power systems’ resilience to extreme weather events. We consider both proactive redispatch and transmission line hardening as alternatives for mitigating expected load shed due to extreme weather, resulting in large-scale stochastic linear programs (LPs) and mixed-integer linear programs (MILPs). We solve these stochastic optimization problems with progressive hedging (PH), a parallel, scenario-based decomposition algorithm. Our computational experiments indicate that our proposed method for enhancing power system resilience can provide high-quality solutions efficiently. With up to 128 scenarios on a 2,000-bus network, the operations (redispatch) and investment (hardening) resilience problems can be solved in approximately 6 min and 2 h of wall-clock time, respectively. Additionally, we solve the investment problems with up to 512 scenarios, demonstrating that the approach scales very well with the number of scenarios. Moreover, the method produces high quality solutions that result in statistically significant reductions in expected load shed. Our proposed approach can be augmented to incorporate a variety of other operational and investment resilience strategies, or a combination of such strategies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Designing workflows for materials characterization

Experimental science is enabled by the combination of synthesis, imaging, and functional characterization organized into evolving discovery loop. Synthesis of new material is typically followed by a set of characterization steps aiming to provide feedback for optimization or discover fundamental mechanisms. However, the sequence of synthesis and characterization methods and their interpretation, or research workflow, has traditionally been driven by human intuition and is highly domain specific. Here, we explore concepts of scientific workflows that emerge at the interface between theory, characterization, and imaging. In this study, we discuss the criteria by which these workflows can be constructed for special cases of multiresolution structural imaging and functional characterization, as a part of more general material synthesis workflows. Some considerations for theory–experiment workflows are provided. We further pose that the emergence of user facilities and cloud labs disrupts the classical progression from ideation, orchestration, and execution stages of workflow development. To accelerate this transition, we propose the framework for workflow design, including universal hyperlanguages describing laboratory operation, ontological domain matching, reward functions and their integration between domains, and policy development for workflow optimization. These tools will enable knowledge-based workflow optimization; enable lateral instrumental networks, sequential and parallel orchestration of characterization between dissimilar facilities; and empower distributed research.

36 MATERIALS SCIENCE↗

Scaling Ultrahigh-Resolution E3SM Land Model for Leadership-Class Supercomputers

This paper presents advancements in scaling the ultrahigh-resolution E3SM Land Model (uELM) for deployment on leadership-class supercomputers, addressing the increased demand for km-scale Earth system modeling. By focusing on km-scale ELM simulations, we enhance predictive capabilities for climate interactions, facilitating improved responses to climate change impacts on energy systems, agriculture, and water resources. Our approach leverages innovative software architecture optimizations, sophisticated data handling techniques, and advanced parallel processing, achieving strong scalability on two leadership supercomputers (2400 nodes (105,600 cores) on Summit, and 1200 nodes (76,800 cores) on Frontier). Results from extensive scalability assessments on the Summit and Frontier also demonstrate outstanding I/O performance (close to 400 GB/s write throughput) and the model's ability to efficiently handle increasing computational demands. This study not only establishes uELM's capability for high-resolution simulations over vast geographical domains, but also sets a foundation for future Earth system modeling breakthroughs.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Scalable Comparative Visualization of Ensembles of Call Graphs

Optimizing the performance of large-scale parallel codes is critical for efficient utilization of computing resources. Code developers often explore various execution parameters, such as hardware configurations, system software choices, and application parameters, and are interested in detecting and understanding bottlenecks in different executions. They often collect hierarchical performance profiles represented as call graphs, which combine performance metrics with their execution contexts. The crucial task of exploring multiple call graphs together is tedious and challenging because of the many structural differences in the execution contexts and significant variability in the collected performance metrics (e.g., execution runtime). In this paper, we present Ensemble CallFlow to support the exploration of ensembles of call graphs using new types of visualizations, analysis, graph operations, and features. We introduce ensemble-Sankey , a new visual design that combines the strengths of resource-flow (Sankey) and box-plot visualization techniques. Whereas the resource-flow visualization can easily and intuitively describe the graphical nature of the call graph, the box plots overlaid on the nodes of Sankey convey the performance variability within the ensemble. Our interactive visual interface provides linked views to help explore ensembles of call graphs, e.g., by facilitating the analysis of structural differences, and identifying similar or distinct call graphs. Finally, we demonstrate the effectiveness and usefulness of our design through case studies on large-scale parallel codes.

97 MATHEMATICS AND COMPUTING↗

FutureTense

Protective vaccines and reliable diagnostics are essential tools for controlling viral diseases. However, the efficacy of these tools can be diminished by mutations in viral genomes. The delay between the emergence of new viral strains and the redesign of vaccines and diagnostics allows for continued viral transmission. Is it possible to address this challenge by computationally predicting viral genome sequence evolution? Can we “future-proof” vaccines and diagnostics by targeting both current and anticipated future sequence variants? While predicting viral evolution is still an unsolved, “grand challenge” problem in biology, the large, and rapidly growing, number of SARS-CoV-2 genome sequences provide an opportunity to quantify the ability of machine learning to predict viral genome sequence evolution. Towards this end, we have developed a simple computational model for predicting viral evolution at the level of individual nucleotides. The key metric for quantifying the per-base, prediction accuracy for viral evolution is the Mann-Whitney U statistic (or, equivalently, the area under the receiver operator curve). Since the Mann-Whitney U statistic is not a differentiable function, existing deep leaning packages (like Pytorch and Keras/TensorFlow) are not useful, as they require that the accuracy metric/objective function be analytically differentiable with respect to the model parameters. To overcome this challenge, we have implemented custom software, “FutureTense”, that can train a machine learning model by maximizing the non-differentiable Mann-Whitney U statistic. This software trains a machine learning model by exploring along the direction of the discrete gradient of the Mann-Whitney U statistic in the model parameter space. Parallel computing and genome sequence-specific optimizations are used to accelerate model training. The resulting machine learning model learns the observed high C->U mutation rates in the SARS-CoV-2 genome (which are potentially induced by host defenses) and provides prediction accuracies that are significantly better than one would expect from random chance. While predicting viral evolution is still quite far from a solved problem, the surprising performance of this simple model gives hope that the accuracy of predicting viral genome evolution can be further increased by more sophisticated approaches.

Gans, Jason↗

MultifidelityOpt- bohydra

Multifidelity Bayesian optimization with serial and MPI-enabled (parallel, asynchronous) workflows.

Grosskopf, Mike [Los Alamos National Laboratory]↗

High Throughput Source-less Plasma Deposition of Structured Silicon Anodes for Lithium-Ion Batteries

Amprius developed a manufacturing solution for silicon nanowire anode that relies on an inexpensive, high throughput, and high gas precursor utilization plasma deposition method that uses the anode foils as electrodes for plasma generation. The capacitively couple plasma (CCP) method is used in semiconductor and photovoltaic industry and Amprius modified existing high throughput equipment to use anode foils and to deposit amorphous silicon. The equipment was installed ahead of the program at Amprius site. The rest of the tasks included foil handling and process development. The equipment passed site acceptance tests (SAT) and the process parameter mapping was completed, indicating that the target process window limits produce output materials at the rate and with yield and specifications that meet the manufacturing target criteria. Amprius has hired supporting personnel to optimize processes and run the equipment. A parallel task verified the baseline performance of the silicon anode material, to be used as reference for the new manufacturing method.

25 ENERGY STORAGE↗

Deconvolute individual genomes from metagenome sequences through short read clustering

Metagenome assembly from short next-generation sequencing data is a challenging process due to its large scale and computational complexity. Clustering short reads by species before assembly offers a unique opportunity for parallel downstream assembly of genomes with individualized optimization. However, current read clustering methods suffer either false negative (under-clustering) or false positive (over-clustering) problems. Here we extended our previous read clustering software, SpaRC, by exploiting statistics derived from multiple samples in a dataset to reduce the under-clustering problem. Using synthetic and real-world datasets we demonstrated that this method has the potential to cluster almost all of the short reads from genomes with sufficient sequencing coverage. The improved read clustering in turn leads to improved downstream genome assembly quality.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-level Hierarchical Poly Tree computer architectures

Based on the concept of hierarchical substructuring, this paper develops an optimal multi-level Hierarchical Poly Tree (HPT) parallel computer architecture scheme which is applicable to the solution of finite element and difference simulations. Emphasis is given to minimizing computational effort, in-core/out-of-core memory requirements, and the data transfer between processors. In addition, a simplified communications network that reduces the number of I/O channels between processors is presented. HPT configurations that yield optimal superlinearities are also demonstrated. Moreover, to generalize the scope of applicability, special attention is given to developing: (1) multi-level reduction trees which provide an orderly/optimal procedure by which model densification/simplification can be achieved, as well as (2) methodologies enabling processor grading that yields architectures with varying types of multi-level granularity.

Padovan, Joe↗

Searching for patterns in remote sensing image databases using neural networks

We have investigated a method, based on a successful neural network multispectral image classification system, of searching for single patterns in remote sensing databases. While defining the pattern to search for and the feature to be used for that search (spectral, spatial, temporal, etc.) is challenging, a more difficult task is selecting competing patterns to train against the desired pattern. Schemes for competing pattern selection, including random selection and human interpreted selection, are discussed in the context of an example detection of dense urban areas in Landsat Thematic Mapper imagery. When applying the search to multiple images, a simple normalization method can alleviate the problem of inconsistent image calibration. Another potential problem, that of highly compressed data, was found to have a minimal effect on the ability to detect the desired pattern. The neural network algorithm has been implemented using the PVM (Parallel Virtual Machine) library and nearly-optimal speedups have been obtained that help alleviate the long process of searching through imagery.

Paola, Justin D.↗

Data Understanding Applied to Optimization

The goal of this research is to explore and develop software for supporting visualization and data analysis of search and optimization. Optimization is an ever-present problem in science. The theory of NP-completeness implies that the problems can only be resolved by increasingly smarter problem specific knowledge, possibly for use in some general purpose algorithms. Visualization and data analysis offers an opportunity to accelerate our understanding of key computational bottlenecks in optimization and to automatically tune aspects of the computation for specific problems. We will prototype systems to demonstrate how data understanding can be successfully applied to problems characteristic of NASA's key science optimization tasks, such as central tasks for parallel processing, spacecraft scheduling, and data transmission from a remote satellite.

Buntine, Wray↗

Practical Aerodynamic Design Optimization Based on the Navier-Stokes Equations and a Discrete Adjoint Method

Compressible and incompressible versions of a three-dimensional unstructured mesh Reynolds-averaged Navier-Stokes flow solver have been differentiated and resulting derivatives have been verified by comparisons with finite differences and a complex-variable approach. In this implementation, the turbulence model is fully coupled with the flow equations in order to achieve this consistency. The accuracy demonstrated in the current work represents the first time that such an approach has been successfully implemented. The accuracy of a number of simplifying approximations to the linearizations of the residual have been examined. A first-order approximation to the dependent variables in both the adjoint and design equations has been investigated. The effects of a "frozen" eddy viscosity and the ramifications of neglecting some mesh sensitivity terms were also examined. It has been found that none of the approximations yielded derivatives of acceptable accuracy and were often of incorrect sign. However, numerical experiments indicate that an incomplete convergence of the adjoint system often yield sufficiently accurate derivatives, thereby significantly lowering the time required for computing sensitivity information. The convergence rate of the adjoint solver relative to the flow solver has been examined. Inviscid adjoint solutions typically require one to four times the cost of a flow solution, while for turbulent adjoint computations, this ratio can reach as high as eight to ten. Numerical experiments have shown that the adjoint solver can stall before converging the solution to machine accuracy, particularly for viscous cases. A possible remedy for this phenomenon would be to include the complete higher-order linearization in the preconditioning step, or to employ a simple form of mesh sequencing to obtain better approximations to the solution through the use of coarser meshes. An efficient surface parameterization based on a free-form deformation technique has been utilized and the resulting codes have been integrated with an optimization package. Lastly, sample optimizations have been shown for inviscid and turbulent flow over an ONERA M6 wing. Drag reductions have been demonstrated by reducing shock strengths across the span of the wing. In order for large scale optimization to become routine, the benefits of parallel architectures should be exploited. Although the flow solver has been parallelized using compiler directives. The parallel efficiency is under 50 percent. Clearly, parallel versions of the codes will have an immediate impact on the ability to design realistic configurations on fine meshes, and this effort is currently underway.

Grossman, Bernard↗

Managing MDO Software Development Projects

Over the past decade, the NASA Langley Research Center developed a series of 'grand challenge' applications demonstrating the use of parallel and distributed computation and multidisciplinary design optimization. All but the last of these applications were focused on the high-speed civil transport vehicle; the final application focused on reusable launch vehicles. Teams of discipline experts developed these multidisciplinary applications by integrating legacy engineering analysis codes. As teams became larger and the application development became more complex with increasing levels of fidelity and numbers of disciplines, the need for applying software engineering practices became evident. This paper briefly introduces the application projects and then describes the approaches taken in project management and software engineering for each project; lessons learned are highlighted.

Townsend, J. C.↗

Understanding the Cray X1 System

This paper helps the reader understand the characteristics of the Cray X1 vector supercomputer system, and provides hints and information to enable the reader to port codes to the system. It provides a comparison between the basic performance of the X1 platform and other platforms that are available at NASA Ames Research Center. A set of codes, solving the Laplacian equation with different parallel paradigms, is used to understand some features of the X1 compiler. An example code from the NAS Parallel Benchmarks is used to demonstrate performance optimization on the X1 platform.

Cheung, Samson↗

Unity power factor switching regulator

A single or multiphase boost chopper regulator operating with unity power factor, for use such as to charge a battery is comprised of a power section for converting single or multiphase line energy into recharge energy including a rectifier (10), one inductor (L.sub.1) and one chopper (Q.sub.1) for each chopper phase for presenting a load (battery) with a current output, and duty cycle control means (16) for each chopper to control the average inductor current over each period of the chopper, and a sensing and control section including means (20) for sensing at least one load parameter, means (22) for producing a current command signal as a function of said parameter, means (26) for producing a feedback signal as a function of said current command signal and the average rectifier voltage output over each period of the chopper, means (28) for sensing current through said inductor, means (18) for comparing said feedback signal with said sensed current to produce, in response to a difference, a control signal applied to the duty cycle control means, whereby the average inductor current is proportionate to the average rectifier voltage output over each period of the chopper, and instantaneous line current is thereby maintained proportionate to the instantaneous line voltage, thus achieving a unity power factor. The boost chopper is comprised of a plurality of converters connected in parallel and operated in staggered phase. For optimal harmonic suppression, the duty cycles of the switching converters are evenly spaced, and by negative coupling between pairs 180.degree. out-of-phase, peak currents through the switches can be reduced while reducing the inductor size and mass.

Rippel, Wally E.↗

Improving Fidelity of Launch Vehicle Liftoff Acoustic Simulations

Launch vehicles experience high acoustic loads during ignition and liftoff affected by the interaction of rocket plume generated acoustic waves with launch pad structures. Application of highly parallelized Computational Fluid Dynamics (CFD) analysis tools optimized for application on the NAS computer systems such as the Loci/CHEM program now enable simulation of time-accurate, turbulent, multi-species plume formation and interaction with launch pad geometry and capture the generation of acoustic noise at the source regions in the plume shear layers and impingement regions. These CFD solvers are robust in capturing the acoustic fluctuations, but they are too dissipative to accurately resolve the propagation of the acoustic waves throughout the launch environment domain along the vehicle. A hybrid Computational Fluid Dynamics and Computational Aero-Acoustics (CFD/CAA) modeling framework has been developed to improve such liftoff acoustic environment predictions. The framework combines the existing highly-scalable NASA production CFD code, Loci/CHEM, with a high-order accurate discontinuous Galerkin (DG) solver, Loci/THRUST, developed in the same computational framework. Loci/THRUST employs a low dissipation, high-order, unstructured DG method to accurately propagate acoustic waves away from the source regions across large distances. The DG solver is currently capable of solving up to 4th order solutions for non-linear, conservative acoustic field propagation. Higher order boundary conditions are implemented to accurately model the reflection and refraction of acoustic waves on launch pad components. The DG solver accepts generalized unstructured meshes, enabling efficient application of common mesh generation tools for CHEM and THRUST simulations. The DG solution is coupled with the CFD solution at interface boundaries placed near the CFD acoustic source regions. Both simulations are executed simultaneously with coordinated boundary condition data exchange.

Liever, Peter↗