Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Test Driven Development of Scientific Models

Test-Driven Development (TDD) is a software development process that promises many advantages for developer productivity and has become widely accepted among professional software engineers. As the name suggests, TDD practitioners alternate between writing short automated tests and producing code that passes those tests. Although this overly simplified description will undoubtedly sound prohibitively burdensome to many uninitiated developers, the advent of powerful unit-testing frameworks greatly reduces the effort required to produce and routinely execute suites of tests. By testimony, many developers find TDD to be addicting after only a few days of exposure, and find it unthinkable to return to previous practices. Of course, scientific/technical software differs from other software categories in a number of important respects, but I nonetheless believe that TDD is quite applicable to the development of such software and has the potential to significantly improve programmer productivity and code quality within the scientific community. After a detailed introduction to TDD, I will present the experience within the Software Systems Support Office (SSSO) in applying the technique to various scientific applications. This discussion will emphasize the various direct and indirect benefits as well as some of the difficulties and limitations of the methodology. I will conclude with a brief description of pFUnit, a unit testing framework I co-developed to support test-driven development of parallel Fortran applications.

Clune, Thomas L.↗

Using the GeoFEST Faulted Region Simulation System

GeoFEST (the Geophysical Finite Element Simulation Tool) simulates stress evolution, fault slip and plastic/elastic processes in realistic materials, and so is suitable for earthquake cycle studies in regions such as Southern California. Many new capabilities and means of access for GeoFEST are now supported. New abilities include MPI-based cluster parallel computing using automatic PYRAMID/Parmetis-based mesh partitioning, automatic mesh generation for layered media with rectangular faults, and results visualization that is integrated with remote sensing data. The parallel GeoFEST application has been successfully run on over a half-dozen computers, including Intel Xeon clusters, Itanium II and Altix machines, and the Apple G5 cluster. It is not separately optimized for different machines, but relies on good domain partitioning for load-balance and low communication, and careful writing of the parallel diagonally preconditioned conjugate gradient solver to keep communication overhead low. Demonstrated thousand-step solutions for over a million finite elements on 64 processors require under three hours, and scaling tests show high efficiency when using more than (order of) 4000 elements per processor. The source code and documentation for GeoFEST is available at no cost from Open Channel Foundation. In addition GeoFEST may be used through a browser-based portal environment available to approved users. That environment includes semi-automated geometry creation and mesh generation tools, GeoFEST, and RIVA-based visualization tools that include the ability to generate a flyover animation showing deformations and topography. Work is in progress to support simulation of a region with several faults using 16 million elements, using a strain energy metric to adapt the mesh to faithfully represent the solution in a region of widely varying strain.

Geophyical Finite Element Simulation Tool (GeoFEST↗

Six Years of Parallel Computing at NAS (1987 - 1993): What Have we Learned?

In the fall of 1987 the age of parallelism at NAS began with the installation of a 32K processor CM-2 from Thinking Machines. In 1987 this was described as an "experiment" in parallel processing. In the six years since, NAS acquired a series of parallel machines, and conducted an active research and development effort focused on the use of highly parallel machines for applications in the computational aerosciences. In this time period parallel processing for scientific applications evolved from a fringe research topic into the one of main activities at NAS. In this presentation I will review the history of parallel computing at NAS in the context of the major progress, which has been made in the field in general. I will attempt to summarize the lessons we have learned so far, and the contributions NAS has made to the state of the art. Based on these insights I will comment on the current state of parallel computing (including the HPCC effort) and try to predict some trends for the next six years.

Simon, Horst D.↗

Fixing Amdahl's Law within the Limits of Accelerated Systems: FALLACY

Closeout report for FALLACY project. The performance of Data Model Convergence Initiative (DMC) applications on parallel machines is far below the limit set by Amdahl’s law. Whether the machine is based on many-core, GPUs, FPGAs, or a heterogeneous combination, usually the most significant bottleneck is accessing data from the memory system. Aligning with DMC’s HW/architecture thrust, this project developed a set of memory-centric tools called ‘MemGaze’ that inform the HW/SW stack about an application’s memory behavior, including data access latency and diagnosing poor data layout and data composition. Our approach uses architectural modeling and analysis of workload data accesses.

97 MATHEMATICS AND COMPUTING↗

Layer-Parallel Training of Deep Residual Neural Networks

Residual neural networks (ResNets) are a promising class of deep neural networks that have shown excellent performance for a number of learning tasks, e.g., image classification and recognition. Mathematically, ResNet architectures can be interpreted as forward Euler discretizations of a nonlinear initial value problem whose time-dependent control variables represent the weights of the neural network. Hence, training a ResNet can be cast as an optimal control problem of the associated dynamical system. For similar time-dependent optimal control problems arising in engineering applications, parallel-in-time methods have shown notable improvements in scalability. This paper demonstrates the use of those techniques for efficient and effective training of ResNets. The proposed algorithms replace the classical (sequential) forward and backward propagation through the network layers with a parallel nonlinear multigrid iteration applied to the layer domain. This adds a new dimension of parallelism across layers that is attractive when training very deep networks. From this basic idea, we derive multiple layer-parallel methods. The most efficient version employs a simultaneous optimization approach where updates to the network parameters are based on inexact gradient information in order to speed up the training process. Finally, using numerical examples from supervised classification, we demonstrate that the new approach achieves a training performance similar to that of traditional methods, but enables layer-parallelism and thus provides speedup over layer-serial methods through greater concurrency.

97 MATHEMATICS AND COMPUTING↗

Application of advanced multidisciplinary analysis and optimization methods to vehicle design synthesis

Advanced multidisciplinary analysis and optimization methods, namely system sensitivity analysis and non-hierarchical system decomposition, are applied to reduce the cost and improve the visibility of an automated vehicle design synthesis process. This process is inherently complex due to the large number of functional disciplines and associated interdisciplinary couplings. Recent developments in system sensitivity analysis as applied to complex non-hierarchic multidisciplinary design optimization problems enable the decomposition of these complex interactions into sub-processes that can be evaluated in parallel. The application of these techniques results in significant cost, accuracy, and visibility benefits for the entire design synthesis process.

Consoli, Robert David↗

An evaluation of machine processing techniques of ERTS-1 data for user applications

A broad study is described to evaluate a set of machine analysis and processing techniques applied to ERTS-1 data. Based on the analysis results in urban land use analysis and soil association mapping together with previously reported results in general earth surface feature identification and crop species classification, a profile of general applicability of this procedure is beginning to emerge. Put in the hands of a user who knows well the information needed from the data and also is familiar with the region to be analyzed it appears that significantly useful information can be generated by these methods. When supported by preprocessing techniques such as the geometric correction and temporal registration capabilities, final products readily useable by user agencies appear possible. In parallel with application, through further research, there is much potential for further development of these techniques both with regard to providing higher performance and in new situations not yet studied.

Landgrebe, D.↗

Algorithms and programming tools for image processing on the MPP, part 2

A number of algorithms were developed for image warping and pyramid image filtering. Techniques were investigated for the parallel processing of a large number of independent irregular shaped regions on the MPP. In addition some utilities for dealing with very long vectors and for sorting were developed. Documentation pages for the algorithms which are available for distribution are given. The performance of the MPP for a number of basic data manipulations was determined. From these results it is possible to predict the efficiency of the MPP for a number of algorithms and applications. The Parallel Pascal development system, which is a portable programming environment for the MPP, was improved and better documentation including a tutorial was written. This environment allows programs for the MPP to be developed on any conventional computer system; it consists of a set of system programs and a library of general purpose Parallel Pascal functions. The algorithms were tested on the MPP and a presentation on the development system was made to the MPP users group. The UNIX version of the Parallel Pascal System was distributed to a number of new sites.

Reeves, Anthony P.↗

Compact Ferroelectric Programmable Majority Gate for Compute-in-Memory Applications

In this study, a compact and novel ferroelectric (FE) programmable majority gate is proposed and its novel application in Binary Neural Network (BNNs) is investigated. We demonstrate: i) by integrating N metal-ferroelectric-metal (MFM) capacitors on the gate of a transistor (1T-N-MFM structure), a nonvolatile and programmable majority (MAJ) gate that performs MAJ of AND between the gate input and polarization is realized; ii) validation the functionality of our 3-input MAJ of AND gate through comprehensive theoretical and experimental investigations; iii) a compact implementation of 3-input MAJ of XNOR gate that leverages only five of our 3-input MAJ of AND gates connected in parallel; iv) application of MAJ of XNOR gates to replace the XNOR gates and the first layer of the adder tree in the BNNs for up to 21x area saving on top of eliminating the energy-hungry memory accesses due to the compute-in-memory nature.

97 MATHEMATICS AND COMPUTING↗

I/O Bottleneck Detection and Tuning: Connecting the Dots using Interactive Log Analysis

Using parallel file systems efficiently is a tricky problem due to inter-dependencies among multiple layers of I/O software, including high-level I/O libraries (HDF5, netCDF, etc.), MPI-IO, POSIX, and file systems (GPFS, Lustre, etc.). Profiling tools such as Darshan collect traces to help understand the I/O performance behavior. However, there are significant gaps in analyzing the collected traces and then applying tuning options offered by various layers of I/O software. Seeking to connect the dots between I/O bottleneck detection and tuning, we propose DXT Explorer, an interactive log analysis tool. In this paper, we present a case study using our interactive log analysis tool to identify and apply various I/O optimizations. We report an evaluation of performance improvement achieved for four I/O kernels extracted from science applications.

Bez, Jean Luca↗

Nonlinear optical control of chiral charge pumping in a topological Weyl semimetal

Solids with topologically robust electronic states exhibit unusual electronic and optical properties that do not exist in other materials. A particularly interesting example is chiral charge pumping, the so-called chiral anomaly, in recently discovered topological Weyl semimetals, where simultaneous application of parallel DC electric and magnetic fields creates an imbalance in the number of carriers of opposite topological charge (chirality). Here, using time-resolved terahertz measurements on the Weyl semimetal TaAs in a magnetic field, we optically interrogate the chiral anomaly by dynamically pumping the chiral charges and monitoring their subsequent relaxation of the nonequilibrium state. Theory based on Boltzmann transport shows that the observed effects originate from an optical nonlinearity in the chiral charge pumping process. Our measurements reveal that the nonequilibrium chiral excitation relaxation time is much greater than 1 ns. The observation of terahertz-controlled chiral carriers with long coherence times and topological protection suggests the application of Weyl semimetals for quantum optoelectronic technology.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A distributed parallel storage architecture and its potential application within EOSDIS

We describe the architecture, implementation, use of a scalable, high performance, distributed-parallel data storage system developed in the ARPA funded MAGIC gigabit testbed. A collection of wide area distributed disk servers operate in parallel to provide logical block level access to large data sets. Operated primarily as a network-based cache, the architecture supports cooperation among independently owned resources to provide fast, large-scale, on-demand storage to support data handling, simulation, and computation.

Johnston, William E.↗

Insights From Dayflow: A Historical Streamflow Reanalysis Dataset for the Conterminous United States

Abstract Reconstructed historical streamflow time series can supplement limited streamflow gauge observations. However, there are common challenges of typical modeling approaches: process‐based hydrologic models can be data/computation‐intensive, and statistics‐based models can be region/stream‐specific. Here we present a nationally scalable modeling framework integrating the simulated runoff from the Variable Infiltration Capacity (VIC) model with the Routing Application for Parallel computatIon of Discharge (RAPID) routing model leveraging high‐performance computing. We demonstrate an efficient method of assimilating streamflow at US Geological Survey (USGS) streamflow monitoring sites using a simple hierarchical approach in the VIC‐RAPID framework. The result is a reconstructed 36‐year (1980–2015) daily and monthly streamflow dataset (Dayflow) at ∼2.7 million NHDPlusV2 stream reaches in the conterminous US (CONUS). We perform a comprehensive evaluation at 7,526 USGS sites and characterize their error statistics. The results demonstrate that 49% of the USGS sites demonstrate Kling–Gupta Efficiency (KGE) > 0.5 and 58% of the sites show percentage bias within ±20% for the daily naturalized streamflow. Streamflow data assimilation across CONUS shows an overall improvement over naturalized streamflow, notably in the western semiarid‐to‐arid regions. Comparison to other national and global streamflow reanalysis datasets such as the National Water Model and Global Reach‐scale A priori Discharge Estimates for SWOT demonstrates improved KGE, reduced bias, and directions for Dayflow improvements. Investigations of error statistics with key hydrologic, hydroclimatic, and geomorphologic basin characteristics reveal region‐specific patterns which may help improve future framework applications. Overall, Dayflow may enable a better understanding of hydrologic conditions in a changing environment, especially in locations currently not represented by streamflow monitoring networks.

54 ENVIRONMENTAL SCIENCES↗

An approach to real-time simulation using parallel processing

Current applications of real-time simulations to the development of complex aircraft propulsion system controls have demonstrated the need for accurate, portable, and low-cost simulators. This paper presents a preliminary simulator design that uses a parallel computer organization to provide these features. The hardware and software for this prototype simulator are discussed. A detailed discussion of the inter-computer data transfer mechanism is also presented.

Blech, R. A.↗

Automated Generation of Message-Passing Programs: An Evaluation Using CAPTools

Scientists at NASA Ames Research Center have been developing computational aeroscience applications on highly parallel architectures over the past ten years. During that same time period, a steady transition of hardware and system software also occurred, forcing us to expend great efforts into migrating and re-coding our applications. As applications and machine architectures become increasingly complex, the cost and time required for this process will become prohibitive. In this paper, we present the first set of results in our evaluation of interactive parallelization tools. In particular, we evaluate CAPTool's ability to parallelize computational aeroscience applications. CAPTools was tested on serial versions of the NAS Parallel Benchmarks and ARC3D, a computational fluid dynamics application, on two platforms: the SGI Origin 2000 and the Cray T3E. This evaluation includes performance, amount of user interaction required, limitations and portability. Based on these results, a discussion on the feasibility of computer aided parallelization of aerospace applications is presented along with suggestions for future work.

Hribar, Michelle R.↗

Automated Generation of Message-Passing Programs: An Evaluation Using CAPTools

Scientists at NASA Ames Research Center have been developing computational aeroscience applications on highly parallel architectures over the past ten years. During that same time period, a steady transition of hardware and system software also occurred, forcing us to expend great efforts into migrating and re-coding our applications. As applications and machine architectures become increasingly complex, the cost and time required for this process will become prohibitive. In this paper, we present the first set of results in our evaluation of interactive parallelization tools. In particular, we evaluate CAPTool's ability to parallelize computational aeroscience applications. CAPTools was tested on serial versions of the NAS Parallel Benchmarks and ARC3D, a computational fluid dynamics application, on two platforms: the SGI Origin 2000 and the Cray T3E. This evaluation includes performance, amount of user interaction required, limitations and portability. Based on these results, a discussion on the feasibility of computer aided parallelization of aerospace applications is presented along with suggestions for future work.

Hribar, Michelle R.↗

An exploration of online-simulation-driven portfolio scheduling in Workflow Management Systems

Workflow Management Systems used to automate the execution of scientific workflow applications on parallel and distributed computing platforms must make scheduling decisions at runtime. A large number of workflow scheduling algorithms have been proposed in the literature, but often these algorithms are evaluated based on simplifying assumptions that may not hold in practice. Furthermore, published algorithm evaluation and/or comparison results are necessarily only for a subset of all possible scenarios, and thus may not include scenarios relevant to particular use-cases. Consequently, it is difficult for Workflow Management Systems (WMSs) developers to decide which scheduling algorithm should be implemented. To obviate this difficulty, one possible approach is to implement a portfolio of scheduling algorithms and select the most effective algorithm at runtime. One method for performing this selection is to run an online simulation for each algorithm in the portfolio. The algorithm that leads to the best performance, in simulation, is selected for future use. The above simulation-driven portfolio scheduling (SDPS) approach has been proposed in a few parallel and distributed computing contexts. The main objective of this work is to evaluate the feasibility and potential merit of SDPS if implemented in WMSs. Here we perform this evaluation using simulated WMS executions, where the simulations are instantiated from real-world platform and workflow configurations. Our main finding is that SDPS is on par with or outperforms an approach in which a single algorithm is used, where this algorithm is the one that performs best on average across all our experimental scenarios. Furthermore, we find that SDPS remains an attractive proposition even in the presence of high levels of simulation error and for simulators with relatively low levels of sophistication. In many of our experimental scenarios we find that mitigating simulation error at runtime can further improve performance. Finally, we show that simulation overhead can be made sufficiently low for SDPS to be feasible in practice.

97 MATHEMATICS AND COMPUTING↗

Unraveling the 2021 Central Tennessee flood event using a hierarchical multi-model inundation modeling framework

Flood prediction systems need hierarchical atmospheric, hydrologic, and hydraulic models to predict rainfall, runoff, streamflow, and floodplain inundation. The accuracy of such systems depends on the error propagation through the modeling chain, sensitivity to input data, and choice of models. In this study, we used multiple precipitation forcings (hindcast and forecast) to drive hydrologic and hydrodynamic models to analyze the impacts of various drivers on the estimates of flood inundation depth and extent. We implement this framework to unravel the August 2021 extreme flooding event that occurred in Central Tennessee, USA. We used two radar-based quantitative precipitation estimates (STAGE4 and MRMS) as well as quantitative precipitation forecasts (QPF) from the National Weather Service Weather Prediction Center (WPC) to drive a series of models in the hierarchical framework, including the Variable Infiltration Capacity (VIC) land surface model, the Routing Application for Parallel Computation of Discharge (RAPID) river routing model, and the AutoRoute and TRITON inundation models. An evaluation with observed high-water marks demonstrates that the framework can reasonably simulate flood inundation. Despite the complex error propagation mechanism of the modeling chain, we show that inundation estimates are most sensitive to rainfall estimates. Most notably, QPF significantly underestimates flood magnitudes and inundations leading to unanticipated severe flooding for all stakeholders involved in the event. Finally, we discuss the implications of the hydrodynamic modeling framework for real-time flood forecasting.

54 ENVIRONMENTAL SCIENCES↗