Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Performance Assessment of OVERFLOW on Distributed Computing Environment

The aerodynamic computer code, OVERFLOW, with a multi-zone overset grid feature, has been parallelized to enhance its performance on distributed and shared memory paradigms. Practical application benchmarks have been set to assess the efficiency of code's parallelism on high-performance architectures. The code's performance has also been experimented with in the context of the distributed computing paradigm on distant computer resources using the Information Power Grid (IPG) toolkit, Globus. Two parallel versions of the code, namely OVERFLOW-MPI and -MLP, have developed around the natural coarse grained parallelism inherent in a multi-zonal domain decomposition paradigm. The algorithm invokes a strategy that forms a number of groups, each consisting of a zone, a cluster of zones and/or a partition of a large zone. Each group can be thought of as a process with one or multithreads assigned to it and that all groups run in parallel. The -MPI version of the code uses explicit message-passing based on the standard MPI library for sending and receiving interzonal boundary data across processors. The -MLP version employs no message-passing paradigm; the boundary data is transferred through the shared memory. The -MPI code is suited for both distributed and shared memory architectures, while the -MLP code can only be used on shared memory platforms. The IPG applications are implemented by the -MPI code using the Globus toolkit. While a computational task is distributed across multiple computer resources, the parallelism can be explored on each resource alone. Performance studies are achieved with some practical aerodynamic problems with complex geometries, consisting of 2.5 up to 33 million grid points and a large number of zonal blocks. The computations were executed primarily on SGI Origin 2000 multiprocessors and on the Cray T3E. OVERFLOW's IPG applications are carried out on NASA homogeneous metacomputing machines located at three sites, Ames, Langley and Glenn. Plans for the future will exploit the distributed parallel computing capability on various homogeneous and heterogeneous resources and large scale benchmarks. Alternative IPG toolkits will be used along with sophisticated zonal grouping strategies to minimize the communication time across the computer resources.

Djomehri, M. Jahed↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngaart, Rob F.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngarrt, Rob F.↗

A Fast and Scalable Genetic Algorithm-Based Approach for Planning of Microgrids in Distribution Networks: Preprint

As a result of climate change, extreme weather events are occurring more frequently and with increasing impact. This trend poses a significant challenge for distribution utilities and system operators to ensure that there is uninterrupted power supply to critical loads in their networks; thus, the level of proactive preparation of the distribution system to be able to handle severe impacts of extreme weather events represents the system's resilience. One method that distribution systems use to prepare for extreme events is to form multiple microgrids and thereby isolate themselves from the grid supply by using local generation as much as possible to supply critical loads. But partitioning an existing system into multiple feasible islands capable of supporting critical loads is still challenging for distribution systems - first, because of the size of the graph partitioning problem and, second, because of the difficulty in properly formulating the desired attributes of such islands or microgrids. Therefore, this paper presents a genetic algorithmbased approach that facilitates incorporating multiple objectives for grid partitioning by formulating two types of problems - node allocation and edge elimination - and it considers multiple topological and resilience-enhancing objectives. The performance of the proposed genetic algorithm-based approach is numerically evaluated on multiple test systems as well as on a real distribution feeder in Colorado, USA.

edge elimination↗

Development of mixed time partition procedures for thermal analysis of structures

The computational methods used to predict and optimize the thermal-structural behavior of aerospace vehicle structures are reviewed. In general, two classes of algorithms, implicit and explicit, are used in transient thermal analysis of structures. Each of these two methods has its own merits. Due to the different time scales of the mechanical and thermal responses, the selection of a time integration method can be a difficult yet critical factor in the efficient solution of such problems. Therefore mixed time integration methods for transient thermal analysis of structures are being developed. This proposed methodology would be readily adaptable to existing computer programs for structural thermal analysis.

Liu, W. K.↗

A high-speed lossless data compression system for space applications

This paper reports on the integration of a lossless data compression/decompression chipset into a space data system architecture. For its compression engine, the data system incorporates the Universal Source Encoder (USE) designed for the NASA/Goddard Space Flight Center. Currently, the data compression testbed generates video frames consisting of 512 lines of 512 pixels having 8-bit resolution. Each image is passed through the USE where the lines are internally partitioned into 16-word blocks. These blocks are adaptively encoded across widely varying entropy levels using a Rice 12-option set coding algorithm. The current system operates at an Input/Output rate of 10 Msamples/s or 80 Mbits/s for each buffered input line. Frame and line synchronization for each image are maintained through the use of uniquely decodable command words. Length information of each variable length compressed image line is also included in the output stream. The data and command information are passed to the next stage of the system architecture through a serial fiber-optic transmitter. The initial segment of this stage consists of packetizer hardware which adds an appropriate CCSDS header to the received source data. An uncompressed mode is optionally available to pass image lines directly to the packetizer hardware. A data decompression testbed has also been developed to confirm the data compression operation.

Miko, Joe↗

Guidance of Nonlinear Systems

The paper describes a method for guiding a dynamic system through a given set of points. The paradigm is a fully automatic aircraft subject to air traffic control (ATC). The ATC provides a sequence of way points through which the aircraft trajectory must pass. The way points typically specify time, position, and velocity. The guidance problem is to synthesize a system state trajectory which satisfies both the ATC and aircraft constraints. Complications arise because the controlled process is multi-dimensional, multi-axis, nonlinear, highly coupled, and the state space is not flat. In addition, there is a multitude of possible operating modes, which may number in the hundreds. Each such mode defines a distinct state space model of the process by specifying the state space coordination, the partition of the controls into active controls and configuration controls, and the output map. Furthermore, mode transitions must be smooth. The guidance algorithm is based on the inversion of the pure feedback approximations, which is followed by iterative corrections for the effects of zero dynamics. The paper describes the structure and modules of the algorithm, and the performance is illustrated by several example aircraft maneuvers.

Meyer, George↗

Progress in Parallel Schur Complement Preconditioning for Computational Fluid Dynamics

We consider preconditioning methods for nonself-adjoint advective-diffusive systems based on a non-overlapping Schur complement procedure for arbitrary triangulated domains. The ultimate goal of this research is to develop scalable preconditioning algorithms for fluid flow discretizations on parallel computing architectures. In our implementation of the Schur complement preconditioning technique, the triangulation is first partitioned into a number of subdomains using the METIS multi-level k-way partitioning code. This partitioning induces a natural 2X2 partitioning of the p.d.e. discretization matrix. By considering various inverse approximations of the 2X2 system, we have developed a family of robust preconditioning techniques. A computer code based on these ideas has been developed and tested on the IBM SP2 and the SGI Power Challenge array using MPI message passing protocol. A number of example CFD calculations will be presented to illustrate and assess various Schur complement approximations.

Barth, Timothy J.↗

Guidance of Nonlinear Systems

The paper describes a method for guiding a dynamic system through a given set of points. The paradigm is a fully automatic aircraft subject to air traffic control (ATC). The ATC provides a sequence of way points through which the aircraft trajectory must pass. The way points typically specify time, position, and velocity. The guidance problem is to synthesize a system state trajectory which satisfies both the ATC and aircraft constraints. Complications arise because the controlled process is multi-dimensional, multi-axis, nonlinear, highly coupled, and the state space is not flat. In addition, there is a multitude of possible operating modes, which may number in the hundreds. Each such mode defines a distinct state space model of the process by specifying the state space coordinatization, the partition of the controls into active controls and configuration controls, and the output map. Furthermore, mode transitions must be smooth. The guidance algorithm is based on the inversion of the pure feedback approximations, which is followed by iterative corrections for the effects of zero dynamics. The paper describes the structure and modules of the algorithm, and the performance is illustrated by several example aircraft maneuvers.

Meyer, George↗

A Parallel Non-Overlapping Domain-Decomposition Algorithm for Compressible Fluid Flow Problems on Triangulated Domains

This paper considers an algebraic preconditioning algorithm for hyperbolic-elliptic fluid flow problems. The algorithm is based on a parallel non-overlapping Schur complement domain-decomposition technique for triangulated domains. In the Schur complement technique, the triangulation is first partitioned into a number of non-overlapping subdomains and interfaces. This suggests a reordering of triangulation vertices which separates subdomain and interface solution unknowns. The reordering induces a natural 2 x 2 block partitioning of the discretization matrix. Exact LU factorization of this block system yields a Schur complement matrix which couples subdomains and the interface together. The remaining sections of this paper present a family of approximate techniques for both constructing and applying the Schur complement as a domain-decomposition preconditioner. The approximate Schur complement serves as an algebraic coarse space operator, thus avoiding the known difficulties associated with the direct formation of a coarse space discretization. In developing Schur complement approximations, particular attention has been given to improving sequential and parallel efficiency of implementations without significantly degrading the quality of the preconditioner. A computer code based on these developments has been tested on the IBM SP2 using MPI message passing protocol. A number of 2-D calculations are presented for both scalar advection-diffusion equations as well as the Euler equations governing compressible fluid flow to demonstrate performance of the preconditioning algorithm.

Barth, Timothy J.↗

Further results on finite-state codes

A general construction for finite-state (FS) codes is applied to some well-known block codes. Subcodes of the (24,12) Golay code are used to generate two optimal FS codes with d sub free = 12 and 16. A partition of the (16,8) Nordstrom-Robinson code yields a d sub free = 10 FS code. Simulation results are shown and decoding algorithms are briefly discussed.

Pollara, F.↗

Slices: A Scalable Partitioner for Finite Element Meshes

A parallel partitioner for partitioning unstructured finite element meshes on distributed memory architectures is developed. The element based partitioner can handle mixtures of different element types. All algorithms adopted in the partitioner are scalable, including a communication template for unpredictable incoming messages, as shown in actual timing measurements.

partitioner finite element meshes parallel computi↗

Assimilation of Gridded Terrestrial Water Storage Observations from GRACE into a Land Surface Model

Observations of terrestrial water storage (TWS) from the Gravity Recovery and Climate Experiment (GRACE) satellite mission have a coarse resolution in time (monthly) and space (roughly 150,000 km(sup 2) at midlatitudes) and vertically integrate all water storage components over land, including soil moisture and groundwater. Data assimilation can be used to horizontally downscale and vertically partition GRACE-TWS observations. This work proposes a variant of existing ensemble-based GRACE-TWS data assimilation schemes. The new algorithm differs in how the analysis increments are computed and applied. Existing schemes correlate the uncertainty in the modeled monthly TWS estimates with errors in the soil moisture profile state variables at a single instant in the month and then apply the increment either at the end of the month or gradually throughout the month. The proposed new scheme first computes increments for each day of the month and then applies the average of those increments at the beginning of the month. The new scheme therefore better reflects submonthly variations in TWS errors. The new and existing schemes are investigated here using gridded GRACE-TWS observations. The assimilation results are validated at the monthly time scale, using in situ measurements of groundwater depth and soil moisture across the U.S. The new assimilation scheme yields improved (although not in a statistically significant sense) skill metrics for groundwater compared to the open-loop (no assimilation) simulations and compared to the existing assimilation schemes. A smaller impact is seen for surface and root-zone soil moisture, which have a shorter memory and receive smaller increments from TWS assimilation than groundwater. These results motivate future efforts to combine GRACE-TWS observations with observations that are more sensitive to surface soil moisture, such as L-band brightness temperature observations from Soil Moisture Ocean Salinity (SMOS) or Soil Moisture Active Passive (SMAP). Finally, we demonstrate that the scaling parameters that are applied to the GRACE observations prior to assimilation should be consistent with the land surface model that is used within the assimilation system.

GRACE↗

A proximity‐based image‐processing algorithm for colloid assignment in segmented multiphase flow datasets

Summary Colloidal transport and deposition are of both environmental and engineering importance. Easier access to x‐ray microtomography (XMT) coupled with improved imaging resolution has made XMT a unique and viable tool for visualizing and quantifying these processes. Currently, there is scant information in the literature addressing colloid segmentation and analysis in saturated and unsaturated porous media, in particular related to spatial partitioning of colloids. To support this need, an approach to assign segmented colloidal particles and aggregates to different partitioning classes based on their proximity to different phases is presented here. The method uses different markers for each attachment site (e.g. wetting‐nonwetting phase interfaces). An example XMT dataset from a drainage experiment is used to demonstrate the efficacy of the image processing algorithms. Flow conditions, and fluid and colloid properties, can thus be compared to the behaviour of colloids within the porous medium. This algorithm can help elucidate colloidal deposition mechanisms and the importance of different attachment sites, explore the importance of fluid properties, as well as the arrangement and shape of the colloids.

BRUECK, C. L.↗

A Fully Implicit Time Accurate Method for Hypersonic Combustion: Application to Shock-induced Combustion Instability

A new fully implicit, time accurate algorithm suitable for chemically reacting, viscous flows in the transonic-to-hypersonic regime is described. The method is based on a class of Total Variation Diminishing (TVD) schemes and uses successive Gauss-Siedel relaxation sweeps. The inversion of large matrices is avoided by partitioning the system into reacting and nonreacting parts, but still maintaining a fully coupled interaction. As a result, the matrices that have to be inverted are of the same size as those obtained with the commonly used point implicit methods. In this paper we illustrate the applicability of the new algorithm to hypervelocity unsteady combustion applications. We present a series of numerical simulations of the periodic combustion instabilities observed in ballistic-range experiments of blunt projectiles flying at subdetonative speeds through hydrogen-air mixtures. The computed frequencies of oscillation are in excellent agreement with experimental data.

Yungster, Shaye↗

A Scalable Space-Time Domain Decomposition Approach for Solving Large Scale Nonlinear Regularized Inverse Ill Posed Problems in 4D Variational Data Assimilation

We address the development of innovative algorithms designed to solve the strong-constraint Four Dimensional Variational Data Assimilation (4DVar DA) problems in large scale applications. We present a space-time decomposition approach which employs the whole domain decomposition, i.e. both along the spacial and temporal direction in the overlapping case, and the partitioning of both the solution and the operator. Starting from the global functional defined on the entire domain, we get to a sort of regularized local functionals on the set of sub domains providing the order reduction of both the predictive and the Data Assimilation models. The algorithm convergence is developed. Performance in terms of reduction of time complexity and algorithmic scalability is discussed on the Shallow Water Equations on the sphere. The number of state variables in the model, the number of observations in an assimilation cycle, as well as numerical parameters as the discretization step in time and in space domain are defined on the basis of discretization grid used by data available at repository Ocean Synthesis/Reanalysis Directory of Hamburg University.

97 MATHEMATICS AND COMPUTING↗

Scalable implementation of polynomial filtering for density functional theory calculation in PARSEC

In this work, we present an efficient implementation of polynomial filtering methods in PARSEC, a real-space pseudopotential based Kohn–Sham density functional theory solver. The implementation described here improves upon a Chebyshev-filtered subspace iteration algorithm used in the previous version of PARSEC. We present a hybrid polynomial filtering scheme that combines Chebyshev-filtered subspace iteration and a spectrum slicing method that partitions the spectrum into several spectral slices and uses bandpass-filtered subspace iteration to compute approximate eigenpairs within each interior slice simultaneously. We describe a procedure to partition a spectrum and construct polynomial filters. We also discuss a number of practical issues such as the use of appropriate data layouts for carrying out the computation on a two-dimensional process grid and how to achieve good load balance by allocating an appropriate number of process groups to each spectral slice. Numerical examples are presented to demonstrate the effectiveness of the hybrid polynomial filtering method as well as the superior parallel scalability of spectrum slicing in comparison to that of Chebyshev-filtered subspace iteration.

97 MATHEMATICS AND COMPUTING↗

Ensemble transfer learning for the prediction of anti-cancer drug response

Abstract Transfer learning, which transfers patterns learned on a source dataset to a related target dataset for constructing prediction models, has been shown effective in many applications. In this paper, we investigate whether transfer learning can be used to improve the performance of anti-cancer drug response prediction models. Previous transfer learning studies for drug response prediction focused on building models to predict the response of tumor cells to a specific drug treatment. We target the more challenging task of building general prediction models that can make predictions for both new tumor cells and new drugs. Uniquely, we investigate the power of transfer learning for three drug response prediction applications including drug repurposing, precision oncology, and new drug development, through different data partition schemes in cross-validation. We extend the classic transfer learning framework through ensemble and demonstrate its general utility with three representative prediction algorithms including a gradient boosting model and two deep neural networks. The ensemble transfer learning framework is tested on benchmark in vitro drug screening datasets. The results demonstrate that our framework broadly improves the prediction performance in all three drug response prediction applications with all three prediction algorithms.

60 APPLIED LIFE SCIENCES↗