Engineering PapersSearch

Engineering topics

Biegel, Bryan

Publications and source records attributed to Biegel, Bryan.

At least 37 records · Page 2

High-Order Semi-Discrete Central-Upwind Schemes for Multi-Dimensional Hamilton-Jacobi Equations

We present the first fifth order, semi-discrete central upwind method for approximating solutions of multi-dimensional Hamilton-Jacobi equations. Unlike most of the commonly used high order upwind schemes, our scheme is formulated as a Godunov-type scheme. The scheme is based on the fluxes of Kurganov-Tadmor and Kurganov-Tadmor-Petrova, and is derived for an arbitrary number of space dimensions. A theorem establishing the monotonicity of these fluxes is provided. The spacial discretization is based on a weighted essentially non-oscillatory reconstruction of the derivative. The accuracy and stability properties of our scheme are demonstrated in a variety of examples. A comparison between our method and other fifth-order schemes for Hamilton-Jacobi equations shows that our method exhibits smaller errors without any increase in the complexity of the computations.

Bryson, Steve

Experiences Using OpenMP Based on Compiler Directed Software DSM on a PC Cluster

In this work we report on our experiences running OpenMP (message passing) programs on a commodity cluster of PCs (personal computers) running a software distributed shared memory (DSM) system. We describe our test environment and report on the performance of a subset of the NAS (NASA Advanced Supercomputing) Parallel Benchmarks that have been automatically parallelized for OpenMP. We compare the performance of the OpenMP implementations with that of their message passing counterparts and discuss performance differences.

Hess, Matthias

A De-Centralized Scheduling and Load Balancing Algorithm for Heterogeneous Grid Environments

In the past two decades, numerous scheduling and load balancing techniques have been proposed for locally distributed multiprocessor systems. However, they all suffer from significant deficiencies when extended to a Grid environment: some use a centralized approach that renders the algorithm unscalable, while others assume the overhead involved in searching for appropriate resources to be negligible. Furthermore, classical scheduling algorithms do not consider a Grid node to be N-resource rich and merely work towards maximizing the utilization of one of the resources. In this paper we propose a new scheduling and load balancing algorithm for a generalized Grid model of N-resource nodes that not only takes into account the node and network heterogeneity, but also considers the overhead involved in coordinating among the nodes. Our algorithm is de-centralized, scalable, and overlaps the node coordination time of the actual processing of ready jobs, thus saving valuable clock cycles needed for making decisions. The proposed algorithm is studied by conducting simulations using the Message Passing Interface (MPI) paradigm.

Arora, Manish

Conductance of AFM Deformed Carbon Nanotubes

This viewgraph presentation provides information on the electrical conductivity of carbon nanotubes upon deformation by atomic force microscopy (AFM). The density of states and conductance were computed using four orbital tight-binding method with various parameterizations. Different chiralities develop bandgap that varies with chirality.

Svizhenko, Alexei

Automatic Multilevel Parallelization Using OpenMP

In this paper we describe the extension of the CAPO parallelization support tool to support multilevel parallelism based on OpenMP directives. CAPO generates OpenMP directives with extensions supported by the NanosCompiler to allow for directive nesting and definition of thread groups. We report first results for several benchmark codes and one full application that have been parallelized using our system.

Jin, Hao-Qiang

A System for Monitoring and Management of Computational Grids

As organizations begin to deploy large computational grids, it has become apparent that systems for observation and control of the resources, services, and applications that make up such grids are needed. Administrators must observe the operation of resources and services to ensure that they are operating correctly and they must control the resources and services to ensure that their operation meets the needs of users. Users are also interested in the operation of resources and services so that they can choose the most appropriate ones to use. In this paper we describe a prototype system to monitor and manage computational grids and describe the general software framework for control and observation in distributed environments that it is based on.

Smith, Warren

Accessing Wind Tunnels From NASA's Information Power Grid

The NASA Ames wind tunnel customers are one of the first users of the Information Power Grid (IPG) storage system at the NASA Advanced Supercomputing Division. We wanted to be able to store their data on the IPG so that it could be accessed remotely in a secure but timely fashion. In addition, incorporation into the IPG allows future use of grid computational resources, e.g., for post-processing of data, or to do side-by-side CFD validation. In this paper, we describe the integration of grid data access mechanisms with the existing DARWIN web-based system that is used to access wind tunnel test data. We also show that the combined system has reasonable performance: wind tunnel data may be retrieved at 50Mbits/s over a 100 base T network connected to the IPG storage server.

Becker, Jeff

Computational Nanomechanics of Carbon Nanotubes and Composites

Nanomechanics of individual carbon and boron-nitride nanotubes and their application as reinforcing fibers in polymer composites has been reviewed with interplay of theoretical modeling, computer simulations and experimental observations. The emphasis in this work is on elucidating the multi-length scales of the problems involved, and of different simulation techniques that are needed to address specific characteristics of individual nanotubes and nanotube polymer-matrix interfaces. Classical molecular dynamics simulations are shown to be sufficient to describe the generic behavior such as strength and stiffness modulus but are inadequate to describe elastic limit and nature of plastic buckling at large strength. Quantum molecular dynamics simulations are shown to bring out explicit atomic nature dependent behavior of these nanoscale materials objects that are not accessible either via continuum mechanics based descriptions or through classical molecular dynamics based simulations. As examples, we discus local plastic collapse of carbon nanotubes under axial compression and anisotropic plastic buckling of boron-nitride nanotubes. Dependence of the yield strain on the strain rate is addressed through temperature dependent simulations, a transition-state-theory based model of the strain as a function of strain rate and simulation temperature is presented, and in all cases extensive comparisons are made with experimental observations. Mechanical properties of nanotube-polymer composite materials are simulated with diverse nanotube-polymer interface structures (with van der Waals interaction). The atomistic mechanisms of the interface toughening for optimal load transfer through recycling, high-thermal expansion and diffusion coefficient composite formation above glass transition temperature, and enhancement of Young's modulus on addition of nanotubes to polymer are discussed and compared with experimental observations.

Srivastava, Deepak

Automatic Multilevel Parallelization Using OpenMP

In this paper we describe the extension of the CAPO (CAPtools (Computer Aided Parallelization Toolkit) OpenMP) parallelization support tool to support multilevel parallelism based on OpenMP directives. CAPO generates OpenMP directives with extensions supported by the NanosCompiler to allow for directive nesting and definition of thread groups. We report some results for several benchmark codes and one full application that have been parallelized using our system.

Jin, Hao-Qiang

Cellular and Network Mechanisms Underlying Information Processing in a Simple Sensory System

Realistic, biophysically-based compartmental models were constructed of several primary sensory interneurons in the cricket cercal sensory system. A dynamic atlas of the afferent input to these cells was used to set spatio-temporal parameters for the simulated stimulus-dependent synaptic inputs. We examined the roles of dendritic morphology, passive membrane properties, and active conductances on the frequency tuning of the neurons. The sensitivity of narrow-band low pass interneurons could be explained entirely by the electronic structure of the dendritic arbors and the dynamic sensitivity of the SIZ. The dynamic characteristics of interneurons with higher frequency sensitivity required models with voltage-dependent dendritic conductances.

Jacobs, Gwen

Tensile Strength of Carbon Nanotubes Under Realistic Temperature and Strain Rate

Strain rate and temperature dependence of the tensile strength of single-wall carbon nanotubes has been investigated with molecular dynamics simulations. The tensile failure or yield strain is found to be strongly dependent on the temperature and strain rate. A transition state theory based predictive model is developed for the tensile failure of nanotubes. Based on the parameters fitted from high-strain rate and temperature dependent molecular dynamics simulations, the model predicts that a defect free micrometer long single-wall nanotube at 300 K, stretched with a strain rate of 1%/hour, fails at about 9 plus or minus 1% tensile strain. This is in good agreement with recent experimental findings.

Wei, Chen-Yu

Minimizing Cache Misses Using Minimum-Surface Bodies

A number of known techniques for improving cache performance in scientific computations involve the reordering of the iteration space. Some of these reorderings can be considered as coverings of the iteration space with the sets having good surface-to-volume ratio. Use of such sets reduces the number of cache misses in computations of local operators having the iteration space as a domain. First, we derive lower bounds which any algorithm must suffer while computing a local operator on a grid. Then we explore coverings of iteration spaces represented by structured and unstructured grids which allow us to approach these lower bounds. For structured grids we introduce a covering by successive minima tiles of the interference lattice of the grid. We show that the covering has low surface-to-volume ratio and present a computer experiment showing actual reduction of the cache misses achieved by using these tiles. For planar unstructured grids we show existence of a covering which reduces the number of cache misses to the level of structured grids. On the other hand, we present a triangulation of a 3-dimensional cube such that any local operator on the corresponding grid has significantly larger number of cache misses than a similar operator on a structured grid.

Frumkin, Michael

Analysis of Carbon Nanotube Metal-Semiconductor Diode Device

We study recently reported drain current Id-drain voltage Vd characteristics of a carbon nanotube metal semiconductor diode device with the gate voltage Vg applied to modulate the carrier density in the nanotube. The diode was kink-shaped at the metal-semiconductor interface. It was shown that (1) larger negative Vg blocked Id more effectively in the negative Vd region, resulting in the rectifying Id-Vd characteristics, and that (2) positive Vg allowed Id in the both Vd polarities, resulting in the non-rectifying characteristics. The negative Vd was the Schottky reverse direction, judging from the negligible Id behavior for a wide region of -4 V less than Vd less than 0 V, with Vg = -4 V. Such negative Vg would attract positive charges from the metallic electrodes (charge reservoir) to the nanotube and lower the nanotube Fermi energy (EF). With larger negative Vg, the experiment showed that the Schottky forward direction (Vd greater than 0) had a smaller turn-on voltage and the Schottky reverse direction (Vd less than 0) was more resistant to the tunneling breakdown. Therefore, the majority carriers in the transport would be electrons since they can see a lower tunneling barrier (shallower built-in potential) in the forward direction when EF is lowered, and a thicker tunneling barrier (Schottky barrier) in the reverse direction due to the reduction in the electron density when EF is lowered.

Yamada, Toshishige

Hybrid MPI+OpenMP Programming of an Overset CFD Solver and Performance Investigations

This report describes a two level parallelization of a Computational Fluid Dynamic (CFD) solver with multi-zone overset structured grids. The approach is based on a hybrid MPI+OpenMP programming model suitable for shared memory and clusters of shared memory machines. The performance investigations of the hybrid application on an SGI Origin2000 (O2K) machine is reported using medium and large scale test problems.

Djomehri, M. Jahed

Engineering DNA Conductance

This poster presentation details the small-scale engineering production of DNA based molecular devices, the theoretical feasibility of which is being assessed.

Addesi, C.

Relative Debugging of Automatically Parallelized Programs

We describe a system that simplifies the process of debugging programs produced by computer-aided parallelization tools. The system uses relative debugging techniques to compare serial and parallel executions in order to show where the computations begin to differ. If the original serial code is correct, errors due to parallelization will be isolated by the comparison. One of the primary goals of the system is to minimize the effort required of the user. To that end, the debugging system uses information produced by the parallelization tool to drive the comparison process. In particular, the debugging system relies on the parallelization tool to provide information about where variables may have been modified and how arrays are distributed across multiple processes. User effort is also reduced through the use of dynamic instrumentation. This allows us to modify, the program execution with out changing the way the user builds the executable. The use of dynamic instrumentation also permits us to compare the executions in a fine-grained fashion and only involve the debugger when a difference has been detected. This reduces the overhead of executing instrumentation.

Jost, Gabriele

Graph Partitioning for Parallel Applications in Heterogeneous Grid Environments

The problem of partitioning irregular graphs and meshes for parallel computations on homogeneous systems has been extensively studied. However, these partitioning schemes fail when the target system architecture exhibits heterogeneity in resource characteristics. With the emergence of technologies such as the Grid, it is imperative to study the partitioning problem taking into consideration the differing capabilities of such distributed heterogeneous systems. In our model, the heterogeneous system consists of processors with varying processing power and an underlying non-uniform communication network. We present in this paper a novel multilevel partitioning scheme for irregular graphs and meshes, that takes into account issues pertinent to Grid computing environments. Our partitioning algorithm, called MiniMax, generates and maps partitions onto a heterogeneous system with the objective of minimizing the maximum execution time of the parallel distributed application. For experimental performance study, we have considered both a realistic mesh problem from NASA as well as synthetic workloads. Simulation results demonstrate that MiniMax generates high quality partitions for various classes of applications targeted for parallel execution in a distributed heterogeneous environment.

Bisws, Rupak

Effects of Ordering Strategies and Programming Paradigms on Sparse Matrix Computations

The Conjugate Gradient (CG) algorithm is perhaps the best-known iterative technique to solve sparse linear systems that are symmetric and positive definite. For systems that are ill-conditioned, it is often necessary to use a preconditioning technique. In this paper, we investigate the effects of various ordering and partitioning strategies on the performance of parallel CG and ILU(O) preconditioned CG (PCG) using different programming paradigms and architectures. Results show that for this class of applications: ordering significantly improves overall performance on both distributed and distributed shared-memory systems, that cache reuse may be more important than reducing communication, that it is possible to achieve message-passing performance using shared-memory constructs through careful data ordering and distribution, and that a hybrid MPI+OpenMP paradigm increases programming complexity with little performance gains. A implementation of CG on the Cray MTA does not require special ordering or partitioning to obtain high efficiency and scalability, giving it a distinct advantage for adaptive applications; however, it shows limited scalability for PCG due to a lack of thread level parallelism.

Oliker, Leonid