Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “performance portable algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ProperCAD: A portable object-oriented parallel environment for VLSI CAD

Most parallel algorithms for VLSI CAD proposed to date have one important drawback: they work efficiently only on machines that they were designed for. As a result, algorithms designed to date are dependent on the architecture for which they are developed and do not port easily to other parallel architectures. A new project under way to address this problem is described. A Portable object-oriented parallel environment for CAD algorithms (ProperCAD) is being developed. The objectives of this research are (1) to develop new parallel algorithms that run in a portable object-oriented environment (CAD algorithms using a general purpose platform for portable parallel programming called CARM is being developed and a C++ environment that is truly object-oriented and specialized for CAD applications is also being developed); and (2) to design the parallel algorithms around a good sequential algorithm with a well-defined parallel-sequential interface (permitting the parallel algorithm to benefit from future developments in sequential algorithms). One CAD application that has been implemented as part of the ProperCAD project, flat VLSI circuit extraction, is described. The algorithm, its implementation, and its performance on a range of parallel machines are discussed in detail. It currently runs on an Encore Multimax, a Sequent Symmetry, Intel iPSC/2 and i860 hypercubes, a NCUBE 2 hypercube, and a network of Sun Sparc workstations. Performance data for other applications that were developed are provided: namely test pattern generation for sequential circuits, parallel logic synthesis, and standard cell placement.

Ramkumar, Balkrishna↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngaart, Rob F.↗

Charon Message-Passing Toolkit for Scientific Computations

The Charon toolkit for piecemeal development of high-efficiency parallel programs for scientific computing is described. The portable toolkit, callable from C and Fortran, provides flexible domain decompositions and high-level distributed constructs for easy translation of serial legacy code or design to distributed environments. Gradual tuning can subsequently be applied to obtain high performance, possibly by using explicit message passing. Charon also features general structured communications that support stencil-based computations with complex recurrences. Through the separation of partitioning and distribution, the toolkit can also be used for blocking of uni-processor code, and for debugging of parallel algorithms on serial machines. An elaborate review of recent parallelization aids is presented to highlight the need for a toolkit like Charon. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability. Some performance results of parallelizing the NAS Parallel Benchmark SP program using Charon are given, showing good scalability.

VanderWijngarrt, Rob F.↗

Automated Subpixel Snow Parameter Mapping with AVIRIS Data

We describe an automated algorithm (MEMSCAG) for mapping subpixel snow covered area (SCA) and snow grain size with AVIRIS data. The algorithm is based on the multiple endmember approach to spectral mixture analysis in which the spectral endmembers and the number of endmembers can vary on a pixel-by-pixel basis. This approach accounts for surface cover heterogeneity within a scene. The mixture analysis runs on endmembers from a spectral library of snow, vegetation, rock, soil, and lake ice spectra. Snow endmembers of varying grain size were produced with a radiative transfer model. All non-snow endmembers were collected with a portable field spectrometer. Mapping is performed through sequential 2-endmember, 3-endmember, and 4- endmember mixture model runs, each subject to constraints on RMS, residuals, fractions and priority. Grain size is determined by the grain size of the snow endmember used in the optimal mixture model. We apply MEMSCAG to AVIRIS data collected over Mammoth Mountain, CA and the northern site of the BOREAS in Manitoba, Canada. MEMSCAG produces appropriate snow covered area estimates in all regions. A preliminary comparison of grain size estimates from MEMSCAG with field measurements demonstrates high accuracy.

Painter, Thomas H.↗

TROJID: A portable software package for upper-stage trajectory optimization

Performance optimization for upper-stage exoatmospheric vehicles often is performed within the framework of a full capability trajectory simulation package requiring either a large mainframe computer or powerful work-station. Since these software packages tend to include capabilities providing for high-fidelity boost and reentry simulations, the programs usually are quite large and not very portable. The program TROJID is an attempt to provide an environment for the optimization of upper-stage trajectories within a small package capable of being run on a standard desktop microcomputer. Utilizing a state-of-the-art nonlinear programming algorithm and a trajectory simulator implementing impulsive burns and an analytic coast phase propagator, TROJID is capable of producing trajectories for optimal multi-burn upper-stage orbit transfers. The package has been designed to allow full generality in definition of both the trajectory simulator and the parameter optimization problem.

Hammes, Steven M.↗

A Novel 24 GHz One-Shot, Rapid and Portable Microwave Imaging System

Development of microwave and millimeter wave imaging systems has received significant attention in the past decade. Signals at these frequencies penetrate inside of dielectric materials and have relatively small wavelengths. Thus. imaging systems at these frequencies can produce images of the dielectric and geometrical distributions of objects. Although there are many different approaches for imaging at these frequencies. they each have their respective advantageous and limiting features (hardware. reconstruction algorithms). One method involves electronically scanning a given spatial domain while recording the coherent scattered field distribution from an object. Consequently. different reconstruction or imaging techniques may be used to produce an image (dielectric distribution and geometrical features) of the object. The ability to perform this accurate~v and fast can lead to the development of a rapid imaging system that can be used in the same manner as a video camera. This paper describes the design of such a system. operating at 2-1 GHz. using modulated scatterer technique applied to 30 resonant slots in a prescribed measurement domain.

Ghasr, M. T.↗

Statistical results from the Virginia Tech propagation experiment using the Olympus 12, 20, and 30 GHz satellite beacons

Virginia Tech has performed a comprehensive propagation experiment using the Olympus satellite beacons at 12.5, 19.77, and 29.66 GHz (which we refer to as 12, 20, and 30 GHz). Four receive terminals were designed and constructed, one terminal at each frequency plus a portable one with 20 and 30 GHz receivers for microscale and scintillation studies. Total power radiometers were included in each terminal in order to set the clear air reference level for each beacon and also to predict path attenuation. More details on the equipment and the experiment design are found elsewhere. Statistical results for one year of data collection were analyzed. In addition, the following studies were performed: a microdiversity experiment in which two closely spaced 20 GHz receivers were used; a comparison of total power and Dicke switched radiometer measurements, frequency scaling of scintillations, and adaptive power control algorithm development. Statistical results are reported.

Stutzman, Warren L.↗

A New Electromagnetic Instrument for Thickness Gauging of Conductive Materials

Eddy current techniques are widely used to measure the thickness of electrically conducting materials. The approach, however, requires an extensive set of calibration standards and can be quite time consuming to set up and perform. Recently, an electromagnetic sensor was developed which eliminates the need for impedance measurements. The ability to monitor the magnitude of a voltage output independent of the phase enables the use of extremely simple instrumentation. Using this new sensor a portable hand-held instrument was developed. The device makes single point measurements of the thickness of nonferromagnetic conductive materials. The technique utilized by this instrument requires calibration with two samples of known thicknesses that are representative of the upper and lower thickness values to be measured. The accuracy of the instrument depends upon the calibration range, with a larger range giving a larger error. The measured thicknesses are typically within 2-3% of the calibration range (the difference between the thin and thick sample) of their actual values. In this paper the design, operational and performance characteristics of the instrument along with a detailed description of the thickness gauging algorithm used in the device are presented.

Fulton, J. P.↗

Load Balancing Sequences of Unstructured Adaptive Grids

Mesh adaption is a powerful tool for efficient unstructured grid computations but causes load imbalance on multiprocessor systems. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. This paper makes several important additions to our previous work. First, a new remapping cost model is presented and empirically validated on an SP2. Next, our load balancing strategy is applied to sequences of dynamically adapted unstructured grids. Results indicate that our framework is effective on many processors for both steady and unsteady problems with several levels of adaption. Additionally, we demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required for a fine initial mesh. Finally, we show that the data remapping overhead can be significantly reduced by applying our heuristic processor reassignment algorithm.

Biswas, Rupak↗

NASA Tech Briefs, November 2007

Topics include: Wireless Measurement of Contact and Motion Between Contact Surfaces; Wireless Measurement of Rotation and Displacement Rate; Portable Microleak-Detection System; Free-to-Roll Testing of Airplane Models in Wind Tunnels; Cryogenic Shrouds for Testing Thermal-Insulation Panels; Optoelectronic System Measures Distances to Multiple Targets; Tachometers Derived From a Brushless DC Motor; Algorithm-Based Fault Tolerance for Numerical Subroutines; Computational Support for Technology- Investment Decisions; DSN Resource Scheduling; Distributed Operations Planning; Phase-Oriented Gear Systems; Freeze Tape Casting of Functionally Graded Porous Ceramics; Electrophoretic Deposition on Porous Non- Conductors; Two Devices for Removing Sludge From Bioreactor Wastewater; Portable Unit for Metabolic Analysis; Flash Diffusivity Technique Applied to Individual Fibers; System for Thermal Imaging of Hot Moving Objects; Large Solar-Rejection Filter; Improved Readout Scheme for SQUID-Based Thermometry; Error Rates and Channel Capacities in Multipulse PPM; Two Mathematical Models of Nonlinear Vibrations; Simpler Adaptive Selection of Golomb Power-of- Two Codes; VCO PLL Frequency Synthesizers for Spacecraft Transponders; Wide Tuning Capability for Spacecraft Transponders; Adaptive Deadband Synchronization for a Spacecraft Formation; Analysis of Performance of Stereoscopic-Vision Software; Estimating the Inertia Matrix of a Spacecraft; Spatial Coverage Planning for Exploration Robots; and Increasing the Life of a Xenon-Ion Spacecraft Thruster.

Source record↗

Sampling Technique for Robust Odorant Detection Based on MIT RealNose Data

This technique enhances the detection capability of the autonomous Real-Nose system from MIT to detect odorants and their concentrations in noisy and transient environments. The lowcost, portable system with low power consumption will operate at high speed and is suited for unmanned and remotely operated long-life applications. A deterministic mathematical model was developed to detect odorants and calculate their concentration in noisy environments. Real data from MIT's NanoNose was examined, from which a signal conditioning technique was proposed to enable robust odorant detection for the RealNose system. Its sensitivity can reach to sub-part-per-billion (sub-ppb). A Space Invariant Independent Component Analysis (SPICA) algorithm was developed to deal with non-linear mixing that is an over-complete case, and it is used as a preprocessing step to recover the original odorant sources for detection. This approach, combined with the Cascade Error Projection (CEP) Neural Network algorithm, was used to perform odorant identification. Signal conditioning is used to identify potential processing windows to enable robust detection for autonomous systems. So far, the software has been developed and evaluated with current data sets provided by the MIT team. However, continuous data streams are made available where even the occurrence of a new odorant is unannounced and needs to be noticed by the system autonomously before its unambiguous detection. The challenge for the software is to be able to separate the potential valid signal from the odorant and from the noisy transition region when the odorant is just introduced.

Duong, Tuan A.↗

Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex

Functional near infrared spectroscopy (fNIRS) is an emerging optical neuroimaging technology that indirectly measures neuronal activity in the cortex via neurovascular coupling. It quantifies hemoglobin concentration ([Hb]) and thus measures the same hemodynamic response as functional magnetic resonance imaging (fMRI), but is portable, non-confining, relatively inexpensive, and is appropriate for long-duration monitoring and use at the bedside. Like fMRI, it is noninvasive and safe for repeated measurements. Patterns of [Hb] changes are used to classify cognitive state. Thus, fNIRS technology offers much potential for application in operational contexts. For instance, the use of fNIRS to detect the mental state of commercial aircraft operators in near real time could allow intelligent flight decks of the future to optimally support human performance in the interest of safety by responding to hazardous mental states of the operator. However, many opportunities remain for improving robustness and reliability. It is desirable to reduce the impact of motion and poor optical coupling of probes to the skin. Such artifacts degrade signal quality and thus cognitive state classification accuracy. Field application calls for further development of algorithms and filters for the automation of bad channel detection and dynamic artifact removal. This work introduces a novel adaptive filter method for automated real-time fNIRS signal quality detection and improvement. The output signal (after filtering) will have had contributions from motion and poor coupling reduced or removed, thus leaving a signal more indicative of changes due to hemodynamic brain activations of interest. Cognitive state classifications based on these signals reflect brain activity more reliably. The filter has been tested successfully with both synthetic and real human subject data, and requires no auxiliary measurement. This method could be implemented as a real-time filtering option or bad channel rejection feature of software used with frequency domain fNIRS instruments for signal acquisition and processing. Use of this method could improve the reliability of any operational or real-world application of fNIRS in which motion is an inherent part of the functional task of interest. Other optical diagnostic techniques (e.g., for NIR medical diagnosis) also may benefit from the reduction of probe motion artifact during any use in which motion avoidance would be impractical or limit usability.

Harrivel, Angela↗

Functional Near-Infrared Spectroscopy Signals Measure Neuronal Activity in the Cortex

Functional near infrared spectroscopy (fNIRS) is an emerging optical neuroimaging technology that indirectly measures neuronal activity in the cortex via neurovascular coupling. It quantifies hemoglobin concentration ([Hb]) and thus measures the same hemodynamic response as functional magnetic resonance imaging (fMRI), but is portable, non-confining, relatively inexpensive, and is appropriate for long-duration monitoring and use at the bedside. Like fMRI, it is noninvasive and safe for repeated measurements. Patterns of [Hb] changes are used to classify cognitive state. Thus, fNIRS technology offers much potential for application in operational contexts. For instance, the use of fNIRS to detect the mental state of commercial aircraft operators in near real time could allow intelligent flight decks of the future to optimally support human performance in the interest of safety by responding to hazardous mental states of the operator. However, many opportunities remain for improving robustness and reliability. It is desirable to reduce the impact of motion and poor optical coupling of probes to the skin. Such artifacts degrade signal quality and thus cognitive state classification accuracy. Field application calls for further development of algorithms and filters for the automation of bad channel detection and dynamic artifact removal. This work introduces a novel adaptive filter method for automated real-time fNIRS signal quality detection and improvement. The output signal (after filtering) will have had contributions from motion and poor coupling reduced or removed, thus leaving a signal more indicative of changes due to hemodynamic brain activations of interest. Cognitive state classifications based on these signals reflect brain activity more reliably. The filter has been tested successfully with both synthetic and real human subject data, and requires no auxiliary measurement. This method could be implemented as a real-time filtering option or bad channel rejection feature of software used with frequency domain fNIRS instruments for signal acquisition and processing. Use of this method could improve the reliability of any operational or real-world application of fNIRS in which motion is an inherent part of the functional task of interest. Other optical diagnostic techniques (e.g., for NIR medical diagnosis) also may benefit from the reduction of probe motion artifact during any use in which motion avoidance would be impractical or limit usability.

Harrivel, Angela↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 kin or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed- shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

Automated clustering-based workload characterization

The demands placed on the mass storage systems at various federal agencies and national laboratories are continuously increasing in intensity. This forces system managers to constantly monitor the system, evaluate the demand placed on it, and tune it appropriately using either heuristics based on experience or analytic models. Performance models require an accurate workload characterization. This can be a laborious and time consuming process. It became evident from our experience that a tool is necessary to automate the workload characterization process. This paper presents the design and discusses the implementation of a tool for workload characterization of mass storage systems. The main features of the tool discussed here are: (1)Automatic support for peak-period determination. Histograms of system activity are generated and presented to the user for peak-period determination; (2) Automatic clustering analysis. The data collected from the mass storage system logs is clustered using clustering algorithms and tightness measures to limit the number of generated clusters; (3) Reporting of varied file statistics. The tool computes several statistics on file sizes such as average, standard deviation, minimum, maximum, frequency, as well as average transfer time. These statistics are given on a per cluster basis; (4) Portability. The tool can easily be used to characterize the workload in mass storage systems of different vendors. The user needs to specify through a simple log description language how the a specific log should be interpreted. The rest of this paper is organized as follows. Section two presents basic concepts in workload characterization as they apply to mass storage systems. Section three describes clustering algorithms and tightness measures. The following section presents the architecture of the tool. Section five presents some results of workload characterization using the tool.Finally, section six presents some concluding remarks.

Pentakalos, Odysseas I.↗

High-Efficiency High-Resolution Global Model Developments at the NASA Goddard Data Assimilation Office

The Data Assimilation Office (DAO) has been developing a new generation of ultra-high resolution General Circulation Model (GCM) that is suitable for 4-D data assimilation, numerical weather predictions, and climate simulations. These three applications have conflicting requirements. For 4-D data assimilation and weather predictions, it is highly desirable to run the model at the highest possible spatial resolution (e.g., 55 km or finer) so as to be able to resolve and predict socially and economically important weather phenomena such as tropical cyclones, hurricanes, and severe winter storms. For climate change applications, the model simulations need to be carried out for decades, if not centuries. To reduce uncertainty in climate change assessments, the next generation model would also need to be run at a fine enough spatial resolution that can at least marginally simulate the effects of intense tropical cyclones. Scientific problems (e.g., parameterization of subgrid scale moist processes) aside, all three areas of application require the model's computational performance to be dramatically improved as compared to the previous generation. In this talk, I will present the current and future developments of the "finite-volume dynamical core" at the Data Assimilation Office. This dynamical core applies modem monotonicity preserving algorithms and is genuinely conservative by construction, not by an ad hoc fixer. The "discretization" of the conservation laws is purely local, which is clearly advantageous for resolving sharp gradient flow features. In addition, the local nature of the finite-volume discretization also has a significant advantage on distributed memory parallel computers. Together with a unique vertically Lagrangian control volume discretization that essentially reduces the dimension of the computational problem from three to two, the finite-volume dynamical core is very efficient, particularly at high resolutions. I will also present the computational design of the dynamical core using a hybrid distributed-shared memory programming paradigm that is portable to virtually any of today's high-end parallel super-computing clusters.

Lin, Shian-Jiann↗

Space Suit Portable Life Support System (PLSS) 2.0 Pre-Installation Acceptance (PIA) Testing

Following successful completion of the space suit Portable Life Support System (PLSS) 1.0 development and testing in 2011, the second system-level prototype, PLSS 2.0, was developed in 2012 to continue the maturation of the advanced PLSS design which is intended to reduce consumables, improve reliability and robustness, and incorporate additional sensing and functional capabilities over the current Space Shuttle/International Space Station Extravehicular Mobility Unit (EMU) PLSS. PLSS 2.0 represents the first attempt at a packaged design comprising first generation or later component prototypes and medium fidelity interfaces within a flight-like representative volume. Pre-Installation Acceptance (PIA) is carryover terminology from the Space Shuttle Program referring to the series of test sequences used to verify functionality of the EMU PLSS prior to installation into the Space Shuttle airlock for launch. As applied to the PLSS 2.0 development and testing effort, PIA testing designated the series of 27 independent test sequences devised to verify component and subsystem functionality, perform in situ instrument calibrations, generate mapping data to define set-points for control algorithms, evaluate hardware performance against advanced PLSS design requirements, and provide quantitative and qualitative feedback on evolving design requirements and performance specifications. PLSS 2.0 PIA testing was carried out from 3/20/13 - 3/15/14 using a variety of test configurations to perform test sequences that ranged from stand-alone component testing to system-level testing, with evaluations becoming increasingly integrated as the test series progressed. Each of the 27 test sequences was vetted independently, with verification of basic functionality required before completion. Because PLSS 2.0 design requirements were evolving concurrently with PLSS 2.0 PIA testing, the requirements were used as guidelines to assess performance during the tests; after the completion of PIA testing, test data served to improve the fidelity and maturity of design requirements as well as plans for future advanced PLSS functional testing.

Watts, Carly↗

Space Suit Portable Life Support System (PLSS) 2.0 Pre-Installation Acceptance (PIA) Testing

Following successful completion of the space suit Portable Life Support System (PLSS) 1.0 development and testing in 2011, the second system-level prototype, PLSS 2.0, was developed in 2012 to continue the maturation of the advanced PLSS design. This advanced PLSS is intended to reduce consumables, improve reliability and robustness, and incorporate additional sensing and functional capabilities over the current Space Shuttle/International Space Station Extravehicular Mobility Unit (EMU) PLSS. PLSS 2.0 represents the first attempt at a packaged design comprising first generation or later component prototypes and medium fidelity interfaces within a flight-like representative volume. Pre-Installation Acceptance (PIA) is carryover terminology from the Space Shuttle Program referring to the series of test sequences used to verify functionality of the EMU PLSS prior to installation into the Space Shuttle airlock for launch. As applied to the PLSS 2.0 development and testing effort, PIA testing designated the series of 27 independent test sequences devised to verify component and subsystem functionality, perform in situ instrument calibrations, generate mapping data, define set-points, evaluate control algorithms, evaluate hardware performance against advanced PLSS design requirements, and provide quantitative and qualitative feedback on evolving design requirements and performance specifications. PLSS 2.0 PIA testing was carried out in 2013 and 2014 using a variety of test configurations to perform test sequences that ranged from stand-alone component testing to system-level testing, with evaluations becoming increasingly integrated as the test series progressed. Each of the 27 test sequences was vetted independently, with verification of basic functionality required before completion. Because PLSS 2.0 design requirements were evolving concurrently with PLSS 2.0 PIA testing, the requirements were used as guidelines to assess performance during the tests; after the completion of PIA testing, test data served to improve the fidelity and maturity of design requirements as well as plans for future advanced PLSS functional testing.

Anchondo, Ian↗