Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data dependencies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Efficacy of Code Optimization on Cache-based Processors

The current common wisdom in the U.S. is that the powerful, cost-effective supercomputers of tomorrow will be based on commodity (RISC) micro-processors with cache memories. Already, most distributed systems in the world use such hardware as building blocks. This shift away from vector supercomputers and towards cache-based systems has brought about a change in programming paradigm, even when ignoring issues of parallelism. Vector machines require inner-loop independence and regular, non-pathological memory strides (usually this means: non-power-of-two strides) to allow efficient vectorization of array operations. Cache-based systems require spatial and temporal locality of data, so that data once read from main memory and stored in high-speed cache memory is used optimally before being written back to main memory. This means that the most cache-friendly array operations are those that feature zero or unit stride, so that each unit of data read from main memory (a cache line) contains information for the next iteration in the loop. Moreover, loops ought to be 'fat', meaning that as many operations as possible are performed on cache data-provided instruction caches do not overflow and enough registers are available. If unit stride is not possible, for example because of some data dependency, then care must be taken to avoid pathological strides, just ads on vector computers. For cache-based systems the issues are more complex, due to the effects of associativity and of non-unit block (cache line) size. But there is more to the story. Most modern micro-processors are superscalar, which means that they can issue several (arithmetic) instructions per clock cycle, provided that there are enough independent instructions in the loop body. This is another argument for providing fat loop bodies. With these restrictions, it appears fairly straightforward to produce code that will run efficiently on any cache-based system. It can be argued that although some of the important computational algorithms employed at NASA Ames require different programming styles on vector machines and cache-based machines, respectively, neither architecture class appeared to be favored by particular algorithms in principle. Practice tells us that the situation is more complicated. This report presents observations and some analysis of performance tuning for cache-based systems. We point out several counterintuitive results that serve as a cautionary reminder that memory accesses are not the only factors that determine performance, and that within the class of cache-based systems, significant differences exist.

VanderWijngaart, Rob F.↗

Automatic Relative Debugging of OpenMP Programs

In this work we show how automatic relative debugging can be used to find differences in computation between a serial program and an OpenMP parallel version of that program. Backtracking and re-execution are used to determine the first OpenMP parallel region that produces a difference in computation that may lead to an incorrect value the user has indicated. Tool-parallelized programs are addressed by utilizing static analysis and directive information from the parallelization tool. Manually-parallelized programs are addressed as well by performing data dependence and directive analysis.

Matthews, Gregory↗

Solving Large Problems Quickly: Progress in 2001-2003

This document describes the progress we have made and the lessons we have learned in 2001 through 2003 under the NASA grant entitled "Solving Important Problems Faster". The long-term goal of this research is to accelerate large, irregular scientific applications which have enormous data sets and which are difficult to parallelize. To accomplish this goal, we are exploring two complementary techniques: (i) using compiler-inserted prefetching to automatically hide the I/O latency of accessing these large data sets from disk; and (ii) using thread-level data speculation to enable the optimistic parallelization of applications despite uncertainty as to whether data dependences exist between the resulting threads which would normally make them unsafe to execute in parallel. Overall, we made significant progress in 2001 through 2003, and the project has gone well.

Mowry, Todd C.↗

Ground processing of Cassini RADAR imagery of Titan

This paper describes the ground processing of Cassini SAR data. We focus upon the unusual features of the data and how these features impact the processing. We exhibit a data dependent mechanism we have implemented for eliminating artifacts due to attitude and ephemeris knowledge error. Finally we describe how we trade-off SAR performance vs. area of coverage when we design our spacecraft pointing profiles.

remote sensing↗

Array processor architecture

A high speed parallel array data processing architecture fashioned under a computational envelope approach includes a data base memory for secondary storage of programs and data, and a plurality of memory modules interconnected to a plurality of processing modules by a connection network of the Omega gender. Programs and data are fed from the data base memory to the plurality of memory modules and from hence the programs are fed through the connection network to the array of processors (one copy of each program for each processor). Execution of the programs occur with the processors operating normally quite independently of each other in a multiprocessing fashion. For data dependent operations and other suitable operations, all processors are instructed to finish one given task or program branch before all are instructed to proceed in parallel processing fashion on the next instruction. Even when functioning in the parallel processing mode however, the processors are not locked-step but execute their own copy of the program individually unless or until another overall processor array synchronization instruction is issued.

Barnes, George H.↗

Synthesis Methods, Microscopy Characterization and Device Integration of Nanoscale Metal Oxide Semiconductors for Gas Sensing in Aerospace Applications

A comparison is made between SnO2, ZnO, and TiO2 single-crystal nanowires and SnO2 polycrystalline nanofibers for gas sensing. Both nanostructures possess a one-dimensional morphology. Different synthesis methods are used to produce these materials: thermal evaporation-condensation (TEC), controlled oxidation, and electrospinning. Advantages and limitations of each technique are listed. Practical issues associated with harvesting, purification, and integration of these materials into sensing devices are detailed. For comparison to the nascent form, these sensing materials are surface coated with Pd and Pt nanoparticles. Gas sensing tests, with respect to H2, are conducted at ambient and elevated temperatures. Comparative normalized responses and time constants for the catalyst and noncatalyst systems provide a basis for identification of the superior metal-oxide nanostructure and catalyst combination. With temperature-dependent data, Arrhenius analyses are made to determine an activation energy for the catalyst-assisted systems.

VanderWal, Randy L.↗

PSR J0007+7303 in the CTA1 SNR: New Gamma-ray Results from Two Years of Fermi-LAT Observations

One of the main results of the Fermi Gamma-Ray Space Telescope is the discovery of gamma-ray selected pulsars. The high magnetic field pulsar, PSR J0007+7303 in CTA1, was the first ever to be discovered through its gamma-ray pulsations. Based on analysis of 2 years of LAT survey data, we report on the discovery of I-ray emission in the off-pulse phase interval at the approx. 6sigma level. The flux from this emission in the energy range E > or =::: 100 MeV is F(sub 100) = (1.73+/-0.40) x 10(exp -8) photons/sq cm/s and is best fitted by a power law with a photon index of Gamma = 2.54+/-0.14. The pulsed gamma-ray flux in the same energy range is F(sub 100) = (3.95+/-0.07) x 10(exp -7) photons/sq cm/s and is best fitted by an exponentially-cutoff power-law spectrum with a photon index of Gamma = 1.41+/-0.23 and a cutoff energy E(sub c) = 4.04+/-0.20 GeV. We find no flux variability neither at the 2009 May glitch nor in the long term behavior. We model the gamma-ray light curve with two high-altitude emission models, the outer gap and slot gap, and find that the model that best fits the data depends strongly on the assumed origin of the off-pulse emission. Both models favor a large angle between the magnetic axis and observer line of sight, consistent with the nondetection of radio emission being a geometrical effect. Finally we discuss how the LAT results bear on the understanding of the cooling of this neutron star.

Abdo, A. A.↗

Scheduling Operations for Massive Heterogeneous Clusters

High-performance computing (HPC) programming has become increasingly difficult with the advent of hybrid supercomputers consisting of multicore CPUs and accelerator boards such as the GPU. Manual tuning of software to achieve high performance on this type of machine has been performed by programmers. This is needlessly difficult and prone to being invalidated by new hardware, new software, or changes in the underlying code. A system was developed for task-based representation of programs, which when coupled with a scheduler and runtime system, allows for many benefits, including higher performance and utilization of computational resources, easier programming and porting, and adaptations of code during runtime. The system consists of a method of representing computer algorithms as a series of data-dependent tasks. The series forms a graph, which can be scheduled for execution on many nodes of a supercomputer efficiently by a computer algorithm. The schedule is executed by a dispatch component, which is tailored to understand all of the hardware types that may be available within the system. The scheduler is informed by a cluster mapping tool, which generates a topology of available resources and their strengths and communication costs. Software is decoupled from its hardware, which aids in porting to future architectures. A computer algorithm schedules all operations, which for systems of high complexity (i.e., most NASA codes), cannot be performed optimally by a human. The system aids in reducing repetitive code, such as communication code, and aids in the reduction of redundant code across projects. It adds new features to code automatically, such as recovering from a lost node or the ability to modify the code while running. In this project, the innovators at the time of this reporting intend to develop two distinct technologies that build upon each other and both of which serve as building blocks for more efficient HPC usage. First is the scheduling and dynamic execution framework, and the second is scalable linear algebra libraries that are built directly on the former.

Humphrey, John↗

Regression Verification Using Impact Summaries

Regression verification techniques are used to prove equivalence of syntactically similar programs. Checking equivalence of large programs, however, can be computationally expensive. Existing regression verification techniques rely on abstraction and decomposition techniques to reduce the computational effort of checking equivalence of the entire program. These techniques are sound but not complete. In this work, we propose a novel approach to improve scalability of regression verification by classifying the program behaviors generated during symbolic execution as either impacted or unimpacted. Our technique uses a combination of static analysis and symbolic execution to generate summaries of impacted program behaviors. The impact summaries are then checked for equivalence using an o-the-shelf decision procedure. We prove that our approach is both sound and complete for sequential programs, with respect to the depth bound of symbolic execution. Our evaluation on a set of sequential C artifacts shows that reducing the size of the summaries can help reduce the cost of software equivalence checking. Various reduction, abstraction, and compositional techniques have been developed to help scale software verification techniques to industrial-sized systems. Although such techniques have greatly increased the size and complexity of systems that can be checked, analysis of large software systems remains costly. Regression analysis techniques, e.g., regression testing [16], regression model checking [22], and regression verification [19], restrict the scope of the analysis by leveraging the differences between program versions. These techniques are based on the idea that if code is checked early in development, then subsequent versions can be checked against a prior (checked) version, leveraging the results of the previous analysis to reduce analysis cost of the current version. Regression verification addresses the problem of proving equivalence of closely related program versions [19]. These techniques compare two programs with a large degree of syntactic similarity to prove that portions of one program version are equivalent to the other. Regression verification can be used for guaranteeing backward compatibility, and for showing behavioral equivalence in programs with syntactic differences, e.g., when a program is refactored to improve its performance, maintainability, or readability. Existing regression verification techniques leverage similarities between program versions by using abstraction and decomposition techniques to improve scalability of the analysis [10, 12, 19]. The abstractions and decomposition in the these techniques, e.g., summaries of unchanged code [12] or semantically equivalent methods [19], compute an over-approximation of the program behaviors. The equivalence checking results of these techniques are sound but not complete-they may characterize programs as not functionally equivalent when, in fact, they are equivalent. In this work we describe a novel approach that leverages the impact of the differences between two programs for scaling regression verification. We partition program behaviors of each version into (a) behaviors impacted by the changes and (b) behaviors not impacted (unimpacted) by the changes. Only the impacted program behaviors are used during equivalence checking. We then prove that checking equivalence of the impacted program behaviors is equivalent to checking equivalence of all program behaviors for a given depth bound. In this work we use symbolic execution to generate the program behaviors and leverage control- and data-dependence information to facilitate the partitioning of program behaviors. The impacted program behaviors are termed as impact summaries. The dependence analyses that facilitate the generation of the impact summaries, we believe, could be used in conjunction with other abstraction and decomposition based approaches, [10, 12], as a complementary reduction technique. An evaluation of our regression verification technique shows that our approach is capable of leveraging similarities between program versions to reduce the size of the queries and the time required to check for logical equivalence. The main contributions of this work are: - A regression verification technique to generate impact summaries that can be checked for functional equivalence using an off-the-shelf decision procedure. - A proof that our approach is sound and complete with respect to the depth bound of symbolic execution. - An implementation of our technique using the LLVMcompiler infrastructure, the klee Symbolic Virtual Machine [4], and a variety of Satisfiability Modulo Theory (SMT) solvers, e.g., STP [7] and Z3 [6]. - An empirical evaluation on a set of C artifacts which shows that the use of impact summaries can reduce the cost of regression verification.

Backes, John↗

Detecting and Characterizing Semantic Inconsistencies in Ported Code

Adding similar features and bug fixes often requires porting program patches from reference implementations and adapting them to target implementations. Porting errors may result from faulty adaptations or inconsistent updates. This paper investigates (I) the types of porting errors found in practice, and (2) how to detect and characterize potential porting errors. Analyzing version histories, we define five categories of porting errors, including incorrect control- and data-flow, code redundancy, inconsistent identifier renamings, etc. Leveraging this categorization, we design a static control- and data-dependence analysis technique, SPA, to detect and characterize porting inconsistencies. Our evaluation on code from four open-source projects shows thai SPA can dell-oct porting inconsistencies with 65% to 73% precision and 90% recall, and identify inconsistency types with 58% to 63% precision and 92% to 100% recall. In a comparison with two existing error detection tools, SPA improves precision by 14 to 17 percentage points

Ray, Baishakhi↗

Detecting and Characterizing Semantic Inconsistencies in Ported Code

Adding similar features and bug fixes often requires porting program patches from reference implementations and adapting them to target implementations. Porting errors may result from faulty adaptations or inconsistent updates. This paper investigates (1) the types of porting errors found in practice, and (2) how to detect and characterize potential porting errors. Analyzing version histories, we define five categories of porting errors, including incorrect control- and data-flow, code redundancy, inconsistent identifier renamings, etc. Leveraging this categorization, we design a static control- and data-dependence analysis technique, SPA, to detect and characterize porting inconsistencies. Our evaluation on code from four open-source projects shows that SPA can detect porting inconsistencies with 65% to 73% precision and 90% recall, and identify inconsistency types with 58% to 63% precision and 92% to 100% recall. In a comparison with two existing error detection tools, SPA improves precision by 14 to 17 percentage points.

Semantic Errors↗

Insulation Resistance Degradation in Ni-BaTiO3 Multilayer Ceramic Capacitors

Insulation resistance (IR) degradation in NiBaTiO3 multilayer ceramic capacitors has been characterized by the measurement of both time to failure (TTF) and direct current leakage current as a function of stress time under highly accelerated life test conditions. The measured leakage current time dependence data fit well to an exponential form, and a characteristic growth time tau (sub SD) can be determined. A greater value of tau (sub SD) represents a slower IR degradation process. Oxygen vacancy migration and localization at the grain boundary region results in the reduction of the Schottky barrier height and has been found to be the main reason for IR degradation in NiBaTiO3 capacitors. The reduction of barrier height as a function oftime follows an exponential relation of phi (t ) = phi (0) e (exp -2Kt), where 13 the degradation rate constant K Koe (Ek/kT) is inversely proportional to the mean TTF (MTTF) and can be determined using an Arrhenius plot. For oxygen vacancy electromigration, a lower barrier height phi (0) will favor a slow IR degradation process, but a lower phi (0) will also promote electronic carrier conduction across the barrier and decrease the IR. As a result, a moderate barrier height phi (0) (and therefore a moderate IR value) with a longer MTTF (smaller degradation rate constant K) will result in a minimized IR degradation process and the most improved reliability in NiBaTiO3 multilayer ceramic capacitors.

reliability↗

A Monte Carlo Approach to Modeling the Breakup of the Space Launch System EM-1 Core Stage with an Integrated Blast and Fragment Catalogue

The Liquid Propellant Fragment Overpressure Acceleration Model (L-FOAM) is a tool developed by Bangham Engineering Incorporated (BEi) that produces a representative debris cloud from an exploding liquid-propellant launch vehicle. Here it is applied to the Core Stage (CS) of the National Aeronautics and Space Administration (NASA) Space Launch System (SLS launch vehicle). A combination of Probability Density Functions (PDF) based on empirical data from rocket accidents and applicable tests, as well as SLS specific geometry are combined in a MATLAB script to create unique fragment catalogues each time L-FOAM is run-tailored for a Monte Carlo approach for risk analysis. By accelerating the debris catalogue with the BEi blast model for liquid hydrogen / liquid oxygen explosions, the result is a fully integrated code that models the destruction of the CS at a given point in its trajectory and generates hundreds of individual fragment catalogues with initial imparted velocities. The BEi blast model provides the blast size (radius) and strength (overpressure) as probabilities based on empirical data and anchored with analytical work. The coupling of the L-FOAM catalogue with the BEi blast model is validated with a simulation of the Project PYRO S-IV destruct test. When running a Monte Carlo simulation, L-FOAM can accelerate all catalogues with the same blast (mean blast, 2 σ blast, etc.), or vary the blast size and strength based on their respective probabilities. L-FOAM then propagates these fragments until impact with the earth. Results from L-FOAM include a description of each fragment (dimensions, weight, ballistic coefficient, type and initial location on the rocket), imparted velocity from the blast, and impact data depending on user desired application. LFOAM application is for both near-field (fragment impact to escaping crew capsule) and far-field (fragment ground impact footprint) safety considerations. The user is thus able to use statistics from a Monte Carlo set of L-FOAM catalogues to quantify risk for a multitude of potential CS destruct scenarios. Examples include the effect of warning time on the survivability of an escaping crew capsule or the maximum fragment velocities generated by the ignition of leaking propellants in internal cavities.

Richardson, Erin↗

Insulation Resistance Degradation in Ni-BaTiO3 Multilayer Ceramic Capacitors

Insulation resistance (IR) degradation in Ni-BaTiO3 multilayer ceramic capacitors has been characterized by the measurement of both time to failure and direct-current (DC) leakage current as a function of stress time under highly accelerated life test conditions. The measured leakage current-time dependence data fit well to an exponential form, and a characteristic growth time SD can be determined. A greater value of tau(sub SD) represents a slower IR degradation process. Oxygen vacancy migration and localization at the grain boundary region results in the reduction of the Schottky barrier height and has been found to be the main reason for IR degradation in Ni-BaTiO3 capacitors. The reduction of barrier height as a function of time follows an exponential relation of phi (𝑡)=phi (0)e(exp -2Κt), where the degradation rate constant 𝐾=𝐾o𝑒(𝐸𝑘/𝑘𝑇) is inversely proportional to the mean time to failure (MTTF) and can be determined using an Arrhenius plot. For oxygen vacancy electromigration, a lower barrier height phi(0) will favor a slow IR degradation process, but a lower phi(0) will also promote electronic carrier conduction across the barrier and decrease the insulation resistance. As a result, a moderate barrier height phi(0) (and therefore a moderate IR value) with a longer MTTF (smaller degradation rate constant 𝐾) will result in a minimized IR degradation process and the most improved reliability in Ni-BaTiO3 multilayer ceramic capacitors.

dielectric degradation↗

Infrared Transmission Spectroscopy of the Exoplanets HD 209458b and XO-1b Using the Wide Field Camera-3 on the Hubble Space Telescope

Exoplanetary transmission spectroscopy in the near-infrared using the Hubble Space Telescope (HST) NICMOS (Near Infrared Camera and Multi-Object Spectrometer) is currently ambiguous because different observational groups claim different results from the same data, depending on their analysis methodologies. Spatial scanning with HST/WFC3 (Wide Field Camera-3) provides an opportunity to resolve this ambiguity.We here report WFC3 spectroscopy of the giant planets HD 209458b and XO-1b in transit, using spatial scanning mode for maximum photon-collecting efficiency. We introduce an analysis technique that derives the exoplanetary transmission spectrum without the necessity of explicitly decorrelating instrumental effects, and achieves nearly photon-limited precision even at the high flux levels collected in spatial scan mode. Our errors are within 6 percent (XO-1) and 26 percent (HD 209458b) of the photon-limit at a resolving power of lambda divided by delta times lambda approximating 70, and are better than 0.01 percent per spectral channel. Both planets exhibit water absorption of approximately 200 ppm at the water peak near 1.38 m. Our result for XO-1b contradicts the much larger absorption derived from NICMOS spectroscopy. The weak water absorption we measure for HD209458b is reminiscent of the weakness of sodium absorption in the first transmission spectroscopy of an exoplanet atmosphere by Charbonneau et al. Model atmospheres having uniformly distributed extra opacity of 0.012 square centimeters per gram account approximately for both our water measurement and the sodium absorption. Our results for HD 209458b support the picture advocated by Pont et al. in which weak molecular absorptions are superposed on a transmission spectrum that is dominated by continuous opacity due to haze and or dust. However,the extra opacity needed for HD 209458b is grayer than for HD 189733b, with a weaker Rayleigh component.

planetary systems↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

LSAFE: a Lightweight Static Analysis Framework for binary Executables

Static analysis is a widely used technique for analyzing various aspects of programs. However, as programs become more complex, static analysis tools require larger resources, such as CPU time and memory, to perform the same tasks. Moreover, the source code of programs may not always be accessible, requiring static analysis to be performed on the binary executable code directly. To overcome these challenges, we propose a lightweight static analysis framework called LSAFE, which constructs control flow graphs (CFGs) and data dependency graphs (DDGs) of target programs with optimized performance in terms of CPU and memory usage. We evaluated the proposed framework using both Spec benchmark programs and real-world industrial applications, and found that it outperformed Angr, an existing state-of-the-art static analysis tool. Additionally, we demonstrate a case study that utilizes the CFG generated by LSAFE to detect memory leaks.

Qu, Guangzhi↗