Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

STGT program: Ada coding and architecture lessons learned

STGT (Second TDRSS Ground Terminal) is currently halfway through the System Integration Test phase (Level 4 Testing). To date, many software architecture and Ada language issues have been encountered and solved. This paper, which is the transcript of a presentation at the 3 Dec. meeting, attempts to define these lessons plus others learned regarding software project management and risk management issues, training, performance, reuse, and reliability. Observations are included regarding the use of particular Ada coding constructs, software architecture trade-offs during the prototyping, development and testing stages of the project, and dangers inherent in parallel or concurrent systems, software, hardware, and operations engineering.

Usavage, Paul↗

Understanding Lustre Internals. Second Edition

The Lustre file system has become a preferred storage resource for systems on the Top500 list, and it is often the file system of choice for small- to medium-sized HPC systems that require parallel shared access to data. Several resources exist to help users deploy and configure Lustre, but the same cannot be said for resources that explain the inner workings of the Lustre source code. A previous ORNL technical report entitled "Understanding Lustre Filesystem Internals" (ORNL/TM-2009/117) provided an excellent summary of Lustre subsystem operations. However, that report is over a decade old and is based on Lustre version 1.6. Since that report was published, Lustre has evolved significantly. Several subsystems underwent significant code changes and many new features have been added to the file system, bringing the current Lustre version up to 2.15.This report aims to document and explain the internal workings of the latest version of the Lustre file system. It will provide more complete and up-to-date information than the previous technical report and should serve as a foundational document for anyone interested in Lustre software development. Key data structures will be described along with the APIs used for interaction among the various Lustre subsystems. Although the Lustre software is constantly being developed, the details in this document should remain relevant for the forseeable future.

97 MATHEMATICS AND COMPUTING↗

High Performance Parallel Computational Nanotechnology

At a recent press conference, NASA Administrator Dan Goldin encouraged NASA Ames Research Center to take a lead role in promoting research and development of advanced, high-performance computer technology, including nanotechnology. Manufacturers of leading-edge microprocessors currently perform large-scale simulations in the design and verification of semiconductor devices and microprocessors. Recently, the need for this intensive simulation and modeling analysis has greatly increased, due in part to the ever-increasing complexity of these devices, as well as the lessons of experiences such as the Pentium fiasco. Simulation, modeling, testing, and validation will be even more important for designing molecular computers because of the complex specification of millions of atoms, thousands of assembly steps, as well as the simulation and modeling needed to ensure reliable, robust and efficient fabrication of the molecular devices. The software for this capacity does not exist today, but it can be extrapolated from the software currently used in molecular modeling for other applications: semi-empirical methods, ab initio methods, self-consistent field methods, Hartree-Fock methods, molecular mechanics; and simulation methods for diamondoid structures. In as much as it seems clear that the application of such methods in nanotechnology will require powerful, highly powerful systems, this talk will discuss techniques and issues for performing these types of computations on parallel systems. We will describe system design issues (memory, I/O, mass storage, operating system requirements, special user interface issues, interconnects, bandwidths, and programming languages) involved in parallel methods for scalable classical, semiclassical, quantum, molecular mechanics, and continuum models; molecular nanotechnology computer-aided designs (NanoCAD) techniques; visualization using virtual reality techniques of structural models and assembly sequences; software required to control mini robotic manipulators for positional control; scalable numerical algorithms for reliability, verifications and testability. There appears no fundamental obstacle to simulating molecular compilers and molecular computers on high performance parallel computers, just as the Boeing 777 was simulated on a computer before manufacturing it.

Saini, Subhash↗

Note on unit tangent vector computation for homotopy curve tracking on a hypercube

Probability-one homotopy methods are a class of methods for solving nonlinear systems of equations that are globally convergent from an arbitrary starting point. The essence of all such algorithms is the construction of an appropriate homotopy map and subsequent tracking of some smooth curve in the zero set of the homotopy map. Tracking a homotopy curve involves finding the unit tangent vector at different points along the zero curve, which amounts to calculating the kernel of the n x (n + 1) Jacobian matrix. While computing the tangent vector is just one part of the curve tracking algorithm, it can require a significant percentage of the total tracking time. This note presents computational results showing the performance of several different parallel orthogonal factorization/triangular system solving algorithms for the tangent vector computation on a hypercube.

Chakraborty, A.↗

Capillary Pumped Loop 3 Flight Experiment Overview

The Capillary Pumped Loop 3 (CAPL 3) Experiment is a follow on to the CAPL 1 and CAPL 2 experiments which flew on STS-60 (2/94) and STS-69 (9/95), respectively. CAPL 3 is tentatively scheduled to fly on the Space Shuttle in late 2000 as part of the Hitchhiker Experiments Advancing Technology (HEAT) payload. The experiment is a joint Naval Research Laboratory (NRL)/Goddard Space Flight Center (GSFC) payload which will meet technology objectives for both the Department of Defense and NASA. The primary objective of CAPL 3 is to demonstrate in space a multiple evaporator capillary pumped loop system, capable of reliable start-up, reliable continuous operation, and at least 50% heat load sharing with hardware for a deployable radiator. CAPL 3 is a full scale CPL system with four parallel capillary evaporators. The loop also contains a capillary starter pump, 8 parallel direct condensation condensers with associated flow regulators, a back pressure regulator, a two-phase reservoir, and various headers and transport tubing. A variable conductance heat pipe is located between one of the evaporators and the experiment radiator to provide a cooling source for the demonstration of heat load sharing. The experiment has an operating power range of 100 W to approximately 1400 W. The experiment ammonia charge will cause it to transition to a fixed conductance mode of operation if the radiator usage reaches 85%. Ambient functional tests have been performed on the experiment. Tests performed included start-up, low power, power cycles, high power, heat load sharing, variable/fixed conductance transition, saturation temperature changes, and pressure primes while the system was operating. The majority of the testing was performed at an ammonia saturation temperature of 30C, but a few tests were done at temperatures above and below this. The testing was highly successful. Details of the tests performed and a discussion of the results will be given in the presentation.

Ottenstein, Laura↗

New Horizons for High-Performance Computing

Here we provide an overview of the past, present, and a diverse collection of future computer architecture alternatives for HPC. The end of Moore’s Law influenced the current HPC architecture focus on accelerated compute nodes composed of CPU and GPU computing components integrated into massively parallel processor architecture systems. There are many alternatives for future HPC directions, with different technologies, computing ecosystems, opportunities for lead user application-driven customization, and the role of open innovation business models. This paper provides an overview of these different new horizons for HPC, an organizing principle to focus future computing research, different public-private partnership models, and the critical role of workforce development.

97 MATHEMATICS AND COMPUTING↗

Cross-Cutting Flight Infrastructure Improvements on M2020

Mars2020 (M2020) was formulated as a mission that leveraged as much Mars Science Laboratory (MSL) heritage as possible, while focusing major new development efforts on the original and unique elements needed to accomplish the different mission objectives. Well publicized examples of high profile new developments include precision landing, the sampling and caching system, the specific instrument suite, improved mobility via Autonomous Navigation, and later the addition of the Ingenuity helicopter. Less well known are the refinements to the core flight infrastructure, primarily in the cross-cutting functions of Telecom, Avionics, Data Management, Communications Behaviors, and Parameter Management. These enhancements are introduced predominately via flight software, and represent increases in capability that justified their inclusion in an otherwise heritage-focused project environment.Perseverance’s cross-cutting flight infrastructure improvements fall into and across the following five categories. First is a trimming of the software footprint of infrastructure modules, in order to make room for memory demands elsewhere in the system. Second is the minimization of data volume to be downlinked, through various methods such as the incorporation of new compression options. Third is the maximization of the available downlink bandwidth for data, by curtailing content-less data (fill) and introducing an improved UHF proximity link protocol. Fourth is a reduction in vulnerabilities, through increased file system redundancy, robustness, and software process monitoring. Fifth is an increase in operations efficiency by lowering file system mount times, improving parallelism between simultaneous events, minimizing the time to recover from file system errors, streamlining the purging of obsolete data, and reducing the number of commands to service parameters by a factor of 100.Individually, none of the cross-cutting infrastructure improvements are likely to garner headlines, but collectively they appreciably improve the safety and operability of Perseverance over its predecessor. This paper will describe the improvements, their promise, and where applicable, their actual impact in operations.

Bohannon, Emily↗

System and method of storing and analyzing information

A system and method of storing and analyzing information is disclosed. The system includes a compiler layer to convert user queries to data parallel executable code. The system further includes a library of multithreaded algorithms, processes, and data structures. The system also includes a multithreaded runtime library for implementing compiled code at runtime. The executable code is dynamically loaded on computing elements and contains calls to the library of multithreaded algorithms, processes, and data structures and the multithreaded runtime library.

Feo, John T.↗

Neptunium mononitride as a target material for Pu-238 production

Deep space exploration requires specialized sources for both thermal and power applications. Radioactive decay heat of plutonium-238 (238Pu) provides these sources in the form of radioisotope thermoelectric generators (RTGs). The 238 Pu is produced via neutron capture reaction involving neptunium-237 ( 237 Np) target material. Continual optimization of 237 Np target materials and evaluation of potential alternative targets for production of 238 Pu RTGs are advantageous for meeting ongoing space power system resource requirements. Current production of 238 Pu for RTGs for the United States space program utilizes neptunium dioxide ( 237 NpO 2 ) targets; however, the use of neptunium mononitride ( 237 NpN) presents an opportunity to increase the mass of 237 Np per target compared to the dioxide form, as well as increase the thermal conductivity of the target. To assess the viability of a 237 NpN target material, the material chemistry must be thoroughly evaluated, including synthesis methods and dissolution and reprocessing schemes. This review presents a summary of synthesis pathways for 237 NpN based on published literature on actinide mononitrides. Specific literature on 237 NpN is limited, necessitating evaluation of other actinide systems to gather parallels. This suggests a need for additional experimental studies on 237 NpN. A particular limitation in the existing literature is a lack of information on the differences in material characteristics, such as morphology, particle size, and trace chemical impurities, as a function of synthesis method. These parameters may affect subsequent reactor performance or dissolution of irradiated targets. The evaluation of existing literature is presented with a focus on the efficacy of 237 NpN targets for 238 Pu production.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Performance and Feature Improvements in Parareal-based Power System Dynamic Simulation

In recent years, a novel Parareal-based approach has been developed for fast transient simulations of large power system interconnections. Parareal belongs to the class of Parallel-in-time algorithms for solution of systems of differential-algebraic equations in parallel over an interval of time. The selection of a reasonably fast and accurate coarse solution is crucial to improve the performance of Parareal algorithm. Semi-analytical solution methods are one promising approach to achieve this goal. They have been investigated, and some preliminary results are presented here. In addition, Parareal-based simulator has been expanded to enable co-simulation with OpenDSS, a widely used open-source distribution system simulator. Preserving the parallel nature of the Parareal approach and taking advantage of the parallel capabilities of the latest versions of OpenDSS, each distribution system can be solved in their entirety on different processors in parallel within the main Parareal simulator. This paper also presents the structure of the transmission and distribution co-simulation and some results with different dynamic models of inverter-based resources in the distribution systems.

Park, Byungkwon↗

Electrical Capacitance Volume Tomography with High-Contrast Dielectrics

The Electrical Capacitance Volume Tomography (ECVT) system has been designed to complement the tools created to sense the presence of water in nonconductive spacecraft materials, by helping to not only find the approximate location of moisture but also its quantity and depth. The ECVT system has been created for use with a new image reconstruction algorithm capable of imaging high-contrast dielectric distributions. Rather than relying solely on mutual capacitance readings as is done in traditional electrical capacitance tomography applications, this method reconstructs high-resolution images using only the self-capacitance measurements. The image reconstruction method assumes that the material under inspection consists of a binary dielectric distribution, with either a high relative dielectric value representing the water or a low dielectric value for the background material. By constraining the unknown dielectric material to one of two values, the inverse math problem that must be solved to generate the image is no longer ill-determined. The image resolution becomes limited only by the accuracy and resolution of the measurement circuitry. Images were reconstructed using this method with both synthetic and real data acquired using an aluminum structure inserted at different positions within the sensing region. The cuboid geometry of the system has two parallel planes of 16 conductors arranged in a 4 4 pattern. The electrode geometry consists of parallel planes of copper conductors, connected through custom-built switch electronics, to a commercially available capacitance to digital converter. The figure shows two 4 4 arrays of electrodes milled from square sections of copper-clad circuit-board material and mounted on two pieces of glass-filled plastic backing, which were cut to approximately square shapes, 10 cm on a side. Each electrode is placed on 2.0-cm centers. The parallel arrays were mounted with the electrode arrays approximately 3 cm apart. The open ends were surrounded by a metal guard to reduce the sensitivity of the electrodes to outside interference and to help maintain the spacing between the arrays. Other uses for this innovation potentially include quantifying the amount of commodity remaining in the fuel and oxidizer tanks while on-orbit without having to fire spacecraft engines. Another orbit application is moisture sensing in plant-growth experiments because microgravity causes moisture in soil to distribute itself in unusual ways. At the moment, the hardware and image reconstruction technique may only be of interest to people involved in nondestructive evaluation. The reconstructed image takes almost a full week to reproduce with existing computer power. However, because computer power and speeds follows Moore s Law, execution times are likely to become acceptable within the next five to eight years. The code was written in Mathematica for dedicated use with the ECVT system. In its present form, it is not suitable to be used directly as a consumer product. However, the code could be likely improved by rewriting it in a compiled language such as C or Fortran.

Nurge, Mark↗

Temporal Precedence Checking for Switched Models and its Application to a Parallel Landing Protocol

This paper presents an algorithm for checking temporal precedence properties of nonlinear switched systems. This class of properties subsume bounded safety and capture requirements about visiting a sequence of predicates within given time intervals. The algorithm handles nonlinear predicates that arise from dynamics-based predictions used in alerting protocols for state-of-the-art transportation systems. It is sound and complete for nonlinear switch systems that robustly satisfy the given property. The algorithm is implemented in the Compare Execute Check Engine (C2E2) using validated simulations. As a case study, a simplified model of an alerting system for closely spaced parallel runways is considered. The proposed approach is applied to this model to check safety properties of the alerting logic for different operating conditions such as initial velocities, bank angles, aircraft longitudinal separation, and runway separation.

Duggirala, Parasara Sridhar↗

Data flow modeling techniques

There have been a number of simulation packages developed for the purpose of designing, testing and validating computer systems, digital systems and software systems. Complex analytical tools based on Markov and semi-Markov processes have been designed to estimate the reliability and performance of simulated systems. Petri nets have received wide acceptance for modeling complex and highly parallel computers. In this research data flow models for computer systems are investigated. Data flow models can be used to simulate both software and hardware in a uniform manner. Data flow simulation techniques provide the computer systems designer with a CAD environment which enables highly parallel complex systems to be defined, evaluated at all levels and finally implemented in either hardware or software. Inherent in data flow concept is the hierarchical handling of complex systems. In this paper we will describe how data flow can be used to model computer system.

Kavi, K. M.↗

Partitioning problems in parallel, pipelined, and distributed computing

The problem of optimally assigning the modules of a parallel program over the processors of a multiple-computer system is addressed. A sum-bottleneck path algorithm is developed that permits the efficient solution of many variants of this problem under some constraints on the structure of the partitions. In particular, the following problems are solved optimally for a single-host, multiple-satellite system: partitioning multiple chain-structured parallel programs, multiple arbitrarily structured serial programs, and single-tree structured parallel programs. In addition, the problem of partitioning chain-structured parallel programs across chain-connected systems is solved under certain constraints. All solutions for parallel programs are equally applicable to pipelined programs. These results extend prior research in this area by explicitly taking concurrency into account and permit the efficient utilization of multiple-computer architectures for a wide range of problems of practical interest.

Bokhari, Shahid H.↗

Efficient Use of Distributed Systems for Scientific Applications

Distributed computing has been regarded as the future of high performance computing. Nationwide high speed networks such as vBNS are becoming widely available to interconnect high-speed computers, virtual environments, scientific instruments and large data sets. One of the major issues to be addressed with distributed systems is the development of computational tools that facilitate the efficient execution of parallel applications on such systems. These tools must exploit the heterogeneous resources (networks and compute nodes) in distributed systems. This paper presents a tool, called PART, which addresses this issue for mesh partitioning. PART takes advantage of the following heterogeneous system features: (1) processor speed; (2) number of processors; (3) local network performance; and (4) wide area network performance. Further, different finite element applications under consideration may have different computational complexities, different communication patterns, and different element types, which also must be taken into consideration when partitioning. PART uses parallel simulated annealing to partition the domain, taking into consideration network and processor heterogeneity. The results of using PART for an explicit finite element application executing on two IBM SPs (located at Argonne National Laboratory and the San Diego Supercomputer Center) indicate an increase in efficiency by up to 36% as compared to METIS, a widely used mesh partitioning tool. The input to METIS was modified to take into consideration heterogeneous processor performance; METIS does not take into consideration heterogeneous networks. The execution times for these applications were reduced by up to 30% as compared to METIS. These results are given in Figure 1 for four irregular meshes with number of elements ranging from 30,269 elements for the Barth5 mesh to 11,451 elements for the Barth4 mesh. Future work with PART entails using the tool with an integrated application requiring distributed systems. In particular this application, illustrated in the document entails an integration of finite element and fluid dynamic simulations to address the cooling of turbine blades of a gas turbine engine design. It is not uncommon to encounter high-temperature, film-cooled turbine airfoils with 1,000,000s of degrees of freedom. This results because of the complexity of the various components of the airfoils, requiring fine-grain meshing for accuracy. Additional information is contained in the original.

Taylor, Valerie↗

Grafted nickel-promoter catalysts for dry reforming of methane identified through high-throughput experimentation

High-throughput synthesis of a series of monometallic and bimetallic catalysts (45 bimetallic and 50 monometallic samples) consisting of nickel and one of nine different metal promoters (B, Co, Cu, Fe, Mg, Mn, Sn, V and Zn) supported on one of six different metal oxides alumina, ceria, magnesia, silica and titania) is carried out via organometallic grafting using a robotic platform. The catalysts are evaluated for their activity and selectivity for the dry reforming of methane at a feed ratio of CH 4 :CO 2 of 1 at 650–800 °C in a parallel flow reactor system. The type of oxide support prevails over the type of additive for both catalyst activity and stability. On Al 2 O 3 and MgO, Fe was found to be the best promoter; on SiO 2 , Cu is the best promoter at 700 °C and higher, while on TiO 2 , Mn is found to enhance the conversion at 800 °C. On CeO 2 , all additives except Fe have beneficial effects. Twenty-five catalysts show > 90% methane conversion with ten catalysts showing > 95% conversion at 800 °C with the H 2 :CO ratios ranging from 0.8 to 1.2. Amongst the ten highest performers, NiFe/Al 2 O 3 and NiFe/MgO are more active than Ni/Al 2 O 3 and Ni/MgO, respectively and were stable over a period of 25 h at 800 °C. Characterization on the as-prepared samples reveals highly dispersed phase, while after reduction in H 2 , highly dispersed and reduced nickel particles up to 10 nm are formed. The particles do not increase in size under dry reforming reaction conditions at 800 °C. An increased hydrogen consumption observed during H 2 -TPR of the nickel particles is positively correlated with methane conversion for Al 2 O 3 -based catalysts. The resistance to deactivation by coking and variation in coke structure are investigated by spectroscopic and microscopic methods to identify the relationship between metal promoters, alloy formation, and type of surface carbon deposits. Carbon whiskers were observed on the ten selected spent samples and are preferentially deposited on Ni rather than on the promoters. Carbon nanotube formation and metal particle removal from support were not observed to cause deactivation while amorphous carbon formation was clearly linked to catalyst deactivation, as amorphous carbon could encapsulate nickel, either on the support or at the end of the carbon nanotube. Furthermore, the organometallic grafting technique is an efficient and suitable technique for synthesizing highly dispersed and homogeneous phases which lead to high conversion and high durability for dry reforming of methane.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electromechanical memcapacitor model offering biologically plausible spiking.

In this article, we introduce a new nanoscale electromechanical device - a leaky memcapacitor - and show that it may be useful for the hardware implementation of spiking neurons. The leaky memcapacitor is a movableplate capacitor that becomes quite conductive when the plates come close to each other. The equivalent circuit of the leaky memcapacitor involves a memcapacitive and memristive system connected in parallel. In the leaky memcapacitor, resistance and capacitance depend on the same internal state variable, which is the displacement of the movable plate. We have performed a comprehensive analysis showing that several types of spiking observed in biological neurons can be implemented with the leaky memcapacitor. Significant attention is paid to the dynamic properties of the model. As in leaky memcapacitors the capacitive, leaking resistive, and reset functionalities are implemented naturally within the same device structure, their use will simplify the creation of spiking neural networks.

Zhang, Zixi↗