Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Arranging computer architectures to create higher-performance controllers

Techniques for integrating microprocessors, array processors, and other intelligent devices in control systems are reviewed, with an emphasis on the (re)arrangement of components to form distributed or parallel processing systems. Consideration is given to the selection of the host microprocessor, increasing the power and/or memory capacity of the host, multitasking software for the host, array processors to reduce computation time, the allocation of real-time and non-real-time events to different computer subsystems, intelligent devices to share the computational burden for real-time events, and intelligent interfaces to increase communication speeds. The case of a helicopter vibration-suppression and stabilization controller is analyzed as an example, and significant improvements in computation and throughput rates are demonstrated.

Jacklin, Stephen A.↗

Biofilm formation and control in a simulated spacecraft water system - Three year results

Two simulated spacecraft water systems are being used to evaluate the effectiveness of iodine for controlling microbial contamination within such systems. An iodine concentration of about 2.0 mg/L is maintained in one system by passing ultrapure water through an iodinated ion exchange resin. Stainless steel coupons with electropolished and mechanically-polished sides are being used to monitor biofilm formation. Results after three years of operation show a single episode of significant bacterial growth in the iodinated system when the iodine level dropped to 1.9 mg/L. This growth was apparently controlled by replacing the iodinated ion exchange resin, thereby increasing the iodine level. The second batch of resin has remained effective in controlling microbial growth down to an iodine level of 1.0 mg/L. SEM indicates that the iodine has impeded but may have not completely eliminated the formation of biofilm. Metals analyses reveal some corrosion in the iodinated system after 3 years of continuous exposure. Significant microbial contamination has been present continuously in a parallel noniodinated system since the third week of operation.

Schultz, John R.↗

Computational Inference of Vibratory System with Incomplete Modal Information Using Parallel, Interactive and Adaptive Markov Chains

Inverse analysis of vibratory system is an important subject in fault identification, model updating, and robust design and control. It is challenging subject because 1) the problem is oftentimes underdetermined while the measurements are limited and/or incomplete; 2) many combinations of parameters may yield results that are similar with respect to actual response measurements; and 3) uncertainties inevitably exist. The aim of this research is to leverage upon computational intelligence through statistical inference to facilitate an enhanced, probabilistic framework using incomplete modal response measurement. This new framework is built upon efficient inverse identification through optimization, whereas Bayesian inference is employed to account for the effect of uncertainties. To overcome the computational cost barrier, we adopt Markov chain Monte Carlo (MCMC) to characterize the target function/distribution. Instead of using single Markov chain in conventional Bayesian approach, we develop a new sampling theory with multiple parallel, interactive and adaptive Markov chains and incorporate it into Bayesian inference. This can harness the collective power of these Markov chains to realize the concurrent search of multiple local optima. The number of required Markov chains and their respective initial model parameters are automatically determined via Monte Carlo simulation-based sample pre-screening followed by K-means clustering analysis. These enhancements can effectively address the aforementioned challenges in finite element inverse analysis. The validity of this framework is systematically demonstrated through case studies.

K Zhou↗

Performance limitations in parallel processor simulations

A jet-engine model is partitioned and simulated on a parallel processor system consisting of five 8086/8087 floating-point computers. The simulation uses Heun's integration method. A near-optimal parallel simulation (in the sense of minimum execution time) achieves speedup of only 2.13 and efficiency of 42.6 percent, in effect wasting 57.4 percent of the available processing power. A detailed analysis identifies and graphically demonstrates why the system fails to achieve ideal performance (viz., speedup of 5 and efficiency of 100 percent). Inherent characteristics of the problem equations and solution algorithm account for the loss of nearly half of the available processing power. Overheads associated with interprocessor communication and processor synchronization account for only a small fraction of the lost processing power. The effects of these and other factors which limit parallel processor performance are illustrated through real-time timing-analyzer tracers describing the run/idle status of the parallel processors during the simulation.

O'Grady, E. Pearse↗

Reducing software mass through behavior control

Attention is given to the tradeoff between communication and computation as regards a planetary rover (both these subsystems are very power-intensive, and both can be the major driver of the rover's power subsystem, and therefore the minimum mass and size of the rover). Software techniques that can be used to reduce the requirements on both communciation and computation, allowing the overall robot mass to be greatly reduced, are discussed. Novel approaches to autonomous control, called behavior control, employ an entirely different approach, and for many tasks will yield a similar or superior level of autonomy to traditional control techniques, while greatly reducing the computational demand. Traditional systems have several expensive processes that operate serially, while behavior techniques employ robot capabilities that run in parallel. Traditional systems make extensive world models, while behavior control systems use minimal world models or none at all.

Miller, David P.↗

Parallel Optimization of an Earth System Model (100 Gigaflops and Beyond?)

We are developing an Earth System Model (ESM) to be used in research aimed to better understand the interactions between the components of the Earth System and to eventually predict their variations. Currently, our ESM includes models of the atmosphere, oceans and the important chemical tracers therein.

Parallel Optimization Earth System Model↗

Development of a Wake Vortex Spacing System for Airport Capacity Enhancement and Delay Reduction

The Terminal Area Productivity project has developed the technologies required (weather measurement, wake prediction, and wake measurement) to determine the aircraft spacing needed to prevent wake vortex encounters in various weather conditions. The system performs weather measurements, predicts bounds on wake vortex behavior in those conditions, derives safe wake spacing criteria, and validates the wake predictions with wake vortex measurements. System performance to date indicates that the potential runway arrival rate increase with Aircraft VOrtex Spacing System (AVOSS), considering common path effects and ATC delivery variance, is 5% to 12% depending on the ratio of large and heavy aircraft. The concept demonstration system, using early generation algorithms and minimal optimization, is performing the wake predictions with adequate robustness such that only 4 hard exceedances have been observed in 1235 wake validation cases. This performance demonstrates the feasibility of predicting wake behavior bounds with multiple uncertainties present, including the unknown aircraft weight and speed, weather persistence between the wake prediction and the observations, and the location of the weather sensors several kilometers from the approach location. A concept for the use of the AVOSS system for parallel runway operations has been suggested, and an initial study at the JFK International Airport suggests that a simplified AVOSS system can be successfully operated using only a single lidar as both the weather sensor and the wake validation instrument. Such a selfcontained AVOSS would be suitable for wake separation close to the airport, as is required for parallel approach concepts such as SOIA.

Hinton, David A.↗

High Performance Programming Using Explicit Shared Memory Model on Cray T3D1

The Cray T3D system is the first-phase system in Cray Research, Inc.'s (CRI) three-phase massively parallel processing (MPP) program. This system features a heterogeneous architecture that closely couples DEC's Alpha microprocessors and CRI's parallel-vector technology, i.e., the Cray Y-MP and Cray C90. An overview of the Cray T3D hardware and available programming models is presented. Under Cray Research adaptive Fortran (CRAFT) model four programming methods (data parallel, work sharing, message-passing using PVM, and explicit shared memory model) are available to the users. However, at this time data parallel and work sharing programming models are not available to the user community. The differences between standard PVM and CRI's PVM are highlighted with performance measurements such as latencies and communication bandwidths. We have found that the performance of neither standard PVM nor CRI s PVM exploits the hardware capabilities of the T3D. The reasons for the bad performance of PVM as a native message-passing library are presented. This is illustrated by the performance of NAS Parallel Benchmarks (NPB) programmed in explicit shared memory model on Cray T3D. In general, the performance of standard PVM is about 4 to 5 times less than obtained by using explicit shared memory model. This degradation in performance is also seen on CM-5 where the performance of applications using native message-passing library CMMD on CM-5 is also about 4 to 5 times less than using data parallel methods. The issues involved (such as barriers, synchronization, invalidating data cache, aligning data cache etc.) while programming in explicit shared memory model are discussed. Comparative performance of NPB using explicit shared memory programming model on the Cray T3D and other highly parallel systems such as the TMC CM-5, Intel Paragon, Cray C90, IBM-SP1, etc. is presented.

Simon, Horst D.↗

Status report on HFIR irradiation of optimized alumina forming alloys

Properties of FeCrAl alloys under neutron irradiation are of interest because of these materials’ potential application as accident-tolerant fuel cladding in nuclear systems. In parallel, alumina-forming austenitic (AFA) alloys are of interest for use as structural materials in advanced nuclear systems for their potential higher resistance to embrittlement and high-temperature steam oxidation resistance. An irradiation campaign for fiscal year 2024 has been developed under the Advanced Fuels Campaign to perform irradiation testing of various FeCrAl and AFA alloys in Oak Ridge National Laboratory’s High Flux Isotope Reactor (HFIR). The goals of this irradiation campaign are to (1) study the impact of minor alloying elements on the neutron-irradiated mechanical properties of FeCrAl alloys and (2) collect neutron-irradiated mechanical properties on AFA alloys for comparison with those of FeCrAl alloys. This campaign will include both tensile and fracture toughness specimens tested following HFIR irradiation at temperatures representative of normal operating conditions in light-water reactors. The pre-irradiation characterization to date, the irradiation plan for the FeCrAl and AFA specimens, and the subsequent post-irradiation experimental test plan are presented in this report, along with the status of HFIR builds and scheduled insertion dates.

36 MATERIALS SCIENCE↗

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability↗

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability↗

Unobtrusive Software and System Health Management with R2U2 on a Parallel MIMD Coprocessor

Dynamic monitoring of software and system health of a complex cyber-physical system requires observers that continuously monitor variables of the embedded software in order to detect anomalies and reason about root causes. There exists a variety of techniques for code instrumentation, but instrumentation might change runtime behavior and could require costly software re-certification. In this paper, we present R2U2E, a novel realization of our real-time, Realizable, Responsive, and Unobtrusive Unit (R2U2). The R2U2E observers are executed in parallel on a dedicated 16-core EPIPHANY co-processor, thereby avoiding additional computational overhead to the system under observation. A DMA-based shared memory access architecture allows R2U2E to operate without any code instrumentation or program interference.

Schumann, Johann↗

Operationally Efficient Propulsion System Study (OEPSS) Data Book: Integrated Booster Propulsion Module (BPM) Engine Start Dynamics - Volume 8

A fluid-dynamic, digital-transient computer model of an integrated, parallel propulsion system was developed for the CDC mainframe and the SUN workstation computers. Since all STME component designs were used for the integrated system, computer subroutines were written characterizing the performance and geometry of all the components used in the system, including the manifolds. Three transient analysis reports were completed. The first report evaluated the feasibility of integrated engine systems in regards to the start and cutoff transient behavior. The second report evaluated turbopump out and combined thrust chamber/turbopump out conditions. The third report presented sensitivity study results in staggered gas generator spin start and in pump performance characteristics.

Kemp, Victoria R.↗

Multiple Coarse Grid Multigrid Methods for Solving Elliptic Problems

In this paper we describe some classes of multigrid methods for solving large linear systems arising in the solution by finite difference methods of certain boundary value problems involving Poisson's equation on rectangular regions. If parallel computing systems are used, then with standard multigrid methods many of the processors will be idle when one is working at the coarsest grid levels. We describe the use of Multiple Coarse Grid MultiGrid (MCGMG) methods. Here one first constructs a periodic set of equations corresponding to the given system. One then constructs a set of coarse grids such that for each grid corresponding to the grid size h there are four grids corresponding to the grid size 2*h. Multigrid operations such as restriction of residuals and interpolation of corrections are done in parallel at each grid level. For suitable choices of the multigrid operators the MCGMG method is equivalent to the Parallel Superconvergent MultiGrid (PSMG) method of Frederickson and McBryan. The convergence properties of MCGMG methods can be accurately analyzed using spectral methods.

Xiao, Shengyou↗

Exploiting user activeness for data retention in HPC systems

HPC systems typically rely on the fixed-lifetime (FLT) data retention strategy, which only considers temporal locality of data accesses to parallel file systems. However, our extensive analysis based on the leadership-class HPC system traces suggests that the FLT approach often fails to capture the dynamics in users' behavior and leads to undesired data purge. In this study, we propose an activeness-based data retention (ActiveDR) solution, which advocates considering the data retention approach from a holistic activeness-based perspective. By evaluating the frequency and impact of users' activities, ActiveDR prioritizes the file purge process for inactive users and rewards active users with extended file lifetime on parallel storage. Our extensive evaluations based on the traces of the prior Titan supercomputer show that, when reaching the same purge target, ActiveDR achieves up to 37% file miss reduction as compared to the current FLT retention methodology.

Zhang, Wei↗

Cable Tester Box

Cables are very important electrical devices that carry power and signals across multiple instruments. Any fault in a cable can easily result in a catastrophic outcome. Therefore, verifying that all cables are built to spec is a very important part of Electrical Integration Procedures. Currently, there are two methods used in lab for verifying cable connectivity. (1) Using a Break-Out Box and an ohmmeter this method is time-consuming but effective for custom cables and (2) Commercial Automated Cable Tester Boxes this method is fast, but to test custom cables often requires pre-programmed configuration files, and cables used on spacecraft are often uniquely designed for specific purposes. The idea is to develop a semi-automatic continuity tester that reduces human effort in cable testing, speeds up the electrical integration process, and ensures system safety. The JPL-Cable Tester Box is developed to check every single possible electrical connection in a cable in parallel. This system indicates connectivity through LED (light emitting diode) circuits. Users can choose to test any pin/shell (test node) with a single push of a button, and any other nodes that are shorted to the test node, even if they are in the same connector, will light up with the test node. The JPL-Cable Tester Boxes offers the following advantages: 1. Easy to use: The architecture is simple enough that it only takes 5 minutes for anyone to learn how operate the Cable Tester Box. No pre-programming and calibration are required, since this box only checks continuity. 2. Fast: The cable tester box checks all the possible electrical connections in parallel at a push of a button. If a cable normally takes half an hour to test, using the Cable Tester Box will improve the speed to as little as 60 seconds to complete. 3. Versatile: Multiple cable tester boxes can be used together. As long as all the boxes share the same electrical potential, any number of connectors can be tested together.

Lee, Jason H.↗

Parallel discrete-event simulation of FCFS stochastic queueing networks

Physical systems are inherently parallel. Intuition suggests that simulations of these systems may be amenable to parallel execution. The parallel execution of a discrete-event simulation requires careful synchronization of processes in order to ensure the execution's correctness; this synchronization can degrade performance. Largely negative results were recently reported in a study which used a well-known synchronization method on queueing network simulations. Discussed here is a synchronization method (appointments), which has proven itself to be effective on simulations of FCFS queueing networks. The key concept behind appointments is the provision of lookahead. Lookahead is a prediction on a processor's future behavior, based on an analysis of the processor's simulation state. It is shown how lookahead can be computed for FCFS queueing network simulations, give performance data that demonstrates the method's effectiveness under moderate to heavy loads, and discuss performance tradeoffs between the quality of lookahead, and the cost of computing lookahead.

Nicol, David M.↗