Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Parallel Multi-Step/Multi-Rate Integration of Two-Time Scale Dynamic Systems

Increasing demands on the fidelity of simulations for real-time and high-fidelity simulations are stressing the capacity of modern processors. New integration techniques are required that provide maximum efficiency for systems that are parallelizable. However many current techniques make assumptions that are at odds with non-cascadable systems. A new serial multi-step/multi-rate integration algorithm for dual-timescale continuous state systems is presented which applies to these systems, and is extended to a parallel multi-step/multi-rate algorithm. The superior performance of both algorithms is demonstrated through a representative example.

dynamics↗

Parallel processing methods for space based power systems

This report presents a method for doing load-flow analysis of a power system by using a decomposition approach. The power system for the Space Shuttle is used as a basis to build a model for the load-flow analysis. To test the decomposition method for doing load-flow analysis, simulations were performed on power systems of 16, 25, 34, 43, 52, 61, 70, and 79 nodes. Each of the power systems was divided into subsystems and simulated under steady-state conditions. The results from these tests have been found to be as accurate as tests performed using a standard serial simulator. The division of the power systems into different subsystems was done by assigning a processor to each area. There were 13 transputers available, therefore, up to 13 different subsystems could be simulated at the same time. This report has preliminary results for a load-flow analysis using a decomposition principal. The report shows that the decomposition algorithm for load-flow analysis is well suited for parallel processing and provides increases in the speed of execution.

Berry, F. C.↗

File-access characteristics of parallel scientific workloads

Phenomenal improvements in the computational performance of multiprocessors have not been matched by comparable gains in I/O system performance. This imbalance has resulted in I/O becoming a significant bottleneck for many scientific applications. One key to overcoming this bottleneck is improving the performance of parallel file systems. The design of a high-performance parallel file system requires a comprehensive understanding of the expected workload. Unfortunately, until recently, no general workload studies of parallel file systems have been conducted. The goal of the CHARISMA project was to remedy this problem by characterizing the behavior of several production workloads, on different machines, at the level of individual reads and writes. The first set of results from the CHARISMA project describe the workloads observed on an Intel iPSC/860 and a Thinking Machines CM-5. This paper is intended to compare and contrast these two workloads for an understanding of their essential similarities and differences, isolating common trends and platform-dependent variances. Using this comparison, we are able to gain more insight into the general principles that should guide parallel file-system design.

Nieuwejaar, Nils↗

GEOS Atmospheric Model: Challenges at Exascale

The Goddard Earth Observing System (GEOS) model at NASA's Global Modeling and Assimilation Office (GMAO) is used to simulate the multi-scale variability of the Earth's weather and climate, and is used primarily to assimilate conventional and satellite-based observations for weather forecasting and reanalysis. In addition, assimilations coupled to an ocean model are used for longer-term forecasting (e.g., El Nino) on seasonal to interannual times-scales. The GMAO's research activities, including system development, focus on numerous time and space scales, as detailed on the GMAO website, where they are tabbed under five major themes: Weather Analysis and Prediction; Seasonal-Decadal Analysis and Prediction; Reanalysis; Global Mesoscale Modeling, and Observing System Science. A brief description of the GEOS systems can also be found at the GMAO website. GEOS executes as a collection of earth system components connected through the Earth System Modeling Framework (ESMF). The ESMF layer is supplemented with the MAPL (Modeling, Analysis, and Prediction Layer) software toolkit developed at the GMAO, which facilitates the organization of the computational components into a hierarchical architecture. GEOS systems run in parallel using a horizontal decomposition of the Earth's sphere into processing elements (PEs). Communication between PEs is primarily through a message passing framework, using the message passing interface (MPI), and through explicit use of node-level shared memory access via the SHMEM (Symmetric Hierarchical Memory access) protocol. Production GEOS weather prediction systems currently run at 12.5-kilometer horizontal resolution with 72 vertical levels decomposed into PEs associated with 5,400 MPI processes. Research GEOS systems run at resolutions as fine as 1.5 kilometers globally using as many as 30,000 MPI processes. Looking forward, these systems can be expected to see a 2 times increase in horizontal resolution every two to three years, as well as less frequent increases in vertical resolution. Coupling these resolution changes with increases in complexity, the computational demands on the GEOS production and research systems should easily increase 100-fold over the next five years. Currently, our 12.5 kilometer weather prediction system narrowly meets the time-to-solution demands of a near-real-time production system. Work is now in progress to take advantage of a hybrid MPI-OpenMP parallelism strategy, in an attempt to achieve a modest two-fold speed-up to accommodate an immediate demand due to increased scientific complexity and an increase in vertical resolution. Pursuing demands that require a 10- to 100-fold increases or more, however, would require a detailed exploration of the computational profile of GEOS, as well as targeted solutions using more advanced high-performance computing technologies. Increased computing demands of 100-fold will be required within five years based on anticipated changes in the GEOS production systems, increases of 1000-fold can be anticipated over the next ten years.

ESMF↗

Study of a hybrid multispectral processor

A hybrid processor is described offering enough handling capacity and speed to process efficiently the large quantities of multispectral data that can be gathered by scanner systems such as MSDS, SKYLAB, ERTS, and ERIM M-7. Combinations of general-purpose and special-purpose hybrid computers were examined to include both analog and digital types as well as all-digital configurations. The current trend toward lower costs for medium-scale digital circuitry suggests that the all-digital approach may offer the better solution within the time frame of the next few years. The study recommends and defines such a hybrid digital computing system in which both special-purpose and general-purpose digital computers would be employed. The tasks of recognizing surface objects would be performed in a parallel, pipeline digital system while the tasks of control and monitoring would be handled by a medium-scale minicomputer system. A program to design and construct a small, prototype, all-digital system has been started.

Marshall, R. E.↗

Using Virtual Reality to Envision Deployment of Spacesuit-Compatible Augmented Reality Displays for Lunar Surface Operations

The National Aeronautics and Space Administration (NASA) will soon land crew on the lunar surface to establish a sustainable presence and develop operational concepts for future long-duration missions. New technologies will be necessary to extend planning and execution capabilities for lunar surface activities. The Joint Augmented Reality Visual Informatics System (Joint AR) at NASA Johnson Space Center (JSC) is one such technology. Joint AR is a suit-mounted augmented reality (AR) display and compute system which facilitates unprecedented information exchange and data visualization capabilities between mission support operators and suited crew. This paper describes challenges associated with developing an AR technology for an envisioned work domain by applying a sociotechnical lens to the iterative testing and development of novel AR technology through virtual reality (VR). A foundational VR testbed was established, providing a high-fidelity approximation of the lunar surface and enabling testing of envisioned AR system features in parallel with real world product development. Using this environment, a series of human-in-the-loop experiments were conducted using VR to deploy a notional AR system supporting use cases envisioned for lunar exploration extravehicular activity (xEVA). Our findings indicate VR is a powerful and immersive tool for testing capabilities which extend beyond the limitations of current AR technology. Our VR testbed enables early testing of proposed system features, advancing identification of high-value features and driving present-day development of Joint AR. Future work directions are discussed, including lessons learned and a roadmap for an iterative research and development approach applying VR (among other testbeds) to accelerate Joint AR system maturation.

Augmented reality↗

Rhomboid prism pair for rotating the plane of parallel light beams

An optical system is described for rotating the plane defined by a pair of parallel light beams. In one embodiment a single pair of rhomboid prisms have their respective input faces disposed to receive the respective input beams. Each prism is rotated about an axis of revolution coaxial with each of the respective input beams by means of a suitable motor and gear arrangement to cause the plane of the parallel output beams to be rotated relative to the plane of the input beams. In a second embodiment, two pairs of rhomboid prisms are provided. In a first angular orientation of the output beams, the prisms merely decrease the lateral displacement of the output beams in order to keep in the same plane as the input beams. In a second angular orientation of the prisms, the input faces of the second pair of prisms are brought into coincidence with the input beams for rotating the plane of the output beams by a substantial angle such as 90 deg.

Orloff, K. L.↗

Extensible Adaptable Simulation Systems: Supporting Multiple Fidelity Simulations in a Common Environment

Common practice in the development of simulation systems is meeting all user requirements within a single instantiation. The Joint Polar Satellite System (JPSS) presents a unique challenge to establish a simulation environment that meets the needs of a diverse user community while also spanning a multi-mission environment over decades of operation. In response, the JPSS Flight Vehicle Test Suite (FVTS) is architected with an extensible infrastructure that supports the operation of multiple observatory simulations for a single mission and multiple mission within a common system perimeter. For the JPSS-1 satellite, multiple fidelity flight observatory simulations are necessary to support the distinct user communities consisting of the Common Ground System development team, the Common Ground System Integration & Test team, and the Mission Rehearsal Team/Mission Operations Team. These key requirements present several challenges to FVTS development. First, the FVTS must ensure all critical user requirements are satisfied by at least one fidelity instance of the observatory simulation. Second, the FVTS must allow for tailoring of the system instances to function in diverse operational environments from the High-security operations environment at NOAA Satellite Operations Facility (NSOF) to the ground system factory floor. Finally, the FVTS must provide the ability to execute sustaining engineering activities on a subset of the system without impacting system availability to parallel users. The FVTS approach of allowing for multiple fidelity copies of observatory simulations represents a unique concept in simulator capability development and corresponds to the JPSS Ground System goals of establishing a capability that is flexible, extensible, and adaptable.

McLaughlin, Brian J.↗

TEAM Project Review, Year 2

This report summarizes our research activities within the TEAM project between December 2020 and December 2021, funded by the ASCR Advanced Research in Quantum Computing program. During the reporting period the LLNL-MSU team has made progress on several fronts. An overarching goal of the team is to provide a comprehensive suite of software tools that can be used for the Characterize-Optimize-Compute loop needed to implement and execute algorithms on quantum devices. We are concurrently developing lightweight solvers that can be used on desktop computers to find optimal control pulses and to characterize small quantum systems (consisting of a few transmons and cavities). However, desktop computers are insufficient for simulating and characterizing larger quantum systems. We have therefore also developed parallel, distributed memory, simulators and optimization solvers, both for open and closed quantum systems. These parallel solvers have, for example, been used to study quantum optimal control for pure-state preparation, utilizing 1000’s of cores on a modern high-performance computing (HPC) platform.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Operation and control of hybrid HVDC system with LCC and full-bridge MMC connected in parallel

In this study, a new kind of hybrid high-voltage direct current (HVDC) system is proposed. Each terminal of the proposed system consists of one line commutated converter (LCC) and one full-bridge modular multilevel converter (FB-MMC). The LCC and FB-MMC are connected in parallel so that they can share the same transmission line. The active–reactive power capability of the hybrid HVDC system is extended compared with the conventional LCC-HVDC system, and power reversal control without power interruption can be achieved by the coordination control of LCC and FB-MMC. Besides, the proposed hybrid HVDC system is capable of handling DC fault, because both LCC and FB-MMC have DC fault blocking capability. Moreover, the power rating of FB-MMC can be designed to low value while keeping the bulk-power transmission capability of LCC. A two-terminal bipolar hybrid HVDC system is built in PSCAD/EMTDC. The simulation results verify the effectiveness and feasibility of the proposed hybrid topology and corresponding control strategies.

42 ENGINEERING↗

Parallelized direct execution simulation of message-passing parallel programs

As massively parallel computers proliferate, there is growing interest in findings ways by which performance of massively parallel codes can be efficiently predicted. This problem arises in diverse contexts such as parallelizing computers, parallel performance monitoring, and parallel algorithm development. In this paper we describe one solution where one directly executes the application code, but uses a discrete-event simulator to model details of the presumed parallel machine such as operating system and communication network behavior. Because this approach is computationally expensive, we are interested in its own parallelization specifically the parallelization of the discrete-event simulator. We describe methods suitable for parallelized direct execution simulation of message-passing parallel programs, and report on the performance of such a system, Large Application Parallel Simulation Environment (LAPSE), we have built on the Intel Paragon. On all codes measured to date, LAPSE predicts performance well typically within 10 percent relative error. Depending on the nature of the application code, we have observed low slowdowns (relative to natively executing code) and high relative speedups using up to 64 processors.

Dickens, Phillip M.↗

Scanning System for Laser Velocimeter

Interference fringes remain parallel and focus-spot diameter same. Scanning system proposed for laser velocimeter (laser Doppler anemometer) to maintain constant beam-crossing angle and beam-waist diameter maintaining beam waist locations at crossing points. As target fluid scanned, interference fringes formed by crossing beams remain parallel and the focus-spot diameter same. System allows accurate velocity profiles obtained in wind tunnels and other fluid flow systems.

Gunter, William D.↗

Dynamic coordination of a self-reconfigurable manipulator system

The authors present the dynamic coordination of a self-reconfigurable manipulator system capable of changing its mechanical structure according to given task requirements. The self-reconfiguration is achieved by reconfiguring the topology of a dual-arm system through serial, parallel, and bracing structures. Particular emphasis is placed on the dynamic coordination of two arms having three different dual-arm topologies. The authors develop the Cartesian space dynamic models of a dual-arm system of three dual-arm topologies and derive the kinematic and dynamic constraints imposed on two arms in cooperation. Dual-arm dynamic manipulabilities are defined to quantify the dynamic performance of three dual-arm topologies in terms of the efficiency of generating Cartesian accelerations. A methodology of selecting serial, parallel, and bracing structures based on dual-arm dynamic manipulabilities is provided.

Kim, Sungbok↗

Development, construction and tests of the Mu2e electromagnetic calorimeter mechanical structures

The “muon-to-electron conversion” (Mu2e) experiment at Fermilab will search for the charged lepton flavour violating neutrino-less coherent conversion of a muon into an electron in the field of an aluminum nucleus. The observation of this process would be the unambiguous evidence of the existence of physics beyond the standard model. Mu2e detectors comprise a straw-tracker, an electromagnetic calorimeter and an external veto for cosmic rays. In particular, the calorimeter provides excellent electron identification, a fast calorimetric online trigger, and complementary information to aid pattern recognition and track reconstruction. The detector has been designed as a state-of-the-art crystal calorimeter and employs 1348 pure Cesium Iodide (CsI) crystals readout by UV-extended silicon photosensors and fast front-end and digitization electronics. A design consisting of two identical annular matrices (named “disks”) positioned at the relative distance of 70 cm downstream the aluminum target along the muon beamline satisfies the Mu2e physics requirements. The hostile Mu2e operational conditions, in terms of radiation levels (total expected ionizing dose of 12 krad and a neutron fluence of 5 × 10$^{10}$ n/cm$^{2}$ @ 1 MeV$_{eq}$ (Si)/y), magnetic field intensity (1 T) and vacuum level (10$^{-4}$ Torr) have posed tight constraints on scintillating materials, sensors, electronics and on the design of the detector mechanical structures and material choice. The support structure of each 674 crystal matrix is composed of an aluminum hollow ring and parts made of open-cell vacuum-compatible carbon fiber. The photosensors and front-end electronics for the readout of each crystal are inserted in a machined copper holder and make a unique mechanical unit. The resulting 674 mechanical units are supported by a machined plate of vacuum-compatible plastic material. The plate also integrates the cooling system made of a network of copper lines flowing a low temperature radiation-hard fluid and placed in thermal contact with the copper holders to constitute a low resistance thermal bridge. The data acquisition electronics are hosted in aluminum custom crates positioned on the external lateral surface of the disks. The crates also integrate the electronics cooling system as lines running in parallel to the front-end system. In this paper we report on the calorimeter mechanical structure design, the mechanical and thermal simulations that have determined the design technological choices, and the status of component production, quality assurance tests and plans for assembly at Fermilab.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Differential Draining of Parallel-Fed Propellant Tanks in Morpheus and Apollo Flight

Parallel-fed propellant tanks are an advantageous configuration for many spacecraft. Parallel-fed tanks allow the center of gravity (cg) to be maintained over the engine(s), as opposed to serial-fed propellant tanks which result in a cg shift as propellants are drained from tank one tank first opposite another. Parallel-fed tanks also allow for tank isolation if that is needed. Parallel tanks and feed systems have been used in several past vehicles including the Apollo Lunar Module. The design of the feedsystem connecting the parallel tank is critical to maintain balance in the propellant tanks. The design must account for and minimize the effect of manufacturing variations that could cause delta-p or mass flowrate differences, which would lead to propellant imbalance. Other sources of differential draining will be discussed. Fortunately, physics provides some self-correcting behaviors that tend to equalize any initial imbalance. The question concerning whether or not active control of propellant in each tank is required or can be avoided or not is also important to answer. In order to provide data on parallel-fed tanks and differential draining in flight for cryogenic propellants (as well as any other fluid), a vertical test bed (flying lander) for terrestrial use was employed. The Morpheus vertical test bed is a parallel-fed propellant tank system that uses passive design to keep the propellant tanks balanced. The system is operated in blow down. The Morpheus vehicle was instrumented with a capacitance level sensor in each propellant tank in order to measure the draining of propellants in over 34 tethered and 12 free flights. Morpheus did experience an approximately 20 lb/m imbalance in one pair of tanks. The cause of this imbalance will be discussed. This paper discusses the analysis, design, flight simulation vehicle dynamic modeling, and flight test of the Morpheus parallel-fed propellant. The Apollo LEM data is also examined in this summary report of the flight data.

Hurlbert, Eric↗

Testing New Programming Paradigms with NAS Parallel Benchmarks

Over the past decade, high performance computing has evolved rapidly, not only in hardware architectures but also with increasing complexity of real applications. Technologies have been developing to aim at scaling up to thousands of processors on both distributed and shared memory systems. Development of parallel programs on these computers is always a challenging task. Today, writing parallel programs with message passing (e.g. MPI) is the most popular way of achieving scalability and high performance. However, writing message passing programs is difficult and error prone. Recent years new effort has been made in defining new parallel programming paradigms. The best examples are: HPF (based on data parallelism) and OpenMP (based on shared memory parallelism). Both provide simple and clear extensions to sequential programs, thus greatly simplify the tedious tasks encountered in writing message passing programs. HPF is independent of memory hierarchy, however, due to the immaturity of compiler technology its performance is still questionable. Although use of parallel compiler directives is not new, OpenMP offers a portable solution in the shared-memory domain. Another important development involves the tremendous progress in the internet and its associated technology. Although still in its infancy, Java promisses portability in a heterogeneous environment and offers possibility to "compile once and run anywhere." In light of testing these new technologies, we implemented new parallel versions of the NAS Parallel Benchmarks (NPBs) with HPF and OpenMP directives, and extended the work with Java and Java-threads. The purpose of this study is to examine the effectiveness of alternative programming paradigms. NPBs consist of five kernels and three simulated applications that mimic the computation and data movement of large scale computational fluid dynamics (CFD) applications. We started with the serial version included in NPB2.3. Optimization of memory and cache usage was applied to several benchmarks, noticeably BT and SP, resulting in better sequential performance. In order to overcome the lack of an HPF performance model and guide the development of the HPF codes, we employed an empirical performance model for several primitives found in the benchmarks. We encountered a few limitations of HPF, such as lack of supporting the "REDISTRIBUTION" directive and no easy way to handle irregular computation. The parallelization with OpenMP directives was done at the outer-most loop level to achieve the largest granularity. The performance of six HPF and OpenMP benchmarks is compared with their MPI counterparts for the Class-A problem size in the figure in next page. These results were obtained on an SGI Origin2000 (195MHz) with MIPSpro-f77 compiler 7.2.1 for OpenMP and MPI codes and PGI pghpf-2.4.3 compiler with MPI interface for HPF programs.

Jin, H.↗

Flow of GE90 Turbofan Engine Simulated

The objective of this task was to create and validate a three-dimensional model of the GE90 turbofan engine (General Electric) using the APNASA (average passage) flow code. This was a joint effort between GE Aircraft Engines and the NASA Lewis Research Center. The goal was to perform an aerodynamic analysis of the engine primary flow path, in under 24 hours of CPU time, on a parallel distributed workstation system. Enhancements were made to the APNASA Navier-Stokes code to make it faster and more robust and to allow for the analysis of more arbitrary geometry. The resulting simulation exploited the use of parallel computations by using two levels of parallelism, with extremely high efficiency.The primary flow path of the GE90 turbofan consists of a nacelle and inlet, 49 blade rows of turbomachinery, and an exhaust nozzle. Secondary flows entering and exiting the primary flow path-such as bleed, purge, and cooling flows-were modeled macroscopically as source terms to accurately simulate the engine. The information on these source terms came from detailed descriptions of the cooling flow and from thermodynamic cycle system simulations. These provided boundary condition data to the three-dimensional analysis. A simplified combustor was used to feed boundary conditions to the turbomachinery. Flow simulations of the fan, high-pressure compressor, and high- and low-pressure turbines were completed with the APNASA code.

Veres, Joseph P.↗

Diagnostics of the vibrations of complex rotor systems

The parameters of the imbalance of a complex rotor system, having n parallel rotors and having six degrees of freedom, can be determined from the parameters of the vibrations of two appropriate degrees of freedom. This considerably simplifies diagnostics of the vibrations of complex rotor systems.

Yugraytis, I. Y.↗