Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed and parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Distributed Finite Element Analysis Using a Transputer Network

The principal objective of this research effort was to demonstrate the extraordinarily cost effective acceleration of finite element structural analysis problems using a transputer-based parallel processing network. This objective was accomplished in the form of a commercially viable parallel processing workstation. The workstation is a desktop size, low-maintenance computing unit capable of supercomputer performance yet costs two orders of magnitude less. To achieve the principal research objective, a transputer based structural analysis workstation termed XPFEM was implemented with linear static structural analysis capabilities resembling commercially available NASTRAN. Finite element model files, generated using the on-line preprocessing module or external preprocessing packages, are downloaded to a network of 32 transputers for accelerated solution. The system currently executes at about one third Cray X-MP24 speed but additional acceleration appears likely. For the NASA selected demonstration problem of a Space Shuttle main engine turbine blade model with about 1500 nodes and 4500 independent degrees of freedom, the Cray X-MP24 required 23.9 seconds to obtain a solution while the transputer network, operated from an IBM PC-AT compatible host computer, required 71.7 seconds. Consequently, the $80,000 transputer network demonstrated a cost-performance ratio about 60 times better than the $15,000,000 Cray X-MP24 system.

Watson, James↗

Heating and acceleration of ions with Kappa distribution functions by low‐frequency Alfvén wave

Abstract Heating and acceleration of ions with Kappa distribution functions (with parameter ) in low‐beta plasmas, by a low‐frequency Alfvén wave, is investigated using test‐particle simulations, yielding interesting new results. As long as the Alfvén wave amplitude is sufficiently large, the computed net heating energy of ions becomes independent of the wave frequency and amplitude, always approaching the same value of . The eventual energy of ions is dictated only by the initial ion energy and the ratio of the magnetic field energy density to the plasma density. The heating effect of the Kappa ions increases with . During the heating process, the ions are picked up by the Alfvén wave and pitch angle scattered, forming a quasi‐isotropic spherical shell velocity distribution. The Kappa ions are accelerated in the parallel direction, reaching a bulk flow speed roughly equal to the local Alfvén speed. Higher ‐value in the initial Kappa distribution leads to faster saturation. The above results may explain certain features of the ion heating and acceleration in the solar wind and corona.

Li, Kehua↗

A Robust Parallel Distributed State Estimation for Large Scale Distribution Systems

The growing need and interest in real-time monitoring of large distribution networks motivated by the rapid population of renewable sources, EVs and etc. demand a computationally efficient state estimation framework. Furthermore, this paper presents an improved computational framework for implementing a robust state estimator using a multi-core processor. The main contribution of the paper is the proposed computational framework along with two partitioning strategies which enable fast and robust state estimation for large scale radial and/or meshed distribution systems. Formulation of the proposed method and its implementation are described in detail. Performance of the estimator is tested by simulations first using a small 84-bus radial distribution system. Then the method’s scalability is demonstrated by simulations on two very large scale distribution networks one configured radially and the other meshed each containing over 12,500 buses.

42 ENGINEERING↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

A Parallel Pipelined Renderer for the Time-Varying Volume Data

This paper presents a strategy for efficiently rendering time-varying volume data sets on a distributed-memory parallel computer. Time-varying volume data take large storage space and visualizing them requires reading large files continuously or periodically throughout the course of the visualization process. Instead of using all the processors to collectively render one volume at a time, a pipelined rendering process is formed by partitioning processors into groups to render multiple volumes concurrently. In this way, the overall rendering time may be greatly reduced because the pipelined rendering tasks are overlapped with the I/O required to load each volume into a group of processors; moreover, parallelization overhead may be reduced as a result of partitioning the processors. We modify an existing parallel volume renderer to exploit various levels of rendering parallelism and to study how the partitioning of processors may lead to optimal rendering performance. Two factors which are important to the overall execution time are re-source utilization efficiency and pipeline startup latency. The optimal partitioning configuration is the one that balances these two factors. Tests on Intel Paragon computers show that in general optimal partitionings do exist for a given rendering task and result in 40-50% saving in overall rendering time.

Chiueh, Tzi-Cker↗

Enabling Recycling of Composites: Understanding the Impacts of Multiple Thermal Processing Cycles

When considering the utilization of recycled short carbon fiber feedstock materials for advanced manufacturing, understanding the material degradation behavior is essential in determining how many times a composite material can be effectively reprocessed and remanufactured. This study characterizes the degradation behavior of short carbon fiber acrylonitrile butadiene (CF-ABS) that has been reprocessed five times with twin screw extrusion. Parallel plate rheology was completed to observe the degradation in complex viscosity of the recycled feedstock materials. Gel permeation chromatography (GPC) was utilized to characterize the changes in molecular weight distribution of the recycled materials as a result of thermal and mechanical degradation during the re-processing steps. Rheological characterization, GPC, and twin-screw processing data help inform the process optimizations required to process the recycled feedstock material. Successful characterization of the degradation behavior of short fiber composite feedstock materials aids in increased understanding of the lifespan of high value carbon fiber composite materials and aids in process optimization of recycled composite materials.

Walker, Roo↗

Autoplan: A self-processing network model for an extended blocks world planning environment

Self-processing network models (neural/connectionist models, marker passing/message passing networks, etc.) are currently undergoing intense investigation for a variety of information processing applications. These models are potentially very powerful in that they support a large amount of explicit parallel processing, and they cleanly integrate high level and low level information processing. However they are currently limited by a lack of understanding of how to apply them effectively in many application areas. The formulation of self-processing network methods for dynamic, reactive planning is studied. The long-term goal is to formulate robust, computationally effective information processing methods for the distributed control of semiautonomous exploration systems, e.g., the Mars Rover. The current research effort is focusing on hierarchical plan generation, execution and revision through local operations in an extended blocks world environment. This scenario involves many challenging features that would be encountered in a real planning and control environment: multiple simultaneous goals, parallel as well as sequential action execution, action sequencing determined not only by goals and their interactions but also by limited resources (e.g., three tasks, two acting agents), need to interpret unanticipated events and react appropriately through replanning, etc.

Dautrechy, C. Lynne↗

Rocket measurements of electrons in a system of multiple auroral arcs

A Nike-Tomahawk rocket was launched into a system of auroral arcs northward of Poker Flat Research Range, Fairbanks, Alaska. The pitch-angle distribution of electrons was measured at 2.5, 5, and 10 keV and also at 10 keV on a separating forward section of the payload. The auroral activity appeared to be the extension of substorm activity centered to the east. The rocket crossed a westward-propagating fold in the brightest band. The electron spectrum was relatively hard through most of the flight, showing a peak in the range from 2.5 to 10 keV in the weaker aurora and below 5 keV in the brightest arc. The detailed structure of the pitch-angle distribution suggested that, at times, a very selective process was accelerating some electrons in the magnetic field direction, so that a narrow field-aligned component appeared superimposed on a more isotropic distribution. It is concluded that this process could not be a near-ionosphere field-aligned potential drop, although the more isotropic component may have been produced by a parallel electric field extending several thousand kilometers along the field line above the ionosphere.

Boyd, J. S.↗

A data distributed parallel algorithm for ray-traced volume rendering

This paper presents a divide-and-conquer ray-traced volume rendering algorithm and a parallel image compositing method, along with their implementation and performance on the Connection Machine CM-5, and networked workstations. This algorithm distributes both the data and the computations to individual processing units to achieve fast, high-quality rendering of high-resolution data. The volume data, once distributed, is left intact. The processing nodes perform local ray tracing of their subvolume concurrently. No communication between processing units is needed during this locally ray-tracing process. A subimage is generated by each processing unit and the final image is obtained by compositing subimages in the proper order, which can be determined a priori. Test results on both the CM-5 and a group of networked workstations demonstrate the practicality of our rendering algorithm and compositing method.

Ma, Kwan-Liu↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Models of Small-Scale Patchiness

Patchiness is perhaps the most salient characteristic of plankton populations in the ocean. The scale of this heterogeneity spans many orders of magnitude in its spatial extent, ranging from planetary down to microscale. It has been argued that patchiness plays a fundamental role in the functioning of marine ecosystems, insofar as the mean conditions may not reflect the environment to which organisms are adapted. Understanding the nature of this patchiness is thus one of the major challenges of oceanographic ecology. The patchiness problem is fundamentally one of physical-biological-chemical interactions. This interconnection arises from three basic sources: (1) ocean currents continually redistribute dissolved and suspended constituents by advection; (2) space-time fluctuations in the flows themselves impact biological and chemical processes, and (3) organisms are capable of directed motion through the water. This tripartite linkage poses a difficult challenge to understanding oceanic ecosystems: differentiation between the three sources of variability requires accurate assessment of property distributions in space and time, in addition to detailed knowledge of organismal repertoires and the processes by which ambient conditions control the rates of biological and chemical reactions. Various methods of observing the ocean tend to lie parallel to the axes of the space/time domain in which these physical-biological-chemical interactions take place. Given that a purely observational approach to the patchiness problem is not tractable with finite resources, the coupling of models with observations offers an alternative which provides a context for synthesis of sparse data with articulations of fundamental principles assumed to govern functionality of the system. In a sense, models can be used to fill the gaps in the space/time domain, yielding a framework for exploring the controls on spatially and temporally intermittent processes. The following discussion highlights only a few of the multitude of models which have yielded insight into the dynamics of plankton patchiness. In addition, this particular collection of examples is intended to furnish some exposure to the diversity of modeling approaches which can be brought to bear on the problem. These approaches range from abstract theoretical models intended to elucidate specific processes, to complex numerical formulations which can be used to actually simulate observed distributions in detail.

McGillicuddy, D. J.↗

ISEE 3 observations of solar wind thermal electrons with T-perpendicular greater than T-parallel

This study presents ISEE 3 observations of anomalous electron distributions for which T-perpendicular exceeds T-parallel in the solar wind near 1 AU. Twelve anomaly events were identified, lasting from 24 min to 6 hours. These events generally share the following characteristics: (1) high plasma density, (2) low solar wind speed, (3) magnetic field which is nearly transverse to the flow, and (4) low electron and ion temperatures. The processes of solar wind adiabatic expansion and isotropization via Coulomb collisions could be expected to lead to such anomalous anisotropies under conditions similar to those observed. However, these conditions actually produce T-perpendicular greater than T-parallel for only a small fraction of the time, suggesting that other mechanisms are also important in regulating solar wind electron distributions.

Phillips, J. L.↗

Operating Stresses and Their Effects on Degradation of LSM-Based SOFC Cathodes

The performance of solid oxides fuel cells (SOFCs) with four different Mn excess of lanthanum strontium manganite (LSM) -based cathodes were examined under different temperatures (1,000 °C, 900 °C), current densities (0, 380, and 760 mA cm -2 ), and cathode atmospheres (p O 2 )=0.1,0.15,0.21) for durations ranging from 58 to 1,008 h. Each yttria-stabilized zirconia (YSZ) electrolyte-supported “button” cell had a porous Ni/YSZ composite anode and a porous LSM/YSZ composite cathode. The cells’ output voltage versus time were recorded and electrochemical impedance spectroscopy (EIS) and linear sweep voltammetry (LSV) measurements were performed every 24 hours. The total area specific resistance (ASR) was calculated from these measurements. The values of ASR from both EIS and LSV were comparable between each other but lower than that from durability test. Distribution of relaxation times (DRT) analysis was performed to investigate the electrochemical processes and their corresponding relaxation frequencies as well as their attributions to the total ASR. The series resistance and parallel resistance that obtained from the equivalent circuit fit of the Nyquist plot were higher at low temperature and low p O 2 regardless of LSM compositions. The area of the peaks derived from DRT analysis increased with time during individual tests. The microstructures of the tested cells were examined using scanning electron microscopy and energy-dispersive x-ray spectroscopy (SEM/EDS). Manganese oxide particles were observed near the cathode-electrolyte interface after prolonged (1,008 h) testing in air of a cell with LSM of 11% manganese excess, and after a short test (58 h) under low oxygen (p O 2 = 0.10) in a cell with LSM of 5% manganese excess. However, the role of MnOx in cell degradation is still unknown.

08 HYDROGEN↗

Adaptive pattern recognition by mini-max neural networks as a part of an intelligent processor

In this decade and progressing into 21st Century, NASA will have missions including Space Station and the Earth related Planet Sciences. To support these missions, a high degree of sophistication in machine automation and an increasing amount of data processing throughput rate are necessary. Meeting these challenges requires intelligent machines, designed to support the necessary automations in a remote space and hazardous environment. There are two approaches to designing these intelligent machines. One of these is the knowledge-based expert system approach, namely AI. The other is a non-rule approach based on parallel and distributed computing for adaptive fault-tolerances, namely Neural or Natural Intelligence (NI). The union of AI and NI is the solution to the problem stated above. The NI segment of this unit extracts features automatically by applying Cauchy simulated annealing to a mini-max cost energy function. The feature discovered by NI can then be passed to the AI system for future processing, and vice versa. This passing increases reliability, for AI can follow the NI formulated algorithm exactly, and can provide the context knowledge base as the constraints of neurocomputing. The mini-max cost function that solves the unknown feature can furthermore give us a top-down architectural design of neural networks by means of Taylor series expansion of the cost function. A typical mini-max cost function consists of the sample variance of each class in the numerator, and separation of the center of each class in the denominator. Thus, when the total cost energy is minimized, the conflicting goals of intraclass clustering and interclass segregation are achieved simultaneously.

Szu, Harold H.↗

Multiscale and Multiphysics Modeling of Additive Manufacturing of Advanced Materials

The objective of this proposed project is to research and develop a prediction tool for advanced additive manufacturing (AAM) processes for advanced materials and develop experimental methods to provide fundamental properties and establish validation data. Aircraft structures and engines demand materials that are stronger, useable at much higher temperatures, provide less acoustic transmission, and enable more aeroelastic tailoring than those currently used. Significant improvements in properties can only be achieved by processing the materials under nonequilibrium conditions, such as AAM processes. AAM processes encompass a class of processes that use a focused heat source to create a melt pool on a substrate. Examples include Electron Beam Freeform Fabrication and Direct Metal Deposition. These types of additive processes enable fabrication of parts directly from CAD drawings. To achieve the desired material properties and geometries of the final structure, assessing the impact of process parameters and predicting optimized conditions with numerical modeling as an effective prediction tool is necessary. The targets for the processing are multiple and at different spatial scales, and the physical phenomena associated occur in multiphysics and multiscale. In this project, the research work has been developed to model AAM processes in a multiscale and multiphysics approach. A macroscale model was developed to investigate the residual stresses and distortion in AAM processes. A sequentially coupled, thermomechanical, finite element model was developed and validated experimentally. The results showed the temperature distribution, residual stress, and deformation within the formed deposits and substrates. A mesoscale model was developed to include heat transfer, phase change with mushy zone, incompressible free surface flow, solute redistribution, and surface tension. Because of excessive computing time needed, a parallel computing approach was also tested. In addition, after investigating various methods, a Smoothed Particle Hydrodynamics Model (SPH Model) was developed to model wire feeding process. Its computational efficiency and simple architecture makes it more robust and flexible than other models. More research on material properties may be needed to realistically model the AAM processes. A microscale model was developed to investigate heterogeneous nucleation, dendritic grain growth, epitaxial growth of columnar grains, columnar-to-equiaxed transition, grain transport in melt, and other properties. The orientations of the columnar grains were almost perpendicular to the laser motion's direction. Compared to the similar studies in the literature, the multiple grain morphology modeling result is in the same order of magnitude as optical morphologies in the experiment. Experimental work was conducted to validate different models. An infrared camera was incorporated as a process monitoring and validating tool to identify the solidus and mushy zones during deposition. The images were successfully processed to identify these regions. This research project has investigated multiscale and multiphysics of the complex AAM processes thus leading to advanced understanding of these processes. The project has also developed several modeling tools and experimental validation tools that will be very critical in the future of AAM process qualification and certification.

Liou, Frank↗

A Scalable Parallel Hypergraph Generator (HyGen)

Graphs are extensively used to model real-world complex systems. An edge in a graph can model pairwise relationships. However, multiway relationships (connections between three or more vertices) are common in many complex systems such as cellular process, image segmentation, and circuit design. A graph edge cannot model multiway relationships. A hypergraph, which can connect more than two vertices, is thus a better option to model multiway relationships. A large-scale hypergraph analysis has the potential to find useful insights from a complex system and assist in knowledge discovery. Currently a limited number of hypergraphs exists that are representative of real-world datasets. Moreover, real-world hypergraph datasets are small in size and inadequate to incorporate future needs. A graph generator that can produce large-scale synthetic hypergraphs can solve the above mentioned problems. In this paper, we present a scalable parallel hypergraph generator (HyGen) based on the Message Passing Interface (MPI) standard. To generate hypergraphs, HyGen takes the following parameter values as inputs: i) number of vertices, ii) number of hyperedges, iii) number of clusters, iv) vertex distribution, v) hyperedge distribution, vi) local cluster cardinality, and vii) global cluster cardinality. We have demonstrated that HyGen can generate hypergraphs of various sizes in a scalable fashion. HyGen takes approximately four minutes to generate a hypergraph with 4.8 million vertices, 1.6 million hyperedges, and 800 clusters using 1,024 processes on a leadership class computing platform. Our strong and weak scaling experiments on supercomputers demonstrate that HyGen can quickly create large-scale hypergraphs in a parallel manner, thus providing a useful capability for hypergraph analysis.

Hasan, S M Shamimul↗

Basic cluster compression algorithm

Feature extraction and data compression of LANDSAT data is accomplished by BCCA program which reduces costs associated with transmitting, storing, distributing, and interpreting multispectral image data. Algorithm uses spatially local clustering to extract features from image data to describe spectral characteristics of data set. Approach requires only simple repetitive computations, and parallel processing can be used for very high data rates. Program is written in FORTRAN IV for batch execution and has been implemented on SEL 32/55.

Hilbert, E. E.↗

Wave scheduling - Decentralized scheduling of task forces in multicomputers

Decentralized operating systems that control large multicomputers need techniques to schedule competing parallel programs called task forces. Wave scheduling is a probabilistic technique that uses a hierarchical distributed virtual machine to schedule task forces by recursively subdividing and issuing wavefront-like commands to processing elements capable of executing individual tasks. Wave scheduling is highly resistant to processing element failures because it uses many distributed schedulers that dynamically assign scheduling responsibilities among themselves. The scheduling technique is trivially extensible as more processing elements join the host multicomputer. A simple model of scheduling cost is used by every scheduler node to distribute scheduling activity and minimize wasted processing capacity by using perceived workload to vary decentralized scheduling rules. At low to moderate levels of network activity, wave scheduling is only slightly less efficient than a central scheduler in its ability to direct processing elements to accomplish useful work.

Van Tilborg, A. M.↗