Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Disorderly Conduct of Benzamide IV: Crystallographic and Computational Analysis of High Entropy Polymorphs of Small Molecules

Benzamide, a simple derivative of benzoic acid and a common intermediate of pharmaceutical compounds, was reported to form two polymorphs in 1832, but the single crystal structure of the more stable form was not solved until 1959. Nearly 50 years later, the second form was characterized by powder diffraction, followed shortly thereafter by characterization of a third form, a polytype of the most thermodynamically stable Form I. These two new forms, Forms II and III, are metastable. Herein, we describe a fourth polymorph, Form IV, discovered by melt crystallization concurrently with its crystallization under confinement at small length scales (<10 nm), where it is stable indefinitely. Form III exists under confinement in larger pores, and melting point data for different pore sizes corroborate the existence of Form IV below 10 nm. Form IV is highly disordered, precluding indexing of powder diffraction data other than hk0 reflections. Nonetheless, a combination of powder X-ray diffraction and computational crystal structure prediction reveals that Form IV contains a 2D motif resembling that of Form II, but with longrange order in the third dimension masked by ubiquitous stacking faults. This approach relies on distilling a large number of candidate structures to a few possible disorder models based on benzamide tetrads that organize in 2D parquet-like tiles, with organization along the third dimension, that can be modeled with various stacking fault configurations having distinct intermolecular interactions and translations in the dimension orthogonal to the tiling planes. These observations reveal a bewildering crystallographic complexity for such a simple molecule. Nonetheless, the approach described herein demonstrates that challenging structures that may be abandoned prematurely because of poor crystallinity, twinning, or disorder can be solved.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Detection and identification of benthic communities and shoreline features in Biscayne Bay

Progress made in the development of a technique for identifying and delinating benthic and shoreline communities using multispectral imagery is described. Images were collected with a multispectral scanner system mounted in a C-47 aircraft. Concurrent with the overflight, ecological ground- and sea-truth information was collected at 19 sites in the bay and on the shore. Preliminary processing of the scanner imagery with a CDC 1604 digital computer provided the optimum channels for discernment among different underwater and coastal objects. Automatic mapping of the benthic plants by multiband imagery and the mapping of isotherms and hydrodynamic parameters by digital model can become an effective predictive ecological tool when coupled together. Using the two systems, it appears possible to predict conditions that could adversely affect the benthic communities. With the advent of the ERTS satellites and space platforms, imagery data could be obtained which, when used in conjunction with water-level and meteorological data, would provide for continuous ecological monitoring.

Kolipinski, M. C.↗

Analysis of the NASA/MSFC airborne Doppler lidar results from San Gorgonio Pass, California

The NASA/MSFC Airborne Doppler Lidar System was flown in July 1981 aboard the NASA/Ames Convair 990 on the east side of San Gorgonio Pass California, near Palm Springs, to measure and investigate the accelerated atmospheric wind field discharging from the pass. At this region, the maritime layer from the west coast accelerates through the pass and spreads out over the valley floor on the east side of the pass. The experiment was selected in order to study accelerated flow in and at the exit of the canyon. Ground truth wind data taken concurrently with the flight data were available from approximately 12 meteorological towers and 3 tala kites for limited comparison purposes. The experiment provided the first spatial data for ensemble averaging of spatial correlations to compute lateral and longitudinal length scales in the lateral and longitudinal directions for both components, and information on atmospheric flow in this region of interest from wind energy resource considerations.

Cliff, W. C.↗

Space station Ada runtime support for nested atomic transactions

The Space Station Data Management System (DMS), associated computing subsystems, and applications have varying degrees of reliability associated with their operation. A model has been developed (McKay '86) which allows the DMS runtime environment to appear as an Ada virtual machine to applications executing within it. This model is modular, flexible, and dynamically configurable to allow for evolution and growth over time. Support for Fault-tolerant computing is included within this model. The basic primitive involved in this support is based on atomic actions (Grey '78). An atomic action possesses two fundamental properties: (1) it is indivisible with respect to concurrent actions, and (2) it is indivisible with respect to failure. A transaction is a collection of atomic actions which collectively appear to be one action. Transactions may be nested, providing even more powerful support for reliability. A proposed approach is described for providing support for nested atomic transactions within the Ada runtime model developed for the Space Station environment. The level of support is modular, flexible and dynamically configurable just like the overall runtime support environment.

Monteiro, Edward J.↗

A performance analysis method for distributed real-time robotic systems: A case study of remote teleoperation

Robot coordination and control systems for remote teleoperation applications are by necessity implemented on distributed computers. Modeling and performance analysis of these distributed robotic systems is difficult, but important for economic system design. Performance analysis methods originally developed for conventional distributed computer systems are often unsatisfactory for evaluating real-time systems. The paper introduces a formal model of distributed robotic control systems; and a performance analysis method, based on scheduling theory, which can handle concurrent hard-real-time response specifications. Use of the method is illustrated by a case of remote teleoperation which assesses the effect of communication delays and the allocation of robot control functions on control system hardware requirements.

Lefebvre, D. R.↗

Applications of concurrent neuromorphic algorithms for autonomous robots

This article provides an overview of studies at the Oak Ridge National Laboratory (ORNL) of neural networks running on parallel machines applied to the problems of autonomous robotics. The first section provides the motivation for our work in autonomous robotics and introduces the computational hardware in use. Section 2 presents two theorems concerning the storage capacity and stability of neural networks. Section 3 presents a novel load-balancing algorithm implemented with a neural network. Section 4 introduces the robotics test bed now in place. Section 5 concerns navigation issues in the test-bed system. Finally, Section 6 presents a frequency-coded network model and shows how Darwinian techniques are applied to issues of parameter optimization and on-line design.

Barhen, J.↗

ACSYNT inner loop flight control design study

The NASA Ames Research Center developed the Aircraft Synthesis (ACSYNT) computer program to synthesize conceptual future aircraft designs and to evaluate critical performance metrics early in the design process before significant resources are committed and cost decisions made. ACSYNT uses steady-state performance metrics, such as aircraft range, payload, and fuel consumption, and static performance metrics, such as the control authority required for the takeoff rotation and for landing with an engine out, to evaluate conceptual aircraft designs. It can also optimize designs with respect to selected criteria and constraints. Many modern aircraft have stability provided by the flight control system rather than by the airframe. This may allow the aircraft designer to increase combat agility, or decrease trim drag, for increased range and payload. This strategy requires concurrent design of the airframe and the flight control system, making trade-offs of performance and dynamics during the earliest stages of design. ACSYNT presently lacks means to implement flight control system designs but research is being done to add methods for predicting rotational degrees of freedom and control effector performance. A software module to compute and analyze the dynamics of the aircraft and to compute feedback gains and analyze closed loop dynamics is required. The data gained from these analyses can then be fed back to the aircraft design process so that the effects of the flight control system and the airframe on aircraft performance can be included as design metrics. This report presents results of a feasibility study and the initial design work to add an inner loop flight control system (ILFCS) design capability to the stability and control module in ACSYNT. The overall objective is to provide a capability for concurrent design of the aircraft and its flight control system, and enable concept designers to improve performance by exploiting the interrelationships between aircraft and flight control system design parameters.

Bortins, Richard↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Enabling Fortran Standard Parallelism in GAMESS for Accelerated Quantum Chemistry Calculations

The performance of Fortran 2008 DO CONCURRENT (DC) relative to OpenACC and OpenMP target offloading (OTO) with different compilers is studied for the GAMESS quantum chemistry application. Specifically, DC and OTO are used to offload the Fock build, which is a computational bottleneck in most quantum chemistry codes, to GPUs. The DC Fock build performance is studied on NVIDIA A100 and V100 accelerators and compared with the OTO versions compiled by the NVIDIA HPC, IBM XL, and Cray Fortran compilers. The results show that DC can speed up the Fock build by 3.0× compared with that of the OTO model. Finally, with similar offloading efforts, DC is a compelling programming model for offloading Fortran applications to GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Determination of cloud fields from analysis of HIRS2/MSU sounding data

IR and microwave remote sensing data collected with the HIRS2 and MSU sensors on the NOAA polar-orbiting satellites were evaluated for their effectiveness as bases for determining the cloud cover and cloud physical characteristics. Techniques employed to adjust for day-night alterations in the radiance fields are described, along with computational procedures applied to compare scene pixel values with reference values for clear skies. Sample results are provided for the mean cloud coverage detected over South America and Africa June 1979, with attention given to concurrent surface pressure and cloud top pressure values.

Susskind, J.↗

A data distributed parallel algorithm for ray-traced volume rendering

This paper presents a divide-and-conquer ray-traced volume rendering algorithm and a parallel image compositing method, along with their implementation and performance on the Connection Machine CM-5, and networked workstations. This algorithm distributes both the data and the computations to individual processing units to achieve fast, high-quality rendering of high-resolution data. The volume data, once distributed, is left intact. The processing nodes perform local ray tracing of their subvolume concurrently. No communication between processing units is needed during this locally ray-tracing process. A subimage is generated by each processing unit and the final image is obtained by compositing subimages in the proper order, which can be determined a priori. Test results on both the CM-5 and a group of networked workstations demonstrate the practicality of our rendering algorithm and compositing method.

Ma, Kwan-Liu↗

A Machine-Checked Proof of A State-Space Construction Algorithm

This paper presents the correctness proof of Saturation, an algorithm for generating state spaces of concurrent systems, implemented in the SMART tool. Unlike the Breadth First Search exploration algorithm, which is easy to understand and formalise, Saturation is a complex algorithm, employing a mutually-recursive pair of procedures that compute a series of non-trivial, nested local fixed points, corresponding to a chaotic fixed point strategy. A pencil-and-paper proof of Saturation exists, but a machine checked proof had never been attempted. The key element of the proof is the characterisation theorem of saturated nodes in decision diagrams, stating that a saturated node represents a set of states encoding a local fixed-point with respect to firing all events affecting only the node s level and levels below. For our purpose, we have employed the Prototype Verification System (PVS) for formalising the Saturation algorithm, its data structures, and for conducting the proofs.

Catano, Nestor↗

Secondary Science Teachers’ Implementation of a Curricular Intervention When Teaching With Global Climate Models

In the past decade, emphasis on promoting “climate literacy” in K-16 science classrooms has increased. Teachers play a critical role in cultivating these opportunities, especially in secondary science classrooms. However, most prior climate education research has focused on students and student learning; little is known about how teachers implement climate-focused curricular interventions. Here, we report findings from a concurrent mixed methods, multiple-case study of four secondary science teachers’ implementation of a new, NGSS-aligned, model-centric climate curriculum module grounded in the use of a data-driven, computer-based climate modeling tool—Easy Global Climate Model (EzGCM). We employ multiple data sources, including video-recorded classroom observations, interviews, and instructional artifacts, and both qualitative and quantitative analyses, to investigate how teachers implemented the curriculum. Findings show that, overall, teachers implemented the curriculum in ways that were less model-centric than designed, placing greater emphasis on EzGCM itself rather than using the model to investigate Earth’s changing climate. Additionally, we present detailed single-case studies of each participant teacher that highlight differences in teachers’ implementation of the curriculum module and their reasoning for making observed instructional decisions. This research sheds light on the design of secondary science learning environments by illustrating the varied ways teachers implement a climate-focused curriculum to support students’ developing climate literacy. This has important implications for the design of climate-focused curriculum and supports for teachers.

Secondary science teaching↗

Design of object-oriented distributed simulation classes

Distributed simulation of aircraft engines as part of a computer aided design package is being developed by NASA Lewis Research Center for the aircraft industry. The project is called NPSS, an acronym for 'Numerical Propulsion Simulation System'. NPSS is a flexible object-oriented simulation of aircraft engines requiring high computing speed. It is desirable to run the simulation on a distributed computer system with multiple processors executing portions of the simulation in parallel. The purpose of this research was to investigate object-oriented structures such that individual objects could be distributed. The set of classes used in the simulation must be designed to facilitate parallel computation. Since the portions of the simulation carried out in parallel are not independent of one another, there is the need for communication among the parallel executing processors which in turn implies need for their synchronization. Communication and synchronization can lead to decreased throughput as parallel processors wait for data or synchronization signals from other processors. As a result of this research, the following have been accomplished. The design and implementation of a set of simulation classes which result in a distributed simulation control program have been completed. The design is based upon MIT 'Actor' model of a concurrent object and uses 'connectors' to structure dynamic connections between simulation components. Connectors may be dynamically created according to the distribution of objects among machines at execution time without any programming changes. Measurements of the basic performance have been carried out with the result that communication overhead of the distributed design is swamped by the computation time of modules unless modules have very short execution times per iteration or time step. An analytical performance model based upon queuing network theory has been designed and implemented. Its application to realistic configurations has not been carried out.

Schoeffler, James D.↗

Design of Object-Oriented Distributed Simulation Classes

Distributed simulation of aircraft engines as part of a computer aided design package being developed by NASA Lewis Research Center for the aircraft industry. The project is called NPSS, an acronym for "Numerical Propulsion Simulation System". NPSS is a flexible object-oriented simulation of aircraft engines requiring high computing speed. It is desirable to run the simulation on a distributed computer system with multiple processors executing portions of the simulation in parallel. The purpose of this research was to investigate object-oriented structures such that individual objects could be distributed. The set of classes used in the simulation must be designed to facilitate parallel computation. Since the portions of the simulation carried out in parallel are not independent of one another, there is the need for communication among the parallel executing processors which in turn implies need for their synchronization. Communication and synchronization can lead to decreased throughput as parallel processors wait for data or synchronization signals from other processors. As a result of this research, the following have been accomplished. The design and implementation of a set of simulation classes which result in a distributed simulation control program have been completed. The design is based upon MIT "Actor" model of a concurrent object and uses "connectors" to structure dynamic connections between simulation components. Connectors may be dynamically created according to the distribution of objects among machines at execution time without any programming changes. Measurements of the basic performance have been carried out with the result that communication overhead of the distributed design is swamped by the computation time of modules unless modules have very short execution times per iteration or time step. An analytical performance model based upon queuing network theory has been designed and implemented. Its application to realistic configurations has not been carried out.

Schoeffler, James D.↗

Hardware-Accelerated Ray Tracing of CAD-Based Geometry for Monte Carlo Radiation Transport

Monte Carlo radiation transport (MCRT) methods have been used to simulate radiation environments for many decades by tracking individual particles through a model to accumulate statistical information. MCRT geometry is historically formed using the constructive solid geometry (CSG). Recently, significant work has been performed to support simulations using computer-aided design (CAD)-based tessellated surfaces to support highly complex geometries. Ray tracing acceleration data structures from the rendering and visualization community are applied to accelerate particle tracking in CAD-based models. Despite these efforts, CSG representations provide the superior performance in surface intersection operations during particle flight. Concurrently, pseudo Monte Carlo methods have become prevalent in rendering applications to support more realistic models for scattering media, motivating innovations that are advantageous for MCRT simulations. Finally, the authors’ work extends these innovations by employing Intel’s Embree ray tracing kernel within a geometry toolkit for Monte Carlo to improve the simulation performance using CAD-based models by factors of 1.5 to 2.

42 ENGINEERING↗