Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Priority-BF: A Task Manager for Priority-Based Scheduling

The increasing demand for computational resources, particularly in High-Performance Computing environments, necessitates to rethink how we handle job scheduling strategies. This work addresses the challenge of managing concurrent jobs with differing priorities on overloaded parallel systems, where strict QoS constraints are often difficult for users to define. Our solution relies on a qualitative description of priorities and pulls from two key approaches: the Easy-BF algorithm and the Conservative Backfilling algorithms. This solution improves the response time for high-priority jobs by 50% without affecting the overall system utilization. We show its applicability in several critical scenarios such as High-Performance Computing (HPC) resource management and in-situ computing.

Gainaru, Ana [ORNL]↗

Storage media pipelining: Making good use of fine-grained media

This paper proposes a new high-performance paradigm for accessing removable media such as tapes and especially magneto-optical disks. In high-performance computing the striping of data across multiple devices is a common means of improving data transfer rates. Striping has been used very successfully for fixed magnetic disks improving overall system reliability as well as throughput. It has also been proposed as a solution for providing improved bandwidth for tape and magneto-optical subsystems. However, striping of removable media has shortcomings, particularly in the areas of latency to data and restricted system configurations, and is suitable primarily for very large I/Os. We propose that for fine-grained media, an alternative access method, media pipelining, may be used to provide high bandwidth for large requests while retaining the flexibility to support concurrent small requests and different system configurations. Its principal drawback is high buffering requirements in the host computer or file server. This paper discusses the possible organization of such a system including the hardware conditions under which it may be effective, and the flexibility of configuration. Its expected performance is discussed under varying workloads including large single I/O's and numerous smaller ones. Finally, a specific system incorporating a high-transfer-rate magneto-optical disk drive and autochanger is discussed.

Vanmeter, Rodney↗

Multithreaded Model for Dynamic Load Balancing Parallel Adaptive PDE Computations

We present a multithreaded model for the dynamic load-balancing of numerical, adaptive computations required for the solution of Partial Differential Equations (PDE's) on multiprocessors. Multithreading is used as a means of exploring concurrency in the processor level in order to tolerate synchronization costs inherent to traditional (non-threaded) parallel adaptive PDE solvers. Our preliminary analysis for parallel, adaptive PDE solvers indicates that multithreading can be used an a mechanism to mask overheads required for the dynamic balancing of processor workloads with computations required for the actual numerical solution of the PDE's. Also, multithreading can simplify the implementation of dynamic load-balancing algorithms, a task that is very difficult for traditional data parallel adaptive PDE computations. Unfortunately, multithreading does not always simplify program complexity, often makes code re-usability not an easy task, and increases software complexity.

Chrisochoides, Nikos↗

Abisko: Deep codesign of an architecture for spiking neural networks using novel neuromorphic materials

The Abisko project aims to develop an energy-efficient spiking neural network (SNN) computing architecture and software system capable of autonomous learning and operation. The SNN architecture explores novel neuromorphic devices that are based on resistive-switching materials, such as memristors and electrochemical RAM. Equally important, Abisko uses a deep codesign approach to pursue this goal by engaging experts from across the entire range of disciplines: materials, devices and circuits, architectures and integration, software, and algorithms. Here, the key objectives of our Abisko project are threefold. First, we are designing an energy-optimized high-performance neuromorphic accelerator based on SNNs. This architecture is being designed as a chiplet that can be deployed in contemporary computer architectures and we are investigating novel neuromorphic materials to improve its design. Second, we are concurrently developing a productive software stack for the neuromorphic accelerator that will also be portable to other architectures, such as field-programmable gate arrays and GPUs. Third, we are creating a new deep codesign methodology and framework for developing clear interfaces, requirements, and metrics between each level of abstraction to enable the system design to be explored and implemented interchangeably with execution, measurement, a model, or simulation. As a motivating application for this codesign effort, we target the use of SNNs for an analog event detector for a high-energy physics sensor.

97 MATHEMATICS AND COMPUTING↗

Optimizing multigrid reduction-in-time and Parareal coarse-grid operators for linear advection

Parallel-in-time methods, such as multigrid reduction-in-time (MGRIT) and Parareal, provide an attractive option for increasing concurrency when simulating time-dependent partial differential equations (PDEs) in modern high-performance computing environments. While these techniques have been very successful for parabolic equations, it has often been observed that their performance suffers dramatically when applied to advection-dominated problems or purely hyperbolic PDEs using standard rediscretization approaches on coarse grids. In this paper, we apply MGRIT or Parareal to the constant-coefficient linear advection equation, appealing to existing convergence theory to provide insight into the typically nonscalable or even divergent behavior of these solvers for this problem. To overcome these failings, we replace rediscretization on coarse grids with improved coarse-grid operators that are computed by applying optimization techniques to approximately minimize error estimates from the convergence theory. Therefore, one of our main findings is that, in order to obtain fast convergence as for parabolic problems, coarse-grid operators should take into account the behavior of the hyperbolic problem by tracking the characteristic curves. Our approach is tested for schemes of various orders using explicit or implicit Runge–Kutta methods combined with upwind-finite-difference spatial discretizations. In all cases, we obtain scalable convergence in just a handful of iterations, with parallel tests also showing significant speed-ups over sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

Using heterogeneous GPU nodes with a Cabana-based implementation of MPCD

In this study, the Kokkos based library Cabana, which has been developed in the Co-design Center for Particle Applications (CoPA), is used for the implementation of Multi-Particle Collision Dynamics (MPCD), a particle-based description of hydrodynamic interactions. Cabana allows for a function portable implementation, which has been used to study the interplay between CPU and GPU usage on a multi-node system as well as analysis of said interplay with performance analysis tools. As a result, we see most advantages in a homogeneous GPU usage, but we also discuss the extent to which heterogeneous applications might be more performant, using both CPU and GPU concurrently.

97 MATHEMATICS AND COMPUTING↗

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O↗

Library for Evolutionary Algorithms in Python (LEAP)

There are generally three types of scientific software users: users that solve problems using existing science software tools, researchers that explore new approaches by extending existing code, and educators that teach students scientific concepts. Python is a general-purpose programming language that is accessible to beginners, such as students, but also as a language that has a rich scientific programming ecosystem that facilitates writing research software. Additionally, as high-performance computing (HPC) resources become more readily available, software support for parallel processing becomes more relevant to scientific software.There currently are no Python-based evolutionary computation frameworks that support all three types of scientific software users. Moreover, some support synchronous concurrent fitness evaluation that do not efficiently use HPC resources. We pose here a new Python-based EC framework that uses an established generalized unified approach to EA concepts to provide an easy to use toolkit for users wishing to use an EA to solve a problem, for researchers to implement novel approaches, and for providing a low-bar to entry to EA concepts for students. Additionally, this toolkit provides a scalable asynchronous fitness evaluation implementation friendly to HPC that has been vetted on hardware ranging from laptops to the world’s fastest supercomputer, Summit.

Coletti, Mark↗

Community reaction to aircraft noise around smaller city airports

The results are presented of a study of community reaction to jet aircraft noise in the vicinity of airports in Chattanooga, Tennessee, and Reno, Nevada. These cities were surveyed in order to obtain data for comparison with that obtained in larger cities during a previous study. (The cities studied earlier were Boston, Chicago, Dallas, Denver, Los Angeles, Miami, and New York.) The purpose of the present effort was to observe the relative reaction under conditions of lower noise exposure and in less highly urbanized areas, and to test the previously developed predictive equation for annoyance under such circumstances. In Chattanooga and Reno a total of 1960 personal interviews based upon questionnaires were obtained. Aircraft noise measurements were made concurrently and aircraft operations logs were maintained for several weeks in each city to permit computation of noise exposures. The survey respondents were chosen randomly from various exposure zones.

Connor, W. K.↗

Total ozone variations 1970-74 using Backscattered Ultraviolet /BUV/ and ground-based observations

The most long-lived satellite set of ozone observations, to date, is that derived from the Backscatter Ultraviolet (BUV) ozone sensor on Nimbus 4 and extends from April 1970 through 1976. Unfortunately, this experiment suffered spacecraft power limitations which limited the spatial and temporal coverage and also appears to have suffered from long-term drifts which may be associated with changes in the instrument characteristics or the incident solar flux. Techniques have been developed to account for these problems, and this paper presents results of the BUV total ozone variations and compares them with those from ground-based observations, specifically the computations of Angell and Korshover (1978). After adjustments for the spatial gaps and comparison with concurrent Dobson ground-based observations, no significant trend was found in the BUV data over the years 1970-74. This finding is in contrast to a general decrease of about 2% during the same period appearing in the data of Angell and Korshover. The difference in these results is discussed in terms of the geographic sampling and the methods of hemispheric integration.

Miller, A. J.↗

Cooperating knowledge-based systems

This final report covers work performed under Contract NCC2-220 between NASA Ames Research Center and the Knowledge Systems Laboratory, Stanford University. The period of research was from March 1, 1987 to February 29, 1988. Topics covered were as follows: (1) concurrent architectures for knowledge-based systems; (2) methods for the solution of geometric constraint satisfaction problems, and (3) reasoning under uncertainty. The research in concurrent architectures was co-funded by DARPA, as part of that agency's Strategic Computing Program. The research has been in progress since 1985, under DARPA and NASA sponsorship. The research in geometric constraint satisfaction has been done in the context of a particular application, that of determining the 3-D structure of complex protein molecules, using the constraints inferred from NMR measurements.

Feigenbaum, Edward A.↗

xSDK: Building an ecosystem of highly efficient math libraries for exascale

Current efforts to build increasingly powerful computer architectures are opening up new avenues for more complex and higher fidelity simulations coupled with data analytics and learning, leading to new scientific insights and deeper understanding. At one extreme, exascale computers will be much faster than previous computer generations (performing 10 18 operations per second—that is, 1,000 times faster than petascale). To achieve these performance improvements, computer architectures are becoming increasingly complex, with deep memory hierarchies, very high node and core counts, and heterogeneous features such as graphics processing units (GPUs). Such architectural changes impact the full breadth of computing scales, as heterogeneity pervades even current-generation laptops, workstations, and moderate-sized clusters. While emerging advanced architectures provide unprecedented opportunities, they also present significant challenges for developers of scientific applications, such as multiphysics and multiscale codes, who must adapt their software to handle disruptive changes in architectures and new programming models that have not yet stabilized. Developers must consider increasing concurrency while reducing communication and synchronization, and other complexities such as the potential for using mixed precision to leverage the compute power available in low-precision tensor cores. On one hand, developers must implement new scientific capabilities, which in turn increase code complexity. On the other hand, the codes must be ported to new architectures, requiring the inclusion of new programming models and the restructuring of code to achieve good performance. Addressing these issues is beyond the capability of any single person or team—leading to the need for collaboration among many teams, who encapsulate their expertise in reusable software and work together to create sustainable software ecosystems.

97 MATHEMATICS AND COMPUTING↗

A sample computation of kinematic properties from cloud motion vectors.

Distributions of relative vorticity and balanced height have been computed from the cloud velocities associated with the cloud structure of an extratropical cyclone over the continental United States during a three-day period in March 1970. Cloud motions are assigned either to a 'mid-level,' or to a 'high level.' Derived vorticity and balanced height are compared with concurrent National Meteorological Center (NMC) analyses and also with similar kinematic quantities obtained from rawins at three constant-pressure levels. The computations of relative vorticity using mid-level cloud motion vectors show encouraging results. Patterns of computed cyclonic vorticity are related to the development, location, and movement of the surface cyclone. The analyses suggest that the 'mid-level' corresponds best to the 700-mb level. The vorticity analysis from the 'high-level' motion vectors presented difficulties.

Viezee, W.↗

Recent trends in digital human modeling and the concurrent issues that face human modeling approach

Tremendous strides have been made in the recent years to digitally represent human beings in computer simulation models ranging from assembly plant maintenance operations to occupants getting in and out of vehicles to action movie scenarios. While some of these tools are being actively pursued by the engineering communities, there is still a lot of work that remains to be done for the newly planned planetary exploration missions. For example, certain unique and several common challenges are seen in developing computer generated suited human models for designing the next generation space vehicle. The purpose of this presentation is to discuss NASA s potential needs for better human models and to show also many of the inherent yet not too obvious pitfalls that still are left unresolved in this new arena of digital human modeling. As part of NASA s Habitability and Human Factors Branch, the Anthropometry and Biomechanics Facility has been engaged in studying the various facets of computer generated human physical performance models; for instance, it has been engaged in utilizing three-dimensional laser scan data along with three dimensional video based motion and reach data to gather suited anthropometric and shape and size information that are not available yet in the form of computer mannequins. Our goal is to bring in new approaches to deal with heavily clothed humans (such as, suited astronauts) and to overcome the current limitations of wrongly identifying humans (either real or virtual) as univariate percentiles. We are looking at whole-body posture based anthropometric models as a means to identify humans of significantly different shapes and sizes to arrive at mathematically sound computer models for analytical purposes.

Rajulu, Sudhakar↗

Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases

A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.

97 MATHEMATICS AND COMPUTING↗

Efficient human activity recognition with spatio-temporal spiking neural networks

In this study, we explore Human Activity Recognition (HAR), a task that aims to predict individuals' daily activities utilizing time series data obtained from wearable sensors for health-related applications. Although recent research has predominantly employed end-to-end Artificial Neural Networks (ANNs) for feature extraction and classification in HAR, these approaches impose a substantial computational load on wearable devices and exhibit limitations in temporal feature extraction due to their activation functions. To address these challenges, we propose the application of Spiking Neural Networks (SNNs), an architecture inspired by the characteristics of biological neurons, to HAR tasks. SNNs accumulate input activation as presynaptic potential charges and generate a binary spike upon surpassing a predetermined threshold. This unique property facilitates spatio-temporal feature extraction and confers the advantage of low-power computation attributable to binary spikes. We conduct rigorous experiments on three distinct HAR datasets using SNNs, demonstrating that our approach attains competitive or superior performance relative to ANNs, while concurrently reducing energy consumption by up to 94%.

60 APPLIED LIFE SCIENCES↗

Tomography of entangling two-qubit logic operations in exchange-coupled donor electron spin qubits

Scalable quantum processors require high-fidelity universal quantum logic operations in a manufacturable physical platform. Donors in silicon provide atomic size, excellent quantum coherence and compatibility with standard semiconductor processing, but no entanglement between donor-bound electron spins has been demonstrated to date. Here we present the experimental demonstration and tomography of universal one- and two-qubit gates in a system of two weakly exchange-coupled electrons, bound to single phosphorus donors introduced in silicon by ion implantation. We observe that the exchange interaction has no effect on the qubit coherence. We quantify the fidelity of the quantum operations using gate set tomography (GST), and we use the universal gate set to create entangled Bell states of the electrons spins, with fidelity 91.3 ± 3.0%, and concurrence 0.87 ± 0.05. These results form the necessary basis for scaling up donor-based quantum computers.

42 ENGINEERING↗