Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data movements”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

Enhancement of Otolith Specific Ocular Responses Using Vestibular Stochastic Resonance

Introduction: Astronauts experience disturbances in sensorimotor function after spaceflight during the initial introduction to a gravitational environment, especially after long-duration missions. Our goal is to develop a countermeasure based on vestibular stochastic resonance (SR) that could improve central interpretation of vestibular input and mitigate these risks. SR is a mechanism by which noise can assist and enhance the response of neural systems to relevant, imperceptible sensory signals. We have previously shown that imperceptible electrical stimulation of the vestibular system enhances balance performance while standing on an unstable surface. Methods: Eye movement data were collected from 10 subjects during variable radius centrifugation (VRC). Subjects performed 11 trials of VRC that provided equivalent tilt stimuli from otolith and other graviceptor input without the normal concordant canal cues. Bipolar stochastic electrical stimulation, in the range of 0-1500 microamperes, was applied to the vestibular system using a constant current stimulator through electrodes placed over the mastoid process behind the ears. In the VRC paradigm, subjects were accelerated to 216 deg./s. After the subjects no longer sensed rotation, the chair oscillated along a track at 0.1 Hz to provide tilt stimuli of 10 deg. Eye movements were recorded for 6 cycles while subjects fixated on a target in darkness. Ocular counter roll (OCR) movement was calculated from the eye movement data during periods of chair oscillations. Results: Preliminary analysis of the data revealed that 9 of 10 subjects showed an average increase of 28% in the magnitude of OCR responses to the equivalent tilt stimuli while experiencing vestibular SR. The signal amplitude at which performance was maximized was in the range of 100-900 microamperes. Discussion: These results indicate that stochastic electrical stimulation of the vestibular system can improve otolith specific responses. This will have a significant impact on development of vestibular SR delivery systems to aid recovery of function in astronauts after long-duration spaceflight or in people with balance disorders.

Fiedler, Matthew↗

A Web of Data Analytics Services

Cloud Computing has become the ubiquitous approach to our Big Data challenge. However, one will quickly discover that moving (a.k.a. forklifting) existing on-premise data analytics solutions to the Cloud doesn’t always translate to costing saving and performance boost. The Cloud’s elasticity, its availability, and its wide selection of computing options and selections of costing models making Cloud an attractive environment to tackle our Big Data challenge. The fact is Cloud, on its own, is not the silver bullet to our daunting challenge need for analyze and derive scientific inferences through vast collections of multi-sensor measurements. We would like to have all scientific data in one easy to access environment, but getting the world of scientific data in one analytic system is immensely difficult to achieve. This paper describes the data analytics web architecture NASA is developing by infusing instances of Integrated Data Analytics systems next to the data. The goal is to minimize unnecessary data movement through collection of data access and analytics webservices for researchers to interact with and analyze measurements without have to download data to their local computer. These services are RESTful and provisioned by the data centers with the help from subject matter and science experts. These services encapsulate the physical computing infrastructure, which could local computing cluster, on-premise or public Cloud environment.

Huang, Thomas↗

Productive Programming of Distributed Systems with the SHAD C++ Library

High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.

Castellana, Vito G.↗

Data Federation Challenges in Remote Near-Real-Time Fusion Experiment Data Processing

Fusion energy experiments and simulations provide critical information needed to plan future fusion reactors. As next-generation devices like ITER move toward long-pulse experiments, analyses, including AI and ML, should be performed in a wide range of time and computing constraints, from near-real-time constraints, between-shot analysis, and to campaign-wide long-term analysis. However, the data volume, velocity, and variety make it extremely challenging for analyses using only local computational resources. Researchers need the ability to compose and execute workflows spanning edge resources to large-scale high-performance computing facilities.We present Delta, a system to address data analysis challenges, including AI/ML, in fusion science, by leveraging the ADIOS I/O library and middleware, to support executing science workflows over the wide area network for near-real-time streaming. We discuss the data federation challenges in performing remote workflows, focusing on on-going research work in (1) managing, reducing, and streaming data to minimize I/O and data movement overheads, (2) decompressing and reorganizing data for analysis, and (3) executing workflows for automated data analysis. We introduce examples for deep-learning based data analysis for the fusion domain and demonstrate how we use Delta to construct end-to-end workflows for a fusion device in Korea, connecting a remote DOE facility in the USA. The capability demonstrated by this project is the basis for improving the state of the art for near-real-time data federation amongst remote facilities.

Choi, Jong Youl↗

[The Strategic Organization of Skill]

Eye-movement software was developed in addition to several studies that focused on expert-novice differences in the acquisition and organization of skill. These studies focused on how increasingly complex strategies utilize and incorporate visual look-ahead to calibrate action. Software for collecting, calibrating, and scoring eye-movements was refined and updated. Some new algorithms were developed for analyzing corneal-reflection eye movement data that detect the location of saccadic eye movements in space and time. Two full-scale studies were carried out which examined how experts use foveal and peripheral vision to acquire information about upcoming environmental circumstances in order to plan future action(s) accordingly.

Roberts, Ralph↗

Enhanced Video-Oculography System

A previously developed video-oculography system has been enhanced for use in measuring vestibulo-ocular reflexes of a human subject in a centrifuge, motor vehicle, or other setting. The system as previously developed included a lightweight digital video camera mounted on goggles. The left eye was illuminated by an infrared light-emitting diode via a dichroic mirror, and the camera captured images of the left eye in infrared light. To extract eye-movement data, the digitized video images were processed by software running in a laptop computer. Eye movements were calibrated by having the subject view a target pattern, fixed with respect to the subject s head, generated by a goggle-mounted laser with a diffraction grating. The system as enhanced includes a second camera for imaging the scene from the subject s perspective, and two inertial measurement units (IMUs) for measuring linear accelerations and rates of rotation for computing head movements. One IMU is mounted on the goggles, the other on the centrifuge or vehicle frame. All eye-movement and head-motion data are time-stamped. In addition, the subject s point of regard is superimposed on each scene image to enable analysis of patterns of gaze in real time.

Moore, Steven T.↗

Can Tracked Great Frigatebirds (Fregata minor) be Used to Measure the Dynamics of the Planetary Boundary Layer Height?

Observing the dynamics of the planetary boundary layer (PBL) will require creative combinations of space-based and in-situ measurements. Animals fitted with biologgers have been used to gather data and parameterize physical models in various settings and may be a useful addition to the suite of space-based measurements that NASA is considering. For example, great frigatebirds (Fregata minor) generally fly below 700m but have been shown to routinely climb up to 2,000 meters during the day, occasionally reaching heights over 4,000 meters. This suggests that they may be tracking boundary layer height dynamics as they soar and glide on air currents with little energetic cost. We show here that the high-altitude climbing heights of great frigatebirds carrying GPS/accelerometer biologgers (e-obs GmbH Bird Solar 15g) at Palmrya Atoll, Pacific Ocean, are consistent with the long-term climatology (2006-2019) of the local planetary boundary layer height, and that these soaring heights can be reliably extracted from the tag data. Additional study is needed to determine how well movement data from tracked birds can constrain PBL height, but this work demonstrates the potential for animal-borne sensors to complement space-based measurements. Biologging data were contributed by the USGS and the Nature Conservancy's Palmyra Bluewater Research (PBR) study and analyzed as part of NASA’s Internet of Animals project.

Planetary Boundary Layer Height↗

Accelerating Advanced Light Source Science Through Multi-Facility HPC Workflows

Synchrotron light sources support a wide array of techniques to investigate materials, often producing complex, high-volume data that challenge traditional workflows. At the Advanced Light Source (ALS), we developed infrastructure to move microtomography data over ESnet to ALCF and NERSC, where CPU- and GPU-based algorithms generate 3D reconstructed volumes of experimental samples. We employ two data movement and reconstruction models: real-time processing as data streams directly to NERSC compute nodes, and automated file transfer to NERSC and ALCF file systems. The streaming pipeline provides users with feedback in under ten seconds, while the file-based workflow produces high-quality reconstructions suitable for deeper analysis in 20-30 minutes. This infrastructure enables users to utilize HPC resources without direct access to backend systems. We plan to extend this architecture to more endstations, supporting our beamline scientists and users.

Abramov, David↗

Polyamines as Possible Modulators of Gravity-induced Calcium Transport in Plants

Data from various laboratories indicate a probable relationship between calcium movement and some aspects of graviperception and tropistic bending responses. The movement of calcium in response to gravistimulation appears to be rapid, polar and opposite in direction to polar auxin transport. What might be the cause of such rapid Ca(2+) movement? Data from studies on polyamine (PA) metabolism may furnish a clue. A transient increase in the activity of ornithine decarboxylase (ODC) and titers of various PAs occurs within 60 seconds after hormonal stimulation of animal cells, followed by Ca(2+) transport out of the cells. Through the use of specific inhibitors, it was shown that the enhanced PA synthesis from ODC was essential not only for Ca(2+) transport, but also for Ca(2+) transport-dependent endocytosis and the movement of hexoses and amino acids across the plasmalemma. In plants, rapid changes in arginine decarboxylase (ADC) activity occur in response to various plant stresses. Physical stresses associated with gravisensor displacement and reorientation of a plant in the gravitational field could similarly activate ADC and that resultant increases in PA levels might initiate transient perturbations in Ca(2+) homeostasis.

Galston, A. W.↗

Analog Computing for Science

Conventional digital computing faces fundamental physical limits: large scale computing systems already con sume tens of Megawatts of power, Dennard scaling has ended, and data movement costs dominate application performance. Next generation experimental facilities generate data at rates that overwhelm conventional pro cessing and demand real-time analysis at the source. Analog computing, which exploits the continuous dynamics of physical systems to perform computation, promises a transformative path toward orders-of-magnitude gains in energy efficiency and time-to-solution for scientific workloads.

97 MATHEMATICS AND COMPUTING↗

Compiler and Runtime Approaches to Enable Large-Scale Irregular Programs. Final report, July 2013 - July 2019

While regular algorithms, characterized by operations on dense matrices and arrays, have long been the mainstay of scientific, high-performance computing, irregular algorithms, which feature unpredictable accesses to pointer-based data structures, are becoming increasingly common in high performance computing, arising in graph analysis, data mining and visualization, among other domains. Unfortunately, the defining characteristics of irregular applications, their dynamic, unpredictable, data-dependent access patterns and data layouts, make achieving high performance on large scale systems difficult. Scaling applications to peta- and exa-scale requires carefully controlling communication and data movement and placement, an inherently difficult task when access patterns and data layouts are unpredictable! Most irregular applications that attain high performance must be painstakingly hand-written and hand-tuned, with few common principles or paradigms uniting various implementations and easing future development. Despite the increasing importance of irregular applications, there is little programmer knowledge, and even less compiler ability, devoted to optimizing them. This project aims to solve these problems. By allowing programmers to write irregular applications in high level forms, with at most a few annotations highlighting key structural properties, programmers can focus on developing their algorithms and methods. The compiler and run-time system can take on the tedious task of optimizing the application for execution at large scales, and can automatically provide efficient implementations. This will provide portability and ease maintenance for existing irregular applications, but, more importantly, open up whole new domains of computational science to large-scale, high-performance simulation codes.

97 MATHEMATICS AND COMPUTING↗

Vector-Matrix Multiplication Engine for Neuromorphic Computation with a CBRAM Crossbar Array [Slides]

The core function of many neural network algorithms is the dot product, or vector matrix multiply (VMM) operation. Crossbar arrays utilizing resistive memory elements can reduce computational energy in neural algorithms by up to five orders of magnitude compared to conventional CPUs. Moving data between a processor, SRAM, and DRAM dominates energy consumption. By utilizing analog operations to reduce data movement, resistive memory crossbars can enable processing of large amounts of data at lower energy than conventional memory architectures.

97 MATHEMATICS AND COMPUTING↗

A single user efficiency measure for evaluation of parallel or pipeline computer architectures

A precise statement of the relationship between sequential computation at one rate, parallel or pipeline computation at a much higher rate, the data movement rate between levels of memory, the fraction of inherently sequential operations or data that must be processed sequentially, the fraction of data to be moved that cannot be overlapped with computation, and the relative computational complexity of the algorithms for the two processes, scalar and vector, was developed. The relationship should be applied to the multirate processes that obtain in the employment of various new or proposed computer architectures for computational aerodynamics. The relationship, an efficiency measure that the single user of the computer system perceives, argues strongly in favor of separating scalar and vector processes, sometimes referred to as loosely coupled processes, to achieve optimum use of hardware.

Jones, W. P.↗

An Integrated Indexing and Search Service for Distributed File Systems

Data services such as search, discovery, and management in scalable distributed environments have traditionally been decoupled from the underlying file systems, and are often deployed using external databases and indexing services. However, modern data production rates, looming data movement costs, and the lack of metadata, entail revisiting the decoupled file system-data services design philosophy. In this article, we present TagIt, a scalable data management service framework aimed at scientific datasets, which can be integrated into prevalent distributed file system architectures. A key feature of TagIt is a scalable, distributed metadata indexing framework, which facilitates a flexible tagging capability to support data discovery. Furthermore, the tags can also be associated with an active operator, for pre-processing, filtering, or automatic metadata extraction, which we seamlessly offload to file servers in a load-aware fashion. We have integrated TagIt into two popular distributed file systems, i.e., GlusterFS and CephFS. Our evaluation demonstrates that TagIt can expedite data search operation by up to 10× over the extant decoupled approach.

97 MATHEMATICS AND COMPUTING↗

Performance Measurement, Visualization and Modeling of Parallel and Distributed Programs

This paper presents a methodology for debugging the performance of message-passing programs on both tightly coupled and loosely coupled distributed-memory machines. The AIMS (Automated Instrumentation and Monitoring System) toolkit, a suite of software tools for measurement and analysis of performance, is introduced and its application illustrated using several benchmark programs drawn from the field of computational fluid dynamics. AIMS includes (i) Xinstrument, a powerful source-code instrumentor, which supports both Fortran77 and C as well as a number of different message-passing libraries including Intel's NX Thinking Machines' CMMD, and PVM; (ii) Monitor, a library of timestamping and trace -collection routines that run on supercomputers (such as Intel's iPSC/860, Delta, and Paragon and Thinking Machines' CM5) as well as on networks of workstations (including Convex Cluster and SparcStations connected by a LAN); (iii) Visualization Kernel, a trace-animation facility that supports source-code clickback, simultaneous visualization of computation and communication patterns, as well as analysis of data movements; (iv) Statistics Kernel, an advanced profiling facility, that associates a variety of performance data with various syntactic components of a parallel program; (v) Index Kernel, a diagnostic tool that helps pinpoint performance bottlenecks through the use of abstract indices; (vi) Modeling Kernel, a facility for automated modeling of message-passing programs that supports both simulation -based and analytical approaches to performance prediction and scalability analysis; (vii) Intrusion Compensator, a utility for recovering true performance from observed performance by removing the overheads of monitoring and their effects on the communication pattern of the program; and (viii) Compatibility Tools, that convert AIMS-generated traces into formats used by other performance-visualization tools, such as ParaGraph, Pablo, and certain AVS/Explorer modules.

Yan, Jerry C.↗