Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

List Models of Procedure Learning

This paper presents a new theory of the initial stages of skill acquisition and then employs the theory to model current and future training programs for fight management systems (FMSs) in modern commercial airliners like the Boeing 777 and the Airbus A320. The theoretical foundations for the theory are a new synthesis of the literature on human memory and the latest version of the ACT-R theory of skill acquisition.

Matessa, Michael P.↗

Network-Capable Application Process and Wireless Intelligent Sensors for ISHM

Intelligent sensor technology and systems are increasingly becoming attractive means to serve as frameworks for intelligent rocket test facilities with embedded intelligent sensor elements, distributed data acquisition elements, and onboard data acquisition elements. Networked intelligent processors enable users and systems integrators to automatically configure their measurement automation systems for analog sensors. NASA and leading sensor vendors are working together to apply the IEEE 1451 standard for adding plug-and-play capabilities for wireless analog transducers through the use of a Transducer Electronic Data Sheet (TEDS) in order to simplify sensor setup, use, and maintenance, to automatically obtain calibration data, and to eliminate manual data entry and error. A TEDS contains the critical information needed by an instrument or measurement system to identify, characterize, interface, and properly use the signal from an analog sensor. A TEDS is deployed for a sensor in one of two ways. First, the TEDS can reside in embedded, nonvolatile memory (typically flash memory) within the intelligent processor. Second, a virtual TEDS can exist as a separate file, downloadable from the Internet. This concept of virtual TEDS extends the benefits of the standardized TEDS to legacy sensors and applications where the embedded memory is not available. An HTML-based user interface provides a visual tool to interface with those distributed sensors that a TEDS is associated with, to automate the sensor management process. Implementing and deploying the IEEE 1451.1-based Network-Capable Application Process (NCAP) can achieve support for intelligent process in Integrated Systems Health Management (ISHM) for the purpose of monitoring, detection of anomalies, diagnosis of causes of anomalies, prediction of future anomalies, mitigation to maintain operability, and integrated awareness of system health by the operator. It can also support local data collection and storage. This invention enables wide-area sensing and employs numerous globally distributed sensing devices that observe the physical world through the existing sensor network. This innovation enables distributed storage, distributed processing, distributed intelligence, and the availability of DiaK (Data, Information, and Knowledge) to any element as needed. It also enables the simultaneous execution of multiple processes, and represents models that contribute to the determination of the condition and health of each element in the system. The NCAP (intelligent process) can configure data-collection and filtering processes in reaction to sensed data, allowing it to decide when and how to adapt collection and processing with regard to sophisticated analysis of data derived from multiple sensors. The user will be able to view the sensing device network as a single unit that supports a high-level query language. Each query would be able to operate over data collected from across the global sensor network just as a search query encompasses millions of Web pages. The sensor web can preserve ubiquitous information access between the querier and the queried data. Pervasive monitoring of the physical world raises significant data and privacy concerns. This innovation enables different authorities to control portions of the sensing infrastructure, and sensor service authors may wish to compose services across authority boundaries.

Figueroa, Fernando↗

Computational Model of Human and System Dynamics in Free Flight: Studies in Distributed Control Technologies

This paper presents a set of studies in full mission simulation and the development of a predictive computational model of human performance in control of complex airspace operations. NASA and the FAA have initiated programs of research and development to provide flight crew, airline operations and air traffic managers with automation aids to increase capacity in en route and terminal area to support the goals of safe, flexible, predictable and efficient operations. In support of these developments, we present a computational model to aid design that includes representation of multiple cognitive agents (both human operators and intelligent aiding systems). The demands of air traffic management require representation of many intelligent agents sharing world-models, coordinating action/intention, and scheduling goals and actions in a potentially unpredictable world of operations. The operator-model structure includes attention functions, action priority, and situation assessment. The cognitive model has been expanded to include working memory operations including retrieval from long-term store, and interference. The operator's activity structures have been developed to provide for anticipation (knowledge of the intention and action of remote operators), and to respond to failures of the system and other operators in the system in situation-specific paradigms. System stability and operator actions can be predicted by using the model. The model's predictive accuracy was verified using the full-mission simulation data of commercial flight deck operations with advanced air traffic management techniques.

Corker, Kevin M.↗

Indicator-directed Dynamic Power Management for Iterative Workloads on GPU-Accelerated Systems

Modern high-performance and warehouse computing centers show strong interest in minimizing system power consumption while satisfying customers’ quality of service (QoS). Dynamic voltage and frequency scaling (DVFS) is effective for achieving this goal. Nevertheless, automating the process online and making it transparent to users must address three major challenges: (1) Complexity — today’s hardware components (e.g., CPUs, GPUs, memory, network, etc.) can be configured in several or dozens of frequency/voltage states for satisfying divergent system demands. Given their combination and the emergence of heterogeneity, searching the optimal configuration in the design space online can be timing consuming. (2) QoS guarantee — user-defined objectives such as power constraint and performance target must be monitored, predicted and ensured at the best effort. (3) Adaptability — various known and unknown workloads run on systems. Workloads characteristics should be quickly determined and configurations dynamically adjusted in accord with workloads and QoS. In this work, we focus on applications exhibiting an interesting feature – iterative or periodic, which is common among conventional HPC and emerging machine learning workloads. We propose an online dynamic power-performance (ODPP) management framework to dynamically adjust GPU DVFS configurations to meet performance and power objectives and constraints, without any code annotation or intrusion. Particularly, ODPP extracts the performance and power indicators for applications from their resources utilization profiles in a short episode. It further automatically constructs an accurate model that infers from the indicators how the application's performance and power vary with GPU core and memory frequencies. Aided with the model, for both seen and unseen applications, ODPP can quickly determine the most appropriate DVFS configuration for their execution. We evaluate ODPP on an NVIDIA GPU using multiple exascale computing (ECP) and deep learning applications.

Zou, Pengfei↗

Multivariable degradation modeling and life prediction using multivariate fractional Brownian motion

In system prognostics and health management, multivariable degradation models have been widely developed to predict the life of complex systems using degradation data of multiple Performance Characteristics (PCs). Recent studies have detected a Long-Term Memory (LTM) effect among the degradation process of various PCs, implying a strong coupling phenomenon between the future degradation behavior and historical degradation trajectory. Although the LTM has been widely integrated into single-PC-based degradation modeling, it has not been considered in multi-PC-based scenarios. To capture LTM among multiple PCs, this article proposes a novel LTM-integrated Multivariate Degradation Model (MDM) for system life prediction based on multivariate fractional Brownian motion, which simultaneously incorporates the cross-correlation among different PCs. To estimate parameters of the LTM-integrated MDM, a maximum likelihood method is developed. Here, two likelihood-ratio hypothesis tests are developed to test the existence of the overall and individual LTM effect among multiple PCs. Both simulation studies and physical experiments on the performance degradation of solar energy conversion and storage devices are conducted to validate the proposed model. Results reveal that the proposed LTM-integrated MDM significantly outperforms existing MDMs in life prediction, while the lifetime uncertainty is heavily underestimated by those traditional approaches that neglect the LTM.

42 ENGINEERING↗

Hod Carrier

SAND2023-06720O Hod Carrier is a simple proof-of-concept library for moving data from host to NVIDIA BlueField device memory using remote direct memory access. The Hod Carrier software also facilitates research activities related to the potential uses of NVIDIA BlueField data processing unit devices. Using well-established technologies (e.g., RDMA over Infiniband), it will move data from the host to the data processing unit’s memory. This software is a middleware utility for moving data without regard to its semantic meaning. Its operation is fundamentally unaffected by storage allocation. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Twin-page storage management for rapid transaction-undo recovery

This paper presents and evaluates a new twin-page disk-storage management scheme for rapid database transaction-undo recovery. In contrast to previous twin-page schemes, the present approach uses static page mapping and allows dirty pages in the main memory to be written, at any instant, onto disk without the requirement of undo logging. No explicit undo is required when a transaction is aborted. Transaction undo is implicitly performed by not subsequently fetching from disk the invalid pages updated by the aborted transaction. Performance in terms of disk I/O and CPU overhead for transaction-undo recovery is analyzed and compared with a previous approach TWIST. It is shown that the scheme achieves rapid transaction-undo recovery without degrading average system performance for various workloads, and that the scheme is well suited for applications with a large number of updates and frequent transaction aborts.

Wu, Kun-Lung↗

Orbit Operations at 433 Eros: Navigation for the NEAR Shoemaker Mission

NASA's Near Earth Asteroid Rendezvous Mission began its record-setting exploration of the asteroid 433 Eros by inserting the spacecraft into orbit about Eros on February 14, 2000. This is the first spacecraft from any country to orbit an asteroid. The mission has overcome a failed insertion burn attempt on December 20, 1998, an event that would have ended most planetary missions, to return to the same target and successfully begin its science mapping a little more than a year later. Shortly after the successful insertion into orbit, the mission was renamed NEAR Shoemaker (NEAR) in memory of the late astronomer and geologist Eugene Shoemaker. NEAR will gather science data at Eros until February 14, 2001, which is the nominal end of mission. The NEAR mission is managed by the Johns Hopkins University, Applied Physics Laboratory in Laurel, Maryland. Since the initial mission concept in 1992, the design and implementation of the NEAR navigation system have been the responsibility of the Jet Propulsion Laboratory, California Institute of Technology. This presentation will show some of the unique features of navigation and mission design related to orbiting an asteroid and to designing a robust navigation system for the NEAR spacecraft. The problem of navigating a spacecraft about an asteroid is made difficult by the relative uncertainty in the asteroid physical properties which perturb the orbit: i.e., the mass, gravity field, and spin state. To help solve this problem, the navigation system for NEAR uses traditional DSN radio metric Doppler and range tracking, along with new technologies of optical landmark tracking and laser ranging to the asteroid surface. The experiences to date for each of these data types in the navigation solutions will be presented. Plans for the remainder of the NEAR mission will be presented, which include low orbits (down to 35 km radius circular orbits), and close flybys that may pass within 1 km of the surface. In addition, at the end of mission, NASA has approved a controlled descent and hovering phase that will culminate with the spacecraft impacting the surface. The maneuver planning for this final phase will also be presented.

Williams, B. G.↗

BioSentinel: Mission Development of a Radiation Biosensor to Gauge DNA Damage and Repair Beyond Low Earth Orbit on a 6U Nanosatellite.

We are designing and developing a "6U" (10 x 22 x 34 cm; 14 kg) nanosatellite as a secondary payload to fly aboard NASA's Space Launch System (SLS) Exploration Mission (EM) 1, scheduled for launch in late 2017. For the first time in over forty years, direct experimental data from biological studies beyond low Earth orbit (LEO) will be obtained during BioSentinel's 12- to 18- month mission. BioSentinel will measure the damage and repair of DNA in a biological organism and allow us to compare that to information from onboard physical radiation sensors. In order to understand the relative contributions of the space environment's two dominant biological perturbations, reduced gravity and ionizing radiation, results from deep space will be directly compared to data obtained in LEO (on ISS) and on Earth. These data points will be available for validation of existing biological radiation damage and repair models, and for extrapolation to humans, to assist in mitigating risks during future long-term exploration missions beyond LEO. The BioSentinel Payload occupies 4U of the spacecraft and will utilize the monocellular eukaryotic organism Saccharomyces cerevisiae (yeast) to report DNA double-strand-break (DSB) events that result from ambient space radiation. DSB repair exhibits striking conservation of repair proteins from yeast to humans. Yeast was selected because of 1) its similarity to cells in higher organisms, 2) the well-established history of strains engineered to measure DSB repair, 3) its spaceflight heritage, and 4) the wealth of available ground and flight reference data. The S. cerevisiae flight strain will include engineered genetic defects to prevent growth and division until a radiation-induced DSB activates the yeast's DNA repair mechanisms. The triggered culture growth and metabolic activity directly indicate a DSB and its successful repair. The yeast will be carried in the dry state within the 1-atm P/L container in 18 separate fluidics cards with each card having 16 independent culture microwells, with integral microchannels and filters to supply nutrients and reagents, confine the yeast to the wells, and enable optical measurement. The measurement subsystem will monitor each subgroup of culture wells continuously for several weeks, optically tracking DSBtriggered cell growth and metabolism. BioSentinel will also include physical radiation sensors based on the TimePix sensor, as implemented by JSC's RadWorks group, which record individual radiation events including estimates of their linear-energytransfer (LET) values. Radiation-dose and LET data will be compared directly to the rate of DSB-and-repair events measured by the S. cerevisiae biosentinels. The spacecraft bus will operate in a deep space environment with functions that include command and data handling, communications, power generation (via deployable solar panels) and storage, and attitude determination-and-control system with micropropulsion. Development of the BioSentinel spacecraft will mature and prove multiple nanosatellite advances in order to function well beyond LEO: Communications from distances of ≥ 500,000 km; Autonomous attitude control, momentum management, and safe mode of nanosatellites in deep space; Shielding-, hardening-, design-, and software-derived radiation tolerance for electronics; Reliable functionality for 12 - 18 months of key subsystems for biofluidics, memory, communications, power, etc.; Close integration of living biological radiation event monitors with miniature physical radiation spectrometers; Biological measurement of solar particle events beyond Earth orbit In addition to providing the first biological results from beyond LEO in over 4 decades, BioSentinel will provide an adaptable small-satellite instrument platform to perform a range of human-exploration-relevant measurements that characterize the biological consequences of multiple outer space environments. BioSentinel is being developed under NASA's Advanced Exploration Systems program.

DNA damage↗

Continuous Test and Transition Infrastructure for Quantum Networking for Science Complex

We propose a network architecture with a separate quantum dataplane and a conventional control plane: (a) Quantum Data Plane: The Quantum Data plane consists of links of dark fibers connecting quantum switches and repeaters, which in turn, connect to quantum computers, memory and sensors. Since the current reach is limited to local areas, it will begin as a collection of site networks at laboratories, (b) Conventional Control Plane: The Control plane provides management access to quantum devices for configuration and provisioning via control nodes with firewall and encryption capabilities. Continuous Test and Transition Infrastructure: We propose an infrastructure with a control plane connecting multiple sites via encrypted tunnels over conventional networks consisting of the following: (a) Site Quantum Networks: Individual site networks supported by their fiber plants connect laboratories that house and connect to their quantum devices for testing and interoperability.(b) Site Control Planes: Sites are connected over individual control planes with control hosts with Software Defined Networking capabilities, which can be peered with other networks via ESnet.(c) Progressive Expansion: Initially, sites will develop individual data planes with their specific quantum devices and fiber connections, and mechanisms to interface with conventional networks. They progressively expand and interconnect, under a wide-area ecosystem of peered control planes for device testing, interoperability development and roll off into production environments.

Rao, Nageswara S.↗

Nanosecond Gated CMOS Camera (NSGCC) ICD (Rev. 2.1)

The Ultra-Fast X-ray Imager (UXI) program is an ongoing effort at Sandia National Laboratories to create high speed, multi-frame, time-gated Read Out Integrated Circuits (ROICs), and a corresponding suite of photodetectors to image a wide variety of High Energy Density (HED) physics experiments on both Sandia’s Z-Machine and LLNL’s National Ignition Facility (NIF). Several cameras have been designed over the length of the program; one of the most recent is the Icarus, which is an improvement on past imagers (Furi and Hippogriff). A second sensor that can be connected is the Daedalus sensor. The Icarus is a 1024 × 512-pixel array with either 25 μm or 8 µm spatial resolution containing four frames of storage per pixel and has improved timing generation and distribution components while achieved 2 ns time gating. The Daedalus sensor is also a 1024 x 512-pixel array with 25 µm special resolution containing three frames of storage per pixel and has an increased set of features for a wider variety of applications from interlacing of rows in each frame to configurability of all shutters. See Section 14 for details regarding the Icarus implementation of the firmware and Section 15 for details regarding the Daedalus implementation. Due to the unique test environments UXI sensors are targeted for, full custom hardware was required to physically mount an Icarus or Daedalus sensor, manage its various functions, and read out pixel data for transfer to a host computer. Beyond experimental functionality, the hardware also needed to accommodate sensor characterization requirements. Lawrence Livermore National Laboratory’s ‘Version 4.0 Board’ was the result of these efforts. It mounts all the components required to fully utilize the Icarus and Daedalus sensors including analog to digital converters to convert pixel data and various system voltages to digital form for readout and analysis, DAC channels for remote configuration of critical bias voltages, static random-access memories to buffer pixel data, RS422 and Gigabit Ethernet communications for remote access, and an FPGA to tie these components together. This document describes the FPGA electrical interfaces in detail to allow the reader a greater understanding of the device, and to facilitate implementation of custom software to control and manage it. The Version 4.0 Board is a continuation of the Nano-second Gated CMOS hardware design that retains much of the functionality of the Version 1.0 Board while adding features including a DAC instead of digital potentiometers, as well as sensors for pressure and radiation. The Version 4.0 board is intended for applications requiring tight form-factor enclosures. It is composed of two stacking boards; one holds the FPGA and regulators to power the various components of the board while the other contains the mating connector to the sensor, image-readoff ADCs, the DAC, and other components.

42 ENGINEERING↗

Efficient loading of reduced data ensembles produced at ORNL SNS/HFIR neutron time-of-flight facilities

We present algorithmic improvements to the loading operations of certain reduced data ensembles produced from neutron scattering experiments at Oak Ridge National Laboratory (ORNL) facilities. Ensembles from multiple measurements are required to cover a wide range of the phase space of a sample material of interest. They are stored using the standard NeXus schema on individual HDF5 files. This makes it a scalability challenge, as the number of experiments stored increases in a single ensemble file. The present work follows up on our previous efforts on data management algorithms, to address identified input output (I/O) bottlenecks in Mantid, an open-source data analysis framework used across several neutron science facilities around the world. We reuse an in-memory binary-tree metadata index that resembles data access patterns, to provide a scalable search and extraction mechanism. In addition, several memory operations are refactored and optimized for the current common use cases, ranging most frequently from 10 to 180, and up to 360 separate measurement configurations. Results from this work show consistent speed ups in wall-clock time on the Mantid LoadMD routine, ranging from 19% to 23% on average, on ORNL production computing systems. The latter depends on the complexity of the targeted instrument-specific data and the system I/O and compute variability for the shared computational resources available to users of ORNL’s Spallation Neutron Source (SNS) and the High Flux Isotope Reactor (HFIR) instruments. Nevertheless, we continue to highlight the need for more research to address reduction challenges as experimental data volumes, user time and processing costs increase.

Godoy, William↗

Nonlinear Fluid Computations in a Distributed Environment

The performance of a loosely and tightly-coupled workstation cluster is compared against a conventional vector supercomputer for the solution the Reynolds- averaged Navier-Stokes equations. The application geometries include a transonic airfoil, a tiltrotor wing/fuselage, and a wing/body/empennage/nacelle transport. Decomposition is of the manager-worker type, with solution of one grid zone per worker process coupled using the PVM message passing library. Task allocation is determined by grid size and processor speed, subject to available memory penalties. Each fluid zone is computed using an implicit diagonal scheme in an overset mesh framework, while relative body motion is accomplished using an additional worker process to re-establish grid communication.

Atwood, Christopher A.↗

Cognition in Space Workshop: Metrics and Models - 1

"Cognition in Space Workshop I: Metrics and Models" was the first in a series of workshops sponsored by NASA to develop an integrated research and development plan supporting human cognition in space exploration. The workshop was held in Chandler, Arizona, October 25-27, 2004. The participants represented academia, government agencies, and medical centers. This workshop addressed the following goal of the NASA Human System Integration Program for Exploration: to develop a program to manage risks due to human performance and human error, specifically ones tied to cognition. Risks range from catastrophic error to degradation of efficiency and failure to accomplish mission goals. Cognition itself includes memory, decision making, initiation of motor responses, sensation, and perception. Four subgoals were also defined at the workshop as follows: (1) NASA needs to develop a human-centered design process that incorporates standards for human cognition, human performance, and assessment of human interfaces; (2) NASA needs to identify and assess factors that increase risks associated with cognition; (3) NASA needs to predict risks associated with cognition; and (4) NASA needs to mitigate risk, both prior to actual missions and in real time. This report develops the material relating to these four subgoals.

Woolford, Barbara↗

NASA and COTS Electronics: Past Approach and Successes - Future Considerations

NASA has a long history of using commercial grade electronics in space. In this talk, a brief history of NASAâ's trends and approaches to commercial grade electronics focusing on processing and memory systems will be presented. This will include providing summary information on the space hazards to electronics as well as NASA mission trade space. We will also discuss developing recommendations for risk management approaches to Electrical, Electronic and Electromechanical (EEE) parts and reliability in space. The final portion of the talk will discuss emerging aerospace trends and the future for Commercial Off The Shelf (COTS) usage.

LaBel, Kenneth A.↗

High-Performance, Radiation-Hardened Electronics for Space Environments

The Radiation Hardened Electronics for Space Environments (RHESE) project endeavors to advance the current state-of-the-art in high-performance, radiation-hardened electronics and processors, ensuring successful performance of space systems required to operate within extreme radiation and temperature environments. Because RHESE is a project within the Exploration Technology Development Program (ETDP), RHESE's primary customers will be the human and robotic missions being developed by NASA's Exploration Systems Mission Directorate (ESMD) in partial fulfillment of the Vision for Space Exploration. Benefits are also anticipated for NASA's science missions to planetary and deep-space destinations. As a technology development effort, RHESE provides a broad-scoped, full spectrum of approaches to environmentally harden space electronics, including new materials, advanced design processes, reconfigurable hardware techniques, and software modeling of the radiation environment. The RHESE sub-project tasks are: SelfReconfigurable Electronics for Extreme Environments, Radiation Effects Predictive Modeling, Radiation Hardened Memory, Single Event Effects (SEE) Immune Reconfigurable Field Programmable Gate Array (FPGA) (SIRF), Radiation Hardening by Software, Radiation Hardened High Performance Processors (HPP), Reconfigurable Computing, Low Temperature Tolerant MEMS by Design, and Silicon-Germanium (SiGe) Integrated Electronics for Extreme Environments. These nine sub-project tasks are managed by technical leads as located across five different NASA field centers, including Ames Research Center, Goddard Space Flight Center, the Jet Propulsion Laboratory, Langley Research Center, and Marshall Space Flight Center. The overall RHESE integrated project management responsibility resides with NASA's Marshall Space Flight Center (MSFC). Initial technology development emphasis within RHESE focuses on the hardening of Field Programmable Gate Arrays (FPGA)s and Field Programmable Analog Arrays (FPAA)s for use in reconfigurable architectures. As these component/chip level technologies mature, the RHESE project emphasis shifts to focus on efforts encompassing total processor hardening techniques and board-level electronic reconfiguration techniques featuring spare and interface modularity. This phased approach to distributing emphasis between technology developments provides hardened FPGA/FPAAs for early mission infusion, then migrates to hardened, board-level, high speed processors with associated memory elements and high density storage for the longer duration missions encountered for Lunar Outpost and Mars Exploration occurring later in the Constellation schedule.

Keys, Andrew S.↗

The ParaScope parallel programming environment

The ParaScope parallel programming environment, developed to support scientific programming of shared-memory multiprocessors, includes a collection of tools that use global program analysis to help users develop and debug parallel programs. This paper focuses on ParaScope's compilation system, its parallel program editor, and its parallel debugging system. The compilation system extends the traditional single-procedure compiler by providing a mechanism for managing the compilation of complete programs. Thus, ParaScope can support both traditional single-procedure optimization and optimization across procedure boundaries. The ParaScope editor brings both compiler analysis and user expertise to bear on program parallelization. It assists the knowledgeable user by displaying and managing analysis and by providing a variety of interactive program transformations that are effective in exposing parallelism. The debugging system detects and reports timing-dependent errors, called data races, in execution of parallel programs. The system combines static analysis, program instrumentation, and run-time reporting to provide a mechanical system for isolating errors in parallel program executions. Finally, we describe a new project to extend ParaScope to support programming in FORTRAN D, a machine-independent parallel programming language intended for use with both distributed-memory and shared-memory parallel computers.

Cooper, Keith D.↗

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.

George, Anjus [ORNL] (ORCID:0000000179737061)↗