Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Cooperative Data Sharing: Simple Support for Clusters of SMP Nodes

Libraries like PVM and MPI send typed messages to allow for heterogeneous cluster computing. Lower-level libraries, such as GAM, provide more efficient access to communication by removing the need to copy messages between the interface and user space in some cases. still lower-level interfaces, such as UNET, get right down to the hardware level to provide maximum performance. However, these are all still interfaces for passing messages from one process to another, and have limited utility in a shared-memory environment, due primarily to the fact that message passing is just another term for copying. This drawback is made more pertinent by today's hybrid architectures (e.g. clusters of SMPs), where it is difficult to know beforehand whether two communicating processes will share memory. As a result, even portable language tools (like HPF compilers) must either map all interprocess communication, into message passing with the accompanying performance degradation in shared memory environments, or they must check each communication at run-time and implement the shared-memory case separately for efficiency. Cooperative Data Sharing (CDS) is a single user-level API which abstracts all communication between processes into the sharing and access coordination of memory regions, in a model which might be described as "distributed shared messages" or "large-grain distributed shared memory". As a result, the user programs to a simple latency-tolerant abstract communication specification which can be mapped efficiently to either a shared-memory or message-passing based run-time system, depending upon the available architecture. Unlike some distributed shared memory interfaces, the user still has complete control over the assignment of data to processors, the forwarding of data to its next likely destination, and the queuing of data until it is needed, so even the relatively high latency present in clusters can be accomodated. CDS does not require special use of an MMU, which can add overhead to some DSM systems, and does not require an SPMD programming model. unlike some message-passing interfaces, CDS allows the user to implement efficient demand-driven applications where processes must "fight" over data, and does not perform copying if processes share memory and do not attempt concurrent writes. CDS also supports heterogeneous computing, dynamic process creation, handlers, and a very simple thread-arbitration mechanism. Additional support for array subsections is currently being considered. The CDS1 API, which forms the kernel of CDS, is built primarily upon only 2 communication primitives, one process initiation primitive, and some data translation (and marshalling) routines, memory allocation routines, and priority control routines. The entire current collection of 28 routines provides enough functionality to implement most (or all) of MPI 1 and 2, which has a much larger interface consisting of hundreds of routines. still, the API is small enough to consider integrating into standard os interfaces for handling inter-process communication in a network-independent way. This approach would also help to solve many of the problems plaguing other higher-level standards such as MPI and PVM which must, in some cases, "play OS" to adequately address progress and process control issues. The CDS2 API, a higher level of interface roughly equivalent in functionality to MPI and to be built entirely upon CDS1, is still being designed. It is intended to add support for the equivalent of communicators, reduction and other collective operations, process topologies, additional support for process creation, and some automatic memory management. CDS2 will not exactly match MPI, because the copy-free semantics of communication from CDS1 will be supported. CDS2 application programs will be free to carefully also use CDS1. CDS1 has been implemented on networks of workstations running unmodified Unix-based operating systems, using UDP/IP and vendor-supplied high- performance locks. Although its inter-node performance is currently unimpressive due to rudimentary implementation technique, it even now outperforms highly-optimized MPI implementation on intra-node communication due to its support for non-copy communication. The similarity of the CDS1 architecture to that of other projects such as UNET and TRAP suggests that the inter-node performance can be increased significantly to surpass MPI or PVM, and it may be possible to migrate some of its functionality to communication controllers.

DiNucci, David C.↗

Scheduling Operations for Massive Heterogeneous Clusters

High-performance computing (HPC) programming has become increasingly difficult with the advent of hybrid supercomputers consisting of multicore CPUs and accelerator boards such as the GPU. Manual tuning of software to achieve high performance on this type of machine has been performed by programmers. This is needlessly difficult and prone to being invalidated by new hardware, new software, or changes in the underlying code. A system was developed for task-based representation of programs, which when coupled with a scheduler and runtime system, allows for many benefits, including higher performance and utilization of computational resources, easier programming and porting, and adaptations of code during runtime. The system consists of a method of representing computer algorithms as a series of data-dependent tasks. The series forms a graph, which can be scheduled for execution on many nodes of a supercomputer efficiently by a computer algorithm. The schedule is executed by a dispatch component, which is tailored to understand all of the hardware types that may be available within the system. The scheduler is informed by a cluster mapping tool, which generates a topology of available resources and their strengths and communication costs. Software is decoupled from its hardware, which aids in porting to future architectures. A computer algorithm schedules all operations, which for systems of high complexity (i.e., most NASA codes), cannot be performed optimally by a human. The system aids in reducing repetitive code, such as communication code, and aids in the reduction of redundant code across projects. It adds new features to code automatically, such as recovering from a lost node or the ability to modify the code while running. In this project, the innovators at the time of this reporting intend to develop two distinct technologies that build upon each other and both of which serve as building blocks for more efficient HPC usage. First is the scheduling and dynamic execution framework, and the second is scalable linear algebra libraries that are built directly on the former.

Humphrey, John↗

Modular droplet injector for sample conservation providing new structural insight for the conformational heterogeneity in the disease-associated NQO1 enzyme

Droplet injection strategies are a promising tool to reduce the large amount of sample consumed in serial femtosecond crystallography (SFX) measurements at X-ray free electron lasers (XFELs) with continuous injection approaches. Here, we demonstrate a new modular microfluidic droplet injector (MDI) design that was successfully applied to deliver microcrystals of the human NAD(P)H:quinone oxidoreductase 1 (NQO1) and phycocyanin. We investigated droplet generation conditions through electrical stimulation for both protein samples and implemented hardware and software components for optimized crystal injection at the Macromolecular Femtosecond Crystallography (MFX) instrument at the Stanford Linac Coherent Light Source (LCLS). Under optimized droplet injection conditions, we demonstrate that up to 4-fold sample consumption savings can be achieved with the droplet injector. In addition, we collected a full data set with droplet injection for NQO1 protein crystals with a resolution up to 2.7 Å, leading to the first room-temperature structure of NQO1 at an XFEL. NQO1 is a flavoenzyme associated with cancer, Alzheimer's and Parkinson's disease, making it an attractive target for drug discovery. Further, our results reveal for the first time that residues Tyr128 and Phe232, which play key roles in the function of the protein, show an unexpected conformational heterogeneity at room temperature within the crystals. These results suggest that different substates exist in the conformational ensemble of NQO1 with functional and mechanistic implications for the enzyme's negative cooperativity through a conformational selection mechanism. Our study thus demonstrates that microfluidic droplet injection constitutes a robust sample-conserving injection method for SFX studies on protein crystals that are difficult to obtain in amounts necessary for continuous injection, including the large sample quantities required for time-resolved mix-and-inject studies.

59 BASIC BIOLOGICAL SCIENCES↗

TRITON: A Multi-GPU open source 2D hydrodynamic flood model

A new open source multi-GPU 2D flood model called TRITON is presented in this work. The model solves the 2D shallow water equations with source terms using a time-explicit first order upwind scheme based on an Augmented Roe's solver that incorporates a careful estimation of bed strengths and a local implicit formulation of friction terms. Here, the scheme is demonstrated to be first order accurate, robust and able to solve for flows under various conditions. TRITON is implemented such that the model effectively utilizes heterogeneous architectures, from single to multiple CPUs and GPUs. Different test cases are shown to illustrate the capabilities and performance of the model, showing promising runtimes for large spatial and temporal scales when leveraging the computer power of GPUs. Under this hardware configuration, communication and input/output subroutines may impact the scalability. The code is developed under an open source license and can be freely downloaded in https://code.ornl.gov/hydro/triton.

2D flood model↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Programming model for distributed intelligent systems

A programming model and architecture which was developed for the design and implementation of complex, heterogeneous measurement and control systems is described. The Multigraph Architecture integrates artificial intelligence techniques with conventional software technologies, offers a unified framework for distributed and shared memory based parallel computational models and supports multiple programming paradigms. The system can be implemented on different hardware architectures and can be adapted to strongly different applications.

Sztipanovits, J.↗

A Backend-agnostic, Quantum-classical Framework for Simulations of Chemistry in C ++

As quantum computing hardware systems continue to advance, the research and development of performant, scalable, and extensible software architectures, languages, models, and compilers is equally as important to bring this novel coprocessing capability to a diverse group of domain computational scientists. For the field of quantum chemistry, applications and frameworks exist for modeling and simulation tasks that scale on heterogeneous classical architectures, and we envision the need for similar frameworks on heterogeneous quantum-classical platforms. Furthermore, we present the XACC system-level quantum computing framework as a platform for prototyping, developing, and deploying quantum-classical software that specifically targets chemistry applications. We review the fundamental design features in XACC, with special attention to its extensibility and modularity for key quantum programming workflow interfaces and provide an overview of the interfaces most relevant to simulations of chemistry. A series of examples demonstrating some of the state-of-the-art chemistry algorithms currently implemented in XACC are presented, while also illustrating the various APIs that would enable the community to extend, modify, and devise new algorithms and applications in the realm of chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TRIDENT Drill Validation at Mars and Lunar Analog Field Sites

Drilling on Earth is typically a human-intensive activity. Drilling on other planets is further complicated by the lack of prior local field surveys of their target area, hence blindly drilling into uncertain target rocks. Field conditions on the Moon or Mars are also different than for shallow drilling on Earth: lower temperatures and pressures, less power available, low masses (hence less weight-on-bit). Given the cost of transport from Earth, no drilling muds or working fluids are likely to be available to carry away cuttings. And impact-gardened regolith and dust vary mechanically and texturally from most terrestrial soils. The Regolith and Ice Drill for Exploration of New Terrains (TRIDENT) is a rotary-percussive 1m-class drill from Honeybee Robotics. It is low-power (rotary and percussive actuators are 200 W each) and lightweight (<20 kg) with the maximum weight on bit limited to 200 N. TRIDENT has been manifested for the Volatiles Investigating Polar Exploration Rover (VIPER) and PRIME-1 lunar south pole missions in 2024, has previously been field tested at a hot, dry analog site in the Atacama Desert, and in lunar conditions in thermal vacuum chamber tests. TRIDENT was also part of the 2019 Icebreaker Mars Discovery proposal, as well as in the Mars Life Explorer concept. During ARADS tests, drill control and fault recovery automation software enabled hands-off operations of a rover-mounted TRIDENT drill. TRIDENT Drill Analog Site Validation: Past TRIDENT tests in thermal vacuum (TVAC) chambers targeted containers of manufactured lunar simulants with added volatiles. 2022 TRIDENT ambient testing at NASA Ames drilled into cemented lunar simulant materials. Low cuttings-permeability led to cuttings buildup, and drill choking and binding was observed. The Atacama analog site in ARADS had desiccated unconsolidated sediments that did not challenge the TRIDENT design. However, lunar polar regolith is expected to be diverse and heterogeneous with varying clast sizes, with abundant impactites and perhaps subsurface ice deposits. Neither the simulants nor Atacama testing had completely covered the TRIDENT-targeted field characteristics, motivating further analog tests prior to the planned lunar missions. To gain more insight into the behavior of the TRIDENT hardware in diverse impactites and subsurface ice, and to verify the software automation in that environment, in August 2023 TRIDENT was brought to Haughton Crater, a field analog site in the Canadian Arctic. In September 2023 the same drill was brought to the Bishop Tuff in southern California to verify whether drilling binding behaviors previously seen in lunar simulant testing would be observed in naturally occurring fine-grained massive layers. The ~22 Ma Haughton Crater impact structure is located at 75 ̊22’ N, 89 ̊41’ W, on northwestern Devon Island, Nunavut, Canada. Numerous deposits of pale-grey crater-fill polymictic impact-melt breccia are found within the crater with a typical thickness reaching ~125 m or greater and covering ~60 km -2 . An approximately 600m-thick permafrost layer is also present with ice typically found within 0.5-0.6m of the surface. The volcanic tableland north of Bishop, CA exposes densely welded tuff laid down during the eruption that created the Long Valley Caldera at approximately 0.76 Ma. Extensional faults and the Owens River gorge expose cross-sections across the plateau. The area is viewed as an analog site for Mars features believed to be of pyroclastic origin. Results: Haughton Crater.Drilling tests were conducted 8-13 August 2023 at a previously undisturbed area separated 5-10 m from past years’ Drill Hill test sites (75.4208, -089.7613). In six days, TRIDENT drilled 8 holes to nearly 1 m depth each, totaling 7.80 m. The active layer/ice boundary was at ~67 cm depth, with a total of approximately 2.4 m drilled into ice or ice-cemented impact breccia. During drilling, five drill fault states were observed and successfully recovered. Holes 23-1, 4 and 7 were drilled under manual control, using Honeybee’s Thorax user interface. Holes 23-5, 6, and 8 were drilled with the Ames IBexec automated drilling control software. TRIDENT was observed to have little difficulty in the thawed uncemented impact breccia above the active layer boundary, but required percussion to make slower headway in the ice-cemented breccia. In Hole 23-7 (Fig.1), drilling slowed down in a massive unit just above the active-layer boundary (perhaps a large rock extending into the ice-cementation?), with only 7cm progress made in 27 minutes of high auger torque and constant percussion, leading to a choking fault and then a binding fault. A similar pattern had been observed in TRIDENT Rio Tinto test data from 2017 [6] as well as in the 2022laboratory tests.Bishop Tuff.A team from NASA Ames and the US Geological Survey deployed the same TRIDENT drill to Bishop Sites 1B and 1C (37.4203, -118.4289; 37.4265, -118.4215) on 13-16 September 2023, on the Bishop Tuff plateau. A third drill site was used 17-18 September 2023(37.4598, -118.3667) in an abandoned pumice mine. Four holes (totaling 2.5m depth) were drilled into the fine-grained, meters-thick tuff units at Sites 1B and 1C, and a further two boreholes (totaling 1.98m depth) were made at the pumice site. Drill behavior in the tuff below 10 cm depth was similar to that seen at 65-74 cm depth in Haughton Hole 23-7 (Fig. 1) and that seen in the 2022 lab simulant drilling. Drill safety torque limits were exceeded multiple times resulting in drill stops downhole. These freezes then required external added torques (with a pipe wrench) to resume rotation, to unstick the drill for withdrawal. To prevent this choking and binding behavior we found that more-frequent cuttings removal was necessary, e.g., reducing the “drill bite” size from the nominal 10 cm to 2 cm per bite --bringing the auger up to the surface more frequently, as seen in Bishop Site 1C Hole 2 (Fig.2). This permitted slow progress without drill binding and without external interventions. Conversely, TRIDENT drilling in the more porous pumice target material showed no cuttings buildup issue, and single bites as large as 40 cm were demonstrated. Discussion: We observed that TRIDENT easily penetrated unconsolidated heterogeneous soils (both above the active layer boundary at Haughton and previously in the Atacama). Cemented or consolidated targets that were cuttings-permeable (icy impact breccia, pumice) required more energy applied and percussion. However, in non-cuttings-permeable targets (welded microporous tuff, cemented simulants, boulder) TRIDENT was observed to be prone to excessive cuttings accumulation leading to choking/binding faults and stalling. The wedge cutting bit, used by TRIDENT in field tests and in its flight versions, pulverizes the target rock and creates fine cuttings that ideally are transported up the auger spirals for removal. In porous, fractured or vesicular target materials (such as at the Bishop pumice site) a significant portion of the cuttings are pushed aside, but for non-fractured, microporous targets the cuttings remain in the borehole and accumulate. Rock powder is relatively incompressible as a working fluid at only 100-200N downward force (TRIDENT limits) and hence eventual drilling progress slows or stops. Our recommended strategy for improving TRIDENT cuttings removal in massive target units with low cuttings-permeability is to reduce TRIDENT bite sizes when encountering these units, from 10cm to as little as 1-2cm, to effectively bail the accumulating cuttings. This approach was demonstrated to reduce choking and allowed slow progress to continue in cuttings-impermeable microporous target units (viz. the Bishop tuff in our September 2023 tests or cemented simulants in 2022 ambient tests).

robotic drilling↗

Extending C++ for Heterogeneous Quantum-Classical Computing

In this report we present qcor - a language extension to C++ and compiler implementation that enables heterogeneous quantum-classical programming, compilation, and execution in a single-source context. Our work provides a first-of-its-kind C++ compiler enabling high-level quantum kernel (function) expression in a quantum-language agnostic manner, as well as a hardware-agnostic, retargetable compiler workflow targeting a number of physical and virtual quantum computing backends. qcor leverages novel Clang plugin interfaces and builds upon the XACC system-level quantum programming framework to provide a state-of-the-art integration mechanism for quantum-classical compilation that leverages the best from the community at-large. qcor translates quantum kernels ultimately to the XACC intermediate representation, and provides user-extensible hooks for quantum compilation routines like circuit optimization, analysis, and placement. This work details the overall architecture and compiler workflow for qcor, and provides a number of illuminating programming examples demonstrating its utility for near-term variational tasks, quantum algorithm expression, and feed-forward error correction schemes.

97 MATHEMATICS AND COMPUTING↗

HEP - A semaphore-synchronized multiprocessor with central control

The paper describes the design concept of the Heterogeneous Element Processor (HEP), a system tailored to the special needs of scientific simulation. In order to achieve high-speed computation required by simulation, HEP features a hierarchy of processes executing in parallel on a number of processors, with synchronization being largely accomplished by hardware. A full-empty-reserve scheme of synchronization is realized by zero-one-valued hardware semaphores. A typical system has, besides the control computer and the scheduler, an algebraic module, a memory module, a first-in first-out (FIFO) module, an integrator module, and an I/O module. The architecture of the scheduler and the algebraic module is examined in detail.

Gilliland, M. C.↗

Preliminary Study on Fine-Grained Power and Energy Measurements on Grace Hopper GH200 with Open-Source Performance Tools

The increasing adoption of tightly integrated, heterogeneous architectures, combined with the slowdown of Moore’s law, has made application power and energy-driven optimizations critical to efficiently use high-performance computing systems. This paper introduces a newly developed open-source toolkit that seamlessly integrates the Linux real-time hardware monitoring program hwmon with the Performance Application Programming Interface and the Score-P performance measurement system, thereby enabling fine-grained power and energy measurements for high-performance computing applications. Our primary target platform is the Wombat test bed, which is a system based on the NVIDIA GH200 superchip. The toolkit can capture transient power peaks with high temporal resolution (50 ms) and, thanks to Score-P integration, can map power metrics to specific code regions, thereby providing actionable information on power-intensive operations and inefficiencies. The toolkit also provides a holistic view of both the power and the energy consumption of the entire GH200 superchip by covering all major components: the Grace CPU, the Hopper GPU, and the I/O subsystem. Experiments that use Locally Self-consistent Multiple Scattering, which is an application for first-principles calculations of materials developed at Oak Ridge National Laboratory, have demonstrated the tool’s ability to identify transient power spikes and uncover opportunities for energy-aware optimizations. Additionally, we introduce a Python-based utility for converting Open Trace Format 2 traces to Parquet format, thus enabling advanced data analysis for numerical integration methods applied to power data for accurate energy profiling.

Hernandez Mendoza, Oscar [ORNL] (ORCID:00000002538↗

Software Defined Architectures for Portability and Performance

The Software Defined Architectures for Portability and Performance (SODAPOP) project developed a co-design framework to partition and map converged applications on specialized heterogeneous architectures. We started from key domain applications that combine scientific simulation with data analytics and machine learning as drivers to integrate our framework. The framework includes high-level compilers that interfaces with high-level programming frameworks, domain-specific optimization passes, and hardware-oriented optimizations. The framework leverages hardware generators to enable specialization and facilitate exploration of custom system designs.

97 MATHEMATICS AND COMPUTING↗

Performance Potential of Mixed Data Management Modes for Heterogeneous Memory Systems

Many high-performance systems now include different types of memory devices within the same compute platform to meet strict performance and cost constraints. Such heterogeneous memory systems often include an upper-level tier with better performance, but limited capacity, and lower-level tiers with higher capacity, but less bandwidth and longer latencies for reads and writes. To utilize the different memory layers efficiently, current systems rely on hardware-directed, memory -side caching or they provide facilities in the operating system (OS) that allow applications to make their own data-tier assignments. Since these data management options each come with their own set of trade-offs, many systems also include mixed data management configurations that allow applications to employ hardware- and software-directed management simultaneously, but for different portions of their address space. Despite the opportunity to address limitations of stand-alone data management options, such mixed management modes are under-utilized in practice, and have not been evaluated in prior studies of complex memory hardware. In this work, we develop custom program profiling, configurations, and policies to study the potential of mixed data management modes to outperform hardware- or software-based management schemes alone. Our experiments, conducted on an Intel ® Knights Landing platform with high-bandwidth memory, demonstrate that the mixed data management mode achieves the same or better performance than the best stand-alone option for five memory intensive benchmark applications (run separately and in isolation), resulting in an average speedup compared to the best stand-alone policy of over 10 %, on average.

Effler, Chad↗

A Hierarchical Task Scheduler for Heterogeneous Computing

Heterogeneous computing is one of the future directions of HPC. Task scheduling in heterogeneous computing must balance the challenge of optimizing the application performance and the need for an intuitive interface with the programming run-time to maintain programming portability. The challenge is further compounded by the varying data communication time between tasks. This paper proposes RANGER, a hardware-assisted task-scheduling framework. By integrating RISC-V cores with accelerators, the RANGER scheduling framework divides scheduling into global and local levels. At the local level, RANGER further partitions each task into fine-grained subtasks to reduce the overall makespan. At the global level, RANGER maintains the coarse granularity of the task specification, thereby maintaining programming portability. The extensive experimental results demonstrate that RANGER achieves a 12.7× performance improvement on average, while only requires 2.7% of area overhead.

Miniskar, Narasinga Rao↗

Reimagining Codesign for Advanced Scientific Computing: Report for the ASCR Workshop on Reimagining Codesign

In March 2021, the U.S. Department of Energy’s Advanced Scientific Computing Research program convened the Workshop on Reimagining Codesign. The workshop, also known as ReCoDe, was organized around discussions on eight topic areas: (1) codesign for traditional high-performance computing workloads; (2) codesign of memory/storage systems; (3) codesign of machine learning, neuromorphic, quantum, and other non-von Neumann accelerators; (4) codesign for edge computing and processing at experimental instruments; (5) codesign for security and privacy; (6) hardware design tools and open-source hardware for high-productivity codesign; (7) tools, software stack, and programming languages for high-productivity codesign; and (8) quantitative tools and data collection for modeling and simulation for codesign. The panels identified four Priority Research Directions from these deliberations: (1) breakthrough computing capabilities with targeted heterogeneity and rapid design; (2) software and applications that embrace radical architecture diversity; (3) engineered security and integrity, from transistors to applications; and (4) design with data-rich processes.

97 MATHEMATICS AND COMPUTING↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

X-composer: enabling cross-environments in-situ workflows between HPC and cloud

As large-scale scientific simulations and big data analyses become more popular, it is increasingly more expensive to store huge amounts of raw simulation results to perform post-analysis. To minimize the expensive data I/O, "in-situ" analysis is a promising approach, where data analysis applications analyze the simulation generated data on the fly without storing it first. However, it is challenging to organize, transform, and transport data at scales between two semantically different ecosystems due to the distinct software and hardware difference. To tackle these challenges, we design and implement the X-Composer framework. X-Composer connects cross-ecosystem applications to form an "in-situ" scientific workflow, and provides a unified approach and recipe for supporting such hybrid in-situ workflows on distributed heterogeneous resources. X-Composer reorganizes simulation data as continuous data streams and feeds them seamlessly into the Cloud-based stream processing services to minimize I/O overheads. For evaluation, we use X-Composer to set up and execute a cross-ecosystem workflow, which consists of a parallel Computational Fluid Dynamics simulation running on HPC, and a distributed Dynamic Mode Decomposition analysis application running on Cloud. Our experimental results show that X-Composer can seamlessly couple HPC and Big Data jobs in their own native environments, achieve good scalability, and provide high-fidelity analytics for ongoing simulations in real-time.

Wang, Dali↗

Hybrid classical-quantum communication networks

Over the past several decades, the proliferation of global classical communication networks has transformed various facets of human society. Concurrently, quantum networking has emerged as a dynamic field of research, driven by its potential applications in distributed quantum computing, quantum sensor networks, and secure communications. This prompts a fundamental question: rather than constructing quantum networks from scratch, can we harness the widely available classical fiber-optic infrastructure to establish hybrid quantum–classical networks? This paper aims to provide a comprehensive review of ongoing research endeavors aimed at integrating quantum communication protocols, such as quantum key distribution, into existing lightwave networks. This approach offers the substantial advantage of reducing implementation costs by allowing classical and quantum communication protocols to share optical fibers, communication hardware, and other network control resources—arguably the most pragmatic solution in the near term. In the long run, classical communication will also reap the rewards of innovative quantum communication technologies, such as quantum memories and repeaters. Accordingly, our vision for the future of the Internet is that of heterogeneous communication networks thoughtfully designed for the seamless support of both classical and quantum communications.

Fiber-optic communication↗