Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Extremely Scalable Distributed Computation of Contour Trees via Pre-Simplification

Contour trees offer an abstract representation of the level set topology in scalar fields and are widely used in topological data analysis and visualization. However, applying contour trees to large-scale scientific datasets remains challenging due to scalability limitations. Recent developments in distributed hierarchical contour trees have addressed these challenges by enabling scalable computation across distributed systems. Building on these structures, advanced analytical tasks—such as volumetric branch decomposition and contour extraction—have been introduced to facilitate large-scale scientific analysis. Despite these advancements, such analytical tasks substantially increase memory usage, which hampers scalability. In this paper, we propose a pre-simplification strategy to significantly reduce the memory overhead associated with analytical tasks on distributed hierarchical contour trees. We demonstrate enhanced scalability through strong scaling experiments, constructing the largest known contour tree—comprising over half a trillion nodes with complex topology—in under 15 minutes on a dataset containing 550 billion elements.

Li, Mingzhe [University of Utah]↗

UNITY: Unified Memory and Storage Space

UNITY is a 36-month project focused on providing design and evaluate a new distributed storage paradigm that unifies the traditionally distinct application views of memory- and file-based data storage into a single scalable and resilient environment. The project is a collaboration among Oak Ridge National Laboratory (ORNL, Lead institution), Los Alamos National Labs (LANL), and Georgia Tech (GT). The main contributions of the GT team have been around development of low level systems software for best leveraging the capabilities of new types of persistent memory technologies, and for development of methods for intelligent data management across memory/storage substrates with heterogeneous components. GT contributed the Phoenix library for optimized checkpoint/restart for HPC I/O for systems with non-volatile memory (NVM), the NVStream library for NVM-specialized streaming I/O for HPC workflows, the CoMerge, Mnemo and Kleio solutions for intelligent data management on NVM-based systems. These contributions result in significant improvements in both application performance and system efficiency.

97 MATHEMATICS AND COMPUTING↗

The Development of a Generalized Riser Flow Regime Map Based Upon Higher Moment and Chaotic Statistics Using Electrical Capacitance Volume Tomography (ECVT)

Dynamic analyses have been applied to the temporal signals from an Electro Capacitance Volume Tomography instrument located near mid-height on the riser of an industrial-scale cold-flow circulating fluidized bed to characterize gas-solids flow behavior in the riser. Twelve capacitance electrodes surround the cylindrical riser over a height of 1.3 m. The instrument used a neural network deconvolution algorithm to determine the spatially resolved solids fraction recorded at 52 Hz. Experiments were carried out over a range of gas and solids flows in the transport regime using a Geldart Group B bed material, high density polyethylene with mean particle size of 880 μm. The radial solids distribution was found to vary from one-time step to the next between profiles typical of laminar and turbulent flow. The duration of time spent in each of these flow profiles depended upon the operating regime – dilute, core-annular, or fast fluidized bed. The chaotic structure of the temporal data was characterized using the three conventional approaches: the first 4 moments from the distribution of signal in time, system memory parameters from the autocorrelation function and the Hurst exponent, and analysis of the correlationentropy and correlation dimension of the attractor. These signal analysis techniques were used to clearly distinguish differences between different transport operating regimes. Specifically, it was experimentally observed that a riser transitions from core annular flow profile to dilute and dense regimes via increasing the frequency of short term transients to either dilute or dense flow profiles, respectively. A regime map was generated based upon these dynamics using solids flux and gas velocity axes. Fast fluidized, core annular, and dilute each exhibited different degree of dynamic characteristics typical of fluid dominated or particle compromising behavior. It should be noted that the magnitude for the different statistics was in the same range regardless of the regime, it was the radial profile for the statistic that changed and subsequently identified that there was a change in the regime. Finally, a reduced regime map was developed consisting of plotting the gas velocity normalized by the upper transport velocity versus the solids flux normalized by the saturation carrying capacity. The use of this reduced plot allowed the data from widely different conditions to be plotted and compared on the same<p>graph. Note that in many instances, some of the statistics identified the operating point as being in one regime while others indicated that it was in another indicating a transition region between dilute or core annular regimes and between the core annular and fast fluidization regimes. This now provides a tool that can be used to optimize process performance, identify changes in operating states, or replicate process dynamics during process scaling or changing operating parameters. </p>

Breault, Ronald↗

Impact of Random Spatial Fluctuation in Non-Uniform Crystalline Phases on the Device Variation of Ferroelectric FET

In this work, a comprehensive study of random spatial fluctuation of the ferroelectric (FE) phase and dielectric (DE) phase in FeFETs is conducted to understand its impact on device variation. It is found that: i) there exists a certain DE percentage threshold that below which the increase of the DE phase does not significantly impact the device memory window and variation and only above which evident device degradation can be observed; ii) increasing the DE phase increases the variation in the memory window and the coercive field distribution further exacerbates the variation, hence degrading the sensing margin; iii) decreasing the number of grains degrades the device variation, which calls for further grain size engineering for variation suppression.

42 ENGINEERING↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗

Edge at the Pier: EPCAPE Software-Defined Sensing Field Campaign Report

The Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) was aimed to enhance the understanding of cloud and aerosol properties in the region surrounding La Jolla, California. To address challenges in data collection and processing from various instruments, an edge computing device known as Waggle Sage Node (WSN) was deployed at the Ellen Browning Scripps Memorial Pier. WSN is a distributed-sensing platform designed to collect and analyze environmental data at the edge. Sage is a multi-agency-supported project that designs and builds a new kind of national-scale reusable cyberinfrastructure to enable artificial intelligence (AI) at the edge based on the Waggle platform. Sponsors include the U.S. Department of Energy (DOE) Advanced Scientific Computing Research (ASCR), DOE National Nuclear Security Administration (NNSA), DOE Biological and Environmental Research (BER) through DOE Artificial Intelligence for Earth System Predictability (AI4ESP), Argonne Laboratory-Directed Research and Development (LDRD). Sage (https://sagecontinuum.org/) is funded as a National Science Foundation Mid-Scale Research Infrastructure (MSRI) project (https://www.nsf.gov/awardsearch/showAward?AWD_ID=1935984). This robust, multi-architecture edge computing platform facilitated environmental monitoring during the campaign. This report details the scientific objectives, deployment process, and key results of integrating Waggle into the EPCAPE field campaign.

54 ENVIRONMENTAL SCIENCES↗

An Entropy-Based Test and Development Framework for Uncertainty Modeling in Level-Set Visualizations

We present a simple comparative framework for testing and developing uncertainty modeling in uncertain marching cubes implementations. The selection of a model to represent the probability distribution of uncertain values directly influences the memory use, run time, and accuracy of an uncertainty visualization algorithm. We use an entropy calculation directly on ensemble data to establish an expected result and then compare the entropy from various probability models, including uniform, Gaussian, histogram, and quantile models. Our results verify that models matching the distribution of the ensemble indeed match the entropy. We further show that fewer bins in nonparametric histogram models are more effective whereas large numbers of bins in quantile models approach data accuracy.

Sisneros, Robert↗

Field-Deployable Quantum Memory for Quantum Networking

High-performance quantum memories are an essential component for regulating temporal events in quantum networks. As a component in quantum-repeaters, they have the potential to support the distribution of entanglement beyond the physical limitations of fiber loss. This will enable key applications such as quantum key distribution, network-enhanced quantum sensing, and distributed quantum computing. Here, we present a quantum memory engineered to meet real-world deployment and scaling challenges. The memory technology utilizes a warm rubidium vapor as the storage medium, and operates at room temperature, without the need for vacuum- and/or cryogenic- support. Here, we demonstrate performance specifications of high-fidelity retrieval (95%) and low operation error (10 –2 ) at a storage time of 160 µs for single-photon level quantum memory operations. We further show a substantially improved storage time (with classical-level light) of up to 1ms by suppressing atomic diffusions. The device is housed in an enclosure with a standard 2U rackmount form factor, and can robustly operate on a day scale in a noisy environment. This result marks an important step toward implementing quantum networks in the field.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Performance and usability enhancements for continuous subgraph matching queries on graph-structured data

A query graph, which includes vertices and edges, represents a query on graph-structured data. The query graph is decomposed into query subgraphs. A network analysis tool performs continuous subgraph matching queries to facilitate analysis of computer network traffic, social media events, or other streams of data represented as a dynamic data graph (graph-structured data). This can help identify emerging trends in the data. Some features of the network analysis tool enhance performance by effectively utilizing distributed computing resources (including processing cores and memory at different nodes of a cluster) to speed up the process of updating the dynamic data graph and detecting matches of query subgraphs. Features of a query graph building tool enhance usability by providing intuitive ways to specify query graphs and their subgraphs. Features of a results visualization tool enhance usability by providing an intuitive way to present the results of continuous subgraph matching queries.

Choudhury, Sutanay↗

Model reduction for a power grid model

We examine the complexity of constructing reduced order models for subsets of the variables needed to represent the state of the power grid. In particular, we apply model reduction techniques to the DeMarco-Zheng power grid model. We show that due to the oscillating nature of the solutions and the absence of timescale separation between resolved and unresolved variables, the construction of accurate reduced models becomes highly non-trivial because one has to account for long memory effects. In addition, we show that a reduced model that includes even a short memory is drastically better than a memoryless model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Tuning the Mechanical Properties of Shape Memory Metallic Glass Composites with Brick and Mortar Designs

We use molecular dynamics simulations to characterize the tensile deformation and failure of shape memory alloy-bulk metallic glass composites (BMGCs). As the failure in BMGCs shifts from the MG matrix to the second phase different strategies are required to further improve their mechanical properties. Results show that a staggered brick BMGC displays enhanced ductility with mild strength loss following a homogeneously distributed plasticity. This demonstrates that judiciously designed BMGCs can effectively synergize the mechanical properties of metallic glasses and shape memory alloys by effectively distributing stress and preventing the generation of critical shear bands and premature failure while preserving strength.

36 MATERIALS SCIENCE↗

Userspace Squash Filesystem for Launching Linux Containers on HPC Systems [Thesis]

The demand for user defined software stacks (UDSS) has been increasing in the high-performance computing (HPC) community. Container technology has become popular due to the flexibility and isolation it provides to HPC users. Container images must be available to all nodes involved for use in HPC and can be distributed to compute nodes in a variety of ways. A common method for container image distribution is to simply copy the container image to memory on each compute node, which can be time consuming at scale, and uses valuable memory on each node. The kernel mounted squash filesystem (squashfs) has proven fast and efficient for this task but requires root-level access. This paper will show a user space mounted squashfs is an efficient and secure solution for container image distribution.

97 MATHEMATICS AND COMPUTING↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Scalable low-loss cryogenic packaging of quantum memories in CMOS-foundry processed photonic chips

Optically linked solid-state quantum memories such as color centers in diamond are a promising platform for distributed quantum information processing and networking. Photonic integrated circuits (PICs) have emerged as a crucial enabling technology for these systems, integrating quantum memories with efficient electrical and optical interfaces in a compact and scalable platform. Packaging these hybrid chips into deployable modules while maintaining low optical loss and resiliency to temperature cycling is a central challenge to their practical use. We demonstrate a packaging method for PICs using surface grating couplers and angle-polished fiber arrays that is robust to temperature cycling, offers scalable channel count, applies to a wide variety of PIC platforms and wavelengths, and offers pathways to automated high-throughput packaging. Using this method, we show optically and electrically packaged quantum memory modules integrating all required qubit controls on chip, operating at millikelvin temperatures with <3 dB losses achievable from fiber to quantum memory for the TE 0 mode at a wavelength of 737 nm.

Bernson, Robert [Tyndall National Institute, Cork ↗

HBMax: Optimizing Memory Efficiency for Parallel Influence Maximization on Multicore Architectures

The goal of influence maximization is to select k most-influential vertices or seeds in a network, where influence is defined by a given diffusion process. The problem has a number of important applications such as viral marketing, information spread, and epidemic control. Although computing optimal seed set is NP-Hard, due to the submodular nature of the problem efficient approximation algorithms exist. However, even state-of-the-art parallel implementations are limited by a sampling step that incurs large memory footprints. This in turn limits the problem size reach and approximation quality. In this work, we study the memory footprint of the sampling process collecting reverse reachability information in the IMM algorithm over large real-world social networks. We present an adaptive and memory-efficient optimization approach for a state-of-the-art multi-threaded parallel influence maximization algorithm. Our approach,HuffMax, uses a portion of the reverse reachable (RR) sets collected by the algorithm to learn the characteristics of the graph. Then, it compresses the intermediate reverse reachability information with Huffman coding, and queries directly on the compressed data to preserve the memory savings obtained through compression. We also propose an efficient sampling strategy based on the distribution of RR sets, which can further reduce the computation time for typical social networks with long-tail distributions. Considering a NUMA architecture, we scale up our solution on 128-core CPUs and reduce the memory footprint by up to 45.7% with negligible time overhead (or even faster) and without perceivable loss of accuracy.

Chen, Xinyu↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗