Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “shared memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.

George, Anjus [ORNL] (ORCID:0000000179737061)↗

Towards Scalable 3D Integration of 2T-nC FeRAM with Hundreds of Layer Stacking

In this article, we study the limits of the number of capacitors and read history dependence in a 2T-nC ferroelectric random-access memory (FeRAM) cell, paving the way for its high-density integration toward hundreds of stacked layers. Through a comprehensive experimental and simulation study on the scaling behavior of the 2T-nC FeRAM architecture, we demonstrate: (i) successful fabrication of 2T-64C cells with robust memory operation and clearly distinguishable ‘0’ and ‘1’ states, even in 64- capacitor configurations; (ii) that the parasitic capacitance of the floating node originates predominantly from the linear component of the ferroelectric capacitor, and its impact on n-scaling—due to degraded sense margin—can be mitigated by floating unselected capacitors with enough TΩ isolation; (iii) that sharing write and read transistors among n capacitors introduces a read history dependence issue due to fluctuating floating node voltage (VFN); and (iv) that a proposed FN discharge scheme can effectively eliminate read-sequence dependence, at the cost of reduced read endurance.

36 MATERIALS SCIENCE↗

Increasing Mosquito Abundance Under Global Warming

Mosquitoes are a key virus vector that poses significant health threats globally, affecting 700 million individuals and causing 1 million deaths annually. Accurately predicting mosquito abundance and dispersion remains a challenge. Complex interactions between mosquito dynamics and various environmental factors, notably hydrology, contribute to this challenge. Existing models typically focus on precipitation and temperature and often overlook further impacts of hydrological variables within mosquito modeling. In this study, we developed an artificial intelligence‐based model for mosquito dynamics, explicitly accounting for different hydrological variables, such as precipitation, soil moisture and streamflow. Using Toronto, Canada, as a case study, we identified causal relationships between changes in mosquito populations, hydrological factors, vegetation (e.g., leaf area index), and climate variables (e.g., daylight length, precipitation, and temperature). We embedded these relationships into a Long Short‐Term Memory (LSTM) Neural Network Model capable of accurately detecting mosquito dynamics across annual, seasonal, and monthly time scales. The LSTM is able to explain, on average, approximately 40% of the variance in the observed mosquito abundance data. Using the calibrated model, we predicted that the summer season mosquito abundance would increase by ∼16% and ∼19% under an intermediate greenhouse emission scenario, Shared Socioeconomic Pathway (SSP) 2–4.5, and a high greenhouse emission scenario, SSP5‐8.5, respectively. We expect that this model can serve as a valuable tool and inform science‐based decisions affecting mosquito dynamics and public health. It can also build a foundation for future risk analysis at the regional and larger scales.

54 ENVIRONMENTAL SCIENCES↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

Assessing Tropical Pacific-induced Predictability of Southern California Precipitation Using a Novel Multi-input Multi-output Autoencoder

We construct a novel Multi-Input Multi-Output Autoencoder-decoder (MIMO-AE) to capture the non-linear relationship of Southern California precipitation and tropical Pacific Ocean sea surface temperature. The MIMO-AE is trained on both monthly TP-SST and SC-PRECIP anomalies simultaneously. The co-variability of the two fields in the MIMO-AE shared nonlinear latent space can be condensed into an index, termed the MIMO-AE index. We use a transfer learning approach to train a MIMO-AE on the combined dataset of 100 years of output from a historical simulation with the Energy Exascale Earth Systems Model version 1 and a segment of observational data. We further use Long Short-Term Memory networks to assess sub-seasonal predictability of SC-PRECIP using the MIMO-AE index. We find that the MIMO-AE index provides enhanced predictability of SC-PRECIP for a lead-time of up-to four months as compared to Niño 3.4 index and the El Niño Southern Oscillation Longitudinal Index.

54 ENVIRONMENTAL SCIENCES↗

The Adjoint Petrov–Galerkin method for non-linear model reduction

Here, we formulate a new projection-based reduced-order modeling technique for non-linear dynamical systems. The proposed technique, which we refer to as the Adjoint Petrov–Galerkin (APG) method, is derived by decomposing the generalized coordinates of a dynamical system into a resolved coarse-scale set and an unresolved fine-scale set. A Markovian finite memory assumption within the Mori–Zwanzig formalism is then used to develop a reduced-order representation of the coarse scales. This procedure leads to a closed reduced-order model that displays commonalities with the adjoint stabilization method used in finite elements. The formulation is shown to be equivalent to a Petrov–Galerkin method with a non-linear, time-varying test basis, thus sharing some similarities with the Least-Squares Petrov–Galerkin method. Theoretical analysis examining a priori error bounds and computational cost is presented. Numerical experiments on the compressible Navier–Stokes equations demonstrate that the proposed method can lead to improvements in numerical accuracy, robustness, and computational efficiency over the Galerkin method on problems of practical interest. Improvements in numerical accuracy and computational efficiency over the Least-Squares Petrov–Galerkin method are observed in most cases.

42 ENGINEERING↗

Multifidelity computing for coupling full and reduced order models

Hybrid physics-machine learning models are increasingly being used in simulations of transport processes. Many complex multiphysics systems relevant to scientific and engineering applications include multiple spatiotemporal scales and comprise a multifidelity problem sharing an interface between various formulations or heterogeneous computational entities. To this end, we present a robust hybrid analysis and modeling approach combining a physics-based full order model (FOM) and a data-driven reduced order model (ROM) to form the building blocks of an integrated approach among mixed fidelity descriptions toward predictive digital twin technologies. At the interface, we introduce a long short-term memory network to bridge these high and low-fidelity models in various forms of interfacial error correction or prolongation. The proposed interface learning approaches are tested as a new way to address ROM-FOM coupling problems solving nonlinear advection-diffusion flow situations with a bifidelity setup that captures the essence of a broad class of transport processes.

59 BASIC BIOLOGICAL SCIENCES↗

Day-Ahead Forecasting with Federated LSTM to Plan Energy Sharing in a Community Microgrid

Energy balancing in microgrids is a key enabler of resilience. Community microgrids located close to each other have the added benefit of networking and sharing surplus energy, if available. Such complex decision-making runs on optimization that requires reliable short-term (up to very-short-term) forecasts of energy generation and consumption for scheduling or trading. Each microgrid may also opt to not expose their sensitive data such as consumption patterns of individual businesses or residences. This paper investigates a federated approach to dayahead forecasting that trains naive long short-term memory (LSTM) at each business in a microgrid and aggregates weights at the microgrid controller using proximal regularization. This approach ensures that the controller has access only to energy surplus/deficit and not the actual generation or consumption values, avoiding unwanted exposure of sensitive data. A community microgrid in Adjuntas, Puerto Rico with 3 businesses is selected as a case study with a laboratory-scale computing setup. A central LSTM forecaster, where sensitive data from businesses are aggregated at the controller, is implemented as a baseline for qualifying the results. This work serves as a proof-of-concept for scaling the approach to networked and nested microgrids with more complex control options.

Sundararajan, Aditya [ORNL] (ORCID:000000033577854↗

Leaky Buddies: Cross-Component Covert Channels on Integrated CPU-GPU Systems

Graphics Processing Units (GPUs) are ubiquitous components used across the range of today’s computing plat- forms, from phones and tablets, through personal computers, to high-end server class platforms. With the increasing importance of graphics and video workloads, recent processors are shipped with GPU devices that are integrated on the same chip. Integrated GPUs share some resources with the CPU and as a result, there is a potential for microarchitectural attacks from the GPU to the CPU or vice versa. We consider the potential for covert channel attacks that arise either from shared microarchitectural components (such as caches) or through shared contention domains (e.g., shared buses). We illustrate these two types of channels by developing two reliable covert channel attacks. The first covert channel uses the shared LLC cache in Intel’s integrated GPU architectures. The second is a contention based channel targeting the ring bus connecting the CPU and GPU to the LLC. This is the first demonstrated microarchitectural attack crossing the component boundary (GPU to CPU or vice versa). Cross-component channels introduce a number of new challenges that we had to overcome since they occur across heterogeneous components that use different computation models and are interconnected using asymmetric memory hierarchies. We also exploit GPU parallelism to increase the bandwidth of the communication, even without relying on a common clock. The LLC based channel achieves a bandwidth of 120 kbps with a low error rate of 2%, while the contention based channel delivers up to 400 kbps with a 0.8% error rate. We also demonstrate a proof- of-concept prime-and-probe side channel attack that probes the full LLC from the GPU.

Cyber secruity, high performance comptuing, Hetero↗

Potentials of mean force fail to describe chemical bond-breaking in solution

Many liquid phase studies assume that the potential energy surfaces of reacting molecules are the same as in the gas phase, neglecting complex solvent dynamics that can completely alter the nature of chemical reactivity. Even studies that include solvent effects typically only consider them in an average, equilibrium way as part of a potential of mean force (PMF). In this work, we use mixed quantum/classical simulations to compare how equilibrium and non-equilibrium solvent motions affect the photodissociation of a simple diatomic molecule, NaK + , in liquid tetrahydrofuran. A PMF analysis shows that as the excited-state molecule dissociates with the solvent at equilibrium, the bonding electron remains associated with K + at short bond distances but eventually localizes on Na + at the end of dissociation. When we examine non-equilibrium dynamical photodissociation trajectories, however, we find that they fall into three distinct categories: about a quarter of them have the bonding electron mainly associated with Na + , another quarter stay mainly associated with K + , and about half have the bonding electron shared roughly equally between the two ions. The results show that equilibrium PMFs cannot accurately describe the dynamics of bond-breaking chemical reactions in solution because there is insufficient time for the solvent to reach equilibrium on the time scale over which bond dissociation occurs. Furthermore, our analysis shows that the solvent coupling between the electronic energy surfaces is similar at and away from equilibrium, suggesting that other factors, such as solute velocity-driven solvent memory effects, play a more important role in explaining the failure of the equilibrium PMF to predict the non-equilibrium dynamics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient loading of reduced data ensembles produced at ORNL SNS/HFIR neutron time-of-flight facilities

We present algorithmic improvements to the loading operations of certain reduced data ensembles produced from neutron scattering experiments at Oak Ridge National Laboratory (ORNL) facilities. Ensembles from multiple measurements are required to cover a wide range of the phase space of a sample material of interest. They are stored using the standard NeXus schema on individual HDF5 files. This makes it a scalability challenge, as the number of experiments stored increases in a single ensemble file. The present work follows up on our previous efforts on data management algorithms, to address identified input output (I/O) bottlenecks in Mantid, an open-source data analysis framework used across several neutron science facilities around the world. We reuse an in-memory binary-tree metadata index that resembles data access patterns, to provide a scalable search and extraction mechanism. In addition, several memory operations are refactored and optimized for the current common use cases, ranging most frequently from 10 to 180, and up to 360 separate measurement configurations. Results from this work show consistent speed ups in wall-clock time on the Mantid LoadMD routine, ranging from 19% to 23% on average, on ORNL production computing systems. The latter depends on the complexity of the targeted instrument-specific data and the system I/O and compute variability for the shared computational resources available to users of ORNL’s Spallation Neutron Source (SNS) and the High Flux Isotope Reactor (HFIR) instruments. Nevertheless, we continue to highlight the need for more research to address reduction challenges as experimental data volumes, user time and processing costs increase.

Godoy, William↗

Scalable access to shared files in a distributed system

Logic may pre-populate cache of compute nodes allocated to a user's session. Logic may access an image file of a global filesystem, the image file to comprise a first set of one or more files associated with a user session and a second set of one or more files upon which execution of the first set will depend. Logic may store the image file in local data storage of each compute node of a spanning tree of compute nodes allocated for the user session. Logic may unpack the image file based on the data. Logic may configure the local environment of each compute node in the spanning tree to identify a location in local memory of each of the files. And logic may initiate execution of a local filesystem service in each of the compute nodes of the spanning tree.

Oganezov, Alexander↗

Space-Charge-Limited van der Waals Spin Transistor

Integrating semiconducting and magnetic materials could combine transistorlike operation with nonvolatility and enable architectures such as logic in memory. Here, in this work, we employ correlated electrical transport and scanning nitrogen-vacancy center magnetic imaging to elucidate a spin transistor concept that amalgamates vertical and lateral hopping transport inside a 2D antiferromagnetic semiconductor, mechanistically distinct from vertical tunneling devices. Our device, based on a monolayer-bilayer junction in CrSBr, displays giant, gate-tunable magnetoresistance driven by the dual action of electrostatic doping on space-charge-limited lateral conduction and interlayer exchange coupling. Moreover, we visualize a field-trainable, layer-sharing effect that selects between coherent or domain-wall reversal at the spin-flip transition, enabling multilevel, memristive conductance states. Our layer-dependent space charge mechanism for convergent electrical and magnetic control opens opportunities to address limitations in contemporary computing.

2-dimensional systems↗

Understanding Reliability Trade-Offs in 1T-nC and 2T-nC FeRAM Designs

Ferroelectric random access memory (FeRAM) is a promising candidate for energy-efficient nonvolatile memory, particularly for logic-in-memory and compute-in-memory (CIM) applications. Among the available cell architectures, One-Transistor–n-Capacitor (1T-nC) and two-transistor–n-capacitor (2T-nC) FeRAMs each offer distinct trade-offs in density, scalability, and reliability. In this work, we present a comparative study of these two architectures under both dimensional scaling ( XY/Z shrinkage) and vertical integration (increasing stacked capacitors per cell). Using technology computer-aided design (TCAD) and circuit-level simulations, we analyze how scaling impacts ferroelectric capacitance, parasitic coupling, and floating-node (FN) dynamics, which together dictate sense margin (SM) and read stability. A key mitigation strategy—floating unselected capacitors—is applied to both architectures, effectively decoupling the SM from the number of stacked capacitors and enabling tractable analysis across scaling regimes. Results show that 1T-nC suffers more from charge sharing with the bitline (BL), while 2T-nC benefits from transistor isolation and stronger low-voltage sensing at the cost of increased area. By systematically evaluating these behaviors across scaling directions, this work establishes the reliability trade-offs of 1T-nC and 2T-nC cells and provides design guidelines for high-density, vertically integrated FeRAM systems.

1T-nC↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Bulk Glassy GeTe 2 : A Missing Member of the Tetrahedral GeX 2 Family and a Precursor for the Next Generation of Phase-Change Materials

Vitreous germanium disulfide GeS 2 and diselenide GeSe 2 belong to canonical chalcogenide glasses extensively studied over the past half century. Their high-temperature orthorhombic polymorphs are congruently melting compounds, and the tetrahedral crystal and glass structure is largely preserved in the melt. In contrast, the ditelluride counterpart is absent in the Ge-Te phase diagram, which shows only a single compound, monotelluride GeTe. Phase-change materials based on GeTe have become a technologically important class of solids, and their structure and properties are also widely studied. Surprisingly, very scarce information is available for alloys having GeTe 2 stoichiometry. Using a fast quenching procedure in silica capillaries, high-energy X-ray diffraction, and Raman spectroscopy supported by first-principles simulations, we show that bulk glassy GeTe 2 differs substantially from the lighter GeX 2 members, revealing 46% of trigonal germanium, 31% of three-fold coordinated tellurium, and only 20% of edge-sharing tetrahedra or pyramids. The fraction of homopolar Ge-Ge bonds is low; however, the population of dominant Te-Te dimers and Te n oligomers, n <= 10, appears to be significant. The complex structural and chemical topology of g-GeTe 2 is directly related to the thermodynamic metastability of germanium ditelluride, schematically represented by the following reaction: GeTe 2 $\rightleftarrows$ GeTe + Te. Disproportionation is complete above liquidus in the temperature range of semiconductor-metal transition, and the dense metallic GeTe 2 liquid, mostly consisting of five-fold coordinated Ge species, exhibits high fluidity, strong fragility (m = 99 ± 5), and presumably a fast structural transformation rate combined with low atomic mobility in the vicinity of the glass transition temperature, favorable for reliable long-term data retention in nonvolatile memories. The observed and predicted characteristic features make GeTe 2 a promising precursor for the next generation of phase-change materials, especially coupled with additional metal doping, depolymerizing the tetrahedral interconnected glass network and accelerating (sub)nanosecond crystallization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗