Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “shared memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Comparison of ML-Based Proxy Modeling Strategies: Lessons Learned from the SMART Initiative

Teams of researchers on Task 5 of the SMART project have developed a variety of modeling architectures to predict subsurface behavior during carbon injection and post-injection periods. One important part of this task was to compare the candidate approaches in terms of accuracy, reliability, speed, and memory use, all using a common set of metrics and visualizations for an “apples to apples” comparison. Dr. Jared Schuetter will share the details of this task, the results that were obtained, and the lessons learned.

Schuetter, Jared↗

Evolution of the ATLAS event data model for the HL-LHC

The upcoming high-luminosity run of the CERN Large Hadron Collider (HL-LHC) will yield an unprecedented volume of data. In order to process this data, the ATLAS collaboration is evolving its offline software to be able to use heterogeneous resources such as graphical processing units (GPUs) and field-programmable gate arrays (FPGAs). To reduce conversion overheads, the event data model (EDM) should be compatible with the requirements of these resources. While the ATLAS EDM has long allowed representing data as a structure of arrays, further evolution of the EDM can enable more efficient sharing of data between CPU and GPU resources. Some of this work will be summarized here, including extensions to allow controlling how memory for event data is allocated and the implementation of jagged vectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MIND-MAC: Multi-Level In-memory Quasi Non-Destructive MAC Operation in Compact 2T-nC FeRAM for Efficient DNN Accelerator

We present MIND-MAC, a compact 2T-nC FeRAM architecture that performs multi-level, quasi-non-destructive in-memory multiply–accumulate (MAC) for deep neural networks. By exploiting voltage-controlled partial domain switching in MFM capacitors and read-transistor amplification, the cell stores multi-bit weights and gates bit-serial inputs to produce an accumulated current on shared lines. We combine TCAD-extracted parasitics with experimentally calibrated ferroelectric models in SPICE to validate device-/circuit-level behavior, and validate multi-level sensing and QNRO with measurements on a fabricated 2T-3C test vehicle. An analytical system model maps MIND-MAC to a 6-GB main-memory in-memory compute (IMC) architecture and benchmarks VGG13 inference in 61.08 ms at 964.99 mJ. Results indicate high density, reduced rewrite overhead, and energy efficiency, positioning 2T-nC FeRAM as a promising IMC candidate for next-generation AI hardware.

36 MATERIALS SCIENCE↗

Open-Source Contributions to Arbiter2

Login nodes at High-Performance Computing (HPC) sites are shared resources used by a multitude of users to compile, test and submit jobs to the batch system. Because these nodes are shared by multiple users at any time, if a minority of users are using a significant proportion of the resources of the node (CPU, memory, etc.), other users' tasks may be negatively impacted. Arbiter2 is an open-source project developed by the University of Utah that aims to prevent these occurrences by dynamically limiting the resources of users depending on whether they are excessively using resources. INL has sought to adapt and develop Arbiter2 for a potential deployment at INL and plans on upstreaming the changes and improvements so that other HPC sites running Arbiter2 can benefit from their work.

97 MATHEMATICS AND COMPUTING↗

Virtualized Logical Qubits: A 2.5D Architecture for Error-Corrected Quantum Computing

Current, near-term quantum devices have shown great progress in the last several years culminating recently with a demonstration of quantum supremacy. In the medium-term, however, quantum machines will need to transition to greater reliability through error correction, likely through promising techniques like surface codes which are well suited for near-term devices with limited qubit connectivity. We discover quantum memory, particularly resonant cavities with transmon qubits arranged in a 2.5D architecture, can efficiently implement surface codes with substantial hardware savings and performance/fidelity gains. Specifically, we virtualize logical qubits by storing them in layers of qubit memories connected to each transmon. Surprisingly, distributing each logical qubit across many memories has a minimal impact on fault tolerance and results in substantially more efficient operations. Our design permits fast transversal application of CNOT operations between logical qubits sharing the same physical address (same set of cavities) which are 6x faster than standard lattice surgery CNOTs. We develop a novel embedding which saves approximately 10x in transmons with another 2x savings from an additional optimization for compactness. Although qubit virtualization pays a 10x penalty in serialization, advantages in the transversal CNOT and in area efficiency result in fault-tolerance and performance comparable to conventional 2D transmon-only architectures. Our simulations show our system can achieve fault tolerance comparable to conventional two-dimensional grids while saving substantial hardware. Furthermore, our architecture can produce magic states at 1.22x the baseline rate given a fixed number of transmon qubits. Here, this is a critical benchmark for future fault-tolerant quantum computers as magic states are essential and machines will spend the majority of their resources continuously producing them. This architecture substantially reduces the hardware requirements for fault-tolerant quantum computing and puts within reach a proof-of-concept experimental demonstration of around 10 logical qubits, requiring only 11 transmons and 9 attached cavities in total.

quantum computing↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

The SONATA data format for efficient description of large-scale network models

Increasing availability of comprehensive experimental datasets and of high-performance computing resources are driving rapid growth in scale, complexity, and biological realism of computational models in neuroscience. To support construction and simulation, as well as sharing of such large-scale models, a broadly applicable, flexible, and high-performance data format is necessary. To address this need, we have developed the Scalable Open Network Architecture TemplAte (SONATA) data format. It is designed for memory and computational efficiency and works across multiple platforms. The format represents neuronal circuits and simulation inputs and outputs via standardized files and provides much flexibility for adding new conventions or extensions. SONATA is used in multiple modeling and visualization tools, and we also provide reference Application Programming Interfaces and model examples to catalyze further adoption. SONATA format is free and open for the community to use and build upon with the goal of enabling efficient model building, sharing, and reproducibility.

59 BASIC BIOLOGICAL SCIENCES↗

An interlaboratory comparison of mid-infrared spectra acquisition: Instruments and procedures matter

Diffuse reflectance spectroscopy has been extensively employed to deliver timely and cost-effective predictions of a number of soil properties. However, although several soil spectral laboratories have been established worldwide, the distinct characteristics of instruments and operations still hamper further integration and interoperability across mid-infrared (MIR) soil spectral libraries. In this study, we conducted a large-scale ring trial experiment to understand the lab-to-lab variability of multiple MIR instruments. By developing a systematic evaluation of different mathematical treatments with modeling algorithms, including regular preprocessing and spectral standardization, we quantified and evaluated instruments' dissimilarity and how this impacts internal and shared model performance. We found that all instruments delivered good predictions when calibrated internally using the same instruments' characteristics and standard operating procedures by solely relying on regular spectral preprocessing that accounts for light scattering and multiplicative/additive effects, e.g., using standard normal variate (SNV). When performing model transfer from a large public library (the USDA NSSCKSSL MIR library) to secondary instruments, good performance was also achieved by regular preprocessing (e. g., SNV) if both instruments shared the same manufacturer. However, significant differences between the KSSL MIR library and contrasting ring trial instruments responses were evident and confirmed by a semi-unsupervised spectral clustering. For heavily contrasting setups, spectral standardization was necessary before transferring prediction models. Non-linear model types like Cubist and memory-based learning delivered more precise estimates because they seemed to be less sensitive to spectral variations than global partial least square regression. In summary, the results from this study can assist new laboratories in building spectroscopy capacity utilizing existing MIR spectral libraries and support the recent global efforts to make soil spectroscopy universally accessible with centralized or shared operating procedures.

58 GEOSCIENCES↗

Rapid wavefield forecasting for earthquake early warning via deep sequence to sequence learning

We propose a deep learning model, WaveCastNet, to forecast high-dimensional wavefields. WaveCastNet integrates a convolutional long expressive memory architecture into a sequence-to-sequence forecasting framework, enabling it to model long-term dependencies and multiscale patterns in both space and time. By sharing weights across spatial and temporal dimensions, WaveCastNet requires significantly fewer parameters than more resource-intensive models such as transformers, resulting in faster inference times. Crucially, WaveCastNet also generalizes better than transformers to rare and critical seismic scenarios, such as high-magnitude earthquakes. Here, we show the ability of the model to predict the intensity and timing of destructive ground motions in real time, using simulated data from the San Francisco Bay Area. Furthermore, we demonstrate its zero-shot capabilities by evaluating WaveCastNet on real earthquake data. Our approach does not require estimating earthquake magnitudes and epicenters, steps that are prone to error in conventional methods, nor does it rely on empirical ground-motion models, which often fail to capture strongly heterogeneous wave propagation effects.

Geophysics↗

Using Machine Learning to Predict Future Temperature Outputs in Geothermal Systems

Optimizing the power output, and economic value, of geothermal power plants over decades of operation is a major challenge in renewable energy. Optimizing the output requires the ability to predict the mass flow rates and the output temperatures of production wells based on the inputs of injection wells, as well as the time history of the system. Machine Learning (ML) that incorporates the known physics of geothermal systems is one possible solution to this challenge. In this work, we explore the ability of ML algorithms to predict future temperature outputs based on historical data. Considering the challenges with obtaining an empirical dataset from field data that is large enough to enable reliable ML, we propose an alternate approach: developing a high-fidelity reservoir model and using computational resources to build a dataset that enables ML. As a first step towards achieving this goal, we present preliminary results from applying ML to predict the temperature timeseries of simple modeled geothermal systems. We describe the application of relevant state-of-the-art ML approaches, such as the Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNN), to extract temporal structures in the model data. We assess the accuracy of the forecasts we obtain, compare the selected approaches, and share the lessons learned that would inform the process of training and utilizing ML algorithms for larger and more complex geothermal systems.

GEOTHERMAL ENERGY↗

Synchronization for CXL Based Memory

Compute Express Link (CXL) is an important emerging standard for disaggregated memory. While this standard provisions coherency across numerous hosts and devices, implementing hardware support for type three devices is challenging. In this work, we look at the overhead of software synchronization and using software-based coherency. Moreover, we discuss the limits of software-based coherency in fully expressing modern synchronization techniques for a CXL-based disaggregate memory system. We demonstrate our approach using a CXL hardware prototype and running a version of the famous Peterson Lock (enhanced to run with more than two threads). We analyze its performance and share how more advanced synchronization techniques might interact with software-based coherence CXL hardware and program execution models.

High Performance Computing (HPC)↗

Hybrid classical-quantum communication networks

Over the past several decades, the proliferation of global classical communication networks has transformed various facets of human society. Concurrently, quantum networking has emerged as a dynamic field of research, driven by its potential applications in distributed quantum computing, quantum sensor networks, and secure communications. This prompts a fundamental question: rather than constructing quantum networks from scratch, can we harness the widely available classical fiber-optic infrastructure to establish hybrid quantum–classical networks? This paper aims to provide a comprehensive review of ongoing research endeavors aimed at integrating quantum communication protocols, such as quantum key distribution, into existing lightwave networks. This approach offers the substantial advantage of reducing implementation costs by allowing classical and quantum communication protocols to share optical fibers, communication hardware, and other network control resources—arguably the most pragmatic solution in the near term. In the long run, classical communication will also reap the rewards of innovative quantum communication technologies, such as quantum memories and repeaters. Accordingly, our vision for the future of the Internet is that of heterogeneous communication networks thoughtfully designed for the seamless support of both classical and quantum communications.

Fiber-optic communication↗

Data and scripts associated with the manuscript "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning"

This package contains the data and scripts used in "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning" (Jiang et al., 2022). The data.zip file contains the flux tower and automated chamber observations used for developing the deep learning model for modeling soil respiration. The scripts.zip file contains the Jupyter notebooks and python scripts for preprocessing the data, training the deep learning models, and postprocessing the results. The src.zip contains the source code for training the deep learning model, performing mutual information analysis, and plotting functions. The trained_models.zip contains multiple folders used for hosting the trained deep-learning models and the associated soil respiration predictions. The whole process is performed using python. We include the REAMD.md to document the python package requirements.Soil respiration in dryland ecosystems is challenging to model due to its complex interactions with environmental drivers. Knowledge-guided deep learning provides a much more effective means of accurately representing these complex interactions than traditional Q10-based models. Mutual information analysis revealed that future soil temperature shares more information with soil respiration than past soil temperature, consistent with their clockwise diel hysteresis. We explicitly encoded diel hysteresis, soil drying, and soil rewetting effects on soil respiration dynamics in a newly designed Long Short Term Memory (LSTM) model. The model takes both past and future environmental drivers as inputs to predict soil respiration. The new LSTM model substantially outperformed three Q10-based models and the Community Land Model when reproducing the observed soil respiration dynamics in a semi-arid ecosystem. The new LSTM model clearly demonstrated its superiority for temporally extrapolating soil respiration dynamics, such that the resulting correlation with observational data is up to 0.7 while the correlations of both Q10-based models and the Community Land Model (CLM) are less than 0.4. Our results underscore the high potential for knowledge-guided deep learning to replace Q10-based soil respiration modules in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning

Soil respiration in dryland ecosystems is challenging to model due to its complex interactions with environmental drivers. Knowledge-guided deep learning provides a much more effective means of accurately representing these complex interactions than traditional Q10-based models. Mutual information analysis revealed that future soil temperature shares more information with soil respiration than past soil temperature, consistent with their clockwise diel hysteresis. We explicitly encoded diel hysteresis, soil drying, and soil rewetting effects on soil respiration dynamics in a newly designed Long Short Term Memory (LSTM) model. The model takes both past and future environmental drivers as inputs to predict soil respiration. The new LSTM model substantially outperformed three Q10-based models and the Community Land Model when reproducing the observed soil respiration dynamics in a semi-arid ecosystem. The new LSTM model clearly demonstrated its superiority for temporally extrapolating soil respiration dynamics, such that the resulting correlation with observational data is up to 0.7 while the correlations of the Q10-based models and the Community Land Model (CLM) are less than 0.4. Our results underscore the high potential for knowledge-guided deep learning to replace Q10-based soil respiration modules in Earth system models.

Jiang, Peishi↗

Status of DUNE Offline Computing

We summarize the status of Deep Underground Neutrino Experiment (DUNE) Offline Software and Computing program. We describe plans for the computing infrastructure needed to acquire, catalog, reconstruct, simulate and analyze the data from the DUNE experiment and its prototypes in pursuit of the experiment's physics goals of precision measurements of neutrino oscillation parameters, detection of astrophysical neutrinos, measurement of neutrino interaction properties and searches for physics beyond the Standard Model. In contrast to traditional HEP computational problems, DUNE's Liquid Argon Time Projection Chamber data consist of simple but very large (many GB) data objects which share many characteristics with astrophysical images. We have successfully reconstructed and simulated data from 4% prototype detector runs at CERN. The data volume from the full DUNE detector, when it starts commissioning late in this decade will present memory management challenges in conventional processing but significant opportunities to use advances in machine learning and pattern recognition as a frontier user of High Performance Computing facilities capable of massively parallel processing. Our goal is to develop infrastructure resources that are flexible and accessible enough to support creative software solutions as HEP computing evolves.

43 PARTICLE ACCELERATORS↗

Interlinked tuples in coordination namespace

A system and method for supporting tuple record interlinking in one or more tuple space/coordinated namespace (CNS) extended memory storage systems. A system-wide CNS provides for efficient storing and communicating of data generated by local processes running at the nodes, and coordinated to generate a union/intersection of multiple CNS where tuple records are interlinked in multiple CNS hashtables, and/or share tuple data between two sets of processes that are part of different CNSs. Local node processes further generate multi-key tuples where two or more tuple records are interlinked within the same CNS hash table, thereby permitting a look up of the tuple data by either tuple name/keys. A CNS controller further provides a tuple iterator for a key-value storage in a CNS system that adds more links between tuples enables creation of iterator structures such as linked list or trees etc. of “different” tuples in a tuple database.

Jacob, Philip↗

Oxidative stress is a shared characteristic of ME/CFS and Long COVID

Over 65 million individuals worldwide are estimated to have Long COVID (LC), a complex multisystemic condition marked by fatigue, post-exertional malaise, and other symptoms resembling myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS). With no clinically approved treatments or reliable diagnostic markers, there is an urgent need to define the molecular underpinnings of these conditions. By studying bioenergetic characteristics of peripheral blood lymphocytes in 25 healthy controls, 27 ME/CFS, and 20 LC donors, we find both ME/CFS and LC donors exhibit signs of elevated oxidative stress, especially in the memory subset. Using a combination of flow cytometry, RNA-seq, mass spectrometry, and systems chemistry analysis, we observed aberrations in reactive oxygen species (ROS) clearance pathways including elevated glutathione levels, decreases in mitochondrial superoxide dismutase protein levels, and glutathione peroxidase 4–mediated lipid oxidative damage. Strikingly, these redox pathways changes show sex-specific trends. While ME/CFS females exhibit higher total ROS and mitochondrial calcium levels, males have normal ROS levels, with pronounced mitochondrial lipid oxidative damage. In females, these higher ROS levels correlate with T cell hyperproliferation, consistent with the known role of elevated ROS in initiating proliferation. This hyperproliferation can be attenuated by metformin, suggesting this Food and Drug Administration (FDA)-approved drug as a possible treatment, as also suggested by a recent clinical study of LC patients. Moreover, these results suggest a shared mechanistic basis for the systemic phenotypes of ME/CFS and LC, which can be detected by quantitative blood cell measurements, and that effective, patient-tailored drugs might be discovered using standard lymphocyte stimulation assays.

ME/CFS↗

Visualizing the gas channel of a monofunctional carbon monoxide dehydrogenase

Carbon monoxide dehydrogenase (CODH) plays an important role in the processing of the one–carbon gases carbon monoxide and carbon dioxide. In CODH enzymes, these gases are channeled to and from the Ni-Fe-S active sites using hydrophobic cavities. In this work, we investigate these gas channels in a monofunctional CODH from Desulfovibrio vulgaris, which is unusual among CODHs for its oxygen-tolerance. By pressurizing D. vulgaris CODH protein crystals with xenon and solving the structure to 2.10 Å resolution, we identify 12 xenon sites per CODH monomer, thereby elucidating hydrophobic gas channels. We find that D. vulgaris CODH has one gas channel that has not been experimentally validated previously in a CODH, and a second channel that is shared with Moorella thermoacetica carbon monoxide dehydrogenase/acetyl-CoA synthase (CODH/ACS). This experimental visualization of D. vulgaris CODH gas channels lays groundwork for further exploration of factors contributing to oxygen-tolerance in this CODH, as well as study of channels in other CODHs. We dedicate this publication to the memory of Dick Holm, whose early studies of the Ni-Fe-S clusters of CODH inspired us all.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗