Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A gradient-based deep neural network model for simulating multiphase flow in porous media

We report simulation of multiphase flow in porous media is crucial for the effective management of subsurface energy and environment-related activities. The numerical simulators used for modeling such processes rely on spatial and temporal discretization of the governing mass and energy balance partial-differential equations (PDEs) into algebraic systems via finite-difference/volume/element methods. These simulators usually require dedicated software development and maintenance, and suffer low efficiency from a runtime and memory standpoint for problems with multi-scale heterogeneity, coupled-physics processes or fluids with complex phase behavior. Therefore, developing cost-effective, data-driven models can become a practical choice, and in this work, we choose deep learning approaches as they can handle high dimensional data and accurately predict state variables with strong nonlinearity. In this paper, we describe a gradient-based deep neural network (GDNN) constrained by the physics related to multiphase flow in porous media. We tackle the nonlinearity of flow in porous media induced by rock heterogeneity, fluid properties, and fluid-rock interactions by decomposing the nonlinear PDEs into a dictionary of elementary differential operators. We use a combination of operators to handle rock spatial heterogeneity and fluid flow by advection. Since the augmented differential operators are inherently related to the physics of fluid flow, we treat them as first principles prior knowledge to regularize the GDNN training. We use the example of pressure management at geologic CO 2 storage sites, where CO 2 is injected in saline aquifers and brine is produced, and apply GDNN to construct a predictive model that is trained with physics-based simulation data and emulates the physics process. We demonstrate that GDNN can effectively predict the nonlinear patterns of subsurface responses, including the temporal and spatial evolution of the pressure and saturation plumes. We also successfully extend the GDNN to convolutional neural network (CNN), namely gradient-based CNN (GCNN), and validate its capability to improve the prediction accuracy. GDNN has great potential to tackle challenging problems that are governed by highly nonlinear physics and enable the development of data-driven models with higher fidelity.

42 ENGINEERING↗

Closed‐Loop Recyclable Vitrimer Plastics from PET Waste: A Design for Circularity

Plastics are essential to modern society, but their low recycling rates and inefficient end-of-life management pose a significant environmental challenge. Herein, the efficient strategy for upcycling postconsumer poly(ethylene terephthalate) (PET) waste into robust, closed-loop recyclable vitrimer plastics and composites is presented to address this issue. The catalyst-free aminolysis utilizes readily available amines to deconstruct diverse PET wastes into macromonomers, which are upcycled into vitrimers, exhibiting superior mechanical properties and exceeding the ultimate tensile stress and Young's Modulus of virgin PET by 80% and 150% respectively. These vitrimers exhibit excellent healability, shape memory, thermal reprocessability, and closed-loop chemical recyclability, enabling quantitative macromonomer recovery even from mixed plastic waste streams and glass/carbon fiber reinforced vitrimer (G/CFRV) composites. Furthermore, the vitrimer resin yields robust GFRV and CFRV composites with tensile strengths exceeding those of traditional epoxy composites by 100% and 80%, respectively, while maintaining complete chemical recyclability of both constituent materials. A preliminary technoeconomic analysis confirms the costeffectiveness and competitiveness of the facile PET deconstruction approach, which is potentially adaptable to other condensation polymers. Further, this study presents a facile approach to upcycling plastic waste into circular plastics and composites, offering a sustainable solution to global plastic waste management and fostering a circular economy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep quantum circuit simulations of low-energy nuclear states

Numerical simulation is an important method for verifying the quantum circuits used to simulate low-energy nuclear states. However, real-world applications of quantum computing for nuclear theory often generate deep quantum circuits that place demanding memory and processing requirements on conventional simulation methods. Here, we present advances in high-performance numerical simulations of deep quantum circuits to efficiently verify the accuracy of low-energy nuclear physics applications. Our approach employs novel methods for accelerating the numerical simulation including management of simulated mid-circuit measurements to verify projection based state preparation circuits. In this study, we test these methods across a variety of high-performance computing systems and our results show that circuits up to 21 qubits and more than 115,000,000 gates can be efficiently simulated.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HiFiPDV2 v. 2

SAND2024-01061O The HiFiPDV-2 software is similar to other photonic Doppler velocimetry (PDV) analysis codes in that it converts signals from intensity-time space to frequency-time space using a short-time Fourier transform (STFT). It then analyzes peaks of the resulting power spectral density information to extract the velocity-time history. The unique aspects of HiFiPDV and HiFiPDV-2 are that this reduction process is repeated for an array of reasonable STFT input variables. The resulting population distribution of output velocity-time histories is evaluated to determine the most likely velocity history and its associated uncertainty. This uncertainty is defined as the systematic uncertainty of the PDV signal as it is a function of input values to the reduction process. The HiFiPDV-2 program brings new functionality and efficiency to these calculations—most notably significant reductions in memory usage and calculation times, the ability to evaluate frequency-shifted signals, and the ability to filter out baseline signals. Developers recommend that users have at least 2 GB of RAM per CPU core in their system. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Voorhees, Travis↗

Character Memory- Umbra Package

Character Memory provides visual and acoustic sensing capabilities for characters within the Umbra simulation framework. Character Memory tracks where the character has seen and heard things, and then classifies them based on what they were and what their affiliation might be. The package provides outputs that can drive behaviors in response to those detections. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3053 O

Hart, BrianE↗

Improving the prediction of daily reservoir releases over the CONUS using conditioned LSTM

Reservoirs play a vital role in regulating streamflow timing and variability for hydroelectricity, flood control, water supply, irrigation, and recreation. Despite their importance, many reservoirs lack comprehensive operational guidelines, making their management complex due to conflicting operational objectives. Hence traditional policy-based reservoir models often fail to capture real-world conditions accurately and they depend on perfect streamflow predictions, which are not always available. In contrast, data-driven models like Long Short-Term Memory (LSTM) networks offer a robust alternative. This study introduces an approach that integrates reservoir characteristics—such as main use, climate, and maximum capacity—into the LSTM model to enhance reservoir release predictions. Using data from nearly 200 reservoirs in the contiguous United States (CONUS), our conditioned LSTM model (LSTM_cond) was compared with both the vanila LSTM and a traditional policy-based approach. Furthermore, our results show that while both LSTM_cond and LSTM perfoms better than the policy-based approach, LSTM_cond consistently outperforms LSTM for hydroelectric, water supply, irrigation, and recreation reservoirs. The KGE median values for LSTM_cond for out-sample reservoirs are 0.764, 0.565, 0.821, and 0.779, respectively, for the aforementioned reservoir types, which are consistently higher that the corresponding KGE values of 0.737, 0.413, 0.775, and 0.713 of LSTM, demonstrating its advantages in improving generalizability.

CONUS↗

Enhancing scalability of a matrix-free eigensolver for studying many-body localization

We propose several techniques to enhance the parallel scalability of a matrix-free eigensolver designed for studying many-body localization (MBL) of quantum spin chain models with nearest-neighbor interactions and on-site disorder. This type of problem is computationally challenging because the dimension of the associated Hamiltonian matrix grows exponentially with respect to the number of spins L, and we need to average over different realizations of the random disorder to obtain relevant statistical behavior. For each disorder realization, we need to compute eigenvalues from different regions of the spectrum and their corresponding eigenvectors. In previous work, the interior eigenstates for a single eigenvalue problem are computed via the shift-and-invert Lanczos algorithm. Due to the extremely high memory footprint of the LU factorizations, this technique is not well suited for large L’s. For example, we need thousands of compute nodes on modern high performance computing infrastructures to go beyond L = 24. The matrix-free approach does not suffer from this memory bottleneck, however, its scalability is limited by a computation and communication load imbalance. To reduce this imbalance and to significantly enhance the scalability of the matrix-free eigensolver, we reorder the matrix and leverage the consistent space runtime, CSPACER. We also show its efficiency in managing irregular communication patterns at scale compared to optimized MPI non-blocking two-sided and one-sided RMA implementation variants. This effort enables us to study MBL for spin chains with a larger number of spins. The efficiency and effectiveness of the proposed algorithm is demonstrated by computing eigenstates on a massively parallel many-core high performance computer.

METIS↗

Versatile High-Gain Low-Noise Readout ASIC for Silicon Microstrip Tracking Detectors

This work presents Turpial, a custom-designed low- power front-end readout ASIC for microstrip silicon sensors. Implemented in 130 nm CMOS technology, the chip integrates 64 identical readout channels, each including a configurable charge-sensitive amplifier, a bipolar pulse shaper, a 32-sample 50 Msps analog memory, and a 12-bit RC-hybrid SAR ADC operating at 1 Msps. To satisfy the target power budget of 5 mW per channel, the architecture employs a time-decoupled readout scheme in which fast transient signals are first captured in the analog memory and subsequently digitized at a lower rate. Turpial supports a wide dynamic range from 1 kℎ+ to 1 Mℎ+ while maintaining low noise performance, targeting an equivalent noise charge (ENC) below 200 𝑒−including the sensor, and providing a maximum gain of 1500 mV/fC. A digital block manages slow control, data acquisition, and data serialization through dual CML 300 Mb/s serializers. In addition, an on-chip reference circuit, based on a sub-1 V bandgap reference and an integrated LDO regulator, eliminates the need for external reference circuitry. Experimental results demonstrate that both the individual building blocks and the fully integrated ASIC meet the design specifications.

Hernandez, Hugo [Stanford University] (ORCID:00000↗

Towards Ultra-high-resolution E3SM Land Modeling on Exascale Computers

Here we present an ultra-high-resolution E3SM land model (uELM) for high-fidelity land simulations targeting new Exascale computers. After considering modeling infrastructure compatibility and ELM software features, we designed a parallel model for the uELM development targeting hybrid architectures of new US Exascale computers. We also described a function unit test framework to expedite the piece-wise code porting (with compiler directives), verification, and global variable management. Furthermore, in this study, we report an early uELM model development using OpenACC within a function unit test framework on a pre-Exascale computer, demonstrate the performance of a uLEM submodel with a 3.0-time speedup, and summarize the code porting experience regarding global variable handling, deepcopy, memory reduction, and parallel loop reconstruction.

97 MATHEMATICS AND COMPUTING↗

CRADA Final Report: Open Microgrid Platform

Current microgrid control technology is mostly proprietary, expensive, and inflexible. An open-source, vendor-neutral, publish-subscribe controller will significantly lower the barrier to entry for new technologies and market entrants: inverter, battery, and load control manufacturers; cloud and services providers; and energy aggregators. An OpenFMB controller can significantly reduce the acquisition, integration, and ongoing operating and maintenance costs and increase revenue opportunities for energy producers and consumers. It allows distributed energy resource (DER) owners to replace individual DER components as richer-function, lower-cost devices become available or to incorporate new forecast and optimization algorithms as they are developed. A standard, secure, full-function DER field controller will also provide a low-cost means for grid operators to effectively manage the variability and uncertainty of solar photovoltaics (PV). Under the terms of this CRADA, ORNL has worked with Open Energy Solutions (OES) to advance the transition of ORNL’s existing open-source microgrid controller into a platform geared toward mainstream implementation. Using a more prevalent and memory-safe programming language, Rust, OES has converted ORNL’s current generation on-grid optimization into an application more suitable for industry use. The results of the ported optimization were analyzed by inputting the same arguments in both the current-generation and next-generation optimizers and verifying that the results were equivalent.

24 POWER TRANSMISSION AND DISTRIBUTION↗

PARAFAC_T1

This is a method of performing trilinear analysis on large data sets using a modification of the PARAFAC-ALS algorithm. It iteratively decomposes the data matrix into a core matrix and three loading matrices based on the Tucker1 model. The algorithm is particularly useful for data sets that are too large to upload into a computer?s main memory. While the performance advantage in utilizing our algorithm is dependent on the number of data elements and dimensions of the data array, we have seen a significant performance improvement over operating PARAFAC-ALS on the full data set. SAND2020-12649 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Van Benthem, Mark↗

IceNet for FireBox - A Berkeley Warehouse-Scale Computer

Berkeley’s FireBox is a next-generation warehouse-scale computer (WSC) that utilizes the energy-efficiency and bandwidth density of integrated silicon-photonic interconnects to enable a new high-bandwidth and low-latency network fabric connecting thousands of compute nodes to petabytes of DRAM and Flash storage. The high bandwidth, low latency and high connectivity of FireBox’s WSC network fabric (IceNet) will enable dramatic improvements in the overall system energy efficiency enabling fine-grain power control on system resources (processors, links and memory/storage components). IceNet is a special 3-stage photonic Clos network architected to achieve ultra-low-latency connectivity between processor nodes and memory, drastically cutting down on the energy wasted in resource idling (processors and memory stalled due to pending network requests). This is achieved by integration of the first and last switch stages into processor/memory hub clients and by heavy over-provisioning of the high-radix middle switches (FlareSwitches). A key hardware component developed in this program is an active laser power management photonic integrated circuits called LightSpark. It interacts with the FlareSwitch and provides laser power to a subset of occupied switch ports, increasing the utilization of laser light in the photonic network by an order of magnitude. In addition to guiding the laser power where it is needed, the laser-power management module enables both wavelength and laser redundancy, significantly increasing the robustness of the system. The goal of the IceNet fabric is to enable communication between 1000s of processor nodes and PBs of memory/storage with <100ns latency, <10pJ/b wall-plug energy cost at multiple Pb/s of available connectivity bandwidth. These metrics represent two-orders of magnitude improvement with respect to the status of current data-center technology.

42 ENGINEERING↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

kynema-fmb [SWR-23-07]

Kynema-FMB (FKA: Kynema) is an open-source performance portable flexible multibody (FMB) dynamics solver designed for time-domain simulations. While originally tailored for wind turbine structural dynamics, the formulation and implementation are those of a general flexible-multidbody dynamics solver that can readily be applied to a wide range of systems. Kynema was designed with a narrow focus, namely to provide a lightweight, fast, accurate FMD solver for coupling to computational-fluid-dynamics (CFD) codes, especially the CFD codes in the Kynema suite, for fluid-structure-interaction (FSI) simulations. Kynema-FMB is equipped to model systems that can be represented as a collection of beams and rigid bodies that are connected through constraints. Degrees of freedom are defined in the inertial/global frame of reference and include displacements and rotations (formally as rotation matrices, but stored as quaternions). The underlying formulation is built on a Lie-group time integrator designed for index-3 differential-algebraic equations, which is second-order accurate in time (Bruls et al., 2012). Beam models are based on geometrically exact beam theory and are discretized as high-order spectral finite elements similar to those in BeamDyn (Wang et al., 2017). The governing equations for a FMD system like a wind turbine constitute a highly nonlinear system of constrained partial-differential equations. Kynema-FMB uses analytical Jacobians in the nonlinear-system solves in each time step. Linear systems use sparse storage and several third-party sparse-linear-system solvers are enabled. Ill conditioning of linear systems is mitigated with preconditioning described in Bottasso et al, 2008. Kynema-FMB is integrated with a simple open-source controller (ROSCO). There is an application programming interface (API) for coupling to geometry-resolved CFD (like that in Sharma et al., 2023) and actuator-force CFD (like that in Kuhn et al., 2025). In the latter, for actuator-line models, Kynema-FMB includes an internal blade-element solver that depends on user-provided lookup tables for coefficients of lift and drag, i.e., aerodynamic polars. Kynema-FMB is written in C++ and leverages Kokkos and Kokkos-Kernels (KokkosEcosystem) as its performance portability layer enabling simulations on both CPU and GPU systems. The repository is equipped with extensive automated testing at the unit and regression/system levels. The following describes the high-level development objectives conceived for Kynema: *Kynema will follow modern software development best practices, including test-driven development (TDD), version control, hierarchical automated testing, and continuous integration (CI) for a robust development environment. *The core data structures are memory efficient and enable vectorization and parallelization at multiple levels. *Data structures are data-oriented to exploit methods for accelerated computing including high utilization of chip resources (e.g., single instruction multiple data (SIMD) instruction sets) and parallelization using GP-GPUs. *The computational algorithms incorporate robust open-source libraries for mathematical operations, resource allocation, and data management. *The API design considers multiple stakeholder needs and ensure integration with existing and future ecosystems for data science, machine learning, and AI. *Kynema-FMB is written in modern C++ and leverages Kokkos as its performance-portability library with inspiration from the kynema stack.

Sprague, MichaelA.↗

Verbal Learning and Memory Deficits across Neurological and Neuropsychiatric Disorders: Insights from an ENIGMA Mega Analysis

Deficits in memory performance have been linked to a wide range of neurological and neuropsychiatric conditions. While many studies have assessed the memory impacts of individual conditions, this study considers a broader perspective by evaluating how memory recall is differentially associated with nine common neuropsychiatric conditions using data drawn from 55 international studies, aggregating 15,883 unique participants aged 15–90. The effects of dementia, mild cognitive impairment, Parkinson’s disease, traumatic brain injury, stroke, depression, attention-deficit/hyperactivity disorder (ADHD), schizophrenia, and bipolar disorder on immediate, short-, and long-delay verbal learning and memory (VLM) scores were estimated relative to matched healthy individuals. Random forest models identified age, years of education, and site as important VLM covariates. A Bayesian harmonization approach was used to isolate and remove site effects. Regression estimated the adjusted association of each clinical group with VLM scores. Memory deficits were strongly associated with dementia and schizophrenia (p < 0.001), while neither depression nor ADHD showed consistent associations with VLM scores (p > 0.05). Differences associated with clinical conditions were larger for longer delayed recall duration items. By comparing VLM across clinical conditions, this study provides a foundation for enhanced diagnostic precision and offers new insights into disease management of comorbid disorders.

Neurosciences & Neurology↗