Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited memory method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗

Analysis and Benchmarking of feature reduction for classification under computational constraints

Abstract Machine learning is most often expensive in terms of computational and memory costs due to training with large volumes of data. Current computational limitations of many computing systems motivate us to investigate practical approaches, such as feature selection and reduction, to reduce the time and memory costs while not sacrificing the accuracy of classification algorithms. In this work, we carefully review, analyze, and identify the feature reduction methods that have low costs/overheads in terms of time and memory. Then, we evaluate the identified reduction methods in terms of their impact on the accuracy, precision, time, and memory costs of traditional classification algorithms. Specifically, we focus on the least resource intensive feature reduction methods that are available in Scikit-Learn library. Since our goal is to identify the best performing low-cost reduction methods, we do not consider complex expensive reduction algorithms in this study. In our evaluation, we find that at quadratic-scale feature reduction, the classification algorithms achieve the best trade-off among competitive performance metrics. Results show that the overall training times are reduced 61%, the model sizes are reduced 6×, and accuracy scores increase 25% compared to the baselines on average with quadratic scale reduction.

97 MATHEMATICS AND COMPUTING↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

Machine Learning for Slow Spill Regulation in the Fermilab Delivery Ring for Mu2e

A third-integer resonant slow extraction system is being developed for the Fermilab’s Delivery Ring to deliver protons to the Mu2e experiment. During a slow extraction process, the beam on target is liable to experience small intensity variations due to many factors. Owing to the experiment’s strict requirements in the quality of the spill, a Spill Regulation System (SRS) is currently under design. The SRS primarily consists of three components - slow regulation, fast regulation, and harmonic content tracker. In this presentation, we shall present the investigations of using Machine Learning (ML) in the fast regulation system, including further optimizations of PID controller gains for the fast regulation, prospects of an ML agent completely replacing the PID controller using supervised learning schemes such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) ML models, the simulated impact and limitation of machine response characteristics on the effectiveness of both PID and ML regulation of the spill. We also present here nascent results of Reinforcement Learning efforts, including continuous-action soft actor-critic methods, to regulate the spill rate.

43 PARTICLE ACCELERATORS↗

Processability and Material Behavior of NiTi Shape Memory Alloys Using Wire Laser-Directed Energy Deposition (WL-DED)

Utilizing additive manufacturing (AM) techniques with shape memory alloys (SMAs) like NiTi shows great promise for fabricating highly flexible and functionally superior 3D metallic structures. Compared to methods relying on powder feedstocks, wire-based additive manufacturing processes provide a viable alternative, addressing challenges such as chemical composition instability, material availability, higher feedstock costs, and limitations on part size while simplifying process development. This study presented a novel approach by thoroughly assessing the printability of Ni-rich Ni55.94Ti (Wt. %) SMA using the wire laser-directed energy deposition (WL-DED) technique, addressing the existing knowledge gap regarding the laser wire-feed metal additive manufacturing of NiTi alloys. For the first time, the impact of processing parameters—specifically laser power (400–1000 W) and transverse speed (300–900 mm/min)—on single-track fabrication using NiTi wires in the WL-DED process was examined. An optimal range of process parameters was determined to achieve high-quality prints with minimal defects, such as wire dripping, stubbing, and overfilling. Building upon these findings, we printed five distinct cubes, demonstrating the feasibility of producing nearly porosity-free specimens. Notably, this study investigated the effect of energy density on the printed part density, impurity pick-up, transformation temperature, and hardness of the manufactured NiTi cubes. The results from the cube study demonstrated that varying energy densities (46.66–70 J/mm3) significantly affected the quality of the deposits. Lower to intermediate energy densities achieved high relative densities (>99%) and favorable phase transformation temperatures. In contrast, higher energy densities led to instability in melt pool shape, increased porosity, and discrepancies in phase transformation temperatures. These findings highlighted the critical role of precise parameter control in achieving functional NiTi parts and offer valuable insights for advancing AM techniques in fabricating larger high-quality NiTi components. Additionally, our research highlighted important considerations for civil engineering applications, particularly in the development of seismic dampers for energy dissipation in structures, offering a promising solution for enhancing structural performance and energy management in critical infrastructure.

Dabbaghi, Hediyeh↗

Optimizing Prediction Error for Time-dependent Solar Radiation Modeling

Numerical weather forecasting models and statistical methods have found wide use to help power companies estimate renewable output, but better methods are needed, particularly for extended forecasts. Machine learning approaches have been used here as well, but so far a major limitation is the ability to also predict the corresponding uncertainty in a forecast. Here we show that both can be done and demonstrate this using a long-term-short memory neural network where the difference between predicted and ground truth data are used to train a model for the corresponding forecast uncertainties.

97 MATHEMATICS AND COMPUTING↗

Numerically exact configuration interaction at quadrillion-determinant scale

The combinatorial growth of configuration interaction (CI) has long limited this formally exact quantum chemistry method to only the smallest molecules. Here, we report a numerically exact CI calculation exceeding one quadrillion (10 15 ) determinants, made possible by a lossless categorical compression strategy within the small-tensor-product distributed active space (STP-DAS) framework. This approach overcomes the traditional memory bottlenecks of CI by a numerically exact compression of the wavefunction representation and reformulating the most computationally demanding matrix–vector operations. Using this method, we performed a fully relativistic CI calculation of the ground state of HBrTe with over 10 15 complex-valued determinants in just 34.5 h on 1000 computing nodes—the largest CI calculation ever reported. We further achieved fast computation for systems with hundreds of billions of determinants on only a few compute nodes. Extensive benchmarks confirm that the method retains full numerical exactness while cutting memory and computational cost by orders of magnitude. Compared to previous state-of-the-art CI calculations, this work achieves a 1000 times increase in CI space, a 10 6 -fold increase in floating-point operations performed, and a 10 6 -fold improvement in computational speed.

Computational chemistry↗

Coupling a recurrent neural network to SPAD TCSPC systems for real-time fluorescence lifetime imaging

Fluorescence lifetime imaging (FLI) has been receiving increased attention in recent years as a powerful diagnostic technique in biological and medical research. However, existing FLI systems often suffer from a tradeoff between processing speed, accuracy, and robustness. Inspired by the concept of Edge Artificial Intelligence (Edge AI), we propose a robust approach that enables fast FLI with no degradation of accuracy. This approach couples a recurrent neural network (RNN), which is trained to estimate the fluorescence lifetime directly from raw timestamps without building histograms, to SPAD TCSPC systems, thereby drastically reducing transfer data volumes and hardware resource utilization, and enabling real-time FLI acquisition. We train two variants of the RNN on a synthetic dataset and compare the results to those obtained using center-of-mass method (CMM) and least squares fitting (LS fitting). Results demonstrate that two RNN variants, gated recurrent unit (GRU) and long short-term memory (LSTM), are comparable to CMM and LS fitting in terms of accuracy, while outperforming them in the presence of background noise by a large margin. To explore the ultimate limits of the approach, we derive the Cramer-Rao lower bound of the measurement, showing that RNN yields lifetime estimations with near-optimal precision. To demonstrate real-time operation, we build a FLI microscope based on an existing SPAD TCSPC system comprising a 32 x 32 SPAD sensor named Piccolo. Four quantized GRU cores, capable of processing up to 4 million photons per second, are deployed on the Xilinx Kintex-7 FPGA that controls the Piccolo. Powered by the GRU, the FLI setup can retrieve real-time fluorescence lifetime images at up to 10 frames per second. The proposed FLI system is promising and ideally suited for biomedical applications, including biological imaging, biomedical diagnostics, and fluorescence-assisted surgery, etc.

47 OTHER INSTRUMENTATION↗

Tree tensor network hierarchical equations of motion based on time-dependent variational principle for efficient open quantum dynamics in structured thermal environments

In this work, we introduce an efficient method, TTN-HEOM, for exactly calculating the open quantum dynamics for driven quantum systems interacting with highly structured bosonic baths by combining the tree tensor network (TTN) decomposition scheme with the bexcitonic generalization of the numerically exact hierarchical equations of motion (HEOM). The method yields a series of quantum master equations for all core tensors in the TTN that efficiently and accurately capture the open quantum dynamics for non-Markovian environments to all orders in the system–bath interaction. These master equations are constructed based on the time-dependent Dirac–Frenkel variational principle, which isolates the optimal dynamics for the core tensors given the TTN ansatz. The dynamics converges to the HEOM when increasing the rank of the core tensors, a limit in which the TTN ansatz becomes exact. We introduce TENSO, tensor equations for non-Markovian structured open systems, as a general-purpose Python code to propagate the TTN-HEOM dynamics. We implement three general propagators for the coupled master equations: two fixed-rank methods that require a constant memory footprint during the dynamics and one adaptive-rank method with a variable memory footprint controlled by the target level of computational error. We exemplify the utility of these methods by simulating a two-level system coupled to a structured bath containing one Drude–Lorentz component and eight Brownian oscillators, which is beyond what can presently be computed using the standard HEOM. Our results show that the TTN-HEOM is capable of simulating both dephasing and relaxation dynamics of driven quantum systems interacting with structured baths, even those of chemical complexity, with an affordable computational cost.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Hybrid Finite-Volume, Discontinuous Galerkin Discretization for the Radiative Transport Equation

In this report we propose a hybrid spatial discretization for the radiative transport equation that combines a second-order discontinuous Galerkin (DG) method and a second-order finite-volume (FV) method. The strategy relies on a simple operator splitting that has been used previously to combine different angular discretizations. Unlike standard FV methods with upwind fluxes, the hybrid approach is able to accurately simulate problems in scattering dominated regimes. However, it requires less memory and yields a faster computational time than a uniform DG discretization. In addition, the underlying splitting allows naturally for hybridization in both space and angle. Numerical results are given to demonstrate the efficiency of the hybrid approach in the context of discrete ordinate angular discretizations and Cartesian spatial grids.

97 MATHEMATICS AND COMPUTING↗

Preparation of excited states for nuclear dynamics on a quantum computer

We study two different methods to prepare excited states on a quantum computer, a key initial step to study dynamics within linear response theory. The first method uses unitary evolution for a short time T = O(√1 - F) to approximate the action of an excitation operator Ô with fidelity F and success probability P ≈ 1 – F. The second method probabilistically applies the excitation operator using the Linear Combination of Unitaries (LCU) algorithm. We benchmark these techniques on emulated and real quantum devices, using a toy model for thermal neutron-proton capture. Despite its larger memory footprint, the LCU-based method is efficient even on current generation noisy devices and can be implemented at a lower gate cost than a naive analysis would suggest. Here, these findings show that quantum techniques designed to achieve good asymptotic scaling on fault tolerant quantum devices might also provide practical benefits on devices with limited connectivity and gate fidelity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AENET–LAMMPS and AENET–TINKER : Interfaces for accurate and efficient molecular dynamics simulations with machine learning potentials

Machine-learning potentials (MLPs) trained on data from quantum-mechanics based first-principles methods can approach the accuracy of the reference method at a fraction of the computational cost. To facilitate efficient MLP-based molecular dynamics and Monte Carlo simulations, an integration of the MLPs with sampling software is needed. Here, we develop two interfaces that link the atomic energy network (ænet) MLP package with the popular sampling packages TINKER and LAMMPS. The three packages, ænet, TINKER, and LAMMPS, are free and open-source software that enable, in combination, accurate simulations of large and complex systems with low computational cost that scales linearly with the number of atoms. Scaling tests show that the parallel efficiency of the ænet–TINKER interface is nearly optimal but is limited to shared-memory systems. The ænet–LAMMPS interface achieves excellent parallel efficiency on highly parallel distributed memory systems and benefits from the highly optimized neighbor list implemented in LAMMPS. We demonstrate the utility of the two MLP interfaces for two relevant example applications: the investigation of diffusion phenomena in liquid water and the equilibration of nanostructured amorphous battery materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large-Volume Injection and Assessment of Reference Standards for n -Alkane δD and δ 13 C Analysis via Gas Chromatography Isotope Ratio Mass Spectrometry

Compound-specific stable isotope analysis of hydrogen (δD) and carbon (δ 13 C) in organic compounds is a valuable tool in biogeochemical research. A key limitation of this method is the relatively large amount of sample required to achieve desirable precision. We developed a large-volume (20 μL) injection method that allows for high throughput analysis of less concentrated samples and tested it for δ 13 C and δD measurements of n-alkanes. We also conducted a comparison of reference standards and assessed several methods to normalize and correct n-alkane δD and δ13C measurements. The mean precision of the δD method based on 233 environmental n-alkane samples (two to three replications per sample) is 4.0‰ (1σ, estimated from the weighted mean of the pooled unbiased standard deviations) and 0.46‰ (1σ) for δ 13 C from 37 environmental samples (two to three replications per sample). The evaluation of reference standards shows that the use of n-alkane standards with large offsets in δD values in adjacent n-alkane chains can lead to biases in measurement correction. The large-volume injection method shows good reproducibility of δ 13 C and δD measurements of n-alkanes and reduces the required sample concentration by about 80%. We propose that for δD measurements, a reference standard set should be used in which each reference standard has a limited range of δD values and no adjacent n-alkane chains, to minimize memory effects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving net ecosystem CO2 flux prediction using memory-based interpretable machine learning

Terrestrial ecosystems play a central role in the global carbon cycle and affect climate change. However, our predictive understanding of these systems is still limited due to their complexity and uncertainty about how key drivers and their legacy effects influence carbon fluxes. Here, we propose an interpretable Long Short-Term Memory (iLSTM) network for predicting net ecosystem CO 2 exchange (NEE) and interpreting the influence on the NEE prediction from environmental drivers and their memory effects. We consider five drivers and apply the method to three forest sites in the United States. Besides performing the prediction in each site, we also conduct transfer learning by using the iLSTM model trained in one site to predict at other sites. Results show that the iLSTM model produces good NEE predictions for all three sites and, more importantly, it provides reasonable interpretations on the input driver's importance as well as their temporal importance on the NEE prediction. Additionally, the iLSTM model demonstrates good across-site transferability in terms of both prediction accuracy and interpretability. The transferability can improve the NEE prediction in unobserved forest sites, and the interpretability advances our predictive understanding and guides process-based model development.

Liu, Siyan↗

Data-driven based coordinated smart inverter control for distributed energy resources

Smart inverters (SI) for distributed energy resources (DER) are becoming popular since they have the ability to stabilize as well as restore the voltage and frequency of power systems. Aiming at establishing the mathematical models combined with SI control methods, multiple optimization methods are developed. However, the computational complexity of solving such a mathematical model with various uncertainties limits the real-time application of the SI control. To conquer this challenge, a data-driven-based SI control approach is developed to achieve coordinated control in the high penetration DER system. First, an optimization problem for maximizing the active power generation and minimizing the power loss is designed using the Volt/VAR control. To reduce the time consumption, the recurrent neural network (RNN) is proposed to model the relationship between the uncertainties and control actions during the offline site. The RNN with different sub-structures such as the long short-term memory cell and gated recurrent unit cell are included to enrich the diversity of features. In the last stage, different experiment comparisons, including multiple uncertainties maps and stateof- art machine learning methods, are conducted to verify the effectiveness of the proposed method based on the IEEE 123 bus power system. The results demonstrate that the proposed method can effectively achieve a rapid and coordinated control with a lower error rate.

Qiu, Wei↗

Predicting Kyasanur forest disease in resource-limited settings using event-based surveillance and transfer learning

In recent years, the reports of Kyasanur forest disease (KFD) breaking endemic barriers by spreading to new regions and crossing state boundaries is alarming. Effective disease surveillance and reporting systems are lacking for this emerging zoonosis, hence hindering control and prevention efforts. We compared time-series models using weather data with and without Event-Based Surveillance (EBS) information, i.e., news media reports and internet search trends, to predict monthly KFD cases in humans. We fitted Extreme Gradient Boosting (XGB) and Long Short-Term Memory models at the national and regional levels. We utilized the rich epidemiological data from endemic regions by applying Transfer Learning (TL) techniques to predict KFD cases in new outbreak regions where disease surveillance information was scarce. Overall, the inclusion of EBS data, in addition to the weather data, substantially increased the prediction performance across all models. The XGB method produced the best predictions at the national and regional levels. The TL techniques outperformed baseline models in predicting KFD in new outbreak regions. Novel sources of data and advanced machine-learning approaches, e.g., EBS and TL, show great potential towards increasing disease prediction capabilities in data-scarce scenarios and/or resource-limited settings, for better-informed decisions in the face of emerging zoonotic threats.

60 APPLIED LIFE SCIENCES↗