Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Memory management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Computational Performance Bounds Prediction in Quantum Computing With Unstable Noise

Quantum computing has significantly advanced in recent years, boasting devices with hundreds of quantum bits (qubits), hinting at its potential quantum advantage over classical computing. Yet, noise in quantum devices poses significant barriers to realizing this supremacy. Understanding noise’s impact is crucial for reproducibility and application reuse; moreover, the next-generation quantum-centric supercomputing essentially requires efficient and accurate noise characterization to support system management (e.g., job scheduling), where ensuring correct functional performance (i.e., fidelity) of jobs on available quantum devices can even be higher-priority than traditional objectives. However, noise fluctuates over time, even on the same quantum device, which makes predicting the computational bounds for on-the-fly noise is vital. Noisy quantum simulation can offer insights but faces efficiency and scalability issues. Here, in this work, we propose a data-driven workflow, namely QuBound, to predict computational performance bounds. It decomposes historical performance traces to isolate noise sources and devises a novel encoder to embed circuit and noise information processed by a Long Short-Term Memory (LSTM) network. For evaluation, we compare QuBound with a state-of-the-art learning-based predictor, which only generates a single performance value instead of a bound. Experimental results show that the result of the existing approach falls outside of performance bounds, while all predictions from our QuBound with the assistance of performance decomposition better fit the bounds. Moreover, QuBound can efficiently produce practical bounds for various circuits with over 106 speedup over simulation; in addition, the range from QuBound is over 10× narrower than the state-of-the-art analytical approach.

Li, Jinyang [George Mason Univ., Fairfax, VA (Unit↗

Explainable machine learning model for multi-step forecasting of reservoir inflow with uncertainty quantification

We propose an explainable machine learning (ML) model with uncertainty quantification (UQ) to improve multi-step reservoir inflow forecasting. Traditional ML methods have challenges in forecasting inflows multiple days ahead, and lack explainability and UQ. To address these limitations, we introduce an encoder–decoder long short-term memory (ED-LSTM) network for multi-step forecasting, employ the SHapley Additive exPlanation (SHAP) technique for understanding the influence of hydrometeorological factors on inflow prediction, and develop a novel UQ method for prediction trustworthiness. We apply these methods to forecast 7-day inflow in snow-dominant and rain-driven reservoirs. The results demonstrate the effectiveness of the ED-LSTM model, with high forecasting accuracy for short lead times. Our UQ method provides reliable uncertainty estimates, covering 90% of data with a 90% confidence level. The SHAP analysis reveals the importance of historical inflow and precipitation as influential factors. These findings and methods may support reservoir operators in optimizing water resources management decisions.

54 ENVIRONMENTAL SCIENCES↗

Athena: High-Performance Sparse Tensor Contraction Sequence on Heterogeneous Memory

Sparse tensor contraction sequence has been widely employed in many fields, such as chemistry and physics. However, how to efficiently implement the sequence faces multiple challenges, such as redundant computations and memory operations, massive memory consumption, and inefficient utilization of hardware. To address the above challenges, we introduce Athena, a high-performance framework for SpTC sequences. Athena introduces new data structures, leverages emerging Optane-based heterogeneous memory (HM) architecture, and stage parallelism. In particular, Athena introduces shared hash table-represented sparse accumulator to eliminate unnecessary input processing and data migration; Athena uses a novel data-semantic guided dynamic migration solution to make the best use of the Optane-based HM for high performance; Athena also co-runs execution phases with different characteristics to enable high hardware utilization. Evaluating with 12 datasets, we show that Athena brings 327-7362× speedup over the state-of-the-art SpTC algorithm. With the dynamic data placement guided by data semantics, Athena brings performance improvement on Optane-based HM over a state-of-the-art software-based data management solution, a hardware-based data management solution, and PMM-only by 1.58×, 1.82×, and 2.34× respectively.

Liu, Jiawen↗

Cryogenic energy storage: Standalone design, rigorous optimization and techno-economic analysis

Energy storage allows flexible use and management of excess electricity and intermittently available renewable energy. Cryogenic energy storage (CES) is a promising storage alternative with a high technology readiness level and maturity, but the round-trip efficiency is often moderate and the Levelized Cost of Storage (LCOS) remains high. The complex flowsheets with intricate thermodynamics at cryogenic temperatures as well as the presence of multiple loops and refrigeration cycles pose considerable challenges for rigorous model-based design and optimization of CES systems. We present an optimization strategy that couples rigorous process simulation and Bayesian optimization with flowsheet decomposition and identification of hidden coupling constraints to optimally design standalone CES systems. Further refinement is done via a local search using the limited-memory Broyden–Fletcher–Goldfarb–Shanno algorithm. Here our results indicate that it is possible to achieve more than 52% round-trip efficiency and an LCOS of $153/MWh for a standalone 100 MW/400 MWh CES system limited to short-term storage with daily charging–discharging. However, a detailed techno-economic assessment reveals that the LCOS considering total capital investment may exceed $267/MWh when all direct and indirect costs of installation and operation are considered.

25 ENERGY STORAGE↗

Determinants of residential mobility: an adaptive retrospective survey method

This study introduces a survey instrument to collect retrospective life-course events, focusing on residential relocation, and it utilizes the survey to evaluate determinants of residential mobility. Here, the survey consists of seven modules collecting information about household structure and household demographics, latest residential relocation, current and previous home, employment, education, vehicle ownership, and travel behavior. The time window of the life-course calendar in the survey is customized according to the latest residential relocation as an anchor point to assist with memory recollection and balance the required input from participants. The survey is used to collect data from a sample of 514 respondents in Sydney, Australia, and another sample of 404 respondents in Chicago, Illinois, in the US. The Cox proportional hazard model is used to analyze residential mobility. The results show primary school commencement is a salient determinant of residential relocation, but its impact is significantly higher in Chicago compared to Sydney.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

UPC++ v1.0 Programmer’s Guide (Rev. 2023.9.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide (Revision 2022.3.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2023.3.0

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2022.9.0

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

Identifying Hydrometeorological Factors Influencing Reservoir Releases Using Machine Learning Methods

Simulation of reservoir releases plays a critical role in social-economic functioning and our nation's security. How-ever, it is challenging to predict the reservoir release accurately because of many influential factors from natural environments and engineering controls such as the reservoir inflow and storage. Moreover, climate change and hydrological intensification causing the extreme precipitation and temperature make the accurate prediction of reservoir releases even more challenging. Machine learning (ML) methods have shown some successful applications in simulating reservoir releases. However, previous studies mainly used inflow and storage data as inputs and only considered their short-term influences (e.g, previous one or two days). In this work, we use long short-term memory (LSTM) networks for reservoir release prediction based on four input variables including inflow, storage, precipitation, and temperature and consider their long-term influences. We apply the LSTM model to 30 reservoirs in Upper Colorado River Basin, United States. We analyze the prediction performance using six statistical metrics. More importantly, we investigate the influence of the input hydrometeorological factors, as well as their temporal effects on reservoir release decisions. Results indicate that inflow and storage are the most influential factors but the inclusion of precipitation and temperature can further improve the prediction of release especially in low flows. Additionally, the inflow and storage have a relatively long-term effect on the release. These findings can help optimize the water resources management in the reservoirs.

Fan, Ming↗

A New Evaluation Metric for Demand Response-Driven Real-Time Price Prediction Towards Sustainable Manufacturing

Abstract The increasing industry energy demand highlights the urgency of demand response management, while the emerging smart manufacturing technologies pave the way for the implementation of real-time price (RTP)-based demand response management towards sustainable manufacturing. The demand response management requires scheduling of manufacturing systems based on RTP predictions, and thus the prediction quality can directly alter the effectiveness of demand response. However, since the general price prediction algorithms and prediction evaluation metrics are not specifically designed for RTP in demand response problems, a good RTP prediction obtained and evaluated by these algorithms and metrics may not be suitable for demand response scheduling. Therefore, in this study, the relationships between the effectiveness of demand response for manufacturing systems and evaluation results from six commonly used metrics are investigated. Meanwhile, a new metric called k-peak distance (KPD), considering the characteristics of the demand response problem, is proposed and compared with the other six metrics. Furthermore, an encoder-decoder long short-term memory recurrent neural network with KPD is proposed to provide better RTP prediction for manufacturing demand response problems. The case studies indicate that the proposed KPD metric shows a 1.8–3.6 times higher correlation with the demand response effectiveness compared to the other metrics. In addition, the production schedule based on the RTP prediction obtained from the proposed algorithm can improve the effectiveness of demand response by 23.4% on average.

Engineering↗

Intrinsic and environmental drivers of pairwise cohesion in wild Canis social groups

Animals within social groups respond to costs and benefits of sociality by adjusting the proportion of time they spend in close proximity to other individuals in the group (cohesion). Variation in cohesion between individuals, in turn, shapes important group-level processes such as subgroup formation and fission–fusion dynamics. Although critical to animal sociality, a comprehensive understanding of the factors influencing cohesion remains a gap in our knowledge of cooperative behavior in animals. We tracked 574 individuals from six species within the genus Canis in 15 countries on four continents with GPS telemetry to estimate the time that pairs of individuals within social groups spent in close proximity and test hypotheses regarding drivers of cohesion. Pairs of social canids (Canis spp.) varied widely in the proportion of time they spent together (5%–100%) during seasonal monitoring periods relative to both intrinsic characteristics and environmental conditions. The majority of our data came from three species of wolves (gray wolves, eastern wolves, and red wolves) and coyotes. For these species, cohesion within social groups was greatest between breeding pairs and varied seasonally as the nature of cooperative activities changed relative to annual life history patterns. Across species, wolves were more cohesive than coyotes. For wolves, pairs were less cohesive in larger groups, and when suitable, small prey was present reflecting the constraints of food resources and intragroup competition on social associations. Pair cohesion in wolves declined with increased anthropogenic modification of the landscape and greater climatic variability, underscoring challenges for conserving social top predators in a changing world. We show that pairwise cohesion in social groups varies strongly both within and across Canis species, as individuals respond to changing ecological context defined by resources, competition, and anthropogenic disturbance. Our work highlights that cohesion is a highly plastic component of animal sociality that holds significant promise for elucidating ecological and evolutionary mechanisms underlying cooperative behavior.

59 BASIC BIOLOGICAL SCIENCES↗

Demystifying asynchronous I/O Interference in HPC applications

With increasing complexity of HPC workflows, data management services need to perform expensive I/O operations asynchronously in the background, aiming to overlap the I/O with the application runtime. However, this may cause interference due to competition for resources: CPU, memory/network bandwidth. The advent of multi-core architectures has exacerbated this problem, as many I/O operations are issued concurrently, thereby competing not only with the application but also among themselves. Furthermore, the interference patterns can dynamically change as a response to variations in application behavior and I/O subsystems (e.g. multiple users sharing a parallel file system). Without a thorough understanding, I/O operations may perform suboptimally, potentially even worse than in the blocking case. To fill this gap, here we investigate the causes and consequences of interference due to asynchronous I/O on HPC systems. Specifically, we focus on multi-core CPUs and memory bandwidth, isolating the interference due to each resource. Then, we perform an in-depth study to explain the interplay and contention in a variety of resource sharing scenarios such as varying priority and number of background I/O threads and different I/O strategies: sendfile, read/write, mmap/write underlining trade-offs. The insights from this study are important both to enable guided optimizations of existing background I/O, as well as to open new opportunities to design advanced asynchronous I/O strategies.

97 MATHEMATICS AND COMPUTING↗

GASNet-EX Specification Collection (Rev. 2024.5.0)

GASNet-EX is a portable, open-source, high-performance communication library designed to efficiently support the networking requirements of PGAS runtime systems and other alternative models in emerging exascale systems. It provides network-independent, high-performance communication primitives including Remote Memory Access (RMA) and Active Messages (AM). GASNet-EX is an evolution of the popular GASNet communication system, building upon over 20 years of lessons learned, and the primary goals are high performance, interface portability, and expressiveness. The library has been used to implement parallel programming models and libraries such as UPC, UPC++, Fortran coarrays, Legion, Chapel, and many others. This anthology collects together the four separate volumes that currently comprise the GASNet-EX specification, as of the 2024.5.0 release of GASNet-EX.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Development and Test of a Large Aperture $Nb_3Sn$ Cos-Theta Dipole Coil with Stress Management

An innovative stress-management (SM) concept for cos-theta (CT) coils (SMCT coil concept) has been proposed at Fermilab. A large-aperture two-layer Nb3Sn SMCT dipole coil was designed and manufactured to validate and test the SM concept including coil design, fabrication technology, and performance. The first large-aperture SMCT coil (SMCT1) was fabricated and assembled with a small-aperture Nb3Sn coil inside a dipole mirror magnet. SMCT1 coil tests in a dipole mirror structure was performed in two configurations - SMCTM1a with only powered two-layer SMCT1 coil and SMCTM1b with the SMCT coil connected in series with an inner two-layer dipole coil. The test goals are to prove the SMCT coil concept in two-layer and four-layer mirror configurations; demonstrate that the magnet can reach the targeted quench current at the established preload; study magnet training, training memory after thermal cycle, ramp rate and temperature dependences of the magnet quench current; and test the SMCT1 coil quench protection parameters. This paper summarizes the SMCT1 coil design and parameters, the coil main fabrication steps, its assembly in the dipole mirror structure. The results of the SMCTM1a/b mirror test are presented and discussed.

43 PARTICLE ACCELERATORS↗

Comparison of Machine Learning-Based Predictive Models of the Nutrient Loads Delivered from the Mississippi/Atchafalaya River Basin to the Gulf of Mexico

Predicting nutrient loads is essential to understanding and managing one of the environmental issues faced by the northern Gulf of Mexico hypoxic zone, which poses a severe threat to the Gulf’s healthy ecosystem and economy. The development of hypoxia in the Gulf of Mexico is strongly associated with the eutrophication process initiated by excessive nutrient loads. Due to the complexities in the excessive nutrient loads to the Gulf of Mexico, it is challenging to understand and predict the underlying temporal variation of nutrient loads. The study was aimed at identifying an optimal predictive machine learning model to capture and predict nonlinear behavior of the nutrient loads delivered from the Mississippi/Atchafalaya River Basin (MARB) to the Gulf of Mexico. For this purpose, monthly nutrient loads (N and P) in tons were collected from US Geological Survey (USGS) monitoring station 07373420 from 1980 to 2020. Machine learning models—including autoregressive integrated moving average (ARIMA), gaussian process regression (GPR), single-layer multilayer perceptron (MLP), and a long short-term memory (LSTM) with the single hidden layer—were developed to predict the monthly nutrient loads, and model performances were evaluated by standard assessment metrics—Root Mean Square Error (RMSE) and Correlation Coefficient (R). The residuals of predictive models were examined by the Durbin–Watson statistic. The results showed that MLP and LSTM persistently achieved better accuracy in predicting monthly TN and TP loads compared to GPR and ARIMA. In addition, GPR models achieved slightly better test RMSE score than ARIMA models while their correlation coefficients are much lower than ARIMA models. Moreover, MLP performed slightly better than LSTM in predicting monthly TP loads while LSTM slightly outperformed for TN loads. Furthermore, it was found that the optimizer and number of inputs didn’t show effects on the LSTM performance while they exhibited impacts on MLP outcomes. This study explores the capability of machine learning models to accurately predict nonlinearly fluctuating nutrient loads delivered to the Gulf of Mexico. Further efforts focus on improving the accuracy of forecasting using hybrid models which combine several machine learning models with superior predictive performance for nutrient fluxes throughout the MARB.

54 ENVIRONMENTAL SCIENCES↗

Xanthos-Lake Model Source Code

This repository contains the source code for Xanthos-Lake, a lake-modeling extension of the Xanthos framework that introduces a coupled lake component comprising the Xanthos-Lake Snow and Ice Model (xLSIM) and the Xanthos-Lake Water Balance Model (xLWBM). xLSIM is a basin-aware machine-learning model for lake snow, ice, and thermal conditions. It predicts monthly lake ice thickness, snow depth, snow-cover fraction, mixing-layer temperature, and lake ice fraction from meteorological forcing and lake surface-area information. It uses sequence-based deep-learning architectures, including Transformer and hybrid Long Short-Term Memory–Transformer (LSTM–Transformer) models, together with seasonal encoding, multi-lake learning, physical masking, and basin-level cryospheric and non-cryospheric classification. The training workflow uses Ray for scalable execution and includes optional Ray Tune hyperparameter optimization. Model predictions, observations, diagnostics, and feature-importance outputs are written in NetCDF. xLWBM is the water-balance component of the new lake framework. It simulates monthly lake storage, surface area, evaporation, inflow, outflow, and lake–groundwater exchange. It combines physical water-balance equations with calibrated bathymetric relationships, weir-based outlet flow, modified Penman open-water evaporation, groundwater head relaxation, Penman–Monteith snow and ice sublimation, and snow, ice, and thermal conditions supplied by xLSIM. The model calibrates lake parameters against satellite-derived surface-area data, using evaporation-based calibration where surface-area data are unavailable, and supports small, medium, and large lake classes. For large lakes, xLWBM is integrated with the managed-routing workflow so that lake storage and outflow interact directly with downstream river routing and reservoir operations. Together, xLSIM and xLWBM provide Xanthos with a coupled lake-modeling capability. xLSIM supplies the snow, ice, and thermal conditions that affect lake evaporation and snow- and ice-related water exchanges, while xLWBM translates those conditions into dynamic lake storage, surface area, evaporation, and discharge. In return, xLWBM supplies evolving lake surface area to xLSIM. This coupling enables Xanthos to represent lakes as active hydrologic components within basin-scale water-availability and routing simulations.

Machine Learning↗

UPC++ v1.0 Programmer’s Guide, Revision 2021.9.0

UPC++ is a C++ library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. PGAS additionally provides one-sided Remote Memory Access (RMA) to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. In UPC++, all communication operations are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all communication operations are asynchronous by default, to enable programmers to write code that scales well even on hundreds of thousands of cores.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗