Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

120 records · Page 7

Recent Advances toward Efficient Calculation of Higher Nuclear Derivatives in Quantum Chemistry

In this article, we provide an overview of state-of-the-art techniques that are being developed for efficient calculation of second and higher nuclear derivatives of quantum mechanical (QM) energy. Calculations of nuclear Hessians and anharmonic terms incur high costs and memory and scale poorly with system size. Three emerging classes of methods—machine learning (ML), automatic differentiation (AD), and matrix completion (MC)—have demonstrated promise in overcoming these challenges. We illustrate studies that employ unsupervised ML methods to reduce the need for multiple Hessian calculations in dynamics simulations and those that utilize supervised ML to construct approximate potential energy surfaces and estimate Hessians and anharmonic terms at reduced cost. By extension, if electronic structure operations could be written in a manner similar to functions underlying ML methods, rapid differentiation or AD routines can be employed to inexpensively calculate higher arbitrary-order derivatives. While ML approaches are typically black-box, we describe methods such as compressed sensing (CS) and MC, which explicitly leverage problem-specific mathematical properties of higher derivatives such as sparsity and low-rank, to complete higher derivative information using only a small, incomplete sample. The three classes of methods facilitate reliable predictions of observables ranging from infrared spectra to thermal conductivity and constitute a promising way forward in accurately capturing otherwise intractable higher-order responses of QM energy to nuclear perturbations.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Predicting river turbidity in Pine Island Bayou using machine learning techniques coupled with variational mode decomposition

Elevated turbidity levels pose significant public health risks by facilitating the transport of harmful pollutants, including metals, organic compounds, and pathogenic microorganisms into the surface water. These conditions create serious challenges for public recreational water use and drinking water treatment, leading to economic losses and health risks. This study utilizes water monitoring data in Pine Island Bayou, Texas, and develops a Sequence-to-Sequence (S2S) model to predict turbidity using Attention-based Gated Recurrent Units with Encoder-Decoder (AT-GRU-ED) and Long Short-Term Memory (LSTM), coupled with Variational Mode Decomposition (VMD). Compared to the model without VMD, the model demonstrates satisfactory 72-hour turbidity prediction performance, achieving MAEs of 2.60 and 3.29 NTU (reductions of 53% and 58%), RMSEs of 21.08 and 31.49 NTU (reductions of 82% and 80%), and R² values of 0.96 and 0.84 on the validation and test sets, respectively. Feature importance analysis reveals that water temperature is the dominant factor influencing seasonal turbidity patterns, while real-time hourly rainfall significantly contributes to short-term variability. Turbidity typically peaks within 48 hours after rainfall events due to lagged effects from surface runoff and upstream flow. Findings suggest suspending recreational water use and water supply pumping for three days after heavy rainfall can benefit public health and improve water treatment processes. Discharges above 100 m3/s are found to accelerate sediment dilution and transport, reducing turbidity levels more quickly after the peak. In conclusion, the proposed model demonstrates reliable 72-hour turbidity prediction, supporting decision-making for water treatment plant operations and providing early warning for public recreational water use.

Deep learning↗

Identification and mitigation of memory block timing issue in ITk ABCStar during ASIC production

The ABCStar is a mixed-signal front-end readout ASIC for the strips sensor portion of the ATLAS ITk detector being developed as part of the High-Luminosity LHC upgrade. In pre-production testing, a subtle design flaw was uncovered in the ABCStar that was reducing wafer yields in some manufactured lots from the expected 90% to as low as 2%. The root cause was determined to be a timing issue in the logic synthesized to control previously silicon proven memory blocks re-used for this ASIC. The solutions proposed included manufacturing process changes by the wafer foundry, changes to the operating parameters for the ABCStar in the detector, and the possibility that a redesign might be required. The two mitigation efforts were undertaken in parallel, with the process modification route a less desirable solution since already manufactured wafers would need to be scrapped in favour of the new ones. Based on a knowledge of the existing process, and testing done on the worst performing wafers, it was proposed that raising the core operating voltage of the ABCStar from 1.20V to 1.25V could address the timing issue by sufficiently speeding up its transistors. An extensive testing program that included the effects of temperature and radiation expected over the lifetime of the ITk detector was conducted to validate that approach. Those tests and studies proved that even the worst performing wafers would have yields over 80% with the 1.25V core voltage, and neither the modified process nor redesign would be required for ensuring reliable operation of the ITk. Based on testing, a further timing mitigation was implemented to provide an additional margin of reliability by increasing the duty cycle of the clock to the ABCStar. Testing of all ABCStar wafers has been completed and the production of the detector modules using these ASICs is now well underway as a result of the efforts detailed herein.

FOS: Physical sciences↗

The future of subsurface monitoring: AEC’s breakthroughs in CCS technology

Carbon capture and storage (CCS) has emerged as a key solution in the fight against climate change. However, for CCS to succeed, it is crucial to ensure that the sequestered CO2 stays safely trapped underground. The U.S. Department of Energy (DOE) has emphasized the need for advancements in subsurface monitoring, measurement, reporting, and verification. Aside from caprock integrity failure, the other primary failure points usually involve defective cement in the casing annulus of wellbores or plugged and abandoned wells. In addition, many energy producers (e.g., oil and gas, geothermal) and storage and disposal operators (e.g., H2 and water) must deal with the same issue. Poorly placed or degraded cement can create pathways for gas or fluid to escape from casing annuli and in plugged and abandoned or orphan wells, posing environmental risks. Yet, a reliable and cost-effective way to monitor cement and well integrity over multiple decades is still unavailable. Traditional geophysical methods like 4D seismic imaging and surface-based electromagnetic monitoring lack the resolution and accuracy for detecting these types of failures (Vasco et al., 2022; Fawad and Mondol, 2021). Wireline logging is expensive to run continuously and is obtrusive to the operation. While fiber optics can potentially be a solution, its bulkiness can significantly compromise the cement's integrity. To address these challenges, the Advanced Energy Consortium (AEC) at The University of Texas at Austin’s Bureau of Economic Geology (the Bureau) has been pioneering research in subsurface monitoring using its portfolio of distributed autonomous microfabricated sensors for harsh subsurface environments since 2008. A class of these microsensors [System on a Chip (SoC)] can be mixed in cement and permanently placed without compromising the cement column; the sensors would then communicate with each other or a data acquisition (DAQ) master node. Another class of the AEC microsensors can be fully autonomous, with rechargeable micro-batteries capable of exceeding 100°C, flash memory, and, currently, a pressure and temperature sensor. They are designed to circulate in mud, geothermal fluids, U-loops, or pipelines. They can log data into memory and are unobtrusive to operations. Our team has been working on a multi-year DOE-funded project (DE-FE0031856)—supported by $2.95M in federal funding and $0.75M in cost-matching from the AEC—to demonstrate SoC sensor utility for CO2 leakage monitoring in CCS applications. This multi-institutional collaboration developed a novel sensing architecture utilizing radiofrequency (RF) microsensors embedded within the cement sheath. These sensors detect CO2 migration and are interrogated via a Smart Casing Collar (SCC).

58 GEOSCIENCES↗

Fourier-MIONet: Fourier-enhanced multiple-input neural operators for multiphase modeling of geological carbon sequestration

Geologic carbon sequestration (GCS) is a safety-critical technology that aims to reduce the amount of carbon dioxide in the atmosphere, which also places high demands on reliability. Multiphase flow in porous media is essential to understand CO 2 migration and pressure fields in the subsurface associated with GCS. However, numerical simulation for such problems in 4D is computationally challenging and expensive, due to the multiphysics and multiscale nature of the highly nonlinear governing partial differential equations (PDEs). It prevents us from considering multiple subsurface scenarios and conducting real-time optimization. Here, we develop a Fourier-enhanced multiple-input neural operator (Fourier-MIONet) to learn the solution operator of the problem of multiphase flow in porous media. Fourier-MIONet utilizes the recently developed framework of the multiple-input deep neural operators (MIONet) and incorporates the Fourier neural operator (FNO) in the network architecture. Once Fourier-MIONet is trained, it can predict the evolution of saturation and pressure of the multiphase flow under various reservoir conditions, such as permeability and porosity heterogeneity, anisotropy, injection configurations, and multiphase flow properties. Compared to the enhanced FNO (U-FNO), the proposed Fourier-MIONet has 90% fewer unknown parameters, and it can be trained in significantly less time (about 3.5 times faster) with much lower CPU memory (<15%) and GPU memory (<35%) requirements, to achieve similar prediction accuracy. In addition to the lower computational cost, Fourier-MIONet can be trained with only 6 snapshots of time to predict the PDE solutions for 30 years. Furthermore, we observed that Fourier-MIONet can maintain good accuracy when predicting out-of-distribution (OOD) data. The excellent generalizability of Fourier-MIONet is enabled by its adherence to the physical principle that the solution to a PDE is continuous over time. Furthermore, the developed Fourier-MIONet makes it possible to solve the long-time evolution of geological carbon sequestration in a large-scale three-dimensional space accurately and efficiently.

97 MATHEMATICS AND COMPUTING↗

Precursor Analysis Report: SQL Slammer Worm Infection of Davis-Besse Nuclear Power Plant 2003

The SQL Slammer Worm Infection of Davis-Besse Nuclear Power Plant 2003 Precursor Analysis Report leverages publicly available information about Davis-Besse’s 2003 cyber attack and catalogs anomalous observables for each technique employed in the attack. This analysis is based upon the methodology of the Cybersecurity for the Operational Technology Environment (CyOTE) program. On 25 January 2003, the SQL Slammer worm infected more than 90% of vulnerable hosts and crashed the internet in 10 to 15 minutes, making it one of the fastest spreading worms in history. SQL Slammer is a fileless, memory-resident worm that remotely exploits a stack-based buffer overflow vulnerability on local hosts to intensively scan and rapidly self-propagate across the internet. The worm infected approximately 300,000 unpatched hosts running Microsoft Structured Query Language (SQL) Server 2000 or Microsoft Desktop Engine (MSDE) 2000 with SQL Server Resolution Service. The SQL Slammer worm indirectly infected FirstEnergy’s Davis-Besse nuclear power plant by first infecting a consultant’s company network server and then propagating through an external misconfigured connection into Davis-Besse’s site network. The infection caused major network congestion, slow performance, data overloads, and the inability of local hosts to communicate with each other, which eventually caused a loss of availability and a loss of view when the Safety Parameter Display System (SPDS) and Plant Process Computer (PPC) crashed. At the time of the infection, the plant was already offline, the digital monitoring systems had redundant analog backups, and the plant control and safety functions were not affected, so there were no concerns of a safety breach. However, this incident resulted in many lessons learned and spawned important discussions about cybersecurity’s role in nuclear safety and electric power reliability regulation, policy, and guidance. Researchers and analysts identified 10 unique techniques utilized during the attack with a total of 640 observables using MITRE ATT&CK® for Industrial Control Systems. The CyOTE program assesses observables accompanying techniques used prior to the triggering event to identify opportunities to detect malicious activity. If observables accompanying the attack techniques are perceived and investigated prior to the triggering event, earlier comprehension of malicious activity can take place. Eight of the identified techniques used during Davis-Besse cyber attack were precursors to the triggering event. Analysis identified 596 observables associated with these precursor techniques, 428 of which were assessed to have an increased likelihood of being perceived in the 331 days preceding the triggering event. The response and comprehension time could have been reduced if the observables had been identified earlier. The information gathered in this report contributes to a library of observables tied to a repository of artifacts, data sources, and technique detection references for practitioners and developers to support the comprehension of indicators of attack. Asset owners and operators can use these products if they experience similar observables or to prepare for comparable scenarios.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Electrochemical Random-Access Memory: Progress, Perspectives, and Opportunities

Non-von Neumann computing using neuromorphic systems based on analogue synaptic and neuronal elements has emerged as a potential solution to tackle the growing need for more efficient data processing, but progress toward practical systems has been stymied due to a lack of materials and devices with the appropriate attributes. Recently, solid state electrochemical ion-insertion, also known as electrochemical random access memory (ECRAM) has emerged as a promising approach to realize the needed device characteristics. ECRAM is a three terminal device that operates by tuning electronic conductance in functional materials through solid-state electrochemical redox reactions. This mechanism can be considered as a gate-controlled bulk modulation of dopants and/or phases in the channel. Early work demonstrating that ECRAM can achieve nearly ideal analogue synaptic characteristics has sparked tremendous interest in this approach. More recently, the realization that electrochemical ion insertion can be used to tune the electronic properties of many types of materials including transition metal oxides, layered two-dimensional materials, organic and coordination polymers, and that the changes in conductance can span orders of magnitude has further attracted interest in ECRAM as the basis for analogue synaptic elements for inference accelerators as well as for dynamical devices that can emulate a wide range of neuronal characteristics for implementation in analogue spiking neural networks. At its core, ECRAM shares many fundamental aspects with rechargeable batteries, where ion insertion materials are used extensively for their ability to reversibly store charge and energy. Computing applications, however, present drastically different requirements: systems will require many millions of devices, scaled down to tens of nanometers, all while achieving reliable electronic-state tuning at scaled-up rates and endurances, and with minimal energy dissipation and noise. Further, in this review, we discuss the history, basic concepts, recent progress, as well as the challenges and opportunities for different types of ECRAM, broadly grouped by their primary mobile ionic charge carrier, including Li, protons, and oxygen vacancies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparative Analysis of ANN and LSTM Prediction Accuracy and Cooling Energy Savings through AHU-DAT Control in an Office Building

This paper proposes the optimal algorithm for controlling the HVAC system in the target building. Previous studies have analyzed pre-selected algorithms without considering the unique data characteristics of the target building, such as location, climate conditions, and HVAC system type. To address this, we compare the accuracy of cooling load prediction using ANN and LSTM algorithms, widely used in building energy research, to determine the optimal algorithm for HVAC control in the target building. We develop a simulation model calibrated with actual data to ensure data reliability and compare the energy consumption of the existing HVAC control method and the two algorithms-based methods. Results show that the ANN algorithm, with a CV(RMSE) of 12.7%, has a higher prediction accuracy than the LSTM algorithm, CV(RMSE) of 17.3%, making it a more suitable algorithm for HVAC control. Furthermore, implementing the ANN-based approach results in a 3.2% cooling energy reduction from the optimal control of Air Handling Unit (AHU) Discharge Air Temperature (DAT) compared to the fixed DAT at 12.8 °C in a representative day. This study demonstrates that ML-based HVAC system control can effectively reduce cooling energy consumption in HVAC systems, providing an effective strategy for energy conservation and improved HVAC system efficiency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Hardware and Software Co-design Framework for Energy Efficient Neuromorphic Systems

Neuromorphic systems can be realized by a variety of algorithms and architectures. A common understanding is that spiking neuromorphic designs, which encode information into spatio-temporal spiking events, are both a biologically-accurate and efficient way of processing information. However, representing the information through timing relationships induces sophisticated circuit designs in traditional CMOS-based implementations. In recent years, high-capacity resistive memory (RRAM, aka, memristor) has demonstrated great potential in mimicking synaptic behaviors. Several RRAM-based spiking neuromorphic designs exist, most of which focus on rate coding schemes. These designs simplify circuit implementations of neuron models and explore challenges such as unsatisfactory speed, resolution, and performance. As an alternative, we will explore temporal coding spiking neuromorphic systems that encode information as the relative timing of neuron activations (spikes), which have been proven to be more adaptive and energy-efficient. Developing a neuromorphic system for spiking neural network (SNN) inference and online training, however, faces some major technical challenges: (1) It lacks circuit implementation support for temporal-coding SNN to achieve satisfying power efficiency and accuracy; (2) Although existing research works have investigated memristive synapse and neuron designs for spike-timing-dependent plasticity, the non-ideal conditions in implementation, such as device variations and signal degradation, degrade online learning accuracy of large scale systems; and (3) Non-optimized, inter-layer data traffic in SNNs, leads to unnecessary data communication costs. In this project, we plan to address these challenges by a hardware and software co-design framework that incorporates solutions at the circuit, architecture, and algorithm levels. At the circuit-level, we will elaborate on the in-situ SNN processing element designs for supporting both inference and online training modes. Variation-aware schemes will be studied to improve reliability. At the architecture level, we propose a pipelined, asynchronous architecture to retain the timing resolution of spikes. At the algorithm level, we will investigate an innovative SNN training algorithm for enabling activation sparsification and reducing unnecessary data communication costs. This neuromorphic system will provide an effective solution to real-life energy-constrained applications and significantly contribute to the exploration of next-generation high-performance computing systems under the DOE context.

97 MATHEMATICS AND COMPUTING↗

Forecasting Solar-Thermal Systems Performance under Transient Operation Using a Data-Driven Machine Learning Approach Based on the Deep Operator Network Architecture

Modeling and prediction of the dynamic behavior of thermal systems operating under intermittent energy input and variable load requirements represent one of the greatest challenges in the development of efficient and reliable renewable-based power generation technologies. In this work, a data-driven machine learning modeling framework was developed based on a modified version of the Deep Operator Network architecture where the time coordinate in the trunk net is replaced with historical data of the predicting quantity. The modeling framework can be used to accurately predict the performance of renewable-based energy conversion technologies including wind- and solar-based power plants. This novel framework was applied on a solar-thermal system that consists of a solar collection loop using a flat plate collector, a power generation loop comprising an Organic Rankine Cycle, and a thermal energy storage tank connecting both loops. Variable solar irradiance, air temperature, and power load profiles were used by the Deep Operator Network to predict the State-of-Charge and the efficiency of the thermal system for several days. The results were compared with the State-of-Charge and efficiency functions calculated using a physics-based model. For a simple operation scenario, characterized by a clear sky solar irradiance profile and constant load, the standard deviation in the State-of-Charge prediction by Deep Operator Network is below 0.9% during a seven-day prediction time horizon. For the most realistic operation scenario that considers real solar irradiance and a rough load profile, the maximum standard deviation in the predictions for the State-of-Charge and efficiency are below 6.8% and 2.5%, respectively. A comparison between Deep Operator Network and Long Short Term Memory network was also performed. In general, both networks predict very well the State-of-Charge for different data density conditions; however, a higher accuracy, with a standard deviation below 2.0%, is obtained by the Deep Operator Network during three and half days using sparser training data of 20-minute points. The same accuracy for the State-of-Charge prediction with the Long Short Term Memory network is achieved only for 14 h. Average standard deviations for the State-of-Charge prediction of 1.1% with the Deep Operator Network and 1.5% with the Long Short Term Memory network are obtained for a four-day prediction time using a denser training data of 5-minute points.

DeepONet↗

A Perspective on ferroelectricity in hafnium oxide: Mechanisms and considerations regarding its stability and performance

Ferroelectric hafnium oxides are poised to impact a wide range of microelectronic applications owing to their superior thickness scaling of ferroelectric stability and compatibility with mainstream semiconductors and fabrication processes. For broad-scale impact, long-term performance and reliability of devices using hafnia will require knowledge of the phases present and how they vary with time and use. In this Perspective article, the importance of phases present on device performance is discussed, including the extent to which specific classes of devices can tolerate phase impurities. Following, the factors and mechanisms that are known to influence phase stability, including substituents, crystallite size, oxygen point defects, electrode chemistry, biaxial stress, and electrode capping layers, are highlighted. Herein, discussions will focus on the importance of considering both neutral and charged oxygen vacancies as stabilizing agents, the limited biaxial strain imparted to a hafnia layer by adjacent electrodes, and the strong correlation of biaxial stress with resulting polarization response. Areas needing additional research, such as the necessity for a more quantitative means to distinguish the metastable tetragonal and orthorhombic phases, quantification of oxygen vacancies, and calculation of band structures, including defect energy levels for pure hafnia and stabilized with substituents, are emphasized.

36 MATERIALS SCIENCE↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗