Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Deep Operator Networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Reconstruction of fast neutron direction in segmented organic detectors using deep learning

A method for reconstructing the direction of a fast neutron source using a segmented organic scintillator-based detector and deep learning model is proposed and analyzed. Here, the model is based on recurrent neural network, which can be trained by a sequence of data obtained from an event recorded in the detector and suitably pre-processed. The performance of deep learning-based model is compared with the conventional double-scatter detection algorithm in reconstructing the direction of a fast neutron source. With the deep learning model, the uncertainty in source direction of 0.301 rad is achieved with 100 neutron detection events in a segmented cubic organic scintillator detector with a side length of 46 mm. To reconstruct the source direction with the same angular resolution as the double-scatter algorithm, the deep learning method requires 75% fewer events. Application of this method could augment the operation of segmented detectors operated in the neutron scatter camera configuration for applications such as special nuclear material detection.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Applying Deep Learning for Wildfire Identification: Economical and Accessible Solutions Leveraging Small Datasets

Wildfires significantly impact human health, air quality, visibility, weather, and climate change and cause substantial economic losses. While state and county-operated air quality monitors provide critical insights during wildfires, they are not available in all regions. This highlights the need for affordable, accessible tools that allow the general public to assess air quality impacts. In this study, we apply machine learning with deep neural networks to diagnose air quality rapidly from sky images taken at the Pacific Northwest National Laboratory in Richland, WA, USA. Using a convolutional neural network (CNN) framework, we trained a deep learning model to classify air quality indices based on sky images. By leveraging transfer learning, our approach fine-tunes a pre-trained model on a small dataset of sky images, significantly reducing training time while maintaining high accuracy. Our results demonstrate the potential of deep learning to provide rapid air quality diagnostics during wildfire episodes, offering early warnings to the public and enabling timely mitigation strategies, particularly for vulnerable populations. Additionally, we show that lower respiratory infections pose the highest health risk during acute smoke exposures. Reactive oxygen species (ROS) from wildfire particles further exacerbate health risks by triggering inflammation and other adverse effects.

54 ENVIRONMENTAL SCIENCES↗

Deep learning based event reconstruction for cyclotron radiation emission spectroscopy

The objective of the cyclotron radiation emission spectroscopy (CRES) technology is to build precise particle energy spectra. This is achieved by identifying the start frequencies of charged particle trajectories which, when exposed to an external magnetic field, leave semi-linear profiles (called tracks) in the time–frequency plane. Due to the need for excellent instrumental energy resolution in application, highly efficient and accurate track reconstruction methods are desired. Deep learning convolutional neural networks (CNNs) - particularly suited to deal with information-sparse data and which offer precise foreground localization—may be utilized to extract track properties from measured CRES signals (called events) with relative computational ease. In this work, we develop a novel machine learning based model which operates a CNN and a support vector machine in tandem to perform this reconstruction. A primary application of our method is shown on simulated CRES signals which mimic those of the Project 8 experiment—a novel effort to extract the unknown absolute neutrino mass value from a precise measurement of tritium β - -decay energy spectrum. When compared to a point-clustering based technique used as a baseline, we show a relative gain of 24.1% in event reconstruction efficiency and comparable performance in accuracy of track parameter reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

New data-driven approach to bridging power system protection gaps with deep learning

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this paper, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A combined convolutional neural network and long short-term memory (CNN-LSTM) network is implemented to achieve data translation invariance and capture the temporal correlation of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN-LSTM--based method avoids the complicated, manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on two kinds of protection gaps: high-impedance faults and transformer inter-turn faults. Lastly, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

42 ENGINEERING↗

The optimal use of segmentation for sampling calorimeters

One of the key design choices of any sampling calorimeter is how fine to make the longitudinal and transverse segmentation. Here, to inform this choice, we study the impact of calorimeter segmentation on energy reconstruction. To ensure that the trends are due entirely to hardware and not to a sub-optimal use of segmentation, we deploy deep neural networks to perform the reconstruction. These networks make use of all available information by representing the calorimeter as a point cloud. To demonstrate our approach, we simulate a detector similar to the forward calorimeter system intended for use in the ePIC detector, which will operate at the upcoming Electron Ion Collider. We find that for the energy estimation of isolated charged pion showers, relatively fine longitudinal segmentation is key to achieving an energy resolution that is better than 10% across the full phase space. These results provide a valuable benchmark for ongoing EIC detector optimizations and may also inform future studies involving high-granularity calorimeters in other experiments at various facilities.

47 OTHER INSTRUMENTATION↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

Predicting future well performance for environmental remediation design using deep learning

Here in this study, we developed a deep learning (DL) framework with a multi-channel three-dimensional convolutional neural network (MC3D-CNN) to predict well performance and thereby assist future environmental remediation design. Such prediction of extraction well performance at designated locations is critical for configuring pump-and-treat (P&T) well network design and operation, setting reasonable target closure dates for overall remedying, and estimating remedy costs. The framework is developed with operational and monitoring data routinely collected during P&T remedy operations, including well extraction and injection rates as well as in situ contaminant concentrations. Traditionally, the collected data were rarely used for purposes other than assessing past well performance and the accuracy of the conceptual site model. However, recent advances in data-driven computational approaches enable better use of the large datasets to inform future well performance, enhance site characterization, and improve remediation planning. In this study, we established a DL framework to integrate transient three-dimensional contaminant plumes and multiple aquifer properties (e.g., hydraulic conductivity and hydrostratigraphic maps) to identify characteristic patterns controlling and representing extraction well mass recovery, aiming at providing future mass recovery estimates for existing wells and candidate wells at any proposed locations. We evaluated our framework by using a realistic synthetic dataset generated from a well-calibrated flow and transport model used in the 200 West Area of the U.S. Department of Energy’s Hanford Site in southeastern Washington state. The multi-channel feature in our framework allows integration of various types and temporal densities of training datasets for DL model development. Overall, we found that the trained DL model achieved an accuracy of over 90% in ranking extraction well performance in validation datasets, and over 80% in predicting high-performance-ranking well locations. This data-informed approach provides a flexible tool to support adaptive site management, streamline decision-making, and potentially reduce remediation time and costs. Our DL framework can be used as a filtering tool to improve the current P&T network optimization design by reducing the number of candidate well locations.

54 ENVIRONMENTAL SCIENCES↗

Quantum Ising model on (2+1)-dimensional anti–de Sitter space using tensor networks

We study the quantum Ising model on (2+1)-dimensional anti-de Sitter space using matrix product states (MPS) and matrix product operators (MPOs). We explore the bulk phase diagram of the theory on regular tessellations of hyperbolic space with coordination number seven and find disordered and ordered phases separated by a phase transition. We find that the boundary-boundary spin correlation function exhibits power law scaling deep in the disordered phase of the Ising model consistent with holography. At the critical point, we find the boundary entanglement entropy scales logarithmically with subsystem size but away from this, we see a linear scaling. In comparison, the full system exhibits a volume law scaling, which is expected in chaotic and/or highly connected systems. We also measure out of time ordered correlators (OTOCs) to explore the scrambling behavior of the theory.

Quantum spin models↗

Strictly Enforcing Invertibility and Conservation in CNN-Based Super Resolution for Scientific Datasets

Abstract Recently, deep convolutional neural networks (CNNs) have revolutionized image “super resolution” (SR), dramatically outperforming past methods for enhancing image resolution. They could be a boon for the many scientific fields that involve imaging or any regularly gridded datasets: satellite remote sensing, radar meteorology, medical imaging, numerical modeling, and so on. Unfortunately, while SR-CNNs produce visually compelling results, they do not necessarily conserve physical quantities between their low-resolution inputs and high-resolution outputs when applied to scientific datasets. Here, a method for “downsampling enforcement” in SR-CNNs is proposed. A differentiable operator is derived that, when applied as the final transfer function of a CNN, ensures the high-resolution outputs exactly reproduce the low-resolution inputs under 2D-average downsampling while improving performance of the SR schemes. The method is demonstrated across seven modern CNN-based SR schemes on several benchmark image datasets, and applications to weather radar, satellite imager, and climate model data are shown. The approach improves training time and performance while ensuring physical consistency between the super-resolved and low-resolution data. Significance Statement Recent advancements in using deep learning to increase the resolution of images have substantial potential across the many scientific fields that use images and image-like data. Most image super-resolution research has focused on the visual quality of outputs, however, and is not necessarily well suited for use with scientific data where known physics constraints may need to be enforced. Here, we introduce a method to modify existing deep neural network architectures so that they strictly conserve physical quantities in the input field when “super resolving” scientific data and find that the method can improve performance across a wide range of datasets and neural networks. Integration of known physics and adherence to established physical constraints into deep neural networks will be a critical step before their potential can be fully realized in the physical sciences.

54 ENVIRONMENTAL SCIENCES↗

Generalized Quantum Convolution for Multidimensional Data

The convolution operation plays a vital role in a wide range of critical algorithms across various domains, such as digital image processing, convolutional neural networks, and quantum machine learning. In existing implementations, particularly in quantum neural networks, convolution operations are usually approximated by the application of filters with data strides that are equal to the filter window sizes. One challenge with these implementations is preserving the spatial and temporal localities of the input features, specifically for data with higher dimensions. In addition, the deep circuits required to perform quantum convolution with a unity stride, especially for multidimensional data, increase the risk of violating decoherence constraints. In this work, we propose depth-optimized circuits for performing generalized multidimensional quantum convolution operations with unity stride targeting applications that process data with high dimensions, such as hyperspectral imagery and remote sensing. We experimentally evaluate and demonstrate the applicability of the proposed techniques by using real-world, high-resolution, multidimensional image data on a state-of-the-art quantum simulator from IBM Quantum.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Bilevel optimization, deep learning and fractional Laplacian regularization with applications in tomography

Here we consider a generalized bilevel optimization framework for solving inverse problems. We introduce fractional Laplacian as a regularizer to improve the reconstruction quality, and compare it with the total variation regularization. We emphasize that the key advantage of using fractional Laplacian as a regularizer is that it leads to a linear operator, as opposed to the total variation regularization which results in a nonlinear degenerate operator. Inspired by residual neural networks, to learn the optimal strength of regularization and the exponent of fractional Laplacian, we develop a dedicated bilevel optimization neural network with a variable depth for a general regularized inverse problem. We illustrate how to incorporate various regularizer choices into our proposed network. As an example, we consider tomographic reconstruction as a model problem and show an improvement in reconstruction quality, especially for limited data, via fractional Laplacian regularization. We successfully learn the regularization strength and the fractional exponent via our proposed bilevel optimization neural network. We observe that the fractional Laplacian regularization outperforms total variation regularization. This is specially encouraging, and important, in the case of limited and noisy data.

97 MATHEMATICS AND COMPUTING↗

Physics Informed Reinforcement Learning for Power Grid Control using Augmented Random Search

Wide adoption of deep reinforcement learning need to overcome several challenges in energy system domain, including scalability, learning from limited samples, and high-dimensional continuous state and action spaces. In this paper, we integrated physics-based information from the normal generator operation state formula in the reinforcement learning agent's neural network loss function, and applied an augmented random search agent to optimize the generator control under dynamic contingency. Simulation results demonstrated the reliability performance improvements in training speed, reward convergence, sampling efficiency, scalability, and transferability.

physics informed ML, Physics Informed Neural Netwo↗

NLML: A Deep Neural Network Emulator for the Exact Nonlinear Interactions in a Wind Wave Model

Nonlinear wave interactions describe the resonant energy transfer between wave components, playing a fundamental role in the evolution of ocean wave spectra. Nonlinear wave interactions significantly influence wave growth and development, making them essential for accurate wave modeling. However, resolving the full six-dimensional Boltzmann integral of the exact nonlinear wave interactions (Webb-Resio-Tracy method, WRT) is computationally expensive, limiting its application in real-time operational wave forecasting and for research purposes. Current approximations, such as the Discrete Interaction Approximation (DIA), prioritize computational speed over accuracy, resulting in significant errors in wave mean parameters. Here, we introduce NLML, a machine learning (ML) emulator designed to approximate the exact nonlinear wave interactions within WAVEWATCH III (WW3), with the goal of achieving the accuracy of WRT while maintaining the stability and computational speed of DIA. By leveraging GPU capabilities such as half precision inference, we achieved substantial speedups, up to 136x mathematical equation faster than the WRT and only a modest 1.04x mathematical equation slowdown relative to DIA, while achieving 2x mathematical equation the accuracy of DIA in global wave spectral energy and mean wave parameters, with up to 7x mathematical equation higher accuracy in some regions. Unlike previous ML approaches, NLML maintained inherent stability throughout model integration in a standalone, year-long WW3 simulation, without requiring additional constraints. Our new ML parameterization bridges the gap between accuracy and efficiency, offering a promising alternative for improving wave modeling in operational settings and research purposes.

16 TIDAL AND WAVE POWER↗

Deep Reinforcement Learning for Microgrid Cost Optimization Considering Load Flexibility

This paper proposes a novel Soft-Actor-Critic (SAC) based Deep Reinforcement Learning (DRL) method for optimizing the cost of microgrid operation by leveraging load flexibility. The proposed SAC-DRL method is designed to coordinate the control of distributed energy resources (DERs) and flexible load, addressing practical energy billing formation by power distribution utilities. Key contributions include an innovative reward function to mitigate sparse reward challenges and a mixed control strategy for discrete and continuous variables, ensuring radial network topology and minimizing power loss. We evaluate the proposed method on the model of a real microgrid located in Southern California, U.S.. The SAC-DRL model is tested to demonstrate its efficacy in reducing grid dependence, optimizing resource use, and minimizing costs. The results highlight the potential of DRL in modern energy systems, offering a sustainable and economically efficient solution for energy management in microgrids.

deep reinforcement learning↗

MACHINE LEARNING-ENABLED PREDICTION OF TRANSIENT INJECTION MAP IN AUTOMOTIVE INJECTORS WITH UNCERTAINTY QUANTIFICATION

Accurate prediction of injection profiles is a critical aspect of linking injector operation with engine performance and emissions. However, highly resolved injector simulations can take one to two weeks of wall-clock time, which is incompatible with engine design cycles with desired turnaround times of less than a day. Hence, it is important to reduce the time-to-solution of the internal flow simulations by several orders of magnitude to make it compatible with engine simulations. This work demonstrates a data-driven approach for tackling the computational overhead of injector simulations, whereby the transient injection profiles are emulated for a side-oriented, single-hole diesel injector using a Bayesian machine-learning framework. First, an interpretable Bayesian learning strategy was employed to understand the effect of design parameters on the total void fraction field. Then, autoencoders are utilized for efficient dimensionality reduction of the flowfields. Gaussian process models are finally used to predict the spatiotemporal void fraction field at the injector exit for unknown operating conditions. The Gaussian process models produce principled uncertainty estimates associated with the emulated flowfields, which provide the engine designer with valuable information of where the data-driven predictions can be trusted in the design space. The Bayesian flowfield predictions are compared with the corresponding predictions from a deep neural network, which has been transfer-learned from static needle simulations from a previous work by the authors. The emulation framework can predict the void fraction field at the exit of the orifice within a few seconds, thus achieving a speed-up factor of up to 38 x 10(6) over the traditional simulation-based approach of generating transient injection maps.

machine learning↗

ASCR@40: Highlights and Impacts of ASCR’s Programs

A report compiled by the ASCAC Subcommittee on the 40-year history of ASCR for the U.S. Department of Energy’s Office of Advanced Scientific Computing Research. The Office of Advanced Scientific Computing Research (ASCR) sits within the Office of Science in the Department of Energy (DOE). Per their web pages, “the mission of the ASCR program is to discover, develop, and deploy computational and networking capabilities to analyze, model, simulate, and predict complex phenomena important to the DOE.” This succinct statement encompasses a wide range of responsibilities for computing and networking facilities; for procuring, deploying, and operating high performance computing, networking, and storage resources; for basic research in mathematics and computer science; for developing and sustaining a large body of software; and for partnering with organizations across the Office of Science and beyond. While its mission statement may seem very contemporary, the roots of ASCR are quite deep—long predating the creation of DOE. Applied mathematics and advanced computing were both elements of the Theoretical Division of the Manhattan Project. In the early 1950s, the Manhattan Project scientist and mathematician John von Neumann, then a commissioner for the AEC (Atomic Energy Commission), advocated for the creation of a Mathematics program to support the continued development and applications of digital computing. Los Alamos National Laboratory (LANL) scientist John Pasta created such a program to fund researchers at universities and AEC laboratories. Under several organizational name changes, this program has persisted ever since, and would eventually grow to become ASCR.

97 MATHEMATICS AND COMPUTING↗

Dimensionally Reduced Model for Rapid and Accurate Prediction of Gas Saturation, Pressure, and Brine Production in a CO 2 Storage Application: Case Study Using the SACROC Field as Part of SMART Task 5

This technical report presents work conducted by the sub surface analysis team of the Strategic Systems Analysis & Engineering group at NETL for Task 5 of SMART Phase 1. This study involved the development of deep learning models for CO 2 geologic storage that are capable of accurate prediction of spatio-temporal outputs of CO 2 saturation, pressure, and brine production in three dimensional space over a storage operation's injection and post-injection timeframes. The model framework involves ensembling multi-layer encoder networks that provide dimesionality reduction of geologic inputs with fully connected long short-term memory (LSTM) neural networks that generate time-series prediction This approach offers a means to maximize training time efficiency, reduce computational memory burden, and minimize prediction turnaround.

54 ENVIRONMENTAL SCIENCES↗