Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

Exploiting Multi-Domain Features for Detection of Unclassified Electromagnetic Signals (Presentation)

Deep Learning based classification techniques have shown excellent performance in static environments, where the training and testing samples are drawn from the same distribution. However, real world scenarios often present samples that do not belong to the known set of classes chosen during training. This is quite common for electromagnetic signals, where it is impractical to assume that all possible waveforms are known a-priori, specially in scenarios like warfare. To address this problem, we propose a deep learning based adversarial model where the generator learns to generate waveform features that can deceive the discriminator model as true samples. We introduce domain knowledge of wireless signals by decomposing the signal into a lower dimensional unique feature set, which is used for classifying known versus unknown signals. We further introduce multiple domain representations of the signal to extract features and combine them together to accurately classify new waveforms as an unknown class. Our results show that combined features from multiple domains outperform any single domain representation, especially at low SNR regimes with fewer number of samples to classify.

99 - GENERAL AND MISCELLANEOUS↗

Knowledge-Informed Uncertainty-Aware Machine Learning for Time Series Forecasting of Dynamical Engineered Systems

The high complexity and multiscale nature of many engineered systems—such as those in nuclear power plants—make representing and forecasting their dynamic behavior challenging. Physics-based models can be overly complex and computationally intractable, whereas machine learning (ML) tools are often data-hungry and prone to unphysical solutions. This study proposes a knowledge-informed ML-aided hybrid residual modeling approach that offers accurate and efficient time series forecasting for the operation of dynamical engineered systems. Hybrid residual modeling entails a baseline solution from domain knowledge and known physics expressions about the system dynamics integrated with an ML model to capture undiscovered information from the mismatch (i.e., residuals) between true states from measurements and baseline-predicted outputs. This study further quantifies the ML model uncertainty to provide trustworthy solutions. Real-time operational data from thermal-hydraulic flow loops of the cryogenic moderator system in Oak Ridge National Laboratory’s Spallation Neutron Source facility were used to demonstrate the potential of knowledge-informed uncertainty-aware ML in real-world applications. The state variables of the cryogenic helium loop were modeled with (1) first principles–based system identification (sysID), (2) long short-term memory (LSTM) neural network, and (3) hybrid sysID (baseline) + LSTM (residual). The superior predictive capability of the sysID+LSTM model versus stand-alone sysID and LSTM is confirmed by average performance metrics and individual data points across different prediction horizons. By creating a robust representation of the underlying physical system, the widely applicable hybrid residual modeling approach will enable the future development of digital twins for performance prediction, prognostics, and operation control.

Zhao, Xingang↗

MultiLoad-GAN: A GAN-Based Synthetic Load Group Generation Method Considering Spatial-Temporal Correlations

This paper presents a deep-learning framework, Multi-load Generative Adversarial Network (MultiLoad-GAN), for generating a group of synthetic load profiles (SLPs) simultaneously. The main contribution of MultiLoad-GAN is the capture of spatial-temporal correlations among a group of loads that are served by the same distribution transformer. This enables the generation of a large amount of correlated SLPs required for microgrid and distribution system studies. Here, the novelty and uniqueness of the MultiLoad-GAN framework are three-fold. First, to the best of our knowledge, this is the first method for generating a group of load profiles bearing realistic spatial- temporal correlations simultaneously. Second, two complementary realisticness metrics for evaluating generated load profiles are developed: computing statistics based on domain knowledge and comparing high-level features via a deep-learning classifier. Third, to tackle data scarcity, a novel iterative data augmentation mechanism is developed to generate training samples for enhancing the training of both the classifier and the MultiLoad-GAN model. Simulation results show that MultiLoad- GAN can generate more realistic load profiles than existing approaches, especially in group level characteristics. With little finetuning, MultiLoad-GAN can be readily extended to generate a group of load or PV profiles for a feeder or a service area.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multi-level optimization with the koopman operator for data-driven, domain-aware, and dynamic system security

Cyber-Physical Systems (CPSs) like the power grid are critically important but also increasingly vulnerable; ensuring reliable system operation in the face of disruptions is becoming more and more challenging. Multi-Level Optimization (MLO) is a powerful way to model adversarial interactions, which naturally makes it applicable to studying CPS security. However, MLO typically does not address underlying system dynamics, and incorporating nonlinear dynamics is generally infeasible. In this paper, we show how to combine MLO with the Koopman Operator (KO) to remedy this. The KO maps nonlinear dynamics to a lifted space in which those dynamics are linear, thus making it ideal for use with MLO. Moreover, the structure of the KO also provides convenient ways to incorporate domain knowledge into the data-driven process of learning the KO representation of a given system. Here we then demonstrate the use of MLO-KO on a small example problem taken from the power grid domain, discuss the scalability and computational cost of MLO-KO, and identify future research directions for this work.

42 ENGINEERING↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗

Posterior Regularized Bayesian Neural Network incorporating soft and hard knowledge constraints

Neural Networks (NNs) have been widely used in supervised learning due to their ability to model complex nonlinear patterns, often presented in high-dimensional data such as images and text. However, traditional NNs often lack the ability for uncertainty quantification. Bayesian NNs (BNNS) could help measure the uncertainty by considering the distributions of the NN model parameters. Besides, domain knowledge is commonly available and could improve the performance of BNNs if it can be appropriately incorporated. In this work, we propose a novel Posterior-Regularized Bayesian Neural Network (PR-BNN) model by incorporating different types of knowledge constraints, such as the soft and hard constraints, as a posterior regularization term. Furthermore, we propose to combine the augmented Lagrangian method and the existing BNN solvers for efficient inference. Furthermore, the experiments in simulation and two case studies about aviation landing prediction and solar energy output prediction have shown the knowledge constraints and the performance improvement of the proposed model over traditional BNNs without the constraints.

14 SOLAR ENERGY↗

HydroEcoLSTM: A Python package with graphical user interface for hydro-ecological modeling with long short-term memory neural network

Machine learning (ML) is emerging as a promising tool for modeling hydro-ecological processes due to the increasing availability of large environmental data. However, the use of ML requires sufficient programming knowledge due to a lack of a graphical user interface (GUI). In this study, we introduced a GUI package, named HydroEcoLSTM, with the long short-term memory network (LSTM) as the core model, that allows non-ML experts to utilize their domain knowledge to construct complex ML models. We demonstrated the functionalities of HydroEcoLSTM with two practical examples, including (1) predictions of streamflow in both gauged and ungauged catchments and (2) predictions of multiple outputs (i.e., streamflow and isotope transport from two catchments). The simulation results obtained in both case experiments are satisfactory. In the first example, the average Nash–Sutcliffe Efficiency (NSE) for streamflow simulation during the testing period is 0.79 while the application of the trained model in two assumed ungauged catchments also achieves the average NSE of 0.68. In the second example, the average NSE for streamflow and instream isotope simulation during the testing period is 0.71. Ultimately, applications of HydroEcoLSTM with real-world examples demonstrate its potential use for practical applications and research without requiring extensive coding skills.

54 ENVIRONMENTAL SCIENCES↗

Understanding fission gas bubble distribution, lanthanide transportation, and thermal conductivity degradation in neutron-irradiated α-U using machine learning

U—10Zr based metallic nuclear fuel is the leading candidate for next-generation sodium-cooled fast reactors in the United States. US research reactors have used and tested this fuel type since the 1960s and accumulated considerable experience and knowledge about the fuel performance. Most of the knowledge, however, remains empirical. The lack of mechanistic understanding of fuel performance puts a large burden on proof through experimental verification for the qualification of U—10Zr fuel for commercial use. Further, this paper proposes an image data-driven machine learning approach, coupled with domain knowledge provided by advanced post irradiation examination, to provide unprecedented quantified insights into the morphology, size, density and the connectivity of fission gas bubbles and their effects on the fission product transportation and thermal conductivity. Specifically, we developed a method to automatically detect, extract statistics, and classify ~19,000 fission gas bubbles into different categories, and quantitatively link the data to lanthanide transportation through connected bubbles and degradation of thermal conductivity along the radial temperature gradient in a neutron irradiated U—10Zr annular fuel. Results indicate the approach can be modified to study other irradiation effects, such as secondary phase redistribution and gaseous fuel swelling in other irradiated nuclear fuels.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Machine learning in materials science: From explainable predictions to autonomous design

The advent of big data and algorithmic developments in the field of machine learning (and artificial intelligence, in general) have greatly impacted the entire spectrum of physical sciences, including materials science. Materials data, measured or computed, combined with various techniques of machine learning have been employed to address a myriad of challenging problems, such as, development of efficient and predictive surrogate models for a range of materials properties, screening and down-selection of novel candidate materials for targeted applications, new methodologies to improve and further expedite molecular and atomistic simulations, with likely many more important developments to come in the foreseeable future. While the applications thus far have provided a glimpse of the true potential data-enabled routes have to offer, it has also become clear that further progress in this direction hinges on our ability to understand, explain and rationalize findings of a machine learning model in light of the domain-knowledge. This focused review provides an overview of the main areas where machine learning has been widely and successfully used in materials science. Subsequently, a brief discussion of several techniques that have been helpful in extracting physically-meaningful insights, causal relationships and design-centric knowledge from materials data is provided. Finally, we identify some of the imminent opportunities and challenges that materials community faces in this exciting and rapidly growing field.

36 MATERIALS SCIENCE↗

Advanced characterization-informed machine learning framework and quantitative insight to irradiated annular U-10Zr metallic fuels

Abstract U-10Zr Metal fuel is a promising nuclear fuel candidate for next-generation sodium-cooled fast spectrum reactors. Since the Experimental Breeder Reactor-II in the late 1960s, researchers accumulated a considerable amount of experience and knowledge on fuel performance at the engineering scale. However, a mechanistic understanding of fuel microstructure evolution and property degradation during in-reactor irradiation is still missing due to a lack of appropriate tools for rapid fuel microstructure assessment and property prediction based on post irradiation examination. This paper proposed a machine learning enabled workflow, coupled with domain knowledge and large dataset collected from advanced post-irradiation examination microscopies, to provide rapid and quantified assessments of the microstructure in two reactor irradiated prototypical annular metal fuels. Specifically, this paper revealed the distribution of Zr-bearing secondary phases and constitutional redistribution across different radial locations. Additionally, the ratios of seven different microstructures at various locations along the temperature gradient were quantified. Moreover, the distributions of fission gas pores on two types of U-10Zr annular fuels were quantitatively compared.

36 MATERIALS SCIENCE↗

Data Fusion via Neural Network Entropy Minimization for Target Detection and Multi-Sensor Event Classification

Broadly applicable solutions to multimodal and multisensory fusion problems across domains remain a challenge because effective solutions often require substantive domain knowledge and engineering. The chief questions that arise for data fusion are in when to share information from different data sources, and how to accomplish the integration of information. The solutions explored in this work remain agnostic to input representation and terminal decision fusion approaches by sharing information through the learning objective as a compound objective function. The objective function this work uses assumes a one-to-one learning paradigm within a one-to-many domain which allows the assumption that consistency can be enforced across the one-to-many dimension. The domains and tasks we explore in this work include multi-sensor fusion for seismic event location and multimodal hyperspectral target discrimination. We find that our domain- informed consistency objectives are challenging to implement in stable and successful learning because of intersections between inherent data complexity and practical parameter optimization. While multimodal hyperspectral target discrimination was not enhanced across a range of different experiments by the fusion strategies put forward in this work, seismic event location benefited substantially, but only for label-limited scenarios.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine learning with knowledge constraints for process optimization of open-air perovskite solar cell manufacturing

Perovskite photovoltaics (PV) have achieved rapid development in the past decade in terms of power conversion efficiency of small-area lab-scale devices; however, successful commercialization still requires further development of low-cost, scalable, and high-throughput manufacturing techniques. One of the critical challenges of developing a new fabrication technique is the high-dimensional parameter space for optimization, but machine learning (ML) can readily be used to accelerate perovskite PV scaling. Herein, we present an ML-guided framework of sequential learning for manufacturing process optimization. We apply our methodology to the Rapid Spray Plasma Processing (RSPP) technique for perovskite thin films in ambient conditions. With a limited experimental budget of screening 100 process conditions, we demonstrated an efficiency improvement to 18.5% as the best-in-our-lab device fabricated by RSPP, and we also experimentally found 10 unique process conditions to produce the top-performing devices of more than 17% efficiency, which is 5 times higher rate of success than the control experiments with pseudo-random Latin hypercube sampling. Our model is enabled by three innovations: (a) flexible knowledge transfer between experimental processes by incorporating data from prior experimental data as a probabilistic constraint; (b) incorporation of both subjective human observations and ML insights when selecting next experiments; (c) adaptive strategy of locating the region of interest using Bayesian optimization first, and then conducting local exploration for high-efficiency devices. Furthermore, in virtual benchmarking, our framework achieves faster improvements with limited experimental budgets than traditional design-of-experiments methods (e.g., one-variable-at-a-time sampling). This framework shows the capability of incorporating researchers’ domain knowledge into the ML-guided optimization loop; therefore, it has the potential to facilitate the wider adoption of ML in scaling to perovskite PV manufacturing.

14 SOLAR ENERGY↗

A Tale from the Trenches: Applying Metamorphic and Differential Testing to Bioinformatics Software

Metamorphic and differential testing have been proposed as best practices for testing software that is difficult to test, such as for programs in scientific domains. An assumption is that these approaches can be easily customized and applied to almost any domain. However, scientific software is often data-driven, and metamorphic relations may require significant domain knowledge to develop. In addition, tools are often written for ad-hoc experimentation by the scientists and often embed many assumptions about the importance and representation of different natural phenomena. In this paper, we present our experience applying both metamorphic and differential testing to a set of four computational biology tools that predict the growth of an organism. While our original goal was to evaluate these techniques to improve our system-level testing, we encountered multiple roadblocks along the way. Although we did find faults (some confirmed by developers), we also uncovered a set of challenges, including the considerable manual effort required for (a) defining domain-specific tests, (b) validating correctness, and (c) distinguishing between issues stemming from poor data and those arising from incorrect software.

Marsh, Alexis L [Iowa State University/Ames Labora↗

What Machine Learning Can and Cannot Do for Inertial Confinement Fusion

Machine learning methodologies have played remarkable roles in solving complex systems with large data, well-defined input–output pairs, and clearly definable goals and metrics. The methodologies are effective in image analysis, classification, and systems without long chains of logic. Recently, machine-learning methodologies have been widely applied to inertial confinement fusion (ICF) capsules and the design optimization of OMEGA (Omega Laser Facility) capsule implosion and NIF (National Ignition Facility) ignition capsules, leading to significant progress. As machine learning is being increasingly applied, concerns arise regarding its capabilities and limitations in the context of ICF. ICF is a complicated physical system that relies on physics knowledge and human judgment to guide machine learning. Additionally, the experimental database for ICF ignition is not large enough to provide credible training data. Most researchers in the field of ICF use simulations, or a mix of simulations and experimental results, instead of real data to train machine learning models and related tools. They then use the trained learning model to predict future events. This methodology can be successful, subject to a careful choice of data and simulations. However, because of the extreme sensitivity of the neutron yield to the input implosion parameters, physics-guided machine learning for ICF is extremely important and necessary, especially when the database is small, the uncertain-domain knowledge is large, and the physical capabilities of the learning models are still being developed. In this work, we identify problems in ICF that are suitable for machine learning and circumstances where machine learning is less likely to be successful. This study investigates the applications of machine learning and highlights fundamental research challenges and directions associated with machine learning in ICF.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SNAD transient miner: Finding missed transient events in ZTF DR4 using k-D trees

Here we report the automatic detection of 11 transients (7 possible supernovae and 4 active galactic nuclei candidates) within the Zwicky Transient Facility fourth data release (ZTF DR4), all of them observed in 2018 and absent from public catalogs. Among these, three were not part of the ZTF alert stream. Our transient mining strategy employs 41 physically motivated features extracted from both real light curves and four simulated light curve models (SN Ia, SN II, TDE, SLSN-I). These features are input to a k-D tree algorithm, from which we calculate the 15 nearest neighbors. After pre-processing and selection cuts, our dataset contained approximately a million objects among which we visually inspected the 105 closest neighbors from seven of our brightest, most well-sampled simulations, comprising 89 unique ZTF DR4 sources. Our result illustrates the potential of coherently incorporating domain knowledge and automatic learning algorithms, which is one of the guiding principles directing the SNAD team. It also demonstrates that the ZTF DR is a suitable testing ground for data mining algorithms aiming to prepare for the next generation of astronomical data.

79 ASTRONOMY AND ASTROPHYSICS↗