Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

Coupled Lattice Boltzmann Modeling Framework for Pore-Scale Fluid Flow and Reactive Transport

In this paper, we propose a modeling framework for pore-scale fluid flow and reactive transport based on a coupled lattice Boltzmann model (LBM). We develop a modeling interface to integrate the LBM modeling code parallel lattice Boltzmann solver and the PHREEQC reaction solver using multiple flow and reaction cell mapping schemes. The major advantage of the proposed workflow is the high modeling flexibility obtained by coupling the geochemical model with the LBM fluid flow model. Consequently, the model is capable of executing one or more complex reactions within desired cells while preserving the high data communication efficiency between the two codes. Meanwhile, the developed mapping mechanism enables the flow, diffusion, and reactions in complex pore-scale geometries. We validate the coupled code in a series of benchmark numerical experiments, including 2D single-phase Poiseuille flow and diffusion, 2D reactive transport with calcite dissolution, as well as surface complexation reactions. The simulation results show good agreement with analytical solutions, experimental data, and multiple other simulation codes. In addition, we design an AI-based optimization workflow and implement it on the surface complexation model to enable increased capacity of the coupled modeling framework. Compared to the manual tuning results proposed in the literature, our workflow demonstrates fast and reliable model optimization results without incorporating pre-existing domain knowledge.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Distilling nanoscale heterogeneity of amorphous silicon using tip-enhanced Raman spectroscopy (TERS) via multiresolution manifold learning

Abstract Accurately identifying the local structural heterogeneity of complex, disordered amorphous materials such as amorphous silicon is crucial for accelerating technology development. However, short-range atomic ordering quantification and nanoscale spatial resolution over a large area on a-Si have remained major challenges and practically unexplored. We resolve phonon vibrational modes of a-Si at a lateral resolution of <60 nm by tip-enhanced Raman spectroscopy. To project the high dimensional TERS imaging to a two-dimensional manifold space and categorize amorphous silicon structure, we developed a multiresolution manifold learning algorithm. It allows for quantifying average Si-Si distortion angle and the strain free energy at nanoscale without a human-specified physical threshold. The multiresolution feature of the multiresolution manifold learning allows for distilling local defects of ultra-low abundance (< 0.3%), presenting a new Raman mode at finer resolution grids. This work promises a general paradigm of resolving nanoscale structural heterogeneity and updating domain knowledge for highly disordered materials.

36 MATERIALS SCIENCE↗

MolDiscovery: learning mass spectrometry fragmentation of small molecules

Identification of small molecules is a critical task in various areas of life science. Recent advances in mass spectrometry have enabled the collection of tandem mass spectra of small molecules from hundreds of thousands of environments. To identify which molecules are present in a sample, one can search mass spectra collected from the sample against millions of molecular structures in small molecule databases. The existing approaches are based on chemistry domain knowledge, and they fail to explain many of the peaks in mass spectra of small molecules. Here, we present molDiscovery, a mass spectral database search method that improves both efficiency and accuracy of small molecule identification by learning a probabilistic model to match small molecules with their mass spectra. A search of over 8 million spectra from the Global Natural Product Social molecular networking infrastructure shows that molDiscovery correctly identify six times more unique small molecules than previous methods.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Combining data and theory for derivable scientific discovery with AI-Descartes

Abstract Scientists aim to discover meaningful formulae that accurately describe experimental data. Mathematical models of natural phenomena can be manually created from domain knowledge and fitted to data, or, in contrast, created automatically from large datasets with machine-learning algorithms. The problem of incorporating prior knowledge expressed as constraints on the functional form of a learned model has been studied before, while finding models that are consistent with prior knowledge expressed via general logical axioms is an open problem. We develop a method to enable principled derivations of models of natural phenomena from axiomatic knowledge and experimental data by combining logical reasoning with symbolic regression. We demonstrate these concepts for Kepler’s third law of planetary motion, Einstein’s relativistic time-dilation law, and Langmuir’s theory of adsorption. We show we can discover governing laws from few data points when logical reasoning is used to distinguish between candidate formulae having similar error on the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Autonomous and dynamic precursor selection for solid-state materials synthesis

Abstract Solid-state synthesis plays an important role in the development of new materials and technologies. While in situ characterization and ab-initio computations have advanced our understanding of materials synthesis, experiments targeting new compounds often still require many different precursors and conditions to be tested. Here we introduce an algorithm (ARROWS 3 ) designed to automate the selection of optimal precursors for solid-state materials synthesis. This algorithm actively learns from experimental outcomes to determine which precursors lead to unfavorable reactions that form highly stable intermediates, preventing the target material’s formation. Based on this information, ARROWS 3 proposes new experiments using precursors it predicts to avoid such intermediates, thereby retaining a larger thermodynamic driving force to form the target. We validate this approach on three experimental datasets, containing results from over 200 synthesis procedures. In comparison to black-box optimization, ARROWS 3 identifies effective precursor sets for each target while requiring substantially fewer experimental iterations. These findings highlight the importance of domain knowledge in optimization algorithms for materials synthesis, which are critical for the development of fully autonomous research platforms.

36 MATERIALS SCIENCE↗

AtomSets as a hierarchical transfer learning framework for small and large materials datasets

Abstract Predicting properties from a material’s composition or structure is of great interest for materials design. Deep learning has recently garnered considerable interest in materials predictive tasks with low model errors when dealing with large materials data. However, deep learning models suffer in the small data regime that is common in materials science. Here we develop the AtomSets framework, which utilizes universal compositional and structural descriptors extracted from pre-trained graph network deep learning models with standard multi-layer perceptrons to achieve consistently high model accuracy for both small compositional data (<400) and large structural data (>130,000). The AtomSets models show lower errors than the graph network models at small data limits and other non-deep-learning models at large data limits. They also transfer better in a simulated materials discovery process where the targeted materials have property values out of the training data limits. The models require minimal domain knowledge inputs and are free from feature engineering. The presented AtomSets model framework can potentially accelerate machine learning-assisted materials design and discovery with less data restriction.

Chen, Chi (ORCID:0000000180087043)↗

SANE: strategic autonomous non-smooth exploration for multiple optima discovery in multi-modal and non-differentiable black-box functions

Both computational and experimental material discovery bring forth the challenge of exploring multidimensional and multimodal parameter spaces, such as phase diagrams of Hamiltonians with multiple interactions, composition spaces of combinatorial libraries, material structure image spaces, and molecular embedding spaces. Often these systems are black-boxes and time-consuming to evaluate, which resulted in strong interest towards active learning methods such as Bayesian optimization (BO). However, these systems are often noisy which make the black box function severely multi-modal and non-differentiable, where a vanilla BO can get overly focused near a single or faux optimum, deviating from the broader goal of scientific discovery. To address these limitations, here we developed Strategic Autonomous Non-Smooth Exploration (SANE) to facilitate an intelligent Bayesian optimized navigation with a proposed cost-driven probabilistic acquisition function to find multiple global and local optimal regions, avoiding the tendency to becoming trapped in a single optimum. To distinguish between a true and false optimal region due to noisy experimental measurements, a human (domain) knowledge driven dynamic surrogate gate is integrated with SANE. We implemented the gate-SANE into pre-acquired piezoresponse spectroscopy data of a ferroelectric combinatorial library with high noise levels in specific regions, and piezoresponse force microscopy (PFM) hyperspectral data. SANE demonstrated better performance than classical BO to facilitate the exploration of multiple optimal regions and thereby prioritized learning with higher coverage of scientific values in autonomous experiments. Our work showcases the potential application of this method to real-world experiments, where such combined strategic and human intervening approaches can be critical to unlocking new discoveries in autonomous research.

Biswas, Arpan [University of Tennessee, Knoxville,↗

Synergy of semiempirical models and machine learning in computational chemistry

Catalyzed by enormous success in the industrial sector, many research programs have been exploring data-driven, machine learning approaches. Performance can be poor when the model is extrapolated to new regions of chemical space, e.g., new bonding types, new many-body interactions. Another important limitation is the spatial locality assumption in model architecture, and this limitation cannot be overcome with larger or more diverse datasets. The outlined challenges are primarily associated with the lack of electronic structure information in surrogate models such as interatomic potentials. Given the fast development of machine learning and computational chemistry methods, we expect some limitations of surrogate models to be addressed in the near future; nevertheless spatial locality assumption will likely remain a limiting factor for their transferability. Here, we suggest focusing on an equally important effort—design of physics-informed models that leverage the domain knowledge and employ machine learning only as a corrective tool. In the context of material science, we will focus on semi-empirical quantum mechanics, using machine learning to predict corrections to the reduced-order Hamiltonian model parameters. The resulting models are broadly applicable, retain the speed of semiempirical chemistry, and frequently achieve accuracy on par with much more expensive ab initio calculations. These early results indicate that future work, in which machine learning and quantum chemistry methods are developed jointly, may provide the best of all worlds for chemistry applications that demand both high accuracy and high numerical efficiency.

36 MATERIALS SCIENCE↗

Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled datasets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed reinforcement learning model, PIRL-AD: Physics-Informed Reinforcement Learning for Anomaly Detection, that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a non-physics informed RL benchmark and a support vector machine (SVM) on data collected from a forced flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous datasets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% vs. 91.06% vs. 83.01%) and F1-score (respectively, 89.23% vs. 86.98% vs. 77.01%). Furthermore, the findings suggest the potential of physics-integrated reinforcement learning models for enhanced anomaly detection performance in NPPs.

Reinforcement learning↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Augmented signal processing in Liquid Argon Time Projection Chambers with a deep neural network

The Liquid Argon Time Projection Chamber (LArTPC) is an advanced neutrino detector technology widely used in recent and upcoming accelerator neutrino experiments. It features a low energy threshold and high spatial resolution that allow for comprehensive reconstruction of event topologies. In current-generation LArTPCs, the recorded data consist of digitized waveforms on wires produced by induced signal on wires of drifting ionization electrons, which can also be viewed as two-dimensional (2D) (time versus wire) projection images of charged-particle trajectories. For such an imaging detector, one critical step is the signal processing that reconstructs the original charge projections from the recorded 2D images. For the first time, we introduce a deep neural network in LArTPC signal processing to improve the signal region of interest detection. By combining domain knowledge (e.g., matching information from multiple wire planes) and deep learning, this method shows significant improvements over traditional methods. This work details the method, software tools, and performance evaluated with realistic detector simulations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Square Kilometre Array Science Data Challenge 1: analysis and results

ABSTRACT As the largest radio telescope in the world, the Square Kilometre Array (SKA) will lead the next generation of radio astronomy. The feats of engineering required to construct the telescope array will be matched only by the techniques developed to exploit the rich scientific value of the data. To drive forward the development of efficient and accurate analysis methods, we are designing a series of data challenges that will provide the scientific community with high-quality data sets for testing and evaluating new techniques. In this paper, we present a description and results from the first such Science Data Challenge 1 (SDC1). Based on SKA MID continuum simulated observations and covering three frequencies (560, 1400, and 9200 MHz) at three depths (8, 100, and 1000 h), SDC1 asked participants to apply source detection, characterization, and classification methods to simulated data. The challenge opened in 2018 November, with nine teams submitting results by the deadline of 2019 April. In this work, we analyse the results for eight of those teams, showcasing the variety of approaches that can be successfully used to find, characterize, and classify sources in a deep, crowded field. The results also demonstrate the importance of building domain knowledge and expertise on this kind of analysis to obtain the best performance. As high-resolution observations begin revealing the true complexity of the sky, one of the outstanding challenges emerging from this analysis is the ability to deal with highly resolved and complex sources as effectively as the unresolved source population.

Bonaldi, A.↗

Anomaly detection in the Zwicky Transient Facility DR3

We present results from applying the SNAD anomaly detection pipeline to the third public data release of the Zwicky Transient Facility (ZTF DR3). The pipeline is composed of three stages: feature extraction, search of outliers with machine learning algorithms, and anomaly identification with followup by human experts. Our analysis concentrates in three ZTF fields, comprising more than 2.25 million objects. A set of four automatic learning algorithms was used to identify 277 outliers, which were subsequently scrutinized by an expert. From these, 188 (68 per cent) were found to be bogus light curves – including effects from the image subtraction pipeline as well as overlapping between a star and a known asteroid, 66 (24 per cent) were previously reported sources whereas 23 (8 per cent) correspond to non-catalogued objects, with the two latter cases of potential scientific interest (e.g. one spectroscopically confirmed RS Canum Venaticorum star, four supernovae candidates, one red dwarf flare). Moreover, using results from the expert analysis, we were able to identify a simple bi-dimensional relation that can be used to aid filtering potentially bogus light curves in future studies. We provide a complete list of objects with potential scientific application so they can be further scrutinised by the community. These results confirm the importance of combining automatic machine learning algorithms with domain knowledge in the construction of recommendation systems for astronomy. Our code is publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning of Kondo physics using variational autoencoders and symbolic regression

In this work, we employ variational autoencoders to extract physical insight from a dataset of one-particle Anderson impurity model spectral functions. Autoencoders are trained to find a low-dimensional, latent space representation that faithfully characterizes each element of the training set, as measured by a reconstruction error. Variational autoencoders, a probabilistic generalization of standard autoencoders, further condition the learned latent space to promote highly interpretable features. In our study, we find that the learned latent variables strongly correlate with well known, but nontrivial, parameters that characterize emergent behaviors in the Anderson impurity model. In particular, one latent variable correlates with particle-hole asymmetry, while another is in near one-to-one correspondence with the Kondo temperature, a dynamically generated low-energy scale in the impurity model. Using symbolic regression, we model this variable as a function of the known bare physical input parameters and "rediscover" the nonperturbative formula for the Kondo temperature. The machine learning pipeline we develop suggests a general purpose approach, which opens opportunities to discover new domain knowledge in other physical systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

ROI-Finder : machine learning to guide region-of-interest scanning for X-ray fluorescence microscopy

The microscopy research at the Bionanoprobe (currently at beamline 9-ID and later 2-ID after APS-U) of Argonne National Laboratory focuses on applying synchrotron X-ray fluorescence (XRF) techniques to obtain trace elemental mappings of cryogenic biological samples to gain insights about their role in critical biological activities. The elemental mappings and the morphological aspects of the biological samples, in this instance, the bacterium Escherichia coli ( E. Coli ), also serve as label-free biological fingerprints to identify E. coli cells that have been treated differently. The key limitations of achieving good identification performance are the extraction of cells from raw XRF measurements via binary conversion, definition of features, noise floor and proportion of cells treated differently in the measurement. Automating cell extraction from raw XRF measurements across different types of chemical treatment and the implementation of machine-learning models to distinguish cells from the background and their differing treatments are described. Principal components are calculated from domain knowledge specific features and clustered to distinguish healthy and poisoned cells from the background without manual annotation. The cells are ranked via fuzzy clustering to recommend regions of interest for automated experimentation. The effects of dwell time and the amount of data required on the usability of the software are also discussed.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

DOC-DICAM: Domain Aware One Class Defect Identification in Composite Aerostructure Material

Fiber-reinforced composites are a common material used in the design of aircraft structures due to their good tensile strength and resistance to compression. During the manufacturing process, these structures are thoroughly inspected for flaws and defects to ensure structural integrity during commercial use. Non-destructive testing (NDT) is a collection of inspection methods that allow inspectors to evaluate material without altering it. Due to the high safety standards in aerospace manufacturing, the NDT process is done manually and can be a significant bottleneck in the development workflow. In this paper, we develop an AI-based assistance tool to drastically reduce inspection time. Typical AI workflows require large amounts of annotated data, but defects rarely occur resulting in strong class imbalance. To overcome this, we formulate the problem of defect identification as an anomaly detection task in which our primary focus is learning non-defect characteristics. To do this, we develop a multi-task self-supervised learning framework that embeds problem specific domain knowledge into the deep learning model. We verify our method using fuselage data generated in a production environment. As a result, we show that our method can effectively identify defects and requires minimal training and inference time.

anomaly detection↗