Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Knowledge graph”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

py-boomer v0.1.0

Py-BOOMER (Python Bayesian OWL Ontology MErgER in Python) is a probabilistic reasoning system for knowledge representation and ontological reasoning with uncertainty. Itnables reasoning over probabilistic facts and taxonomic relationships, finding the most likely consistent interpretation of potentially conflicting assertions. It uses a combination of graph-based reasoning and Bayesian probabilistic inference. Key features: Represent probabilistic ontological statements Reason over class subsumption hierarchies Evaluate class equivalence relationships Detect and resolve logical inconsistencies Calculate posterior probabilities for each assertion

Mungall, Chris [Lawrence Berkeley National Laborat↗

Projecting the Thermal Response in a HTGR-Type System during Conduction Cooldown Using Graph-Laplacian Based Machine Learning

Accurate prediction of an off-normal event in a nuclear reactor is dependent upon the availability of sensory data, reactor core physical condition, and understanding of the underlying phenomenon. This work presents a method to project the data from some discrete sensory locations to the overall reactor domain during conduction cooldown scenarios similar to High Temperature Gas-cooled Reactors (HTGRs). The existing models for conductive cooldown in a heterogeneous multi-body system, such as an assembly of prismatic blocks or pebble beds relies on knowledge of the thermal contact conductance, requiring significant knowledge of local thermal contacts and heat transport possibilities across those contacts. With a priori knowledge of bulk geometry features and some discrete sensors, a machine learning approach was devised. The presented work uses an experimental facility to mimic conduction cooldown with an assembly of 68 cylindrical rods initially heated to 1200 K. High-fidelity temperature data were collected using an infrared (IR) camera to provide training data to the model and validate the predicted temperature data. The machine learning approach used here first converts the macroscopic bulk geometry information into Graph-Laplacian, and then uses the eigenvectors of the Graph-Laplacian to develop Kernel functions. Support vector regression (SVR) was implemented on the obtained Kernels and used to predict the thermal response in a packed rod assembly during a conduction cooldown experiment. The usage of SVR modeling differs from most models today because of its representation of thermal coupling between rods in the core. When trained with thermographic data, the average normalized error is less than 2% over 400 s, during which temperatures of the assembly have dropped by more than 500 K. The rod temperature prediction performance was significantly better for rods in the interior of the assembly compared to those near the exterior, likely due to the model simplification of the surroundings.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A unified understanding of minimum lattice thermal conductivity

Here, we propose a first-principles model of minimum lattice thermal conductivity ($κ^{min}_L$) based on a unified theoretical treatment of thermal transport in crystals and glasses. We apply this model to thousands of inorganic compounds and find a universal behavior of $κ^{min}_L$ in crystals in the high-temperature limit: The isotropically averaged $κ^{min}_L$ is independent of structural complexity and bounded within a range from ~0.1 to ~2.6 W/(m K), in striking contrast to the conventional phonon gas model which predicts no lower bound. We unveil the underlying physics by showing that for a given parent compound, $κ^{min}_L$ is bounded from below by a value that is approximately insensitive to disorder, but the relative importance of different heat transport channels (phonon gas versus diffuson) depends strongly on the degree of disorder. Moreover, we propose that the diffuson-dominated $κ^{min}_L$ in complex and disordered compounds might be effectively approximated by the phonon gas model for an ordered compound by averaging out disorder and applying phonon unfolding. With these insights, we further bridge the knowledge gap between our model and the well-known Cahill–Watson–Pohl (CWP) model, rationalizing the successes and limitations of the CWP model in the absence of heat transfer mediated by diffusons. Finally, we construct graph network and random forest machine learning models to extend our predictions to all compounds within the Inorganic Crystal Structure Database (ICSD), which were validated against thermoelectric materials possessing experimentally measured ultralow κ L . Our work offers a unified understanding of $κ^{min}_L$, which can guide the rational engineering of materials to achieve .

42 ENGINEERING↗

Meta-Learning Enhanced Physics-Informed Graph Attention Convolutional Network for Distribution Power System State Estimation

Promptly perceiving distribution system states is challenged by frequent topology changes and uncertain power injections. To address these issues, a Meta-learning enhanced physics-informed graph attention convolutional network (Meta-PIGACN) model is proposed to handle topological variability in distribution system state estimation (DSSE). Specifically, physics information is integrated into the graph convolutional network, enabling a physics-informed edge-weighting process that incorporates physical information to control the aggregation of neighboring nodes. Besides, the graph attention mechanism automatically adjusts the importance of different neighboring nodes, allowing the capture and preservation of inherent system features across varying topologies, thereby improving state estimation accuracy. Furthermore, meta-learning is proposed to acquire empirical knowledge across multiple topologies so that the model can rapidly adapt to new configurations through iterative gradient descent updates even in large-scale systems. In conclusion, the simulation results based on the 33/118/1746-node distribution systems show the high accuracy and efficiency of the proposed model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning Distribution Grid Topologies: A Tutorial

Unveiling feeder topologies from data is of paramount importance to advance situational awareness and proper utilization of smart resources in power distribution grids. This tutorial summarizes, contrasts, and establishes useful links between recent works on topology identification and detection schemes that have been proposed for power distribution grids. The primary focus is to highlight methods that overcome the limited availability of measurement devices in distribution grids, while enhancing topology estimates using conservation laws of power-flow physics and structural properties of feeders. Grid data from phasor measurement units or smart meters can be collected either passively in the traditional way, or actively, upon actuating grid resources and measuring the feeder's voltage response. Analytical claims on feeder identifiability and detectability are reviewed under disparate meter placement scenarios. Such topology learning claims can be attained exactly or approximately so via algorithmic solutions with various levels of computational complexity, ranging from least-squares fits to convex optimization problems, and from polynomial-time searches over graphs to mixed-integer programs. Although the emphasis is on radial single-phase feeders, extensions to meshed and/or multiphase circuits are sometimes possible and discussed. Here this tutorial aspires to provide researchers and engineers with knowledge of the current state-of-the-art in tractable distribution grid learning and insights into future directions of work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

graphenv: a Python library for reinforcement learning on graph search spaces

Many important and challenging problems in combinatorial optimization (CO) can be expressed as graph search problems, in which graph vertices represent full or partial solutions and edges represent decisions that connect them. Graph structure not only introduces strong relational inductive biases for learning (Battaglia et al., 2018) - in this context, by providing a way to explicitly model the value of transitioning (along edges) between one search state (vertex) and the next - but lends itself to problems both with and without clearly defined algebraic structure. For example, classic CO problems on graphs such as the Traveling Salesman Problem (TSP) can be expressed as either pure graph search or integer programs. Other problems, however, such as molecular optimization, do no have concise algebraic formulations and yet are readily implemented as a graph search (V. et al., 2022; Zhou et al., 2019). Such "model-free" problems constitute a large fraction of modern reinforcement learning (RL) research owing to the fact that it is often much easier to write a forward simulation that expresses all of the state transitions and rewards, than to write down the precise mathematical expression of the full optimization problem. In the case of molecular optimization, for example, one can use domain knowledge alongside existing software libraries to model the effect of adding a single bond or atom to an existing but incomplete molecule, and let the RL algorithm build a model of how good a given decision is by "experiencing" the simulated environment many times through. In contrast, a model-based mathematical formulation that fully expresses all the chemical and physical constraints is intractable. In recent years, RL has emerged as an effective paradigm for optimizing searches over graphs and led to state-of-the-art heuristics for games like Go and chess, as well as for classical CO problems such as the TSP. This combination of graph search and RL, while powerful, requires non-trivial software to execute, especially when combining advanced state representations such as Graph Neural Networks (GNN) with scalable RL algorithms.

97 MATHEMATICS AND COMPUTING↗

Track Seeding and Labelling with Embedded-space Graph Neural Networks

To address the unprecedented scale of HL-LHC data, the Exa.TrkX project is investigating a variety of machine learning approaches to particle track reconstruction. The most promising of these solutions, graph neural networks (GNN), process the event as a graph that connects track measurements (detector hits corresponding to nodes) with candidate line segments between the hits (corresponding to edges). Detector information can be associated with nodes and edges, enabling a GNN to propagate the embedded parameters around the graph and predict node-, edge- and graph-level observables. Previously, message-passing GNNs have shown success in predicting doublet likelihood, and we here report updates on the state-of-the-art architectures for this task. In addition, the Exa.TrkX project has investigated innovations in both graph construction, and embedded representations, in an effort to achieve fully learned end-to-end track finding. Hence, we present a suite of extensions to the original model, with encouraging results for hitgraph classification. In addition, we explore increased performance by constructing graphs from learned representations which contain non-linear metric structure, allowing for efficient clustering and neighborhood queries of data points. We demonstrate how this framework fits in with both traditional clustering pipelines, and GNN approaches. The embedded graphs feed into high-accuracy doublet and triplet classifiers, or can be used as an end-to-end track classifier by clustering in an embedded space. A set of post-processing methods improve performance with knowledge of the detector physics. Finally, we present numerical results on the TrackML particle tracking challenge dataset, where our framework shows favorable results in both seeding and track finding.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Multi-objective goal-directed optimization of de novo stable organic radicals for aqueous redox flow batteries

Abstract Advances in the field of goal-directed molecular optimization offer the promise of finding feasible candidates for even the most challenging molecular design applications. One example of a fundamental design challenge is the search for novel stable radical scaffolds for an aqueous redox flow battery that simultaneously satisfy redox requirements at the anode and cathode, as relatively few stable organic radicals are known to exist. To meet this challenge, we develop a new open-source molecular optimization framework based on AlphaZero coupled with a fast, machine-learning-derived surrogate objective trained with nearly 100,000 quantum chemistry simulations. The objective function comprises two graph neural networks: one that predicts adiabatic oxidation and reduction potentials and a second that predicts electron density and local three-dimensional environment, previously shown to be correlated with radical persistence and stability. With no hard-coded knowledge of organic chemistry, the reinforcement learning agent finds molecule candidates that satisfy a precise combination of redox, stability and synthesizability requirements defined at the quantum chemistry level, many of which have reasonable predicted retrosynthetic pathways. The optimized molecules show that alternative stable radical scaffolds may offer a unique profile of stability and redox potentials to enable low-cost symmetric aqueous redox flow batteries.

25 ENERGY STORAGE↗

CheKiPEUQ

Parameter estimation for complex physical problems often suffers from finding 'solutions' that are not physically realistic. The CheKiPEUQ software provides tools for finding physically realistic parameter estimates.and CheKiPEUQ provide tools for making graphs of the parameter positions within parameter space as well as plots of the final simulation results. The primary purpose of the CheKiPEUQ software is to enable more physically realistic parameter estimation from comparing simulations to experiments. Specifically, when prior knowledge is available about the region of parameter space which is physically realistic and when the level of uncertainty from experiment can be estimated. The software is intentionally general and can be used for almost any type of simulation. The software is made in a user friendly manner so that users can simply enter the required data and then run the program, following guidelines provided by the authors along with some trial and error. Examples are provided. Users do not need to understand the methodology that will be namedi in the following sentences. While CheKiPEUQ can be used for conventional parameter estimation and other uses, the primary uses for CheKiPEUQ are: 1) Bayesian Parameter Estimation, 2) Bayesian Model Discrimination, 3) Bayesian Design of Experiments. For more information see the project website, documentation, examples, and related publications.

Savara, Aditya [Oak Ridge National Lab. (ORNL), Oa↗

ChemGraph as an agentic framework for computational chemistry workflows

Atomistic simulations are essential in chemistry and materials science but remain challenging to run due to the expert knowledge required for the setup, execution, and validation stages of these calculations. We present ChemGraph, an agentic framework powered by artificial intelligence and state-of-the-art simulation tools to streamline and automate computational chemistry and materials science workflows. ChemGraph leverages graph neural network-based foundation models for accurate yet computationally efficient calculations and large language models (LLMs) for natural language understanding, task planning, and scientific reasoning to provide an intuitive and interactive interface. We evaluate ChemGraph across 13 benchmark tasks and demonstrate that smaller LLMs (GPT-4o-mini, Claude-3.5-haiku, Qwen-2.5-14B) perform well on simple workflows, while more complex tasks benefit from using larger models. Importantly, we show that decomposing complex tasks into smaller subtasks through a multi-agent framework enables GPT-4o to reach perfect accuracy and smaller LLMs to match or exceed single-agent GPT-4o's performance in these benchmarks.

Computational chemistry↗

A Mass‐Conserving‐Perceptron for Machine‐Learning‐Based Modeling of Geoscientific Systems

Although decades of effort have been devoted to building Physical-Conceptual (PC) models for predicting the time-series evolution of geoscientific systems, recent work shows that Machine Learning (ML) based Gated Recurrent Neural Network technology can be used to develop models that are much more accurate. However, the difficulty of extracting physical understanding from ML-based models complicates their utility for enhancing scientific knowledge regarding system structure and function. Here, we propose a physically interpretable Mass-Conserving-Perceptron (MCP) as a way to bridge the gap between PC-based and ML-based modeling approaches. The MCP exploits the inherent isomorphism between the directed graph structures underlying both PC models and GRNNs to explicitly represent the mass-conserving nature of physical processes while enabling the functional nature of such processes to be directly learned (in an interpretable manner) from available data using off-the-shelf ML technology. As a proof of concept, we investigate the functional expressivity (capacity) of the MCP, explore its ability to parsimoniously represent the rainfall-runoff (RR) dynamics of the Leaf River Basin, and demonstrate its utility for scientific hypothesis testing. To conclude, we discuss extensions of the concept to enable ML-based physical-conceptual representation of the coupled nature of mass-energy-information flows through geoscientific systems.

58 GEOSCIENCES↗

Computational Estimation by Scientific Data Mining with Classical Methods to Automate Learning Strategies of Scientists

Experimental results are often plotted as 2-dimensional graphical plots (aka graphs) in scientific domains depicting dependent versus independent variables to aid visual analysis of processes. Repeatedly performing laboratory experiments consumes significant time and resources, motivating the need for computational estimation. The goals are to estimate the graph obtained in an experiment given its input conditions, and to estimate the conditions that would lead to a desired graph. Existing estimation approaches often do not meet accuracy and efficiency needs of targeted applications. We develop a computational estimation approach called AutoDomainMine that integrates clustering and classification over complex scientific data in a framework so as to automate classical learning methods of scientists. Knowledge discovered thereby from a database of existing experiments serves as the basis for estimation. Challenges include preserving domain semantics in clustering, finding matching strategies in classification, striking a good balance between elaboration and conciseness while displaying estimation results based on needs of targeted users, and deriving objective measures to capture subjective user interests. These and other challenges are addressed in this work. The AutoDomainMine approach is used to build a computational estimation system, rigorously evaluated with real data in Materials Science. Our evaluation confirms that AutoDomainMine provides desired accuracy and efficiency in computational estimation. It is extendable to other science and engineering domains as proved by adaptation of its sub-processes within fields such as Bioinformatics and Nanotechnology.

Computer Science↗

AtomSets as a hierarchical transfer learning framework for small and large materials datasets

Abstract Predicting properties from a material’s composition or structure is of great interest for materials design. Deep learning has recently garnered considerable interest in materials predictive tasks with low model errors when dealing with large materials data. However, deep learning models suffer in the small data regime that is common in materials science. Here we develop the AtomSets framework, which utilizes universal compositional and structural descriptors extracted from pre-trained graph network deep learning models with standard multi-layer perceptrons to achieve consistently high model accuracy for both small compositional data (<400) and large structural data (>130,000). The AtomSets models show lower errors than the graph network models at small data limits and other non-deep-learning models at large data limits. They also transfer better in a simulated materials discovery process where the targeted materials have property values out of the training data limits. The models require minimal domain knowledge inputs and are free from feature engineering. The presented AtomSets model framework can potentially accelerate machine learning-assisted materials design and discovery with less data restriction.

Chen, Chi (ORCID:0000000180087043)↗

Optimizing FPGA-based Accelerator Design for Large-Scale Molecular Similarity Search (Special Session Paper)

Molecular similarity search has been widely used in drug discovery to rapidly identify structurally similar compounds from large molecular databases. With the increasing size of chemical libraries, there is growing interest in the efficient ac- celeration of large-scale similarity search. Existing works mainly focus on CPU and GPU to accelerate the computation of Tatimoto coefficient in measuring the pairwise similarity between different molecular fingerprints. In this paper, we propose and optimize an FPGA-based accelerator design on exhaustive and approximate search algorithms. On exhaustive search using BitBound & fold- ing, we analyze the similarity cutoff and folding level relationship with search speedup and accuracy, and propose a scalable on- the-fly query engine on FPGAs to reduce the resource utilization and pipeline interval. We achieve a 450 million compounds-per- second processing throughput for a single query engine. On approximate search using hierarchical navigable small world (HNSW), a popular algorithm with high recall and query speed, we propose an FPGA-based graph traversal engine to utilize high throughput register array based priority queue and fine- grained distance calculation engine to increase the processing capability. Experimental results show that the proposed FPGA- based HNSW implementation achieves a 35× speedup than existing works on CPU. To the best of our knowledge, our FPGA- based implementation is the first attempt to accelerate molecular similarity search on FPGA and has the highest performance among existing approaches.

Peng, Hongwu↗

3D-Reconstruction of Tau Neutrinos in LArTPC Detectors

The Deep Underground Neutrino Experiment (DUNE) is a next-generation neutrino experiment currently under construction. DUNE will consist of two high-resolution neutrino interaction imaging detectors exposed to the world’s most intense neutrino beam, with the Near Detector at Fermilab and the Far Detector 1,300 km away in the Sanford Underground Research Facility in South Dakota, US. The high statistics and excellent resolution capabilities of DUNE's $^{40}$Ar detector will allow us to make precision studies of oscillation parameters capable of searching for CP violation in the lepton sector, testing interaction models, and studying phenomena that have until now, seemed too complex to measure, like $\nu_\tau$ detection and therefore, providing the completion of the 3-flavor neutrino paradigm. Knowledge of the $\nu_\tau$ detection can impact a broad spectrum of open questions. These include searching for non-standard neutrino interactions, constraining the unitarity of the PMNS matrix, searching for sterile neutrinos, and studying neutrino interactions. In the case of LArTPC data, the detector hits can be considered nodes in a graph, and the edges represent the spatial and temporal relationships between them. By using graph neural networks, it is possible to exploit these relationships and improve the accuracy of particle identification and reconstruction. During my presentation and specifically for tau neutrino reconstruction, I will show the effectiveness and reliability of our in-house developed graph neural network (GNN), NuGraph. This GNN classifies detector hits based on the particle type responsible for their production, assuring that the system accurately identifies and categorizes information based on its unique characteristics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Semi-Automated, Object-Based Tomography of Dislocation Structures

The characterization of the three-dimensional arrangement of dislocations is important for many analyses in materials science. Dislocation tomography in transmission electron microscopy is conventionally accomplished through intensity-based reconstruction algorithms. Although such methods work successfully, a disadvantage is that they require many images to be collected over a large tilt range. Here, we present an alternative, semi-automated object-based approach that reduces the data collection requirements by drawing on the prior knowledge that dislocations are line objects. Our approach consists of three steps: (1) initial extraction of dislocation line objects from the individual frames, (2) alignment and matching of these objects across the frames in the tilt series, and (3) tomographic reconstruction to determine the full three-dimensional configuration of the dislocations. Drawing on innovations in graph theory, we employ a node-line segment representation for the dislocation lines and a novel arc-length mapping scheme to relate the dislocations to each other across the images in the tilt series. We demonstrate the method for a dataset collected from a dislocation network imaged by diffraction-contrast scanning transmission electron microscopy. Based on these results and a detailed uncertainty analysis for the algorithm, we discuss opportunities for optimizing data collection and further automating the method.

47 OTHER INSTRUMENTATION↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗