Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Comparing Individualized Survival Predictions From Random Survival Forests and Multistate Models in the Presence of Missing Data: A Case Study of Patients With Oropharyngeal Cancer

Background: In recent years, interest in prognostic calculators for predicting patient health outcomes has grown with the popularity of personalized medicine. These calculators, which can inform treatment decisions, employ many different methods, each of which has advantages and disadvantages. Methods: We present a comparison of a multistate model (MSM) and a random survival forest (RSF) through a case study of prognostic predictions for patients with oropharyngeal squamous cell carcinoma. The MSM is highly structured and takes into account some aspects of the clinical context and knowledge about oropharyngeal cancer, while the RSF can be thought of as a black-box non-parametric approach. Key in this comparison are the high rate of missing values within these data and the different approaches used by the MSM and RSF to handle missingness. Results: We compare the accuracy (discrimination and calibration) of survival probabilities predicted by both approaches and use simulation studies to better understand how predictive accuracy is influenced by the approach to (1) handling missing data and (2) modeling structural/disease progression information present in the data. We conclude that both approaches have similar predictive accuracy, with a slight advantage going to the MSM. Conclusions: Although the MSM shows slightly better predictive ability than the RSF, consideration of other differences are key when selecting the best approach for addressing a specific research question. These key differences include the methods’ ability to incorporate domain knowledge, and their ability to handle missing data as well as their interpretability, and ease of implementation. Ultimately, selecting the statistical method that has the most potential to aid in clinical decisions requires thoughtful consideration of the specific goals.

60 APPLIED LIFE SCIENCES↗

Charting the chemical space of Zintl phases with graph neural networks and bonding insights

A large number of Zintl phases have been discovered by solid-state chemists driven by empirical knowledge, chemical intuition and in some cases, through serendipitous accidents. These discoveries have only scratched the surface, given the vast compositional and structural diversity that Zintl phases can accommodate. The large chemical space of Zintl phases, as well as intermetallic compounds in general, remain under-explored. Here, we use graph neural networks and the upper bound energy minimization approach to efficiently scan a large chemical space of >90 000 hypothetical Zintl phases and accurately discover 1810 new thermodynamically stable phases with 90% precision, as validated with first-principles calculations. We show that our approach is more than 2× more accurate in predicting DFT stability than M3GNet (40% precision) on the same dataset. Using a random forest model and SHAP analysis, we demonstrate the critical role of ionic bonding in the thermodynamic stability of Zintl phases. Our results not only expand the known chemical landscape of Zintl phases but also highlight the efficacy of machine learning frameworks combined with domain knowledge in uncovering chemically meaningful insights across complex intermetallics.

36 MATERIALS SCIENCE↗

Combustion machine learning: Principles, progress and prospects

Progress in combustion science and engineering has led to the generation of large amounts of data from large-scale simulations, high-resolution experiments, and sensors. This corpus of data offers enormous opportunities for extracting new knowledge and insights—if harnessed effectively. Machine learning (ML) techniques have demonstrated remarkable success in data analytics, thus offering a new paradigm for data-intense analyses and scientific investigations through combustion machine learning (CombML). While data-driven methods are utilized in various combustion areas, recent advances in algorithmic developments, the accessibility of open-source software libraries, the availability of computational resources, and the abundance of data have together rendered ML techniques ubiquitous in scientific analysis and engineering. This article examines ML techniques for applications in combustion science and engineering. Starting with a review of sources of data, data-driven techniques, and concepts, we examine supervised, unsupervised, and semi-supervised ML methods. Various combustion examples are considered to illustrate and to evaluate these methods. Next, we review past and recent applications of ML approaches to problems in combustion, spanning fundamental combustion investigations, propulsion and energy-conversion systems, and fire and explosion hazards. Challenges unique to CombML are discussed and further opportunities are identified, focusing on interpretability, uncertainty quantification, robustness, consistency, creation and curation of benchmark data, and the augmentation of ML methods with prior combustion-domain knowledge.

33 ADVANCED PROPULSION SYSTEMS↗

A deep learning approach to identify missing is-a relations in SNOMED CT

Abstract Objective SNOMED CT is the largest clinical terminology worldwide. Quality assurance of SNOMED CT is of utmost importance to ensure that it provides accurate domain knowledge to various SNOMED CT-based applications. In this work, we introduce a deep learning-based approach to uncover missing is-a relations in SNOMED CT. Materials and Methods Our focus is to identify missing is-a relations between concept-pairs exhibiting a containment pattern (ie, the set of words of one concept being a proper subset of that of the other concept). We use hierarchically related containment concept-pairs as positive instances and hierarchically unrelated containment concept-pairs as negative instances to train a model predicting whether an is-a relation exists between 2 concepts with containment pattern. The model is a binary classifier leveraging concept name features, hierarchical features, enriched lexical attribute features, and logical definition features. We introduce a cross-validation inspired approach to identify missing is-a relations among all hierarchically unrelated containment concept-pairs. Results We trained and applied our model on the Clinical finding subhierarchy of SNOMED CT (September 2019 US edition). Our model (based on the validation sets) achieved a precision of 0.8164, recall of 0.8397, and F1 score of 0.8279. Applying the model to predict actual missing is-a relations, we obtained a total of 1661 potential candidates. Domain experts performed evaluation on randomly selected 230 samples and verified that 192 (83.48%) are valid. Conclusions The results showed that our deep learning approach is effective in uncovering missing is-a relations between containment concept-pairs in SNOMED CT.

97 MATHEMATICS AND COMPUTING↗

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE↗

Nonnegative canonical tensor decomposition with linear constraints: nnCANDELINC

Abstract There is an emerging interest for tensor factorization applications in big‐data analytics and machine learning. To speed up the factorization of extra‐large datasets, organized in multidimensional arrays (also known as tensors), easy to compute compression‐based tensor representations, such as, Tucker and tensor train formats, are used to approximate the initial large‐tensor. Further, tensor factorization is used to extract latent features that can facilitate discoveries of new mechanisms and signatures hidden in the data, where the explainability of the latent features is of principal importance. Nonnegative tensor factorization extracts latent features that are naturally sparse and parts of the data, which makes them easily interpretable. However, to take into account available domain knowledge and subject matter expertise, often additional constraints need to be imposed, which lead us to canonical decomposition with linear constraints (CANDELINC), a canonical polyadic decomposition with rank deficient factors. In CANDELINC, Tucker compression is used as a preprocessing step, which lead to a larger residual error but to more explainable latent features. Here, we propose a nonnegative CANDELINC (nnCANDELINC) accomplished via a specific nonnegative Tucker decomposition; we refer to as minimal or canonical nonnegative Tucker. We derive several results required to understand the specificity of nnCANDELINC, focusing on the difficulties of preserving the nonnegative rank of a tensor to its Tucker core and comparing the real valued to nonnegative case. Finally, we demonstrate nnCANDELINC performance on synthetic and real‐world examples.

97 MATHEMATICS AND COMPUTING↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Image Processing Pipeline for Fluoroelastomer Crystallite Detection in Atomic Force Microscopy Images

Phase transformations in materials systems can be tracked using atomic force microscopy (AFM), enabling the examination of surface properties and macroscale morphologies. In situ measurements investigating phase transformations generate large datasets of time-lapse image sequences. The interpretation of the resulting image sequences, guided by domain-knowledge, requires manual image processing using handcrafted masks. Here this approach is time-consuming and restricts the number of images that can be processed. Her in this study, we developed an automated image processing pipeline which integrates image detection and segmentation methods. We examine five time-series AFM videos of various fluoroelastomer phase transformations. The number of image sequences per video ranges from a hundred to a thousand image sequences. The resulting image processing pipeline aims to automatically classify and analyze images to enable batch processing. Using this pipeline, the growth of each individual fluoroelastomer crystallite can be tracked through time. We incorporated statistical analysis into the pipeline to investigate trends in phase transformations between different fluoroelastomer batches. Understanding these phase transformations is crucial, as it can provide valuable insights into manufacturing processes, improve product quality, and possibly lead to the development of more advanced fluoroelastomer formulations.

36 MATERIALS SCIENCE↗

Physics-based hybrid machine learning for critical heat flux prediction with uncertainty quantification

Critical heat flux (CHF) is a key quantity in nuclear system modeling due to its impact on heat transfer, safety margins, and reactor performance. This study develops and validates an uncertainty-aware hybrid modeling approach that combines machine learning with physics-based models to predict CHF in cases of dryout. The Biasi and Bowring empirical correlations were paired with three ML uncertainty quantification (UQ) techniques: deep neural network (DNN) ensembles, Bayesian neural networks (BNNs), and deep Gaussian processes (DGPs). A pure ML model without a base model was evaluated for comparison. Model performance was assessed under plentiful (7,350 points) and limited (9 points) training data scenarios using parity, uncertainty distributions, and calibration curves. Results show that the Biasi hybrid DNN ensemble achieved the best overall performance, with a mean absolute relative error of 1.846%, and well-calibrated uncertainty estimates. The BNN-based hybrids showed slightly higher error (2.14%) but superior uncertainty calibration. DGP models underperformed, with over 6% error and poor uncertainty calibration. All hybrid models outperformed pure machine learning configurations, demonstrating resistance against data scarcity. These findings indicate that hybrid modeling significantly improves predictive accuracy, interpretability, and resilience to data scarcity. The integration of uncertainty awareness provides actionable confidence in CHF predictions, which is vital for safety-critical decisions in nuclear applications. This hybrid approach offers a viable pathway for deploying ML models in reactor analysis tools while preserving domain knowledge and physical consistency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

edxia: Microstructure characterisation from quantified SEM-EDS hypermaps

The characterisation of cement paste microstructure is an important step towards understanding durability mechanisms in cementitious materials. Scanning electron microscopy (SEM) coupled with energy dispersive spectroscopy (EDS) is a widely used technique to analyse the microstructure at the micron-scale. However, it is challenging, notably because the characteristic size of many phases is found on a scale smaller than the EDS interaction volume. This work presents a new image analysis framework to identify phases and quantify the microstructure of cementitious materials from SEM-EDS hypermaps. By leveraging domain knowledge, representative points are attributed to phases and mixtures of phases based on ratio plots. Then, quantitative analysis of the microstructure can be carried out (chemical composition, particle size distributions, volume fractions, …). We demonstrate the abilities of the framework, and we present possible applications and extensions of the method. The framework is available as both a graphical interface and a Python code.

36 MATERIALS SCIENCE↗

Bayesian Optimization for Anything (BOA): An open-source framework for accessible, user-friendly Bayesian optimization

We introduce Bayesian Optimization for Anything (BOA), a high-level Bayesian Optimization (BO) framework and model wrapping toolkit, which presents a novel approach to simplifying BO, with the goal of making it more accessible and user-friendly, particularly for those with limited expertise in the field. BOA addresses common barriers in implementing BO, focusing on ease of use, reducing the need for deep domain knowledge, and cutting down on extensive coding requirements. A notable feature of BOA is its language-agnostic architecture, which facilitates broader application in various fields and to a wider audience. We showcase BOA's application through three examples: a high-dimensional optimization with parameters of the SWAT+ watershed model, a highly parallelized optimization of this intrinsically non-parallel model, and a multi-objective optimization of the FETCH Tree-Crown Hydrodynamics model. Furthermore, these test cases illustrate BOA's effectiveness in addressing complex optimization challenges in diverse scenarios.

54 ENVIRONMENTAL SCIENCES↗

"Hidden" hydrothermal technical potential & technoeconomics: Revealing permeability & fluids with more data

Historical hydrothermal estimates have largely relied on temperature or heat flow estimates ignoring the need for natural flowing fluids. More accurate hydrothermal estimates require some indication of permeability and fluids that naturally exist in the subsurface. This paper describes a novel approach that includes proxies of permeability and fluids in hydrothermal estimates by leveraging the relatively data-rich Great Basin. Specifically, nameplate capacities (megawatts) of operating geothermal plants, negative (0 megawatt) locations and 48 geophysical and geologic features are used to used in eXtreme Gradient Boosting (XGBoost) regression to make hydrothermal capacity predictions. Additionally, this work inputs the XGBoost-based hydrothermal predictions into the Renewable Energy Potential (reV) model to quantify technical capacity, its uncertainty and techno-economics. Compared to historical hydrothermal estimates, these predictions adhere to the 37 operating geothermal plants and negative locations. We present a method for subsampling the negative sites to bring the labels into balance that uses the geologic domain knowledge to proportionally represent negatives. Overall, the distributions of the hydrothermal technical capacity and the site levelized cost of energy are respectively much tighter, lower and more accurate than the previous estimates for the Great Basin, as they include geological and geophysical surrogates for permeability and fluids. Percentile (50th and 90th, median and high estimate, respectively) models provide bookends for these metrics.

13 HYDRO ENERGY↗

Deep Learning Explicit Differentiable Predictive Control Laws for Buildings

We present a differentiable predictive control (DPC) methodology for learning constrained control laws for unknown nonlinear systems. DPC poses an approximate solution to multiparametric programming problems emerging from explicit nonlinear model predictive control (MPC). Contrary to approximate MPC, DPC does not require supervision by an expert controller. Instead, a system dynamics model is learned from a small dataset of recorded observations of the perturbed system's dynamics and the control law is optimized offline by interaction with the learned system model. The DPC method is based on two sequential steps, i) system identification using a constrained neural state-space model, and ii) optimization of an explicit control law parametrized by another neural network in closed-loop simulation with the identified neural state-space model. The combination of a differentiable closed-loop system and penalty methods for constraint handling of system outputs and inputs allows us to optimize the control law's parameters directly by backpropagating economic MPC loss through the learned system model. By incorporating domain knowledge and leveraging established techniques from optimal control, our method leverages deep neural networks as nonlinear function approximators for system identification and control while avoiding concomitant costs of intractably large datasets, and computationally expensive over-parametrized models. The scalability, data efficiency, and constrained optimal control capability of the proposed DPC method are demonstrated in simulation using a multi-zone building emulator.

Drgona, Jan↗

A machine learning pipeline for membrane segmentation of cryo-electron tomograms

We describe how to use several machine learning techniques organized in a learning pipeline to segment and identify cell membrane structures from cryo electron tomograms. These tomograms are difficult to analyze with traditional segmentation tools. The learning pipeline in our approach starts from supervised learning via a special convolutional neural network trained with simulated data. It continues with semi-supervised reinforcement learning and/or a region merging technique that tries to piece together disconnected components belonging to the same membrane structure. A parametric or non-parametric fitting procedure is then used to enhance the segmentation results and quantify uncertainties in the fitting. Domain knowledge is used in generating the training data for the neural network and in guiding the fitting procedure through the use of appropriately chosen priors and constraints. We demonstrate that the approach proposed here works well for extracting membrane surfaces in two real tomogram datasets.

97 MATHEMATICS AND COMPUTING↗

Systematic feature design for cycle life prediction of lithium-ion batteries during formation

Optimization of the formation step in lithium-ion battery manufacturing is challenging due to limited physical understanding of solid-electrolyte interphase formation and the long testing time (∼100 days) for cells to reach the end of life. We propose a systematic feature-design framework that requires minimal domain knowledge for accurate cycle life prediction during formation. By only using two simple Q (V) features designed from our framework, extracted from formation data without any additional diagnostic cycles, we achieved an average of 9.87% error for cycle life prediction. Here, the physics-based investigation guided by the two designed features shows that the voltage ranges identified by our framework capture the effects of formation temperature and microscopic-particle resistance heterogeneity. By designing highly predictive, robust, and interpretable features, our approach can accelerate industrial battery formation research, leveraging the interplay between data-driven feature design and mechanistic understanding.

25 ENERGY STORAGE↗

Improved departure from nucleate boiling prediction in rod bundles using a physics-informed machine learning-aided framework

The critical heat flux (CHF) corresponding to the departure from nucleate boiling (DNB) crisis is a regulatory limit for the licensing of pressurized water reactors (PWRs) worldwide. Despite the abundance of predictive tools available to the reactor thermal-hydraulics community, the path for an accurate CHF model remains elusive. This work approaches the prediction of DNB through a physics-informed machine learning-aided framework (PIMLAF) with the objective of achieving superior predictive capabilities for a rod bundle. In view of the limitations in existing macro-scale physics-driven tools, an improved mechanistic model is first proposed, leveraging key concepts in the liquid sublayer dryout and bubble crowding mechanisms. Furthermore, the proposed mechanistic model is able to predict DNB in different heater geometries for a broad range of flow conditions without the need for recalibration. This model is then incorporated as the physics-informed component of the hybrid framework PIMLAF, which takes advantage of established understanding in the field (i.e., domain knowledge [DK]) and uses machine learning (ML) to capture undiscovered information from the mismatch between the actual and DK-predicted output. Two bundle-related case studies using the PWR subchannel and bundle tests (PSBT) database are carried out to illustrate the PIMLAF’s improved performance over traditional approaches for both interpolation and extrapolation purposes. In light of the PIMLAF’s promising potential to reduce prediction error, reactor vendors are encouraged to leverage their in-house experimental efforts and apply the hybrid framework to potentially achieve margin reductions in the minimum DNB ratio (MDNBR) for the designs of interest.

42 ENGINEERING↗

A Universal Machine Learning Model for Elemental Grain Boundary Energies

The grain boundary (GB) energy has a profound influence on the grain growth and properties of polycrystalline metals. Here, we show that the energy of a GB, normalized by the bulk cohesive energy, can be described purely by four geometric features. By machine learning on a large computed database of 361 small Σ (Σ<10) GBs of more than 50 metals, we develop a model that can predict the grain boundary energies to within a mean absolute error of 0.13 J m –2 . More importantly, this universal GB energy model can be extrapolated to the energies of high Σ GBs without loss in accuracy. These results highlight the importance of capturing fundamental scaling physics and domain knowledge in the design of interpretable, extrapolatable machine learning models for materials science.

36 MATERIALS SCIENCE↗

Importance of Engineered and Learned Molecular Representations in Predicting Organic Reactivity, Selectivity, and Chemical Properties

Machine-readable chemical structure representations are foundational in all attempts to harness machine learning for the prediction of reactivities, selectivities, and chemical properties directly from molecular structure. The featurization of discrete chemical structures into a continuous vector space is a critical phase undertaken before model selection, and the development of new ways to quantitatively encode molecules is an active area of research. Here, we highlight the application and suitability of different representations, from expert-guided “engineered” descriptors to automatically “learned” features, in different prediction tasks relevant to organic and organometallic chemistry, where differing amounts of training data are available. These tasks include statistical models of stereo- and enantioselectivity, thermochemistry, and kinetics developed using experimental and quantum chemical data. The use of expert-guided molecular descriptors provides an opportunity to incorporate chemical knowledge, domain expertise, and physical constraints into statistical modeling. In applications to stereoselective organic and organometallic catalysis, where data sets may be relatively small and 3D-geometries and conformations play an important role, mechanistically informed features can be used successfully to obtain predictive statistical models that are also chemically interpretable. We provide an overview of several recent applications of this approach to obtain quantitative models for reactivity and selectivity, where topological descriptors, quantum mechanical calculations of electronic and steric properties, along with conformational ensembles, all feature as essential ingredients of the molecular representations used. Alternatively, more flexible, general-purpose molecular representations such as attributed molecular graphs can be used with machine learning approaches to learn the complex relationship between a structure and prediction target. This approach has the potential to out-perform more traditional representation methods such as “hand-crafted” molecular descriptors, particularly as data set sizes grow. One area where this is particularly relevant is in the use of large sets of quantum mechanical data to train quantitative structure–property relationships. A general approach toward curating useful data sets and training highly accurate graph neural network models is discussed in the context of organic bond dissociation enthalpies, where this strategy outperforms regression using precomputed descriptors. Finally, we describe how graph neural network predictions can be incorporated into mechanistically informed statistical models of chemical reactivity and selectivity. Once trained, this approach avoids the expensive computational overhead associated with quantum mechanical calculations, while maintaining chemical interpretability. We illustrate examples for which fast predictions of bond dissociation enthalpy and of the identities of radicals formed through cleavage of a molecule’s weakest bond are used in simple physical models of site-selectivity and reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗