Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Knowledge constrained learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Posterior Regularized Bayesian Neural Network

Traditional NNs often lack the ability for uncertainty quantification. Bayesian NNs(BNNs) could help measure the confidence level by using distributions in NNs modeling. Besides, knowledge is commonly available and could improve the performance of BNNs if it can be properly incorporated. In this work, we propose a novel Posterior-Regularized BNN(PR-BNN) model by incorporating soft and hard constraints as a posterior regularization term. We also propose an augmented Lagrangian method and stochastic optimization algorithm for efficient updating via Monte Carlo sampling. The simulations and case studies for solar PV plants have shown the performance improvement of the proposed model over traditional BNNs.

97 MATHEMATICS AND COMPUTING↗

High‐Asymmetry Metasurface: A New Solution for Terahertz Resonance via Active Learning‐Augmented Diffusion Model

Terahertz (THz) metamaterials with high‐figure‐of‐merit (high‐FoM) performance resonance are essential for advancing sensors, detectors, and imagers. Conventional designs focus on symmetric or low‐asymmetry geometric structures, leaving high‐asymmetry designs largely unexplored due to the inefficiency of trial‐and‐error‐based rational design. Recent deep learning techniques offer automation and acceleration but are constrained by the need for large datasets inherent to their data‐driven nature. Here, a novel prior knowledge‐guided generative model augmented by a physics‐constrained active learning mechanism to design high‐asymmetry metamaterials. An advanced diffusion model learns features from a small set of classical structures with high‐FoM THz resonance and generates new high‐asymmetry structures. To mitigate the limited number of classical structures, the generated high‐asymmetry structures are actively selected and integrated into the initial training dataset based on their physical characteristics. Experimental results demonstrate the superior resonance performance of the generated high‐asymmetry metamaterials over classical designs, exhibiting improvements exceeding 30% in key resonance metrics. Remarkably, this performance is attained using only 68 classical structures as the initial training dataset, significantly reducing the data requirements for deep learning‐based metamaterial design. The proposed scheme for generating high‐asymmetry structures provides a new effective and efficient solution for high‐FoM resonance, expanding applications in high‐sensitivity THz metadevices.

diffusion model↗

Continental-Scale Controls on Hyporheic Respiration Revealed by Knowledge-Guided Machine Learning

Hyporheic zone sediments regulate organic matter turnover and in-stream respiration, yet controls on sediment respiration remain poorly constrained across heterogeneous river networks, limiting prediction of stream metabolism and carbon processing at continental scales. Here, we integrate observations from ~90 river corridors across the United States in the WHONDRS consortium with a knowledge-guided machine learning (KGML) framework that couples thermodynamic rate theory with machine learning to identify dominant controls on hyporheic respiration. Diagnostic analyses show that organic matter concentration and thermodynamic favorability define an upper bound on respiration potential, whereas biological catalytic capacity and physical accessibility jointly govern realized respiration rates through interaction effects. To represent unmeasurable accessibility constraints, we use the mechanistic model as a scaffold for KGML, allowing machine learning to target residual structure not explained by process theory. This hybrid framework improves predictive skill relative to both the mechanistic model alone and fully data-driven models while preserving interpretability. These results indicate that variability in hyporheic respiration is largely mechanistically structured and demonstrate how integrating process theory with explainable AI enhances predictive performance while enabling scalable synthesis of river corridor observations.

Zheng, Jianqiu↗

Knowledge-guided machine learning can improve carbon cycle quantification in agroecosystems

Abstract Accurate and cost-effective quantification of the carbon cycle for agroecosystems at decision-relevant scales is critical to mitigating climate change and ensuring sustainable food production. However, conventional process-based or data-driven modeling approaches alone have large prediction uncertainties due to the complex biogeochemical processes to model and the lack of observations to constrain many key state and flux variables. Here we propose a Knowledge-Guided Machine Learning (KGML) framework that addresses the above challenges by integrating knowledge embedded in a process-based model, high-resolution remote sensing observations, and machine learning (ML) techniques. Using the U.S. Corn Belt as a testbed, we demonstrate that KGML can outperform conventional process-based and black-box ML models in quantifying carbon cycle dynamics. Our high-resolution approach quantitatively reveals 86% more spatial detail of soil organic carbon changes than conventional coarse-resolution approaches. Moreover, we outline a protocol for improving KGML via various paths, which can be generalized to develop hybrid models to better predict complex earth system dynamics.

54 ENVIRONMENTAL SCIENCES↗

Coherency-Aware Learning Control of Inverter-Dominated Grids: A Distributed Risk-Constrained Approach

Here, this letter investigates the importance of integrating the coherency knowledge for designing controllers to dampen sustained oscillations in wide-area power networks with significant penetration of inverter-interfaced resources. Coherency is a fundamental property of power systems, where time-scale separation in frequency dynamics leads to clustered behavior among generators of different groups. Large-scale penetration of inverter-driven low inertia resources replacing conventional synchronous generators (SGs) can lead to perturbation in the coherent partitioning; hence, integrating such information is of utmost importance for oscillation control designs. We present the coherency-aware design of a distributed output feedback-based reinforcement learning method that additionally incorporates risk constraints to capture the uncertainties related to net-load fluctuations. The use of domain-aware coherency information has produced improved training and oscillation performance than the coherency-agnostic control design, hence proving to be effective in controller design. Finally, we validated the proposed method with numerical experiments on the benchmark IEEE 68-bus test system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Leveraging transfer learning and leaf spectroscopy for leaf trait prediction with broad spatial, species, and temporal applicability

Accurate and reliable prediction of leaf traits is crucial for understanding plant adaptations to environmental variation, monitoring terrestrial ecosystems, and enhancing comprehension of functional diversity and ecosystem functioning. Currently, various approaches (e.g., statistical, physical models) have been developed to estimate leaf traits through hyperspectral remote sensing and leaf spectroscopy. However, the absence of high-performing, transferable, and stable models across various domains of space, plant functional types (PFTs) and seasons hinder our ability to quantify and comprehend spatiotemporal variations in leaf traits. This study proposes robust and highly transferable models for better predicting leaf traits with hyperspectral reflectance. Initially, three datasets were assembled, pairing common leaf traits — chlorophyll (Chla+b), carotenoids (Ccar), leaf mass per area (LAM), equivalent water thickness (EWT) — with leaf spectra measurements collected across diverse geographic locations in the U.S. and Europe, PFTs, and seasons. Measurements were acquired using spectroradiometers (e.g., ASD FieldSpec 3/4/Pro and SVC HR-1024i) with integrating spheres, leaf clips, and contact probes. Here, we then developed transfer learning-based hybrid models that incorporated the domain knowledge of radiative transfer models (RTMs) through pretraining processes and were well-constrained by fine-tuning with field measurements. Through comparison with other state-of-the-art statistical models, including partial-least squares regression (PLSR) and Gaussian Process Regression (GPR), as well as pure physical models, we found that the proposed transfer learning models achieved better predictive performance and higher transferability. Specifically, compared to other statistical models and pure RTMs, the transfer learning model exhibited higher coefficient of determination (R 2 ) values with range of 0.01 to 0.79, lower normalized root mean square error (NRMSE) with range of 0.06 % to 33.25 % in model performance. Additionally, the models exhibited improved transferability, with higher R 2 values range from 0.04 to 0.32, lower NRMSE range from 0.08 % to 30.81 %. The findings underscore that transfer learning models through integrating domain knowledge from RTMs and limited observations, can harness the advantages of both RTMs and statistical models and serve as a promising approach for effectively predicting leaf traits.

59 BASIC BIOLOGICAL SCIENCES↗

Distilling Knowledge from Ensembles of Cluster-Constrained-Attention Multiple-Instance Learners for Whole Slide Image Classification

The peculiar nature of whole slide imaging (WSI), digitizing conventional glass slides to obtain multiple high resolution images which capture microscopic details of a patient’s histopathological features, has garnered increased interest from the computer vision research community over the last two decades. Given the unique computational space and time complexity inherent to gigapixel-size whole slide image data, researchers have proposed novel machine learning algorithms to aid in the performance of diagnostic tasks in clinical pathology. One effective algorithm represents a Whole slide image as a bag of smaller image patches, which can be represented as low-dimension image patch embeddings. Weakly supervised deep-learning methods, such as cluster-constrained-attention multiple instance learning (CLAM), have shown promising results when combined with image patch embeddings. While traditional ensemble classifiers yield improved task performance, such methods come with a steep cost in model complexity. Through knowledge distillation, it is possible to retain some performance improvements from an ensemble, while minimizing costs to model complexity. In this work, we implement a weakly supervised ensemble using clustering-constrained-attention multiple-instance learners (CLAM), which uses attention and instance-level clustering to identify task salient regions and feature extraction in whole slides. By applying logit-based and attention-based knowledge distillation, we show it is possible to retain some performance improvements resulting from the ensemble at zero cost to model complexity.

Alamudun, Folami↗

Knowledge Distillation for Anomaly Detection

Unsupervised deep learning techniques are widely used to identify anomalous behaviour. The performance of such methods is a product of the amount of training data and the model size. However, the size is often a limiting factor for the deployment on resource-constrained devices. Here, we present a novel procedure based on knowledge distillation for compressing an unsupervised anomaly detection model into a supervised deployable one and we suggest a set of techniques to improve the detection sensitivity. Compressed models perform comparably to their larger counterparts while significantly reducing the size and memory footprint.

Pol, Adrian Alan↗

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

An evaluation of GPT models for phenotype concept recognition

Clinical deep phenotyping and phenotype annotation play a critical role in both the diagnosis of patients with rare disorders as well as in building computationally-tractable knowledge in the rare disorders field. These processes rely on using ontology concepts, often from the Human Phenotype Ontology, in conjunction with a phenotype concept recognition task (supported usually by machine learning methods) to curate patient profiles or existing scientific literature. With the significant shift in the use of large language models (LLMs) for most NLP tasks, we examine the performance of the latest Generative Pre-trained Transformer (GPT) models underpinning ChatGPT as a foundation for the tasks of clinical phenotyping and phenotype annotation. The experimental setup of the study included seven prompts of various levels of specificity, two GPT models (gpt-3.5-turbo and gpt-4.0) and two established gold standard corpora for phenotype recognition, one consisting of publication abstracts and the other clinical observations. The best run, using in-context learning, achieved 0.58 document-level F1 score on publication abstracts and 0.75 document-level F1 score on clinical observations, as well as a mention-level F1 score of 0.7, which surpasses the current best in class tool. Without in-context learning, however, performance is significantly below the existing approaches. Our experiments show that gpt-4.0 surpasses the state of the art performance if the task is constrained to a subset of the target ontology where there is prior knowledge of the terms that are expected to be matched. While the results are promising, the non-deterministic nature of the outcomes, the high cost and the lack of concordance between different runs using the same prompt and input make the use of these LLMs challenging for this particular task.

59 BASIC BIOLOGICAL SCIENCES↗

DiffLense: a conditional diffusion model for super-resolution of gravitational lensing data

Abstract Gravitational lensing data is frequently collected at low resolution due to instrumental limitations and observing conditions. Machine learning-based super-resolution techniques offer a method to enhance the resolution of these images, enabling more precise measurements of lensing effects and a better understanding of the matter distribution in the lensing system. This enhancement can significantly improve our knowledge of the distribution of mass within the lensing galaxy and its environment, as well as the properties of the background source being lensed. Traditional super-resolution techniques typically learn a mapping function from lower-resolution to higher-resolution samples. However, these methods are often constrained by their dependence on optimizing a fixed distance function, which can result in the loss of intricate details crucial for astrophysical analysis. In this work, we introduce DiffLense , a novel super-resolution pipeline based on a conditional diffusion model specifically designed to enhance the resolution of gravitational lensing images obtained from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). Our approach adopts a generative model, leveraging the detailed structural information present in Hubble space telescope (HST) counterparts. The diffusion model, trained to generate HST data, is conditioned on HSC data pre-processed with denoising techniques and thresholding to significantly reduce noise and background interference. This process leads to a more distinct and less overlapping conditional distribution during the model’s training phase. We demonstrate that DiffLense outperforms existing state-of-the-art single-image super-resolution techniques, particularly in retaining the fine details necessary for astrophysical analyses.

Computer Science↗

Predicting Atomistic Transitions with Transformers

Accurate knowledge of the atomistic transition pathways in materials and material surfaces is crucial for many material science problems. However, conventional simulation techniques used to find these transitions are extremely computationally intensive. Even with large-scale, accelerated material simulations, the computational cost constrains the applicable domain in practice. Machine learning models, with the potential to learn the complex emergent behaviors governing atomistic transitions as a fast surrogate model, have great promise to predict transitions with a vastly reduced computational cost. Here, we demonstrate how transformers can be trained to predict atomistic transitions in nano-clusters. We show how we evaluate physical validity of the predictions and how a multitude of additional, different microstates can be generated by slightly varying the data provided to the model.

36 MATERIALS SCIENCE↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Domain-aware Control-oriented Neural Models for Autonomous Underwater Vehicles

Conventional physics-based modeling is a time-consuming bottleneck in control design for complex nonlinear systems like autonomous underwater vehicles (AUVs). In contrast, purely data-driven models, require a large number of observations and lack operational guarantees for safety-critical systems. Data-driven models leveraging available partially characterized dynamics have potential to provide reliable systems models in a typical data-limited scenario for high value complex systems, thereby avoiding months of expensive expert modeling time. In this work we explore this middle-ground between expert-modeled and pure data-driven modeling. We present control-oriented parametric models with varying levels of domain-awareness that exploit known system structure and prior physics knowledge to create constrained deep neural dynamical system models. We employ universal differential equations to construct data-driven blackbox and graybox representations of the AUV dynamics. In addition, we explore a hybrid formulation that explicitly models the residual error related to imperfect graybox models. We compare the prediction performance of the learned models for different distributions of initial conditions and control inputs to assess their suitability for control.

Shaw Cortez, Wenceslao E.↗

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving multiwell petrophysical interpretation from well logs via machine learning and statistical models

Well-log interpretation estimates in situ rock properties along well trajectory, such as porosity, water saturation, and permeability, to support reserve-volume estimation, production forecasts, and decision making in reservoir development. However, due to measurement errors, variability of well logs caused by multiple measurement vendors, different borehole tools, and nonuniform drilling/borehole conditions, estimations of rock properties with original well logs without proper preprocessing may not be accurate, especially in the context of multiwell estimation. Well-log normalization techniques such as two-point scaling and mean-variance normalization are commonly used to improve the robustness of multiwell rock-property estimation. However, these techniques do not consider the correlation between well logs and require subjective knowledge for their effective implementation. To reduce uncertainties and processing time associated with multiwell rock-property estimation from well logs, we develop discriminative adversarial (DA) and linear constraint models for well-log normalization and rock-property estimation. The DA neural network model developed for well-log normalization and interpretation can perform linear and nonlinear well-log normalization while considering the joint distribution of each well log and rock properties. However, the linear constraint model uses an ensemble of predictions from linear models to constrain well-log normalization and rock-property estimation. We also develop a divergence-based type well identification method to select type (training) wells for a test well based on the statistical similarity of associated well-log distributions instead of the interwell distance. We apply the DA model to perform well-log normalization and prediction of permeability for the Seminole San Andres Unit carbonate reservoir. Compared with the permeability predicted with the classical machine learning model without well-log normalization and models with two-point scaling normalization, the DA model yields the most accurate permeability prediction by decreasing the mean-squared error of permeability prediction by 20%–50%.

Geochemistry & Geophysics↗

Stability-Constrained Learning for Frequency Regulation in Power Grids With Variable Inertia

The increasing penetration of converter-based renewable generation has resulted in faster frequency dynamics, and low and variable inertia. As a result, there is a need for frequency control methods that are able to stabilize a disturbance in the power system at timescales comparable to the fast converter dynamics. This paper proposes a combined linear and neural network controller for inverter-based primary frequency control that is stable at time-varying levels of inertia. We model the time-variance in inertia via a switched affine hybrid system model. We derive stability certificates for the proposed controller via a quadratic candidate Lyapunov function. We test the proposed control on a 12-bus 3-area test network, and compare its performance with a base case linear controller, optimized linear controller, and finite-horizon Linear Quadratic Regulator (LQR). Our proposed controller achieves faster mean settling time and over 50% reduction in average control cost across 100 inertia scenarios compared to the optimized linear controller. Unlike LQR which requires complete knowledge of the inertia trajectories and system dynamics over the entire control time horizon, our proposed controller is real-time tractable, and achieves comparable performance to LQR.

data-driven control↗

A gradient-based deep neural network model for simulating multiphase flow in porous media

We report simulation of multiphase flow in porous media is crucial for the effective management of subsurface energy and environment-related activities. The numerical simulators used for modeling such processes rely on spatial and temporal discretization of the governing mass and energy balance partial-differential equations (PDEs) into algebraic systems via finite-difference/volume/element methods. These simulators usually require dedicated software development and maintenance, and suffer low efficiency from a runtime and memory standpoint for problems with multi-scale heterogeneity, coupled-physics processes or fluids with complex phase behavior. Therefore, developing cost-effective, data-driven models can become a practical choice, and in this work, we choose deep learning approaches as they can handle high dimensional data and accurately predict state variables with strong nonlinearity. In this paper, we describe a gradient-based deep neural network (GDNN) constrained by the physics related to multiphase flow in porous media. We tackle the nonlinearity of flow in porous media induced by rock heterogeneity, fluid properties, and fluid-rock interactions by decomposing the nonlinear PDEs into a dictionary of elementary differential operators. We use a combination of operators to handle rock spatial heterogeneity and fluid flow by advection. Since the augmented differential operators are inherently related to the physics of fluid flow, we treat them as first principles prior knowledge to regularize the GDNN training. We use the example of pressure management at geologic CO 2 storage sites, where CO 2 is injected in saline aquifers and brine is produced, and apply GDNN to construct a predictive model that is trained with physics-based simulation data and emulates the physics process. We demonstrate that GDNN can effectively predict the nonlinear patterns of subsurface responses, including the temporal and spatial evolution of the pressure and saturation plumes. We also successfully extend the GDNN to convolutional neural network (CNN), namely gradient-based CNN (GCNN), and validate its capability to improve the prediction accuracy. GDNN has great potential to tackle challenging problems that are governed by highly nonlinear physics and enable the development of data-driven models with higher fidelity.

42 ENGINEERING↗