Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “transfer learning.”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

A pre-training and self-training approach for biomedical named entity recognition

Named entity recognition (NER) is a key component of many scientific literature mining tasks, such as information retrieval, information extraction, and question answering; however, many modern approaches require large amounts of labeled training data in order to be effective. This severely limits the effectiveness of NER models in applications where expert annotations are difficult and expensive to obtain. In this work, we explore the effectiveness of transfer learning and semi-supervised self-training to improve the performance of NER models in biomedical settings with very limited labeled data (250-2000 labeled samples). We first pre-train a BiLSTM-CRF and a BERT model on a very large general biomedical NER corpus such as MedMentions or Semantic Medline, and then we fine-tune the model on a more specific target NER task that has very limited training data; finally, we apply semi-supervised self-training using unlabeled data to further boost model performance. We show that in NER tasks that focus on common biomedical entity types such as those in the Unified Medical Language System (UMLS), combining transfer learning with self-training enables a NER model such as a BiLSTM-CRF or BERT to obtain similar performance with the same model trained on 3x-8x the amount of labeled data. We further show that our approach can also boost performance in a low-resource application where entities types are more rare and not specifically covered in UMLS.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

HydraGNN_OPF_GFM_2026 - Ensemble of predictive graph foundation models for power grid applications

This dataset supports research on graph foundation models for optimal power flow (OPF) on electric grids using HydraGNN. It contains heterogeneous graph representations of PGLib-OPF cases spanning systems from 14 to 13,659 buses, together with packed HDF5 datasets for pretraining, feasibility classification, and N-1 contingency analysis. The release includes OPF solution data, downstream fine-tuning datasets, pretrained HeteroSAGE and HeteroHEAT model checkpoints, hyperparameter-optimization summaries across multiple heterogeneous GNN architectures, and aggregated fine-tuning results for sample-efficiency studies. The dataset is designed to enable scalable training, evaluation, and transfer-learning studies for OPF surrogate modeling, including node-level AC-OPF solution prediction, graph-level prediction, feasibility classification, operating-condition generalization, and contingency-response tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Subaru Hyper Suprime-Cam revisits the large-scale environmental dependence on galaxy morphology over 360 deg2 at z = 0.3–0.6

This study investigates the role of large-scale environments on the fraction of spiral galaxies at z = 0.3–0.6 sliced to three redshift bins of Δz = 0.1. Here, we sample 276220 massive galaxies in a limited stellar mass of 5 × 1010 solar mass (~M*) over 360 deg2, as obtained from the Second Public Data Release of the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). By combining projected two-dimensional density information (Shimakawa et al. 2021, MNRAS, 503, 3896) and the CAMIRA cluster catalog (Oguri et al. 2018, PASJ, 70, S20), we investigate the spiral fraction across large-scale overdensities and in the vicinity of red sequence clusters. We adopt transfer learning to reduce the cost of labeling spiral galaxies significantly and then perform stacking analysis across the entire field to overcome the limitations of sample size. Here we employ a morphological classification catalog by the Galaxy Zoo Hubble (Willett et al., 2017, MNRAS, 464, 4176) to train the deep learning model. Based on 74103 sources classified as spirals, we find moderate morphology–density relations on a 10 comoving Mpc scale, thanks to the wide-field coverage of HSC-SSP. Clear deficits of spiral galaxies have also been confirmed, in and around 1136 red sequence clusters. Furthermore, we verify whether there is a large-scale environmental dependence on rest-frame u - r colors of spiral galaxies; such a tendency was not observed in our sample.

Astronomy & Astrophysics↗

PemNet: A Transfer Learning-Based Modeling Approach of High-Temperature Polymer Electrolyte Membrane Electrochemical Systems

Widespread adoption of high-temperature electrochemical systems such as polymer electrolyte membrane fuel cells (HT-PEMFCs) requires models and computational tools for accurate optimization and guiding new materials for enhancing fuel cell performance and durability. Furthermore, while robust and better suited for extrapolation, knowledge-based modeling has limitations as it is time-consuming and requires information about the system that is not always available (e.g., material properties and interfacial behavior between different materials). Data-driven modeling, on the other hand, is easier to implement but often necessitates large datasets that could be difficult to obtain. In this contribution, knowledge-based modeling and data-driven modeling are combined by implementing a few-shot learning (FSL) approach. A knowledge-based model originally developed for a HT-PEMFCs was used to generate simulated data (887,735 points) and used to pretrain a neural network source model tuned via a genetic algorithm-based AutoML. Then, experimental datasets from HT-PEMFCs with different materials and operating conditions (~50 points each) were used to train six target models via FSL. Models for the unseen data reached high accuracies in all cases (rRMSE < 10%).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,↗

SMALE: Enhancing Scalability of Machine Learning Algorithms on Extreme-Scale Computing Platforms

Deployment and execution of machine learning tasks on extreme-scale computing platforms face several significant technical challenges: 1) High computing cost incurred by dense networks – The computing workload of deep networks with densely-connected topology increases rapidly with the network size, imposing a non-scalable computing model of extreme-scale computing platforms; 2) Non-optimized workload distribution – Many advanced deep learning algorithms, e.g., sparsification and irregular net-work topology, produce very unbalanced workload distribution on extreme-scale computing platforms. The computation efficiency is greatly hindered by the incurred data and computation redundancies as well as long tails of the node with extensive workload; 3) Constraints in data movement and I/O bottle-neck – Inter-node data movement in extreme-scale computing platforms are associated with high energy and latency costs, and subject to the constraints of I/O bandwidth; and 4) Generalization of algorithm realization and acceleration on computing platforms – The large varieties of machine learning algorithms and structures of extreme-scale computing platforms make the derivation of a generalized algorithm realization and acceleration method very challenging, which, however, is the requirement by domain scientists and interested users. We call the above challenges Smale’s Problems in Machine Learning and Understanding for High-Performance Computing Scientific Discovery. The objective of our three-year research project is to develop a holistic innovation set at structure, assembly, and acceleration layers of machine learning algorithms to address the above challenges in algorithm deployment and execution. Three tasks are particularly performed, including: At the algorithm structure level, we investigate the techniques that can structurally sparsify on the topology of deep networks for computing workload reduction. We also study clustering and pruning techniques that can optimize the workload distributions over the extreme-scale computing platforms; At the algorithm assembly level, we derive a unified learning framework for unsupervised transfer learning and dynamic growing capabilities. Novel training methods are also exploited to enhance the training efficiency of the proposed framework; At the algorithm acceleration level, we will develop a series of techniques that can accelerate the computation of sparse matrix operations, which are one of the core executions in deep learning and optimize memory access of the concerned platforms. Our proposed techniques attack the fundamental problems in machine learning algorithms running on extreme-scale computing platforms by vertically integrating the solutions at three closely entangled layers, paving the long-term scaling path of machine learning applications under DOE context. Three tasks corresponding to the above respective research orientations are performed during the three-year project period with our collaborators at ORNL. The outcome of the proposed project is anticipated to form a holistic solution set of novel algorithms and network topologies, efficient training techniques, and fast acceleration methods to promote the computing scalability of the machine learning applications of particular interest to DOE.

97 MATHEMATICS AND COMPUTING↗

New data-driven approach to bridging power system protection gaps with deep learning

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this paper, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A combined convolutional neural network and long short-term memory (CNN-LSTM) network is implemented to achieve data translation invariance and capture the temporal correlation of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN-LSTM--based method avoids the complicated, manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on two kinds of protection gaps: high-impedance faults and transformer inter-turn faults. Lastly, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

42 ENGINEERING↗

Leveraging Multitime Hamilton–Jacobi PDEs for Certain Scientific Machine Learning Problems

Hamilton-Jacobi partial differential equations (HJ PDEs) have deep connections with a wide range of fields, including optimal control, differential games, and imaging sciences. By considering the time variable to be a higher dimensional quantity, HJ PDEs can be extended to the multi-time case. In this paper, we establish a novel theoretical connection between specific optimization problems arising in machine learning and the multi-time Hopf formula, which corresponds to a representation of the solution to certain multi-time HJ PDEs. Through this connection, we increase the interpretability of the training process of certain machine learning applications by showing that when we solve these learning problems, we also solve a multi-time HJ PDE and, by extension, its corresponding optimal control problem. As a first exploration of this connection, we develop the relation between the regularized linear regression problem and the Linear Quadratic Regulator (LQR). We then leverage our theoretical connection to adapt standard LQR solvers (namely, those based on the Riccati ordinary differential equations) to design new training approaches for machine learning. Lastly, we provide some numerical examples that demonstrate the versatility and possible computational advantages of our Riccati-based approach in the context of continual learning, post-training calibration, transfer learning, and sparse dynamics identification.

97 MATHEMATICS AND COMPUTING↗

Technology Transfer

The objective of this summer's work was to attempt to enhance Technology Application Group (TAG) ability to measure the outcomes of its efforts to transfer NASA technology. By reviewing existing literature, by explaining the economic principles involved in evaluating the economic impact of technology transfer, and by investigating the LaRC processes our William & Mary team has been able to lead this important discussion. In reviewing the existing literature, we identified many of the metrics that are currently being used in the area of technology transfer. Learning about the LaRC technology transfer processes and the metrics currently used to track the transfer process enabled us to compare other R&D facilities to LaRC. We discuss and diagram impacts of technology transfer in the short run and the long run. Significantly, it serves as the basis for analysis and provides guidance in thinking about what the measurement objectives ought to be. By focusing on the SBIR Program, valuable information regarding the strengths and weaknesses of this LaRC program are to be gained. A survey was developed to ask probing questions regarding SBIR contractors' experience with the program. Specifically we are interested in finding out whether the SBIR Program is accomplishing its mission, if the SBIR companies are providing the needed innovations specified by NASA and to what extent those innovations have led to commercial success. We also developed a survey to ask COTR's, who are NASA employees acting as technical advisors to the SBIR contractors, the same type of questions, evaluating the successes and problems with the SBIR Program as they see it. This survey was developed to be implemented interactively on computer. It is our hope that the statistical and econometric studies that can be done on the data collected from all of these sources will provide insight regarding the direction to take in developing systematic evaluations of programs like the SBIR Program so that they can reach their maximum effectiveness.

Smith, Nanette R.↗

ABF DFO with Technology Holding, Inc.

This Agile BioFoundry Directed Funding Opportunity project with Technology Holding and partners focuses on the development of both a strain of Pseudomonas putida KT2440 and a corresponding bioprocess to convert cellulosic sugars to beta-ketoadipic acid, which can be used in performance nylons and polyesters. Our approach follows the Design-Build-Test-Learn cycle wherein we have transferred learnings from muconic acid production in P. putida to develop a glucose and xylose-utilizing beta-ketoadipic acid production strain. This strain achieves 65 g/L of beta-ketoadipic acid at 0.7 g/L/hr and a C-mol yield of 0.40. We are currently on-boarding arabinose utilization as well. To identify non-intuitive strain modifications as well, we are deploying a beta-ketoadipic acid biosensor and building randomly barcoded transposon insertion sequencing (RB-TnSeq) libraries and gene over-expression libraries in beta-ketoadipic acid production strains. Moreover, we are using global metabolomics and other systems biology tools to identify off-target pathways. Lastly, we are scaling up beta-ketoadipic acid production to kg-scale production for Technology Holding to evaluate in performance polymers with their partners.

beta-ketoadipic acid↗

Deep reinforcement learning control of hydraulic fracturing

Hydraulic fracturing is a technique to extract oil and gas from shale formations, and obtaining a uniform proppant concentration along the fracture is key to its productivity. Recently, various model predictive control schemes have been proposed to achieve this objective. But such controllers require an accurate and computationally efficient model which is difficult to obtain given the complexity of the process and uncertainties in the rock formation properties. In this article, we design a model-free data-based reinforcement learning controller which learns an optimal control policy through interactions with the process. Deep reinforcement learning (DRL) controller is based on the Deep Deterministic Policy Gradient algorithm that combines Deep-Q-network with actor-critic framework. In addition, we utilize dimensionality reduction and transfer learning to quicken the learning process. We show that the controller learns an optimal policy to obtain uniform proppant concentration despite the complex nature of the process while satisfying various input constraints.

42 ENGINEERING↗

Identifying common stored product insects using automated deep learning methods

Monitoring stored product insect pests is a common practice for post-harvest management of stored grain and grain-based commodities, which helps ensure product quality from harvest to final consumer. Current methods of sampling and monitoring can be time-consuming, labor-intensive, expensive and require expertise in insect identification. Therefore, this study aims to develop an image-based automated identification system for common stored product insect species using deep-learning methods. Top-down images of the common stored product adult insect species of Rhyzopertha dominica, Cryptolestes ferrugineus, Tribolium castaneum, Sitophilus oryzae, and Oryzaephilus surinamensis were acquired and analyzed. Deep learning-based, state-of-the-art Convolutional Neural Networks (CNN) models (ResNet-50, MobileNet-v2, DarkNet-53, and EfficientNet-b0) were fine-tuned with a transfer learning approach to classify the insect species. All models were able to correctly identify the insect species with at least 96% accuracy and with few misclassifications. One issue with trained CNNs is that they do not explain the reasoning for the classification and are often called a “black box”. Therefore, visualization methods called Gradient-weighted Class Activation Mapping (Grad-CAM) were implemented to explore the black box network. The Grad-CAM uses heat maps to highlight the major image features that the network focused on to make insect species predictions. The Grad-CAM verifies the network's prediction and also helps improve network performance. This study contributes to the overall goal of developing a camera-based system for monitoring stored grain insects. As a result, the developed system would empower warehouse, flour mills, and other food facilities with a tool to quickly and accurately identify insect species in stored product environments and could be implemented as part of a close to real-time monitoring system.

60 APPLIED LIFE SCIENCES↗