Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep transfer learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Accurate Machine Learning for Predicting the Viscosities of Deep Eutectic Solvents

Deep eutectic solvents (DESs) are emerging as environmentally friendly designer solvents for mass transport and heat transfer processes in industrial applications; however, the lack of accurate tools to predict and thus control their viscosities under both a range of environmental factors and formulations hinders their general application. While DESs may serve as designer solvents, with nearly unlimited combinations, this unfortunately makes it experimentally infeasible to comprehensively measure the viscosities of all DESs of potential industrial interest. To assist in the design of DESs, we have developed several new machine learning (ML) models that accurately and rapidly predict the viscosities of a diverse group of DESs at different temperatures and molar ratios using, to date, one of the most comprehensive data sets containing the properties of over 670 DESs over a wide range of temperatures (278.15–385.25 K). Three ML models, including support vector regression (SVR), feed forward neural networks (FFNNs), and categorical boosting (CatBoost), were developed to predict DES viscosity as a function of temperature and molar ratio and contrasted with multilinear and two-factor polynomial regression baselines. Further, quantum chemistry-based, COSMO-RS-derived sigma profile (σ-profile) features were used as inputs for the ML models. The CatBoost model is excellent at externally predicting DES viscosity, as indicated by high R 2 (0.99) and low root-mean-square-error (RMSE) and average absolute relative deviations (AARD) (5.22%) values for the testing data sets, and 98% of the data points lie within the 15% of AARD deviations. Furthermore, SHapley additive explanation (SHAP) analysis was employed to interpret the ML results and rationalize the viscosity predictions. The result is an ML approach that accurately predicts viscosity and will aid in accelerating the design of appropriate DESs for industrial applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ensemble learning-iterative training machine learning for uncertainty quantification and automated experiment in atom-resolved microscopy

Deep learning has emerged as a technique of choice for rapid feature extraction across imaging disciplines, allowing rapid conversion of the data streams to spatial or spatiotemporal arrays of features of interest. However, applications of deep learning in experimental domains are often limited by the out-of-distribution drift between the experiments, where the network trained for one set of imaging conditions becomes sub-optimal for different ones. This limitation is particularly stringent in the quest to have an automated experiment setting, where retraining or transfer learning becomes impractical due to the need for human intervention and associated latencies. Here we explore the reproducibility of deep learning for feature extraction in atom-resolved electron microscopy and introduce workflows based on ensemble learning and iterative training to greatly improve feature detection. This approach allows incorporating uncertainty quantification into the deep learning analysis and also enables rapid automated experimental workflows where retraining of the network to compensate for out-of-distribution drift due to subtle change in imaging conditions is substituted for human operator or programmatic selection of networks from the ensemble. This methodology can be further applied to machine learning workflows in other imaging areas including optical and chemical imaging.

36 MATERIALS SCIENCE↗

Overcoming small minirhizotron datasets using transfer learning

Minirhizotron technology is widely used to study root growth and development. Yet, standard approaches for tracing roots in minirhiztron imagery is extremely tedious and time consuming. Machine learning approaches can help to automate this task. However, lack of enough annotated training data is a major limitation for the application of machine learning methods. Transfer learning is a useful technique to help with training when available datasets are limited. In this paper, we investigated the effect of pre-trained features from the massives-cale, irrelevant ImageNet dataset and a relatively moderate-scale, but relevant peanut root dataset on switchgrass root imagery segmentation applications. We compiled two minirhizotron image datasets to accomplish this study: one with 17,550 peanut root images and another with 28 switchgrass root images. Both datasets were paired with manually labeled ground truth masks. Deep neural networks based on the U-net architecture were used with different pre-trained features as initialization for automated, precise pixel-wise root segmentation in minirhizotron imagery. We observed that features pre-trained on a closely related but relatively moderate size dataset like our peanut dataset were more effective than features pre-trained on the large but unrelated ImageNet dataset. Here, we achieved high quality segmentation on peanut root dataset with 99.04% accuracy at the pixel-level and overcame errors in human-labeled ground truth masks. By applying transfer learning technique on limited switchgrass dataset with features pre-trained on peanut dataset, we obtained 99% segmentation accuracy in switchgrass imagery using only 21 images for training (fine tuning). Furthermore, the peanut pre-trained features can help the model converge faster and have much more stable performance.

59 BASIC BIOLOGICAL SCIENCES↗

Deep Learning for Automated Identification of Eels in Sonar Data

Freshwater eels, such as the American eel (Anguilla rostrata) present numerous challenges related to safe downstream fish passage at hydroelectric facilities. One of those challenges is effective monitoring of their abundance, movements, and behavior to facilitate design and operation of eel protection and passage facilities. A previous EPRI study documented the ability of human analysts to reliably identify American eels in data obtained with a 1100/1800 kHz, multibeam sonar. This report describes a project to develop deep learning (a subset of artificial intelligence) tools to automate the time-consuming, subjective process of eel identification in multibeam sonar data. The project exploited new data collected in the laboratory and the existing data from the prior EPRI field study to develop and test deep learning and other data analytic tools, including wavelet filtering, differencing for static object removal, and convolutional neural network analysis. The analysis of the laboratory data demonstrated feasibility of the approach, revealed object characteristics observed with the sonar that distinguish eels from similarly sized and shaped acoustic targets, and provided additional data for algorithm selection and training. Deep learning algorithms trained and tested on the laboratory data alone achieved accuracy rates of greater than 98% when classifying acoustic images of eels and similar-sized neutrally buoyant sticks. The algorithm trained and tested on the pre-existing field data alone, and yielded classification accuracy of 9.3% false positives and 13.3% false negatives when distinguishing between eels and sticks/PVC pipes based on video clips (i.e., multiple, consecutive images). This performance is comparable to the classification accuracy achieved by human analysts in the prior study. The deep learning algorithm trained on a combination of video clips obtained in the laboratory and the field and tested on video clips from the field, was able to distinguish eels from sticks and PVC pipes (a river debris analog) of similar size with 100% accuracy. Outreach to the hardware, software, and end-user communities early in the project helped to identify needs and specify the application space. Outreach to those communities at the end of the project communicated project results and opportunities for further development. The project achieved proof of concept for automated identification of eel in multibeam sonar data. Future work should focus on acquisition of additional data for more robust algorithm training and testing; modification of the software tools to accommodate multiple acoustic targets in the acoustic field at a given time; identification of additional object classes; incorporation of motion in the object identification and classification algorithms; operationalizing the software tools, including integration with other existing sonar data analysis tools; and partnering with hardware and software providers for distribution of the software tools with their commercial products.

13 HYDRO ENERGY↗

Transfer-Learnt Energy Models for Predicting Electricity Consumption in Buildings with Limited and Sparse Field Data

Modeling energy consumption is critical for energy-efficient utilization of the electric appliances in a building, smart grid programs (like demand-response), and many other smart home applications. State-of-the-art energy modeling techniques either rely on theoretical models, or extensive instrumentation of the building envelope to gather ``big" data to train a deep neural network. While theoretical models are often limited by their estimation accuracy, it is not always feasible to gather a significant amount of field data. In this paper, we explore transfer learning-based strategies to train much more accurate model for energy estimation when using a sparse field data. We transferred knowledge, in the form of data and parameters, from the simulation framework to the field data. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that transfer learning-based models trained over one month data can perform comparative (and in some cases better) than the state-of-the-art machine learning and deep learning solutions.

Jain, Milan↗

Distilling particle knowledge for fast reconstruction at high-energy physics experiments

Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Analytical Modeling of Exoplanet Transit Spectroscopy with Dimensional Analysis and Symbolic Regression

Abstract The physical characteristics and atmospheric chemical composition of newly discovered exoplanets are often inferred from their transit spectra, which are obtained from complex numerical models of radiative transfer. Alternatively, simple analytical expressions provide insightful physical intuition into the relevant atmospheric processes. The deep-learning revolution has opened the door for deriving such analytical results directly with a computer algorithm fitting to the data. As a proof of concept, we successfully demonstrate the use of symbolic regression on synthetic data for the transit radii of generic hot-Jupiter exoplanets to derive a corresponding analytical formula. As a preprocessing step, we use dimensional analysis to identify the relevant dimensionless combinations of variables and reduce the number of independent inputs, which improves the performance of the symbolic regression. The dimensional analysis also allowed us to mathematically derive and properly parameterize the most general family of degeneracies among the input atmospheric parameters that affect the characterization of an exoplanet atmosphere through transit spectroscopy.

79 ASTRONOMY AND ASTROPHYSICS↗

Ground State Energy Functional with Hartree–Fock Efficiency and Chemical Accuracy

We introduce the deep post Hartree–Fock (DeePHF) method, a machine learning-based scheme for constructing accurate and transferable models for the ground-state energy of electronic structure problems. DeePHF predicts the energy difference between results of highly accurate models such as the coupled cluster method and low accuracy models such as the Hartree–Fock (HF) method, using the ground-state electronic orbitals as the input. It preserves all the symmetries of the original high accuracy model. The added computational cost is less than that of the reference HF or DFT and scales linearly with respect to system size. We examine the performance of DeePHF on organic molecular systems using publicly available data sets and obtain the state-of-art performance, particularly on large data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

New data-driven approach to bridging power system protection gaps with deep learning

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this paper, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A combined convolutional neural network and long short-term memory (CNN-LSTM) network is implemented to achieve data translation invariance and capture the temporal correlation of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN-LSTM--based method avoids the complicated, manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on two kinds of protection gaps: high-impedance faults and transformer inter-turn faults. Lastly, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

42 ENGINEERING↗

Improving radiograph analysis throughput through transfer learning and object detection

SIGN Fracture Care International partners with surgeons in low-resource hospitals worldwide to provide access to effective orthopedic care by donating educational materials and innovatively designed surgical implants. Over two decades, SIGN’s Online Surgical Database (SOSD) has grown to contain over 500,000 medical images, with radiographs holding the majority share. One challenge in working with hospitals worldwide is that both the radiographs uploaded to the SOSD and the data entry accompanying the uploads vary in quality. To improve the accuracy of data in the SOSD, we trained a model to detect surgical implants in radiographs. We first developed a tool to automatically detect radiographs, then trained an object detection model to determine the number and placement of surgical implants visible in the radiograph. Active learning was used to generate a training set containing 2,510 radiographs with screws, nails, and plates labeled by bounding boxes. Training a model to simultaneously recognize all three classes of implants gave a low average precision (AP) for the plate class, likely due to the low number of plate instances in our training set and the large variety of surgical plates used by SIGN-partnered surgeons. Applying standard image augmentation techniques to increase the plate count in our training set did not appreciably increase the AP of plate detection. To improve plate detection, we redrew the bounding boxes to account for correlations between the screw and plate classes. Training one model to detect nails and screws and a separate model to detect plates increased the AP of plate detection by 78.8 percentage points. The AP of each class was 80.7% for screws, 93.6% for nails, and 92.6% for plates; meanwhile, the sensitivity was 92% for screws, 86% for nails, and 81% for plates. We show that object detection methods can be used to detect surgical implants in radiographs of varying quality; however, the detection ability is dependent on the type of implant, and some implants, in our case plates, must be treated differently than others. Such tools can improve the throughput of radiograph analysis, assisting physicians and surgeons with the treatment of bone fractures.

60 APPLIED LIFE SCIENCES↗

Suppressing simulation bias in multi-modal data using transfer learning

Abstract Many problems in science and engineering require making predictions based on few observations. To build a robust predictive model, these sparse data may need to be augmented with simulated data, especially when the design space is multi-dimensional. Simulations, however, often suffer from an inherent bias. Estimation of this bias may be poorly constrained not only because of data sparsity, but also because traditional predictive models fit only one type of observed outputs, such as scalars or images, instead of all available output data modalities, which might have been acquired and simulated at great cost. To break this limitation and open up the path for multi-modal calibration, we propose to combine a novel, transfer learning technique for suppressing the bias with recent developments in deep learning, which allow building predictive models with multi-modal outputs. First, we train an initial neural network model on simulated data to learn important correlations between different output modalities and between simulation inputs and outputs. Then, the model is partially retrained, or transfer learned, to fit the experiments; a method that has never been implemented in this type of architecture. Using fewer than 10 inertial confinement fusion experiments for training, transfer learning systematically improves the simulation predictions while a simple output calibration, which we design as a baseline, makes the predictions worse. We also offer extensive cross-validation with real and carefully designed synthetic data. The method described in this paper can be applied to a wide range of problems that require transferring knowledge from simulations to the domain of experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A fine pore-preserved deep neural network for porosity analytics of a high burnup U-10Zr metallic fuel

Abstract U-10 wt.% Zr (U-10Zr) metallic fuel is the leading candidate for next-generation sodium-cooled fast reactors. Porosity is one of the most important factors that impacts the performance of U-10Zr metallic fuel. The pores generated by the fission gas accumulation can lead to changes in thermal conductivity, fuel swelling, Fuel-Cladding Chemical Interaction (FCCI) and Fuel-Cladding Mechanical Interaction (FCMI). Therefore, it is crucial to accurately segment and analyze porosity to understand the U-10Zr fuel system to design future fast reactors. To address the above issues, we introduce a workflow to process and analyze multi-source Scanning Electron Microscope (SEM) image data. Moreover, an encoder-decoder-based, deep fully convolutional network is proposed to segment pores accurately by integrating the residual unit and the densely-connected units. Two SEM 250 × field of view image datasets with different formats are utilized to evaluate the new proposed model’s performance. Sufficient comparison results demonstrate that our method quantitatively outperforms two popular deep fully convolutional networks. Furthermore, we conducted experiments on the third SEM 2500 × field of view image dataset, and the transfer learning results show the potential capability to transfer the knowledge from low-magnification images to high-magnification images. Finally, we use a pre-trained network to predict the pores of SEM images in the whole cross-sectional image and obtain quantitative porosity analysis. Our findings will guide the SEM microscopy data collection efficiently, provide a mechanistic understanding of the U-10Zr fuel system and bridge the gap between advanced characterization to fuel system design.

36 MATERIALS SCIENCE↗

Hybrid deep learning architecture for general disruption prediction across tokamaks

In this paper, we present a new deep learning disruption prediction algorithm based on important findings from explorative data analysis which effectively allows knowledge transfer from existing devices to new ones, thereby predicting disruptions using very limited disruptive data from the new devices. Here, the explorative data analysis conducted via unsupervised clustering techniques confirms that time-sequence data are much better separators of disruptive and non-disruptive behavior than the instantaneous plasma state data with further advantageous implications for a sequence-based predictor. Based on such important findings, we have designed a new algorithm for multi-machine disruption prediction that achieves high predictive accuracy on the C-Mod (AUC=0.801), DIII-D (AUC=0.947) and EAST (AUC=0.973) tokamaks with limited hyperparameter tuning. Through numerical experiments, we show that boosted accuracy (AUC=0.959) is achieved on EAST predictions by including in the training only 20 disruptive discharges, thousands of non-disruptive discharges from EAST, and combining this with more than a thousand discharges from DIII-D and C-Mod. The improvement of predictive ability obtained by combining disruptive data from other devices is found to be true for all permutations of the three devices. Furthermore, by comparing the predictive performance of each individual numerical experiment, we find that non-disruptive data are machine-specific while disruptive data from multiple devices contain device-independent knowledge that can be used to inform predictions for disruptions occurring on a new device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Transfer Learning Trained LSTM Models for Household Load Profile Forecasting

Grid edge renewable energy resources, such as rooftop solar photovoltaics, closely interact with consumer load profiles. Therefore, forecasting future electricity demand, ideally at the individual household level, is indispensable. In this paper, we present a transfer learning enhanced household load profile forecasting method. First, we tune a long short-term memory forecasting model to perform day-ahead prediction of household electricity load profiles. Then we improve these individualized models using transfer learning, and we use k-means clustering to create optimal source data sets. We find average improvements of 4.38% (largest improvement of 10.71%) when the entire data set was used to train the source model and 2.45% (largest improvement of 11.57%) in the mean absolute error when households were first clustered and used to train separate source models for each cluster. We find that transfer learning with clustered data can effectively boost the forecasting performance of the LSTM models. We use realistic household power measurements for 148 real residential households in Austin, Texas.

deep learning↗

Multiscale Data-Driven Seismic Full-Waveform Inversion With Field Data Study

Seismic full-waveform inversion (FWI), which uses iterative methods to estimate high-resolution subsurface models from seismograms, is a powerful imaging technique in exploration geophysics. In recent years, the computational cost of FWI has grown exponentially due to the increasing size and resolution of seismic data. Moreover, it is a nonconvex problem and can encounter local minima due to the limited accuracy of the initial velocity models or the absence of low frequencies in the measurements. To overcome these computational issues, we develop a multiscale data-driven FWI method based on fully convolutional networks (FCNs). In preparing the training data, we first develop a real-time style transform method to create a large set of synthetic subsurface velocity models from natural images. We then develop two convolutional neural networks with encoder-decoder structures to reconstruct the low- and high-frequency components of the subsurface velocity models, separately. To validate the performance of our data-driven inversion method and the effectiveness of the synthesized training set, we compare it with conventional physics-based waveform inversion approaches using both synthetic and field data. Finally, these numerical results demonstrate that, once our model is fully trained, it can significantly reduce the computation time and yield more accurate subsurface velocity models in comparison with conventional FWI.

58 GEOSCIENCES↗

Hybrid deep learning architecture for general disruption prediction across tokamaks

In this paper, we present a new deep learning disruption prediction algorithm based on important findings from explorative data analysis which effectively allows knowledge transfer from existing devices to new ones, thereby predicting disruptions using very limited disruptive data from the new devices. The explorative data analysis conducted via unsupervised clustering techniques confirms that time-sequence data are much better separators of disruptive and non-disruptive behavior than the instantaneous plasma state data with further advantageous implications for a sequence-based predictor. Based on such important findings, we have designed a new algorithm for multi-machine disruption prediction that achieves high predictive accuracy on the C-Mod (AUC=0.801), DIII-D (AUC=0.947) and EAST (AUC=0.973). tokamaks with limited hyperparameter tuning. Through numerical experiments, we show that boosted accuracy (AUC=0.959) is achieved on EAST predictions by including in the training only 20 disruptive discharges, thousands of non-disruptive discharges from EAST, and combining this with more than a thousand discharges from DIII-D and C-Mod. The improvement of predictive ability obtained by combining disruptive data from other devices is found to be true for all permutations of the three devices. Furthermore, by comparing the predictive performance of each individual numerical experiment, we find that non-disruptive data are machine-specific while disruptive data from multiple devices contain device-independent knowledge that can be used to inform predictions for disruptions occurring on a new device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Physics-informed neural network with transfer learning (TL-PINN) based on domain similarity measure for prediction of nuclear reactor transients

Nuclear reactor safety and efficiency can be enhanced through the development of accurate and fast methods for prediction of reactor transient (RT) states. Physics informed neural networks (PINNs) leverage deep learning methods to provide an alternative approach to RT modeling. Applications of PINNs in monitoring of RTs for operator support requires near real-time model performance. However, as with all machine learning models, development of a PINN involves time-consuming model training. Here, we show that a transfer learning (TL-PINN) approach achieves significant performance gain, as measured by reduction of the number of iterations for model training. Using point kinetic equations (PKEs) model with six neutron precursor groups, constructed with experimental parameters of the Purdue University Reactor One (PUR-1) research reactor, we generated different RTs with experimentally relevant range of variables. The RTs were characterized using Hausdorff and Fréchet distance. We have demonstrated that pre-training TL-PINN on one RT results in up to two orders of magnitude acceleration in prediction of a different RT. The mean error for conventional PINN and TL-PINN models prediction of neutron densities is smaller than 1%. We have developed a correlation between TL-PINN performance acceleration and similarity measure of RTs, which can be used as a guide for application of TL-PINNs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Identifying common stored product insects using automated deep learning methods

Monitoring stored product insect pests is a common practice for post-harvest management of stored grain and grain-based commodities, which helps ensure product quality from harvest to final consumer. Current methods of sampling and monitoring can be time-consuming, labor-intensive, expensive and require expertise in insect identification. Therefore, this study aims to develop an image-based automated identification system for common stored product insect species using deep-learning methods. Top-down images of the common stored product adult insect species of Rhyzopertha dominica, Cryptolestes ferrugineus, Tribolium castaneum, Sitophilus oryzae, and Oryzaephilus surinamensis were acquired and analyzed. Deep learning-based, state-of-the-art Convolutional Neural Networks (CNN) models (ResNet-50, MobileNet-v2, DarkNet-53, and EfficientNet-b0) were fine-tuned with a transfer learning approach to classify the insect species. All models were able to correctly identify the insect species with at least 96% accuracy and with few misclassifications. One issue with trained CNNs is that they do not explain the reasoning for the classification and are often called a “black box”. Therefore, visualization methods called Gradient-weighted Class Activation Mapping (Grad-CAM) were implemented to explore the black box network. The Grad-CAM uses heat maps to highlight the major image features that the network focused on to make insect species predictions. The Grad-CAM verifies the network's prediction and also helps improve network performance. This study contributes to the overall goal of developing a camera-based system for monitoring stored grain insects. As a result, the developed system would empower warehouse, flour mills, and other food facilities with a tool to quickly and accurately identify insect species in stored product environments and could be implemented as part of a close to real-time monitoring system.

60 APPLIED LIFE SCIENCES↗