Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “convolutional neural network model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Fault Detection Utilizing Convolution Neural Network on Timeseries Synchrophasor Data From Phasor Measurement Units

An end-to-end supervised learning method is proposed for fault detection in the electric grid using Big Data from multiple Phasor Measurement Units (PMUs). The approach consists of preprocessing steps aimed at reducing data noise and dimensionality, followed by utilization of six classification models considered for detecting faults. Three of the models were variants of Convolutional Neural Network (CNN) architectures that consider a single type of measurement (voltage, current or frequency) at all PMUs or all types together also at all PMUs. CNN based models were compared to traditional methods of Logistic Regression (LR), Multi-layer Perceptron (MLP) and Support Vector Machine (SVM). Evaluation was conducted on two-year data measured by PMUs at 37 locations in a large electric grid. Here, the response variable for classification were extracted from the grid-wide outage event log. Experiments show that CNN-based models outperformed traditional methods on one year out-of-sample outage detection over the entire grid.

42 ENGINEERING↗

Autonomous fabrication of tailored defect structures in 2D materials using machine learning-enabled scanning transmission electron microscopy

Materials with tailored quantum properties can be engineered from atomic-scale assembly techniques, but existing methods often lack the agility and accuracy to precisely and intelligently control the manufacturing process. Here, we demonstrate a fully autonomous approach for fabricating atomic-level defects using electron beams in scanning transmission electron microscopy (STEM) that combines advanced machine learning and automated beam control. As a proof of concept, we achieved controlled fabrication of MoS-nanowire (MoS-NW) edge structures by iterative and targeted exposure of MoS 2 monolayer to a focused electron beam to selectively eject sulfur atoms, utilizing high-angle annular dark-field (HAADF) imaging for feedback-controlled monitoring of structural evolution of defects. A machine learning framework combining a random forest model and a convolutional neural network (CNN) was developed to decode the HAADF image and accurately identify atomic positions and species. This atomic-level information was then integrated into an autonomous decision-making platform, which applied predefined fabrication strategies to instruct beam control about atomic sites to be ejected. The selected sites were subsequently exposed to a localized electron beam using an FPGA-controlled scan routine with precise control over beam positioning and duration. While the MoS-NW edge structures produced exhibit promising mechanical and electronic properties, the proposed methods to build the autonomous fabrication framework is material-agnostic and can be extended to other 2D materials for the creation of diverse defect structures and heterostructures beyond Mo S2 .

Engineering↗

Convolutional Neural Networks Based Remote Sensing Scene Classification under Clear and Cloudy Environments

Remote sensing (RS) scene classification has wide applications in the environmental monitoring and geological survey. In the real-world applications, the RS scene images taken by the satellite might have two scenarios: clear and cloudy environments. However, most of existing methods did not consider these two environments simultaneously. In this paper, we assume that the global and local features are iscriminative in either clear or cloudy environments. Many existing Convolution Neural Networks (CNN) based models have made excellent achievements in the image classification, however they somewhat ignored the global and local features in their network structure. In this paper, we propose a new CNN based network (named GLNet) with the Global Encoder and Local Encoder to extract the discriminative global and local features for the RS scene classification, where the constraints for inter-class dispersion and intra-class compactness are embedded in the GLNet training. The experimental results on two publicized RS scene classification datasets show that the proposed GLNet could achieve better performance based on many existing CNN backbones under both clear and cloudy environments.

97 MATHEMATICS AND COMPUTING↗

Improve the Search of Very Metal-poor Stars Using the Deep Learning Method

Very metal-poor (VMP) stars have [Fe/H] < -2.0 dex. They are among the oldest stars in the universe, and their unique metallicity can help explore the enrichment mechanism and evolutionary history of the chemical elements of stars in the early universe. However, most current stellar parameter estimation methods do not perform well in determining the stellar parameters of VMP stars, which limits our ability to discover and exploit the properties of VMP stars. In this study, we propose a new model based on a convolutional neural network to determine the stellar atmospheric parameters of VMP stars. We tested our model on the LAMOST spectra; our model can determine the effective temperature (T {sub eff}), surface gravity (log g), metallicity ([Fe/H]), and carbon abundance ([C/Fe]) of the LAMOST spectra with precisions σ(T {sub eff}) = 134.82 K, σ(log g) = 0.33 dex, σ ([Fe/H]) = 0.20 dex, and σ([C/Fe]) = 0.35 dex. Furthermore, our model can distinguish VMP stars from normal stars with an accuracy of 88.65%. We also compared this model with other widely used methods, and found that this method performs better than other methods. It can be applied to the stellar parameter pipelines of upcoming large surveys such as 4MOST, WEAVES, and MOONS to search for VMP stars and identify carbon-enhanced metal-poor stars.

36 MATERIALS SCIENCE↗

An algorithm for physics informed scan path optimization in additive manufacturing

Site specific microstructure control is a critical research area within the field of additive manufacturing due to its potential to revolutionize part performance. One way to achieve site specific microstructure control is through control of the solidification conditions via the construction of intricate scan paths; however, the search space for such a problem is large. Previous attempts only considered the solidification conditions at the top surface while also requiring either lots of manual-fine tuning or large amounts of computational resources. This paper introduces a general method for scan path optimization which considers the solidification conditions in the bulk of the material without an increase in computational expense. This method consists of three core components:1. A heat transfer model for simulating the temperature field at a given time.2. A surrogate model which takes scan pattern information and temperature data and predicts the solidification conditions of the bulk as well as the meltpool depths for a spot melt.3. A decision algorithm to decide which spot melt should be printed next based on the outputs of the surrogate model.Each of these components can be changed without changing the overall method. Within this work, this method is applied in the creation of an algorithm containing a semi-analytic heat transfer model to simulate the temperature field, a fully convolutional neural network (FCNN) as the surrogate model, and a greedy decision algorithm. The resulting algorithm produced complex scan patterns which gave strong results for simulated microstructure control.

36 MATERIALS SCIENCE↗

Optimal vocabulary selection approaches for privacy-preserving deep NLP model training for information extraction and cancer epidemiology

With the use of artificial intelligence and machine learning techniques for biomedical informatics, security and privacy concerns over the data and subject identities have also become an important issue and essential research topic. Without intentional safeguards, machine learning models may find patterns and features to improve task performance that are associated with private personal information. The privacy vulnerability of deep learning models for information extraction from medical textural contents needs to be quantified since the models are exposed to private health information and personally identifiable information. The objective of the study is to quantify the privacy vulnerability of the deep learning models for natural language processing and explore a proper way of securing patients’ information to mitigate confidentiality breaches. The target model is the multitask convolutional neural network for information extraction from cancer pathology reports, where the data for training the model are from multiple state population-based cancer registries. This study proposes the following schemes to collect vocabularies from the cancer pathology reports; (a) words appearing in multiple registries, and (b) words that have higher mutual information. We performed membership inference attacks on the models in high-performance computing environments. The comparison outcomes suggest that the proposed vocabulary selection methods resulted in lower privacy vulnerability while maintaining the same level of clinical task performance.

59 BASIC BIOLOGICAL SCIENCES↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗

CNN-Encoder-Decoder Model

Code and data for training CNN-Encoder-Decoder model described in the publication 'Noise reduction in X-ray photon correlation spectroscopy with convolutional neural networks encoder-decoder models.'

Konstantinova, Tatiana [Brookhaven National Lab. (↗

Laser Wakefield Accelerator modelling with Variational Neural Networks

A machine learning model was created to predict the electron spectrum generated by a GeV-class laser wakefield accelerator. The model was constructed from variational convolutional neural networks, which mapped the results of secondary laser and plasma diagnostics to the generated electron spectrum. An ensemble of trained networks was used to predict the electron spectrum and to provide an estimation of the uncertainty of that prediction. It is anticipated that this approach will be useful for inferring the electron spectrum prior to undergoing any process that can alter or destroy the beam. In addition, the model provides insight into the scaling of electron beam properties due to stochastic fluctuations in the laser energy and plasma electron density.

43 PARTICLE ACCELERATORS↗

Improving the Transportability of a Deep Learning Denoising Model Using Transfer Learning Techniques

The adoption of machine learning techniques in the seismology community has led to great performance improvements in several areas, including signal processing. Specifically, the development of deep learning–based seismic waveform denoising models has the potential to yield improvements in signal detection capabilities for networks operating in particularly noisy environments. Recent advancements in the design of these deep learning denoising models have included the incorporation of continuous and discrete wavelet transform functions into the network architecture to improve the learning capabilities and efficiency of said models. These wavelet transform–based seismic denoising models have shown improved denoising capabilities in regions where there is good agreement between the data features present in the training and evaluation datasets. However, questions remain about the overall transportability of these models to other monitoring regions. Here, in this study, we will determine the baseline transportability of a newly developed multilevel wavelet‐transform convolutional neural network (MWCNN) seismic denoising model. We accomplish this by taking a version of the MWCNN denoising model trained on data collected from the Utah region and evaluating its denoising performance on datasets collected from the neighboring Nevada region, which differ with regard to monitoring sensor types and event histories. We find that there is a notable variability in denoising performance related to the degree of similarity between the initial and new target datasets. The most notable difference in denoising performance is the ability of the denoising model to preserve accurate amplitude information associated with the signal energy present in the waveform data. Finally, we evaluate the ability of transfer learning techniques to improve the transportability of the MWCNN denoising model. We find that although there is still a performance gap present in the denoising results of the MWCNN model, transfer learning did yield improved results.

Quinones, Louis [Sandia National Laboratories (SNL↗

Identifying Critical Infrastructure in Imagery Data Using Explainable Convolutional Neural Networks

To date, no method utilizing satellite imagery exists for detailing the locations and functions of critical infrastructure across the United States, making response to natural disasters and other events challenging due to complex infrastructural interdependencies. This paper presents a repeatable, transferable, and explainable method for critical infrastructure analysis and implementation of a robust model for critical infrastructure detection in satellite imagery. This model consists of a DenseNet-161 convolutional neural network, pretrained with the ImageNet database. The model was provided additional training with a custom dataset, containing nine infrastructure classes. The resultant analysis achieved an overall accuracy of 90%, with the highest accuracy for airports (97%), hydroelectric dams (96%), solar farms (94%), substations (91%), potable water tanks (93%), and hospitals (93%). Critical infrastructure types with relatively low accuracy are likely influenced by data commonality between similar infrastructure components for petroleum terminals (86%), water treatment plants (78%), and natural gas generation (78%). Local interpretable model-agnostic explanations (LIME) was integrated into the overall modeling pipeline to establish trust for users in critical infrastructure applications. The results demonstrate the effectiveness of a convolutional neural network approach for critical infrastructure identification, with higher than 90% accuracy in identifying six of the critical infrastructure facility types.

97 MATHEMATICS AND COMPUTING↗

Acoustic-based monitoring and machine learning of component status for microreactor applications

This report provides a description and assessment of recent efforts to couple acoustic-based experimental measurements and characterization with machine learning models in order to enhance structural health monitoring capabilities for nuclear microreactors. With resilient embedded sensors in development by others supported by programs funded by the US Department of Energy’s Office of Nuclear Energy, the work described herein builds upon ongoing efforts to improve non-destructive testing technology that relates measured acoustic signatures to component stresses and/or structural defects, using a combination of new experimental measurements and machine learning architectures. The experimental procedure remained similar to that developed for the previous year’s demonstration of damage detection by the authors, with the same damaged sample tested under similar applied stress conditions. Notably, a new mounting fixture was designed and implemented to improve measurement consistency and a more sophisticated laser Doppler vibrometer was employed to make high-fidelity vibration measurements. Two nominally identical sets of training data were collected for each experimental setup to better understand the repeatability of the experiment and to better test the generality of trained neural network models. Additionally, we obtained new high-quality 3D mode shapes of the damaged test article at various stress and excitation levels, providing greater insights into the physical response of the sample during testing. Previously, we demonstrated that a machine learning model based on a convolutional neural network can predict structural details of an artificially introduced interface (intact, rough cut, smooth cut), and the applied torque level. In this study, we have transitioned to graph-based neural network architectures to better develop and test a flexible framework that is more suitable to being transferred away from controlled benchtop experiments and into more applied settings where less-structured data inputs may be expected. In general, performance testing of a graph neural network on frequency-domain representations of the data indicates strong and consistent identification of test conditions for datasets recorded on damaged components. With goals of predicting damage location and other changing experimental conditions using limited datasets, predictive models using a graph neural network architecture correctly predicted the applied torque level with an accuracy of 85% using only a single measurement point and predicted within one torque level in 95% of test windows. Predictions of damage location had limited success due to the symmetry and minimal number of the damage scenarios presented during model training. Results were ambiguous as to whether the model could detect the location of the artificial damage, or if it was instead learning the location of a given measurement point on the part and subsequently detecting which points were closest to the location of the damage. This finding will be factored into upcoming planned work on damaged graphite components, where new experimental tests with a larger number and variety of damage scenarios are expected to provide improved validation of recent developments in monitoring methodology.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Observations and Machine-Learned Models of Near-Surface Permafrost along the Koyukuk River, Alaska, USA

This dataset contains GeoTIFs (raster) and GeoPackages (vector) that map observations of near-surface permafrost and not-permafrost from a field campaign conducted near the village of Huslia, AK along the Koyukuk River and its floodplain in July 2018. These data were collected as part of a campaign to understand if and how permafrost impacts riverbank erosion. This problem cannot be assessed without knowing where permafrost exists. Permafrost was observed via frost probing (to a maximum depth of one meter), coring (to a maximum depth of two meters) and bank/bar excavations. An additional boat survey was performed wherein expert (Joel Rowland) judgment assessed the presence or absence of distinctive permafrost features (e.g., overhanging tundra mats, thermoerosional niching, ice wedges, active drainage of ice melt from soils). This dataset also contains the input features and results of two machine learning models (random forest and convolutional neural network) that extrapolate the observations to the full floodplain that may be useful for building, testing, or validating other machine-learned permafrost models. Permafrost data are provided as georasters of the same shape and geovectors (polylines/polygons) and are all projected into EPSG:32605. All data can be visualized with a GIS (QGIS, ArcGIS, etc.).

54 ENVIRONMENTAL SCIENCES↗

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting bulge to total luminosity ratio of galaxies using deep learning

ABSTRACT We present a deep learning model to predict the r-band bulge-to-total luminosity ratio (B/T) of nearby galaxies using their multiband JPEG images alone. Our Convolutional Neural Network (CNN) based regression model is trained on a large sample of galaxies with reliable decomposition into the bulge and disc components. The existing approaches to estimate the B/T ratio use galaxy light-profile modelling to find the best fit. This method is computationally expensive, prohibitively so for large samples of galaxies, and requires a significant amount of human intervention. Machine learning models have the potential to overcome these shortcomings. In our CNN model, for a test set of 20 000 galaxies, 85.7 per cent of the predicted B/T values have absolute error (AE) less than 0.1. We see further improvement to 87.5 per cent if, while testing, we only consider brighter galaxies (with r-band apparent magnitude <17) with no bright neighbours. Our model estimates the B/T ratio for the 20 000 test galaxies in less than a minute. This is a significant improvement in inference time from the conventional fitting pipelines, which manage around 2–3 estimates per minute. Thus, the proposed machine learning approach could potentially save a tremendous amount of time, effort, and computational resources while predicting B/T reliably, particularly in the era of next-generation sky surveys such as the Legacy Survey of Space and Time (LSST) and the Euclid sky survey which will produce extremely large samples of galaxies.

Grover, Harsh (ORCID:0000000321338142)↗

Reduced-order modeling of advection-dominated systems with recurrent neural networks and convolutional autoencoders

A common strategy for the dimensionality reduction of nonlinear partial differential equations (PDEs) relies on the use of the proper orthogonal decomposition (POD) to identify a reduced subspace and the Galerkin projection for evolving dynamics in this reduced space. However, advection-dominated PDEs are represented poorly by this methodology since the process of truncation discards important interactions between higher-order modes during time evolution. In this study, we demonstrate that encoding using convolutional autoencoders (CAEs) followed by a reduced-space time evolution by recurrent neural networks overcomes this limitation effectively. We demonstrate that a truncated system of only two latent space dimensions can reproduce a sharp advecting shock profile for the viscous Burgers equation with very low viscosities, and a six-dimensional latent space can recreate the evolution of the inviscid shallow water equations. Additionally, the proposed framework is extended to a parametric reduced-order model by directly embedding parametric information into the latent space to detect trends in system evolution. Furthermore, our results show that these advection-dominated systems are more amenable to low-dimensional encoding and time evolution by a CAE and recurrent neural network combination than the POD-Galerkin technique.

97 MATHEMATICS AND COMPUTING↗

Understanding Twinning and Deformation in High Entropy Alloys

A combination of high strength and high ductility has been observed in multi-principal element alloys due to twin formation attributed to low stacking fault energy (SFE). In the pursuit of low SFE alloys, a key bottleneck is the lack of understanding of the composition–SFE cor- relations that would guide tailoring SFE via alloy composition. Using density functional theory (DFT), we show that dopant radius, which have been postulated as a key descriptor for SFE in dilute alloys, does not fully explain SFE trends across different host metals. Instead, charge density is a much more central descriptor. It allows us to (1) explain contrasting SFE trends in Ni and Cu host metals due to various dopants in dilute concentrations, (2) explain the large SFE variations observed in the literature even within a given alloy composition due to the nearest neighbor environments in “model” concentrated alloys, and (3) develop a machine learning model that can be used to predict SFEs in multi-elemental alloys. This model opens a possibility to use charge density as a descriptor for predicting SFE in alloys. Furthermore, a descriptor-less machine learning (ML) model based only on charge density images extracted from density functional theory (DFT) is developed to predict stacking fault energies (SFE) in concentrated alloys. The model is based on convolutional neural networks (CNNs) as one of the promising ML techniques for dealing with complex images and data. Identification of correct descriptors is a key bottleneck to develop ML models for predicting materials properties. Often, in most ML models, textbook physical descriptors such as atomic radius, valence charge and electronegativity are used as descriptors which have limitations because these properties change in concentrated alloys when multiple elements are mixed to form a solid solution. We illustrate that, within the scope of DFT, the search for descriptors can be circumvented by electronic charge density, which is the backbone of the Kohn-Sham DFT and describes the system completely. The performance of our model is demonstrated by predicting SFE of concentrated alloys with an RMSE and R2 of 6.18 mJ/m2 and 0.87, respectively, validating the accuracy of the proposed approach.

36 MATERIALS SCIENCE↗