Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

Dynamic Ride-Matching for Large-Scale Transportation Systems

Efficient dynamic ride-matching (DRM) in large-scale transportation systems is a key driver in transport simulations to yield answers to challenging problems. Although the DRM problem is simple to solve, it quickly becomes a computationally challenging problem in large-scale transportation system simulations. Therefore, this study thoroughly examines the DRM problem dynamics and proposes an optimization-based solution framework to solve the problem efficiently. To benefit from parallel computing and reduce computational times, the problem’s network is divided into clusters utilizing a commonly used unsupervised machine learning algorithm along with a linear programming model. Then, these sub-problems are solved using another linear program to finalize the ride-matching. At the clustering level, the framework allows users adjusting cluster sizes to balance the trade-off between the computational time savings and the solution quality deviation. A case study in the Chicago Metropolitan Area, U.S., illustrates that the framework can reduce the average computational time by 58% at the cost of increasing the average pick up time by 26% compared with a system optimum, that is, non-clustered, approach. Another case study in a relatively small city, Bloomington, Illinois, U.S., shows that the framework provides quite similar results to the system-optimum approach in approximately 62% less computational time.

33 ADVANCED PROPULSION SYSTEMS↗

Optimal dimensionality selection for independent component analysis of transcriptomic data

Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Anomaly Detection and Identification Using a Leave-One-Variable-Out Method

At nuclear power plants (NPPs), anomaly detection and identification (i.e., determining the causes of anomalies) are important tasks for ensuring the safe and efficient operation of NPPs. These tasks are currently labor-intensive and costly, and are made more difficult by the size and complexity of NPP systems. An alternative approach to conducting these tasks is to automate them, such as via the reconstruction-based contribution method, which is a well-researched unsupervised machine learning method that uses a data-driven model of anomaly-free behavior to detect events and then identify each variable’s contributions to those events. The present effort developed a novel contribution approach that utilized a leave-one-variable-out (LOVO) model, with which each variable is predicted using all the other variables. The novelty lay in transforming this model into a reconstruction model and modifying the identification algorithm to work with the new reconstruction model. To evaluate this method in a controlled environment, a synthetic dataset based on spring-mass-damper (SMD) systems (commonly found in mechanical engineering references) was used, with known anomalies introduced into the system. The proposed method successfully detected the anomalies and afforded insights into their causes, thus enabling the appropriate identifications to be made.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Performance of Compact Pulsed Thermal Imaging System for In-Service Applications. Pulsed thermal tomography nondestructive examination of additively manufactured reactor materials and components

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of complex topology nuclear reactor parts from high-strength corrosion resistance alloys, such as stainless steel and Inconel. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which has the capability of melting metallic powder and net shaping the structures with relatively high precision. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures due to creep in high temperature nuclear reactor environment. Currently, there exist limited capabilities to evaluate actual AM structures nondestructively. Pulsed Thermography (PT) imaging provides a capability for non-destructive evaluation (NDE) of sub-surface defects in arbitrary size structures. The PT method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PT method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. The data cube of PT measurements consists of surface temperature taken at sequential time intervals T(x,y,t). Material defects can be detected either by analyzing the thermograms T(x,y,t) data cube, or by using thermal tomography (TT) algorithm to obtain 3D spatial reconstruction of thermal effusivity e(x,y,z). To reduce the cost and enable in-service NDE in spatially constrained environment, it is highly desirable to develop PT with compact and inexpensive IR camera. Following initial qualification of an AM component for deployment in a nuclear reactor, a compact PT system can also be used for in-service nondestructive evaluation (NDE) applications. However, data cube obtained with PT based on compact IR camera suffers from strong thermal noises and loss of features due to relatively low sampling rate. In this report we describe two unsupervised machine learning (ML) algorithms for enhancement of PT images obtained with compact IR camera. In one approach, we introduce Sparse Coding Discrete Cosine Transform (SC/DCT) algorithm to remove additive white Gaussian noise (AWGN) from spatial thermal effusivity reconstructions. In another approach we introduce a Spatial Temporal Denoised Thermal Source Separation (STDTSS) ML algorithm to process thermograms. The STDTSS algorithm consists of spatial and temporal denoising using Gaussian and Savitzky–Golay filtering, followed by the matrix decomposition using Principal Component Analysis (PCA), and Independent Component Analysis (ICA) to automatically detect flaws. In the work described in this report, we constructed a compact PT system using a relatively small and low-cost FLIR A65 camera, consisting on uncooled microbolometer detector. Performance of SC/DCT algorithm was demonstrated on enhancing TT images of Inconel 718 AM plate. Performance of the STDTSS methods was investigated using thermography data obtained from imaging stainless steel 316L specimens produced with LPBF method with imprinted calibrated porosity defects.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Performance Validation of Pulsed Thermal Imaging System for In-Service Applications

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of complex topology nuclear reactor parts from high-strength corrosion resistance alloys, such as stainless steel and Inconel. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which has the capability of melting metallic powder and net shaping the structures with relatively high precision. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures in high temperature nuclear reactor environment. Currently, there exist limited capabilities to evaluate actual AM structures non-destructively. Pulsed Thermography Imaging (PTI) provides a capability for non-destructive evaluation (NDE) of subsurface defects in arbitrary size structures. The PTI method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PTI method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. Following initial qualification of an AM component for deployment in a nuclear reactor, a PTI system can also be used for in-service nondestructive evaluation (NDE) applications. In this report, we describe recent progress in enhancing PTI capabilities in detecting microscopic defects in metallic specimens. SS316 and IN718 specimens were developed with a pattern of subsurface calibrated flat bottom hole (FBH) defects with diameters from 500µm to 200µm. FBH’s were created with EDM (electron discharge machining) drill. PTI imaging data was processed Spatial Temporal Denoised Thermal Source Separation (STDTSS) unsupervised machine learning (ML) algorithm. We show that defects as small as 200µm in SS316 and IN718 can be detected with STDTSS algorithm. To the best of our knowledge, these are the smallest detected defects which are reported in literature.

42 ENGINEERING↗

Pulsed Thermal Tomography Nondestructive Examination of Additively Manufactured Reactor Materials and Components. Third Annual Progress Report

Additive manufacturing (AM) of high-strength corrosion resistance alloys for nuclear energy applications, such as stainless steel and Inconel, is currently based on laser powder bed fusion (LPBF) process. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Probability of crack initiation at the pore depends on size, shape, and orientation of the defect. Pulsed Infrared Thermography Imaging (PIT) provides a capability for non-destructive evaluation (NDE) of sub-surface defects in arbitrary size structures. The PIT method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PIT method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. Following initial qualification of an AM component for deployment in a nuclear reactor, a PIT system can also be used for in-service nondestructive evaluation (NDE) applications. In this report, we describe recent progress in enhancing PIT capabilities in detecting microscopic subsurface defects in metals, and classifying shapes and orientation of pores in thermal images. For detection of microscopic defects in PIT imaging data, we have developed Spatial Temporal Denoised Thermal Source Separation (STDTSS) unsupervised machine learning (ML) image processing algorithm. We show that flat bottom hole (FBH) defects as small as 200µm in SS316 and IN718 specimens, can be detected with STDTSS algorithm. To the best of our knowledge, these are the smallest detected defects which are reported in literature. For classification of defects shapes, we have previously developed thermal tomography (TT) algorithm to obtain depth reconstructions of material defects from data cube of sequentially recorded surface temperatures. However, interpretation of TT images is non-trivial because of blurring with increasing depth. To address this challenge, we have developed a deep learning convolutional neural network (CNN) to classify size and orientation subsurface defects in simulated TT images.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Detection of Anomalies in Gamma Background Radiation Data with K-Means and Self-Organizing Map Clustering Algorithms (Consortium on Nuclear Security Technologies (CONNECT) Q1 Report)

Environmental screening of gamma radiation consists of detecting weak nuisance and anomaly signal in the presence of strong and highly varying background. In a typical scenario, a mobile detector-spectrometer continuously measures gamma radiation spectra in short, e.g., one-second, signal acquisition intervals. The measurement data is a 2D matrix, where one dimension is gamma ray energy, and the other dimension is the number of measurements or total time. In principle, gamma radiation sources can be detected and identified from the measured data by their unique spectral lines. Detecting sources from data measured in a search scenario is difficult due to the highly varying background because of naturally occurring radioactive material (NORM), and low signal-to-noise ratio (S/N) of spectral signal measured during one-second acquisition intervals. The objective of this work is to explore unsupervised machine learning (ML) algorithms for detection and identification of weak nuisances and anomalies events in the presence of highly fluctuating background. The challenge is that spectral lines of isotopes are difficult to observe in one-second measurements. Averaging over the entire measurement campaign data set reveals spectral lines of most common background isotopes. Spectral lines of orphan sources, which might appear only in a few measurements during the campaign, will be washed out if averaging is performed over the entire measurement data set. The approach we have explored consists of extracting one-second measurements containing weak spectral features through data clustering. Averaging one-second spectra in a cluster should reveal the presence of anomaly sources. We created two ML models using K-means clustering and Neural Network Self-organizing Map (SOM). Performance of these ML models was benchmarked using search data. One data set contained 137 Cs source, and another dataset contained 131 I source.

61 RADIATION PROTECTION AND DOSIMETRY↗

Exploring New Ways to Classify Industries for Energy Analysis and Modeling

As the US moves closer to embracing a net zero greenhouse gas emissions position, combustion processes outside the power sector are becoming urgent concerns. Industry is an important end user of energy and relies on fossil fuels used directly for process heating and as feedstocks for a diverse range of applications. Fuel and energy use by industry is heterogeneous, meaning that even a single product group can vary broadly in its production routes and associated energy usage. In the US, the North American Industry Classification System (NAICS) serves as the basis for data collection and reporting. In turn, data based on NAICS is the foundation of most US energy modeling. Thus, the effectiveness of NAICS at representing energy use is a limiting condition for plans to improve energy efficiency and alternatives to fossil fuels in industry. Facility-level data to build more detail into heterogeneous sectors is scarce. This work explores alternative classification schemes for industry based on energy use characteristics, and provides a validation of an approach to make facility-level energy use estimates based on publicly available data from the greenhouse gas reporting program. First, several approaches to industrial taxonomies and their usefulness for industrial energy modeling are summarized. Data from Industrial Assessment Centers is analyzed using unsupervised machine learning techniques to detect clusters. Cladistics, an approach from biology, is adapted to energy and process characteristics of industries. A cladogram is presented for evolutionary directions in the iron and steel sector. Cladograms are a promising tool for constructing scenarios and summarizing directions of sectoral innovation. Finally, validation is performed for facility-level energy estimates from the US EPA Greenhouse Gas Reporting Program. This validation assists in making this data source available for use in energy modeling. Together, this work explores alternative approaches for categorizing industries in a way that aids understanding energy use, and presenting pathways for the future.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Behavioral Segmentation and Clustering of Geospatial Trajectories

The rapid growth of global positioning system (GPS) devices has led to a corresponding increase in the size of GPS datasets. While these large GPS datasets contain a wealth of information about the behaviors of the moving objects in them, manual classification and anomaly detection are prohibitively time consuming. We utilize unsupervised machine learning techniques to first identify the behaviors for individual moving objects and then cluster those objects by their behavioral sequences. In this way, trajectories behaving unusually as well as common patterns of behavior are both detectable in large datasets without requiring an a priori definition of "unusual" or "common."

97 MATHEMATICS AND COMPUTING↗

Flat and Level Analysis Tool (FLAT) for real-time automated segmentation and analysis of concrete slab point clouds

In the United States, the flatness and levelness of concrete floors during construction is traditionally specified by a maximum allowable gap under a 3 meter straightedge. However, the straightedge method is inexact and rarely representative of the entire floor since the technician is free to choose any location on the floor to perform the measurement. In cases requiring a higher degree of precision and repeatability, concrete floor flatness and levelness can be measured using the standard test method ASTM E1155. With the recent introduction of advanced surveying instruments such as robotic theodolites and terrestrial laser scanners (TLS), the means now exist to modernize and expedite the measurement of floor flatness and levelness. This paper details the development and demonstration of a digital tool, named the Flat and Level Analysis Tool (FLAT), to automate and expedite the segmentation and analysis of flatness and levelness from dense point cloud data of concrete floor slabs. Segmentation algorithms were developed using unsupervised machine learning to extract the set of points belonging to the concrete floor slab from a full 360 scan of a construction site. After segmentation, automated analysis algorithms report the results according to the standard method. The developed algorithms were demonstrated on a dense point cloud captured from a concrete slab-on-grade at a construction site. Results show that the digital tool can quickly provide estimates for floor flatness and levelness with minimal human involvement with comparable accuracy to manual methods.

Hayes, Nolan↗

Seismic Characterization of the Blue Mountain Geothermal Field

Subsurface characterization is crucial for geothermal energy exploration and production. Yet hydrothermal reservoirs usually reside in highly fractured and faulted zones where accurate characterization is very challenging because of low signal-to-noise ratios of land seismic data and lack of coherent reflection signals. We perform an active-source seismic characterization for the Blue Mountain geothermal field in Nevada using active seismic data to reveal the elastic medium property complexity and fault distribution at this field. We first employ an unsupervised machine learning method to attenuate groundroll and near-surface guided-wave noise and enhance coherent reflection and scattering signals from noisy seismic data. We then build a smooth initial P-wave velocity model based on an existing magnetotellurics survey result, and use 3D first-arrival traveltime tomography to refine the initial velocity model. We then derive a set of elastic wave velocities and anisotropic parameters using elastic full-waveform inversion, and obtain PP and PS images using elastic reverse-time migration. We identify major faults by analyzing the variations of seismic velocities and anisotropy parameters, and reveal mid- to small-scale faults by applying a supervised machine learning method to the seismic migration images. Our characterization reveals complex velocity heterogeneities and anisotropies, as well as faults, with a high spatial resolution. These results can provide valuable information for optimal placement of future injection and production wells to increase geothermal energy production at the Blue Mountain geothermal power plant.

58 GEOSCIENCES↗

Rheological Properties of Small-Molecular Liquids at High Shear Strain Rates

Molecular-scale understanding of rheological properties of small-molecular liquids and polymers is critical to optimizing their performance in practical applications such as lubrication and hydraulic fracking. We combine nonequilibrium molecular dynamics simulations with two unsupervised machine learning methods: principal component analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE), to extract the correlation between the rheological properties and molecular structure of squalane sheared at high strain rates (10 6 –10 10 s -1 ) for which substantial shear thinning is observed under pressures P ϵ 0.1–955 MPa at 293 K. Intramolecular atom pair orientation tensors of 435 × 6 dimensions and the intermolecular atom pair orientation tensors of 61 × 6 dimensions are reduced and visualized using PCA and t-SNE to assess the changes in the orientation order during the shear thinning of squalane. Dimension reduction of intramolecular orientation tensors at low pressures P = 0.1,100 MPa reveals a strong correlation between changes in strain rate and the orientation of the side-backbone atom pairs, end-backbone atom pairs, short backbone-backbone atom pairs, and long backbone-backbone atom pairs associated with a squalane molecule. At high pressures P ≥ 400 MPa, the orientation tensors are better classified by these different pair types rather than strain rate, signaling an overall limited evolution of intramolecular orientation with changes in strain rate. Dimension reduction also finds no clear evidence of the link between shear thinning at high pressures and changes in the intermolecular orientation. The alignment of squalane molecules is found to be saturated over the entire range of rates during which squalane exhibits substantial shear thinning at high pressures.

36 MATERIALS SCIENCE↗

Formation and Reconnection of Electron Scale Current Layers in the Turbulent Outflows of a Primary Reconnection Site

We simulate with 3D particle in cell, the spontaneous formation of turbulent outflows in an initially laminar 3D reconnecting current layer. We observe the formation of many secondary current layers and reconnection sites in the outflow. The approach we follow is to study each individual feature within the turbulent outflow. To identify all clusters of current in the outflow we use a clustering technique widely used in unsupervised machine learning: density-based spatial clustering of applications with noise. Once the clusters are identified we measure their size and compute reconnection indicators to establish which are undergoing reconnection. With this analysis we establish that the size of the current clusters reaches all the way from its initial system scale down to subelectron skin depth scale. We observe that the smaller current clusters are more prone to reconnecting and to releasing energy. We then find the process of reconnection of the smaller current cluster to be of the recently observed electron-only type that leaves the ions essentially unaffected.

79 ASTRONOMY AND ASTROPHYSICS↗

Synoptic Weather Regime Classifications for June, July, August, and September, 2022

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A). This dataset includes the data in June, July, August, and September; the last year of the data is 2022.

54 ENVIRONMENTAL SCIENCES↗

Synoptic Weather Regime Classifications for the whole year, from 2014 to 2015

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A).

54 ENVIRONMENTAL SCIENCES↗

Synoptic Weather Regime Classifications for June, July, August, from 2000 to 2024

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A).

node_som_ml↗