Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical feature extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Noise reduction in X-ray photon correlation spectroscopy with convolutional neural networks encoder–decoder models

Abstract Like other experimental techniques, X-ray photon correlation spectroscopy is subject to various kinds of noise. Random and correlated fluctuations and heterogeneities can be present in a two-time correlation function and obscure the information about the intrinsic dynamics of a sample. Simultaneously addressing the disparate origins of noise in the experimental data is challenging. We propose a computational approach for improving the signal-to-noise ratio in two-time correlation functions that is based on convolutional neural network encoder–decoder (CNN-ED) models. Such models extract features from an image via convolutional layers, project them to a low dimensional space and then reconstruct a clean image from this reduced representation via transposed convolutional layers. Not only are ED models a general tool for random noise removal, but their application to low signal-to-noise data can enhance the data’s quantitative usage since they are able to learn the functional form of the signal. We demonstrate that the CNN-ED models trained on real-world experimental data help to effectively extract equilibrium dynamics’ parameters from two-time correlation functions, containing statistical noise and dynamic heterogeneities. Strategies for optimizing the models’ performance and their applicability limits are discussed.

36 MATERIALS SCIENCE↗

Speech Acquisition and Automatic Speech Recognition for Integrated Spacesuit Audio Systems

A voice-command human-machine interface system has been developed for spacesuit extravehicular activity (EVA) missions. A multichannel acoustic signal processing method has been created for distant speech acquisition in noisy and reverberant environments. This technology reduces noise by exploiting differences in the statistical nature of signal (i.e., speech) and noise that exists in the spatial and temporal domains. As a result, the automatic speech recognition (ASR) accuracy can be improved to the level at which crewmembers would find the speech interface useful. The developed speech human/machine interface will enable both crewmember usability and operational efficiency. It can enjoy a fast rate of data/text entry, small overall size, and can be lightweight. In addition, this design will free the hands and eyes of a suited crewmember. The system components and steps include beam forming/multi-channel noise reduction, single-channel noise reduction, speech feature extraction, feature transformation and normalization, feature compression, model adaption, ASR HMM (Hidden Markov Model) training, and ASR decoding. A state-of-the-art phoneme recognizer can obtain an accuracy rate of 65 percent when the training and testing data are free of noise. When it is used in spacesuits, the rate drops to about 33 percent. With the developed microphone array speech-processing technologies, the performance is improved and the phoneme recognition accuracy rate rises to 44 percent. The recognizer can be further improved by combining the microphone array and HMM model adaptation techniques and using speech samples collected from inside spacesuits. In addition, arithmetic complexity models for the major HMMbased ASR components were developed. They can help real-time ASR system designers select proper tasks when in the face of constraints in computational resources.

Huang, Yiteng↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

Cyber-Attack Identification of Synchrophasor Data Via VMD and Multifusion SVM

A large amount of synchrophasor data in the wide area measurement system (WAMS) needs to be collected and transmitted to the phasor data concentrator, thereby increasing the possibility of being attacked by hackers. The attacked data are therefore hidden into the normal synchrophasor data so that the synchrophasor data based application will be affected. To remedy this problem, an identification framework is proposed to detect the data cyber-attack in WAMS utilizing variational mode decomposition (VMD) and multifusion support vector machine (MSVM). First, VMD is used to transform the attacked data into multiple modal components. Thereafter, a novel MSVM is employed to classify the deterministic features using the proposed linear combined multikernel (LCM). Further, this LCM can fuse multiple types of features, including the time, frequency, and statistical domains of the synchrophasor data. Utilizing the actual data from FNET/GridEye, different experiments are conducted under multiple attack strengths and types. The results demonstrate that the identification framework has higher precision and robustness compared with other conventional classifiers.

97 MATHEMATICS AND COMPUTING↗

Algorithms for Spectral Decomposition with Applications to Optical Plume Anomaly Detection

The analysis of spectral signals for features that represent physical phenomenon is ubiquitous in the science and engineering communities. There are two main approaches that can be taken to extract relevant features from these high-dimensional data streams. The first set of approaches relies on extracting features using a physics-based paradigm where the underlying physical mechanism that generates the spectra is used to infer the most important features in the data stream. We focus on a complementary methodology that uses a data-driven technique that is informed by the underlying physics but also has the ability to adapt to unmodeled system attributes and dynamics. We discuss the following four algorithms: Spectral Decomposition Algorithm (SDA), Non-Negative Matrix Factorization (NMF), Independent Component Analysis (ICA) and Principal Components Analysis (PCA) and compare their performance on a spectral emulator which we use to generate artificial data with known statistical properties. This spectral emulator mimics the real-world phenomena arising from the plume of the space shuttle main engine and can be used to validate the results that arise from various spectral decomposition algorithms and is very useful for situations where real-world systems have very low probabilities of fault or failure. Our results indicate that methods like SDA and NMF provide a straightforward way of incorporating prior physical knowledge while NMF with a tuning mechanism can give superior performance on some tests. We demonstrate these algorithms to detect potential system-health issues on data from a spectral emulator with tunable health parameters.

Srivastava, Askok N.↗

Accelerating phase field simulations through a hybrid adaptive Fourier neural operator with U-net backbone

Prolonged contact between a corrosive liquid and metal alloys can cause progressive dealloying. For one such process as liquid-metal dealloying (LMD), phase field models have been developed to understand the mechanisms leading to complex morphologies. However, the LMD governing equations in these models often involve coupled non-linear partial differential equations (PDE), which are challenging to solve numerically. In particular, numerical stiffness in the PDEs requires an extremely refined time step size (on the order of 10 -12 s or smaller). This computational bottleneck is especially problematic when running LMD simulation until a late time horizon is required. This motivates the development of surrogate models capable of leaping forward in time, by skipping several consecutive time steps at-once. In this paper, we propose a U-shaped adaptive Fourier neural operator (U-AFNO), a machine learning (ML) based model inspired by recent advances in neural operator learning. U-AFNO employs U-Nets for extracting and reconstructing local features within the physical fields, and passes the latent space through a vision transformer (ViT) implemented in the Fourier space (AFNO). We use U-AFNOs to learn the dynamics of mapping the field at a current time step into a later time step. We also identify global quantities of interest (QoI) describing the corrosion process (e.g., the deformation of the liquid-metal interface, lost metal, etc.) and show that our proposed U-AFNO model is able to accurately predict the field dynamics, in spite of the chaotic nature of LMD. Most notably, our model reproduces the key microstructure statistics and QoIs with a level of accuracy on par with the high-fidelity numerical solver, while achieving a significant 11, 200 × speed-up on a high-resolution grid when comparing the computational expense per time step. Finally, we also investigate the opportunity of using hybrid simulations, in which we alternate forward leaps in time using the U-AFNO with high-fidelity time stepping. We demonstrate that while advantageous for some surrogate model design choices, our proposed U-AFNO model in fully auto-regressive settings consistently outperforms hybrid schemes.

36 MATERIALS SCIENCE↗

Active Missions and the VxOs with THEMIS as an Example

The Virtual Observatories (VxOs) provide a host of services to data producers and researchers. They help data producers to describe their data in standard Space Physics Archive Search and Extract (SPASE) terms that enable scientists to understand data products from a wide range of missions. They offer search interfaces based on specified criteria that help researchers discover conjunctions, prominent events, and intervals of interest. In this talk, we show how VMO services can be used with Time History of Events and Macroscale Interactions during Substorms (THEMIS) observations to identify magnetotail intervals marked by high speed flows, enhanced densities, or high temperatures. We present statistical surveys of when and where these phenomena occur. We then show how the VMO services can be used to identify events in which two or more THEMIS spacecraft observe specified features for more detailed analysis. We conclude by discussing the current limitations of VMO tools and outline plans for the future.

Sibeck, D. G.↗

Rotor blade imbalance fault detection for variable-speed marine current turbines via generator power signal analysis

Marine hydrokinetic (MHK) turbines extract renewable energy from oceanic environments. However, due to the harsh conditions that these turbines operate in, system performance naturally degrades over time. Thus, ensuring efficient condition-based maintenance is imperative towards guaranteeing reliable operation and reduced costs for marine hydrokinetic power. This work proposes a novel framework aimed at identifying and classifying the severity of rotor blade pitch imbalance faults experienced by marine current turbines (MCTs). In the framework, a Continuous Morlet Wavelet Transform (CMWT) is first utilized to acquire the wavelet coefficients encompassed within the 1P frequency range of the turbine's rotor shaft. From these coefficients, several statistical indices are tabulated into a six-dimensional feature space. Next, Principle Component Analysis (PCA) is employed on the resulting feature space for dimensionality reduction, and then the application of a K-Nearest Neighbor (KNN) machine learning algorithm is utilized for fault detection and severity classification. The effectiveness of the proposed framework is validated using a high-fidelity MCT numerical simulation platform, where results demonstrate that the presence of a pitch imbalance fault can be accurately detected 100% of the time and correctly classified based upon severity more than 97% of the time.

42 ENGINEERING↗

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

New constraints on sodium production in globular clusters from the Na 23 ( He 3 , d ) Mg 24 reaction

The star-to-star anticorrelation of sodium and oxygen is a defining feature of globular clusters, but, to date, the astrophysical site responsible for this unique chemical signature remains unknown. Sodium enrichment within these clusters depends sensitively on reaction rate of the sodium destroying reactions 23 Na(p, γ) and 23 Na(p,α). In this paper, we report the results of a 23 Na( 3 He,d) 24 Mg transfer reaction carried out at Triangle Universities Nuclear Laboratory using a 21 MeV 3 He beam. Astrophysically relevant states in 24 Mg between 11 < E x < 12 MeV were studied using high-resolution magnetic spectroscopy, thereby allowing the extraction of excitation energies and spectroscopic factors. Bayesian methods are combined with the distorted wave Born approximation to assign statistically meaningful uncertainties to the extracted spectroscopic factors. For the first time, these uncertainties are propagated through to the estimation of proton partial widths. Our experimental data are used to calculate the reaction rate. The impact of the new rates are investigated using asymptotic giant branch star models. Furthermore, it is found that while the astrophysical conditions still dominate the total uncertainty, intramodel variations on sodium production from the 23 Na(p, γ) and 23 Na(p,α) reaction channels are a lingering source of uncertainty.

20 ≤ A ≤ 38↗

Exploring diversion-pathway analysis of a generic molten-salt fast reactor using multiphysics informed signatures

Molten salt reactors are being explored by multiple commercial ventures due to their inherent safety features, flexibility in fuel sources, and high fuel utilization and thermal efficiency. The continual flow of fuel salt, large fissile quantities present, and ability to add or divert material due to the liquid nature introduces new challenges for international safeguards. To understand how international safeguards should be applied, it is important to capture the inherent multi-physics nature of a molten salt reactor. This work examines a generic molten salt fast reactor to understand how potential diversion scenarios would affect the concentration of radionuclides in the primary and auxiliary systems. Three types of diversion were examined: a slow drip of fuel salt, gaseous plutonium extraction, and uranium metal plating. The analysis determined that several key isotopes become statistically significant once diversion begins, indicating that detection of such diversion cases would be possible through measuring specific signatures such as gamma spectra.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

MaTableGPT: GPT‐Based Table Data Extractor from Materials Science Literature

Abstract Efficiently extracting data from tables in the scientific literature is pivotal for building large‐scale databases. However, the tables reported in materials science papers exist in highly diverse forms; thus, rule‐based extractions are an ineffective approach. To overcome this challenge, the study presents MaTableGPT, which is a GPT‐based table data extractor from the materials science literature. MaTableGPT features key strategies of table data representation and table splitting for better GPT comprehension and filtering hallucinated information through follow‐up questions. When applied to a vast volume of water splitting catalysis literature, MaTableGPT achieves an extraction accuracy (total F1 score) of up to 96.8%. Through comprehensive evaluations of the GPT usage cost, labeling cost, and extraction accuracy for the learning methods of zero‐shot, few‐shot, and fine‐tuning, the study presents a Pareto‐front mapping where the few‐shot learning method is found to be the most balanced solution owing to both its high extraction accuracy (total F1 score >95%) and low cost (GPT usage cost of 5.97 US dollars and labeling cost of 10 I/O paired examples). The statistical analyses conducted on the database generated by MaTableGPT revealed valuable insights into the distribution of the overpotential and elemental utilization across the reported catalysts in the water splitting literature.

Yi, Gyeong Hoon [Computational Science Research Ce↗

Investigation of process history and underlying phenomena associated with the synthesis of plutonium oxides using Vector Quantizing Variational Autoencoder

Accurate, high throughput, and unbiased analysis of plutonium oxide particles is needed for analysis of the phenomenology associated with process parameters in their synthesis. Compared to qualitative and taxonomic descriptors, quantitative descriptors of particle morphology through scanning electron microscopy (SEM) have shown success in analyzing process parameters of uranium oxides. Among other candidates, a neural network called a Vector Quantizing Variational Autoencoder (VQ-VAE) has shown the ability to quantitatively describe particle morphology to attain >85% accuracy in identifying uranium oxide processing routes. We utilize a VQ-VAE to quantitatively describe plutonium dioxide (PuO 2 ) particles created in a designed experiment and investigate their phenomenology and prediction of their process parameters. PuO 2 was calcined from Pu(III) oxalates that were precipitated under varying synthetic conditions that related to concentrations, temperature, addition and digestion times, precipitant feed, and strike order; the surface morphology of the resulting PuO 2 powders were analyzed by SEM. A pipeline was developed to extract and quantify useful image representations for individual particles with the VQ-VAE, then further reduce the dimensionality of the feature space using a bottlenecking neural network fit to perform multiple classification tasks simultaneously. The reduced feature space could predict process parameters with greater than 80% accuracies for some parameters with a single particle. They also showed utility for grouping particles with similar surface morphology characteristics together. Both the clustering and classification results reveal valuable information regarding which chemical process parameters chiefly influence the PuO 2 particle morphologies: strike order and oxalic acid feedstock. Doing the same analysis with multiple particles was shown to improve the classification accuracy on each process parameter over the use of a single particle, with statistically significant results generally seen with as few as four particles in a sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Different CT slice thickness and contrast‐enhancement phase in radiomics models on the differential performance of lung adenocarcinoma

Abstract Background To investigate the effects of computed tomography (CT) reconstruction slice thickness and contrast‐enhancement phase on the differential diagnosis performance of radiomic signature in lung adenocarcinoma. Methods A total of 187 patients who had been pathologically confirmed with lung adenocarcinoma and nonadenocarcinoma were divided into a training cohort ( n = 149) and validation cohort ( n = 38). All the patients underwent contrast‐enhanced CT and the images were reconstructed with different slice thickness. The radiomic features were extracted from different slice thickness and scan phase. The logistic regression (LR) algorithm was used to build a machine learning model for each group. The area under the curve (AUC) obtained from the receiver operating characteristic (ROC) curve and DeLong test was used to evaluate its discriminating performance. Results Finally, 34 image features and five semantic features were selected to establish a radiomics model. Based on the three contrast‐enhanced CT phases and four reconstruction slice thickness, 12 groups of radiomics models showed good discrimination ability with the AUCs range from 0.9287 to 0.9631, sensitivity range from 0.8349 to 0.9083, specificity range from 0.825 to 0.925 in the training group. Similar results were observed in the validation group. However, there was no statistical significance between the different CT scan phase groups and different slice thickness ( p > 0.05). Conclusions The radiomic analysis of contrast‐enhanced CT can be used for the differential diagnosis of lung adenocarcinoma. Moreover, different slice thickness and contrast‐enhanced scan phase did not affect the discriminating ability in the radiomics models.

Wang, Yang↗

An IDL-based analysis package for COBE and other skycube-formatted astronomical data

UIMAGE is a data analysis package written in IDL for the Cosmic Background Explorer (COBE) project. COBE has extraordinarily stringent accuracy requirements: 1 percent mid-infrared absolute photometry, 0.01 percent submillimeter absolute spectrometry, and 0.0001 percent submillimeter relative photometry. Thus, many of the transformations and image enhancements common to analysis of large data sets must be done with special care. UIMAGE is unusual in this sense in that it performs as many of its operations as possible on the data in its native format and projection, which in the case of COBE is the quadrilateralized sphereical cube ('skycube'). That is, after reprojecting the data, e.g., onto an Aitoff map, the user who performs an operation such as taking a crosscut or extracting data from a pixel is transparently acting upon the skycube data from which the projection was made, thereby preserving the accuracy of the result. Current plans call for formatting external data bases such as CO maps into the skycube format with a high-accuracy transformation, thereby allowing Guest Investigators to use UIMAGE for direct comparison of the COBE maps with those at other wavelengths from other instruments. It is completely menu-driven so that its use requires no knowledge of IDL. Its functionality includes I/O from the COBE archives, FITS files, and IDL save sets as well as standard analysis operations such as smoothing, reprojection, zooming, statistics of areas, spectral analysis, etc. One of UIMAGE's more advanced and attractive features is its terminal independence. Most of the operations (e.g., menu-item selection or pixel selection) that are driven by the mouse on an X-windows terminal are also available using arrow keys and keyboard entry (e.g., pixel coordinates) on VT200 and Tektronix-class terminals. Even limited grey scales of images are available this way. Obviously, image processing is very limited on this type of terminal, but it is nonetheless surprising how much analysis can be done on that medium. Such flexibility has the virtue of expanding the user community to those who must work remotely on non-image terminals, e.g., via modem.

Ewing, J. A.↗

Limited-view Cone Beam CT reconstruction using 3D Patch-based Supervised and Adversarial Learning [Slides]

We present a novel machine learning CNN architecture that can learn from limited data combined appropriately with physics and statistical priors (e.g., forward models and noise models). To address the limited availability of training data we adopt a 3D patch-based approach for our models. Patch-based learning is central to several image reconstruction methods and demands fewer training data than DL approaches, as a single data volume can be broken into several millions of overlapping 3D sub-volumes or patches. This creates a very large number of training sub-volumes from a limited number of overall image volumes. A 3D Generative Adversarial Networks (GAN) is then trained to remove artifacts at the sub-volume level. The combination of a sub-volume-based approach with DL allows us to exploit the richness of the latter in extracting and representing image features, while avoiding risks associated with overfitting due to limited training data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Stress and Strain Heterogeneity and Persistence in Uniaxially‐ and Triaxially‐Loaded Sandstone

Two critical questions in brittle rock mechanics are how rocks developed localized strains and to what extent internal stress heterogeneity controls this localization and subsequent macroscopic failure. Definitive answers have not yet been found, but would provide insight into rock fracture mechanics as relevant to hydrocarbon extraction and sequestration. Here, we use synchrotron X‐ray tomography (XRT) and 3D X‐ray diffraction (3DXRD) during uniaxial and triaxial tests on Nugget and Bentheimer sandstones to examine strain and stress localization prior to mechanical failure. 3DXRD was used to measure intra‐granular lattice strains which were used to compute elastic stress tensors of each grain. Digital volume correlation (DVC) was applied to XRT images to determine the strain field in the sample. Both samples featured marked spatial heterogeneity, localization, and temporal persistence of elevated stresses and strains during their mechanical deformation toward failure. Both samples featured a majority of grains with at least one principal stress component that was tensile, a signature of the influence of heterogeneity on stress transmission. Measurements further revealed that compressive stress orientations and statistics evolved in a similar manner to those of inter‐particle forces in loose granular materials, with triaxially‐compressed rock exhibiting enhanced grain stress heterogeneity compared to uniaxially‐compressed rock. Our results complement recent work by others who employed XRT and scanning 3DXRD to study triaxially‐compressed sandstone, but extend those results to uniaxial compression, sandstones of varied porosity, and grain stress measurements throughout the 3D full extent of the samples rather than in a single layer examined with scanning 3DXRD.

58 GEOSCIENCES↗