Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical feature extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Estimating Leaf Area Index in Row Crops Using Wheel-Based and Airborne Discrete Return Light Detection and Ranging Data

Leaf area index (LAI) is an important variable for characterizing plant canopy in crop models. It is traditionally defined as the total one-sided leaf area per unit ground area and is estimated by both direct and indirect methods. This paper explores the effectiveness of using light detection and ranging (LiDAR) data to estimate LAI for sorghum and maize with different treatments at multiple times during the growing season from both a wheeled vehicle and Unmanned Aerial Vehicles. Linear and nonlinear regression models are investigated for prediction utilizing statistical and plant structure-based features extracted from the LiDAR point cloud data with ground reference obtained from an in-field plant canopy analyzer (indirect method). Results based on the value of the coefficient of determination ( R 2 ) and root mean squared error for predictive models ranged from ∼0.4 in the early season to ∼0.6 for sorghum and ∼0.5 to 0.80 for maize from 40 Days after Sowing to harvest.

59 BASIC BIOLOGICAL SCIENCES↗

Robust Matching of Wavelet Features for Sub-Pixel Registration of Landsat Data

For many Earth and Space Science applications, automatic geo-registration at sub-pixel accuracy has become a necessity. In this work, we are focusing on building an operational system, which will provide a sub-pixel accuracy registration of Landsat-5 and Landsat-7 data. The input to our registration method consists of scenes that have been geometrically and radiometrically corrected. Such pre-processed scenes are then geo-registered relative to a database of Landsat chips. The method assumes a transformation composed of a rotation and a translation, and utilizes rotation- and translation-invariant wavelets to extract image features that are matched using statistically robust feature matching and a generalized Hausdorff distance metric. The registration process is described and results on four Landsat input scenes of the Washington, D.C. area are presented.

LeMoigne, Jacqueline↗

Feasibility of an Accelerometer-Based Structural Health Monitoring System for the LANL Blast Tube

A modeling- and simulation-based study was conducted on the feasibility of implementing an accelerometer-based SHM system on the Los Alamos National Laboratory blast tube. A blast tube experiment was modeled using the Abaqus explicit finite element solver. A custom user subroutine was written to apply test-like pressure loading to the inside surface of the blast tube. The subroutine applies analytically defined pressure loads derived from tracer output taken from a Compressible Flow Computational Fluid Dynamics Solver model of the blast tube. Five unique versions of the model were created: an undamaged reference model at 65°F was used as the baseline and compared to equivalent models at 10°F and 100°F. These three models were compared to models with small damage at the reference temperature. The two types of damage considered were a radial (circumferential) crack in the main tube body and a longitudinal crack in the supports. Acceleration outputs were extracted from accelerometer bodies included in the model and were post processed using a variety of standard SHM techniques. Different potential features signaling failure were extracted and compared using statistical methods in the time and frequency domains. A method was identified that clearly shows that differences in structural response resulting from the modeled damage can be differentiated from the structural response resulting from changing environmental conditions. However, the amount of damage applied to create observable differences in the accelerometer data was so large that simpler methods of damage detection would be more cost effective in locating damage.

42 ENGINEERING↗

Visual Data Analysis for Satellites

The Visual Data Analysis Package is a collection of programs and scripts that facilitate visual analysis of data available from NASA and NOAA satellites, as well as dropsonde, buoy, and conventional in-situ observations. The package features utilities for data extraction, data quality control, statistical analysis, and data visualization. The Hierarchical Data Format (HDF) satellite data extraction routines from NASA's Jet Propulsion Laboratory were customized for specific spatial coverage and file input/output. Statistical analysis includes the calculation of the relative error, the absolute error, and the root mean square error. Other capabilities include curve fitting through the data points to fill in missing data points between satellite passes or where clouds obscure satellite data. For data visualization, the software provides customizable Generic Mapping Tool (GMT) scripts to generate difference maps, scatter plots, line plots, vector plots, histograms, timeseries, and color fill images.

Lau, Yee↗

Effects of preprocessing Landsat MSS data on derived features

Important to the use of multitemporal Landsat MSS data for earth resources monitoring, such as agricultural inventories, is the ability to minimize the effects of varying atmospheric and satellite viewing conditions, while extracting physically meaningful features from the data. In general, the approaches to the preprocessing problem have been derived from either physical or statistical models. This paper compares three proposed algorithms; XSTAR haze correction, Color Normalization, and Multiple Acquisition Mean Level Adjustment. These techniques represent physical, statistical, and hybrid physical-statistical models, respectively. The comparisons are made in the context of three feature extraction techniques; the Tasseled Cap, the Cate Color Cube. and Normalized Difference.

Parris, T. M.↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine

03 NATURAL GAS↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

A Topological Approach for Motion Track Discrimination

Detecting small targets at range is difficult because there is not enough spatial information present in an image sub-region containing the target to use correlation-based methods to differentiate it from dynamic confusers present in the scene. Moreover, this lack of spatial information also disqualifies the use of most state-of-the-art deep learning image-based classifiers. Here, we use characteristics of target tracks extracted from video sequences as data from which to derive distinguishing topological features that help robustly differentiate targets of interest from confusers. In particular, we calculate persistent homology from time-delayed embeddings of dynamic statistics calculated from motion tracks extracted from a wide field-of-view video stream. In short, we use topological methods to extract features related to target motion dynamics that are useful for classification and disambiguation and show that small targets can be detected at range with high probability.

Emerson, Tegan H.↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Statistical White-Line Analysis in High-Throughput TXM-XANES for Chemical State Quantification

The transmission X-ray microscopy (TXM) based X-ray absorption near-edge structure (XANES) technique provides three-dimensional mapping of element-specific chemical states at nanometer-scale spatial resolution and micrometer-scale fields of view. However, compared to conventional volume-averaged XANES (VA-XANES) measurements, the inherently small voxel size in TXM-XANES leads to a lower signal-to-noise ratio, making full-spectrum analysis computationally demanding and less robust. Here, we present the structural and compositional conditions for a statistical white-line analysis framework under which chemical state information can be directly extracted from the white-line peak position in voxel spectra without the need for voxel-wise background subtraction or normalization, under well-defined structural and compositional conditions. The method is validated on layered oxide cathode materials, where low-order polynomial fitting accurately reproduces white-line features, and the extracted energy distributions correlate strongly with VA-XANES results. This statistical approach enables high-throughput, dose-efficient, and noise-robust chemical state quantification in TXM-XANES, offering broad applicability to functional materials requiring nanoscale oxidation-state mapping.

TXM↗

Evidence of chaotic pattern in solar flux through a reproducible sequence of period-doubling-type bifurcations

Presented here is a preliminary study of the limits to solar flux intensity prediction, and of whether the general lack of predictability in the solar flux arises from the nonlinear chaotic nature of the Sun's physical activity. Statistical analysis of a chaotic signal can extract only its most gross features, and detailed physical models fail, since even the simplest equations of motion for a nonlinear system can exhibit chaotic behavior. A recent theory by Feigenbaum suggests that nonlinear systems that can be led into chaotic behavior through a sequence of period-doubling bifurcations will exhibit a universal behavior. As the control parameter is increased, the bifurcation points occur in such a way that a proper ratio of these will approach the universal Feigenbaum number. Experimental evidence supporting the applicability of the Feigenbaum scenario to solar flux data is sparse. However, given the hypothesis that the Sun's convection zones are similar to a Rayleigh-Bernard mechanism, we can learn a great deal from the remarkable agreement observed between the prediction by theory (period doubling - a universal route to chaos) and the amplitude decrease of the signal's regular subharmonics. The authors show that period-doubling-type bifurcation is a possible route to a chaotic pattern of solar flux that is distinguishable from the logarithm of its power spectral density. This conclusion is the first positive step toward a reformulation of solar flux by a nonlinear chaotic approach. The ultimate goal of this research is to be able to predict an estimate of the upper and lower bounds for solar flux within its predictable zones. Naturally, it is an important task to identify the time horizons beyond which predictability becomes incompatible with computability.

Ashrafi, S.↗

Evidence of chaotic pattern in solar flux through a reproducible sequence of period-doubling-type bifurcations

A preliminary study of the limits to solar flux intensity prediction, and of whether the general lack of predictability in the solar flux arises from the nonlinear chaotic nature of the Sun's physical activity is presented. Statistical analysis of a chaotic signal can extract only its most gross features, and detailed physical models fail, since even the simplest equations of motion for a nonlinear system can exhibit chaotic behavior. A recent theory by Feigenbaum suggests that nonlinear systems that can be led into chaotic behavior through a sequence of period-doubling bifurcations will exhibit a universal behavior. As the control parameter is increased, the bifurcation points occur in such a way that a proper ratio of these will approach the universal Feigenbaum number. Experimental evidence supporting the applicability of the Feigenbaum scenario to solar flux data is sparse. However, given the hypothesis that the Sun's convection zones are similar to a Rayleigh-Bernard mechanism, we can learn a great deal from the remarkable agreement observed between the prediction by theory (period doubling - a universal route to chaos) and the amplitude decrease of the signal's regular subharmonics. It is shown that period-doubling-type bifurcation is a possible route to a chaotic pattern of solar flux that is distinguishable from the logarithm of its power spectral density. This conclusion is the first positive step toward a reformulation of solar flux by a nonlinear chaotic approach. The ultimate goal of this research is to be able to predict an estimate of the upper and lower bounds for solar flux within its predictable zones. Naturally, it is an important task to identify the time horizons beyond which predictability becomes incompatible with computability.

Ashrafi, S.↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Decoding the Mechanisms of Phase Transitions from In Situ Microscopy Observations

Abstract Analysis of the temperature‐ and stimulus‐dependent imaging data toward elucidation of the physical transformations is an ubiquitous problem in multiple fields. Here, temperature‐induced phase transition in BaTiO 3 is explored using the machine learning analysis of domain morphologies visualized via variable‐temperature scanning transmission electron microscopy (STEM) imaging data. This approach is based on the multivariate statistical analysis of the time or temperature dependence of the statistical descriptors of the system, derived in turn from the categorical classification of observed domain structures or projection on the continuous parameter space of the feature extraction‐dimensionality reduction transform. The proposed workflow offers a powerful tool for the exploration of the dynamic data based on the statistics of image representation as a function of the external control variable to visualize the transformation pathways during phase transitions and chemical reactions. This can include the mesoscopic STEM data as demonstrated here, but also optical, chemical imaging, etc., data. It can further be extended to the higher dimensional spaces, for example, analysis of the combinatorial libraries of materials compositions.

Valleti, Sai Mani Prudhvi↗

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

Low-level processing for real-time image analysis

A system that detects object outlines in television images in real time is described. A high-speed pipeline processor transforms the raw image into an edge map and a microprocessor, which is integrated into the system, clusters the edges, and represents them as chain codes. Image statistics, useful for higher level tasks such as pattern recognition, are computed by the microprocessor. Peak intensity and peak gradient values are extracted within a programmable window and are used for iris and focus control. The algorithms implemented in hardware and the pipeline processor architecture are described. The strategy for partitioning functions in the pipeline was chosen to make the implementation modular. The microprocessor interface allows flexible and adaptive control of the feature extraction process. The software algorithms for clustering edge segments, creating chain codes, and computing image statistics are also discussed. A strategy for real time image analysis that uses this system is given.

Eskenazi, R.↗