Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical feature extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

Structured light scanning artifact-based performance study

Structured light scanning is used to create a digital twin of a manufactured part, where features are extracted to determine if the part meets the designer’s intent and required tolerances. This paper describes repeatability and reproducibility analyses for a commercially-available structured light scanning system and measurement artifact. The repeatability study used five repeated scans at 15 measurement positions. Repeatability was assessed by randomly selecting one of the five scans at each of the 15 positions and creating a part mesh. This process was performed 50 times and the statistics for the dimension variations were calculated to isolate the scanning effects only. The same sequence was then performed for 10 of the 15 positions and five of the 15 positions to evaluate the repeatability sensitivity to the number of measurement positions. Reproducibility was assessed by selecting 15 positions to create a mesh and repeating the 15-position measurement sequence 10 times using different positions for each mesh construction. The statistics for the dimension variations were then calculated. This incorporated the effects of both scanning and the position and orientation of the part relative to the scanner. This sequence was repeated for 10-position and five-position scans to evaluate the corresponding sensitivity. Finally, the artifact dimensions from structured light scanning were compared to coordinate measuring machine measurements of the same features.

42 ENGINEERING↗

Mapathons versus automated feature extraction: a comparative analysis for strengthening immunization microplanning

Background: Social instability and logistical factors like the displacement of vulnerable populations, the difficulty of accessing these populations, and the lack of geographic information for hard-to-reach areas continue to serve as barriers to global essential immunizations (EI). Microplanning, a population-based, healthcare intervention planning method has begun to leverage geographic information system (GIS) technology and geospatial methods to improve the remote identification and mapping of vulnerable populations to ensure inclusion in outreach and immunization services, when feasible. We compare two methods of accomplishing a remote inventory of building locations to assess their accuracy and similarity to currently employed microplan line-lists in the study area. Methods: The outputs of a crowd-sourced digitization effort, or mapathon, were compared to those of a machine-learning algorithm for digitization, referred to as automatic feature extraction (AFE). The following accuracy assessments were employed to determine the performance of each feature generation method: (1) an agreement analysis of the two methods assessed the occurrence of matches across the two outputs, where agreements were labeled as “befriended” and disagreements as “lonely”; (2) true and false positive percentages of each method were calculated in comparison to satellite imagery; (3) counts of features generated from both the mapathon and AFE were statistically compared to the number of features listed in the microplan line-list for the study area; and (4) population estimates for both feature generation method were determined for every structure identified assuming a total of three households per compound, with each household averaging two adults and 5 children. Results: The mapathon and AFE outputs detected 92,713 and 53,150 features, respectively. A higher proportion (30%) of AFE features were befriended compared with befriended mapathon points (28%). The AFE had a higher true positive rate (90.5%) of identifying structures than the mapathon (84.5%). The difference in the average number of features identified per area between the microplan and mapathon points was larger (t = 3.56) than the microplan and AFE (t = -2.09) (alpha = 0.05). Conclusions: Our findings indicate AFE outputs had higher agreement (i.e., befriended), slightly higher likelihood of correctly identifying a structure, and were more similar to the local microplan line-lists than the mapathon outputs. These findings suggest AFE may be more accurate for identifying structures in high-resolution satellite imagery than mapathons. However, they both had their advantages and the ideal method would utilize both methods in tandem.

59 BASIC BIOLOGICAL SCIENCES↗

Maximum likelihood classification of synthetic aperture radar imagery

Classification of synthetic aperture radar (SAR) images has important applications in geology, agriculture, and the military. A statistical model for SAR images is reviewed and a maximum likelihood classification algorithm developed for the classification of agricultural fields based on the model. It is first assumed that the target feature information is known a priori. The performance of the algorithm is then evaluated in terms of the probability of incorrect classification. A technique is also presented to extract the needed feature information from a SAR image; then both the feature extraction and the maximum likelihood classification algorithms are tested on a SEASAT-A SAR image.

Frost, V. S.↗

IUE data reduction: Wavelength determinations and line identifications using a VAX/750 computer

A fully automated, interactive system for determining the wavelengths of features in extracted IUE spectra is described. Wavelengths are recorded from video displays of expanded plots of individual orders using a movable cursor, and then corrected for IUE wavelength scale errors. The estimated accuracy of an individual wavelength in the final tabulation is 0.050 A. Such lists are ideally suited for line identification work using the method of wavelength coincidence statistics (WCS). The results of WCS studies of the ultraviolet spectra of the chemically peculiar (CP) stars iota Coronae Borealis and kappa Camcri. Aside from confirming a number of previously reported aspects of the abundance patterns in these stars, the searches produced some interesting, new discoveries, notably the presence of Hf in the spectrum of kappa Camcri. The implications of this work for theories designed to account for anomalous abundances in chemically peculiar stars are discussed.

Davidson, J. P.↗

Continental Spatio-Temporal Data Analysis with Linear Spectral Mixture Model Using FOSS

This work demonstrates the development and implementation of a Fully Constrained Least Squares (FCLS) unmixing model developed in C++ programming language with OpenCV package and boost C++ libraries in the NASA Earth Exchange (NEX). Visualization of the results is supported by GRASS GIS and statistical analysis is carried in R in a Linux system environment. FCLS was first tested on computer simulated data with Gaussian noise of various signal-to-noise ratio, and Landsat data of an agricultural scenario and an urban environment using a set of global end members of substrate (soils, sediments, rocks, and non-photosynthetic vegetation), vegetation that includes green photosynthetic plants and dark objects which encompasses absorptive substrate materials, clear water, deep shadows, etc. For the agricultural scenario, a spectrally diverse collection of 11 scenes of Level 1 terrain corrected, cloud free Landsat-5 TM data of Fresno, California, USA were unmixed and the results were validated with the corresponding ground data. To study an urbanized landscape, a clear sky Landsat-5 TM data were unmixed and validated with coincident World View-2 abundance maps (of 2 m spatial resolution) for an area of San Francisco, California, USA. The results were evaluated using descriptive statistics, correlation coefficient, RMSE, probability of success, boxplot and bivariate distribution function. Finally, FCLS was used for sub-pixel land cover analysis of the monthly WELD (Wen-enabled Landsat data) repository from 2008 to 2011 of North America. The abundance maps in conjunction with DMSP-OLS nighttime lights data were used to extract the urban land cover features and analyze their spatial-temporal growth.

Landsat Satellites↗

Dark Energy Survey Year 3 results: $w$CDM cosmology from simulation-based inference with persistent homology on the sphere

We present cosmological constraints from Dark Energy Survey Year 3 (DES Y3) weak lensing data using persistent homology, a topological data analysis technique that tracks how features like clusters and voids evolve across density thresholds. For the first time, we apply spherical persistent homology to galaxy survey data through the algorithm TopoS2, which is optimized for curved-sky analyses and HEALPix compatibility. Employing a simulation-based inference framework with the Gower Street simulation suite, specifically designed to mimic DES Y3 data properties, we extract topological summary statistics from convergence maps across multiple smoothing scales and redshift bins. After neural network compression of these statistics, we estimate the likelihood function and validate our analysis against baryonic feedback effects, finding minimal biases (under $0.3σ$) in the $Ω_\mathrm{m}-S_8$ plane. Assuming the $w$CDM model, our combined Betti numbers and second moments analysis yields $S_8 = 0.821 \pm 0.018$ and $Ω_\mathrm{m} = 0.304\pm0.037$-constraints 70% tighter than those from cosmic shear two-point statistics in the same parameter plane. Our results demonstrate that topological methods provide a powerful and robust framework for extracting cosmological information, with our spherical methodology readily applicable to upcoming Stage IV wide-field galaxy surveys.

Prat, J. [Nordita; Royal Inst. Tech., Sodertalje; ↗

Simulation-trained machine learning models for Lorentz transmission electron microscopy

Understanding the collective behavior of complex spin textures, such as lattices of magnetic skyrmions, is of fundamental importance for exploring and controlling the emergent ordering of these spin textures and inducing phase transitions. It is also critical to understand the skyrmion–skyrmion interactions for applications such as magnetic skyrmion-enabled reservoir or neuromorphic computing. Magnetic skyrmion lattices can be studied using in situ Lorentz transmission electron microscopy (LTEM), but quantitative and statistically robust analysis of the skyrmion lattices from LTEM images can be difficult. In this work, we show that a convolutional neural network, trained on simulated data, can be applied to perform segmentation of spin textures and to extract quantitative data, such as spin texture size and location, from experimental LTEM images, which cannot be obtained manually. This includes quantitative information about skyrmion size, position, and shape, which can, in turn, be used to calculate skyrmion–skyrmion interactions and lattice ordering. We apply this approach to segmenting images of Néel skyrmion lattices so that we can accurately identify skyrmion size and deformation in both dense and sparse lattices. The model is trained using a large set of micromagnetic simulations as well as simulated LTEM images. This entirely open-source training pipeline can be applied to a wide variety of magnetic features and materials, enabling large-scale statistical studies of spin textures using LTEM.

McCray, Arthur R. C. (ORCID:0000000160774698)↗

Improving the Quasi‐Biennial Oscillation via a Surrogate‐Accelerated Multi‐Objective Optimization

Accurate simulation of the quasi-biennial oscillation (QBO) is challenging due to uncertainties in representing convectively generated gravity waves. We develop an end-to-end uncertainty quantification workflow that calibrates these gravity wave processes in E3SM for a realistic QBO. Central to our approach is a domain knowledge-informed, compressed representation of high-dimensional spatio-temporal wind fields. By employing a parsimonious statistical model that learns the fundamental frequency from complex observations, we extract interpretable and physically meaningful quantities capturing key attributes. Building on this, we train a probabilistic surrogate model that approximates the fundamental characteristics of the QBO as functions of critical physics parameters governing gravity wave generation. Leveraging the Karhunen–Loève decomposition, our surrogate efficiently represents these characteristics as a set of orthogonal features, capturing cross-correlations among multiple physics quantities evaluated at different pressure levels and enabling rapid surrogate-based inference at a fraction of the computational cost of full-scale simulations. Finally, we analyze the inverse problem using a multi-objective approach. Our study reveals a tension between amplitude and period that constrains the QBO representation, precluding a single optimal solution. To navigate this, we quantify the bi-criteria trade-off and generate a set of Pareto optimal parameter values that balance the conflicting objectives. This integrated workflow improves the fidelity of QBO simulations and offers a versatile template for uncertainty quantification in complex geophysical models.

54 ENVIRONMENTAL SCIENCES↗

DeepSAT: A Deep Learning Approach to Tree-Cover Delineation in 1-m NAIP Imagery for the Continental United States

High resolution tree cover classification maps are needed to increase the accuracy of current land ecosystem and climate model outputs. Limited studies are in place that demonstrates the state-of-the-art in deriving very high resolution (VHR) tree cover products. In addition, most methods heavily rely on commercial softwares that are difficult to scale given the region of study (e.g. continents to globe). Complexities in present approaches relate to (a) scalability of the algorithm, (b) large image data processing (compute and memory intensive), (c) computational cost, (d) massively parallel architecture, and (e) machine learning automation. In addition, VHR satellite datasets are of the order of terabytes and features extracted from these datasets are of the order of petabytes. In our present study, we have acquired the National Agriculture Imagery Program (NAIP) dataset for the Continental United States at a spatial resolution of 1-m. This data comes as image tiles (a total of quarter million image scenes with ~60 million pixels) and has a total size of ~65 terabytes for a single acquisition. Features extracted from the entire dataset would amount to ~8-10 petabytes. In our proposed approach, we have implemented a novel semi-automated machine learning algorithm rooted on the principles of "deep learning" to delineate the percentage of tree cover. Using the NASA Earth Exchange (NEX) initiative, we have developed an end-to-end architecture by integrating a segmentation module based on Statistical Region Merging, a classification algorithm using Deep Belief Network and a structured prediction algorithm using Conditional Random Fields to integrate the results from the segmentation and classification modules to create per-pixel class labels. The training process is scaled up using the power of GPUs and the prediction is scaled to quarter million NAIP tiles spanning the whole of Continental United States using the NEX HPC supercomputing cluster. An initial pilot over the state of California spanning a total of 11,095 NAIP tiles covering a total geographical area of 163,696 sq. miles has produced true positive rates of around 88 percent for fragmented forests and 74 percent for urban tree cover areas, with false positive rates lower than 2 percent for both landscapes.

Imagery↗

Compression of Solar Spectroscopic Observations: a Case Study of MgII k Spectral Line Profiles Observed by NASA’s IRIS Satellite

In this study we extract the deep features and investigate the compression of the MgII k spectral line profiles observed in quiet Sun regions by NASA’s IRIS satellite. The data set of line profiles used for the analysis was obtained on April 20th, 2020, at the center of the solar disc, and contains almost 300,000 individual MgII k line profiles after data cleaning. The data are separated into train and test subsets. The train subset was used to train the autoencoder of the varying embedding layer size. The early stopping criterion was implemented on the test subset to prevent the model from overfitting. Our results indicate that it is possible to compress the spectral line profiles more than 27 times (which corresponds to the reduction of the data dimensionality from 110 to 4) while having a 4DN average reconstruction error, which is comparable to the variations in the line continuum. The mean squared error and the reconstruction error of even statistical moments sharply decrease when the dimensionality of the embedding layer increases from 1 to 4 and almost stop decreasing for higher numbers. The observed occasional improvements in training for values higher than 4 indicate that a better compact embedding may potentially be obtained if other training strategies and longer training times are used. The features learned for the critical four-dimensional case can be interpreted. In particular, three of these four features mainly control the line width, line asymmetry, and line dip formation respectively. The presented results are the first attempt to obtain a compact embedding for spectroscopic line profiles and confirm the value of this approach, in particular for feature extraction, data compression, and denoising.

SMD↗

The PAU Survey: Photometric redshifts using transfer learning from simulations

In this paper, we introduce the DEEPZ deep learning photometric redshift (photo-z) code. As a test case, we apply the code to the PAU survey (PAUS) data in the COSMOS field. DEEPZ reduces the σ68 scatter statistic by 50 percent at iAB = 22.5 compared to existing algorithms. This improvement is achieved through various methods, including transfer learning from simulations where the training set consists of simulations as well as observations, which reduces the need for training data. The redshift probability distribution is estimated with a mixture density network (MDN), which produces accurate redshift distributions. Our code includes an autoencoder to reduce noise and extract features from the galaxy SEDs. It also benefits from combining multiple networks, which lowers the photo-z scatter by 10 percent. Furthermore, training with randomly constructed coadded fluxes adds information about individual exposures, reducing the impact of photometric outliers. In addition to opening up the route for higher redshift precision with narrow bands, these machine learning techniques can also be valuable for broad-band surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Automatic Detection of Defects in High-Reliability Components

Disastrous consequences can result from defects in manufactured parts—particularly the high consequence parts developed at Sandia. Identifying flaws in as-built parts can be done with nondestructive means, such as X-ray Computed Tomography (CT). However, due to artifacts and complex imagery, the task of analyzing the CT images falls to humans. Human analysis is inherently unreproducible, unscalable, and can easily miss subtle flaws. We hypothesized that deep learning methods could improve defect identification, increase the number of parts that can effectively be analyzed, and do it in a reproducible manner. We pursued two methods: 1) generating a defect-free version of a scan and looking for differences (PandaNet), and 2) using pre-trained models to develop a statistical model of normality (Feature-based Anomaly Detection System: FADS). Both PandaNet and FADS provide good results, are scalable, and can identify anomalies in imagery. In particular, FADS enables zero-shot (training-free) identification of defects for minimal computational cost and expert time. It significantly outperforms prior approaches in computational cost while achieving comparable results. FADS’ core concept has also shown utility beyond anomaly detection by providing feature extraction for downstream tasks.

47 OTHER INSTRUMENTATION↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Spatial cross-correlation of Antarctic Sea ice and seabed topography

A time series of derived sea ice concentrations as observed about Antarctica by the Nimbus-7 Scanning Multichannel Microwave Radiometer (SMMR) satellite in 1983 is considered. The degree of spatial cross correlation between these data and seabed topography is quantified. The approach is to implement a statistical image processing filter designed to extract local patterns of spatial cross correlation over the entire sea ice field as it undergoes daily changes. Throughout the sea ice, it was found that large scale variations in sea ice concentration correlate systematically with variations in the topography of the seabed. Generally speaking, high concentrations of sea ice occur over deep ocean, whereas areas of encavement, early dissipation and polynya formation develop over topographic features of high elevation. The latter was studied in detail with respect to the features Maud Rise, Astrid Ridge and the continental shelf in the Cosmonaut and Ross Seas. In each case, it is shown that an encavement in sea ice, a polynya, or both develops in the vicinity of the feature in question. As these results are quantified in terms of spatial cross correlation, a potential role is inferred for seabed topography in such fluctuations in the sea ice about Antarctica.

Deveaux, Richard D.↗

Recurrent neural networks for short-term and long-term prediction of geothermal reservoirs

Accurate prediction of geothermal reservoir responses to alternative energy production scenarios is critical for optimizing the development of the underlying resources. While the conventional physics-based models offer a comprehensive prediction tool, data-driven models provide an efficient alternative to build fit-for-purpose predictive models by extracting and using the statistical patterns in the collected data to make predictions. The recurrent neural network (RNN) is a data-driven model that is commonly applied to predict time series sequences. This paper presents a variant of RNN that also utilizes the efficiency of convolutional neural networks (CNN) for the prediction of energy production from geothermal reservoirs. Specifically, a CNN–RNN architecture is developed that takes historical well controls as input (features) and their corresponding production response data as output (labels) to learn an input-output mapping that can predict the future well production responses/performance for any given future well control inputs. The model is paired with a labeling scheme to handle real field disturbances that create data gaps. In addition to the model structure, we introduce a thorough workflow for applying the model, which includes data pre-processing, feature selection, as well as different training strategies for short-term and long-term prediction. Finally, the performance and accuracy of the model are evaluated by applying it to multiple datasets, including a field reservoir model.

15 GEOTHERMAL ENERGY↗

A boundary finding algorithm and its applications

An algorithm for locating gray level and/or texture edges in digitized pictures is presented. The algorithm is based on the concept of hypothesis testing. The digitized picture is first subdivided into subsets of picture elements, e.g., 2 x 2 arrays. The algorithm then compares the first- and second-order statistics of adjacent subsets; adjacent subsets having similar first- and/or second-order statistics are merged into blobs. By continuing this process, the entire picture is segmented into blobs such that the picture elements within each blob have similar characteristics. The boundaries between the blobs comprise the boundaries. The algorithm always generates closed boundaries. The algorithm was developed for multispectral imagery of the earth's surface. Application of this algorithm to various image processing techniques such as efficient coding, information extraction (terrain classification), and pattern recognition (feature selection) are included.

Gupta, J. N.↗

Decision rules for unbiased inventory estimates

An efficient and accurate procedure for estimating inventories from remote sensing scenes is presented. In place of the conventional and expensive full dimensional Bayes decision rule, a one-dimensional feature extraction and classification technique was employed. It is shown that this efficient decision rule can be used to develop unbiased inventory estimates and that for large sample sizes typical of satellite derived remote sensing scenes, resulting accuracies are comparable or superior to more expensive alternative procedures. Mathematical details of the procedure are provided in the body of the report and in the appendix. Results of a numerical simulation of the technique using statistics obtained from an observed LANDSAT scene are included. The simulation demonstrates the effectiveness of the technique in computing accurate inventory estimates.

Argentiero, P. D.↗