Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Feature selection with distance correlation

Choosing which properties of the data to use as input to multivariate decision algorithms—also known as feature selection—is an important step in solving any problem with machine learning. While there is a clear trend towards training sophisticated deep networks on large numbers of relatively unprocessed inputs (so-called automated feature engineering), for many tasks in physics, sets of theoretically well-motivated and well-understood features already exist. Working with such features can bring many benefits, including greater interpretability, reduced training and run time, and enhanced stability and robustness. We develop a new feature selection method based on distance correlation, and demonstrate its effectiveness on the tasks of boosted top- and W -tagging. Using our method to select features from a set of over 7,000 energy flow polynomials, we show that we can match the performance of much deeper architectures, by using only ten features and two orders-of-magnitude fewer model parameters. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Exploring Robust Features for Improving Adversarial Robustness

While deep neural networks (DNNs) have revolutionized many fields, their fragility to carefully designed adversarial attacks impedes the usage of DNNs in safety-critical applications. In this article, we strive to explore the robust features that are not affected by the adversarial perturbations, that is, invariant to the clean image and its adversarial examples (AEs), to improve the model’s adversarial robustness. Specifically, we propose a feature disentanglement model to segregate the robust features from nonrobust features and domain-specific features. Here, the extensive experiments on five widely used datasets with different attacks demonstrate that robust features obtained from our model improve the model’s adversarial robustness compared to the state-of-the-art approaches. Moreover, the trained domain discriminator is able to identify the domain-specific features from the clean images and AEs almost perfectly. This enables AE detection without incurring additional computational costs. With that, we can also specify different classifiers for clean images and AEs, thereby avoiding any drop in clean image accuracy.

97 MATHEMATICS AND COMPUTING↗

A Comparative Study of the Perceptual Sensitivity of Topological Visualizations to Feature Variations

Color maps are a commonly used visualization technique in which data are mapped to optical properties, e.g., color or opacity. Color maps, however, do not explicitly convey structures (e.g., positions and scale of features) within data. Topology-based visualizations reveal and explicitly communicate structures underlying data. Although our understanding of what types of features are captured by topological visualizations is good, our understanding of people's perception of those features is not. Further, this paper evaluates the sensitivity of topology-based isocontour, Reeb graph, and persistence diagram visualizations compared to a reference color map visualization for synthetically generated scalar fields on 2-manifold triangular meshes embedded in 3D. In particular, we built and ran a human-subject study that evaluated the perception of data features characterized by Gaussian signals and measured how effectively each visualization technique portrays variations of data features arising from the position and amplitude variation of a mixture of Gaussians. For positional feature variations, the results showed that only the Reeb graph visualization had high sensitivity. For amplitude feature variations, persistence diagrams and color maps demonstrated the highest sensitivity, whereas isocontours showed only weak sensitivity. These results take an important step toward understanding which topology-based tools are best for various data and task scenarios and their effectiveness in conveying topological variations as compared to conventional color mapping.

97 MATHEMATICS AND COMPUTING↗

Code Artifact for: Clustering Analysis of Commercial Vehicles Using Automatically Extracted Features from Time Series Data [SWR-21-96]

This repository contains data ingestion, feature extraction, and analysis code used in NREL Technical report "Clustering Analysis of Commercial Vehicles Using Automatically Extracted Features from Time Series Data." The code is written in Python. The ETL and feature extraction code must be run in a Spark context. The analysis code can be run without Spark, provided you have pre-computed features in a CSV file. Analysis code related to the NREL Technical Report NREL/TP-2C00-74212. Includes PySpark functions to perform trip segmentation and feature extraction over big time series data in Apache Spark. Includes "domain specific" features such as Aerodynamic Speed (ft/s), Characteristic Acceleration (ft/s2), Percent Below 55 (%), Percent Zero (%), Stops Per Mile, Average Speed (mph), Maximum Speed (mph), and Speed Standard Deviation (mph). Includes Pyspark UDF to compute "domain agnostic" features using the TSFresh library. This software record also includes the analysis notebooks and code to generate the results in the previously mentioned technical report.

Perr-Sauer, Jordan↗

Mining Product Reviews for Important Product Features of Refurbished iPhones

Problem: Remanufacturers want to increase consumer interest in refurbished products, which motivates the need to understand which product features are important to buyers of refurbished products such as mobile phones. Research Questions: This study addresses two questions. First, which product features are most important for buyers of refurbished iPhones? Second, how do those preferences differ from the preferences of buyers of new iPhones? Methods: Online reviews of iPhones are obtained and converted into a document–term matrix. Using this text model, three subsets of features are identified using statistical analysis of frequency of mention: most frequent, average, and least frequent. A logistic regression (LR) model is then used to identify which features are most predictive of whether a review is for a new or refurbished phone. Results: Buyers of refurbished phones mention battery health, screen/display, shell condition, and brand significantly more often than other features. Directly contrasting reviews of refurbished versus new phones shows that shell condition, brand, speaker, and charger are found to be the most predictive product features indicated in reviews for refurbished phones. Of those, the shell condition is significantly more predictive than the others. Implications: The results identify product features that remanufacturers of iPhones can emphasize to increase customer demand.

Anisi, Atefeh↗

Background-Aware 3-D Point Cloud Segmentation With Dynamic Point Feature Aggregation

With the proliferation of LiDAR sensors and 3-D vision cameras, 3-D point cloud analysis has attracted significant attention in recent years. In this article, we propose a novel 3-D point cloud learning network, referred to as dynamic point feature aggregation network (DPFA-Net), by selectively performing the neighborhood feature aggregation (FA) with dynamic pooling and an attention mechanism. DPFA-Net has two variants for semantic segmentation and classification of 3-D point clouds. As the core module of the DPFA-Net, we propose an FA layer, in which features of the dynamic neighborhood of each point are aggregated via a self-attention mechanism. In contrast to other segmentation models, which aggregate features from fixed neighborhoods, our approach can aggregate features from different neighbors in different layers providing a more selective and broader view to the query points and focusing more on the relevant features in a local neighborhood. In addition, to further improve the performance of semantic segmentation, we exploit the background–foreground (BF) information and present two novel approaches, namely, two-stage BF-Net and BF regularization. Experimental results show that the proposed DPFA-Net achieves the state-of-the-art overall accuracy score of 89.22% for semantic segmentation on the Stanford large-scale 3-D Indoor Spaces (S3DIS) dataset and provides consistently satisfactory performance across different tasks of semantic segmentation, part segmentation, and 3-D object classification. Furthermore, our model achieves 93.1% accuracy on the ModelNet40 dataset and provides a mean shape intersection-over-union (IoU) value of 85.5% for part segmentation on the ShapeNet-Part dataset. It is a also computationally more efficient compared to other methods.

3-D↗

Feature Selection Techniques for a Machine Learning Model to Detect Autonomic Dysreflexia

Feature selection plays a crucial role in the development of machine learning algorithms. Understanding the impact of the features on a model, and their physiological relevance can improve the performance. This is particularly helpful in the healthcare domain wherein disease states need to be identified with relatively small quantities of data. Autonomic Dysreflexia (AD) is one such example, wherein mismanagement of this neurological condition could lead to severe consequences for individuals with spinal cord injuries. We explore different methods of feature selection needed to improve the performance of a machine learning model in the detection of the onset of AD. We present different techniques used as well as the ideal metrics using a dataset of thirty-six features extracted from electrocardiograms, skin nerve activity, blood pressure and temperature. The best performing algorithm was a 5-layer neural network with five relevant features, which resulted in 93.4% accuracy in the detection of AD. The techniques in this paper can be applied to a myriad of healthcare datasets allowing forays into deeper exploration and improved machine learning model development. Through critical feature selection, it is possible to design better machine learning algorithms for detection of niche disease states using smaller datasets.

electrocardiography↗

Deep convolutional autoencoders as generic feature extractors in seismological applications

The idea of using a deep autoencoder to encode seismic waveform features and then use them in different seismological applications is appealing. In this paper, we designed tests to evaluate this idea of using autoencoders as feature extractors for different seismological applications, such as event discrimination (i.e., earthquake vs. noise waveforms, earthquake vs. explosion waveforms), and phase picking. These tests involve training an autoencoder, either undercomplete or overcomplete, on a large amount of earthquake waveforms, and then using the trained encoder as a feature extractor with subsequent application layers (either a fully connected layer, or a convolutional layer plus a fully connected layer) to make the decision. By comparing the performance of these newly designed models against the baseline models trained from scratch, we conclude that the autoencoder feature extractor approach may only outperform the baseline under certain conditions, such as when the target problems require features that are similar to the autoencoder encoded features, when a relatively small amount of training data is available, and when certain model structures and training strategies are utilized. The model structure that works best in all these tests is an overcomplete autoencoder with a convolutional layer and a fully connected layer to make the estimation.

58 GEOSCIENCES↗

A systematic method for selecting molecular descriptors as features when training models for predicting physiochemical properties

Machine learning has proven to be a powerful tool for accelerating biofuel development. Although numerous models are available to predict a range of properties using chemical descriptors, there is a trade-off between interpretability and performance. Neural networks provide predictive models with high accuracy at the expense of some interpretability, while simpler models such as linear regression often lack in accuracy. In addition to model architecture, feature selection is also critical for developing interpretable and accurate predictive models. We present a method for systematically selecting molecular descriptor features and developing interpretable machine learning models without sacrificing accuracy. Our method simplifies the process of selecting features by reducing feature multicollinearity and enables discoveries of new relationships between global properties and molecular descriptors. To demonstrate our approach, we developed models for predicting melting point, boiling point, flash point, yield sooting index, and net heat of combustion with the help of the Tree-based Pipeline Optimization Tool (TPOT). For training, we used publicly available experimental data for up to 8351 molecules. Our models accurately predict various molecular properties for organic molecules (mean absolute percent error (MAPE) ranges from 3.3% to 10.5%) and provide a set of features that are well-correlated to the property. This method enables researchers to explore sets of features that significantly contribute to the prediction of the property, offering new scientific insights. To help accelerate early stage biofuel research and development, we also integrated the data and models into a open-source, interactive web tool.

09 BIOMASS FUELS↗

Bioinspired aligned magnetic features in aerogels for humidity sensing

Many natural plant tissues, such as pinecones, utilize structures to transduce changes in humidity into movement or strain. These often employ layered structures with aligned features in order to create differential stress across the material as different portions swell at different rates. Here, by replicating these types of structures in synthetic materials we can create structures that respond to humidity in a similar manner. Taking these aligned features as inspiration, silica aerogels with tailored humidity reactive properties were created. This behavior was achieved by creating aligned features in these aerogels using a Helmholtz coil to produce a uniform magnetic field that aligned ferrofluid droplets into chains and needle like features. These magnetite-based structures were shown to retain their superparamagnetic properties, allowing them to easily be removed from the aerogels, leaving a large number of aligned channels. Aerogels with the removed needles experienced significantly different strains than both unmodified aerogels and aerogels with the aligned features still in place. Aerogels with these tunable aligned features could be used in humidity sensors that can transduce changes in humidity into strain and could be used in applications such as volatile organic compound capture.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Main and Satellite Features in the Ni 2p XPS of NiO

The origin and assignment of the complex main and satellite XPS features of the cations in ionic compounds has been the subject of extensive theoretical studies using different methods. There is agreement that within a molecular orbital model one needs to take into account different types of configurations. Specifically, those where a core electron is removed but no other configuration changes are made and those where in addition to ionization there are also shake or charge transfer changes to the ionic configuration. However, there are strong disagreements about the assignment of XPS features to these configurations. The present work is directed toward resolving the origin of main and satellite features for the Ni 2p XPS of NiO based on ab initio molecular orbital wavefunctions for a cluster model of NiO. A major problem in earlier ab initio XPS studies of ionic compounds has been the use of a common set of orbitals that was not able to properly describe all the ionic configurations that contribute to the full XPS spectra. This is resolved in the present work by using orbitals that are optimized for averages of the occupations of the different configurations that contribute to the XPS. The approach of using State Averaged orbitals is validated through comparisons between different averages and through use of higher order excitations in the wavefunctions for the ionic states. It represents a major extension of our earlier work on the main and satellite features of the Fe 2p XPS of Fe 2 O 3 and proves the reliability and the generality of the assignments of the character and origin of the different features of the XPS obtained with orbitals optimized for State Averages. These molecular orbital methods permit the characterization of the ionic states in terms of the importance of shake excitations and of the coupling of ionization of 2p 1/2 and 2p 3/2 spin-orbit split sub shells. Further, the work lays the foundation for definitive assignments of the character of main and satellite XPS features and points to their origin in the electronic structure of the material.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Overview of the Morphology and Chemistry of Diagenetic Features in the Clay-Rich Glen Torridon Unit of Gale Crater, Mars

The clay-rich Glen Torridon region of Gale crater, Mars, was explored between sols 2300 and 3007. Here, we analyzed the diagenetic features observed by Curiosity, including veins, cements, nodules, and nodular bedrock, using the ChemCam, Mastcam, and Mars Hand Lens Imager instruments. We discovered many diagenetic features in Glen Torridon, including dark-toned iron- and manganese-rich veins, magnesium- and fluorine-rich linear features, Ca-sulfate cemented bedrock, manganese-rich nodules, and iron-rich strata. We have characterized the chemistry and morphology of these features, which are most widespread in the higher stratigraphic members in Glen Torridon, and exhibit a wide range of chemistries. These discoveries are strong evidence for multiple generations of fluids from multiple chemical endmembers that likely underwent redox reactions to form some of these features. In a few cases, we may be able to use mineralogy and chemistry to constrain formation conditions of the diagenetic features. For example, the dark-toned veins likely formed in warmer, highly alkaline, and highly reducing conditions, while manganese-rich nodules likely formed in oxidizing and circumneutral conditions. We also hypothesize that an initial enrichment of soluble elements, including fluorine, occurred during hydrothermal alteration early in Gale crater history to account for elemental enrichment in nodules and veins. The presence of redox-active elements, including Fe and Mn, and elements required for life, including P and S, in these fluids is strong evidence for habitability of Gale crater groundwater. Hydrothermal alteration also has interesting implications for prebiotic chemistry during the earliest stages of the crater’s evolution and early Mars.

58 GEOSCIENCES↗

Classical cosmological collider physics and primordial features

Features in the inflationary landscape can inject extra energies to inflation models and produce on-shell particles with masses much larger than the Hubble scale of inflation. This possibility extends the energy reach of the program of cosmological collider physics, in which signals associated with these particles are generically Boltzmann-suppressed. Here we study the mechanisms of this classical cosmological collider in two categories of primordial features. In the first category, the primordial feature is classical oscillation, which includes the case of coherent oscillation of a massive field and the case of oscillatory features in the inflationary potential. The second category includes any sharp feature in the inflation model. All these classical features can excite unsuppressed quantum modes of other heavy fields which leave observational signatures in primordial non-Gaussianities, including the information about the particle spectra of these heavy degrees of freedom.

79 ASTRONOMY AND ASTROPHYSICS↗

Hyperparameter Optimization and Feature Inclusion in Graph Neural Networks for Spiking Implementation

Graph convolutional networks leverage both graph structures and features on nodes and edges for improved learning performance in comparison with classical machine learning approaches. Spiking neuromorphic computers natively implement network-like computation and have been shown to be successful at implementing graph learning without features. Incorporating graph features brings the challenge of efficient feature representation and balancing the contribution of topology and features in learning. In this work, we present our design of a simulated network of spiking neurons to perform semi-supervised learning on graph data using both the graph structure and the node features. We explore various design choices, present preliminary results, and discuss the opportunities for using neuromorphic computers for this task in the future.

Cong, Guojing↗

A Flang Plugin for Fortran Feature Characterization

As new compute systems are developed, there is still a need to compile and execute codes authored in Fortran on these leading edge systems. In order to achieve this, development of compilers that support the latest hardware is continuously under development. Though the specification of Fortran is extensive, it is helpful to compiler authors to be able to prioritize the development of key features in order to get certain codes deemed important, e.g., applications of interest to leadership computing facilities, executable on leading edge compute systems. Identifying key features though is largely done through querying software experts or users of the Fortran applications of interest, who then manually report what features are and are not present. This exercise can both time consuming and error prone. To automate this process, we present a compiler plugin to Flang, the Fortran frontend for LLVM. This plugin is a tool that operates on the parse tree representation generated by Flang and detects key features based on walking parse tree nodes that correspond to features of interest. We show the result of our tool on four applications, three of which were manually profiled by software experts. We show the discrepancies between our tool and the manual characterization of the three applications, as well as generate a characterization for an application not yet profiled. We intend to open-source our tool in order to invite the community to benefit from the tool and make contributions for other features.

Cabrera, Anthony [ORNL]↗

Evaluating generative networks using Gaussian mixtures of image features

We develop a measure for evaluating the performance of generative networks given two sets of images. A popular performance measure currently used to do this is the Fréchet Inception Distance (FID). However, FID assumes that images featurized using the penultimate layer of Inception follow a Gaussian distribution. This assumption allows FID to be easily computed, since FID uses the 2-Wasserstein distance of two Gaussian distributions fitted to the featurized images. However, we show that Inception features of the ImageNet dataset are not Gaussian; in particular, each marginal is not Gaussian. To remedy this problem, we model the featurized images using Gaussian mixture models (GMMs) and compute the 2-Wasserstein distance restricted to GMMs. We define a performance measure, which we call WaM, on two sets of images by using inception (or another classifier) to featurize the images, estimate two GMMs, and use the restricted 2-Wasserstein distance to compare the GMMs. We experimentally show the advantages of WaM over FID, including how FID is more sensitive than WaM to image perturbations. By modelling the non-Gaussian features obtained from inception as GMMs and using a GMM metric, we can more accurately evaluate generative network performance.

machine learning, genrative adversarial networks↗

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

Ranking Biological Features in Soil-Based Microbial Multi-Omics Data with Integration Modeling

Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.

54 ENVIRONMENTAL SCIENCES↗