Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Co-Occurring Atmospheric Features and Their Contributions to Precipitation Extremes

Object-based identification algorithms for atmospheric features are commonly utilized to attribute global precipitation. This study employs a systematic approach to examine feature co-occurrences and their relationships to mean and extreme precipitation. Four features are identified using existing data sets for atmospheric rivers (ARs), mesoscale convective systems (MCSs), low-pressure systems (LPSs), and fronts (FTs). Often, a single atmospheric phenomenon satisfies the criteria set by multiple feature identification algorithms, yielding an association between precipitation and multiple features. Over the extra-tropics, the number of features attributed to a single event typically increases with precipitation intensity. Over two-thirds of the precipitation is from co-occurring features, with a considerable fraction related to AR-FT co-occurrences. Over the tropics, about one-quarter of precipitation is associated with co-occurring features, with LPS-MCS co-occurrences contributing substantially in monsoon regions. MCSs are the leading single-feature contributors over tropical land and oceans. In the extra-tropics, FTs, ARs, and their co-occurrences account for over half of the total precipitation over oceans. AR-FT-MCS and FT-MCS co-occurrences contribute to extremes (precipitation exceeding the 95th percentile) over both oceans (over 30%) and land (over 20%). Any combination of features involving MCSs shows a larger contribution to high percentiles of precipitation intensity. A case analysis indicates that AR-FT-MCS co-occurrences exhibit convective instability and deep vertical motion, suggesting that the feature trackers and reanalysis are capturing physics relevant to both convective and frontal systems. The results here emphasize the need for simultaneous identifications of multiple features when attributing precipitation to atmospheric phenomena.

54 ENVIRONMENTAL SCIENCES↗

WinnowML: Stable feature selection for maximizing prediction accuracy of time-based system modeling

Online deep learning (ODL) has become an important methodology for modeling time-based performance of computer systems. An open problem is the intelligent selection of features from raw workload traces of computer systems. The best methods are overly sensitive to noisy data, causing frequent feature changes and re-training. Using all available features inflates training time and introduces model artifacts if some features should have been dropped. We present WinnowML, a method for automatically determining the most relevant feature subset for a predictive time-series model. WinnowML combines existing feature ranking algorithms and a history of each feature's ranking to iteratively rank a feature set to lower prediction error and maximize long term relevance. From this ranked feature set, the most relevant and stable subset is selected to train a model. Experimentally, we show how WinnowML can lower a model's mean absolute relative error up to 42% on average compared to the closest performing approach. Additionally, we lower the fluctuation in feature ranking and selection up to 65%. We also demonstrate how to combine WinnowML and a model search tool to provide improvements in performance of up to 14.5% when compared to using all the feature available.

Bel, Oceane MS↗

Performance evaluation of automated data-driven feature extraction and selection methods for practical and scalable building energy consumption prediction models

Here, this study quantifies the impact of automated feature engineering methods (feature extraction and selection) on the quality and accuracy of machine learning models that predict building energy consumption. The case study compares model performance for three main scenarios: baseline (no feature extraction and selection), feature extraction only, and feature extraction combined with feature selection (filter and/or wrapper methods) for fully trained machine learning models for 200 metered/sub-metered energy measurements across 118 real buildings. For consistency, the same machine learning model architecture (a black box deep learning neural network with probabilistic forecast output) was used for all scenarios. Based on results, all feature engineering methods provided noticeable prediction accuracy improvements (e.g., 29%-68% median prediction improvement) compared to baseline scenarios. However, in this application, feature selection methods provide little practical value due to their limited performance gains and high computational cost. Smarter algorithm development supported by better computational environments will be needed before feature selection methods can reliably and efficiently improve predictive model performance.

97 MATHEMATICS AND COMPUTING↗

Improving materials property predictions for graph neural networks with minimal feature engineering *

Graph neural networks (GNNs) have been employed in materials research to predict physical and functional properties, and have achieved superior performance in several application domains over prior machine learning approaches. Recent studies incorporate features of increasing complexity such as Gaussian radial functions, plane wave functions, and angular terms to augment the neural network models, with the expectation that these features are critical for achieving a high performance. Here, we propose a GNN that adopts edge convolution where hidden edge features evolve during training and extensive attention mechanisms, and operates on simple graphs with atoms as nodes and distances between them as edges. As a result, the same model can be used for very different tasks as no other domain-specific features are used. With a model that uses no feature engineering, we achieve performance comparable with state-of-the-art models with elaborate features for formation energy and band gap prediction with standard benchmarks; we achieve even better performance when the dataset size increases. Although some domain-specific datasets still require hand-crafted features to achieve state-of-the-art results, our selected architecture choices greatly reduce the need for elaborate feature engineering and still maintain predictive power in comparison.

42 ENGINEERING↗

Feature selection and causal analysis for microbiome studies in the presence of confounding using standardization

Abstract Background Microbiome studies have uncovered associations between microbes and human, animal, and plant health outcomes. This has led to an interest in developing microbial interventions for treatment of disease and optimization of crop yields which requires identification of microbiome features that impact the outcome in the population of interest. That task is challenging because of the high dimensionality of microbiome data and the confounding that results from the complex and dynamic interactions among host, environment, and microbiome. In the presence of such confounding, variable selection and estimation procedures may have unsatisfactory performance in identifying microbial features with an effect on the outcome. Results In this manuscript, we aim to estimate population-level effects of individual microbiome features while controlling for confounding by a categorical variable. Due to the high dimensionality and confounding-induced correlation between features, we propose feature screening, selection, and estimation conditional on each stratum of the confounder followed by a standardization approach to estimation of population-level effects of individual features. Comprehensive simulation studies demonstrate the advantages of our approach in recovering relevant features. Utilizing a potential-outcomes framework, we outline assumptions required to ascribe causal, rather than associational, interpretations to the identified microbiome effects. We conducted an agricultural study of the rhizosphere microbiome of sorghum in which nitrogen fertilizer application is a confounding variable. In this study, the proposed approach identified microbial taxa that are consistent with biological understanding of potential plant-microbe interactions. Conclusions Standardization enables more accurate identification of individual microbiome features with an effect on the outcome of interest compared to other variable selection and estimation procedures when there is confounding by a categorical variable.

59 BASIC BIOLOGICAL SCIENCES↗

Mapathons versus automated feature extraction: a comparative analysis for strengthening immunization microplanning

Background: Social instability and logistical factors like the displacement of vulnerable populations, the difficulty of accessing these populations, and the lack of geographic information for hard-to-reach areas continue to serve as barriers to global essential immunizations (EI). Microplanning, a population-based, healthcare intervention planning method has begun to leverage geographic information system (GIS) technology and geospatial methods to improve the remote identification and mapping of vulnerable populations to ensure inclusion in outreach and immunization services, when feasible. We compare two methods of accomplishing a remote inventory of building locations to assess their accuracy and similarity to currently employed microplan line-lists in the study area. Methods: The outputs of a crowd-sourced digitization effort, or mapathon, were compared to those of a machine-learning algorithm for digitization, referred to as automatic feature extraction (AFE). The following accuracy assessments were employed to determine the performance of each feature generation method: (1) an agreement analysis of the two methods assessed the occurrence of matches across the two outputs, where agreements were labeled as “befriended” and disagreements as “lonely”; (2) true and false positive percentages of each method were calculated in comparison to satellite imagery; (3) counts of features generated from both the mapathon and AFE were statistically compared to the number of features listed in the microplan line-list for the study area; and (4) population estimates for both feature generation method were determined for every structure identified assuming a total of three households per compound, with each household averaging two adults and 5 children. Results: The mapathon and AFE outputs detected 92,713 and 53,150 features, respectively. A higher proportion (30%) of AFE features were befriended compared with befriended mapathon points (28%). The AFE had a higher true positive rate (90.5%) of identifying structures than the mapathon (84.5%). The difference in the average number of features identified per area between the microplan and mapathon points was larger (t = 3.56) than the microplan and AFE (t = -2.09) (alpha = 0.05). Conclusions: Our findings indicate AFE outputs had higher agreement (i.e., befriended), slightly higher likelihood of correctly identifying a structure, and were more similar to the local microplan line-lists than the mapathon outputs. These findings suggest AFE may be more accurate for identifying structures in high-resolution satellite imagery than mapathons. However, they both had their advantages and the ideal method would utilize both methods in tandem.

59 BASIC BIOLOGICAL SCIENCES↗

ASDFL: An adaptive super‐pixel discriminative feature‐selective learning for vehicle matching

Abstract There are a large number of cameras in modern transportation system that capture numerous vehicle images continuously. Therefore, automatic analysis of these vehicle images is helpful for traffic flow management, criminal investigations and vehicle inspections. Vehicle matching, which aims to determine whether two input images depict an identical vehicle, is one of the core tasks in vehicle analysis. Recent relevant studies have focused on local feature extraction instead of global extraction, since local details can provide crucial cues to distinguish between cars. However, these methods do not select local features; that is, they do not assign weights to local features. Therefore, in this research, we systematically study the vehicle matching task, and present a novel annotation‐free local‐based deep learning method called Adaptive super‐pixel discriminative feature‐selective learning (ASDFL) to address this issue. In ASDFL, vehicle images are segmented into clusters of super‐pixels of similar size by considering the location and colour similarities of pixels without using any component‐level annotation. These super‐pixels are deemed to be the virtual components of vehicles. Moreover, a convolutional neural network is used to extract the deep features of these virtual components. Thereafter, an instance‐specific mask generation module driven by the extracted global features is enhanced to produce a mask to select the most distinctive virtual components of each vehicle image pair in the feature space. Finally, the vehicle matching task is accomplished by classifying the selected virtual component features of each imaged vehicle pair. Extensive experiments on two popular vehicle identification benchmarks demonstrate that our method is 1.57% and 0.8% more accurate than the previous baselines in a vehicle matching task on the VeRi and VehicleID datasets, respectively, which demonstrates the effectiveness of our method.

Qin, Rong↗

Feature Extraction: Improving Remote Sensor Classification of Non-Proliferation

This research focuses on developing algorithms for nuclear non-proliferation detection using remote sensor modeling. To improve the performance of classification models, we implemented a data pipeline with feature extraction. This pipeline takes raw data and transforms it into smaller data points called features that still describe the model. Improving this classification works towards the departments of energy’s missions of ensuring American’s security and prosperity by creating technology that addresses nuclear challenges. To conduct this analysis, we used the Python programming language and some key packages, including tsfresh and TSFEL. Originally tsfresh was selected because it has the most statistical features out of all the packages. Later TSFEL was incorporated due to the additional features it can extract from data, such as temporal and spectral. However, feature extraction becomes challenging in the presence of missing values. In this case, two additional Python packages were added to our workflow, NumPy and pandas, allowing for the feature extraction process to handle unknown values. Our data pipeline was tested on data collected from a simulation that describes the process state of a physical example. The results show the pipeline’s capability to consume and extract a total 17 features from tabular data. Future work includes producing classifications using decision tree-based models such as XGBoost and improving data collection by analyzing feature importance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TempestExtremes v2.1: a community framework for feature detection, tracking, and analysis in large datasets

TempestExtremes (TE) is a multifaceted framework for feature detection, tracking, and scientific analysis of regional or global Earth system datasets on either rectilinear or unstructured/native grids. Version 2.1 of the TE framework now provides extensive support for examining both nodal (i.e., pointwise) and areal features, including tropical and extratropical cyclones, monsoonal lows and depressions, atmospheric rivers, atmospheric blocking, precipitation clusters, and heat waves. Available operations include nodal and areal thresholding, calculations of quantities related to nodal features such as accumulated cyclone energy and azimuthal wind profiles, filtering data based on the characteristics of nodal features, and stereographic compositing. This paper describes the core algorithms (kernels) that have been added to the TE framework since version 1.0, including algorithms for editing pointwise trajectory files, composition of fields around nodal features, generation of areal masks via thresholding and nodal features, and tracking of areal features in time. Several examples are provided of how these kernels can be combined to produce composite algorithms for evaluating and understanding common atmospheric features and their underlying processes. These examples include analyzing the fraction of precipitation from tropical cyclones, compositing meteorological fields around extratropical cyclones, calculating fractional contribution to poleward vapor transport from atmospheric rivers, and building a climatology of atmospheric blocks.

58 GEOSCIENCES↗

Investigation of near-field jet stability of a single-hole injector based on fast X-ray phase contrast imaging and image feature matching

Fuel spray is very effective in controlling the combustion process to improve engine performance and reduce emissions. Understanding the spray unstability during the injection process is of importance to improve the control of spray characteristics and engine operation. In this study, near-field biodiesel jets were recorded using fast X-ray phase contrast imaging and the flow features inside the jet were extracted using Speeded Up Robust Features (SURF) method. Here, the image similarity by feature matching was successfully used to represent the near-field jet stability. Based on the jet stability, an injection process can be divided into five stages: an unstable stage at needle opening, a partially stable transition stage at needle opening, a stable stage, a stable transition stage at needle closing and an unstable stage at needle closing. The ranges of needle lift for five stages were also determined. The jet unstability is highly related to the cavitation formation and gas purging process during needle opening. The variation of stable jet feature is dependent on the needle lift at needle closing. Higher needle lift for the similar jet feature at needle opening indicates a hydraulic delay compared to needle closing. Finally, the possible reasons of jet feature formation and feature detection used on the multi-hole injector are discussed.

33 ADVANCED PROPULSION SYSTEMS↗

Experimental platforms for investigating feature-driven jets for HED mix model validation

High-energy-density (HED) systems, such as inertial confinement fusion (ICF), are susceptible to hydrodynamic instabilities that can significantly affect both experimental results and modeling predictions. Isolated features, such as fill tubes or divots in the capsule, can cause material to jet as a result of the compressive shock exciting the Richtmyer–Meshkov instability, and serve as one of the primary degradation mechanisms in ICF yield. Simulations of feature-driven jets and how they mix require extensive experimental validation, particularly for understanding to what degree the initial size and shape of a feature influence jet dynamics, and how much instability feeds through downstream layers. A better understanding of feature-driven jetting can improve our mix modeling capabilities and increase hydrodynamic simulation accuracy. This manuscript describes a series of experimental platforms fielded by Los Alamos National Laboratory as a part of the Mshock Omega 60 and ModCons Omega EP campaigns to explore feature-driven jetting. These platforms are designed to benchmark jet evolution and growth as a function of initial feature size and shape, investigate jet-layer interactions leading to instability feedthrough, and will be used to characterize jet-jet interactions resulting from clusters of features. In conclusion, preliminary results for both platforms are shown. The ModCons experiments are on-going, and a discussion of future work directions is included.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A bioinspired approach for adaptive solid-solid phase change material coatings with optimized surface features for passive thermal regulation

The necessity to reduce global energy consumption calls for innovative strategies in building thermal management. Passive thermal regulation, particularly through bio-inspired designs, offers a promising avenue by mimicking nature's efficient control of optical properties. This research introduces a novel, climate-responsive coating that integrates optimized bio-inspired surface features with a solid-solid phase change material (SS-PCM) to dynamically manage solar absorptivity without adding additional thickness, enabling both heating and cooling as needed. Drawing on the photonic architectures of the Saharan silver ant and Morpho Didius butterfly, we employed a modeling and multi-objective optimization framework to tailor these surface features. Simulations reveal that surface texture, rather than the intrinsic phase transition of the SS PCM, dominates optical control. Relative to a flat SS PCM coating, optimized isotropic random roughness and broader range features yielded the highest passive heating power increase of about 144 % and 319 % respectively suitable for cold climates. Saharan ant-inspired features enhanced passive cooling for hot climates, achieving a 21.8 % improvement. For moderate climates, Butterfly-wing-inspired surface features provided a balanced enhancement of 19 % for heating and 7 % for cooling. Across all cases, the optimized surface features reduced combined heating and cooling energy demand more effectively than the baseline coating, while preserving material thickness. These findings demonstrate that climate-adaptive, optimized bio-inspired surface features can unlock the full potential of SS PCM coatings, providing a versatile pathway to significant energy savings in buildings and other applications. The methodology establishes a framework for designing next-generation adaptive envelopes that leverage natural photonic principles for high-impact, low-cost thermal regulation.

36 MATERIALS SCIENCE↗

Basin-Scale Structural Features Database

The Basin-Scale Structural Features database provides spatial datasets of faults, fractures, folds, and earthquakes compiled from public, authoritative sources (e.g., U.S. Geological Survey and State Geological Surveys) and aggregated into derivative forms to support subsurface assessments. Recognizing that characterizing basin-scale structural features requires interpreting data that are often ambiguous or lack key information, the source data were evaluated using a knowledge-data framework and geospatial fuzzy logic method (Justman et al., 2020) to represent both measured (observed) and predicted (inferred or potential) structural features as derivative datasets. This workflow employs conceptual models for known structural features and predicted structural features, incorporating geospatial data to estimate potential, even with limited data. The aim is to aid and support an understanding of basin-scale features and identify potential gaps in data and knowledge. As of 4/30/2025, the database includes resources for nine sedimentary basins: Appalachian, Denver, U.S. Gulf Coast, Illinois, Michigan, Permian, Sacramento, San Joquin and Williston. The database is organized by basin and then data category: 1) Faults, fractures, folds, 2) Earthquakes, 3) Topographic, 4) Structural contours and isopachs, 5) Geophysical, and 6) Structural feature density assessment maps.

basin scale↗

Evaluating causal‐based feature selection for fuel property prediction models

Abstract In‐silico screening of novel biofuel molecules based on chemical and fuel properties is a critical first step in the biofuel evaluation process due to the significant volumes of samples required for experimental testing, the destructive nature of engine tests, and the costs associated with bench‐scale synthesis of novel fuels. Predictive models are limited by training sets of few existing measurements, often containing similar classes of molecules that represent just a subset of the potential molecular fuel space. Software tools can be used to generate every possible molecular descriptor for use as input features, but most of these features are largely irrelevant and training models on datasets with higher dimensionality than size tends to yield poor predictive performance. Feature selection has been shown to improve machine learning models, but correlation‐based feature selection fails to provide scientific insight into the underlying mechanisms that determine structure–property relationships. The implementation of causal discovery in feature selection could potentially inform the biofuel design process while also improving model prediction accuracy and robustness to new data. In this study, we investigate the benefits causal‐based feature selection might have on both model performance and identification of key molecular substructures. We found that causal‐based feature selection performed on par with alternative filtration methods, and that a structural causal model provides valuable scientific insights into the relationships between molecular substructures and fuel properties.

Nguyen, Bernard↗

Two-Dimensional Energy Histograms as Features for Machine Learning to Predict Adsorption in Diverse Nanoporous Materials

A major obstacle for machine learning (ML) in chemical science is the lack of physically informed feature representations that provide both accurate prediction and easy interpretability of the ML model. In this work, we describe adsorption systems using novel two-dimensional energy histogram (2D-EH) features, which are obtained from the probe-adsorbent energies and energy gradients at grid points located throughout the adsorbent. The 2D-EH features encode both energetic and structural information of the material and lead to highly accurate ML models (coefficient of determination R2 ~ 0.94–0.99) for predicting single-component adsorption capacity in metal–organic frameworks (MOFs). Here, we consider the adsorption of spherical molecules (Kr and Xe), linear alkanes with a wide range of aspect ratios (ethane, propane, n-butane, and n-hexane), and a branched alkane (2,2-dimethylbutane) over a wide range of temperatures and pressures. The interpretable 2D-EH features enable the ML model to learn the basic physics of adsorption in pores from the training data. We show that these MOF-data-trained ML models are transferrable to different families of amorphous nanoporous materials. We also identify several adsorption systems where capillary condensation occurs, and ML predictions are more challenging. Nevertheless, our 2D-EH features still outperform structural features including those derived from persistent homology. The novel 2D-EH features may help accelerate the discovery and design of advanced nanoporous materials using ML for gas storage and separation in the future.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Atomic-Level Features for Kinetic Monte Carlo Models of Complex Chemistry from Molecular Dynamics Simulations

The high computational cost of evaluating atomic interactions recently motivated the development of computationally inexpensive kinetic models, which can be parametrized from MD simulations of complex chemistry of thousands of species or other processes and accelerate the prediction of the chemical evolution by up to four order of magnitude. Such models go beyond the commonly employed potential energy surface fitting methods in that they are aimed purely at describing kinetic effects. So far, such kinetic models utilize molecular descriptions of reactions and have been constrained to only reproduce molecules previously observed in MD simulations. Therefore, these descriptions fail to predict the reactivity of unobserved molecules, for example in the case of large molecules or solids. In this work, we propose a new approach for the extraction of reaction mechanisms and reaction rates from MD simulations, namely the use of atomic-level features. Using the complex chemical network of hydrocarbon pyrolysis as example, it is demonstrated that kinetic models built using atomic features are able to explore chemical reaction pathways never observed in the MD simulations used to parametrize them, a critical feature to describe rare events. Atomic-level features are shown to construct reaction mechanisms and estimate reaction rates of unknown molecular species from elementary atomic events. Through comparisons of the model ability to extrapolate to longer simulation timescales and different chemical compositions than the ones used for parameterization, it is demonstrated that kinetic models employing atomic features retain the same level of accuracy and transferability as the use of features based on molecular species, while being more compact and parametrized with less data. We also find that atomic features can better describe the formation of large molecules enabling the simultaneous description of small molecules and condensed phases.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Feature selection with distance correlation

Choosing which properties of the data to use as input to multivariate decision algorithms—also known as feature selection—is an important step in solving any problem with machine learning. While there is a clear trend towards training sophisticated deep networks on large numbers of relatively unprocessed inputs (so-called automated feature engineering), for many tasks in physics, sets of theoretically well-motivated and well-understood features already exist. Working with such features can bring many benefits, including greater interpretability, reduced training and run time, and enhanced stability and robustness. We develop a new feature selection method based on distance correlation, and demonstrate its effectiveness on the tasks of boosted top- and W -tagging. Using our method to select features from a set of over 7,000 energy flow polynomials, we show that we can match the performance of much deeper architectures, by using only ten features and two orders-of-magnitude fewer model parameters. Published by the American Physical Society 2024

Astronomy & Astrophysics↗