Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A reproducible study design for the MIMIC-IV in-hospital mortality task

Open, tabular electronic health record (EHR) datasets such as MIMIC-III and MIMIC-IV have become critical resources for developing machine learning (ML) models addressing clinical prediction tasks, including hospital readmission, length of stay, and in-hospital mortality (IHM). While MIMIC-III has benefited from well-established preprocessing pipelines and standardized feature sets, MIMIC-IV remains comparatively challenging to work with because there are no standardized benchmarks to support reproducibility and comparability across studies. To address this limitation, we present a rigorously curated MIMIC-IV custom feature set optimized for IHM prediction, constructed through a reproducible preprocessing pipeline and feature selection strategy.

97 MATHEMATICS AND COMPUTING↗

Single Cell RNA-Seq and Machine Learning Reveal Novel Subpopulations in Low-Grade Inflammatory Monocytes With Unique Regulatory Circuits

Subclinical doses of LPS (SD-LPS) are known to cause low-grade inflammatory activation of monocytes, which could lead to inflammatory diseases including atherosclerosis and metabolic syndrome. Sodium 4-phenylbutyrate is a potential therapeutic compound which can reduce the inflammation caused by SD-LPS. To understand the gene regulatory networks of these processes, we have generated scRNA-seq data from mouse monocytes treated with these compounds and identified 11 novel cell clusters. We have developed a machine learning method to integrate scRNA-seq, ATAC-seq, and binding motifs to characterize gene regulatory networks underlying these cell clusters. Using guided regularized random forest and feature selection, our method achieved high performance and outperformed a traditional enrichment-based method in selecting candidate regulatory genes. Our method is particularly efficient in selecting a few candidate genes to explain observed expression pattern. In particular, among 531 candidate TFs, our method achieves an auROC of 0.961 with only 10 motifs. Finally, we found two novel subpopulations of monocyte cells in response to SD-LPS and we confirmed our analysis using independent flow cytometry experiments. Our results suggest that our new machine learning method can select candidate regulatory genes as potential targets for developing new therapeutics against low grade inflammation.

60 APPLIED LIFE SCIENCES↗

classLog: Logistic regression for the classification of genetic sequences

Introduction Sequencing and phylogenetic classification have become a common task in human and animal diagnostic laboratories. It is routine to sequence pathogens to identify genetic variations of diagnostic significance and to use these data in realtime genomic contact tracing and surveillance. Under this paradigm, unprecedented volumes of data are generated that require rapid analysis to provide meaningful inference. Methods We present a machine learning logistic regression pipeline that can assign classifications to genetic sequence data. The pipeline implements an intuitive and customizable approach to developing a trained prediction model that runs in linear time complexity, generating accurate output rapidly, even with incomplete data. Our approach was benchmarked against porcine respiratory and reproductive syndrome virus (PRRSv) and swine H1 influenza A virus (IAV) datasets. Trained classifiers were tested against sequences and simulated datasets that artificially degraded sequence quality at 0, 10, 20, 30, and 40%. Results When applied to a poor-quality sequence data, the classifier achieved between >85% to 95% accuracy for the PRRSv and the swine H1 IAV HA dataset and this increased to near perfect accuracy when using the full dataset. The model also identifies amino acid positions used to determine genetic clade identity through a feature selection ranking within the model. These positions can be mapped onto a maximum-likelihood phylogenetic tree, allowing for the inference of clade defining mutations. Discussion Our approach is implemented as a python package with code available at https://github.com/flu-crew/classLog .

Zeller, Michael A.↗

Quantized Information in Spectral Cyberspace

The constant-Q Gabor atom is developed for spectral power, information, and uncertainty quantification from time–frequency representations. Stable multiresolution spectral entropy algorithms are constructed with continuous wavelet and Stockwell transforms. The recommended processing and scaling method will depend on the signature of interest, the desired information, and the acceptable levels of uncertainty of signal and noise features. Selected Lamb wave signatures and information spectra from the 2022 Tonga eruption are presented as representative case studies. Resilient transformations from physical to information metrics are provided for sensor-agnostic signal processing, pattern recognition, and machine learning applications.

74 ATOMIC AND MOLECULAR PHYSICS↗

Inter-Domain Fusion for Enhanced Intrusion Detection in Power Systems: An Evidence Theoretic and Meta-Heuristic Approach

False alerts due to misconfigured or compromised intrusion detection systems (IDS) in industrial control system (ICS) networks can lead to severe economic and operational damage. However, research using deep learning to reduce false alerts often requires the physical and cyber sensor data to be trustworthy. Implicit trust is a major problem for artificial intelligence or machine learning (AI/ML) in cyber-physical system (CPS) security, because when these solutions are most urgently needed is also when they are most at risk (e.g., during an attack). To address this, the Inter-Domain Evidence theoretic Approach for Inference (IDEA-I) is proposed that reframes the detection problem as how to make good decisions given uncertainty. Specifically, an evidence theoretic approach leveraging Dempster–Shafer (DS) combination rules and their variants is proposed for reducing false alerts. A multi-hypothesis mass function model is designed that leverages probability scores obtained from supervised-learning classifiers. Using this model, a location-cum-domain-based fusion framework is proposed to evaluate the detector’s performance using disjunctive, conjunctive, and cautious conjunctive rules. The approach is demonstrated in a cyber-physical power system testbed, and the classifiers are trained with datasets from Man-In-The-Middle attack emulation in a large-scale synthetic electric grid. For evaluating the performance, we consider plausibility, belief, pignistic, and general Bayesian theorem-based metrics as decision functions. To improve the performance, a multi-objective-based genetic algorithm is proposed for feature selection considering the decision metrics as the fitness function. Finally, we present a software application to evaluate the DS fusion approaches with different parameters and architectures.

42 ENGINEERING↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Decreasing wind speed extrapolation error via domain-specific feature extraction and selection

Abstract. Model uncertainty is a significant challenge in the wind energy industry and can lead to mischaracterization of millions of dollars' worth of wind resources. Machine learning methods, notably deep artificial neural networks (ANNs), are capable of modeling turbulent and chaotic systems and offer a promising tool to produce high-accuracy wind speed forecasts and extrapolations. This paper uses data collected by profiling Doppler lidars over three field campaigns to investigate the efficacy of using ANNs for wind speed vertical extrapolation in a variety of terrains, and it quantifies the role of domain knowledge in ANN extrapolation accuracy. A series of 11 meteorological parameters (features) are used as ANN inputs, and the resulting output accuracy is compared with that of both standard log-law and power-law extrapolations. It is found that extracted nondimensional inputs, namely turbulence intensity, current wind speed, and previous wind speed, are the features that most reliably improve the ANN's accuracy, providing up to a 65 % and 52 % increase in extrapolation accuracy over log-law and power-law predictions, respectively. The volume of input data is also deemed important for achieving robust results. One test case is analyzed in depth using dimensional and nondimensional features, showing that the feature nondimensionalization drastically improves network accuracy and robustness for sparsely sampled atmospheric cases.

17 WIND ENERGY↗

ML Pipeline (Machine Learning Pipeline) [SWR-22-35]

The Machine Learning Pipeline (ML Pipeline) allows the user to fit a model to predict a dependent variable, y, based on a feature set, X. ML Pipeline gives the ability to automatically generate features with its Feature Engineering function, automatically select the most important features with its Feature Selection function, fit the model with its Fit function, and evaluate the model with its Evaluate function. ML Pipeline's functionality can not only help the user predict values for the desired variable it also gives the user a better understanding of which features are important and how effective the model is. ML Pipeline accomplishes all this, while remaining simple to use and interpret.

Schiek, Andrew↗

Selective Enhancement of Spectroscopic Features by Quantum Optimal Control

Tailored light can be used to steer atomic motions into selected quantum pathways. In optimal control theory (OCT), the target is usually expressed in terms of the molecular wave function, a quantity that is not directly observable in experiment. In this study, we present simulations using OCT that optimize the spectroscopic signal itself. By shaping the optical pump, the x-ray stimulated Raman signal, which occurs solely during the passage through conical intersections, is temporally controlled and amplified by up to 2 orders of magnitude. This enhancement can be crucial in order to bring small coherence-based signatures above the detectable threshold. Our approach is applicable to any signal that depends on the expectation value of a positive definite operator.

74 ATOMIC AND MOLECULAR PHYSICS↗

Effects of multiple simultaneous faults on characteristic fault detection features of a heat pump in cooling mode

Faults in air-cooled vapor compression air-conditioning systems are known to reduce performance, including efficiency, capacity, and lifespan. Their effects have been studied, and fault detection and diagnostic (FDD) methods have been developed as tools for field technicians to install or repair systems, or for monitoring to alert operators to the fault’s presence. Most of this work has focused on faults that occur singly. It is likely that in some systems, multiple faults occur simultaneously, but it is uncertain what effects this may have on diagnostics. Here, this paper describes a laboratory study of a split system air source heat pump in which combinations of two, three, and four simultaneous faults occur. The study includes all combinations of: improper evaporator airflow; overcharge or undercharge of refrigerant; liquid line restrictions; and non-condensable gas in the refrigerant, each at multiple fault intensities. Fault features – those characteristics that can be determined from measurements, for use in diagnostics – are analyzed, and the key fault features are presented. A robust existing method for determining refrigerant charge, the virtual refrigerant charge sensor (VRC) is tested using the multiple fault data, in order to understand how its performance is impacted by the combined faults. The VRC performs well, typically able to correctly determine whether a system is undercharged or overcharged, but the magnitude estimates are impacted. The results suggest that simple subcooling-based methods of charging a system are likely to provide unsatisfactory results when other faults are present.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Species-Selective Detection of Volatile Organic Compounds by Ionic Liquid-Based Electrolyte Using Electrochemical Methods

The detection of volatile organic compounds (VOCs) is an important topic for environmental safety and public health. However, the current commercial VOC detectors suffer from cross-sensitivity and low reproducibility. In this work, we present species-selective detection for VOCs using an electrochemical cell based on ionic liquid (IL) electrolytes with features of high selectivity and reliability. The voltammograms measured with the IL-based electrolyte absorbing different VOCs exhibited species-selective features that were extracted and classified by linear discriminant analysis (LDA). The detection system could identify as many as four types of VOCs, including methanol, ethanol, acetone, formaldehyde, and additional water. A mixture of methanol and formaldehyde was detected as well. The sample required for the VOCs classification system was 50 μL, or 1.164 mmol, on average. The response time for each VOC measurement is as fast as 24 s. The volume of VOCs such as formaldehyde in solution could also be quantified by LDA and electrochemical impedance spectroscopy techniques, respectively. The system showed a tunable detection range for 1.6 and 16% (w/v) CH 2 O solution by adjusting the composition of the electrolyte. The limit of detection was as low as 1 μL. For the 1.6% CH 2 O solution, the linearity calibration range was determined to be from 5.30 to 53.00 μmol with a limit of detection at 0.53 μmol. The mechanisms for VOCs determination and quantification are also thoroughly discussed. Finally, it is expected that this work could provide a new insight into the concept of electrochemical detection of VOCs with machine learning analysis and be applied to both VOCs gas monitoring and fluid detection.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genomic Features and Pervasive Negative Selection in Rhodanobacter Strains Isolated from Nitrate and Heavy Metal Contaminated Aquifer

Despite the dominance of Rhodanobacter species in the subsurface of the contaminated Oak Ridge Reservation (ORR) site, very little is known about the mechanisms underlying their adaptions to the various stressors present at ORR. Recently, multiple Rhodanobacter strains have been isolated from the ORR groundwater samples from several wells with varying geochemical properties.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of Differential Equations by Dynamics-Guided Weighted Weak Form with Voting

In the identification of differential equations from data, significant progresses have been made with the weak/integral formulation. In this paper, we explore the direction of finding more efficient and robust test functions adaptively given the observed data. While this is a difficult task, we propose weighting a collection of localized test functions for better identification of differential equations from a single trajectory of noisy observations on the differential equation. We find that using high dynamic regions is effective in finding the equation as well as the coefficients, and propose a dynamics indicator per differential term and weight the weak form accordingly. For stable identification against noise, we further introduce a voting strategy to identify the active features from an ensemble of recovered results by selecting the features that frequently occur in different weighting of test functions. Systematic numerical experiments are provided to demonstrate the robustness of our method.

97 MATHEMATICS AND COMPUTING↗

Physiochemical Machine Learning Models Predict Operational Lifetimes of CH3NH3PbI3 Perovskite Solar Cells

Halide perovskites are promising photovoltaic (PV) materials with the potential to lower the cost of electricity and greatly expand the penetration of PV if they can demonstrate long-term stability under illumination in the presence of moisture and oxygen. The solar cell service lifetime as quantified by the T80 (the time required for the power conversion efficiency to drop to 80% of its starting value) is a useful metric to assess stability. The T80 for utility, commercial, or residential PV systems needs to be several decades in order to yield low-cost electricity, and thus it is not practical to directly measure the T80. It would be useful if T80 could be predicted from the initial dynamics of a solar cell’s performance, but until now no models have been developed to forecast T80. In this work, we report the development of machine learning models to predict T80 of ITO/NiOx/CH3NH3PbI3/C60/BCP/Ag solar cells operating at maximum power point under 1-sun equivalent photon flux in air at varying temperatures and relative humidities. Efficiency losses are driven by short-circuit current and fill factor, indicating that chemical decomposition of the perovskite is a major contributor to degradation. Spatial patterns evident from in situ dark field optical microscopy suggest that the electric field gradient at device edges plays a significant role in perovskite decomposition, along with photochemical reactions with O2 and H2O. Models are trained using a menu of features from three distinct categories: (i) features based on measurements of the initial rates of change of device parameters, (ii) features based on the ambient conditions during operation (temperature, & partial pressure of H2O), and (iii) features based on underlying physics and chemistry. We show that a theory-based physiochemical feature derived from a model of the chemical reaction kinetics of the rate of degradation of the CH3NH3PbI3 is particularly valuable for prediction. This physiochemical feature was selected as the first or second most dominant feature in the best performing models. With a dataset consisting of 45 accelerated degradation experiments with T80 that range over a factor of 30, the model predicts T80 with an accuracy of about 40% (|predicted T80 - observed T80| / observed T80) on samples not used in training. This hybrid ML approach should be effective when applied to other compositions, device architectures, and advanced packaging schemes.

14 SOLAR ENERGY↗

Physiochemical Machine Learning Models Predict Operational Lifetimes of CH3NH3PbI3 Perovskite Solar Cells

Halide perovskites are promising photovoltaic (PV) materials with the potential to lower the cost of electricity and greatly expand the penetration of PV if they can demonstrate long-term stability under illumination in the presence of moisture and oxygen. The solar cell service lifetime as quantified by the T80 (the time required for the power conversion efficiency to drop to 80% of its starting value) is a useful metric to assess stability. The T80 for utility, commercial, or residential PV systems needs to be several decades in order to yield low-cost electricity, and thus it is not practical to directly measure the T80. It would be useful if T80 could be predicted from the initial dynamics of a solar cell’s performance, but until now no models have been developed to forecast T80. In this work, we report the development of machine learning models to predict T80 of ITO/NiOx/CH3NH3PbI3/C60/BCP/Ag solar cells operating at maximum power point under 1-sun equivalent photon flux in air at varying temperatures and relative humidities. Efficiency losses are driven by short-circuit current and fill factor, indicating that chemical decomposition of the perovskite is a major contributor to degradation. Spatial patterns evident from in situ dark field optical microscopy suggest that the electric field gradient at device edges plays a significant role in perovskite decomposition, along with photochemical reactions with O2 and H2O. Models are trained using a menu of features from three distinct categories: (i) features based on measurements of the initial rates of change of device parameters, (ii) features based on the ambient conditions during operation (temperature, & partial pressure of H2O), and (iii) features based on underlying physics and chemistry. We show that a theory-based physiochemical feature derived from a model of the chemical reaction kinetics of the rate of degradation of the CH3NH3PbI3 is particularly valuable for prediction. This physiochemical feature was selected as the first or second most dominant feature in the best performing models. With a dataset consisting of 45 accelerated degradation experiments with T80 that range over a factor of 30, the model predicts T80 with an accuracy of about 40% (|predicted T80 - observed T80| / observed T80) on samples not used in training. This hybrid ML approach should be effective when applied to other compositions, device architectures, and advanced packaging schemes.

14 SOLAR ENERGY↗

Collective Risk Ranking of Highway Segments on the Basis of Severity-Weighted Crash Rates

This study is intended to focus on the major factors affecting traffic crash rates and severity levels, in addition to identifying crash-prone locations (i.e., black spots) based on the two indicators. The available crash data for different road segments used for the analysis were obtained from the Washington state database provided by the Highway Safety Information System (HSIS) for the years 2006 to 2011. A Random Forest (RF) classifier was used to predict the outcome level of crash severity, while crash rates were predicted by applying RF regressor. Certain features were selected for each model besides the abstraction of new features to check if there are unobserved correlations affecting the independent variables, such as accounting for the number and weight of crashes within 1 km2 area by implementing the Getis-Ord Gi∗ index. Moreover, to calculate the collective risk (CR) score, crash rates were adjusted to incorporate crash severity weights (cost per severity type) and regression-to-the-mean (RTM) bias via Empirical Bayes (EB) method. Finally, segments were ranked according to their CR score.

Li, Dawei↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bia. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Ref. [2]. The prerequisite for applying machine learning techniques is casting the metadata into a format that can be parsed by the algorithm. This step might seem trivial but requires to find a unique language where metadata that carry the same physics meaning across several experiments must have the same identifier. One example is, for instance, the neutron detector. As seen in Figure 1, the machine learning code identified the use of 6 Li detectors as being related to bias in some datasets of the AIACHNE 252 Cf PFNS experimental database. In fact, here are several experiments that used neutron detectors containing 6Li in the database, for instance for the example below. EXFOR format has a unique keywords describing detectors such as “SCIN” or “GLASD”. One may think that these keywords are already sufficient descriptors for ML to uniquely find an issue. However, “SCIN” (used for [3, 4]) and “GLASD” (used for [5]) fail to inform the algorithm what is the active material in the detector. And, the key common issue leading to bias in 252 Cf related to neutron detectors is not whether it is a glass detector or a scintillator. No, the issue is that 6 Li was within both detector types and that even small mistakes in the detector response functions around approximately 200 keV are amplified by the 6 Li(n,α) resonance there leading to bias in data as highlighted in Fig. 1 and Ref. [1]. Hence, the features describing the neutron detector must call out the active material in the detector, rather than the existing EXFOR detector keyword, that the ML algorithm can find physically meaningful features related to bias. The AIACHNE team used a precursor of the WPEC (Working Party on International Nuclear Data Evaluation Co-operation) SG(Subgroup)-50 format to store the metadata for the ML analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗