Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A new μ -high energy resolution fluorescence detection microprobe imaging spectrometer at the Stanford Synchrotron Radiation Lightsource beamline 6-2

In this report we describe a new synchrotron X-ray Fluorescence (XRF) imaging instrument with an integrated High Energy Fluorescence Detection X-ray Absorption Spectroscopy (HERFD-XAS) spectrometer at the Stanford Synchrotron Radiation Lightsource at beamline 6-2. The X-ray beam size on the sample can be defined via a range of pinhole apertures or focusing optics. XRF imaging is performed using a continuous rapid scan system with sample stages covering a travel range of 250 × 200 mm 2 , allowing for multiple samples and/or large samples to be mounted. The HERFD spectrometer is a Johann-type with seven spherically bent 100 mm diameter crystals arranged on intersecting Rowland circles of 1 m diameter with a total solid angle of about 0.44% of 4π sr. A wide range of emission lines can be studied with the available Bragg angle range of ~64.5°–82.6°. With this instrument, elements in a sample can be rapidly mapped via XRF and then selected features targeted for HERFD-XAS analysis. Furthermore, utilizing the higher spectral resolution of HERFD for XRF imaging provides better separation of interfering emission lines, and it can be used to select a much narrower emission bandwidth, resulting in increased image contrast for imaging specific element species, i.e., sparse excitation energy XAS imaging. This combination of features and characteristics provides a highly adaptable and valuable tool in the study of a wide range of materials.

47 OTHER INSTRUMENTATION↗

Prediction of Self-Diffusion in Binary Fluid Mixtures Using Artificial Neural Networks

Artificial neural networks (ANNs) were developed to accurately predict the self-diffusion constants for individual components in binary fluid mixtures. The ANNs were tested on an experimental database of 4328 self-diffusion constants from 131 mixtures containing 75 unique compounds. The presence of strong hydrogen bonding molecules may lead to clustering or dimerization resulting in non-linear diffusive behavior. To address this, self- and binary association energies were calculated for each molecule and mixture to provide information on intermolecular interaction strength and were used as input features to the ANN. An accurate, generalized ANN model was developed with an overall average absolute deviation of 4.1%. Forward input feature selection reveals the importance of critical properties and self-association energies along with other fluid properties. Additional ANNs were developed with subsets of the full input feature set to further investigate the impact of various properties on model performance. The results from two specific mixtures are discussed in additional detail: one providing an example of strong hydrogen bonding and the other an example of extreme pressure changes, with the ANN models predicting self-diffusion well in both cases.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Combining artificial intelligence and physics-based modeling to directly assess atomic site stabilities: from sub-nanometer clusters to extended surfaces

The performance of functional materials is dictated by chemical and structural properties of individual atomic sites. In catalysts, for instance, the thermodynamic stability of constituting atomic sites is a key descriptor from which more complex properties, such as molecular adsorption energies and reaction rates, can be derived. In this study, we present a widely applicable machine learning (ML) approach to instantaneously compute the stability of individual atomic sites in structurally and electronically complex nano-materials. Conventionally, we determine such site stabilities using computationally intensive first-principles calculations. With our approach, we predict the stability of atomic sites in sub-nanometer metal clusters of 3–55 atoms with mean absolute errors in the range of 0.11–0.14 eV. To extract physical insights from the ML model, we introduce a genetic algorithm (GA) for feature selection. This algorithm distills the key structural and chemical properties governing the stability of atomic sites in size-selected nanoparticles, allowing for physical interpretability of the models and revealing structure–property relationships. The results of the GA are generally model and materials specific. In the limit of large nanoparticles, the GA identifies features consistent with physics-based models for metal–metal interactions. By combining the ML model with the physics-based model, we predict atomic site stabilities in real time for structures ranging from sub-nanometer metal clusters (3–55 atom) to larger nanoparticles (147 to 309 atoms) to extended surfaces using a physically interpretable framework. Finally, we present a proof of principle showcasing how our approach can determine stable and active nanocatalysts across a generic materials space of structure and composition.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems

Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.We address such issues by building an end-to-end machine learn- ing (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and gener- ative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.

Paul, Arnab↗

Support Vector Machines for Classification of Direct Energy Deposition Standoff Distance for Improved Process Control

A critical factor in the implementation of direct energy deposition is the ability to maintain the standoff distance between the nozzle and the build surface, as this influences powder capture efficiency and overall part quality. Due to process-related variations, layer height may vary, causing unintended variation in standoff distance and poor build quality. While prior work has utilized contact probing to qualify standoff distance during processing, in situ methods for qualification of standoff distance are of major interest. The present work seeks to understand efficacy of image-based methods for classifying standoff distance variation in real-time using support vector machines (SVMs). It was hypothesized that the size of the melt pool and the amount of spatter will have significant correlations with deviations in the standoff distance; thus, SVMs were used on a dataset that is comprised of morphological features of melt pool size and image entropy. The SVM model was used to classify melt pool images into categories according to standoff distance variation from nominal. K-folds cross validation was used to find the optimal hyperparameters for the SVM model. To understand the impact of the selected features on the classification performance and inference speed, multiple models were trained with differing numbers of included features. Results for classification score, inference time, and image preprocessing/feature extraction from these data are reported. The present results show that the SVM model was able to predict the standoff distance classification with an accuracy of 97 percent and a speed of 0.122 s per image, making it a viable solution for real-time control of standoff distance.

Klesmith, Zoe↗

Time‐Dependent Cation Selectivity of Titanium Carbide MXene in Aqueous Solution

Abstract Electrochemical ion separation is a promising technology to recover valuable ionic species from water. Pseudocapacitive materials, especially 2D materials, are up‐and‐coming electrodes for electrochemical ion separation. For implementation, it is essential to understand the interplay of the intrinsic preference of a specific ion (by charge/size), kinetic ion preference (by mobility), and crystal structure changes. Ti 3 C 2 T z MXene is chosen here to investigate its selective behavior toward alkali and alkaline earth cations. Utilizing an online inductively coupled plasma system, it is found that Ti 3 C 2 T z shows a time‐dependent selectivity feature. In the early stage of charging (up to about 50 min), K + is preferred, while ultimately Ca 2+ and Mg 2+ uptake dominate; this unique phenomenon is related to dehydration energy barriers and the ion exchange effect between divalent and monovalent cations. Given the wide variety of MXenes, this work opens the door to a new avenue where selective ion‐separation with MXene can be further engineered and optimized.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep Learning for In-Situ Layer Quality Monitoring during Laser-Based Directed Energy Deposition (LB-DED) Additive Manufacturing Process

Defects are a leading issue for the rejection of parts manufactured through the Directed Energy Deposition (DED) Additive Manufacturing (AM) process. In an attempt to illuminate and advance in situ quality monitoring and control of workpieces, we present an innovative data-driven method that synchronously collects sensing data and AM process parameters with a low sampling rate during the DED process. The proposed data-driven technique determines the important influences that individual printing parameters and sensing features have on prediction at the inter-layer qualification to perform feature selection. Three Machine Learning (ML) algorithms including Random Forest (RF), Support Vector Machine (SVM), and Convolutional Neural Network (CNN) are used. During post-production, a threshold is applied to detect low-density occurrences such as porosity sizes and quantities from CT scans that render individual layers acceptable or unacceptable. This information is fed to the ML models for training. Training/testing are completed offline on samples deemed “high-quality” and “low-quality”, utilizing only features recorded from the build process. CNN results show that the classification of acceptable/unacceptable layers can reach between 90% accuracy while training/testing on a “high-quality” sample and dip to 65% accuracy when trained/tested on “low-quality”/“high-quality” (respectively), indicating over-fitting but showing CNN as a promising inter-layer classifier.

36 MATERIALS SCIENCE↗

Light-Duty Vehicle Trip Classification Using One-Class Novelty Detection and Exhaustive Feature Extraction

Travel mode classification within travel survey data sets, especially light-duty vehicle (LDV) trips, is foundational, though nontrivial, to emerging mobility systems, travel behavior analysis, and fuel consumption estimation. Current travel mode detection approaches require well-sampled and balanced data sets with ground truth travel mode labels. The detection approaches are rarely applied and validated on large-scale, real-world data sets, which may not satisfy the dataset requirements. This work proposes an LDV trip detection model as a supplement to current travel mode detection methods, for the case when the training set is highly (and/or completely) unbalanced, to the extent that classical machine-learning approaches become difficult or impossible to deploy. The proposed model uses a novelty detection technique - one-class support vector machines (OCSVMs) - and a novel exhaustive feature extraction (EFE) technique on continuous time series data (i.e., Global Positioning System [GPS] speed profiles) for single-mode trip trajectories. Training and validation of the model are conducted on a large-scale, real-world data set. The proposed method accurately identifies LDV trips from a broad set of multimodal trips by leveraging a wealth of preexisting in-vehicle GPS travel data. Additional sensitivity analysis sheds light on the optimal training size and feature selection, which will benefit applications limited by highly imbalanced data. The paper also discusses performance comparison with regular machine-learning approaches, the model's robustness, and the potential to extend the proposed model to multi-modal trip prediction.

33 ADVANCED PROPULSION SYSTEMS↗

Chatter detection in simulated machining data: a simple refined approach to vibration data

Vibration monitoring is a critical aspect of assessing the health and performance of machinery and industrial processes. This study explores the application of machine learning techniques, specifically the Random Forest (RF) classification model, to predict and classify chatter—a detrimental self-excited vibration phenomenon—during machining operations. While sophisticated methods have been employed to address chatter, this research investigates the efficacy of a novel approach to an RF model. The study leverages simulated vibration data, bypassing resource-intensive real-world data collection, to develop a versatile chatter detection model applicable across diverse machining configurations. The feature extraction process combines time-series features and Fast Fourier Transform (FFT) data features, streamlining the model while addressing challenges posed by feature selection. By focusing on the RF model’s simplicity and efficiency, this research advances chatter detection techniques, offering a practical tool with improved generalizability, computational efficiency, and ease of interpretation. The study demonstrates that innovation can reside in simplicity, opening avenues for wider applicability and accelerated progress in the machining industry.

42 ENGINEERING↗

Development of an end state vision to implement digital monitoring in nuclear plants

Transitioning from an onsite Maintenance & Diagnostics Center to cloud-based services offers many new opportunities with computing power and storage, but also new challenges in terms of networking and security. This report will cover everything required for that transition including data processing and uploading to cloud services, feature selection, model creation, and result visualization for decision making. Although there are several other cloud-based services (e.g. Amazon Web Services and Google Cloud), this report explores Microsoft Azure to simplify nomenclature and maintain a consistent focus. Many of the services offered by Microsoft Azure are also available in the other cloud-based services, and their differences have been recorded in other literature. The Azure services most important to a nuclear power plant including networking & security, storage & databases, and Artificial Intelligence (AI) are reviewed here. Networking covers all aspects related to communication to Azure resources including security, privacy, and redundancy. Storage & databases includes data storage, upgrading, patching, backups, and monitoring. The AI services allows the user access to the machine learning (ML) techniques developed with Azure including automated ML, anomaly detection, computer vision, and natural language processing. This report summaries the features, capabilities, and challenges when using cloud-based services in a user-friendly manner.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TROPHY: A Topologically Robust Physics-Informed Tracking Framework for Tropical Cyclones

Tropical cyclones (TCs) are among the most destructive weather systems. Realistically and efficiently detecting and tracking TCs are critical for assessing their impacts and risks. In particular, the eye is a signature feature of a mature TC. Therefore, knowing the eyes’ locations and movements is crucial for both operational weather forecasts and climate risk assessments. Recently, a multilevel robustness framework has been introduced to study the critical points of time-varying vector fields. The framework quantifies the robustness (i.e., structural stability) of critical points across varying neighborhoods. By relating the multilevel robustness with critical point tracking, the framework has demonstrated its potential in cyclone tracking. An advantage is that it identifies cyclonic features using only 2D wind vector fields, which is encouraging as most tracking algorithms require multiple dynamic and thermodynamic variables at different altitudes. A disadvantage is that the framework does not scale well computationally for datasets containing a large number of cyclones. Herein this paper introduces a topologically robust physics-informed tracking framework (TROPHY) for TC tracking. The main idea is to integrate physical knowledge of TC to drastically improve the computational efficiency of multilevel robustness framework for large-scale climate datasets. First, during preprocessing, we propose a physics-informed feature selection strategy to filter 90% of critical points that are short-lived and have low stability, thus preserving good candidates for TC tracking. Second, during in-processing, we impose constraints during the multilevel robustness computation to focus only on physics-informed neighborhoods of TCs. We apply TROPHY to 30 years of 2D wind fields from reanalysis data in ERA5 and generate a number of TC tracks. In comparison with the observed tracks, we demonstrate that TROPHY can capture TC characteristics (e.g., frequency, intensity, duration, latitudes with maximum intensity, and genesis) that are comparable to and sometimes even better than a well-validated TC tracking algorithm that requires multiple dynamic and thermodynamic scalar fields.

97 MATHEMATICS AND COMPUTING↗

Calcium–Based Metal–Organic Frameworks and Their Potential Applications

Metal–organic frameworks (MOFs) built on calcium metal (Ca-MOFs) represent a unique subclass of MOFs featuring high stability, low toxicity, and relatively low density. Ca-MOFs show considerable potential for molecular separations, electronic, magnetic, and biomedical applications, although they are not investigated as extensively as transition metal-based MOFs. Compared to MOFs made of other groups of metals, Ca-MOFs may be particularly advantageous for certain applications such as adsorption and storage of light molecules because of their gravimetric benefit, and drug delivery due to their high biocompatibility. This review intends to provide an overview on the recent development of Ca-MOFs, including their synthesis, crystal structures, important properties, and related applications. Here, various synthetic methods and techniques, types of building blocks, structure and porosity features, selected physical properties, and potential uses will be discussed and summarized. Representative examples will be illustrated for each type of important applications with a focus on their structure–property relations.

36 MATERIALS SCIENCE↗

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION↗

Resolving extreme jet substructure

We study the effectiveness of theoretically-motivated high-level jet observables in the extreme context of jets with a large number of hard sub-jets (up to N = 8). Previous studies indicate that high-level observables are powerful, interpretable tools to probe jet substructure for N ≤ 3 hard sub-jets, but that deep neural networks trained on low-level jet constituents match or slightly exceed their performance. We extend this work for up to N = 8 hard sub-jets, using deep particle-flow networks (PFNs) and Transformer based networks to estimate a loose upper bound on the classification performance. A fully-connected neural network operating on a standard set of high-level jet observables, 135 N-subjetiness observables and jet mass, reach classification accuracy of 86.90%, but fall short of the PFN and Transformer models, which reach classification accuracies of 89.19% and 91.27% respectively, suggesting that the constituent networks utilize information not captured by the set of high-level observables. We then identify additional high-level observables which are able to narrow this gap, and utilize LASSO regularization for feature selection to identify and rank the most relevant observables and provide further insights into the learning strategies used by the constituent-based neural networks. The final model contains only 31 high-level observables and is able to match the performance of the PFN and approximate the performance of the Transformer model to within 2%.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Key predictors of soil organic matter vulnerability to mineralization differ with depth at a continental scale

Abstract Soil organic matter (SOM) is the largest terrestrial pool of organic carbon, and potential carbon-climate feedbacks involving SOM decomposition could exacerbate anthropogenic climate change. However, our understanding of the controls on SOM mineralization is still incomplete, and as such, our ability to predict carbon-climate feedbacks is limited. To improve our understanding of controls on SOM decomposition, A and upper B horizon soil samples from 26 National Ecological Observatory Network (NEON) sites spanning the conterminous U.S. were incubated for 52 weeks under conditions representing site-specific mean summer temperature and sample-specific field capacity (−33 kPa) water potential. Cumulative carbon dioxide respired was periodically measured and normalized by soil organic C content to calculate cumulative specific respiration (CSR), a metric of SOM vulnerability to mineralization. The Boruta algorithm, a feature selection algorithm, was used to select important predictors of CSR from 159 variables. A diverse suite of predictors was selected (12 for A horizons, 7 for B horizons) with predictors falling into three categories corresponding to SOM chemistry, reactive Fe and Al phases, and site moisture availability. The relationship between SOM chemistry predictors and CSR was complex, while sites that had greater concentrations of reactive Fe and Al phases or were wetter had lower CSR. Only three predictors were selected for both horizon types, suggesting dominant controls on SOM decomposition differ by horizon. Our findings contribute to the emerging consensus that a broad array of controls regulates SOM decomposition at large scales and highlight the need to consider changing controls with depth.

59 BASIC BIOLOGICAL SCIENCES↗

Explaining and predicting human behavior and social dynamics in simulated virtual worlds: reproducibility, generalizability, and robustness of causal discovery methods

Ground Truth program was designed to evaluate social science modeling approaches using simulation test beds with ground truth intentionally and systematically embedded to understand and model complex Human Domain systems and their dynamics Lazer et al. (Science 369:1060–1062, 2020). Our multidisciplinary team of data scientists, statisticians, experts in Artificial Intelligence (AI) and visual analytics had a unique role on the program to investigate accuracy, reproducibility, generalizability, and robustness of the state-of-the-art (SOTA) causal structure learning approaches applied to fully observed and sampled simulated data across virtual worlds. In addition, we analyzed the feasibility of using machine learning models to predict future social behavior with and without causal knowledge explicitly embedded. In this paper, we first present our causal modeling approach to discover the causal structure of four virtual worlds produced by the simulation teams—Urban Life, Financial Governance, Disaster and Geopolitical Conflict. Our approach adapts the state-of-the-art causal discovery (including ensemble models), machine learning, data analytics, and visualization techniques to allow a human-machine team to reverse-engineer the true causal relations from sampled and fully observed data. We next present our reproducibility analysis of two research methods team’s performance using a range of causal discovery models applied to both sampled and fully observed data, and analyze their effectiveness and limitations. We further investigate the generalizability and robustness to sampling of the SOTA causal discovery approaches on additional simulated datasets with known ground truth. Our results reveal the limitations of existing causal modeling approaches when applied to large-scale, noisy, high-dimensional data with unobserved variables and unknown relationships between them. We show that the SOTA causal models explored in our experiments are not designed to take advantage from vasts amounts of data and have difficulty recovering ground truth when latent confounders are present; they do not generalize well across simulation scenarios and are not robust to sampling; they are vulnerable to data and modeling assumptions, and therefore, the results are hard to reproduce. Finally, when we outline lessons learned and provide recommendations to improve models for causal discovery and prediction of human social behavior from observational data, we highlight the importance of learning data to knowledge representations or transformations to improve causal discovery and describe the benefit of causal feature selection for predictive and prescriptive modeling.

97 MATHEMATICS AND COMPUTING↗

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka↗

Sensor impact evaluation and verification for fault detection and diagnostics in building energy systems: A review

Sensors are the key information source for fault detection and diagnostics (FDD) in buildings. However, sensors are often not properly designed, installed, calibrated, located, and maintained, which negatively impacts FDD performance. Several sensor-related FDD topics have been widely studied, covering a wide range of fault types and applications. However, it is difficult to get a clear picture of the technical development of sensor-related topics in FDD. A systematic review of sensor topics is needed to summarize the existing research in a logical way, draw conclusions on the current development, and predict the future development of sensors in building FDD. To address this gap, we conducted a comprehensive literature review of more than 100 FDD-sensor-related papers. In this article, we subdivide the FDD tasks into building-level, system-level, and component-level FDD, and review sensor-related topics in each category. Our major conclusions are: (a) current data-driven FDD research focuses more on FDD algorithms than sensors, (b) sensor “hardware” research topics are less studied than sensor “software” topics, (c) very few papers focus on sensor engineering as an integral aspect of FDD development, and (d) some important sensor topics, such as sensor cost-effectiveness and sensor schema/layout/location, are not well studied. Finally, we discuss the need for a systematic framework of FDD sensors and models to integrate sensor design/selection, sensor data analysis/mining, feature selection, physics-based or data-driven algorithm development, sensor fault detection, sensor calibration, and sensor maintenance. Finally, expert interviews are conducted to validate the above findings and conclusions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗