Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Multispectral data acquisition and classification - Statistical models for system design

In this paper we relate the statistical processes that are involved in multispectral data acquisition and classification to a simple radiometric model of the earth surface and atmosphere. If generalized, these formulations could provide an analytical link between the steadily improving models of our environment and the performance characteristics of rapidly advancing device technology. This link is needed to bring system analysis tools to the task of optimizing remote sensing and (real-time) signal processing systems as a function of target and atmospheric properties, remote sensor spectral bands and system topology (e.g., image-plane processing), radiometric sensitivity and calibration accuracy, compensation for imaging conditions (e.g., atmospheric effects), and classification rates and errors.

Huck, F. O.↗

Multispectral data acquisition and classification - Computer modeling for smart sensor design

In this paper a model of the processes involved in multispectral remote sensing and data classification is developed as a tool for designing and evaluating smart sensors. The model has both stochastic and deterministic elements and accounts for solar radiation, atmospheric radiative transfer, surface reflectance, sensor spectral reponses, and classification algorithms. Preliminary results are presented which indicate the validity and usefulness of this approach. Future capabilities of smart sensors will ultimately be limited by the accuracy with which multispectral remote sensing processes and their error sources can be computationally modeled.

Park, S. K.↗

Epidural anesthesia needle guidance by forward-view endoscopic optical coherence tomography and deep learning

Epidural anesthesia requires injection of anesthetic into the epidural space in the spine. Accurate placement of the epidural needle is a major challenge. To address this, we developed a forward-view endoscopic optical coherence tomography (OCT) system for real-time imaging of the tissue in front of the needle tip during the puncture. We tested this OCT system in porcine backbones and developed a set of deep learning models to automatically process the imaging data for needle localization. A series of binary classification models were developed to recognize the five layers of the backbone, including fat, interspinous ligament, ligamentum flavum, epidural space, and spinal cord. The classification models provided an average classification accuracy of 96.65%. During puncture, it is important to maintain a safe distance between the needle tip and the dura mater. Regression models were developed to estimate that distance based on the OCT imaging data. Based on the Inception architecture, our models achieved a mean absolute percentage error of 3.05% ± 0.55%. Overall, our results validated the technical feasibility of using this novel imaging strategy to automatically recognize different tissue structures and measure the distances ahead of the needle tip during the epidural needle placement.

60 APPLIED LIFE SCIENCES↗

Data-driven search for promising intercalating ions and layered materials for metal-ion batteries

The rise in demand for lithium-ion batteries has led to a large-scale search for electrode materials and intercalating ion species to meet the demands of next-generation energy technologies. Recent efforts largely focus on searching for cathodes that can accommodate large amounts of intercalating ions, but similar work on anodes is relatively limited. This study utilizes machine learning methods to find alternative two-dimensional (2D) materials and intercalating ions beyond Li for metal-ion batteries with high-power efficiencies. The approach first uses density functional theory (DFT) calculations to estimate the theoretical capacities and voltages of various metal ions on 2D materials. The DFT-generated data also provide insights into the local structural accommodation upon ion intercalation on various 2D materials. Significant changes to the lattice can result in irreversible changes to the bonding environments in the anode material, resulting in poor cycling stability. Next, this study develops a binding energy and structural accommodation-based classification model to screen anode materials for next-generation batteries. The classification model selects intercalating ions and 2D material pairs suitable for batteries based on the calculated voltage and volumetric changes in the 2D material upon intercalation. Finally, this study builds a regression model to accurately predict the binding energies of the various intercalating ions on 2D materials. The approach highlights the importance of different elemental and structural features for classification and regression tasks. In conclusion, the insights gained from this study on the role of involved features, such as electronegativities of the constituent ions and the presence of unfilled electronic levels, will help to streamline further studies towards the search for future layered battery materials.

36 MATERIALS SCIENCE↗

Solar forecasting using machine learned cloudiness classification

Methods and systems for predicting irradiance include learning a classification model using unsupervised learning based on historical irradiance data. The classification model is updated using supervised learning based on an association between known cloudiness states and historical weather data. A cloudiness state is predicted based on forecasted weather data. An irradiance is predicted using a regression model associated with the cloudiness state.

Hamann, Hendrik F.↗

Quantifying leaf symptoms of sorghum charcoal rot in images of field‐grown plants using deep neural networks

Abstract Charcoal rot of sorghum (CRS) is a significant disease affecting sorghum crops, with limited genetic resistance available. The causative agent, Macrophomina phaseolina (Tassi) Goid, is a highly destructive fungal pathogen that targets over 500 plant species globally, including essential staple crops. Utilizing field image data for precise detection and quantification of CRS could greatly assist in the prompt identification and management of affected fields and thereby reduce yield losses. The objective of this work was to implement various machine learning algorithms to evaluate their ability to accurately detect and quantify CRS in red‐green‐blue images of sorghum plants exhibiting symptoms of infection. EfficientNet‐B3 and a fully convolutional network emerged as the top‐performing models for image classification and segmentation tasks, respectively. Among the classification models evaluated, EfficientNet‐B3 demonstrated superior performance, achieving an accuracy of 86.97%, a recall rate of 0.71, and an F1 score of 0.73. Of the segmentation models tested, FCN proved to be the most effective, exhibiting a validation accuracy of 97.76%, a recall rate of 0.68, and an F1 score of 0.66. As the size of the image patches increased, both models’ validation scores increased linearly, and their inference time decreased exponentially. This trend could be attributed to larger patches containing more information, improving model performance, and fewer patches reducing the computational load, thus decreasing inference time. The models, in addition to being immediately useful for breeders and growers of sorghum, advance the domain of automated plant phenotyping and may serve as a foundation for drone‐based or other automated field phenotyping efforts. Additionally, the models presented herein can be accessed through a web‐based application where users can easily analyze their own images.

Gonzalez, Emmanuel M.↗

Explaining word embeddings with perfect fidelity: a case study in predicting research impact

The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models can also provide predictions even for words that were not contained in the labelled training data for the classification model, which is important in the context of the ever-evolving research terminology. Although model-agnostic explanation methods, such as Local interpretable model-agnostic explanations, can be applied to explain machine learning classifiers trained on embedding models, these produce results with questionable correspondence to the model. We introduce a new feature importance method, Self-Model Entities Rated (SMER), for logistic regression-based classification models trained on word embeddings. We show that SMER has theoretically perfect fidelity with the explained model, as the average of logits of SMER scores for individual words (SMER explanation) exactly corresponds to the logit of the prediction of the explained model. Quantitative and qualitative evaluation is performed through five diverse experiments conducted on 50,000 research articles (papers) from the CORD-19 corpus. In conclusion, through an AOPC curve analysis, we experimentally demonstrate that SMER produces better explanations than LIME, SHAP and global tree surrogates.

Coarse-grained models↗

A system for verifying models and classification maps by extraction of information from a variety of data sources

Recent updates to a geographical information system (GIS) called VICAR (Video Image Communication and Retrieval)/IBIS are described. The system is designed to handle data from many different formats (vector, raster, tabular) and many different sources (models, radar images, ground truth surveys, optical images). All the data are referenced to a single georeference plane, and average or typical values for parameters defined within a polygonal region are stored in a tabular file, called an info file. The info file format allows tracking of data in time, maintenance of links between component data sets and the georeference image, conversion of pixel values to `actual' values (e.g., radar cross-section, luminance, temperature), graph plotting, data manipulation, generation of training vectors for classification algorithms, and comparison between actual measurements and model predictions (with ground truth data as input).

Norikane, L.↗

Application of Orthogonal Defect Classification for Software Reliability Analysis

The modernization of existing and new nuclear power plants with digital instrumentation and control systems (DI&C) is a recent and highly trending topic. However, there lacks strong consensus on best-estimate reliability methodologies by both the United States (U.S.) Nuclear Regulatory Commission (NRC) and the industry. This has resulted in hesitation for further modernization projects until a more unified methodology is realized. In this work, we develop an approach called Orthogonal-defect Classification for Assessing Software Reliability (ORCAS) to quantify probabilities of various software failure modes in a DI&C system. The method utilizes accepted industry methodologies for software quality assurance that are also verified by experimental or mathematical formulations. In essence, the approach combines a semantic failure classification model with a reliability growth model to predict (and quantify) the potential failure modes of a DI&C software system. The semantic classification model is used to address the question: How do latent defects in software contribute to different software failure root causes? The use of reliability growth models is then used to address the question: Given the connection between latent defects and software failure root causes, how can we quantify the reliability of the software? A case study was conducted on a representative I&C platform (ChibiOS) running a smart sensor acquisition software developed by Virginia Commonwealth University (VCU). The testing and evidence collection guidance in ORCAS was applied, and defects were uncovered in the software. Qualitative evidence, such as condition coverage, was used to gauge the completeness and trustworthiness of the assessment while quantitative evidence was used to determine the software failure probabilities. The reliability of the software was then estimated and compared to existing operational data of the sensor device. It is demonstrated that by using ORCAS, a semantic reasoning framework can be developed to justify software reliability (or unreliability) while still leveraging the strength of the existing methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Application of Orthogonal Defect Classification for Software Reliability Analysis

The modernization of existing and new nuclear power plants with digital instrumentation and control systems (DI&C) is a recent and highly trending topic. However, there lacks strong consensus on best-estimate reliability methodologies by both the United States (U.S.) Nuclear Regulatory Commission (NRC) and the industry. This has resulted in hesitation for further modernization projects until a more unified methodology is realized. In this work, we develop an approach called Orthogonal-defect Classification for Assessing Software Reliability (ORCAS) to quantify probabilities of various software failure modes in a DI&C system. The method utilizes accepted industry methodologies for software quality assurance that are also verified by experimental or mathematical formulations. In essence, the approach combines a semantic failure classification model with a reliability growth model to predict (and quantify) the potential failure modes of a DI&C software system. The semantic classification model is used to address the question: How do latent defects in software contribute to different software failure root causes? The use of reliability growth models is then used to address the question: Given the connection between latent defects and software failure root causes, how can we quantify the reliability of the software? A case study was conducted on a representative I&C platform (ChibiOS) running a smart sensor acquisition software developed by Virginia Commonwealth University (VCU). The testing and evidence collection guidance in ORCAS was applied, and defects were uncovered in the software. Qualitative evidence, such as condition coverage, was used to gauge the completeness and trustworthiness of the assessment while quantitative evidence was used to determine the software failure probabilities. The reliability of the software was then estimated and compared to existing operational data of the sensor device. It is demonstrated that by using ORCAS, a semantic reasoning framework can be developed to justify if the software is reliable (or unreliable) while still leveraging the strength of the existing methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Identifying COVID-19 cases and extracting patient reported symptoms from Reddit using natural language processing

We used social media data from “covid19positive” subreddit, from 03/2020 to 03/2022 to identify COVID-19 cases and extract their reported symptoms automatically using natural language processing (NLP). We trained a Bidirectional Encoder Representations from Transformers classification model with chunking to identify COVID-19 cases; also, we developed a novel QuadArm model, which incorporates Question-answering, dual-corpus expansion, Adaptive rotation clustering, and mapping, to extract symptoms. Our classification model achieved a 91.2% accuracy for the early period (03/2020-05/2020) and was applied to the Delta (07/2021–09/2021) and Omicron (12/2021–03/2022) periods for case identification. We identified 310, 8794, and 12,094 COVID-positive authors in the three periods, respectively. The top five common symptoms extracted in the early period were coughing (57%), fever (55%), loss of sense of smell (41%), headache (40%), and sore throat (40%). During the Delta period, these symptoms remained as the top five symptoms with percent authors reporting symptoms reduced to half or fewer than the early period. During the Omicron period, loss of sense of smell was reported less while sore throat was reported more. Our study demonstrated that NLP can be used to identify COVID-19 cases accurately and extracted symptoms efficiently.

60 APPLIED LIFE SCIENCES↗

Comparison of Expert Vocabulary Usage Patterns Between Mental Health and Nonmental Health Clinicians When Diagnosing Pediatric Anxiety Disorders

Objective: To compare the utilization patterns of expert vocabulary (EVo) in diagnosing pediatric anxiety between mental health and non-mental health clinical notes from electronic health records to understand the role of Evo in informing classification and decision-making in anxiety diagnoses. Study design: We conducted a retrospective study using a cohort less than age 25 from Cincinnati Children's Hospital including 897 685 patients with 61 586 446 notes. We analyzed EVo, collected from mental health clinicians, in both mental and nonmental health notes. We compared classification accuracy using EVo-based patient-level embedding from all clinical notes, mental-health notes, and nonmental health notes for 2 tasks: 1) pre-vs postdiagnosis anxiety patients, and 2) prediagnosis anxiety vs nonanxiety patients. Results: EVo usage was highest in prediagnosis anxiety, lower in nonanxiety, and lowest in post-diagnosis. Classification models using EVo features from all, mental-health, and non-mental health notes showed similar F1 scores for prediagnosis anxiety (0.70 ± 0.2 for 2 categories). For anxiety vs nonanxiety classification, all clinical and nonmental health notes had better F1 scores than mental-health notes (above 0.90 for 3 categories). There was a notable difference in class-wise performance across both tasks. Conclusions: There are significant differences in anxiety EVo use between mental health and nonmental health clinicians. Despite less anxiety-specific terminology, non-mental health notes still captured key aspects of patient presentations, emphasizing the importance of including all clinicians' notes in analysis. EVo's utility for anxiety classification is most effective in prediagnostic phases, suggesting the need for a dedicated diagnostic lexicon and further study before incorporating EVo into classification models.

feature engineering↗

A Taxonomy-Based Approach to Shed Light on the Babel of Mathematical Models for Rice Simulation

For most biophysical domains, differences in model structures are seldom quantified. Here, we used a taxonomy-based approach to characterise thirteen rice models. Classification keys and binary attributes for each key were identified, and models were categorised into five clusters using a binary similarity measure and the unweighted pair-group method with arithmetic mean. Principal component analysis was performed on model outputs at four sites. Results indicated that (i) differences in structure often resulted in similar predictions and (ii) similar structures can lead to large differences in model outputs. User subjectivity during calibration may have hidden expected relationships between model structure and behaviour. This explanation, if confirmed, highlights the need for shared protocols to reduce the degrees of freedom during calibration, and to limit, in turn, the risk that user subjectivity influences model performance.

model parameterisation↗

Generalization of Deep-Learning Models for Classification of Local Distance Earthquakes and Explosions across Various Geologic Settings

Although accurately classifying signals from earthquakes and explosions at local distance (<250 km) remains an important task for seismic network operations, the growing volume of available seismic data presents a challenge for analysts using traditional source discrimination techniques. In recent years, deep-learning models have proven effective at discriminating between low-magnitude earthquakes and explosions measured at local distances, but it is not clear how well these models are capable of generalizing across different geological settings. To address the issue of generalization between regions, we train deep-learning models (convolutional neural networks [CNNs]) on time–frequency representations (scalograms) of three-component earthquake and explosion signals from eight different regions in the continental United States. We explore scenarios where models are trained on data from all regions, individual regions, or all but one region. We find that although CNN models trained on individual regions do not necessarily generalize well across different settings, models trained on multiple regions that include diverse path coverage generalize to new regions, with station-level accuracy of up to 90% or more for data sets from unseen regions. In general, CNN-based discrimination models significantly outperform models based on uncorrected P/S ratio (measured in the 10–18 Hz frequency band), even when CNN models are tested on data from entirely unseen regions.

58 GEOSCIENCES↗

Benchmark Models for Classification of Radiation Type Induced in Immune Cells

NASA Biological and Physical Sciences and the Science Mission Directorate have published a benchmark dataset of mouse immune cells subjected to radiation-induced DNA damage. The dataset comprises ML-ready microscopic imagery of said cells, including labels indicating radiation type and dose. The machine learning team at NASA Interagency Implementation and Advanced Concept Team (IMPACT) created multiple benchmark models. Initially, we conducted a preliminary analysis using thresholding. The algorithm used thresholds on average brightness of the available images to classify them into their respective radiation type. We also tested machine learning approaches. Convolutional Neural Networks (CNN) emerged as the best-performing model. This poster presents the benchmark scores obtained by the models.

Vishal Perekadan↗

Modeling and analysis of several classes of self-oscillating inverters. I - State-plane representations. II - Model extension, classification, and duality relationships

The present investigation is concerned with an important class of power conditioning networks, taking into account self-oscillating dc-to-square-wave transistor inverters. The considered circuits are widely used both as the principal power converting and processing means in many systems and as low-power analog-to-discrete-time converters for controlling the switching of the output-stage semiconductors in a variety of power conditioning systems. Aspects of piecewise-linear modeling are discussed, taking into consideration component models, and an equivalent-circuit model. Questions of singular point analysis and state plane representation are also investigated, giving attention to limit cycles, starting circuits, the region of attraction, a hard oscillator, and a soft oscillator.

Lee, F. C. Y.↗

Understanding oxidation of Fe-Cr-Al alloys through explainable artificial intelligence

Abstract The oxidation resistance of FeCrAl based on alloying composition and oxidizing conditions is predicted using a combinatorial experimental and artificial intelligence approach. A neural network (NN) classification model was trained on the experimental FeCrAl dataset produced at GE Research. Furthermore, using the SHapley Additive exPlanations (SHAP) explainable artificial intelligence (XAI) tool, we explore how the NN can showcase further material insights that are unavailable directly from a black-box model. We report that high Al and Cr content forms protective oxide layer, while Mo in FeCrAl creates thick unprotective oxide scale that is vulnerable to spallation due to thermal expansion. Graphical abstract

Materials Science↗

Detecting Reactive Products in Carbon Capture Polymers with Chemical Shift Anisotropy and Machine Learning

Aminopolymers are attractive sorbents for CO 2 direct air capture applications due to their high density of amine groups, which can readily react with atmospheric levels of CO 2 to form chemisorbed species. The identity of these chemisorbed species and the functional groups that form upon oxidative degradation depends on both material properties and processing conditions, forming a variety of carbonyl-type sites such as ammonium carbamates, bicarbonates, carbonates, carbamic acids, ureas, and amides. 13 C solid-state nuclear magnetic resonance (NMR) is often used to help elucidate the identity of these reacted species, but it is challenging due to the narrow chemical shift range of carbonyl sites. Herein, we demonstrate the application of a two-dimensional (2D) chemical shift anisotropy (CSA) recoupling pulse sequence (ROCSA) to obtain CSA tensor values at each isotropic chemical shift, overcoming limitations of isotropic peak resolution. CSA tensor values describe the local chemical environment and can readily differentiate between the chemisorbed and degradation products. To aid identification, we also developed a k-nearest neighbor (kNN) classification model to distinguish the functional groups via their CSA tensor parameters. This methodology was demonstrated on poly(ethylenimine) in γ-Al 2 O 3 exposed to CO 2 and showed that the chemisorbed products are ammonium carbamate and a mixed carbamate–carbamic acid species. The sample was analyzed again after desorption at 100 °C inducing mild degradation, and the remaining products were strongly bound carbamate and urea species. In conclusion, the combination of 2D CSA measurements coupled with a kNN classification model enhances the ability to accurately identify chemisorbed or degradation products in complex carbon capture materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗