Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Autonomous Detection and Classification of Lunar Minerals Using a Convolutional Neural Network Based Framework for the SUCR DALI Project

NASA’s long-term goal is to deploy humans to the Moon and, from there, advance human exploration to Mars, with Artemis missions as pivotal milestones. Raman spectroscopy can uniquely identify minerals, compounds, water states, and other materials, providing distinctive fingerprints for classification. A Raman instrument has been successfully deployed and utilized on the Mars surface via the Perseverance rover, but has not yet been utilized at the lunar surface The SUCR DALI project is working towards developing a Raman spectroscopy instrument to be applied in various lunar mission concepts, including within the Artemis program. The objective of my research is to assist in the maturation of the proposed SUCR DALI lunar Raman instrument through the development of an autonomous detection and classification model capable of identifying minerals and water states on the Moon’s surface.

Convolutional Neural Networks↗

Fusion of Experiments and Simulations for Real-Time Identification of Pipeline Defects

In this study, we explored fusion of experiments and simulations for real time identification of pipeline defects across physical and non-physical domains. The challenges associated to data processing were addressed and a combined classification models was presented via CNN models. In addition, regression model based on XGBOOST is built to determine the defect location and defect dimension from data-driven features of guided wave signals captured by SMS fiber optic sensor.

deep learning↗

Classification and regression models of audio and vibration signals for machine state monitoring in precision machining systems

Here we present a data-driven method for monitoring machine status in manufacturing processes. Audio and vibration data from precision machining are used for inference in two operating scenarios: (a) variable machine health states (anomaly detection); and (b) settings of machine operation (state estimation). Audio and vibration signals are first processed through Fast Fourier Transform and Principal Component Analysis to extract transformed and informative features. These features are then used in the training of classification and regression models for machine state monitoring. Specifically, three classifiers (K-nearest neighbors, convolutional neural networks and support vector machines) and two regressors (support vector regression and neural network regression) were explored, in terms of their accuracy in machine state prediction. It is shown that the audio and vibration signals are sufficiently rich in information about the machine that 100% state classification accuracy could be accomplished. Data fusion was also explored, showing overall superior accuracy of data-driven regression models.

42 ENGINEERING↗

Physics-guided logistic classification for tool life modeling and process parameter optimization in machining

This paper describes a physics-guided logistic classification method for tool life modeling and process parameter optimization in machining. Tool life is modeled using a classification method since the exact tool life cannot be measured in a typical production environment where tool wear can only be directly measured when the tool is replaced. Here, in this study, laboratory tool wear experiments are used to simulate tool wear data normally collected during part production. Two states are defined: tool not worn (class 0) and tool worn (class 1). The non-linear reduction in tool life with cutting speed is modeled by applying a logarithmic transformation to the inputs for the logistic classification model. A method for interpretability of the logistic model coefficients is provided by comparison with the empirical Taylor tool life model. The method is validated using tool wear experiments for milling. Results show that the physics-guided logistic classification method can predict tool life using limited datasets. A method for pre-process optimization of machining parameters using a probabilistic machining cost model is presented. The proposed method offers a robust and practical approach to tool life modeling and process parameter optimization in a production environment.

Machine learning↗

Pre-training Vision Models for the Classification of Alerts from Wide-field Time-domain Surveys

Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.

79 ASTRONOMY AND ASTROPHYSICS↗

Land Covering Classifications of Boreas Modeling Grid Using AIRSAR Images

Mapping forest types in the boreal ecosystem in an integrated part of any modeling excercise of biogeophysical processes characterizing the interaction of forest with the atmosphere. In this paper, we report the results of the land cover classification of the SAR data acquired during the BOREAS (BOReal Ecosystem Atmospheric Study) intensive field campaigns over the modeling sub-grid of the southern study area in Saskatchewan , Canada. A Bayesian-maximum-a-posteriori classifier has been applied on the NASA/JPL AIRSAR images covering the region during the peak of the growing season in July, 1994.

Atmospheric Study↗

Analytical models and system topologies for remote multispectral data acquisition and classification

Simple analytical models are presented of the radiometric and statistical processes that are involved in multispectral data acquisition and classification. Also presented are basic system topologies which combine remote sensing with data classification. These models and topologies offer a preliminary but systematic step towards the use of computer simulations to analyze remote multispectral data acquisition and classification systems.

Huck, F. O.↗

Logistic classification for tool life modeling in machining

This paper describes the application of logistic classification for tool life modeling and prediction in an industrial setting using shop floor data. Tool life is treated as a classification problem since tool wear can only be measured at the time of tool replacement in a production environment. Laboratory tool wear experiments are used to simulate shop floor wear data by two states: not worn (class 0); and worn (class 1). To incorporate non-linearity in logistic classification, a log-transformation of input features is performed. The logistic classification approach, results, and interpretability of the logistic model are presented.

Karandikar, Jaydeep↗

The Dark Machines Anomaly Score Challenge: Benchmark Data and Model Independent Event Classification for the Large Hadron Collider

We describe the outcome of a data challenge conducted as part of the Dark Machines (https://www.darkmachines.org) initiative and the Les Houches 2019 workshop on Physics at TeV colliders. The challenged aims to detect signals of new physics at the Large Hadron Collider (LHC) using unsupervised machine learning algorithms. First, we propose how an anomaly score could be implemented to define model-independent signal regions in LHC searches. We define and describe a large benchmark dataset, consisting of >1 billion simulated LHC events corresponding to 10\, fb^{-1} 10 f b − 1 of proton-proton collisions at a center-of-mass energy of 13 TeV. We then review a wide range of anomaly detection and density estimation algorithms, developed in the context of the data challenge, and we measure their performance in a set of realistic analysis environments. We draw a number of useful conclusions that will aid the development of unsupervised new physics searches during the third run of the LHC, and provide our benchmark dataset for future studies at https://www.phenoMLdata.org. Code to reproduce the analysis is provided at https://github.com/bostdiek/DarkMachines-UnsupervisedChallenge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Using ensembles and distillation to optimize the deployment of deep learning models for the classification of electronic cancer pathology reports

One of the goals of the Surveillance, Epidemiology, and End Results (SEER) program is to estimate incidence, prevalence, and mortality of all cancers. To that end, cancer registries across the country maintain a massive database of cancer pathology reports which contain rich information to understand cancer trends. However, these reports are stored in the form of unstructured text, and human annotators are required to read and extract relevant information. In this article, we show that existing deep learning models for automating information extraction from cancer pathology reports can be significantly improved by using ensemble model distillation. We found that by training multiple predictive models and transferring their knowledge to a single, low-resource model, we can reduce the number of highly confident wrong predictions. Our results show that our implemented methods could save 1000s of manual annotation hours.

60 APPLIED LIFE SCIENCES↗

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗

Apollo 16 rocks - Classification and petrogenetic model

The Apollo 16 rocks include cataclastic anorthosites, two varieties of unequilibrated breccia, two varieties of partly to fully equilibrated breccia, and a sequence of partially melted breccias. The latter, which dominate the Apollo 16 collection, include glass, divitrified glass, mesostasis-olivine-plagioclase rock, mesostasis-rich basalt, basalt, and poikilitic rocks. All sequence members contain vesicles and relics of plagioclase, olivine, pink spinel, and lithic fragments. Their equilibrated matrices define a series from glass, to a plagioclase-olivine-mesostasis assemblage displaying spherulitic and skeletal shaped crystals, through a plagioclase-pyroxene-olivine assemblage displaying euhedral shaped crystals. Such data suggest that the sequence lithologies were derived from breccias or soils that were partially melted in an impact event.

Warner, J. L.↗

Machine learning models for segmentation and classification of cyanobacterial cells

Abstract Timelapse microscopy has recently been employed to study the metabolism and physiology of cyanobacteria at the single-cell level. However, the identification of individual cells in brightfield images remains a significant challenge. Traditional intensity-based segmentation algorithms perform poorly when identifying individual cells in dense colonies due to a lack of contrast between neighboring cells. Here, we describe a newly developed software package called Cypose which uses machine learning (ML) models to solve two specific tasks: segmentation of individual cyanobacterial cells, and classification of cellular phenotypes. The segmentation models are based on the Cellpose framework, while classification is performed using a convolutional neural network named Cyclass. To our knowledge, these are the first developed ML-based models for cyanobacteria segmentation and classification. When compared to other methods, our segmentation models showed improved performance and were able to segment cells with varied morphological phenotypes, as well as differentiate between live and lysed cells. We also found that our models were robust to imaging artifacts, such as dust and cell debris. Additionally, the classification model was able to identify different cellular phenotypes using only images as input. Together, these models improve cell segmentation accuracy and enable high-throughput analysis of dense cyanobacterial colonies and filamentous cyanobacteria.

Huffine, Clair A.↗

FEW questions, many answers: using machine learning to assess how students connect food–energy–water (FEW) concepts

There is growing support and interest in postsecondary interdisciplinary environmental education which integrate concepts and disciplines in addition to providing varied perspectives. There is a need to assess student learning in these programs as well as rigorous evaluation of educational practices, especially of complex synthesis concepts. This work tests a text classification machine learning model as a tool to assess student systems thinking capabilities using two questions anchored by the Food-Energy-Water (FEW) Nexus phenomena by answering two questions (1) Can machine learning models be used to identify instructor-determined important concepts in student responses? (2) What do college students know about the interconnections between food, energy and water, and how have students assimilated systems thinking into their constructed responses about FEW? Reported here are a broad range of model performances across 26 text classification models associated with two different assessment items, with model accuracy ranging from 0.755 to 0.992. Expert-like responses were infrequent in our dataset compared to responses providing simpler, incomplete explanations of the systems presented in the question. For those students moving from describing individual effects to multiple effects, their reasoning about the mechanism behind the system indicates advanced systems thinking ability. Specifically, students exhibit higher expertise for explaining changing water usage than discussing tradeoffs for such changing usage. This research represents one of the first attempts to assess the links between foundational, discipline-specific concepts and systems thinking ability. These text classification approaches to scoring student FEW Nexus Constructed Responses (CR) indicate how these approaches can be used, in addition to several future research priorities for interdisciplinary, practice-based education research. Development of further complex question items using machine learning would allow evaluation of the relationship between foundational concept understanding and integration of those concepts as well as more nuanced understanding of student comprehension of complex interdisciplinary concepts.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Machine Learning Models for Binary Molecular Classification using VUV Absorption Spectra

Machine learning methods were combined with differential absorption spectroscopy measurements in the vacuum-ultraviolet region (5.167 – 9.920 eV) in order to develop predictive capabilities for inferring molecular structure from the spectra. Several types of species were analyzed and, for modeling purposes, were defined using a single classification: (1) alkane, (2) conjugation with oxygen (e.g. diacetyl, ethyl vinyl ether), (3) non-conjugated alkene (e.g. 1-butene, 1,4-cyclohexadiene), (4) oxygen-containing (e.g. 1-butanol, tetrahydrofuran), or (5) cyclic (e.g. cyclopentane, cyclohexanone). The latter molecular classification excluded cyclic ethers. Several modeling methods were employed in the analysis of 102 absorption spectra, 24 of which were measured for the first time. The primary objective was to identify suitable methods that enable accurate predictions of molecular structure classifications with minimized statistical uncertainties. Rather than identifying a single, unifying method to reliably predict molecular structure contributions to VUV absorption spectra, coordination is required among a particular method, the type of molecular structure detail (e.g. conjugation), and absorption region of interest. The latter is accomplished using a binning approach, wherein absorption regions of ~0.5 eV were utilized rather than the entire ~4.8 eV range. Photon energy binning enabled analysis of region-specific predictions of accuracy, precision, and recall. The outcome from the binning approach is that, rather than utilizing the entire spectrum, optimal determination of molecular structure using machine learning methods depends on the absorption region. Furthermore, the present work provides separate machine learning models for each molecular classification, which enables the identification of multi-functional species relevant to atmospheric chemistry and combustion chemistry, where isomer-resolved speciation is critical to understanding complex reaction networks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DeepAdversaries: examining the robustness of deep learning models for galaxy morphology classification

With increased adoption of supervised deep learning methods for work with cosmological survey data, the assessment of data perturbation effects (that can naturally occur in the data processing and analysis pipelines) and the development of methods that increase model robustness are increasingly important. In the context of morphological classification of galaxies, we study the effects of perturbations in imaging data. In particular, we examine the consequences of using neural networks when training on baseline data and testing on perturbed data. We consider perturbations associated with two primary sources: (a) increased observational noise as represented by higher levels of Poisson noise and (b) data processing noise incurred by steps such as image compression or telescope errors as represented by one-pixel adversarial attacks. We also test the efficacy of domain adaptation techniques in mitigating the perturbation-driven errors. We use classification accuracy, latent space visualizations, and latent space distance to assess model robustness in the face of these perturbations. For deep learning models without domain adaptation, we find that processing pixel-level errors easily flip the classification into an incorrect class and that higher observational noise makes the model trained on low-noise data unable to classify galaxy morphologies. On the other hand, we show that training with domain adaptation improves model robustness and mitigates the effects of these perturbations, improving the classification accuracy up to 23% on data with higher observational noise. Domain adaptation also increases up to a factor of ${\approx}2.3$ the latent space distance between the baseline and the incorrectly classified one-pixel perturbed image, making the model more robust to inadvertent perturbations. Successful development and implementation of methods that increase model robustness in astronomical survey pipelines will help pave the way for many more uses of deep learning for astronomy.

79 ASTRONOMY AND ASTROPHYSICS↗

Linear dimensionality of Landsat agricultural data with implications for classification

A model for the Landsat multispectral scanner data, representing a generalization of the commonly used Gaussian model, has been formulated and analyzed. The model hypothesizes that the data for different crop types essentially lie on distinct hyperplanes in the feature space. Tests of this model reveal that: (1) the agricultural data from any single acquisition (i.e., four-channel) of Landsat are essentially two dimensional, regardless of the crop type; and (2) the data from different sites and different stages of crop development all lie on planes which are parallel. These findings have significant implications for data display, classification, feature extraction, and signature extension.

Wheeler, S. G.↗