Quantifying Uncertainty in Machine Learning Models for Time Series Classification.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Quark model status examined in classifying low lying meson states
Context. The advent of next-generation survey instruments, such as theVera C. RubinObservatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. Aims. Our aim is to develop a next-generation model for the classification of variable astronomical objects. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. Methods. ATAT consists of two transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a balanced hierarchical random forest (BHRF) trained on human-engineered features derived from light curves and metadata. Results. When trained on light curves and metadata, ATAT achieves a macro F1 score of 82.9 ± 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1 score of 79.4 ± 0.1. Conclusions. The use of transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as theVera C. RubinObservatory, in real-world brokering scenarios.
Classification, decomposition and modeling of polarimetric SAR data has received a great deal of attention in the recent literature. The objective behind these efforts is to better understand the scattering mechanisms which give rise to the polarimetric signatures seen in SAR image data. In this Paper an approach is described, which involves the fit of a combination of two simple scattering mechanisms to polarimetric SAR observations. The mechanisms am canopy scatter from a cloud of randomly oriented oblate spheroids, and a ground scatter term, which can represent double-bounce scatter from a pair of orthogonal surfaces with different dielectric constants or Bragg scatter from a moderately rough surface, seen through a layer of vertically oriented scatterers. An advantage of this model fit approach is that the scattering contributions from the two basic scattering mechanisms can be estimated for clusters of pixels in polarimetric SAR images. The solution involves the estimation of four parameters from four separate equations. The model fit can be applied to polarimetric AIRSAR data at C-, L- and P-Band.
Classification, decomposition and modeling of polarimetric SAR data has received a great deal of attention in the recent literature. The objective behind these efforts is to better understand the scattering mechanisms which give rise to the polarimetric signatures seen in SAR image data.
Species distribution models are increasing in popularity for mapping suitable habitat for species of management concern. Many investigators now recognize that extrapolations of these models with geographic information systems (GIS) might be sensitive to the environmental bounds of the data used in their development, yet there is no recommended best practice for "clamping" model extrapolations. We relied on two commonly used modeling approaches: classification and regression tree (CART) and maximum entropy (Maxent) models, and we tested a simple alteration of the model extrapolations, bounding extrapolations to the maximum and minimum values of primary environmental predictors, to provide a more realistic map of suitable habitat of hybridized Africanized honey bees in the southwestern United States. Findings suggest that multiple models of bounding, and the most conservative bounding of species distribution models, like those presented here, should probably replace the unbounded or loosely bounded techniques currently used [Current Zoology 57 (5): 642-647, 2011].
In some embodiments, a system model construction platform may receive, from a system node data store, system node data associated with an industrial asset. The system model construction platform may automatically construct a data-driven, dynamic system model for the industrial asset based on the received system node data. A synthetic attack platform may then inject at least one synthetic attack into the data-driven, dynamic system model to create, for each of a plurality of monitoring nodes, a series of synthetic attack monitoring node values over time that represent simulated attacked operation of the industrial asset. The synthetic attack platform may store, in a synthetic attack space data source, the series of synthetic attack monitoring node values over time that represent simulated attacked operation of the industrial asset. This information may then be used, for example, along with normal operational data to construct a threat detection model for the industrial asset.
Earth system models (ESMs) are a common tool for estimating local and global greenhouse gas emissions under current and projected future conditions. Efforts are underway to expand the representation of wetlands in the Energy Exascale Earth System Model (E3SM) Land Model (ELM) by resolving the simultaneous contributions to greenhouse gas fluxes from multiple, different, sub-grid-scale patch-types, representing different eco-hydrological patches within a wetland. However, for this effort to be effective, it should be coupled with the detection and mapping of within-wetland eco-hydrological patches in real-world wetlands, providing models with corresponding information about vegetation cover. In this short communication, we describe the application of a recently developed NDVI-based method for within-wetland vegetation classification on a coastal wetland in Louisiana and the use of the resulting yearly vegetation cover as input for ELM simulations. Processed Harmonized Landsat and Sentinel-2 (HLS) datasets were used to drive the sub-grid composition of simulated wetland vegetation each year, thus tracking the spatial heterogeneity of wetlands at sufficient spatial and temporal resolutions and providing necessary input for improving the estimation of methane emissions from wetlands. Our results show that including NDVI-based classification in an ELM reduced the uncertainty in predicted methane flux by decreasing the model’s RMSE when compared to Eddy Covariance measurements, while a minimal bias was introduced due to the resampling technique involved in processing HLS data. Our study shows promising results in integrating the remote sensing-based classification of within-wetland vegetation cover into earth system models, while improving their performances toward more accurate predictions of important greenhouse gas emissions.
This research focuses on developing algorithms for nuclear non-proliferation detection using remote sensor modeling. To improve the performance of classification models, we implemented a data pipeline with feature extraction. This pipeline takes raw data and transforms it into smaller data points called features that still describe the model. Improving this classification works towards the departments of energy’s missions of ensuring American’s security and prosperity by creating technology that addresses nuclear challenges. To conduct this analysis, we used the Python programming language and some key packages, including tsfresh and TSFEL. Originally tsfresh was selected because it has the most statistical features out of all the packages. Later TSFEL was incorporated due to the additional features it can extract from data, such as temporal and spectral. However, feature extraction becomes challenging in the presence of missing values. In this case, two additional Python packages were added to our workflow, NumPy and pandas, allowing for the feature extraction process to handle unknown values. Our data pipeline was tested on data collected from a simulation that describes the process state of a physical example. The results show the pipeline’s capability to consume and extract a total 17 features from tabular data. Future work includes producing classifications using decision tree-based models such as XGBoost and improving data collection by analyzing feature importance.
With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized.
With new emerging technologies in the field of NLP, we explore their applications to digitize and analyze heritage Air Traffic Management (ATM) documents for planning and optimizing airspace operations. Specifically, this research focuses on harvesting semi-structured or un-structured information contained in Notices to Airmen (NOTAMs). Using NLP and other advanced data analytics, we will construct a data-driven framework which facilitates finding language patterns and the use of pretrained language models for classification and extraction of useful airspace constraints and restrictions. These may lead to tools that assist airspace users in understanding the constraints more efficiently, contributing to better route planning and safer execution. This paper explores three workflows entailing different NLP tasks. First, unsupervised techniques like word embedding and topic modeling are used for pattern finding and document classification. Second, a dataset is created by extracting information from the semi-structured NOTAM format as metadata for categorizing, visualizing, and extracting key entities driving NOTAM content. Third, modern pre-built deep learning based transformer models such as BERT, RoBERTa, and XLNet are evaluated on the question answering task, an even more robust approach to information extraction, as well as their respective fine-tuning tasks. In this work we include various performance metrics for the trained models to evaluate both accuracy and precision and we show that the models can be generalized for their respective tasks. The research work developed shows promise in uncovering trends in digital NOTAMs in the NAS and also offers a new framework for digitizing and inferring insights from free-form legacy NOTAMs, that are yet to be digitized. Video is an mp4 download, with a play time of 9 min 35 secs.
The International Classification of Childhood Cancer (ICCC) facilitates the effective classification of a heterogeneous group of cancers in the important pediatric population. However, there has been no development of machine learning models for the ICCC classification. We developed deep learning-based information extraction models from cancer pathology reports based on the ICD-O-3 coding standard. In this article, we describe extending the models to perform ICCC classification. We developed 2 models, ICD-O-3 classification and ICCC recoding (Model 1) and direct ICCC classification (Model 2), and 4 scenarios subject to the training sample size. We evaluated these models with a corpus consisting of 29206 reports with age at diagnosis between 0 and 19 from 6 state cancer registries. Our findings suggest that the direct ICCC classification (Model 2) is substantially better than reusing the ICD-O-3 classification model (Model 1). Applying the uncertainty quantification mechanism to assess the confidence of the algorithm in assigning a code demonstrated that the model achieved a micro-F1 score of 0.987 while abstaining (not sufficiently confident to assign a code) on only 14.8% of ambiguous pathology reports. Our experimental results suggest that the machine learning-based automatic information extraction from childhood cancer pathology reports in the ICCC is a reliable means of supplementing human annotators at state cancer registries by reading and abstracting the majority of the childhood cancer pathology reports accurately and reliably.
Contemporary models of pattern, detection and discrimination often employ template matching, but there have been few direct tests of this proposition. Adopting a method developed by Ahumada, we have analyzed how human observers discriminate between two letters of the alphabet ('c' and 'x'). The stimulus consisted of a one degree tall letter plus a four degree field of static white noise, both displayed for 16 frames at a 67 Hz frame rate. Our font and display dimensions approximated those of Solomon and Pelli. The observer identified the letter presented. A QUEST staircase varied letter contrast to maintain a 75% correct rate. For each trial, we preserved the information required to reconstruct the noise field. Possible trial categories based on (signal, response) pairs are: (c,c), (c,x), (x,c), (x,x). Noise fields were averaged separately for each category, and a final classification image was obtained by averaging the four mean images after inverting the sign of categories in which x was the response. If the observer employs a template, it should be revealed in the classification image. The lowpass-filtered classification image derived from 2048 responses of one observer is shown here, along with the corresponding ideal template. An approximation to the ideal template can be seen appropriately located within the classification image. We have also simulated and will discuss the classification images expected from various discrimination models in this experimental context. The construction of classification images appears to be a powerful tool for studying classification strategies used by human observers. Like a Rorschach test, it surreptitiously discovers the inner desires of the visual system.
New data are reported from five previously unanalyzed Apollo 12 mare basalts that are incorporated into an evaluation of previous petrogenetic models and classification schemes for these basalts. This paper proposes a classification for Apollo 12 mare basalts on the basis of whole-rock Mg# (molar 100*(Mg/(Mg+Fe))) and Rb/Sr ratio (analyzed by isotope dilution), whereby the ilmenite, olivine, and pigeonite basalt groups are readily distinguished from each other. Scrutiny of the Apollo 12 feldspathic 'suite' demonstrates that two of the three basalts previously assigned to this group (12031, 12038, 12072) can be reclassified: 12031 is a plagioclase-rich pigeonite basalt; and 12072 is an olivine basalt. Only basalt 12038 stands out as a unique sample to the Apollo 12 site, but whether this represents a single sample from another flow at the Apollo 12 site or is exotic to this site is equivocal. The question of whether the olivine and pigeonite basalt suites are co-magmatic is addressed by incompatible trace-element chemistry: the trends defined by these two suites when Co/Sm and Sm/Eu ratios are plotted against Rb/Sr ratio demonstrate that these two basaltic types cannot be co-magmatic. Crystal fractionation/accumulation paths have been calculated and show that neither the pigeonite, olivine, or ilmenite basalts are related by this process. Each suite requires a distinct and separate source region. This study also examines sample heterogeneity and the degree to which whole-rock analyses are representative, which is critical when petrogenetic interpretation is undertaken. Sample heterogeneity has been investigated petrographically (inhomogeneous mineral distribution) with consideration of duplicate analyses, and whether a specific sample (using average data) plots consistently upon a fractionation trend when a number of different compostional parameters are considered. Using these criteria, four basalts have been identified where reported analyses are not representative of the whole-rock composition: 12005, an ilmenite basalt; 12006 and 12036, olivine basalts; and 12031 previously classified as a feldspathic basalt, but reclassified as part of the pigeonite suite.
Barium (Ba) stars are characterised by an abundance of heavy elements made by the slow neutron capture process (s-process). This peculiar observed signature is due to the mass transfer from a stellar companion, bound in a binary stellar system, to the Ba star observed today. The signature is created when the stellar companion is an asymptotic giant branch (AGB) star. We aim to analyse the abundance pattern of 169 Ba stars using machine learning techniques and the AGB final surface abundances predicted by the FRUITY and Monash stellar models. We developed machine learning algorithms that use the abundance pattern of Ba stars as input to classify the initial mass and metallicity of each Ba star’s companion star using stellar model predictions. We used two algorithms. The first exploits neural networks to recognise patterns, and the second is a nearest-neighbour algorithm that focuses on finding the AGB model that predicts the final surface abundances closest to the observed Ba star values. In the second algorithm, we included the error bars and observational uncertainties in order to find the best-fit model. The classification process was based on the abundances of Fe, Rb, Sr, Zr, Ru, Nd, Ce, Sm, and Eu. We selected these elements by systematically removing s-process elements from our AGB model abundance distributions and identifying the elements whose removal had the biggest positive effect on the classification. We excluded Nb, Y, Mo, and La. Our final classification combined the output of both algorithms to identify an initial mass and metallicity range for each Ba star companion. With our analysis tools, we identified the main properties for 166 of the 169 Ba stars in the stellar sample. The classifications based on both stellar sets of AGB final abundances show similar distributions, with an average initial mass of M = 2.23 M ⊙ and 2.34 M ⊙ and an average [Fe/H] = –0.21 and –0.11, respectively. We investigated why the removal of Nb, Y, Mo, and La improves our classification and identified 43 stars for which the exclusion had the biggest effect. We found that these stars have statistically significant and different abundances for these elements compared to the other Ba stars in our sample. We discuss the possible reasons for these differences in the abundance patterns.
The tremendous backlog of unanalyzed satellite data necessitates the development of improved methods for data cataloging and analysis. Ford Aerospace has developed an image analysis system, SIANN (Satellite Image Analysis using Neural Networks) that integrates the technologies necessary to satisfy NASA's science data analysis requirements for the next generation of satellites. SIANN will enable scientists to train a neural network to recognize image data containing scenes of interest and then rapidly search data archives for all such images. The approach combines conventional image processing technology with recent advances in neural networks to provide improved classification capabilities. SIANN allows users to proceed through a four step process of image classification: filtering and enhancement, creation of neural network training data via application of feature extraction algorithms, configuring and training a neural network model, and classification of images by application of the trained neural network. A prototype experimentation testbed was completed and applied to climatological data.
The main application considered in this paper is predicting true kinases from randomly permuted kinases that share the same length and amino acid distributions as the true kinases. Numerous methods already exist for this classification task, such as HMMs, motif-matchers, and sequence comparison algorithms. We build on some of these efforts by creating a vector from the output of thousands of structurally based HMMs, created offline with Pfam-A seed alignments using SAM-T99, which then must be combined into an overall classification for the protein. Then we use a Support Vector Machine for classifying this large ensemble Pfam-Vector, with a polynomial and chisquared kernel. In particular, the chi-squared kernel SVM performs better than the HMMs and better than the BLAST pairwise comparisons, when predicting true from false kinases in some respects, but no one algorithm is best for all purposes or in all instances so we consider the particular strengths and weaknesses of each.
The Transformer eXplainability and eXploration library is intended to aid in the explorability and explainability of transformer classification networks, or transformer language models with sequence classification heads. The basic function of this library is to take a trained transformer and test/train dataset and produce an ipywidget dashboard which can be displayed in a jupyter notebook or in jupyter lab.