Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Integrating NDVI-Based Within-Wetland Vegetation Classification in a Land Surface Model Improves Methane Emission Estimations

Earth system models (ESMs) are a common tool for estimating local and global greenhouse gas emissions under current and projected future conditions. Efforts are underway to expand the representation of wetlands in the Energy Exascale Earth System Model (E3SM) Land Model (ELM) by resolving the simultaneous contributions to greenhouse gas fluxes from multiple, different, sub-grid-scale patch-types, representing different eco-hydrological patches within a wetland. However, for this effort to be effective, it should be coupled with the detection and mapping of within-wetland eco-hydrological patches in real-world wetlands, providing models with corresponding information about vegetation cover. In this short communication, we describe the application of a recently developed NDVI-based method for within-wetland vegetation classification on a coastal wetland in Louisiana and the use of the resulting yearly vegetation cover as input for ELM simulations. Processed Harmonized Landsat and Sentinel-2 (HLS) datasets were used to drive the sub-grid composition of simulated wetland vegetation each year, thus tracking the spatial heterogeneity of wetlands at sufficient spatial and temporal resolutions and providing necessary input for improving the estimation of methane emissions from wetlands. Our results show that including NDVI-based classification in an ELM reduced the uncertainty in predicted methane flux by decreasing the model’s RMSE when compared to Eddy Covariance measurements, while a minimal bias was introduced due to the resampling technique involved in processing HLS data. Our study shows promising results in integrating the remote sensing-based classification of within-wetland vegetation cover into earth system models, while improving their performances toward more accurate predictions of important greenhouse gas emissions.

54 ENVIRONMENTAL SCIENCES↗

Feature Extraction: Improving Remote Sensor Classification of Non-Proliferation

This research focuses on developing algorithms for nuclear non-proliferation detection using remote sensor modeling. To improve the performance of classification models, we implemented a data pipeline with feature extraction. This pipeline takes raw data and transforms it into smaller data points called features that still describe the model. Improving this classification works towards the departments of energy’s missions of ensuring American’s security and prosperity by creating technology that addresses nuclear challenges. To conduct this analysis, we used the Python programming language and some key packages, including tsfresh and TSFEL. Originally tsfresh was selected because it has the most statistical features out of all the packages. Later TSFEL was incorporated due to the additional features it can extract from data, such as temporal and spectral. However, feature extraction becomes challenging in the presence of missing values. In this case, two additional Python packages were added to our workflow, NumPy and pandas, allowing for the feature extraction process to handle unknown values. Our data pipeline was tested on data collected from a simulation that describes the process state of a physical example. The results show the pipeline’s capability to consume and extract a total 17 features from tabular data. Future work includes producing classifications using decision tree-based models such as XGBoost and improving data collection by analyzing feature importance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automatic information extraction from childhood cancer pathology reports

The International Classification of Childhood Cancer (ICCC) facilitates the effective classification of a heterogeneous group of cancers in the important pediatric population. However, there has been no development of machine learning models for the ICCC classification. We developed deep learning-based information extraction models from cancer pathology reports based on the ICD-O-3 coding standard. In this article, we describe extending the models to perform ICCC classification. We developed 2 models, ICD-O-3 classification and ICCC recoding (Model 1) and direct ICCC classification (Model 2), and 4 scenarios subject to the training sample size. We evaluated these models with a corpus consisting of 29206 reports with age at diagnosis between 0 and 19 from 6 state cancer registries. Our findings suggest that the direct ICCC classification (Model 2) is substantially better than reusing the ICD-O-3 classification model (Model 1). Applying the uncertainty quantification mechanism to assess the confidence of the algorithm in assigning a code demonstrated that the model achieved a micro-F1 score of 0.987 while abstaining (not sufficiently confident to assign a code) on only 14.8% of ambiguous pathology reports. Our experimental results suggest that the machine learning-based automatic information extraction from childhood cancer pathology reports in the ICCC is a reliable means of supplementing human annotators at state cancer registries by reading and abstracting the majority of the childhood cancer pathology reports accurately and reliably.

60 APPLIED LIFE SCIENCES↗

Barium stars as tracers of s -process nucleosynthesis in AGB stars: II. Using machine learning techniques on 169 stars

Barium (Ba) stars are characterised by an abundance of heavy elements made by the slow neutron capture process (s-process). This peculiar observed signature is due to the mass transfer from a stellar companion, bound in a binary stellar system, to the Ba star observed today. The signature is created when the stellar companion is an asymptotic giant branch (AGB) star. We aim to analyse the abundance pattern of 169 Ba stars using machine learning techniques and the AGB final surface abundances predicted by the FRUITY and Monash stellar models. We developed machine learning algorithms that use the abundance pattern of Ba stars as input to classify the initial mass and metallicity of each Ba star’s companion star using stellar model predictions. We used two algorithms. The first exploits neural networks to recognise patterns, and the second is a nearest-neighbour algorithm that focuses on finding the AGB model that predicts the final surface abundances closest to the observed Ba star values. In the second algorithm, we included the error bars and observational uncertainties in order to find the best-fit model. The classification process was based on the abundances of Fe, Rb, Sr, Zr, Ru, Nd, Ce, Sm, and Eu. We selected these elements by systematically removing s-process elements from our AGB model abundance distributions and identifying the elements whose removal had the biggest positive effect on the classification. We excluded Nb, Y, Mo, and La. Our final classification combined the output of both algorithms to identify an initial mass and metallicity range for each Ba star companion. With our analysis tools, we identified the main properties for 166 of the 169 Ba stars in the stellar sample. The classifications based on both stellar sets of AGB final abundances show similar distributions, with an average initial mass of M = 2.23 M ⊙ and 2.34 M ⊙ and an average [Fe/H] = –0.21 and –0.11, respectively. We investigated why the removal of Nb, Y, Mo, and La improves our classification and identified 43 stars for which the exclusion had the biggest effect. We found that these stars have statistically significant and different abundances for these elements compared to the other Ba stars in our sample. We discuss the possible reasons for these differences in the abundance patterns.

79 ASTRONOMY AND ASTROPHYSICS↗

Transformer eXplainability and eXploration

The Transformer eXplainability and eXploration library is intended to aid in the explorability and explainability of transformer classification networks, or transformer language models with sequence classification heads. The basic function of this library is to take a trained transformer and test/train dataset and produce an ipywidget dashboard which can be displayed in a jupyter notebook or in jupyter lab.

Martindale, Nathan [Oak Ridge National Lab. (ORNL)↗

Noise signal identification in time projection chamber data using deep learning model

Deep learning has been employed in various scientific fields and has provided promising results. Here, in this study, a deep learning classifier was implemented to improve the quality of data obtained from a time projection chamber. Digital waveforms of the detected signals were classified into the following three categories: particles, noises, and particles piled up with noises. A simple 1-dimensional convolutional neural network was developed for the classification. The model demonstrated an excellent performance on the test dataset. Its practical performance was also examined using track images and particle identification plots by comparing the original and clean data without the noise signals. The comparison clearly showed that the deep learning model improved the quality of data. The current study presents an effective application of the deep learning model for the time projection chamber data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Particle Track Classification Using Quantum Associative Memory (Final Technical Report)

This project explored the use of quantum-assisted algorithms for pattern matching in sub-atomic physics experiments. Pattern matching algorithms are commonly employed to prune data of random noise and to help discriminate between signals generated by particle tracks of interest and signals generated by background events. The quantum-assisted algorithms explored in this project were based on an Ising formulation of quantum associative model (QAMM) recall and quantum content-addressable memory (QCAM) recall. The recall is performed by comparing a probe pattern with those stored in a library of patterns encoded in the QAMM/QCAM model. The classification accuracy of QAMM and QCAM recall was determined as a function of detector resolution, noise, and efficiency and pattern density, where pattern density is defined as the ratio of the number of reference signal patterns encoded in the library to each pattern’s length. We found that QAMM achieved high classification accuracy when applied to datasets with low pattern density. QCAM achieved high classification accuracy for datasets with high pattern density and was found to be more robust to detector noise. The project methodology and results are described in detail in our arXiv preprint (arXiv:2011.11848) . This project was conducted by scientists at the Johns Hopkins University Applied Physics Laboratory and Oak Ridge National Laboratory from August 2018 to August 2020 and was supported by DOE grant DE-SC0019497.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An extended focused assessment with sonography in trauma ultrasound tissue-mimicking phantom for developing automated diagnostic technologies

Medical imaging-based triage is critical for ensuring medical treatment is timely and prioritized. However, without proper image collection and interpretation, triage decisions can be hard to make. While automation approaches can enhance these triage applications, tissue phantoms must be developed to train and mature these novel technologies. Here, we have developed a tissue phantom modeling the ultrasound views imaged during the enhanced focused assessment with sonography in trauma exam (eFAST). The tissue phantom utilized synthetic clear ballistic gel with carveouts in the abdomen and rib cage corresponding to the various eFAST scan points. Various approaches were taken to simulate proper physiology without injuries present or to mimic pneumothorax, hemothorax, or abdominal hemorrhage at multiple locations in the torso. Multiple ultrasound imaging systems were used to acquire ultrasound scans with or without injury present and were used to train deep learning image classification predictive models. Performance of the artificial intelligent (AI) models trained in this study achieved over 97% accuracy for each eFAST scan site. We used a previously trained AI model for pneumothorax which achieved 74% accuracy in blind predictions for images collected with the novel eFAST tissue phantom. Grad-CAM heat map overlays for the predictions identified that the AI models were tracking the area of interest for each scan point in the tissue phantom. Overall, the eFAST tissue phantom ultrasound scans resembled human images and were successful in training AI models. Tissue phantoms are critical first steps in troubleshooting and developing medical imaging automation technologies for this application that can accelerate the widespread use of ultrasound imaging for emergency triage.

60 APPLIED LIFE SCIENCES↗

Classifying and analyzing small-angle scattering data using weighted k nearest neighbors machine learning techniques

A consistent challenge for both new and expert practitioners of small-angle scattering (SAS) lies in determining how to analyze the data, given the limited information content of said data and the large number of models that can be employed. Machine learning (ML) methods are powerful tools for classifying data that have found diverse applications in many fields of science. Here, ML methods are applied to the problem of classifying SAS data for the most appropriate model to use for data analysis. The approach employed is built around the method of weighted k nearest neighbors (wKNN), and utilizes a subset of the models implemented in the SasView package (https://www.sasview.org/) for generating a well defined set of training and testing data. The prediction rate of the wKNN method implemented here using a subset of SasView models is reasonably good for many of the models, but has difficulty with others, notably those based on spherical structures. A novel expansion of the wKNN method was also developed, which uses Gaussian processes to produce local surrogate models for the classification, and this significantly improves the classification accuracy. Further, by integrating a stochastic gradient descent method during post-processing, it is possible to leverage the local surrogate model both to classify the SAS data with high accuracy and to predict the structural parameters that best describe the data. The linking of data classification and model fitting has the potential to facilitate the translation of measured data into results for both novice and expert practitioners of SAS.

97 MATHEMATICS AND COMPUTING↗

MindSynchro

This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automated classification of big X-ray diffraction data using deep learning models

Abstract In current in situ X-ray diffraction (XRD) techniques, data generation surpasses human analytical capabilities, potentially leading to the loss of insights. Automated techniques require human intervention, and lack the performance and adaptability required for material exploration. Given the critical need for high-throughput automated XRD pattern analysis, we present a generalized deep learning model to classify a diverse set of materials’ crystal systems and space groups. In our approach, we generate training data with a holistic representation of patterns that emerge from varying experimental conditions and crystal properties. We also employ an expedited learning technique to refine our model’s expertise to experimental conditions. In addition, we optimize model architecture to elicit classification based on Bragg’s Law and use evaluation data to interpret our model’s decision-making. We evaluate our models using experimental data, materials unseen in training, and altered cubic crystals, where we observe state-of-the-art performance and even greater advances in space group classification.

Chemistry↗

Evaluating Gaussian process metamodels and sequential designs for noisy level set estimation

Abstract We consider the problem of learning the level set for which a noisy black-box function exceeds a given threshold. To efficiently reconstruct the level set, we investigate Gaussian process (GP) metamodels. Our focus is on strongly stochastic simulators, in particular with heavy-tailed simulation noise and low signal-to-noise ratio. To guard against noise misspecification, we assess the performance of three variants: (i) GPs with Student- t observations; (ii) Student- t processes (TPs); and (iii) classification GPs modeling the sign of the response. In conjunction with these metamodels, we analyze several acquisition functions for guiding the sequential experimental designs, extending existing stepwise uncertainty reduction criteria to the stochastic contour-finding context. This also motivates our development of (approximate) updating formulas to efficiently compute such acquisition functions. Our schemes are benchmarked by using a variety of synthetic experiments in 1–6 dimensions. We also consider an application of level set estimation for determining the optimal exercise policy of Bermudan options in finance.

97 MATHEMATICS AND COMPUTING↗

PAVC: The foundation for a Pan-Arctic Vegetation Cover database

Field-measured Arctic vegetation cover data is essential for creating accurate, high-quality vegetation structure and composition maps. Extrapolating field data into high-resolution cover maps provides detailed, function-specific information for use in Earth System Models, vegetation classifications, and monitoring vegetation change over time and space. However, field campaigns that collect plant cover vary substantially in scope, method, and purpose, which makes them difficult to unify across data stores, and they are often not designed to meet remote sensing needs. In this work, we synthesized and harmonized field-based fractional cover data from various data stores to create a high-quality, consistent repository schema for remote sensing-based vegetation cover mapping applications. We developed a reproducible workflow for synthesizing visual estimate and point-intercept fractional cover data. The resultant Pan-Arctic Vegetation Cover (PAVC) database contains synthesized fractional cover at both the species and plant functional type levels. The latter includes absolute foliar cover for deciduous shrubs and trees, evergreen shrubs and trees, forbs, graminoids, lichen, bryophytes, and “other” vegetation, as well as absolute cover for litter and top cover for water and bare ground.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Classification of compact objects and model comparison using EOS knowledge

Nuclear theory and experiments, alongside astrophysical observations, constrain the equation of state (EOS) of supranuclear-dense matter. Conversely, knowledge of the EOS allows an improved interpretation of nuclear or astrophysical data. In this article, we use several established constraints on the EOS and the new NICER measurement of PSR J0437-4715 to comment on the nature of the primary companion in GW230529 and the companion of PSR J0514-4002E. We find that, with a probability of ≳84% and ≳68%, respectively, both objects are black holes. These likelihoods increase to above 95% when one uses GW170817’s remnant as an upper limit on the TOV mass. We also demonstrate that the current knowledge of the EOS substantially disfavors high masses and radii for PSR J⁢0030+0451, inferred recently when combining NICER with XMM-Newton background data and using particular hot-spot models. Lastly, we also use our obtained EOS knowledge to comment on measurements of the nuclear symmetry energy, finding that the large value predicted by the PREX-II measurement displays some mild tension with other constraints on the EOS.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep Cellular Recurrent Network for Efficient Analysis of Time-Series Data With Spatial Information

Efficient processing of large-scale time series data is an intricate problem in machine learning. Conventional sensor signal processing pipelines with hand engineered feature extraction often involve huge computational cost with high dimensional data. Deep recurrent neural networks have shown promise in automated feature learning for improved time-series processing. However, generic deep recurrent models grow in scale and depth with increased complexity of the data. This is particularly challenging in presence of high dimensional data with temporal and spatial characteristics. Consequently, this work proposes a novel deep cellular recurrent neural network (DCRNN) architecture to efficiently process complex multi-dimensional time series data with spatial information. Here, the cellular recurrent architecture in the proposed model allows for location-aware synchronous processing of time series data from spatially distributed sensor signal sources. Extensive trainable parameter sharing due to cellularity in the proposed architecture ensures efficiency in the use of recurrent processing units with high-dimensional inputs. This study also investigates the versatility of the proposed DCRNN model for classification of multi-class time series data from different application domains. Consequently, the proposed DCRNN architecture is evaluated using two time-series datasets: a multichannel scalp EEG dataset for seizure detection, and a machine fault detection dataset obtained in-house. The results suggest that the proposed architecture achieves state-of-the-art performance while utilizing substantially less trainable parameters when compared to comparable methods in the literature.

60 APPLIED LIFE SCIENCES↗

ORNL-Chi-Geometry

Library for benchmarking neural network models on classification tasks for chirality detection in atomistic structures of organic compounds.

Weaver, Rylie [Oak Ridge National Laboratory (ORNL↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗