Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Remote sensing of physiographic soil units of Bennett County, South Dakota

A study was conducted in Bennett County, South Dakota, to establish a rangeland test site for evaluating the usefulness of ERTS data for mapping soil resources in rangeland areas. Photographic imagery obtained in October, 1970, was analyzed to determine which type of imagery is best for mapping drainage and land use patterns. Imagery of scales ranging from 1:1,000,000 to 1.20,000 was used to delineate soil-vegetative physiographic units. The photo characteristics used to define physiographic units were texture, drainage pattern, tone pattern, land use pattern and tone. These units will be used as test data for evaluating ERTS data. The physiographic units were categorized into a land classification system. The various categories which were delineated at the different scales of imagery were designed to be useful for different levels of land use planning. The land systems are adequate only for planning of large areas for general uses. The lowest category separated was the facet. The facets have a definite soil composition and represent different soil landscapes. These units are thought to be useful for providing natural resource information needed for local planning.

Frazee, C. J.↗

An Impedance-Based Complexity Metric for Unmanned Aircraft System Traffic Scenario Classification

This paper introduces an impedance-based metric to capture the complexity of a given unmanned aircraft system traffic scenario. The metric accounts for both the number of aircraft and the traffic flow pattern. The work presented here extends an earlier approach that introduced another scenario complexity metric based on the number of potential conflicts weighted by the conflict resolution cost associated. Complexity measurements for randomly-generated scenarios were produced through high-fidelity fast-time simulations and treated as baseline. Then the impedance based metric was evaluated, for the same scenarios, without the need for an actual flight simulation and a conflict resolution method. The results show that the impedance-based metric has a strong correlation to the baseline data and performs marginally better than the weighted conflict-based complexity metric introduced in the earlier work. The metric computation generates impedance maps which are useful for identifying high complexity regions in a scenario, where flight plan changes might be necessitated. This metric can therefore be used, in conjunction with other complexity metrics, to inform adequate traffic management strategies and classify a traffic scenario as acceptable, unacceptable or acceptable with changes made to flight plans that pass through the high complexity regions. The metric can also be used as a guidance metric for strategic conflict management methods.

Complexity↗

Unsupervised Clustering of Microseismic Events and Focal Mechanism Analysis at the CO 2 Injection Site in Decatur, Illinois

Characterization of induced microseismicity at a carbon dioxide (CO 2 ) storage site is critical for preserving reservoir integrity and mitigating seismic hazards. We apply a multilevel machine learning (ML) approach that combines the nonnegative matrix factorization and hidden Markov model to extract spectral representations of microseismic events and cluster them to identify seismic patterns at the Illinois Basin-Decatur Project. Unlike traditional waveform correlation methods, this approach leverages spectral characteristics of first arrivals to improve event classification and detect previously undetected planes of weakness. By integrating ML-based clustering with focal mechanism analysis, we resolve small-scale fault structures that are below the detection limits of conventional seismic imaging. Our findings reveal temporal bursts of microseismicity associated with brittle failure, providing insights into the spatio-temporal evolution of fault reactivation during CO 2 injection. This approach enhances seismic monitoring capabilities at CO 2 injection sites by improving fault characterization beyond the resolution of standard geophysical surveys.

Willis, Rachel Marie [Sandia National Laboratories↗

Robust Event Classification Using Imperfect Real-world PMU Data

Here, this paper studies robust event classification using imperfect real-world phasor measurement unit (PMU) data. By analyzing the real-world PMU data, we find it is challenging to directly use this dataset for event classifiers due to the low data quality observed in PMU measurements and event logs. To address these challenges, we develop a novel machine learning framework for training robust event classifiers, which consists of three main steps: data preprocessing, fine-grained event data extraction, and feature engineering. Specifically, the data preprocessing step addresses the data quality issues of PMU measurements (e.g., bad data and missing data); in the fine-grained event data extraction step, a model-free event detection method is developed to accurately localize the events from the inaccurate event timestamps in the event logs; and the feature engineering step constructs the event features based on the patterns of different event types, in order to improve the performance and the interpretability of the event classifiers. Based on the proposed framework, we develop a workflow for event classification using the real-world PMU data streaming into the system in real time. Using the proposed framework, robust event classifiers can be efficiently trained based on many off-the-shelf lightweight machine learning models. Numerical experiments using the real-world dataset from the Western Interconnection of the U.S power transmission grid show that the event classifiers trained under the proposed framework can achieve high classification accuracy while being robust against low-quality data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Barium stars as tracers of s -process nucleosynthesis in AGB stars: II. Using machine learning techniques on 169 stars

Barium (Ba) stars are characterised by an abundance of heavy elements made by the slow neutron capture process (s-process). This peculiar observed signature is due to the mass transfer from a stellar companion, bound in a binary stellar system, to the Ba star observed today. The signature is created when the stellar companion is an asymptotic giant branch (AGB) star. We aim to analyse the abundance pattern of 169 Ba stars using machine learning techniques and the AGB final surface abundances predicted by the FRUITY and Monash stellar models. We developed machine learning algorithms that use the abundance pattern of Ba stars as input to classify the initial mass and metallicity of each Ba star’s companion star using stellar model predictions. We used two algorithms. The first exploits neural networks to recognise patterns, and the second is a nearest-neighbour algorithm that focuses on finding the AGB model that predicts the final surface abundances closest to the observed Ba star values. In the second algorithm, we included the error bars and observational uncertainties in order to find the best-fit model. The classification process was based on the abundances of Fe, Rb, Sr, Zr, Ru, Nd, Ce, Sm, and Eu. We selected these elements by systematically removing s-process elements from our AGB model abundance distributions and identifying the elements whose removal had the biggest positive effect on the classification. We excluded Nb, Y, Mo, and La. Our final classification combined the output of both algorithms to identify an initial mass and metallicity range for each Ba star companion. With our analysis tools, we identified the main properties for 166 of the 169 Ba stars in the stellar sample. The classifications based on both stellar sets of AGB final abundances show similar distributions, with an average initial mass of M = 2.23 M ⊙ and 2.34 M ⊙ and an average [Fe/H] = –0.21 and –0.11, respectively. We investigated why the removal of Nb, Y, Mo, and La improves our classification and identified 43 stars for which the exclusion had the biggest effect. We found that these stars have statistically significant and different abundances for these elements compared to the other Ba stars in our sample. We discuss the possible reasons for these differences in the abundance patterns.

79 ASTRONOMY AND ASTROPHYSICS↗

Benchmark Dose Analysis of DNA Damage Biomarker Responses Provides Compound Potency and Adverse Outcome Pathway Information for the Topoisomerase II Inhibitor Class of Compounds

Genetic toxicology data have traditionally been utilized for hazard identification to provide a binary call for a compound's risk. Recent advances in the scientific field, especially with the development of high‐throughput methods to quantify DNA damage, have influenced a change of approach in genotoxicity assessment. The in vitro MultiFlow® DNA Damage Assay is one such method which multiplexes γH2AX, p53, phospho‐histone H3 biomarkers into a single‐flow cytometric analysis (Bryce et al., [2016]: Environ Mol Mutagen 57:546–558). This assay was used to study human TK6 cells exposed to each of eight topoisomerase II poisons for 4 and 24 hr. Using PROAST v65.5, the Benchmark Dose approach was applied to the resulting flow cytometric datasets. With “compound” serving as covariate, all eight compounds were combined into a single analysis, per time point and endpoint. The resulting 90% confidence intervals, plotted in Log scale, were considered as the potency rank for the eight compounds. The in vitro MultiFlow data showed a maximum confidence interval span of 1Log, which indicates data of good quality. Patterns observed in the compound potency rank were scrutinized by using the expert rule‐based software program Derek Nexus, developed by Lhasa Limited. Compound sub‐classification and structural alerts were considered contributory to the potencies observed for the topoisomerase II poisons studied herein. The Topo II poison Adverse Outcome Pathway was evaluated with MultiFlow endpoints serving as Key Events. The step‐wise approach described herein can be considered as a foundation for risk assessment of compounds within a specific mode of action of interest. Environ. Mol. Mutagen. 2020. © 2020 Wiley Periodicals, Inc.

Wheeldon, Ryan P.↗

Attend and Decode: 4D fMRI Task State Decoding Using Attention Models

Source code for Brain Attend and Decode paper. Functional magnetic resonance imaging (fMRI) is a neuroimaging modality that captures the blood oxygen level in a subject's brain while the subject either rests or performs a variety of functional tasks under different conditions. Given fMRI data, the problem of inferring the task, known as task state decoding, is challenging due to the high dimensionality (hundreds of million sampling points per datum) and complex spatio-temporal blood flow patterns inherent in the data. In this work, we propose to tackle the fMRI task state decoding problem by casting it as a 4D spatiotemporal classification problem. We present a novel architecture called Brain Attend and Decode (BAnD), that uses residual convolutional neural networks for spatial feature extraction and self-attention mechanisms for temporal modeling. We achieve significant performance gain compared to previous works on a 7-task benchmark from the large-scale Human Connectome Project-Young Adult (HCP-YA) dataset. We also investigate the transferability of BAnD's extracted features on unseen HCP tasks, either by freezing the spatial feature extraction layers and retraining the temporal model, or finetuning the entire model. The pre-trained features from BAnD are useful on similar tasks while finetuning them yields competitive results on unseen tasks/conditions.

Ng, BrendaM.↗

Analysis of scanner data for crop inventories

Accomplishments for a machine-oriented small grains labeler T&E, and for Argentina ground data collection are reported. Features of the small grains labeler include temporal-spectral profiles, which characterize continuous patterns of crop spectral development, and crop calendar shift estimation, which adjusts for planting date differences of fields within a crop type. Corn and soybean classification technology development for area estimation for foreign commodity production forecasting is reported. Presentations supporting quarterly project management reviews and a quarterly technical interchange meeting are also included.

Horvath, R.↗

The effect of lossy image compression on image classification

We have classified four different images, under various levels of JPEG compression, using the following classification algorithms: minimum-distance, maximum-likelihood, and neural network. The training site accuracy and percent difference from the original classification were tabulated for each image compression level, with maximum-likelihood showing the poorest results. In general, as compression ratio increased, the classification retained its overall appearance, but much of the pixel-to-pixel detail was eliminated. We also examined the effect of compression on spatial pattern detection using a neural network.

Paola, Justin D.↗

Gamma-Ray Burst Class Properties

Guided by the Supervised pattern recognition algorithm C4.5, we examine the three gamma-ray burst classes identified by Mukherjee et al. C4.5 provides strong statistical support for this classification. However, with C4.5 and our knowledge of the BATSE instrument, we demonstrate that Class 3 (intermediate fluence, intermediate duration, soft) does not have to be a distinct source population: statistical/systematic errors in measuring burst attributes combined with the well-known hardness/intensity correlation can cause low peak flux Class I (high fluence, long, intermediate hardness) bursts to take on Class 3 characteristics naturally. Based on our hypothesis that the third class is not a distinct one, we provide rules so that future events can be placed in either Class I or Class 2 (low fluence, short, hard). Using classified bursts from the BATSE 4B Catalog, we plot log(N>P) vs. log(P) curves and study spectral features of each class. We find that the two classes are relatively distinct on the basis of spectral parameters, alpha, Beta, and E(sub peak) alone. Although this does not indicate a better basis for classification, it does suggest that different physical conditions exist for Class I and Class 2 bursts. In the process of studying burst class characteristics, we identify a new bias that affects measurement of burst fluences and durations. Using a simple model of how burst duration can be underestimated, we generally characterize how this fluence duration bias affects BATSE measurements, and demonstrate the type of effect it can have on the BATSE fluence vs. peak flux diagram.

Hakkila, Jon↗

Unsupervised classification for region of interest in X-ray ptychography

X-ray ptychography offers high-resolution imaging of large areas at a high computational cost due to the large volume of data provided. To address the cost issue, we propose a physics-informed unsupervised classification algorithm that is performed prior to reconstruction and removes data outside the region of interest (RoI) based on the multimodal features present in the diffraction patterns. The preprocessing time for the proposed method is inconsequential in contrast to the resource-intensive reconstruction process, leading to an impressive reduction in the data workload to a mere 20% of the initial dataset. This capability consequently reduces computational time dramatically while preserving reconstruction quality. Through further segmentation of the diffraction patterns, our proposed approach can also detect features that are smaller than beam size and correctly classify them as within the RoI.

97 MATHEMATICS AND COMPUTING↗

Interpretable machine learning models classify minerals via spectroscopy

Developing methods to identify mineral species confidently and rapidly from Raman spectral analysis is critical to numerous fields. Traditionally, analysis relies on pattern matching the Raman spectrum of an unknown dataset with a supporting library of well-characterized spectral data, which may prove difficult for environmental samples that are poorly crystalline or phase mixtures. Here, we developed interpretable machine learning models that can classify uranium minerals by secondary oxyanion chemistry and other physicochemical properties based solely on Raman spectra. This new ML method produces a mineral profile of physical and chemical properties for an unknown sample and can rapidly classify or identify unknown minerals from Raman data, without the need for an exact pattern match in a spectral library. Training models are validated by 1. Strong correlation of high confidence model regions with published spectroscopic assignments and 2. Correct classification of a mineral not present in training data. Training data are from the Compendium of Uranium Raman and Infrared Experimental Spectra and available crystallographic information files within the open-source Smart Spectral Matching scientific framework. Physically meaningful classifier models can rapidly identify key structural and chemical information about unknown uranium minerals and the overall methodology is broadly applicable for mineral phases.

Machine learning↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

Feasibility study for locating archaeological village sites by satellite remote sensing techniques

The author has identified the following significant results. The objective is to determine the feasibility of detecting large Alaskan archaeological sites by satellite remote sensing techniques and mapping such sites. The approach used is to develop digital multispectral signatures of dominant surface features including vegetation, exposed soils and rock, hydrological patterns and known archaeological sites. ERTS-1 scenes are then printed out digitally in a map-like array with a letter reflecting the most appropriate classification representing each pixel. Preliminary signatures were developed and tested. It was determined that there was a need to tighten up the archaeological site signature by developing accurate signatures for all naturally-occurring vegetation and surface conditions in the vicinity of the test area. These second generation signatures have been tested by means of computer printouts and classified tape displays on the University of Alaska CDU-200 and by comparison with aerial photography. It has been concluded that the archaeological signatures now in use are as good as can be developed. Plans are to print out signatures for the entire test area and locate on topographic maps the likely locations of archaeological sites within the test area.

Cook, J. P.↗

Performance analysis of image processing algorithms for classification of natural vegetation in the mountains of southern California

The earth's forests fix carbon from the atmosphere during photosynthesis. Scientists are concerned that massive forest removals may promote an increase in atmospheric carbon dioxide, with possible global warming and related environmental effects. Space-based remote sensing may enable the production of accurate world forest maps needed to examine this concern objectively. To test the limits of remote sensing for large-area forest mapping, we use Landsat data acquired over a site in the forested mountains of southern California to examine the relative capacities of a variety of popular image processing algorithms to discriminate different forest types. Results indicate that certain algorithms are best suited to forest classification. Differences in performance between the algorithms tested appear related to variations in their sensitivities to spectral variations caused by background reflectance, differential illumination, and spatial pattern by species. Results emphasize the complexity between the land-cover regime, remotely sensed data and the algorithms used to process these data.

Yool, S. R.↗

The Value of Long-Term (40 years) Airborne Gamma Radiation SWE Record for Evaluating Three Observation-Based Gridded SWE Data Sets by Seasonal Snow and Land Cover Classifications

Observation‐based long‐term gridded snow water equivalent (SWE) products are important assets for hydrological and climate research. However, an evaluation of the currently available SWE products has been limited due to the lack of independent SWE data that extend over a large range of environmental conditions. In this study, three daily long‐term SWE products (Special Sensor Microwave Imager and Sounder [SSMI/S] SWE, GlobSnow‐2 SWE, and University of Arizona [UA] SWE) we reevaluated by seasonal snow cover and land cover classifications over the conterminous United States from1982 to 2017, using the historical airborne gamma radiation SWE observations (20,738 measurements).We found that there are similar patterns in SSMI/S and GlobSnow‐2 SWE when compared against the gamma SWE. However, GlobSnow‐2 SWE had better agreement with gamma SWE than SSMI/S SWE in some forested‐type classes and maritime and prairie snow classes. As compared to SSMI/S and GlobSnow‐2SWE, UA SWE has much better agreement with gamma SWE in all land cover types and snow classes. Tree cover and topographic heterogeneity affect the agreement between the gamma and gridded SWE and accuracy of gamma SWE itself with the largest differences typically occurring when the percent tree cover was 80% or higher, the terrain slope was steeper than 2.5°, and the elevation range exceeded 100 m. The results demonstrate the reliability of the UA SWE products and the benefits of the gamma radiation approach to measure SWE, especially in forested regions.

Eunsang Cho↗

Spatiotemporal Methane Emissions from Global Lakes and Reservoirs

Inland aquatic systems, such as lakes and reservoirs, contribute substantially to global methane (CH4) emissions; yet are among the most uncertain components of the total CH4 budget. Lakes and reservoirs have received recent attention as they may generate high CH4 fluxes. Improved quantification of these CH4 fluxes, particularly their spatiotemporal distribution, is key to realistically incorporating them in CH4 modeling and budget studies. Here we report on a new global, gridded (0.25° lat × 0.25° lon) study of lake and reservoir CH4 emissions, accounting for new knowledge regarding lake and reservoir areal extent and distribution, and spatiotemporal emission patterns influenced by diurnal variability, temperature-dependent seasonality, satellite-derived freeze-thaw dynamics, and eco-climatic and physical CH4-centric type classification. The results of this new data set comprise daily CH4 emissions from lake and reservoirs throughout the full annual cycle and are tightly anchored to field observations, in situ measurements, and remote-sensing observations. Results show that reservoirs cover 297 × 103 km2 globally and emit 10.1 Tg CH4 yr-1 from diffusive (1.2 Tg CH4 yr-1) and ebullitive (8.9 Tg CH4 yr-1) emission pathways. On a global scale, CH4 emitting areas of lakes cover 1853 × 103 km2 and emit 37.1 Tg CH4 yr-1 from diffusive (17.1 Tg CH4 yr-1) and ebullitive (23.0 Tg CH4 yr-1) emission pathways and an additional ebullition flux of 5.2 Tg CH4 yr-1 upon ice-melt due to the accumulation of bubbles during the freeze period. This analysis of lakes and reservoir CH4 emission addresses multiple gaps and uncertainties in previous studies and represents an important contribution to studies of the global CH4 budget. The new data sets and methodologies from this study provide a framework to better understand and model the current and future role of lakes and reservoirs in the global CH4 budget and to guide efforts to mitigate inland aquatic system CH4 emissions. This presentation will describe the methodologies applied and the major results of this study focusing on the spatiotemporal distribution of lake and reservoir type classification, processes driving lake and reservoir emission seasonality, and the contribution of individual lake and reservoir types to the annual cycle of inland aquatic CH4 emissions.

Spatiotemporal↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗