Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Synergy of Urban Heat, Pollution, and Social 1 Vulnerability in One of America’s Most Rapidly 2 Growing Cities: Houston, We Have a Problem

During the first two decades of the twenty-first century, we analyze the expansion of 23urban land cover, urban heat island (UHI), and urban pollution island (UPI) in the Houston 24Metropolitan Area (HMA) using land cover classifications derived from Landsat and land/aerosol 25products from NASA's Moderate Resolution Imaging Spectroradiometer. Our approach involves 26both direct utilization and fusion with in situobservations for a comprehensive characterization.27We also examined how social vulnerability within the HMA changed during the study period and 28whether the synergy of UHI, UPI,and social vulnerability enhances environmental inequalities.

Andrew Blackford↗

Early Season Large-Area Winter Crop Mapping Using MODIS NDVI Data, Growing Degree Days Information and a Gaussian Mixture Model

Knowledge on geographical location and distribution of crops at global, national and regional scales is an extremely valuable source of information applications. Traditional approaches to crop mapping using remote sensing data rely heavily on reference or ground truth data in order to train/calibrate classification models. As a rule, such models are only applicable to a single vegetation season and should be recalibrated to be applicable for other seasons. This paper addresses the problem of early season large-area winter crop mapping using Moderate Resolution Imaging Spectroradiometer (MODIS) derived Normalized Difference Vegetation Index (NDVI) time-series and growing degree days (GDD) information derived from the Modern-Era Retrospective analysis for Research and Applications (MERRA-2) product. The model is based on the assumption that winter crops have developed biomass during early spring while other crops (spring and summer) have no biomass. As winter crop development is temporally and spatially non-uniform due to the presence of different agro-climatic zones, we use GDD to account for such discrepancies. A Gaussian mixture model (GMM) is applied to discriminate winter crops from other crops (spring and summer). The proposed method has the following advantages: low input data requirements, robustness, applicability to global scale application and can provide winter crop maps 1.5-2 months before harvest. The model is applied to two study regions, the State of Kansas in the US and Ukraine, and for multiple seasons (2001-2014). Validation using the US Department of Agriculture (USDA) Crop Data Layer (CDL) for Kansas and ground measurements for Ukraine shows that accuracies of greater than 90% can be achieved in mapping winter crops 1.5-2 months before harvest. Results also show good correspondence to official statistics with average coefficients of determination R(exp. 2) greater than 0.85.

mixture model↗

Adaptive Discovery and Mixed-Variable Optimization of Next Generation Synthesizable Microelectronic Materials

Design of new microelectronic materials is characterized by several challenges such as high-dimensionality of the atomic structure-composition variable space, formidable cost of directly using high-fidelity simulations for design optimization, dispersity in literature-reported similar materials and synthesis methods, complex physical mechanisms, and mixed qualitative and quantitative design variables that lead to a disjointed design space. Even though machine learning (ML) techniques have been employed to expedite materials innovation, existing methods treat ML and design optimization as two separate processes, failing to resolve the fundamental challenges associated with high dimensionality and mixed-variable complexity. We have developed a ML enhanced mixed-variable material design optimization framework to efficiently extract useful information from existing data in literature and physics-based simulations to guide the autonomous search for optimal materials. Our proposed framework is composed of four computational modules: (1) a natural language processing (NLP) based virtual screening module, (2) classification based concept exploration module, (3) a density functional theory (DFT)-based high-fidelity evaluation model, and (4) a novel latent-variable Gaussian process (LVGP) ML model for mixed-variable problems with uncertainty quantification, which seamlessly integrates with Bayesian Optimization (BO) and achieves superb efficiency through embedded physics-based dimension reduction. Our approach is demonstrated and validated using the testbed of functional materials exhibiting metal-insulation transitions (MITs), with the targeted reversible resistivity changes (∼10^5) near room temperature. At the end of the 30-month project, we have developed a series of new ML techniques using NLP, conditional variational autoencoders, active learning, latent-variable Gaussian processes, integrated with Bayesian optimization. Our project has resulted in new predicted MITs compounds and improved understanding of MITs microscopic mechanisms, which in turn will revolutionize microelectronics science to provide energy-saving solutions. Our research has improved both creativity and efficiency in transforming rare-event discoveries of new functional materials to persistent innovations. In addition to open-sourcing the online MIT database and the classification model, the LVGP open source code has been downloaded more than 15,000 times within two years. More than 40 MIT compounds have been identified and many have been pursued experimentally via collaborators. The research results are published in close to 20 collaborative papers in high-impact journals, such as Chem. Mater., Appl. Phys. Rev., Sci. Rep., among others of design space.

36 MATERIALS SCIENCE↗

Predicting Search Task Difficulty through a Discrete‐Time Action Log Representation on Spectrum Kernel

ABSTRACT Predicting perceived difficulty on a web search task is an open problem in the interactive information retrieval field. A common approach to tackle it, is through features obtained from full search sessions, which are then used to train classification models. In this poster we attempt to predict perceived task difficulty at different stages of the search process. To do so, we use the spectrum kernel for support vector machine (SVM) classification. Our preliminary results suggest that by using behavioral data from the first query segment, it is possible to provide timely classifications of whether a search task is perceived as hard or easy.

Gacitúa, Daniel↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

Digital Analytics, Causal Knowledge Acquisition and Reasoning for Technical Language Processing

Complex engineering systems such as nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) data elements that contain information on the status of components, assets, and systems. Some of this information is textual in form and can be found in documents such as incident reports (IRs) and work orders (WOs). Analyses of textual data in current NPPs-using natural language processing (NLP) methods-have been expanded over the last decade, and it is only recently that the true potential of such analyses has emerged. So far, applications of NLP methods have mostly been limited to classification and prediction, the goal being to identify the nature of the textual element (e.g., safety or non-safety related). Here, we target a more complex problem: automatically extracting knowledge from a textual element in order to assist system engineers in conducting system health assessments. Knowledge extraction is a very broad concept, and its definition may vary depending on the application context. Our methods are a blend of both rule-based and machine learning (ML) algorithms. For our purposes, knowledge extraction means identifying the systems or assets mentioned in a given textual element, as well as the type of event described (e.g., component failure or maintenance activity). In addition, we want to capture details such as measured quantities and the temporal/cause-effect relations between events. In this tool, we also demonstrate how textual data elements are preprocessed in order to handle typos, acronyms, and abbreviations. One main feature of these methods is that they are not based solely on data, but are in fact model-based. In other words, they also rely on MBSE models that are designed to capture-from a functional point of view-the architecture of the systems/assets under consideration. The main purpose of such models is to digitally emulate system engineers' knowledge of system and asset architecture and to identify dependencies among systems, assets, and components. Provided these models, analyses of textual and numeric ER data can be performed by first identifying the OPM model elements to which the ER data elements are referring. The relationships between ER data elements are then identified by checking for any temporal or logical dependencies.

Mandelli, Diego [Idaho National Laboratory (INL), ↗

A regional land use survey based on remote sensing and other data: A report on a LANDSAT and computer mapping project, volume 2

The author has identified the following significant results. The project mapped land use/cover classifications from LANDSAT computer compatible tape data and combined those results with other multisource data via computer mapping/compositing techniques to analyze various land use planning/natural resource management problems. Data were analyzed on 1:24,000 scale maps at 1.1 acre resolution. LANDSAT analysis software and linkages with other computer mapping software were developed. Significant results were also achieved in training, communication, and identification of needs for developing the LANDSAT/computer mapping technologies into operational tools for use by decision makers.

Nez, G.↗

Status and Direction of Tribology as a Science in the 80's. Understanding and Prediction

The most challenging research problems in tribology for the next decade or beyond are classified horizontally into two categories: (1) understanding of basic mechanisms and (2) prediction of practical performance. Vertical classifications are in terms of particular themes or fields of interest. Areas where more fundamental work is required are: adhesion and friction of clean and contaminated surfaces; lubrication; new materials; surface characterization at the engineering level (topography) and at the atomic levels (various spectroscopies); and wear.

Tabor, D.↗

Detection of lowland flooding using active microwave systems

The development of radar systems with longer wavelenths (greater than 3 cm) has provided new possibilities regarding the utilization of radar. Thus, it has been found that the interpretation of data from radar images can be a valuable classification aid for applications related to water resources. In the case of an interpreter accustomed to photographic or visible/infrared images, an evaluation of radar images presents some problems, because the radar is sensing a set of surface characteristics which have little influence on visible/infrared systems. Detectable features in radar images caused by differences in dielectric properties are usually associated with the water content of either soils or vegetation. The present paper is concerned with studies which were initiated in 1976. The studies had the objective to define the magnitude of the effects on radar data caused by flood waters under vegetation. The obtained results indicate the feasibility to detect flood conditions beneath a forest canopy, and to obtain an improved definition of the land-water boundary.

Ormsby, J. P.↗

MIMD computing in the USA - 1984

It is often said that the 1980s are becoming the decade of multiinstruction stream or MIMD computers, while the 1970s could be described as the decade of the SIMD (single instruction stream multiple data stream) computers. The availability of microprocessors and VLSI facilities has led to the proposal and construction of novel computer architectures based on linking many hundreds or even thousands of microprocessors, or specially designed VLSI chips. Some of the larger manufacturers offer computers with a small number of CPUs. Because of the variety of the new developments, it was decided to conduct a survey of proposed and existing MIMD computers in the U.S., taking into account a simple classification of the different devices. Particular attention is given to computers which are designed for numerical work with floating-point numbers and the solution of large problems in physics, chemistry, and engineering.

Hockney, R. W.↗

Interaction of the terrestrial and atmospheric hydrological cycles in the context of the North American southwest summer monsoon

Work under this grant has used information on precipitation and water vapor fluxes in the area of the Mexican Monsoon to analyze the regional precipitation climatology, to understand the nature of water vapor transport during the monsoon using model and observational data, and to analyze the ability of the TRMM remote sensing algorithm to characterize precipitation. An algorithm for estimating daily surface rain volumes from hourly GOES infrared images was developed and compared to radar data. Estimates were usually within a factor of two, but different linear relations between satellite reflectances and rainfall rate were obtained for each day, storm type and storm development stage. This result suggests that using TRMM sensors to calibrate other satellite IR will need to be a complex process taking into account all three of the above factors. Another study, this one of the space-time variability of the Mexican Monsoon, indicate that TRMM will have a difficult time, over the course of its expected three year lifetime, identifying the diurnal cycle of precipitation over monsoon region. Even when considering monthly rainfalls, projected satellite estimates of August rainfall show a root mean square error of 38 percent. A related examination of spatial variability of mean monthly rainfall using a novel method for removing the effects of elevation from gridded gauge data, show wide variation from a satellite-based rainfall estimates for the same time and space resolution. One issue addressed by our research, relating to the basic character of the monsoon circulation, is the determination of the source region for moisture. The monthly maps produced from our study of monsoon variability show the presence of two rainfall maxima in the analysis normalized to sea level, one in south-central Arizona associated with the Mexican monsoon maximum and one in southeastern New Mexico associated with the Gulf of Mexico. From the point of view of vertically-integrated fluxes and flux divergence of water vapor from ECMWF data, most moisture at upper levels arrives from the Gulf of Mexico, while low level moisture comes from the northern Gulf of California. Composites of ECMWF analyses for wet and dry periods (classified by rain gauge data) show that both regimes show low level moisture arriving from northern and central Gulf of California. Above 700 MB, moisture comes from both source regions and the Sierra Madre Occidental. During wet periods a longer fetch through the moist air mass above western Mexico results in a greater moisture flux into the Sonoran Desert region, while there is less moisture from the Gulf of Mexico both above and below 700 mb. Work on the grant subcontract at the University of Colorado concentrated on the development of a technique useful to TRMM combining visible, infrared and passive microwave data for measuring precipitation. Two established techniques using either visible or infrared data applied over the US Southwest correlated with gauges at the 0.58 to 0.70 level. The application of some established passive microwave techniques were less successful for a variety of reason, including problems in both the gauge and satellite data quality, sampling problems and weaknesses inherent in the algorithms themselves. A more promising solution for accurate rainfall estimation was explored using visible and infrared data to perform a cloud classification, which when combined with information about the background (e.g. Iand/ocean), was used to select the most appropriate microwave algorithm from a suite of possibilities.

Dickinson, Robert E.↗

Semi-Supervised Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection

In the era of big astronomical surveys, our ability to leverage artificial intelligence algorithms simultaneously for multiple datasets will open new avenues for scientific discovery. Unfortunately, simply training a deep neural network on images from one data domain often leads to very poor performance on any other dataset. Here we develop a Universal Domain Adaptation method DeepAstroUDA, capable of performing semi-supervised domain alignment that can be applied to datasets with different types of class overlap. Extra classes can be present in any of the two datasets, and the method can even be used in the presence of unknown classes. For the first time, we demonstrate the successful use of domain adaptation on two very different observational datasets (from SDSS and DECaLS). We show that our method is capable of bridging the gap between two astronomical surveys, and also performs well for anomaly detection and clustering of unknown data in the unlabeled dataset. We apply our model to two examples of galaxy morphology classification tasks with anomaly detection: 1) classifying spiral and elliptical galaxies with detection of merging galaxies (three classes including one unknown anomaly class); 2) a more granular problem where the classes describe more detailed morphological properties of galaxies, with the detection of gravitational lenses (ten classes including one unknown anomaly class).

79 ASTRONOMY AND ASTROPHYSICS↗

Statistical theory and methodology for remote sensing data analysis

A model is developed for the evaluation of acreages (proportions) of different crop-types over a geographical area using a classification approach and methods for estimating the crop acreages are given. In estimating the acreages of a specific croptype such as wheat, it is suggested to treat the problem as a two-crop problem: wheat vs. nonwheat, since this simplifies the estimation problem considerably. The error analysis and the sample size problem is investigated for the two-crop approach. Certain numerical results for sample sizes are given for a JSC-ERTS-1 data example on wheat identification performance in Hill County, Montana and Burke County, North Dakota. Lastly, for a large area crop acreages inventory a sampling scheme is suggested for acquiring sample data and the problem of crop acreage estimation and the error analysis is discussed.

Odell, P. L.↗

FOCIS: A forest classification and inventory system using LANDSAT and digital terrain data

Accurate, cost-effective stratification of forest vegetation and timber inventory is the primary goal of a Forest Classification and Inventory System (FOCIS). Conventional timber stratification using photointerpretation can be time-consuming, costly, and inconsistent from analyst to analyst. FOCIS was designed to overcome these problems by using machine processing techniques to extract and process tonal, textural, and terrain information from registered LANDSAT multispectral and digital terrain data. Comparison of samples from timber strata identified by conventional procedures showed that both have about the same potential to reduce the variance of timber volume estimates over simple random sampling.

Strahler, A. H.↗

The analysis of polar clouds from AVHRR satellite data using pattern recognition techniques

The cloud cover in a set of summertime and wintertime AVHRR data from the Arctic and Antarctic regions was analyzed using a pattern recognition algorithm. The data were collected by the NOAA-7 satellite on 6 to 13 Jan. and 1 to 7 Jul. 1984 between 60 deg and 90 deg north and south latitude in 5 spectral channels, at the Global Area Coverage (GAC) resolution of approximately 4 km. This data embodied a Polar Cloud Pilot Data Set which was analyzed by a number of research groups as part of a polar cloud algorithm intercomparison study. This study was intended to determine whether the additional information contained in the AVHRR channels (beyond the standard visible and infrared bands on geostationary satellites) could be effectively utilized in cloud algorithms to resolve some of the cloud detection problems caused by low visible and thermal contrasts in the polar regions. The analysis described makes use of a pattern recognition algorithm which estimates the surface and cloud classification, cloud fraction, and surface and cloudy visible (channel 1) albedo and infrared (channel 4) brightness temperatures on a 2.5 x 2.5 deg latitude-longitude grid. In each grid box several spectral and textural features were computed from the calibrated pixel values in the multispectral imagery, then used to classify the region into one of eighteen surface and/or cloud types using the maximum likelihood decision rule. A slightly different version of the algorithm was used for each season and hemisphere because of differences in categories and because of the lack of visible imagery during winter. The classification of the scene is used to specify the optimal AVHRR channel for separating clear and cloudy pixels using a hybrid histogram-spatial coherence method. This method estimates values for cloud fraction, clear and cloudy albedos and brightness temperatures in each grid box. The choice of a class-dependent AVHRR channel allows for better separation of clear and cloudy pixels than does a global choice of a visible and/or infrared threshold. The classification also prevents erroneous estimates of large fractional cloudiness in areas of cloudfree snow and sea ice. The hybrid histogram-spatial coherence technique and the advantages of first classifying a scene in the polar regions are detailed. The complete Polar Cloud Pilot Data Set was analyzed and the results are presented and discussed.

Smith, William L.↗

Multimode squeezing, biphotons and uncertainty relations in polarization quantum optics

The concept of squeezing and uncertainty relations are discussed for multimode quantum light with the consideration of polarization. Using the polarization gauge SU(2) invariance of free electromagnetic fields, we separate the polarization and biphoton degrees of freedom from other ones, and consider uncertainty relations characterizing polarization and biphoton observables. As a consequence, we obtain a new classification of states of unpolarized (and partially polarized) light within quantum optics. We also discuss briefly some interrelations of our analysis with experiments connected with solving some fundamental problems of physics.

Karassiov, V. P.↗

An improved classification tree analysis of high cost modules based upon an axiomatic definition of complexity

Identification of high cost modules has been viewed as one mechanism to improve overall system reliability, since such modules tend to produce more than their share of problems. A decision tree model was used to identify such modules. In this current paper, a previously developed axiomatic model of program complexity is merged with the previously developed decision tree process for an improvement in the ability to identify such modules. This improvement was tested using data from the NASA Software Engineering Laboratory.

Tian, Jianhui↗

Technical Language Processing of Nuclear Power Plants Equipment Reliability Data

Operating nuclear power plants (NPPs) generate and collect large amounts of equipment reliability (ER) element data that contain information about the status of components, assets, and systems. Some of this information is in textual form where the occurrence of abnormal events or maintenance activities are described. Analyses of NPP textual data via natural language processing (NLP) methods have expanded in the last decade, and only recently the true potential of such analyses has emerged. So far, applications of NLP methods have been mostly limited to classification and prediction in order to identify the nature of the given textual element (e.g., safety or non-safety relevant). In this paper, we target a more complex problem: the automatic generation of knowledge based on a textual element in order to assist system engineers in assessing an asset’s historical health performance. The goal is to assist system engineers in the identification of anomalous behaviors, cause–effect relations between events, and their potential consequences, and to support decision-making such as the planning and scheduling of maintenance activities. “Knowledge extraction” is a very broad concept whose definition may vary depending on the application context. In our particular context, it refers to the process of examining an ER textual element to identify the systems or assets it mentions and the type of event it describes (e.g., component failure or maintenance activity). In addition, we wish to identify details such as measured quantities and temporal or cause–effect relations between events. This paper describes how ER textual data elements are first preprocessed to handle typos, acronyms, and abbreviations, then machine learning (ML) and rule-based algorithms are employed to identify physical entities (e.g., systems, assets, and components) and specific phenomena (e.g., failure or degradation). A few applications relevant from an NPP ER point of view are presented as well.

97 MATHEMATICS AND COMPUTING↗