Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Results on the two population feature selection problem using probability of correct classification as a criterion

Variational equations are presented for maximizing the probability of correct classification as a function of a 1xn feature selection matrix B for the two-population problem. For the special case of equal covariance matrices the optimal B is unique up to scalar multiples and rank one sufficient. For equal population means, the best 1xn B is an eigenvector corresponding either to the largest or smallest eigenvalue of sigma sub 2 to the minus 1 power sigma sub 1 where sigma sub 1 and sigma sub 2 are the nxn covariance matrices of the two populations. The transformed probability of correct classification depends only on the eigenvalue. Finally, a procedure is proposed for constructing an optimal or nearly optimal kxn matrix of rank k without solving the k-dimensional variational equation.

Peters, B. C.↗

KIPSE1: A Knowledge-based Interactive Problem Solving Environment for data estimation and pattern classification

A knowledge-based interactive problem solving environment called KIPSE1 is presented. The KIPSE1 is a system built on a commercial expert system shell, the KEE system. This environment gives user capability to carry out exploratory data analysis and pattern classification tasks. A good solution often consists of a sequence of steps with a set of methods used at each step. In KIPSE1, solution is represented in the form of a decision tree and each node of the solution tree represents a partial solution to the problem. Many methodologies are provided at each node to the user such that the user can interactively select the method and data sets to test and subsequently examine the results. Otherwise, users are allowed to make decisions at various stages of problem solving to subdivide the problem into smaller subproblems such that a large problem can be handled and a better solution can be found.

Han, Chia Yung↗

Application of normal mode theory to seismic source and structure problems: Seismic investigations of upper mantle lateral heterogeneity

The theory of the normal modes of the earth is investigated and used to build synthetic seismograms in order to solve source and structural problems. A study is made of the physical properties of spheroidal modes leading to a rational classification. Two problems addressed are the observability of deep isotropic seismic sources and the investigation of the physical properties of the earth in the neighborhood of the Core-Mantle boundary, using SH waves diffracted at the core's surface. Data sets of seismic body and surface waves are used in a search for possible deep lateral heterogeneities in the mantle. In both cases, it is found that seismic data do not require structural differences between oceans and continents to extend deeper than 250 km. In general, differences between oceans and continents are found to be on the same order of magnitude as the intrinsic lateral heterogeneity in the oceanic plate brought about by the aging of the oceanic lithosphere.

Okal, E. A.↗

An Iterative Approach to the Feature Selection Problem

The problem dealt with concerns feature selection or reducing the dimension of the data to be processed from n to k. By reducing the dimension of the data from n to k, classification time is generally reduced. Yet the dimension reduction should not be so great that classification accuracy is impaired. Thus, the general problem is considered of classifying an n-dimensional observation vector x into one of m-distinct classes where each class is normally distributed with mean and covariance. It is shown that the probability of misclassification is minimized if a maximum likelihood classification procedure is used to classify the data. The dimension of each observation vector to be processed is conveniently reduced by performing the transformation y = Bx, where B is a K by n matrix of rank k. Thus, the n-dimensional classification problem transforms into a k-dimensional classification problem.

Decell, H. P., Jr.↗

Minimum distance classification in remote sensing

The utilization of minimum distance classification methods in remote sensing problems, such as crop species identification, is considered. Literature concerning both minimum distance classification problems and distance measures is reviewed. Experimental results are presented for several examples. The objective of these examples is to: (a) compare the sample classification accuracy of a minimum distance classifier, with the vector classification accuracy of a maximum likelihood classifier, and (b) compare the accuracy of a parametric minimum distance classifier with that of a nonparametric one. Results show the minimum distance classifier performance is 5% to 10% better than that of the maximum likelihood classifier. The nonparametric classifier is only slightly better than the parametric version.

Wacker, A. G.↗

Multipole radiation in charged-particle scattering

This paper formulates the general problem of photon emission in particle scattering using a classical and quantum mechanical approach. The connection between the classical short collision time (SCT) and Born results is examined for various special classifications of problems. In the dipole case the two formulations yield results that can be expressed in the same form and for arbitrary scattering potential. For quadrupole emission the SCT and Born results are the same only for a short-range potential, however. The quadrupole problem is more sensitive to details in the process because the calculation requires an expansion of the total amplitude for the process to lowest order in the photon wave number or momentum. The special case of photon emission associated with spin-flip transitions during scattering is considered for spin-1/2 particles. Like classical magnetic dipole radiation, there is no infrared divergence feature for this type of emission.

Gould, Robert J.↗

Software support for irregular and loosely synchronous problems

A large class of scientific and engineering applications may be classified as irregular and loosely synchronous from the perspective of parallel processing. We present a partial classification of such problems. This classification has motivated us to enhance FORTRAN D to provide language support for irregular, loosely synchronous problems. We present techniques for parallelization of such problems in the context of FORTRAN D.

Choudhary, A.↗

A stochastic atmospheric model for remote sensing applications

There are many factors which reduce the accuracy of classification of objects in the satellite remote sensing of Earth's surface. One important factor is the variability in the scattering and absorptive properties of the atmospheric components such as particulates and the variable gases. For multispectral remote sensing of the Earth's surface in the visible and infrared parts of the spectrum the atmospheric particulates are a major source of variability in the received signal. It is difficult to design a sensor which will determine the unknown atmospheric components by remote sensing methods, at least to the accuracy needed for multispectral classification. The problem of spatial and temporal variations in the atmospheric quantities which can affect the measured radiances are examined. A method based upon the stochastic nature of the atmospheric components was developed, and, using actual data the statistical parameters needed for inclusion into a radiometric model was generated. Methods are then described for an improved correction of radiances. These algorithms will then result in a more accurate and consistent classification procedure.

Turner, R. E.↗

Preprocessing for Unintended Conducted Emissions Classification with ResNet

Characterization of Unintended Conducted Emissions (UCE) from electronic devices is important when diagnosing electromagnetic interference, performing nonintrusive load monitoring (NILM) of power systems, and monitoring electronic device health, among other applications. Prior work has demonstrated that UCE analysis can serve as a diagnostic tool for energy efficiency investigations and detailed load analysis. While explaining the feature selection of deep networks with certainty is often not fully comprehensive, or in other applications, quite lacking, additional tools/methods for further corroboration and confirmation can help further the understanding of the researcher. This is true especially in the subject application of the study in this paper. Often the focus of such efforts is the selected features themselves, and there is not as much understanding gained about the noise in the collected data. If selected feature and noise characteristics are known, it can be used to further shape the design of the deep network or associated preprocessing. This is additionally difficult when the available data are limited, as in the case which the authors investigated in this study. Here, the authors present a novel work (which is a proposed complementary portion of the overall solution to the deep network classification explainability problem for this application) by applying a systematic progression of preprocessing and a deep neural network (ResNet architecture) to classify UCE data obtained via current transformers. By using a methodical application of preprocessing techniques prior to a deep classifier, hypotheses can be produced concerning what features the deep network deems important relative to what it perceives as noise. For instance, it is hypothesized in this particular study as a result of execution of the proposed method and periodic inspection of the classifier output that the UCE spectral features are relatively close to each other or to the interferers, as systematically reducing the beta parameter of the Kaiser window produced progressively better classification performance, but only to a point, as going below the Beta of eight produced decreased classifier performance, as well as the hypothesis that further spectral feature resolution was not as important to the classifier as rejection of the leakage from a spectrally distant interference. This can be very important in unpredictable low-FNR applications, where knowing the difference between features and noise is difficult. As a side-benefit, much was learned regarding the best preprocessing to use with the selected deep network for the UCE collected from these low power consumer devices obtained via current transformers. Baseline rectangular windowed FFT preprocessing provided a 62% classification increase versus using raw samples. After performing a more optimal preprocessing, more than 90% classification accuracy was achieved across 18 low-power consumer devices for scenarios in which the in-band features-to-noise ratio (FNR) was very poor.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Quantifying the atomistic free-volume morphology of materials with graph theory

Here, we introduce a new computational methodology for the identification and characterization of free volume within/around atomistic configurations. This scheme employs a three-stage workflow, by which spheres are iteratively grown inside of voxels, and ultimately converted to planar graphs, which are then characterized via a graph-based order parameter. Our approach is computationally efficient, physically intuitive, and universally transferable to any material system. Validation of our methodology is performed on several sets of materials problems: (1) classification of unique free volumes in various crystal phases, (2) autonomous detection and classification of complex surface defects during epitaxial growth simulations, (3) characterization of free volume defects in metals/alloys, and (4) quantification of the spatio-temporal behavior of nano-scale free volume morphologies as a function of both temperature and free-volume size. Our method accurately identifies and characterizes unique free volumes over a multitude of systems and length scales, indicating its potential for future use in understanding the relationship between free volume morphology and material properties under both static and dynamic conditions.

36 MATERIALS SCIENCE↗

Topics in inference and decision-making with partial knowledge

Two essential elements needed in the process of inference and decision-making are prior probabilities and likelihood functions. When both of these components are known accurately and precisely, the Bayesian approach provides a consistent and coherent solution to the problems of inference and decision-making. In many situations, however, either one or both of the above components may not be known, or at least may not be known precisely. This problem of partial knowledge about prior probabilities and likelihood functions is addressed. There are at least two ways to cope with this lack of precise knowledge: robust methods, and interval-valued methods. First, ways of modeling imprecision and indeterminacies in prior probabilities and likelihood functions are examined; then how imprecision in the above components carries over to the posterior probabilities is examined. Finally, the problem of decision making with imprecise posterior probabilities and the consequences of such actions are addressed. Application areas where the above problems may occur are in statistical pattern recognition problems, for example, the problem of classification of high-dimensional multispectral remote sensing image data.

Safavian, S. Rasoul↗

Spectral simulations of reacting turbulent flows

Spectral methods for simulating flows are reviewed, emphasizing their recent applications to reacting flow problems. Various classifications of spectral methods and their convergence properties are described and the 'spectral element' method is presented, highlighting its flexibility in dealing with complex flow geometries. The applications considered include chemical reactions in homogeneous turbulence, temporally evolving mixing layers, variable-density simulations, nonequilibrium chemistry, and spatially evolving mixing layers.

Mcmurtry, Patrick A.↗

Data Science and the Knowledge Discovery Adventure

This talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)

Data science↗

Machine Learning-Based Identification of the Interface Regions for Coupling Local and Nonlocal Models

Local-nonlocal coupling approaches provide a means to combine the computational efficiency of local models and the accuracy of nonlocal models. However, the coupling process can be challenging, requiring expertise to identify the interface between local and nonlocal regions. Here, this study introduces a machine learning-based approach to automatically detect the regions in which the local and nonlocal models should be used. The method uses loading functions evaluated at grid points to decide the model selection at those points. Training of the networks is based on datasets provided by classes of loading functions for which reference coupling configurations are computed using accurate coupled solutions, where accuracy is measured in terms of the relative error between the solution to the coupling approach and the solution to the nonlocal model. We study two approaches that vary in data structure. The first, the full-domain input data approach, uses the entire load vector and outputs a complete label vector, performing a global classification. The second, a window-based approach, processes loads into windows and addresses the problem as a node-wise classification where each window's central point is classified individually. The classification problems are solved via deep learning algorithms based on convolutional neural networks. The performance of these approaches is studied on one-dimensional numerical examples using F1-scores and accuracy metrics. Notably, the windowing approach achieves an accuracy of 0.96 and an F1-score of 0.97, highlighting its potential to automate coupling processes effectively and enhance computational efficiency in material science applications.

97 MATHEMATICS AND COMPUTING↗

Image feature extraction and galaxy classification: a novel and efficient approach with automated machine learning

ABSTRACT In this work, we explore the possibility of applying machine learning methods designed for 1D problems to the task of galaxy image classification. The algorithms used for image classification typically rely on multiple costly steps, such as the point spread function deconvolution and the training and application of complex Convolutional Neural Networks of thousands or even millions of parameters. In our approach, we extract features from the galaxy images by analysing the elliptical isophotes in their light distribution and collect the information in a sequence. The sequences obtained with this method present definite features allowing a direct distinction between galaxy types. Then, we train and classify the sequences with machine learning algorithms, designed through the platform Modulos AutoML. As a demonstration of this method, we use the second public release of the Dark Energy Survey (DES DR2). We show that we are able to successfully distinguish between early-type and late-type galaxies, for images with signal-to-noise ratio greater than 300. This yields an accuracy of $86{{\ \rm per\ cent}}$ for the early-type galaxies and $93{{\ \rm per\ cent}}$ for the late-type galaxies, which is on par with most contemporary automated image classification approaches. The data dimensionality reduction of our novel method implies a significant lowering in computational cost of classification. In the perspective of future data sets obtained with e.g. Euclid and the Vera Rubin Observatory, this work represents a path towards using a well-tested and widely used platform from industry in efficiently tackling galaxy classification problems at the peta-byte scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Uncertainty quantification in machine learning for engineering design and health prognostics: A tutorial

On top of machine learning (ML) models, uncertainty quantification (UQ) functions as an essential layer of safety assurance that could lead to more principled decision making by enabling sound risk assessment and management. The safety and reliability improvement of ML models empowered by UQ has the potential to significantly facilitate the broad adoption of ML solutions in high-stakes decision settings, such as healthcare, manufacturing, and aviation, to name a few. In this tutorial, we aim to provide a holistic lens on emerging UQ methods for ML models with a particular focus on neural networks and the applications of these UQ methods in tackling engineering design as well as prognostics and health management problems. Towards this goal, we start with a comprehensive classification of uncertainty types, sources, and causes pertaining to UQ of ML models. Next, we provide a tutorial-style description of several state-of-the-art UQ methods: Gaussian process regression, Bayesian neural network, neural network ensemble, and deterministic UQ methods focusing on spectral-normalized neural Gaussian process. Established upon the mathematical formulations, we subsequently examine the soundness of these UQ methods quantitatively and qualitatively (by a toy regression example) to examine their strengths and shortcomings from different dimensions. Then, we review quantitative metrics commonly used to assess the quality of predictive uncertainty in classification and regression problems. Afterward, we discuss the increasingly important role of UQ of ML models in solving challenging problems in engineering design and health prognostics. In conclusion, two case studies with source codes available on GitHub are used to demonstrate these UQ methods and compare their performance in the life prediction of lithium-ion batteries at the early stage (case study 1) and the remaining useful life prediction of turbofan engines (case study 2).

97 MATHEMATICS AND COMPUTING↗

Investigation of Error Patterns in Geographical Databases

The objective of the research conducted in this project is to develop a methodology to investigate the accuracy of Airport Safety Modeling Data (ASMD) using statistical, visualization, and Artificial Neural Network (ANN) techniques. Such a methodology can contribute to answering the following research questions: Over a representative sampling of ASMD databases, can statistical error analysis techniques be accurately learned and replicated by ANN modeling techniques? This representative ASMD sample should include numerous airports and a variety of terrain characterizations. Is it possible to identify and automate the recognition of patterns of error related to geographical features? Do such patterns of error relate to specific geographical features, such as elevation or terrain slope? Is it possible to combine the errors in small regions into an error prediction for a larger region? What are the data density reduction implications of this work? ASMD may be used as the source of terrain data for a synthetic visual system to be used in the cockpit of aircraft when visual reference to ground features is not possible during conditions of marginal weather or reduced visibility. In this research, United States Geologic Survey (USGS) digital elevation model (DEM) data has been selected as the benchmark. Artificial Neural Networks (ANNS) have been used and tested as alternate methods in place of the statistical methods in similar problems. They often perform better in pattern recognition, prediction and classification and categorization problems. Many studies show that when the data is complex and noisy, the accuracy of ANN models is generally higher than those of comparable traditional methods.

Dryer, David↗