Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning↗

Modular, Hierarchical Learning By Artificial Neural Networks

Modular and hierarchical approach to supervised learning by artificial neural networks leads to neural networks more structured than neural networks in which all neurons fully interconnected. These networks utilize general feedforward flow of information and sparse recurrent connections to achieve dynamical effects. The modular organization, sparsity of modular units and connections, and fact that learning is much more circumscribed are all attractive features for designing neural-network hardware. Learning streamlined by imitating some aspects of biological neural networks.

Baldi, Pierre F.↗

Design of Materials with Alchemite

Machine learning models that establish the relationships between materials processing and properties can enable inverse design of materials through active learning. Alchemite is a commercial software that can perform inverse materials design on sparse data. Here we evaluate Alchemite’s performance on a dataset of shape memory alloys and a dataset of heat exchangers compared to baseline random forest models. Alchemite had higher accuracy when making predictions on sparse data and was more accurate or nearly as accurate as random forests on complete datasets while also quantifying uncertainty. The software was also used to suggest processing steps and design parameters to optimize properties and performance; however, physical validation of the suggested design parameters was beyond the scope of this work. Several useful design insights were gained about the impact of the design parameters on properties and performance including the importance of dopant choice and amount for shape memory alloys and the importance of height and weight on the thermal resistance of heat exchangers.

Machine learning↗

Sparse Solutions for Single Class SVMs: A Bi-Criterion Approach

In this paper we propose an innovative learning algorithm - a variation of One-class nu Support Vector Machines (SVMs) learning algorithm to produce sparser solutions with much reduced computational complexities. The proposed technique returns an approximate solution, nearly as good as the solution set obtained by the classical approach, by minimizing the original risk function along with a regularization term. We introduce a bi-criterion optimization that helps guide the search towards the optimal set in much reduced time. The outcome of the proposed learning technique was compared with the benchmark one-class Support Vector machines algorithm which more often leads to solutions with redundant support vectors. Through out the analysis, the problem size for both optimization routines was kept consistent. We have tested the proposed algorithm on a variety of data sources under different conditions to demonstrate the effectiveness. In all cases the proposed algorithm closely preserves the accuracy of standard one-class nu SVMs while reducing both training time and test time by several factors.

Das, Santanu↗

Fast Solution in Sparse LDA for Binary Classification

An algorithm that performs sparse linear discriminant analysis (Sparse-LDA) finds near-optimal solutions in far less time than the prior art when specialized to binary classification (of 2 classes). Sparse-LDA is a type of feature- or variable- selection problem with numerous applications in statistics, machine learning, computer vision, computational finance, operations research, and bio-informatics. Because of its combinatorial nature, feature- or variable-selection problems are NP-hard or computationally intractable in cases involving more than 30 variables or features. Therefore, one typically seeks approximate solutions by means of greedy search algorithms. The prior Sparse-LDA algorithm was a greedy algorithm that considered the best variable or feature to add/ delete to/ from its subsets in order to maximally discriminate between multiple classes of data. The present algorithm is designed for the special but prevalent case of 2-class or binary classification (e.g. 1 vs. 0, functioning vs. malfunctioning, or change versus no change). The present algorithm provides near-optimal solutions on large real-world datasets having hundreds or even thousands of variables or features (e.g. selecting the fewest wavelength bands in a hyperspectral sensor to do terrain classification) and does so in typical computation times of minutes as compared to days or weeks as taken by the prior art. Sparse LDA requires solving generalized eigenvalue problems for a large number of variable subsets (represented by the submatrices of the input within-class and between-class covariance matrices). In the general (fullrank) case, the amount of computation scales at least cubically with the number of variables and thus the size of the problems that can be solved is limited accordingly. However, in binary classification, the principal eigenvalues can be found using a special analytic formula, without resorting to costly iterative techniques. The present algorithm exploits this analytic form along with the inherent sequential nature of greedy search itself. Together this enables the use of highly-efficient partitioned-matrix-inverse techniques that result in large speedups of computation in both the forward-selection and backward-elimination stages of greedy algorithms in general.

Moghaddam, Baback↗

A neural network with modular hierarchical learning

This invention provides a new hierarchical approach for supervised neural learning of time dependent trajectories. The modular hierarchical methodology leads to architectures which are more structured than fully interconnected networks. The networks utilize a general feedforward flow of information and sparse recurrent connections to achieve dynamic effects. The advantages include the sparsity of units and connections, the modular organization. A further advantage is that the learning is much more circumscribed learning than in fully interconnected systems. The present invention is embodied by a neural network including a plurality of neural modules each having a pre-established performance capability wherein each neural module has an output outputting present results of the performance capability and an input for changing the present results of the performance capabilitiy. For pattern recognition applications, the performance capability may be an oscillation capability producing a repeating wave pattern as the present results. In the preferred embodiment, each of the plurality of neural modules includes a pre-established capability portion and a performance adjustment portion connected to control the pre-established capability portion.

Baldi, Pierre F.↗

TPSAS-NF1676L-35322-DND

This talk discusses emerging methods that seek to fuse and integrate physics-based modeling with machine learning. With the recent rise of machine learning and artificial intelligence, there has been a huge surge in data-driven approaches to solve computational science and engineering problems. However, neglecting a priori knowledge of established physical laws and relying solely on data-driven methods can yield unreliable, less interpretable, and/or non-physical results, especially when data is sparse or predictions are required outside of the training data domain. This two part talk presents two distinct approaches for accelerating predictions with machine learning that are grounded and constrained by relevant physics and their application to problems at NASA.

Julian Cuevas Paniagua↗

CyberGAN: Generating High-fidelity Cybersecurity Data With Generative Adversarial Networks

Machine learning for cyber defense offers the promise of detecting adversarial activity against the ground data systems managing critical space assets. A fundamental challenge facing machine learning research in cybersecurity is the lack of high-fidelity, shareable datasets for robust evaluation and testing of machine learning-based solutions. High-fidelity, real-world datasets are necessary for reliable benchmarking of nominal system behavior and malicious activity. Unfortunately, such realistic datasets of both nominal and adversarial activity are rarely shared publicly by data owners due to security and privacy concerns. Besides, the available adversarial data is sparse, which makes training models on malicious activity much harder. This situation has impeded and continues to impede the research and successful adoption of machine learning methods for cyber defense. Researchers have dealt with this problem by generating data within a low-fidelity lab environment, using classified and thus unshareable datasets, or downloading low-fidelity public datasets made available by others. We propose an innovative solution to the problem by employing machine learning methods to generate high-fidelity data. Specifically, we propose the use of Generative Adversarial Networks (GANs) to generate high-fidelity data for cybersecurity purposes. GANs have found successful image processing and natural language applications, but have not yet been investigated for cyber data generation. Our proposed approach first involves training the `discriminator' network of the GAN with a sample of real-world data consisting of malicious and nominal samples. We then use the `generator' network to generate new high-fidelity data samples consisting of an appropriate mix of malicious and nominal activity. We demonstrate applications of our architecture by generating high-fidelity cybersecurity data containing both malicious and nominal samples. We thoroughly evaluate the fidelity of our generated data using heuristics and evaluate its usefulness for machine learning applications using three different datasets. Overall, our approach results in high-fidelity, shareable datasets.

Zhang, Yuening↗

An A-Train Climatology of Extratropical Cyclone Clouds

Extratropical cyclones (ETCs) are the main purveyors of precipitation in the mid-latitudes, especially in winter, and have a significant radiative impact through the clouds they generate. However, general circulation models (GCMs) have trouble representing precipitation and clouds in ETCs, and this might partly explain why current GCMs disagree on to the evolution of these systems in a warming climate. Collectively, the A-train observations of MODIS, CloudSat, CALIPSO, AIRS and AMSR-E have given us a unique perspective on ETCs: over the past 10 years these observations have allowed us to construct a climatology of clouds and precipitation associated with these storms. This has proved very useful for model evaluation as well in studies aimed at improving understanding of moist processes in these dynamically active conditions. Using the A-train observational suite and an objective cyclone and front identification algorithm we have constructed cyclone centric datasets that consist of an observation-based characterization of clouds and precipitation in ETCs and their sensitivity to large scale environments. In this presentation, we will summarize the advances in our knowledge of the climatological properties of cloud and precipitation in ETCs acquired with this unique dataset. In particular, we will present what we have learned about southern ocean ETCs, for which the A-train observations have filled a gap in this data sparse region. In addition, CloudSat and CALIPSO have for the first time provided information on the vertical distribution of clouds in ETCs and across warm and cold fronts. We will also discuss how these observations have helped identify key areas for improvement in moist processes in recent GCMs. Recently, we have begun to explore the interaction between aerosol and cloud cover in ETCs using MODIS, CloudSat and CALIPSO. We will show how aerosols are climatologically distributed within northern hemisphere ETCs, and how this relates to cloud cover.

clouds↗

Sparse Regression as a Sparse Eigenvalue Problem

We extend the l0-norm "subspectral" algorithms for sparse-LDA [5] and sparse-PCA [6] to general quadratic costs such as MSE in linear (kernel) regression. The resulting "Sparse Least Squares" (SLS) problem is also NP-hard, by way of its equivalence to a rank-1 sparse eigenvalue problem (e.g., binary sparse-LDA [7]). Specifically, for a general quadratic cost we use a highly-efficient technique for direct eigenvalue computation using partitioned matrix inverses which leads to dramatic x103 speed-ups over standard eigenvalue decomposition. This increased efficiency mitigates the O(n4) scaling behaviour that up to now has limited the previous algorithms' utility for high-dimensional learning problems. Moreover, the new computation prioritizes the role of the less-myopic backward elimination stage which becomes more efficient than forward selection. Similarly, branch-and-bound search for Exact Sparse Least Squares (ESLS) also benefits from partitioned matrix inverse techniques. Our Greedy Sparse Least Squares (GSLS) generalizes Natarajan's algorithm [9] also known as Order-Recursive Matching Pursuit (ORMP). Specifically, the forward half of GSLS is exactly equivalent to ORMP but more efficient. By including the backward pass, which only doubles the computation, we can achieve lower MSE than ORMP. Experimental comparisons to the state-of-the-art LARS algorithm [3] show forward-GSLS is faster, more accurate and more flexible in terms of choice of regularization

Exact Sparse Least Squares (ESLS)↗

Sparse distributed memory

Theoretical models of the human brain and proposed neural-network computers are developed analytically. Chapters are devoted to the mathematical foundations, background material from computer science, the theory of idealized neurons, neurons as address decoders, and the search of memory for the best match. Consideration is given to sparse memory, distributed storage, the storage and retrieval of sequences, the construction of distributed memory, and the organization of an autonomous learning system.

Kanerva, Pentti↗

Evidence of chaotic pattern in solar flux through a reproducible sequence of period-doubling-type bifurcations

Presented here is a preliminary study of the limits to solar flux intensity prediction, and of whether the general lack of predictability in the solar flux arises from the nonlinear chaotic nature of the Sun's physical activity. Statistical analysis of a chaotic signal can extract only its most gross features, and detailed physical models fail, since even the simplest equations of motion for a nonlinear system can exhibit chaotic behavior. A recent theory by Feigenbaum suggests that nonlinear systems that can be led into chaotic behavior through a sequence of period-doubling bifurcations will exhibit a universal behavior. As the control parameter is increased, the bifurcation points occur in such a way that a proper ratio of these will approach the universal Feigenbaum number. Experimental evidence supporting the applicability of the Feigenbaum scenario to solar flux data is sparse. However, given the hypothesis that the Sun's convection zones are similar to a Rayleigh-Bernard mechanism, we can learn a great deal from the remarkable agreement observed between the prediction by theory (period doubling - a universal route to chaos) and the amplitude decrease of the signal's regular subharmonics. The authors show that period-doubling-type bifurcation is a possible route to a chaotic pattern of solar flux that is distinguishable from the logarithm of its power spectral density. This conclusion is the first positive step toward a reformulation of solar flux by a nonlinear chaotic approach. The ultimate goal of this research is to be able to predict an estimate of the upper and lower bounds for solar flux within its predictable zones. Naturally, it is an important task to identify the time horizons beyond which predictability becomes incompatible with computability.

Ashrafi, S.↗

Evidence of chaotic pattern in solar flux through a reproducible sequence of period-doubling-type bifurcations

A preliminary study of the limits to solar flux intensity prediction, and of whether the general lack of predictability in the solar flux arises from the nonlinear chaotic nature of the Sun's physical activity is presented. Statistical analysis of a chaotic signal can extract only its most gross features, and detailed physical models fail, since even the simplest equations of motion for a nonlinear system can exhibit chaotic behavior. A recent theory by Feigenbaum suggests that nonlinear systems that can be led into chaotic behavior through a sequence of period-doubling bifurcations will exhibit a universal behavior. As the control parameter is increased, the bifurcation points occur in such a way that a proper ratio of these will approach the universal Feigenbaum number. Experimental evidence supporting the applicability of the Feigenbaum scenario to solar flux data is sparse. However, given the hypothesis that the Sun's convection zones are similar to a Rayleigh-Bernard mechanism, we can learn a great deal from the remarkable agreement observed between the prediction by theory (period doubling - a universal route to chaos) and the amplitude decrease of the signal's regular subharmonics. It is shown that period-doubling-type bifurcation is a possible route to a chaotic pattern of solar flux that is distinguishable from the logarithm of its power spectral density. This conclusion is the first positive step toward a reformulation of solar flux by a nonlinear chaotic approach. The ultimate goal of this research is to be able to predict an estimate of the upper and lower bounds for solar flux within its predictable zones. Naturally, it is an important task to identify the time horizons beyond which predictability becomes incompatible with computability.

Ashrafi, S.↗

Mars Atmospheric Dynamics

The Martian atmosphere is dynamically similar to the Earth's. Its spin-axis rotation rate is only minutes longer than Earth's so the Coriolois force is nearly identical to Earth's. The inclination of its spin axis is also similar to Earth's giving it similarity in seasonal change. And the Martian atmosphere is nearly transparent to solar radiation (except during dust periods) such that it is heated primarily by upwelling infrared radiation from the surface. These characteristics make Mars an ideal laboratory for studying the dynamics of rapidly rotating differentially heated atmospheres. This talk reviews what we have learned about Mars atmospheric dynamics and how if compares with Earth. The source of information to make such a comparison comes from observations and models. The former are sparse and that the latter have played a major role in shaping our thinking about the general circulation on Mars. However, the models need validation. Fortunately, the first two orbiters in NASA's Mars Surveyor Program have instrumentation to address many of the issues related to the general circulation and climate of Mars. The first, Mars Global Surveyor, is already at Mars gathering data. The second, the Mars 98 Orbiter to be launched later this year, carries a dedicated atmospheric sounder. Thus, much will be learned about Mars' atmosphere in the next few years.

Haberle, Robert↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Prediction of Aerodynamic Coefficients for Wind Tunnel Data using a Genetic Algorithm Optimized Neural Network

A fast, reliable way of predicting aerodynamic coefficients is produced using a neural network optimized by a genetic algorithm. Basic aerodynamic coefficients (e.g. lift, drag, pitching moment) are modelled as functions of angle of attack and Mach number. The neural network is first trained on a relatively rich set of data from wind tunnel tests of numerical simulations to learn an overall model. Most of the aerodynamic parameters can be well-fitted using polynomial functions. A new set of data, which can be relatively sparse, is then supplied to the network to produce a new model consistent with the previous model and the new data. Because the new model interpolates realistically between the sparse test data points, it is suitable for use in piloted simulations. The genetic algorithm is used to choose a neural network architecture to give best results, avoiding over-and under-fitting of the test data.

Rajkumar, T.↗

Biomimetic Models for An Ecological Approach to Massively-Deployed Sensor Networks

Promises of ubiquitous control of the physical environment by massively-deployed wireless sensor networks open avenues for new applications that will redefine the way we live and work. Due to small size and low cost of sensor devices, visionaries promise systems enabled by deployment of massive numbers of sensors ubiquitous throughout our environment working in concert. Recent research has concentrated on developing techniques for performing relatively simple tasks with minimal energy expense, assuming some form of centralized control. Unfortunately, centralized control is not conducive to parallel activities and does not scale to massive size networks. Execution of simple tasks in sparse networks will not lead to the sophisticated applications predicted. We propose a new way of looking at massively-deployed sensor networks, motivated by lessons learned from the way biological ecosystems are organized. We demonstrate that in such a model, fully distributed data aggregation can be performed in a scalable fashion in massively deployed sensor networks, where motes operate on local information, making local decisions that are aggregated across the network to achieve globally-meaningful effects. We show that such architectures may be used to facilitate communication and synchronization in a fault-tolerant manner, while balancing workload and required energy expenditure throughout the network.

Jones, Kennie H.↗