Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distance learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Experimental High Energy Physics at the University of Illinois

This research effort funded by the Office of High Energy Physics in the U.S. Department of Energy is aimed at exploring our universe at its most basic level. The goal of our effort is to learn more about how and why nature behaves the way it does. Within this effort, we are studying the smallest particles and the largest distances we can possibly observe. Our work in particular focuses on the search for new interactions that will help us understand how the universe came into being. Through this work, we collaborate with other scientists and engineers to develop new technologies that can be utilized throughout society. Benefits from our research include advances in medical technology, electronics, transportation, sustainability, information and computing as well as the advanced training of undergraduate and graduate students in science and engineering.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Adaptive fuzzy leader clustering of complex data sets in pattern recognition

A modular, unsupervised neural network architecture for clustering and classification of complex data sets is presented. The adaptive fuzzy leader clustering (AFLC) architecture is a hybrid neural-fuzzy system that learns on-line in a stable and efficient manner. The initial classification is performed in two stages: a simple competitive stage and a distance metric comparison stage. The cluster prototypes are then incrementally updated by relocating the centroid positions from fuzzy C-means system equations for the centroids and the membership values. The AFLC algorithm is applied to the Anderson Iris data and laser-luminescent fingerprint image data. It is concluded that the AFLC algorithm successfully classifies features extracted from real data, discrete or continuous.

Newton, Scott C.↗

Deep Learning Classification of Cheatgrass Invasion in the Western United States Using Biophysical and Remote Sensing Data

Cheatgrass (Bromus tectorum) invasion is driving an emerging cycle of increased fire frequency and irreversible loss of wildlife habitat in the western US. Yet, detailed spatial information about its occurrence is still lacking for much of its presumably invaded range. Deep learning (DL) has demonstrated success for remote sensing applications but is less tested on more challenging tasks like identifying biological invasions using sub-pixel phenomena. We compare two DL architectures and the more conventional Random Forest and Logistic Regression methods to improve upon a previous effort to map cheatgrass occurrence at >2% canopy cover. High-dimensional sets of biophysical, MODIS, and Landsat-7 ETM+ predictor variables are also compared to evaluate different multi-modal data strategies. All model configurations improved results relative to the case study and accuracy generally improved by combining data from both sensors with biophysical data. Cheatgrass occurrence is mapped at 30 m ground sample distance (GSD) with an estimated 78.1% accuracy, compared to 250-m GSD and 71% map accuracy in the case study. Furthermore, DL is shown to be competitive with well-established machine learning methods in a limited data regime, suggesting it can be an effective tool for mapping biological invasions and more broadly for multi-modal remote sensing applications.

54 ENVIRONMENTAL SCIENCES↗

Improving the Accuracy of Clustering Electric Utility Net Load Data using Dynamic Time Warping

Identifying patterns in electric utility net load data in a time-series format is very useful in preparing the operation for next day. Machine learning algorithms have been used in other domains and those concepts are applied in this paper on real-world net load measurement data. Clustering is the practice of grouping data with similar characteristics as determined by the distance measure. The K-means clustering algorithm is utilized here with actual electric utility data. The paper uses the standard distance measure, Euclidean distance (ED), and compares its performance against the dynamic time warping (DTW) measure. An actual case study with real data is presented, and DTW distance measure-based method observed to result better accuracy compared to the ED based method for substation net load measurements predominantly with residential customers.

clustering↗

Characterizing the Spread of COVID-19 from Human Mobility Patterns and SocioDemographic Indicators

Mobility is an indicator of human movement through space and time. With the increasing availability of geolocated data (from GPS, accelerometers, etc.), it is now possible to examine individual as well as group human mobility patterns. Human mobility is influenced by both intrinsic (i.e. personal motivations) and extrinsic (i.e., events like natural hazards or a pandemic like the COVID-19) factors. However, the intricate relationships between human mobility patterns and sociodemographic characteristics in the context of a pandemic are yet to be fully explored. Our goal is to overcome this gap by using human mobility data at the census block group level from mobile phones and combining those with social vulnerability indicators to examine the overall spread of COVID-19 at local spatial scales. We used 585,878 weekly visits to 37,871 points of interests (POIs) from Safegraph to quantify mobility indices and social distancing metrics in 2,820 census block groups in the city of Los Angeles (LA) - before and during lockdown as well as during the phase1 and phase 2 reopening. Finally, using supervised machine learning algorithms, we classified the census block groups in LA into High, Medium and Low categories that represented the vulnerability of these block groups based on the cumulative number of occurrences of COVID-19 cases till July 24, 2020. Our results indicate that the tree-based classifiers performed well in comparison to the Support Vector Machines and Multinomial Logit models. Gradient Boosting had the highest classification accuracy of 97.4% COVID-19 with an AUC score of 0.987. The block groups with high COVID-19 cases also had a high concentration of socially vulnerable populations, high human mobility index and a low social distancing index.

Roy, Avipsa↗

MarCO: Interplanetary Mission Development on a CubeSat Scale

Shortly after JPL’s Interior Exploration using Seismic Investigations, Geodesy and Heat Transport (InSight) mission launches, separates, and commences its cruise phase, two CubeSats will deploy from the launch vehicle’s upper stage and begin independent flight to Mars (Fig. 1). During InSight’s entry, descent, and landing (EDL) sequence, these twin Mars Cube One (MarCO) spacecraft will fly 3,500 km above the Martian surface, recording and relaying InSight UHF radio data to the Deep Space Network (DSN) on Earth1. MarCO is a twin CubeSat mission developed by the NASA Jet Propulsion Laboratory (JPL) to accompany the InSight (Interior Exploration using Seismic Investigations, Geodesy and Heat Transport) Mars mission lander. MarCO's primary mission objective is to launch with InSight and independently fly to Mars to serve as a communications relay during InSight's entry, descent, and landing (EDL) phase. MarCO represents a new type of deep space mission: CubeSats at Mars. Building on the development of JPL's first interplanetary CubeSat project, the Interplanetary Nano-Spacecraft Pathfinder in Relevant Environment (INSPIRE), MarCO further refined the approach to hardware, software, and ground architecture development to solve the challenges of quickly building low-budget spacecraft to fly to Mars. The greatest constraint, beyond others typical of CubeSat missions, was time. The duration between MarCO's conception to completion of spacecraft assembly was less than two years - an unprecedented schedule for any planetary mission to date. Through necessity, MarCO has built on previous experience, procedures, systems, and development methodologies, defining a new niche for supporting larger primary missions. The MarCO spacecraft are poised to write a new chapter in deep space exploration. Originally slated to launch and reach Mars in 2016, the InSight mission schedule subsequently slipped to 2018. During the original landing of InSight, Earth would not be in view, and no orbiter around Mars would have been in position to both receive UHF EDL data and simultaneously relay it back to Earth. It was from this obstacle that MarCO was conceived. Regardless of any changes to InSight’s 2018 EDL configuration geometry, MarCO is still expected to fly and serve in the same capacity as originally designed: the first CubeSat mission to Mars. CubeSats have historically been firmly in the domain of universities and small companies. As first conceived, they served as a platform upon which to teach all aspects of the space mission lifecycle. JPL took on this mission type with Interplanetary Nano-Spacecraft Pathfinder in Relevant Environment2 (INSPIRE), moving the concept into a new domain: deep space. Building from the INSPIRE platform and lessons learned, MarCO addressed new challenges in the domain of planetary missions: independent interplanetary flight and navigation, integration with a large-scale mission, long-distance and long-delay communication, short development time, and a small development team. Of these, the greatest constraint was schedule: only 18 months passed from conception of mission concept until delivery of fully assembled and tested flight hardware. Careful selection of mission team, along with extensive use of off-the-shelf equipment, and streamlining automated processes, was essential. This achievement represents the next step in the evolution of CubeSats beyond low-Earth orbit.

Werne, Thomas↗

BraggNN : fast X-ray Bragg peak analysis using deep learning

X-ray diffraction based microscopy techniques such as high-energy diffraction microscopy (HEDM) rely on knowledge of the position of diffraction peaks with high precision. These positions are typically computed by fitting the observed intensities in detector data to a theoretical peak shape such as pseudo-Voigt. As experiments become more complex and detector technologies evolve, the computational cost of such peak-shape fitting becomes the biggest hurdle to the rapid analysis required for real-time feedback in experiments. To this end, we propose BraggNN, a deep-learning based method that can determine peak positions much more rapidly than conventional pseudo-Voigt peak fitting. When applied to a test dataset, peak center-of-mass positions obtained from BraggNN deviate less than 0.29 and 0.57 pixels for 75 and 95% of the peaks, respectively, from positions obtained using conventional pseudo-Voigt fitting (Euclidean distance). When applied to a real experimental dataset and using grain positions from near-field HEDM reconstruction as ground-truth, grain positions using BraggNN result in 15% smaller errors compared with those calculated using pseudo-Voigt. Recent advances in deep-learning method implementations and special-purpose model inference accelerators allow BraggNN to deliver enormous performance improvements relative to the conventional method, running, for example, more than 200 times faster on a consumer-class GPU card with out-of-the-box software.

36 MATERIALS SCIENCE↗

Multiclass Continuous Correspondence Learning

We extend the Structural Correspondence Learning (SCL) domain adaptation algorithm of Blitzer er al. to the realm of continuous signals. Given a set of labeled examples belonging to a 'source' domain, we select a set of unlabeled examples in a related 'target' domain that play similar roles in both domains. Using these 'pivot samples, we map both domains into a common feature space, allowing us to adapt a classifier trained on source examples to classify target examples. We show that when between-class distances are relatively preserved across domains, we can automatically select target pivots to bring the domains into correspondence.

correspondence learning↗

Genarris 2.0: A Random Structure Generator for Molecular Crystals

Genarris is an open source Python package for generating random molecular crystal structures with physical constraints for seeding crystal structure prediction algorithms and training machine learning models. Here we present a new version of the code, containing several major improvements. A MPI-based parallelization scheme has been implemented, which facilitates the seamless sequential execution of user-defined workflows. A new method for estimating the unit cell volume based on the single molecule structure has been developed using a machine-learned model trained on experimental structures. A new algorithm has been implemented for generating crystal structures with molecules occupying special Wyckoff positions. A new hierarchical structure check procedure has been developed to detect unphysical close contacts efficiently and accurately. New intermolecular distance settings have been implemented for strong hydrogen bonds. To demonstrate these new features, we study two specific cases: benzene and glycine. Genarris finds the experimental structures of the two polymorphs of benzene and the three polymorphs of glycine. Program summary Program Title: Genarris 2.0 Program Files doi: http://dx.doi.org/10.17632/grx6mz4pjn.1 Licensing provisions: BSD-3 Clause Programming language: Python, C External routines/libraries: Spglib, ASE, pymatgen, SciPy, mpi4py, scikit-learn, PyTorch, FHI-aims. Nature of problem: Molecular crystal structure prediction. Solution method: Genarris 2.0 generates molecular crystal structures over the 230 space groups, on general and special Wyckoff positions, using physical constraints. Down-sampling of the generated structures may be performed subsequently, based on molecular crystal packing descriptors and an unsupervised machine learning algorithm. Lastly, ab initio structure relaxation may be performed for the final pool. Depending on the user-defined workflow implemented, Genarris may be used to generate diverse molecular crystal datasets to seed evolutionary algorithms or to train machine learning algorithms or as a standalone crystal structure prediction method. Restrictions: For crystal structure generation, the molecule of interest must be semi-rigid with no bond rotational degrees of freedom. Unusual features: Genarris 2.0 is a highly distributed program, making use of MPI for Python parallelization. The user has the ability to design and implement workflows by executing a user-defined list of procedures. Genarris 2.0 offers new features including a machine learning model for estimating the molecular volume in the solid state from the single molecule structure, structure generation in special Wyckoff positions of space groups, hierarchical structure checks including rigorous treatment of non-orthogonal structures, and clustering and down-selection workflows combining first principles simulations with machine learning. (C) 2020 Elsevier B.V. All rights reserved.

Crystal structure prediction↗

Chasing Accreted Structures within Gaia DR2 Using Deep Learning

In previous work, we developed a deep neural network classifier that only relies on phase-space information to obtain a catalog of accreted stars based on the second data release of Gaia (DR2). In this paper, we apply two clustering algorithms to identify velocity substructure within this catalog. We focus on the subset of stars with line-of-sight velocity measurements that fall in the range of Galactocentric radii $r\in [6.5,9.5]\,{\rm{kpc}}$ and vertical distances $| z| \lt 3\,{\rm{kpc}}$. Known structures such as Gaia Enceladus and the Helmi stream are identified. The largest previously unknown structure, Nyx, is a vast stream consisting of at least 200 stars in the region of interest. This study displays the power of the machine-learning approach by not only successfully identifying known features but also discovering new kinematic structures that may shed light on the merger history of the Milky Way.

Astronomy & Astrophysics↗

Recent Advances in Machine Learning for Fiber Optic Sensor Applications

Over the last three decades, fiber optic sensors (FOS) have gained a lot of attention for their wide range of monitoring applications across many industries, including aerospace, defense, security, civil engineering, and energy. FOS technologies hold great promise to form the backbone for next‐generation intelligent sensing platforms that offer long‐distance, high‐accuracy, distributed measurement capabilities and multiparametric monitoring with resilience to harsh environmental conditions. The major limitations posed by FOS are 1) cross‐sensitivity, 2) enormous volume and large data generation, 3) low data processing speed, 4) degradation of signal‐to‐noise ratio over the fiber length, and 5) overall cost of sensor and interrogator systems. These challenges can be overcome by building advanced data analytics engines enabled by recent breakthroughs in machine learning (ML) and artificial intelligence (AI). This article presents a comprehensive review of recent studies that integrate ML and AI algorithms with FOS technologies. This review also highlights several FOS technology development directions that promise a significant impact on widespread use for several industrial applications, with an emphasis on energy systems monitoring. A perspective on future directions for further research development is also provided.

97 MATHEMATICS AND COMPUTING↗

DiffLense: a conditional diffusion model for super-resolution of gravitational lensing data

Abstract Gravitational lensing data is frequently collected at low resolution due to instrumental limitations and observing conditions. Machine learning-based super-resolution techniques offer a method to enhance the resolution of these images, enabling more precise measurements of lensing effects and a better understanding of the matter distribution in the lensing system. This enhancement can significantly improve our knowledge of the distribution of mass within the lensing galaxy and its environment, as well as the properties of the background source being lensed. Traditional super-resolution techniques typically learn a mapping function from lower-resolution to higher-resolution samples. However, these methods are often constrained by their dependence on optimizing a fixed distance function, which can result in the loss of intricate details crucial for astrophysical analysis. In this work, we introduce DiffLense , a novel super-resolution pipeline based on a conditional diffusion model specifically designed to enhance the resolution of gravitational lensing images obtained from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). Our approach adopts a generative model, leveraging the detailed structural information present in Hubble space telescope (HST) counterparts. The diffusion model, trained to generate HST data, is conditioned on HSC data pre-processed with denoising techniques and thresholding to significantly reduce noise and background interference. This process leads to a more distinct and less overlapping conditional distribution during the model’s training phase. We demonstrate that DiffLense outperforms existing state-of-the-art single-image super-resolution techniques, particularly in retaining the fine details necessary for astrophysical analyses.

Computer Science↗

Process Optimization of Carbon Electrode Materials Manufacturing by Experimental Study and Machine Learning Techniques

Electrospun carbon fibers from coal have been investigated as electrodes for batteries and supercapacitors. Despite the excellent properties of coal-derived carbon fibers (CCNF) for energy storage devices, there still lacks systematic understanding on how various process parameters affect final electrode performances, which poses challenges to scale from pilot to high volume manufacturing. The goals of this project are twofold. First, we focuse on process optimization for converting a new precursor from powder river basin (PRB) coal, referred to as coal-based polyurethane (CPU) to CCNF using electrospinning. Second, different machine learning techniques will be examined using experimental data from this work and open literature. Specifically, for CPU the following process parameters need to be characterized and optimized in order to produce CCNFs with desirable mechanical integrity and physiochemical properties: precursor composition and viscosity, operating voltage and distance, oxidation and carbonization temperature and duration. Consequently, physiochemical properties of the fibers were characterized to correlate these process parameters with desirable electrochemical performance. Given the complex nature of the fiber production process, ML models are assessed for their ability to capture the nonlinear relationship between process parameters and the electrochemical properties in applications including supercapacitors. As such, we applied various machine learning techniques, to determine which technique produces a model that best predicts device function.

Cincotta, Robert E.F.↗

Q-Cluster: Quantum Error Mitigation Through Noise-Aware Unsupervised Learning

Quantum error mitigation (QEM) is critical in reducing the impact of noise in the pre-fault-tolerant era, and is expected to complement error correction in fault-tolerant quantum computing (FTQC). In this work, we propose a novel QEM approach, Q-Cluster, that uses unsupervised learning (clustering) to reshape the measured bit-string distribution. Our approach starts with a simplified bit-flip noise model. It first performs clustering on noisy measurement results, i.e., bit-strings, based on the Hamming distance. The centroid of each cluster is calculated using a qubit-wise majority vote. Next, the noisy distribution is adjusted with the clustering outcomes and the bitflip error rates using Bayesian inference. Our simulation results show that Q-Cluster can mitigate high noise rates (up to 40% per qubit) with the simple bit-flip noise model. However, real quantum computers do not fit such a simple noise model. To address the problem, we (a) apply Pauli twirling to tailor the complex noise channels to Pauli errors, and (b) employ a machine learning model, ExtraTrees regressor, to estimate an effective bit-flip error rate using a feature vector consisting of machine calibration data (gate & measurement error rates), circuit features (number of qubits, numbers of different types of gates, etc.) and the shape of the noisy distribution (entropy). Our experimental results show that our proposed Q-Cluster scheme improves the fidelity by a factor of 1.46x, on average, compared to the unmitigated output distribution, for a set of low-entropy benchmarks on five different IBM quantum machines. Our approach outperforms the state-of-art QEM approaches RZNE [28], M3 [24], Hammer [35], and QBEEP [33] by 1.26x,1.29x,1.47x, and 2.65 x, respectively.

42 ENGINEERING↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Data Production on Past and Future NASA Missions

Data return is a metric that is commonly publicized for all space science missions. In the early days of the Space Program, this figure was small, and could be described in bits or maybe even megabits. But now, missions are capable of returning data volumes two or three orders of magnitude larger. For example, Voyager 1 and 2 combined produced a little over 5 Terabits of data in 39 years of operation. In contrast, the Cassini mission, launched two decades after Voyager, produced about one and a half times those data volumes in half the time. NISAR, an Earth Science Mission currently in implementation, plans to produce over 28 Petabits of raw data in just 3 years. This means that NISAR will produce about as many data in 30 days as the combined data production of nearly all planetary missions to date. These increases in capability are a result of technology enhancements in two main areas: telecommunications architecture (both space and ground segments) and data storage technology. This paper describes the progression of these two technologies over the course of more than three decades of space missions and provides additional insight into the design of the end-to-end NISAR Data System Architecture. Trends in the data are briefly explored and compared to Moore’s Law which provides only a qualitative model for memory growth but not for data production. In summary, early missions are found to be driven by unrefined processes while later missions, having utilized earlier lessons learned, focus more on improvements to flight and ground capabilities. Data return seems to fall into three categories. First, deep space missions are driven by the large distances that limit data return to the Earth. Next, the orbiter infrastructure around Mars helps these missions generate more data than other deep space spacecraft. Finally, near-Earth missions have the greatest capabilities for the studied metrics due to their close proximity to Earth and the ground network availability.

Xaypraseuth, Peter↗

Improving multiwell petrophysical interpretation from well logs via machine learning and statistical models

Well-log interpretation estimates in situ rock properties along well trajectory, such as porosity, water saturation, and permeability, to support reserve-volume estimation, production forecasts, and decision making in reservoir development. However, due to measurement errors, variability of well logs caused by multiple measurement vendors, different borehole tools, and nonuniform drilling/borehole conditions, estimations of rock properties with original well logs without proper preprocessing may not be accurate, especially in the context of multiwell estimation. Well-log normalization techniques such as two-point scaling and mean-variance normalization are commonly used to improve the robustness of multiwell rock-property estimation. However, these techniques do not consider the correlation between well logs and require subjective knowledge for their effective implementation. To reduce uncertainties and processing time associated with multiwell rock-property estimation from well logs, we develop discriminative adversarial (DA) and linear constraint models for well-log normalization and rock-property estimation. The DA neural network model developed for well-log normalization and interpretation can perform linear and nonlinear well-log normalization while considering the joint distribution of each well log and rock properties. However, the linear constraint model uses an ensemble of predictions from linear models to constrain well-log normalization and rock-property estimation. We also develop a divergence-based type well identification method to select type (training) wells for a test well based on the statistical similarity of associated well-log distributions instead of the interwell distance. We apply the DA model to perform well-log normalization and prediction of permeability for the Seminole San Andres Unit carbonate reservoir. Compared with the permeability predicted with the classical machine learning model without well-log normalization and models with two-point scaling normalization, the DA model yields the most accurate permeability prediction by decreasing the mean-squared error of permeability prediction by 20%–50%.

Geochemistry & Geophysics↗