Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distance learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Beyond the hubble sequence – exploring galaxy morphology with unsupervised machine learning

We explore unsupervised machine learning for galaxy morphology analyses using a combination of feature extraction with a vector-quantized variational autoencoder (VQ-VAE) and hierarchical clustering (HC). We propose a new methodology that includes: (1) consideration of the clustering performance simultaneously when learning features from images; (2) allowing for various distance thresholds within the HC algorithm; (3) using the galaxy orientation to determine the number of clusters. This set-up provides 27 clusters created with this unsupervised learning that we show are well separated based on galaxy shape and structure (e.g. Sérsic index, concentration, asymmetry, Gini coefficient). These resulting clusters also correlate well with physical properties such as the colour–magnitude diagram, and span the range of scaling relations such as mass versus size amongst the different machine-defined clusters. When we merge these multiple clusters into two large preliminary clusters to provide a binary classification, an accuracy of $\sim 87{{\ \rm per\ cent}}$ is reached using an imbalanced data set, matching real galaxy distributions, which includes 22.7 per cent early-type galaxies and 77.3 per cent late-type galaxies. Comparing the given clusters with classic Hubble types (ellipticals, lenticulars, early spirals, late spirals, and irregulars), we show that there is an intrinsic vagueness in visual classification systems, in particular galaxies with transitional features such as lenticulars and early spirals. Based on this, the main result in this work is not how well our unsupervised method matches visual classifications and physical properties, but that the method provides an independent classification that may be more physically meaningful than any visually based ones.

79 ASTRONOMY AND ASTROPHYSICS↗

Thermodynamically Optimized Machine-Learned Reaction Coordinates for Hydrophobic Ligand Dissociation

Ligand unbinding is mediated by its free energy change, which has intertwined contributions from both energy and entropy. It is important, but not easy, to quantify their individual contributions to the free energy profile. We model hydrophobic ligand unbinding for two systems, a methane particle and a C 60 fullerene, both unbinding from hydrophobic pockets in all-atom water. Using a modified deep learning framework, we learn a thermodynamically optimized reaction coordinate to describe the hydrophobic ligand dissociation for both systems. Interpretation of these reaction coordinates reveals the roles of entropic and enthalpic forces as the ligand and pocket sizes change. In both cases, we observe that the free-energy barrier to unbinding is dominated by entropy considerations. Furthermore, the process of methane unbinding is driven by methane solvation, while fullerene unbinding is driven first by pocket wetting and then fullerene wetting. For both solutes, the direct importance of the distance from the binding pocket to the learned reaction coordinate is present, but low. Furthermore, our framework and subsequent feature important analysis thus give useful thermodynamic insight into hydrophobic ligand dissociation problems that are otherwise difficult to glean.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Circular Velocity Curve of the Milky Way from 5–25 kpc Using Luminous Red Giant Branch Stars

We present a sample of 254,882 luminous red giant branch (LRGB) stars selected from the APOGEE and LAMOST surveys. By combining photometric and astrometric information from the Two Micron All Sky Survey and Gaia survey, the precise distances of the sample stars are determined by a supervised machine-learning algorithm: the gradient-boosted decision trees. To test the accuracy of the derived distances, member stars of globular clusters (GCs) and open clusters are used. The tests by cluster member stars show a precision of about 10% with negligible zero-point offsets, for the derived distances of our sample stars. The final sample covers a large volume of the Galactic disk(s) and halo of 0 < R < 30 kpc and |Z| ≤ 15 kpc. The rotation curve (RC) of the Milky Way across the radius of 5 ≲ R ≲ 25 kpc has been accurately measured with ~54,000 stars of the thin disk population selected from the LRGB sample. The derived RC shows a weak decline along R with a gradient of -1.83 ± 0.02 (stat.) ± 0.07 (sys.) km s -1 kpc -1 , in excellent agreement with the results measured by previous studies. The circular velocity at the solar position, yielded by our RC is 234.04 ± 0.08 (stat.) ± 1.36 (sys.) km s -1 , again in great consistency with other independent determinations. From the newly constructed RC, as well as constraints from other data, we have constructed a mass model for our Galaxy, yielding a mass of the dark matter halo of M 200 = (8.05 ± 1.15) × 10 11 M ⊙ with a corresponding radius of R 200 = 192.37 ± 9.24 kpc and a local dark matter density of 0.39 ± 0.03 GeV cm -3 .

79 ASTRONOMY AND ASTROPHYSICS↗

Simulated meteorological impacts of offshore wind turbines and sensitivity to the amount of added turbulence kinetic energy

Offshore wind energy projects are currently in development off the east coast of the United States and may influence the local meteorology of the region. Wind power production and other commercial uses in this area are related to atmospheric conditions, and so it is important to understand how future wind plants may change the local meteorology. In the absence of measurements of potential wind plant impacts on meteorology, simulations offer the next-best possible insight into wake effects on boundary layer height, temperature, fluxes, and wind speeds. However, simulation tools that capture these effects offer multiple options for representing the amount of turbine-added turbulence that may impact assessments of micrometeorological effects. To explore this sensitivity, we compare 1 year of simulations from the Weather Research and Forecasting (WRF) model with and without wind plants incorporated, focusing on the lease area south of Massachusetts and Rhode Island. The simulations with wind plants are repeated to include both the maximum and minimum amounts of added turbulence to provide bounds on the potential impacts. We assess changes in wind speeds, 2 m temperature, surface heat flux, turbulence kinetic energy (TKE), and boundary layer height during different stability classifications and ambient wind speeds over the entire year and compare results for the degree of added turbulence in the wind plant simulations. Because the wake behavior may be a function of boundary layer stability, in this paper, we also present a machine learning algorithm to quantify the area and distance of the wake generated by the wind plant. This analysis enables us to identify the relationship between wake extent and boundary layer height. We find that hub-height wind speed is reduced within and downwind of the wind plant, with the strongest impacts occurring during stable conditions and faster wind speeds in region 3 of the turbine power curve, although impacts lessen as wind speeds increase past 15 m s−1. In contrast, wind speeds near the surface decrease when no turbine-added turbulence is included but can increase for stably stratified conditions when 100 % of possible TKE is included in the simulations. TKE increases at hub height in the simulations with added TKE for all stability classes, suggesting that atmospheric stability does not immediately modify the TKE generated by turbines. Negligible changes in hub-height TKE manifest in the simulations without the added TKE. At the surface, TKE increases in the simulations with maximum added turbulence only for unstable conditions. In the no-added-turbulence simulations, surface TKE decreases slightly in neutral and unstable simulations. Differences in 2 m temperatures and surface heat fluxes are small but vary considerably with atmospheric stability and the amount of added TKE. Boundary layer heights increase within the wind plant when turbine-added turbulence is included and decrease slightly downwind during stable conditions. In contrast, with no added turbulence, the boundary layer height is in general reduced in stable conditions with wind speeds less than 15 m s −1 and slightly increased in neutral conditions. Finally, shallower upwind boundary layer heights tend to correlate with larger wake areas and distances, though other factors likely also play a role in determining the extent of the wind plant wake. These simulation-based results provide a bound for micrometeorological impacts of wind plant wakes: simulations that couple the atmosphere to the ocean may reduce these impacts, and we await observational verification.

17 WIND ENERGY↗

Technical Performance and Cost Optimization of Unobtrusive Multi-static Serial LiDAR Imager (UMSLI) for Wide-area Surveillance and Identification of Marine Life at Marine Energy Installations

Florida Atlantic University developed an underwater optical monitoring system prototype - Unobtrusive Multi-static Serial LiDAR Imager (UMSLI) - suitable for marine energy full project lifecycle observation (baseline, commissioning, and decommissioning), with an automated real-time classification of marine animals. With precursor 2014 DOE funding (award DE-EE0006787), a prototype UMSLI was demonstrated in a controlled laboratory environment and achieved a TRL 6. This EERE DE-EE0007828 award aimed to both increase the TRL of the UMSLI by improving the technology performance (e.g., increase distance of marine animal target detection capability, add additional species classification capabilities, improve the system performance during the day, etc.) and reduce the cost. The UMSLI presents a novel application of underwater distributed Light Detection And Ranging (LiDAR), an advanced remote sensing method that uses light in the form of laser pulses for an application, this paired with an algorithm provides 360 degrees underwater detection, imaging, and classification of marine life. This solution for underwater monitoring of biota preserves the advantages of traditional optical and acoustic solutions while overcoming many associated disadvantages for marine energy site environmental monitoring, such as difficulties in species detection and classification in low light or night and turbid environments. This new approach is a purposefully designed, reconfigurable adaptation of an existing class 3B laser technology into one UMSLI instrument that can be easily mounted on or around different classes of marine energy equipment, such as devices to capture ocean current energy or wave energy. The system uses relatively low average power, utilizes far-red (> 635nm) laser illumination to be invisible and eye-safe to marine animals, is compact, and cost-effective. The equipment is designed for long-term, maintenance-free operations (i.e., current design is targeting more than 7 days continuous operation), to inherently generate a sparse primary dataset that only includes detected anomalies, and to allow robust real-time automated animal classification and identification with a low data bandwidth requirement. The technology’s overarching goal for application, is a system that can be deployed to collect pre-installation baseline species observations at a proposed marine energy deployment site with minimal post-processing overhead. The envisioned system will also produce high-resolution imagery of marine animals through a wide range of conditions and support automated tracking and notification of the presence of managed animals within established perimeters of marine energy equipment to satisfy deployed marine energy projects’ endangered and threatened species monitoring requirements. Through the current project, we demonstrated the UMSLI prototype in an operational environment and increased the UMSLI’s Technology Readiness Level from 6 to 7. The project resulted in many novel technologies, including: 1) an eye-safe red laser based, low-cost LiDAR system that can detect targets up to 10 meters distance; 2) the GAN-based machine learning underwater LiDAR image enhancement technique (this is the first known application of GAN technique in underwater LiDAR); 3) LiDAR-based real-time automated detection capabilities; and 4) a template matching based automated classification tool. These technologies build a solid foundation for future efforts to develop an extended range electro-optical monitoring system suitable for marine energy deployments. In addition to technology development, a key lesson learned is that addressing regulatory and safety requirements must be front and center in any marine energy monitoring applications.

16 TIDAL AND WAVE POWER↗

Allosteric prediction via convolutional neural networks and protein structural and dynamical features

Allostery is the phenomenon whereby a binding event or covalent modification at one site in a protein modulates function at a distal site, thus changing a protein’s functional state. As such, it is a ubiquitous aspect of protein functional regulation. Computationally predicting allosteric states is important as part of the broader challenge of functional annotation, but it also has practical implications for drug development, as targeting an allosteric site often affords greater specificity compared with targeting an orthosteric site. This study introduces a machine learning approach to predict the allosteric functional state using the small G-protein KRas as the model system, due to its implication in many types of cancer and being well studied as a result with many x-ray crystallographic structures of KRas available with different mutations and ligands bound. Using structural and dynamical features that can be cast as images, namely interatomic distances, contact maps, covariance, and mutual information, supervised learning was performed using convolutional neural networks. Two pretrained convolutional neural network architectures, GoogLeNet and ResNet18, were fine-tuned to classify KRas into active or inactive states based on these features. Across training regimes, atomic contact maps emerged as the most effective structural feature, whereas linearized mutual information outperformed covariance in capturing dynamical correlations relevant to allostery. Models achieved significant validation accuracy, with atomic contact maps yielding up to 90% accuracy. In conclusion, the findings suggest that integrating global structural rearrangements and correlated motion patterns with deep learning can reliably predict protein allosteric states, offering a promising framework for understanding allosteric regulation and developing targeted therapeutics.

Rajeshwar T., Rajitha [Oak Ridge National Laborato↗

Rocket Launch Detection with Smartphone Audio and Transfer Learning

Rocket launches generate infrasound signatures that have been detected at great distances. Due to the sparsity of the networks that have made these detections, however, most signals are detected tens of minutes to hours after the rocket launch. In this work, a method of near-real-time detection of rocket launches using data from a network of smartphones located 10–70 km from launch sites is presented. A machine learning model is trained and tested on the open-access Aggregated Smartphone Timeseries of Rocket-generated Acoustics (ASTRA), Smartphone High-explosive Audio Recordings Dataset (SHAReD), and ESC-50 datasets, resulting in a final accuracy of 97% and a false positive rate of <1%. The performance and behavior of the model are summarized, and its suitability for persistent monitoring applications is discussed.

acoustics↗

Multi-head attention-based U-Nets for predicting protein domain boundaries using 1D sequence features and 2D distance maps

Abstract The information about the domain architecture of proteins is useful for studying protein structure and function. However, accurate prediction of protein domain boundaries (i.e., sequence regions separating two domains) from sequence remains a significant challenge. In this work, we develop a deep learning method based on multi-head U-Nets (called DistDom) to predict protein domain boundaries utilizing 1D sequence features and predicted 2D inter-residue distance map as input. The 1D features contain the evolutionary and physicochemical information of protein sequences, whereas the 2D distance map includes the structural information of proteins that was rarely used in domain boundary prediction before. The 1D and 2D features are processed by the 1D and 2D U-Nets respectively to generate hidden features. The hidden features are then used by the multi-head attention to predict the probability of each residue of a protein being in a domain boundary, leveraging both local and global information in the features. The residue-level domain boundary predictions can be used to classify proteins as single-domain or multi-domain proteins. It classifies the CASP14 single-domain and multi-domain targets at the accuracy of 75.9%, 13.28% more accurate than the state-of-the-art method. Tested on the CASP14 multi-domain protein targets with expert annotated domain boundaries, the average per-target F1 measure score of the domain boundary prediction by DistDom is 0.263, 29.56% higher than the state-of-the-art method.

59 BASIC BIOLOGICAL SCIENCES↗

In situ Detection of Plasma Induced Surface Interaction based on Deep Learning based Visual Diagnostics (Technical Report)

It is characteristic for many plasma devices to undergo plasma-material interaction leading to surface erosion. These processes, often not easily detectable, lead to changes in device performance and lifespan. State-of-the-art lifetime tests and wear experiments require over 1000s hours. A self-consistent model for accurately predicting the erosion's effects is not available. In situ detection of these processes is not a trivial task since the surface variations at the early stages have a micron scale. Such limitations not only restrict testing and prediction capabilities but also slow the development of new thrusters and limit mission duration. To address these challenges, an in-situ diagnostic for real-time erosion assessment has been developed, aiming to expedite lifetime testing and broaden experimental campaigns. Several works were dedicated to real-time and in situ monitoring of material erosion during plasma exposure using laser holography, microscopy, and with telemicroscopes. However, the applicability of these approaches is limited due to complexity, cost and less flexibility as they often require placing diagnostic equipment inside the vacuum chamber. In collaboration with Princeton Collaborative Research Facility (PCRF), Princeton Plasma Physics Laboratory (PPPL), a new diagnostic approach is developed, where geometry modifications to the ceramic channel walls were introduced that would result in accelerated channel erosion. We employed Long-distance microscope (LDM) imagery, combined with Deep-Learning based Shape from focus or depth from focus (DFF or SFF) approach, that provides an accessible and cost-effective solution. LDM employs focus variation techniques to continuously capture multiple images of the target object at distinct focal planes. DFF, an optical focus variation method, generates a 3D topographical surface depth map from a sequence of variably focused images. Combined with the developed diagnostic, this approach offers a controllable means to study erosion under accelerated conditions. In this work, we develop Neural Network-based DFF algorithm applicable for LDM data to quantitatively evaluate plasma induced surface modification from LDM data. Next, we develop Deep Learning-based super-resolution depth map image reconstruction technique to increase the resolution of depth maps obtained from DFF algorithm to improve the accuracy of erosion measurements. Thirdly, we develop several image processing techniques to remove noise and improve the quality of depth map image. Here we report the results of initial tests for this approach. An experimental setup designed and built in PPPL was employed that consists of a 3-cm gridded ion source that produces a neutralized argon beam with energies up to 600 eV. A hexagonal boron nitride (h-BN) ceramic target, designed based on computational predictions, was used. Tests were conducted to reconstruct the complex geometry of the target under the lighting conditions of the operated ion source.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The Twins Embedding of Type Ia Supernovae. II. Improving Cosmological Distance Estimates

We show how spectra of Type Ia supernovae (SNe Ia) at maximum light can be used to improve cosmological distance estimates. In a companion article, we used manifold learning to build a three-dimensional parameterization of the intrinsic diversity of SNe Ia at maximum light that we call the "Twins Embedding."In this article, we discuss how the Twins Embedding can be used to improve the standardization of SNe Ia. With a single spectrophotometrically calibrated spectrum near maximum light, we can standardize our sample of SNe Ia with an rms of 0.101 0.007 mag, which corresponds to 0.084 0.009 mag if peculiar velocity contributions are removed and to 0.073 0.008 mag if a larger reference sample were obtained. Our techniques can standardize the full range of SNe Ia, including those typically labeled as peculiar and often rejected from other analyses. We find that traditional light-curve width + color standardization such as SALT2 is not sufficient. The Twins Embedding identifies a subset of SNe Ia, including, but not limited to, 91T-like SNe Ia whose SALT2 distance estimates are biased by 0.229 0.045 mag. Standardization using the Twins Embedding also significantly decreases host-galaxy correlations. We recover a host mass step of 0.040 0.020 mag compared to 0.092 0.026 mag for SALT2 standardization on the same sample of SNe Ia. These biases in traditional standardization methods could significantly impact future cosmology analyses if not properly taken into account.

79 ASTRONOMY AND ASTROPHYSICS↗

Cohort organized learning: clustering through agreement

In this article we describe cohort organized learning (CoOL), a method for clustering data without explicit distance or similarity computations. Herein, we will describe CoOL, derive the gradients determined by expectation maximization to train the networks, show how to monitor convergence during training and evaluate the clusters after training, and discuss a series of examples and use cases. We also discuss CoOL’s limitations and future prospects on related tasks. Because CoOL uses neural networks to estimate the clusters, it can be used to cluster any data that can be made compatible and we illustrate this on vector data and images.

clustering↗

Quantum-assisted associative adversarial network: applying quantum annealing in deep learning

Abstract Generative models have the capacity to model and generate new examples from a dataset and have an increasingly diverse set of applications driven by commercial and academic interest. In this work, we present an algorithm for learning a latent variable generative model via generative adversarial learning where the canonical uniform noise input is replaced by samples from a graphical model. This graphical model is learned by a Boltzmann machine which learns low-dimensional feature representation of data extracted by the discriminator. A quantum processor can be used to sample from the model to train the Boltzmann machine. This novel hybrid quantum-classical algorithm joins a growing family of algorithms that use a quantum processor sampling subroutine in deep learning, and provides a scalable framework to test the advantages of quantum-assisted learning. For the latent space model, fully connected, symmetric bipartite and Chimera graph topologies are compared on a reduced stochastically binarized MNIST dataset, for both classical and quantum sampling methods. The quantum-assisted associative adversarial network successfully learns a generative model of the MNIST dataset for all topologies. Evaluated using the Fréchet inception distance and inception score, the quantum and classical versions of the algorithm are found to have equivalent performance for learning an implicit generative model of the MNIST dataset. Classical sampling is used to demonstrate the algorithm on the LSUN bedrooms dataset, indicating scalability to larger and color datasets. Though the quantum processor used here is a quantum annealer, the algorithm is general enough such that any quantum processor, such as gate model quantum computers, may be substituted as a sampler.

Wilson, Max (ORCID:0000000207983391)↗

Vehicular Re-Identification from Uncontrolled Multiple Views

Vehicle re-identification (re-ID) across disparate sensing modalities remains a fundamental challenge for transportation research. In this work, we introduce a deep multi-view vehicle re-ID framework that leverages Siamese networks to compare pairs of vehicle images and produce matching scores, enabling robust association across drastically different viewpoints such as those from UAVs, surveillance cameras, and ground sensors. The model exploits convolutional neural networks to learn features that remain discriminative under changes in angle, distance, and illumination, supporting more generalizable re-ID performance. As part of this effort, we also developed an automated pipeline to synchronize roadside and UAV video streams, producing a multi-perspective dataset that complements preexisting real collections and a synthetic dataset generated in this study. Together, these contributions advance the capability to re-identify vehicles across wide viewing baselines; establish a foundation for scalable, reproducible research in vehicle re-ID; and open pathways for future applications, such as inferring routine behaviors, movement patterns, and daily habits of the individual associated with the vehicle.

convolutional neural networks↗

Learning Quantum States and Unitaries of Bounded Gate Complexity

While quantum state tomography is notoriously hard, most states hold little interest to practically minded tomographers. Given that states and unitaries appearing in nature are of bounded gate complexity, it is natural to ask if efficient learning becomes possible. In this work, we prove that to learn a state generated by a quantum circuit with G two-qubit gates to a small trace distance, a sample complexity scaling linearly in G is necessary and sufficient. We also prove that the optimal query complexity to learn a unitary generated by G gates to a small average-case error scales linearly in G . While sample-efficient learning can be achieved, we show that under reasonable cryptographic conjectures, the computational complexity for learning states and unitaries of gate complexity G must scale exponentially in G . We illustrate how these results establish fundamental limitations on the expressivity of quantum machine-learning models and provide new perspectives on no-free-lunch theorems in unitary learning. Together, our results answer how the complexity of learning quantum states and unitaries relate to the complexity of creating these states and unitaries. Published by the American Physical Society 2024

Zhao, Haimeng (ORCID:0000000166751489)↗

Spatial arrangement of dynamic surface species from solid-state NMR and machine learning-accelerated MD simulations

Here, the surface arrangement of motional organic functionalities is explored by experimental dipolar coupling measurements and the prediction of motionally-averaged coupling constant from molecular dynamics simulations. The use of machine learning potentials was key to reaching the timescale required. The distance between dynamic surface species are important in cooperative heterogeneous catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Residuals-based distributionally robust optimization with covariate information

We consider data-driven approaches that integrate a machine learning prediction model within distributionally robust optimization (DRO) given limited joint observations of uncertain parameters and covariates. Our framework is flexible in the sense that it can accommodate a variety of regression setups and DRO ambiguity sets. We investigate asymptotic and finite sample properties of solutions obtained using Wasserstein, sample robust optimization, and phi-divergence-based ambiguity sets within our DRO formulations, and explore cross-validation approaches for sizing these ambiguity sets. Through numerical experiments, we validate our theoretical results, study the effectiveness of our approaches for sizing ambiguity sets, and illustrate the benefits of our DRO formulations in the limited data regime even when the prediction model is misspecified.

97 MATHEMATICS AND COMPUTING↗