Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distance learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Parallax-corrected VISST-derived pixel-level products from satellite GOES-16

The NASA Langley group led by William Smith produced GOES-16 satellite cloud retrievals over an approximate 10 by 10 degree region over the CACTI field campaign location. These retrievals are described here: https://www.arm.gov/capabilities/vaps/visst and are available for download here . They use algorithms historically called VISST that are now referred to as SatCORPS. More information can be found in Trepte et al. (2019), Minnis et al. (2021), and Yost et al. (2021). If using this dataset, please cite these references, the CACTI VISST dataset DOI found at the download link above, and this dataset’s DOI. The CACTI VISST pixel-level retrievals are on a 2 km spatial grid and available every 15 minutes (every 10 minutes late in the campaign), producing 21,765 files for the entire field campaign between October 2018 and April 2019. They are not corrected for parallax error, which is an offset in the actual geographical location of a cloud above the surface due to the satellite viewing the cloud partly from the side off nadir. This dataset applies a correction for parallax using the location relative to the satellite and the retrieved cloud top height above the surface, which allows the dataset to be geo-located with surface-based observations. The parallax correction for each location depends on the longitude, latitude and cloud top height above ground level (AGL) for that longitude and latitude in the original VISST files. The cloud top height AGL requires first computing the surface elevation at each VISST grid point. Data from the Advanced Spaceborne Thermal Emission and Reflection (ASTER) Global Digital Elevation Map Version 3 at 30-m resolution is projected onto the VISST grid using conservative coarsening (conserving surface elevation) in the xESMF Python package. The surface elevation is then subtracted from the VISST-retrieved cloud top height above mean sea level. These cloud top heights AGL are then combined with longitude and latitude to estimate the latitude and longitude corrections. Due to variability in cloud top height, the parallax shifts produce an irregular grid of values since higher cloud tops are shifted further than lower cloud tops. A ball tree-based neighbor search with Haversine distance is performed using the Python-based scikit-learn library to find the nearest VISST grid point to each parallax correction-shifted point. The data value of the shifted point is then assigned to that VISST grid point. In this manner, the irregular geographical shifts to correct for parallax are projected back to the rectilinear VISST grid. Because relatively higher clouds should obscure lower clouds, the variable values for the highest cloud top are preferentially chosen if two or more values are assigned to a grid point. The parallax correction should be viewed as an improved but still imperfect estimation of the cloud top locations, largely because the cloud top height is an imperfect retrieval. Please see the attached README document for further information. Users are encouraged to contact the authors with any additional questions.

54 ENVIRONMENTAL SCIENCES↗

Comparing storm resolving models and climates via unsupervised machine learning

Global storm-resolving models (GSRMs) have gained widespread interest because of the unprecedented detail with which they resolve the global climate. However, it remains difficult to quantify objective differences in how GSRMs resolve complex atmospheric formations. This lack of comprehensive tools for comparing model similarities is a problem in many disparate fields that involve simulation tools for complex data. To address this challenge we develop methods to estimate distributional distances based on both nonlinear dimensionality reduction and vector quantization. Our approach automatically learns physically meaningful notions of similarity from low-dimensional latent data representations that the different models produce. This enables an intercomparison of nine GSRMs based on their high-dimensional simulation data (2D vertical velocity snapshots) and reveals that only six are similar in their representation of atmospheric dynamics. Furthermore, we uncover signatures of the convective response to global warming in a fully unsupervised way. Our study provides a path toward evaluating future high-resolution simulation data more objectively.

54 ENVIRONMENTAL SCIENCES↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Using Deep Learning to Develop a High Resolution Planetary Boundary Layer Model for Infrasound Propagation

Infrasound, with frequencies less than 20 Hz, is generated by both natural and anthropogenic sources. When one of these sources exerts a force on the atmosphere, infrasonic waves are generated. The propagation of these waves largely depends on temperature, wind speed, and wind direction. Previous work has used deep learning to accurately predict atmospheric specifications to altitudes of ~40 km. However, this model breaks down for local distances because it is too low resolution. Here we use a high-resolution meteorological dataset collected in Las Vegas, Nevada, USA to develop a deep learning model that can predict temperature, wind speed, and wind direction. Predictions are compared to ground truth observations to show that the model performs well at predicting temperature and wind direction but struggles with prediction wind speed. Model limitations and improvements are also discussed.

54 ENVIRONMENTAL SCIENCES↗

DeepAdversaries: examining the robustness of deep learning models for galaxy morphology classification

With increased adoption of supervised deep learning methods for work with cosmological survey data, the assessment of data perturbation effects (that can naturally occur in the data processing and analysis pipelines) and the development of methods that increase model robustness are increasingly important. In the context of morphological classification of galaxies, we study the effects of perturbations in imaging data. In particular, we examine the consequences of using neural networks when training on baseline data and testing on perturbed data. We consider perturbations associated with two primary sources: (a) increased observational noise as represented by higher levels of Poisson noise and (b) data processing noise incurred by steps such as image compression or telescope errors as represented by one-pixel adversarial attacks. We also test the efficacy of domain adaptation techniques in mitigating the perturbation-driven errors. We use classification accuracy, latent space visualizations, and latent space distance to assess model robustness in the face of these perturbations. For deep learning models without domain adaptation, we find that processing pixel-level errors easily flip the classification into an incorrect class and that higher observational noise makes the model trained on low-noise data unable to classify galaxy morphologies. On the other hand, we show that training with domain adaptation improves model robustness and mitigates the effects of these perturbations, improving the classification accuracy up to 23% on data with higher observational noise. Domain adaptation also increases up to a factor of ${\approx}2.3$ the latent space distance between the baseline and the incorrectly classified one-pixel perturbed image, making the model more robust to inadvertent perturbations. Successful development and implementation of methods that increase model robustness in astronomical survey pipelines will help pave the way for many more uses of deep learning for astronomy.

79 ASTRONOMY AND ASTROPHYSICS↗

Improving Spectral Resolution from Real-time Evolution for Correlated Systems

Abstract The quality of numerically simulated spectra using real-time evolution methods for strongly correlated systems is affected by both the length of simulation time and the system size, limiting resolution in both frequency and momentum. In this work, we propose a computationally cheap, linear autoregressive machine learning-based framework to extend short-time and short-distance results over a wider range. We use the proposed method to extend the lesser Green’s function for both the Hubbard model and the much more computationally challenging Hubbard-extended Holstein model. This technique significantly improves both the frequency and momentum resolution of the single-particle removal spectrum $${\mathcal{A}}(k,\omega )$$ A ( k , ω ) , allowing the observation of otherwise obscured spectral features due to electron-phonon coupling.

Tang, Ta↗

X-ray and molecular dynamics study of the temperature-dependent structure of molten NaF-Zr⁢F 4

The local atomic structure of NaF-Zr⁢F 4 (53–47 mol%) molten system and its evolution with temperature are examined with x-ray scattering measurements which are then used to validate the quality of ab initio and neural network-based molecular dynamics (NNMD) calculations in the temperature range 515–700°⁢C. The machine-learning enhanced NNMD calculations offer improved efficiency while maintaining accuracy at higher distances compared to ab initio calculations. Looking at the evolution of the pair distribution function with increasing temperature, a fundamental change in the liquid structure within the selected temperature range, accompanied by a slight decrease in overall correlation is revealed. NNMD calculations indicate the coexistence of three different fluorozirconate complexes: [Zr⁢F 6 ] 2– , [Zr⁢F 7 ] 3– , and [Zr⁢F 8 ] 4– , with a shift in the dominant coordination state from the 7-coordinated Zr cation toward a 6-coordinated cation with increasing temperature. The study also highlights the metastability of different local coordination structures, with frequent interconversions between the 6- and 7-coordinate states. Analysis of the Zr-F-Zr angular distribution function reveals the presence of both “edge-sharing” and “corner-sharing” fluorozirconate complexes with specific bond angles and distances in accord with previous studies, while the next-nearest-neighbor cation-cation correlations demonstrate a clear preference for unlike cations as nearest-neighbor pairs, emphasizing nonrandom arrangement. Finally, these findings contribute to a comprehensive understanding of the complex local structure of the molten salt, providing insights into temperature-dependent preferences and correlations within the molten system.

36 MATERIALS SCIENCE↗

Improving the accuracy of freight mode choice models: A case study using the 2017 CFS PUF data set and ensemble learning techniques

Here, the US Census Bureau has collected two rounds of experimental data from the Commodity Flow Survey, providing shipment-level characteristics of nationwide commodity movements, published in 2012 (i.e., Public Use Microdata) and in 2017 (i.e., Public Use File). With this information, data-driven methods have become increasingly valuable for understanding detailed patterns in freight logistics. In this study, we used the 2017 Commodity Flow Survey Public Use File data set to explore building a high-performance freight mode choice model, considering three main improvements: (1) constructing local models for each separate commodity/industry category; (2) extracting useful geographical features, particularly the derived distance of each freight mode between origin/destination zones; and (3) applying additional ensemble learning methods such as stacking or voting to combine results from local and unified models for improved performance. The proposed method achieved over 92% accuracy without incorporating external information, an over 19% increase compared to directly fitting Random Forests models over 10,000 samples. Furthermore, SHAP (Shapely Additive Explanations) values were computed to explain the outputs and major patterns obtained from the proposed model. The model framework could enhance the performance and interpretability of existing freight mode choice models.

42 ENGINEERING↗

A deep dilated convolutional residual network for predicting interchain contacts of protein homodimers

Abstract Motivation Deep learning has revolutionized protein tertiary structure prediction recently. The cutting-edge deep learning methods such as AlphaFold can predict high-accuracy tertiary structures for most individual protein chains. However, the accuracy of predicting quaternary structures of protein complexes consisting of multiple chains is still relatively low due to lack of advanced deep learning methods in the field. Because interchain residue–residue contacts can be used as distance restraints to guide quaternary structure modeling, here we develop a deep dilated convolutional residual network method (DRCon) to predict interchain residue–residue contacts in homodimers from residue–residue co-evolutionary signals derived from multiple sequence alignments of monomers, intrachain residue–residue contacts of monomers extracted from true/predicted tertiary structures or predicted by deep learning, and other sequence and structural features. Results Tested on three homodimer test datasets (Homo_std dataset, DeepHomo dataset and CASP-CAPRI dataset), the precision of DRCon for top L/5 interchain contact predictions (L: length of monomer in a homodimer) is 43.46%, 47.10% and 33.50% respectively at 6 Å contact threshold, which is substantially better than DeepHomo and DNCON2_inter and similar to Glinter. Moreover, our experiments demonstrate that using predicted tertiary structure or intrachain contacts of monomers in the unbound state as input, DRCon still performs well, even though its accuracy is lower than using true tertiary structures in the bound state are used as input. Finally, our case study shows that good interchain contact predictions can be used to build high-accuracy quaternary structure models of homodimers. Availability and implementation The source code of DRCon is available at https://github.com/jianlin-cheng/DRCon. The datasets are available at https://zenodo.org/record/5998532#.YgF70vXMKsB. Supplementary information Supplementary data are available at Bioinformatics online.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Time Series Classification for Locating Forced Oscillation Sources

Here, this article presents a machine learning based time-series classification method for using synchrophasor measurements to locate the source of forced oscillation (FO) for fast disturbance removal. First, multivariate time series (MTS) matrices are constructed by the most informative measurements selected by sequential feature selection from each power plant. Then, the Mahalanobis matrix is trained such that the Mahalanobis distance between the MTSs from the same class (i.e., with the same FO source location) are minimized and from different classes (i.e., with different FO source locations) are maximized. This allows MTSs to be classified by classifiers with class membership corresponding to the location of each FO source. To meet the runtime requirements of online matching, class templates are constructed to reduce data size and improve matching efficiency. To account for uncertainty in identifying the exact beginning of an FO event, dynamic time warping is used to align the out-of-sync MTSs. IEEE 39bus and WECC 179bus systems are used for algorithm development and validation. Simulation results demonstrate that the algorithm meets online operation runtime requirement with high accuracy using misaligned data sets.

42 ENGINEERING↗

Comparing Mapper Graphs of Artificial Neuron Activations

The mapper graph is a popular tool from topological data analysis that provides a graphical summary of point cloud data. It has been used to study data from cancer research, sports analytics, neurosciences, and machine learning. In particular, mapper graphs have been used recently to visualize the topology of high-dimensional artificial neural activations from convolutional neural networks and large language models. However, a key question that arises from using mapper graphs across applications is how to compare mapper graphs to study their structural differences. In this paper, we introduce a distance between mapper graphs using tools from optimal transport. We demonstrate the utility of such a distance by studying the topological changes of neural activations across convolutional layers in deep learning, as well as by capturing the loss of structural information for multiscale mapper.

mapper graphs, computational topology, machine lea↗

A Geometry-Driven Longitudal Topic Model

A simple and scalable framework for longitudinal analysis of Twitter data is developed that combines latent topic models with computational geometric methods. Dimensionality reduction tools from computational geometry are applied to learn the intrinsic manifold on which the latent, temporal topics reside. Then shortest path distances on the manifold are used to link together these topics. The proposed framework permits visualization of the low-dimensional embedding which provides clear interpretation of the complex, high-dimensional trajectories that may exist among latent topics. Practical application of the proposed framework is demonstrated through its ability to capture and effectively visualize natural progression of latent COVID-19 related topics learned from Twitter data. Interpretability of the trajectories is achieved by comparing to real-world events. In addition, the framework permits study of spatial variation in Twitter behavior for learned topics. The analysis demonstrates that the proposed framework is able to capture granular-level impact of COVID-19 on public discussions. We end by arguing that Twitter data, when analyzed within the proposed framework, can serve as a valuable supplementary data stream for COVID-related studies.

97 MATHEMATICS AND COMPUTING↗

Metric DBSCAN

SAND2025-11725O Metric DBSCAN is an implementation of the popular DBSCAN clustering algorithm that works in general metric spaces. DBSCAN is a clustering algorithm, a fundamental building block in machine learning. It takes a set of objects and, given some notion of distance, identifies coherent groups of objects. With Metric DBSCAN, users can provide an arbitrary function to compute distance. Nearly all existing implementations of DBSCAN restrict distance to one of a few formulations. Metric DBScan accomplishes this cleanly and efficiently. The Python source code is on Github. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

PhotonIDs: ML-Powered Photon Identification System for Dark Count Elimination

Reliable single photon detection is the foundation for practical quantum communication and networking. However, today's superconducting nanowire single photon detector(SNSPD) inherently fails to distinguish between genuine photon events and dark counts, leading to degraded fidelity in long-distance quantum communication. In this work, we introduce PhotonIDs, a machine learning-powered photon identification system that is the first end-to-end solution for real-time discrimination between photons and dark count based on full SNSPD readout signal waveform analysis. PhotonIDs ~demonstrates: 1) an FPGA-based high-speed data acquisition platform that selectively captures the full waveform of signal only while filtering out the background data in real time; 2) an efficient signal preprocessing pipeline, and a novel pseudo-position metric that is derived from the physical temporal-spatial features of each detected event; 3) a hybrid machine learning model with near 98% accuracy achieved on photon/dark count classification. Additionally, proposed PhotonIDs ~ is evaluated on the dark count elimination performance with two real-world case studies: (1) 20 km quantum link, and (2) Erbium ion-based photon emission system. Our result demonstrates that PhotonIDs ~could improve more than 31.2 times of signal-noise-ratio~(SNR) on dark count elimination. PhotonIDs ~ marks a step forward in noise-resilient quantum communication infrastructure.

Linne, Karl C. [Chicago U.] (ORCID:000900091870358↗

REC protein family expansion by the emergence of a new signaling pathway

This report presents multi-genome evidence that REC protein family expansion occurs when the emergence of new pathways gives rise to functional discordance. Specificity between residues in REC domain containing response regulators with paired histidine kinases is under negative purifying selection, constrained by the presence of other bacterial two-component systems signaling cascades that share sequence and structural identity. Presuming that the two-component systems can evolve by neutral amino acid changes (neutral drift) when purifying evolutionary constraints are relaxed, how might the REC protein family expand by amino acid changes when these constraints remain intact? Using an unsupervised machine learning approach to observe the sequence landscape of REC domains across long phylogenetic distances, we find that within-gene recombination, a subcategory of gene conversion, switched the effector domain and, consequently, the regulatory context of a duplicated response regulator from transcriptional regulation by σ54 to that by σ70. We determined that the recombined response regulator diverged from its parent by episodic diversifying selection and neutral drift. Functional experiments of the parent of recombined response regulators in a model Pseudomonas putida KT2440 model system revealed that the parent and recombined response regulators sense and respond to different carboxylic acids. Finally, a residue-switching experiment using structural predictions and functional characterization suggests that the new residues in the recombined regulator could form a new interaction interface and mediate condition-specific phosphotransfer. Overall, our study finds that genetic perturbations can create conditions of functional discordance, whereby the REC protein family can evolve by episodic diversifying selection.

59 BASIC BIOLOGICAL SCIENCES↗

Deep probabilistic direction prediction in 3D with applications to directional dark matter detectors

Abstract We present the first method to probabilistically predict 3D direction in a deep neural network model. The probabilistic predictions are modeled as a heteroscedastic von Mises-Fisher distribution on the sphere S 2 , giving a simple way to quantify aleatoric uncertainty. This approach generalizes the cosine distance loss which is a special case of our loss function when the uncertainty is assumed to be uniform across samples. We develop approximations required to make the likelihood function and gradient calculations stable. The method is applied to the task of predicting the 3D directions of electrons, the most complex signal in a class of experimental particle physics detectors designed to demonstrate the particle nature of dark matter and study solar neutrinos. Using simulated Monte Carlo data, the initial direction of recoiling electrons is inferred from their tortuous trajectories, as captured by the 3D detectors. For 40 keV electrons in a 70% He 30% CO 2 gas mixture at STP, the new approach achieves a mean cosine distance of 0.104 (26 ∘ ) compared to 0.556 (64 ∘ ) achieved by a non-machine learning algorithm. We show that the model is well-calibrated and accuracy can be increased further by removing samples with high predicted uncertainty. This advancement in probabilistic 3D directional learning could increase the sensitivity of directional dark matter detectors.

Computer Science↗

Evaluating Physics-Informed Neural Network Performance for Seismic Discrimination between Earthquakes and Explosions

In this article, we evaluate adding a weak physics constraint, that is, a physics‐based empirical relationship, to the loss function with a physics‐informed manner in local distance explosion discrimination in the hope of improving the generalization capability of the machine learning (ML) model. We compare the proposed model with the two‐branch model we previously developed, as well as with a pure data‐driven model. Unexpectedly, the proposed model did not consistently outperform the pure data‐driven model. By varying the level of inconsistency in the training data, we find this approach is modulated by the strength of the physics relationship. In conclusion, this result has important implications for how to best incorporate physical constraints in ML models.

58 GEOSCIENCES↗

A deep learning approach to real-time HIV outbreak detection using genetic data

Pathogen genomic sequence data are increasingly made available for epidemiological monitoring. A main interest is to identify and assess the potential of infectious disease outbreaks. While popular methods to analyze sequence data often involve phylogenetic tree inference, they are vulnerable to errors from recombination and impose a high computational cost, making it difficult to obtain real-time results when the number of sequences is in or above the thousands. Here, we propose an alternative strategy to outbreak detection using genomic data based on deep learning methods developed for image classification. The key idea is to use a pairwise genetic distance matrix calculated from viral sequences as an image, and develop convolutional neutral network (CNN) models to classify areas of the images that show signatures of active outbreak, leading to identification of subsets of sequences taken from an active outbreak. We showed that our method is efficient in finding HIV-1 outbreaks with R0 ≥ 2.5, and overall a specificity exceeding 98% and sensitivity better than 92%. We validated our approach using data from HIV-1 CRF01 in Europe, containing both endemic sequences and a well-known dual outbreak in intravenous drug users. Our model accurately identified known outbreak sequences in the background of slower spreading HIV. Importantly, we detected both outbreaks early on, before they were over, implying that had this method been applied in real-time as data became available, one would have been able to intervene and possibly prevent the extent of these outbreaks. This approach is scalable to processing hundreds of thousands of sequences, making it useful for current and future real-time epidemiological investigations, including public health monitoring using large databases and especially for rapid outbreak identification.

59 BASIC BIOLOGICAL SCIENCES↗