Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distance learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Lessons from the COVID Era and Visions for the Future

In December 2020, the U.S. Department of Energy Office of Science convened a virtual Roundtable of its 27 operating scientific user facilities to discuss facility challenges and lessons learned during the COVID-19 pandemic as well as facility responses, best practices, and innovations that could be adopted going forward. Roundtable participants included facility staff, users, and user executive committee chairs. This report summarizes their discussions, which encompassed topics such as user research and facility operations in virtual and physically distanced contexts; user training and engagement; computation, data, and network resources; and crosscutting issues.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Field Testing of a Mixed Potential IoT Sensor Platform for Methane Quantification

Emissions of CH 4 from natural gas infrastructure must urgently be addressed to mitigate its effect on global climate. With hundreds of thousands of miles of pipeline in the US used to transport natural gas, current methods of surveying for leaks are inadequate. Mixed potential sensors are a low cost, field deployable technology for remote and continuous monitoring of natural gas infrastructure. We demonstrate for the first time a field trial of a mixed potential sensor device coupled with machine learning and internet-of-things platform at Colorado State University’s Methane Emissions Technology Evaluation Center (METEC). Emissions were detected from a simulated buried underground pipeline source. Sensor data was acquired and transmitted from the field test site to a remote cloud server. Quantification of concentration as a function of vertical distance is consistent with previously reported transport modelling efforts and experimental surveys of methane emissions by more sophisticated CH 4 analyzers.

03 NATURAL GAS↗

Neural Scaling Laws for Jet Generation

Recently observed empirical scaling laws describe the performance of foundation-type models as three independent key quantities -- dataset size, compute, and model parameters -- are modified. Extracting these scaling laws informs the training of large complex models for which the tuning of hyperparameters in traditional ways is not feasible. This work for the first time explores if scaling laws can also be observed for the task of particle jet generation -- both relevant as a pre-training objective for foundation models and as in-situ simulation by itself. We indeed replicate the key logarithmic scaling law behavior for model-size scaling. Beyond studying the next token prediction validation loss of the generative model, we also study the sliced Wasserstein distance of five physical quantities that are not immediately available to the model during training. Our study shows that this quantity is monotonically related to the next token prediction validation loss, meaning that this loss is indeed a good proxy for the physics performance. For the scaling with dataset size and compute, we observe substantially weaker scaling behavior of both the loss and the sliced Wasserstein distance. We analyze this behavior by introducing the concept of a learnable window, and argue that autoregressive next token prediction on jet constituents exhibits comparatively rapid saturation relative to language-model studies. We discuss possible origins of this behavior, including the stochastic nature of QCD radiation and differences between generative and supervised learning tasks in collider physics.

Amram, Oz [Fermilab]↗

Reinforcement learning-based design of shape-changing metamaterials

During the last decade, artificially architected materials have been designed to obtain properties unreachable by naturally occurring materials, whose properties are determined by their atomic structure and chemical composition. In this work, we implement a new reinforcement learning (RL) method able to rationally design unique metamaterial structures at the nano-, micro-, and macroscale, which change shape during operational conditions. As an example, we apply this method to design nanostructured silicon anodes for Li-ion batteries (LIBs). The RL model is designed to apply different actions and predict change during operational conditions. The multi-component reward function comprises an increase in the total storage capacity of the resulting battery electrode and structural parameters, such as the minimum distance between the individual components of the nanostructure. Upon experimental validation using a polymer-based 3D printing technique, we expect that the newly discovered structures improve the current Si-based LIB anodes state-of-the-art by almost three times and almost ten times the current commercial LIB based on a graphitic anode. Furthermore, this RL-based optimization method opens up vast design space for other responsive metamaterials with tailored properties and pre-programmed structural transformation.

25 ENERGY STORAGE↗

Confinement Effects on Proton Transfer in TiO 2 Nanopores from Machine Learning Potential Molecular Dynamics Simulations

Improved understanding of proton transfer in nanopores is critical for a wide range of emerging applications, yet experimentally probing mechanisms and energetics of this process remains a significant challenge. To help reveal details of this process, we developed and applied a machine learning potential derived from first-principles calculations to examine water reactivity and proton transfer in TiO 2 slit-pores. Here, we find that confinement of water within pores smaller than 0.5 nm imposes strong and complex effects on water reactivity and proton transfer. Although the proton transfer mechanism is similar to that at a TiO 2 interface with bulk water, confinement reduces the activation energy of this process, leading to more frequent proton transfer events. This enhanced proton transfer stems from the contraction of oxygen–oxygen distances dictated by the interplay between confinement and hydrophilic interactions. Our simulations also highlight the importance of the surface topology, where faster proton transport is found in the direction where a unique arrangement of surface oxygens enables the formation of an ordered water chain. In a broader context, our study demonstrates that proton transfer in hydrophilic nanopores can be enhanced by controlling pore size, surface chemistry, and topology.

36 MATERIALS SCIENCE↗

Composite Qdrift-product formulas for quantum and classical simulations in real and imaginary time

Recent study has shown that it can be advantageous to implement a composite channel that partitions the Hamiltonian H for a given simulation problem into subsets A and B such that H = A + B , where the terms in A are simulated with a Trotter-Suzuki channel and the B terms are randomly sampled via the Qdrift algorithm. Here we extend Qdrift and composite product formulas to imaginary time, formulating candidate classical algorithms for quantum Monte Carlo calculations. We upper bound the induced Schatten- 1 → 1 norm on both imaginary-time Qdrift and composite channels. Another recent result demonstrated that simulations of lattice Hamiltonians containing geometrically local interactions can be improved using a Lieb-Robinson argument to decompose H into subsets that contain only terms supported on that subset of the lattice. Here, we provide a quantum algorithm by unifying this result with the composite approach into “local composite channels” and we upper bound the diamond distance. We provide exact numerical simulations of algorithmic cost by counting the number of gates of the form e − i H j t and e − H j β to meet a certain error tolerance ε . In doing so, we optimize the partitioning into sets A and B using gradient boosted tree models from machine learning. These numerical studies are important given that product formulas have been historically known to outperform analytic upper bounds. We show constant factor advantages for a variety of interesting Hamiltonians, the maximum of which is a ≈ 20 -fold speedup that occurs in the simulation of Jellium. Published by the American Physical Society 2024

Pocrnic, Matthew (ORCID:0000000203089376)↗

Converting tabular data into images for deep learning with convolutional neural networks

Abstract Convolutional neural networks (CNNs) have been successfully used in many applications where important information about data is embedded in the order of features, such as speech and imaging. However, most tabular data do not assume a spatial relationship between features, and thus are unsuitable for modeling using CNNs. To meet this challenge, we develop a novel algorithm, image generator for tabular data (IGTD), to transform tabular data into images by assigning features to pixel positions so that similar features are close to each other in the image. The algorithm searches for an optimized assignment by minimizing the difference between the ranking of distances between features and the ranking of distances between their assigned pixels in the image. We apply IGTD to transform gene expression profiles of cancer cell lines (CCLs) and molecular descriptors of drugs into their respective image representations. Compared with existing transformation methods, IGTD generates compact image representations with better preservation of feature neighborhood structure. Evaluated on benchmark drug screening datasets, CNNs trained on IGTD image representations of CCLs and drugs exhibit a better performance of predicting anti-cancer drug response than both CNNs trained on alternative image representations and prediction models trained on the original tabular data.

59 BASIC BIOLOGICAL SCIENCES↗

Reducing ridesourcing empty vehicle travel with future travel demand prediction

Ridesourcing services provide alternative mobility options in several cities. Their market share has grown exponentially due to the convenience they provide. The use of such services may be associated with car-light or car-free lifestyles. However, there are growing concerns regarding their impact on urban transportation operations performance due to empty, unproductive miles driven without a passenger (commonly referred to as deadheading). This paper is motivated by the potential to reduce deadhead mileage of ridesourcing trips by providing drivers with information on future ridesourcing trip demand. Future demand information enables the driver to wait in place for the next rider’s request without cruising around and contributing to congestion. A machine learning model is employed to predict hourly and 10-minute future interval travel demand for ridesourcing at a given location. Using future demand information, we propose algorithms to (i) assign drivers to act on received demand information by waiting in place for the next rider, and (ii) match these drivers with riders to minimize deadheading distance. Real-world data from ridesourcing providers in Austin, TX (RideAustin) and Chengdu, China (DiDi Chuxing) are leveraged. Results show that this process achieves 68%–82% and 53%–60% reduction of trip-level deadheading miles for the RideAustin and DiDi Chuxing sample operations respectively, under the assumption of unconstrained availability of short-term parking. Deadheading savings increase slightly as the maximum tolerable waiting time for the driver increases. Further, it is observed that significant deadhead savings per trip are possible, even when a small percent of the ridesourcing driver pool is provided with future ridesourcing demand information.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Accurate prediction of global-density-dependent range-separation parameters based on machine learning

In this work, we develop an accurate and efficient XGBoost machine learning model for predicting the global-density-dependent range-separation parameter, ωGDD, for long-range corrected functional (LRC)-ωPBE. This ωGDDML model has been built using a wide range of systems (11 466 complexes, ten different elements, and up to 139 heavy atoms) with fingerprints for the local atomic environment and histograms of distances for the long-range atomic correlation for mapping the quantum mechanical range-separation values. The promising performance on the testing set with 7046 complexes shows a mean absolute error of 0.001 117 a0−1 and only five systems (0.07%) with an absolute error larger than 0.01 a0−1, which indicates the good transferability of our ωGDDML model. In addition, the only required input to obtain ωGDDML is the Cartesian coordinates without electronic structure calculations, thereby enabling rapid predictions. LRC-ωPBE(ωGDDML) is used to predict polarizabilities for a series of oligomers, where polarizabilities are sensitive to the asymptotic density decay and are crucial in a variety of applications, including the calculations of dispersion corrections and refractive index, and surpasses the performance of all other popular density functionals except for the non-tuned LRC-ωPBE. Finally, LRC-ωPBE (ωGDDML) combined with (extended) symmetry-adapted perturbation theory is used in calculating noncovalent interactions to further show that the traditional ab initio system-specific tuning procedure can be bypassed. The present study not only provides an accurate and efficient way to determine the range-separation parameter for LRC-ωPBE but also shows the synergistic benefits of fusing the power of physically inspired density functional LRC-ωPBE and the data-driven ωGDDML model.

Chemistry↗

Graphical Gaussian Process Regression Model for Aqueous Solvation Free Energy Prediction of Organic Molecules in Redox Flow Battery

The solvation free energy of organic molecules is a critical parameter in determining emergent properties such as solubility, liquid-phase equilibrium constants, and pKa and redox potentials in an organic redox flow battery. In this work, we present a machine learning (ML) model that can learn and predict the aqueous solvation free energy of an organic molecule using Gaussian process regression method based on a new molecular graph kernel. To investigate the performance of the ML model on electrostatic interaction, the nonpolar interaction contribution of solvent and the conformational entropy of solute in solvation free energy, three data sets with implicit or explicit water solvent models, and contribution of conformational entropy of solute are tested. We demonstrate that our ML model can predict the solvation free energy of molecules at chemical accuracy with a mean absolute error of less than 1 kcal/mol for subsets of the QM9 dataset and the Freesolv database. To solve the general data scarcity problem for a graph-based ML model, we propose a dimension reduction algorithm based on the distance between molecular graphs, which can be used to examine the diversity of the molecular data set. It provides a promising way to build a minimum training set to improve prediction for certain test sets where the space of molecular structures is predetermined.

25 ENERGY STORAGE↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Epidural anesthesia needle guidance by forward-view endoscopic optical coherence tomography and deep learning

Epidural anesthesia requires injection of anesthetic into the epidural space in the spine. Accurate placement of the epidural needle is a major challenge. To address this, we developed a forward-view endoscopic optical coherence tomography (OCT) system for real-time imaging of the tissue in front of the needle tip during the puncture. We tested this OCT system in porcine backbones and developed a set of deep learning models to automatically process the imaging data for needle localization. A series of binary classification models were developed to recognize the five layers of the backbone, including fat, interspinous ligament, ligamentum flavum, epidural space, and spinal cord. The classification models provided an average classification accuracy of 96.65%. During puncture, it is important to maintain a safe distance between the needle tip and the dura mater. Regression models were developed to estimate that distance based on the OCT imaging data. Based on the Inception architecture, our models achieved a mean absolute percentage error of 3.05% ± 0.55%. Overall, our results validated the technical feasibility of using this novel imaging strategy to automatically recognize different tissue structures and measure the distances ahead of the needle tip during the epidural needle placement.

60 APPLIED LIFE SCIENCES↗

Daily Forecasting of Regional Epidemics of Coronavirus Disease with Bayesian Uncertainty Quantification, United States

To increase situational awareness and support evidence-based policymaking, we formulated a mathematical model for coronavirus disease transmission within a regional population. This compartmental model accounts for quarantine, self-isolation, social distancing, a nonexponentially distributed incubation period, asymptomatic persons, and mild and severe forms of symptomatic disease. We used Bayesian inference to calibrate region-specific models for consistency with daily reports of confirmed cases in the 15 most populous metropolitan statistical areas in the United States. We also quantified uncertainty in parameter estimates and forecasts. This online learning approach enables early identification of new trends despite considerable variability in case reporting.

59 BASIC BIOLOGICAL SCIENCES↗

Interpretable Machine Learning for Characterizing Electric Vehicle Charging Behavior: Insights from Real-World Data

As electric vehicle (EV) adoption rises globally, concerns about the impact on aging electrical grids grow, particularly regarding the charging behavior of EV drivers. This study analyzes real-world driving and charging data from Ford battery electric vehicles (BEVs) collected between 2018 and 2019 to develop interpretable models that characterize charging behavior and quantify influencing factors. Prior research has relied on assumptions regarding driver behavior, often overlooking actual charging patterns. By employing generalized linear mixed models (GLMMs), this work offers insights into how various elements, such as next trip distance and state of charge (SOC), influence charging decisions. The dataset comprises over three million park-trip pairs from 1,997 vehicles, revealing that features related to driving behavior significantly dictate charging behavior, while infrastructure and regional factors have lesser impacts. The findings suggest that existing simulation models may oversimplify EV charging behavior assumptions. This work utilizes real-world EV driving and charging data to train interpretable models that describe charging behavior and quantify the factors most associated with how drivers use charging infrastructure. This research underscores the need for interpretable, data-driven methodologies to inform future EV infrastructure planning and grid management.

29 - ENERGY PLANNING, POLICY AND ECONOMY↗

The Evaluation of Machine Learning Techniques for Isotope Identification Contextualized by Training and Testing Spectral Similarity

Precise gamma-ray spectral analysis is crucial in high-stakes applications, such as nuclear security. Research efforts toward implementing machine learning (ML) approaches for accurate analysis are limited by the resemblance of the training data to the testing scenarios. The underlying spectral shape of synthetic data may not perfectly reflect measured configurations, and measurement campaigns may be limited by resource constraints. Consequently, ML algorithms for isotope identification must maintain accurate classification performance under domain shifts between the training and testing data. To this end, four different classifiers (Ridge, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron) were trained on the same dataset and evaluated on twelve other datasets with varying standoff distances, shielding, and background configurations. A tailored statistical approach was introduced to quantify the similarity between the training and testing configurations, which was then related to the predictive performance. Wilcoxon signed-rank tests revealed that the OVR-wrapped XGB significantly outperformed the other algorithms, with confidence levels of 99.0% or above for the 133Ba, 60Co, 137Cs, and 152Eu sources. The findings from this work are significant as they outline techniques to promote the development of robust ML-based approaches for isotope identification.

domain adaptation↗

Jacobian-scaled K-means clustering for physics-informed segmentation of reacting flows

This work introduces Jacobian-scaled K-means (JSK-means) clustering, which is a physicsinformed clustering strategy centered on the K-means framework. The method allows for the injection of underlying physical knowledge into the clustering procedure through a distance function modification: instead of leveraging conventional Euclidean distance vectors, the JSKmeans procedure operates on distance vectors scaled by matrices obtained from dynamical system Jacobians evaluated at the cluster centroids. The goal of this work is to show how the JSKmeans algorithm - without modifying the input dataset - produces clusters that capture regions of dynamical similarity, in that the clusters are redistributed towards high-sensitivity regions in phase space and are described by similarity in the source terms of samples instead of the samples themselves. The algorithm is demonstrated on a complex reacting flow simulation dataset (a channel detonation configuration), where the dynamics in the thermochemical composition space are known through the highly nonlinear and stiff Arrhenius-based chemical source terms. Interpretations of cluster partitions in both physical space and composition space reveal how JSK-means shifts clusters produced by standard K-means towards regions of high chemical sensitivity (e.g., towards regions of peak heat release rate near the detonation reaction zone). Furthermore, the findings presented here illustrate the benefits of utilizing Jacobian-scaled distances in clustering techniques, and the JSK-means method in particular displays promising potential for improving former partition-based modeling strategies in reacting flow (and other multi-physics) applications.

Clustering↗

Optimal Coordination of Electric Vehicles for Grid Services using Deep Reinforcement Learning

Recent research has shown the effectiveness of reinforcement learning (RL) in coordinating electric vehicles (EVs) with vehicle-to-grid capabilities for grid services. However, many of these studies rely on lookup table and deep Q-network techniques, which can be impractical when dealing with continuous states and actions. In addition, existing RL designs inadequately account for battery aging effects, EV user satisfaction, uncertain departure and arrival time, and trip distance, which may compromise effective coordination. This paper aims to bridge these gaps by developing an innovative deep deterministic policy gradient-based RL framework for optimal coordination of EVs. Case studies were carried out using a test system with 100 EVs, and numerical analysis results showed that the proposed RL framework can effectively coordinate EVs to maximize economic benefits and user satisfaction while ensuring the expected battery lifespan.

Das, Avijit↗