Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “MACHINE LEARNING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine Learning for Improving Surface-Layer-Flux Estimates

Abstract Flows in the atmospheric boundary layer are turbulent, characterized by a large Reynolds number, the existence of a roughness sublayer and the absence of a well-defined viscous layer. Exchanges with the surface are therefore dominated by turbulent fluxes. In numerical models for atmospheric flows, turbulent fluxes must be specified at the surface; however, surface fluxes are not known a priori and therefore must be parametrized. Atmospheric flow models, including global circulation, limited area models, and large-eddy simulation, employ Monin–Obukhov similarity theory (MOST) to parametrize surface fluxes. The MOST approach is a semi-empirical formulation that accounts for atmospheric stability effects through universal stability functions. The stability functions are determined based on limited observations using simple regression as a function of the non-dimensional stability parameter representing a ratio of distance from the surface and the Obukhov length scale (Obukhov in Trudy Inst Theor Geofiz AN SSSR 1:95–115, 1946), $$z/L$$ z / L . However, simple regression cannot capture the relationship between governing parameters and surface-layer structure under the wide range of conditions to which MOST is commonly applied. We therefore develop, train, and test two machine-learning models, an artificial neural network (ANN) and random forest (RF), to estimate surface fluxes of momentum, sensible heat, and moisture based on surface and near-surface observations. To train and test these machine-learning algorithms, we use several years of observations from the Cabauw mast in the Netherlands and from the National Oceanic and Atmospheric Administration’s Field Research Division tower in Idaho. The RF and ANN models outperform MOST. Even when we train the RF and ANN on one set of data and apply them to the second set, they provide more accurate estimates of all of the fluxes compared to MOST. Estimates of sensible heat and moisture fluxes are significantly improved, and model interpretability techniques highlight the logical physical relationships we expect in surface-layer processes.

Meteorology & Atmospheric Sciences↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

Data-driven wind turbine wake modeling via probabilistic machine learning

Wind farm design primarily depends on the variability of the wind turbine wake flows to the atmospheric wind conditions and the interaction between wakes. Physics-based models that capture the wake flow field with high-fidelity are computationally very expensive to perform layout optimization of wind farms, and, thus, data-driven reduced-order models can represent an efficient alternative for simulating wind farms. In this work, we use real-world light detection and ranging (LiDAR) measurements of wind-turbine wakes to construct predictive surrogate models using machine learning. Specifically, we first demonstrate the use of deep autoencoders to find a low-dimensional latent space that gives a computationally tractable approximation of the wake LiDAR measurements. Then, we learn the mapping between the parameter space and the (latent space) wake flow fields using a deep neural network. Additionally, we also demonstrate the use of a probabilistic machine learning technique, namely, Gaussian process modeling, to learn the parameter-space-latent-space mapping in addition to the epistemic and aleatoric uncertainty in the data. Finally, to cope with training large datasets, we demonstrate the use of variational Gaussian process models that provide a tractable alternative to the conventional Gaussian process models for large datasets. Furthermore, we introduce the use of active learning to adaptively build and improve a conventional Gaussian process model predictive capability. Overall, we find that our approach provides accurate approximations of the wind-turbine wake flow field that can be queried at an orders-of-magnitude cheaper cost than those generated with high-fidelity physics-based simulations.

Deep neural networks↗

Inferring the Isotropic-Nematic Phase Transition with Generative Machine Learning

Generative machine learning models are capable of learning the phase behavior in condensed matter systems such as the Ising model. We utilize a score-based modeling procedure called thermodynamic maps to describe the isotropic-nematic phase transition in a melt of Gay-Berne ellipsoids. When trained on samples from a single temperature on either side of the phase transition, this generative machine learning approach infers effectively the nematic order parameter at intermediate temperatures. Furthermore, these results demonstrate score-based models’ ability to learn the physics of a nontrivial liquid crystal phase transition.

Critical exponents↗

Machine learning for modern power distribution systems: Progress and perspectives

The application of machine learning (ML) to power and energy systems (PES) is being researched at an astounding rate, resulting in a significant number of recent additions to the literature. As the infrastructure of electric power systems evolves, so does interest in deploying ML techniques to PES. However, despite growing interest, the limited number of reported real-world applications suggests that the gap between research and practice is yet to be fully bridged. To help highlight areas where this gap could be narrowed, this article discusses the challenges and opportunities in developing and adapting ML techniques for modern electric power systems, with a particular focus on power distribution systems. These systems play a crucial role in transforming the electric power sector and accommodating emerging distributed technologies to mitigate the impacts of climate change and accelerate the transition to a sustainable energy future. The objective of this article is not to provide an exhaustive overview of the state-of-the-art in the literature, but rather to make the topic accessible to readers with an engineering or computer science background and an interest in the field of ML for PES, thereby encouraging cross-disciplinary research in this rapidly developing field. To this end, the article discusses the ways in which ML can contribute to addressing the evolving operational challenges facing power distribution systems and identifies relevant application areas that exemplify the potential for ML to make near-term contributions. At the same time, key considerations for the practical implementation of ML in power distribution systems are discussed, along with suggestions for several potential future directions.

Marković, Marija (ORCID:0000000247839837)↗

PhILMs: Collaboratory on Mathematics and Physics-Informed Learning Machines for Multiscale and Multiphysics Problems

The landscape of computational science and engineering is continually evolving, with the challenge of high-dimensional regression problems standing as a significant hurdle in numerous scientific endeavors. Addressing this challenge, our research, funded by this award, has led to the development of an innovative computational framework known as Probabilistic Partition of Unity Networks (PPOU-Nets). This initiative represents a collaborative effort to harness the potential of mathematics and physics-informed machine learning in tackling multiscale and multiphysics problems prevalent in high-dimensional spaces. Through this work, we have proposed a novel methodology that seamlessly integrates adaptive dimensionality reduction and a mixture of experts model, thereby facilitating a more efficient and accurate approximation of complex functions. This research effort has not only advanced the state of computational science but also opened new avenues for exploration in quantum computing and beyond. This report outlines the motivation, methodology, key findings, and implications of our work, underscoring our contributions to the broader scientific community and the potential pathways for future research.

97 MATHEMATICS AND COMPUTING↗

Three-dimensional realizations of flood flow in large-scale rivers using the neural fuzzy-based machine-learning algorithms

Machine learning methods have been extensively used to study the dynamics of complex fluid flows. One such algorithm, known as adaptive neural fuzzy inference system (ANFIS), can generate data-driven predictions for flow fields, but has not been applied to natural geophysical flows in large-scale rivers. Herein, we demonstrate the potential of ANFIS to produce three-dimensional (3D) realizations of the instantaneous flood flow field in several large-scale, virtual meandering rivers. The 3D dynamics of flood flow in large-scale rivers were obtained using large-eddy simulation (LES). The LES results, i.e., the 3D velocity components, were employed to train the learnable coefficients of an ANFIS. Further, the trained ANFIS, along with a few time-steps of LES results (precursor data) were then used to produce 3D realizations of flood flow fields in large-scale rivers with geometries other than the one the ANFIS was trained with. We also used the trained ANFIS to generate 3D realizations of river flow at a discharge other than that the ANFIS was trained with. The flow field results obtained from ANFIS were validated using separate LES runs to assess the accuracy of the 3D instantaneous realizations of the machine learning algorithm. An error analysis was conducted to quantify the discrepancies among the ANFIS and LES results for various flood flow predictions in large-scale rivers.

54 ENVIRONMENTAL SCIENCES↗

Hamiltonian learning using machine-learning models trained with continuous measurements

Here, we build upon recent work on the use of machine-learning models to estimate Hamiltonian parameters using continuous weak measurement of qubits as input. We consider two settings for the training of our model: (1) supervised learning, where the weak-measurement training record can be labeled with known Hamiltonian parameters, and (2) unsupervised learning, where no labels are available. The first has the advantage of not requiring an explicit representation of the quantum state, thus potentially scaling very favorably to a larger number of qubits. The second requires the implementation of a physical model to map the Hamiltonian parameters to a measurement record, which we implement using an integrator of the physical model with a recurrent neural network to provide a model-free correction at every time step to account for small effects not captured by the physical model. We test our construction on a system of two qubits and demonstrate accurate prediction of multiple physical parameters in both the supervised context and the unsupervised context. We demonstrate that the model benefits from larger training sets, establishing that it is “learning,” and we show robustness regarding errors in the assumed physical model by achieving accurate parameter estimation in the presence of unanticipated single-particle relaxation.

97 MATHEMATICS AND COMPUTING↗

Methods in PES-Learn: Direct-Fit Machine Learning of Born–Oppenheimer Potential Energy Surfaces

The release of PES-L EARN version 1.0 as an open-source software package for the automatic construction of machine learning models of semi-global molecular potential energy surfaces (PESs) is presented. Improvements to PES-L EARN ’s interoperability are stressed with new Python API that simplifies workflows for PES construction via interaction with QCSchema input and output infrastructure. In addition, a new machine learning method is introduced to PES-L EARN : kernel ridge regression (KRR). The capabilities of KRR are emphasized with examination of select semi-global PESs. All machine learning methods available in PES-L EARN are benchmarked with benzene and ethanol datasets from the rMD17 database to illustrate PES-L EARN ’s performance ability. Fitting performance and timings are assessed for both systems. Finally, the ability to predict gradients with neural network models is presented and benchmarked with ethanol and benzene. PES-L EARN is an active project and welcomes community suggestions and contributions.

kernel ridge regression↗

Use of Physics to Improve Solar Forecast: Part II, Machine Learning and Model Interpretability

Machine learning (ML) models have been applied to forecast solar energy; however, they often lack clarity of interpretability and underlying physics. This work addresses such challenges by developing a hierarchy of ML models that gradually introduce predictors to improve the forecast accuracy based on a physics-based framework. Three ML models (ARIMA, LSTM, and XGBoost) are examined and compared with four physics-informed persistence models reported in Part I and the simple persistence model to assess the improvement of different models. The 7-year measurements at the U.S. Department of Energy's Atmospheric Radiation Measurement's Southern Great Plains Central Facility site are used for forecasts and evaluations. The results reveal that the step-by-step introduction of predictors leads to different improvements for models at different hierarchical levels. Comparison of the ML models with persistence models shows that LSTM and XGBoost outperform all the persistence models, with LSTM having the overall best performance; however, ARIMA underperforms the four physics-informed persistence models. This study demonstrates the importance and utility of incorporating physics into ML models in improving forecast accuracy by introducing a hierarchy of physics-based predictors, distinguishing predictor contributions, and enhancing the ML interpretability. The combined use of Global Horizontal Irradiance (GHI) and Direct Normal Irradiance (DNI) significantly improves the forecast accuracy compared to using individual irradiances alone because the pair contains more information on cloud-radiation interactions.

interpretability↗

Scalable Second Order Optimization for Machine Learning

Many machine learning (ML) training tasks are essentially optimization processes that would at first glance appear eminently parallelizable and scalable. However, effective acceleration of these tasks with scalable parallel hardware has proven to be elusive. While standard methods for machine learning, e.g., stochastic gradient descent (SGD) for DNNs, tend to be resource efficient, they appear to be fundamentally sequential in nature.

97 MATHEMATICS AND COMPUTING↗

Performance Evaluation of Vertical Federated Machine Learning Against Adversarial Threats on Wide-Area Control System: Preprint

Federated machine learning (FL) is gaining significant popularity to develop cybersecurity solutions in power grids because of its advanced capability to support decentralized data handing at local devices, its privacy preservation, and its low-bandwidth requirement. However, the evolving adversarial machine learning (AML) threats raise significant concerns for the cybersecurity of FL architectures. The FL-based split neural network (SplitNN) achieves high performance through the decentralized training of local neural network models while preserving data privacy across multiple entities. In this paper, we propose a methodology for evaluating the performance of a vertical FLbased anomaly detector against different types of AML attacks, including denial-of-service attacks, adversarial data injection attacks, and replay attacks on the trained local models deployed in the grid network. For a case study, we consider the modified IEEE 13-bus system, and we develop SplitNN-based binary and multiclass classification models to detect, locate, and identify different types of data integrity attacks on the volt-watt control with two pooling layers: maximum pooling and AvgPool. Our experimental results, computed through performance metrics, reveal that the severity of these AML attacks varies with the integrated pooling mechanism, the type of classification model, and the nature of the cyberattack. Further, the AML attacks negatively impacted the prediction time per sample for the pretrained SplitNN during the online testing.

adversarial threats↗

Teaching Freight Mode Choice Models New Tricks Using Interpretable Machine Learning Methods

Understanding and forecasting the intricate freight mode choice behavior under various industry, policy, and technology contexts is essential in freight planning and policymaking. Numerous models have been developed in prior studies to provide insights into freight mode selection, the majority of which use discrete choice models such as multinomial logit (MNL) models. However, logit models often rely on linear specifications of independent variables, despite potential nonlinear relationships in the data. Moreover, there often lacks a heuristic and efficient approach to identify such complex relationships to define the logit model specifications. To fill this gap, we developed an MNL model for freight mode choice using the insights from state-of-the- art machine learning (ML) models. ML models can capture the nonlinear nature of the complex decision-making process, and recent advances in 'explainable AI' have greatly improved their interpretability. The interpretable ML methods help enhance the performance of MNL models and advance knowledge of freight mode choice. Specifically, the influential factors and their relationship with individual modes are identified using SHapley Additive exPlanations (SHAP) to improve the MNL's performance. The workflow is demonstrated in a case study of Austin, Texas, and the SHAP results reveal multiple nonlinear relationships predicted by ML models. Incorporating those relationships into MNL model specifications improves the interpretability and accuracy of the MNL model compared to a conventional MNL model. Findings from this study can be used to guide freight planning and inform policymakers and practitioners on how key factors affect freight decision-making.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

Data for A Hybrid Biophysical-Machine Learning Framework for Diurnal Surface Energy Flux Estimation Using Proximal Sensing

Thermal infrared-based remote sensing of surface energy fluxes has traditionally relied on high spatial resolution satellite data with revisit frequencies on the order of weeks. In this study, we evaluate a biophysics-based analytical surface energy balance model for predicting latent energy (LE) and sensible heat (H) fluxes using proximal sensing observations. The Surface Temperature Initiated Closure (STIC1.2) model has been extensively validated across a wide range of spatial and temporal scales using various satellite-derived thermal infrared data sets. Here we extend this validation by applying STIC at sub-hourly temporal resolution over multiple growing seasons for four distinct agricultural systems. We further develop and evaluate novel STIC variants that incorporate machine learning (ML) techniques to eliminate the need for surface energy balance observations, specifically net radiation and soil heat flux, thereby enhancing model applicability in data-sparse settings. The integration of a ML component to estimate surface available energy is shown to have strong predictive performance for both LE (R2 = 0.81–0.94) and H (R2 = 0.46–0.72) across all agricultural systems examined here, demonstrating the potential of hybrid biophysical-machine learning approaches for surface energy balance modeling with minimal data requirements. This study concludes with a novel application of explainable machine learning (exML) to diagnose sources of model error. This exML framework attributes residual prediction errors to both model input variables and environmental drivers not explicitly included in the simulation experiments. This approach provides a new pathway for improving model design and integrating previously overlooked yet influential variables into future model iterations.

AI/ML↗

Learning to Branch with Interpretable Machine Learning Models

Machine learning is being increasingly used in improving decisions made within branch-and-bound algorithms for solving mixed-integer programs (MIPs). Branching is a key component in branch-and-bound algorithms, this work presents IDAES-core project update on building simple and interpretable machine learning models for branching and improving decision-making tools applied for the optimization of advanced energy systems.

Bayramoglu, Selin↗

Python Codebase and Jupyter Notebooks - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

Git archive containing Python modules and resources used to generate machine-learning models used in the "Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada" project. This software is licensed as free to use, modify, and distribute with attribution. Full license details are included within the archive. See "documentation.zip" for setup instructions and file trees annotated with module descriptions.

Brown, Stephen↗

Predicting wind farm operations with machine learning and the P2D‐RANS model: A case study for an AWAKEN site

Abstract The power performance and the wind velocity field of an onshore wind farm are predicted with machine learning models and the pseudo‐2D RANS model, then assessed against SCADA data. The wind farm under investigation is one of the sites involved with the American WAKE experimeNt (AWAKEN). The performed simulations enable predictions of the power capture at the farm and turbine levels while providing insights into the effects on power capture associated with wake interactions that operating upstream turbines induce, as well as the variability caused by atmospheric stability. The machine learning models show improved accuracy compared to the pseudo‐2D RANS model in the predictions of turbine power capture and farm power capture with roughly half the normalized error. The machine learning models also entail lower computational costs upon training. Further, the machine learning models provide predictions of the wind turbulence intensity at the turbine level for different wind and atmospheric conditions with very good accuracy, which is difficult to achieve through RANS modeling. Additionally, farm‐to‐farm interactions are noted, with adverse impacts on power predictions from both models.

17 WIND ENERGY↗

Inferring colloidal interaction from scattering by machine learning

A machine learning solution for the potential inversion problem in elastic scattering is outlined. The inversion scheme consists of two major components, a generative network featuring a variational autoencoder which extracts the targeted static two-point correlation functions from experimentally measured scattering cross sections, and a Gaussian process framework which probabilistically infers the relevant structural parameters from the inverted correlation functions. Via a case study of charged colloidal suspensions, the feasibility of this approach for quantitative study of molecular interaction is critically benchmarked and its merit over existing deterministic approaches, in terms of numerical accuracy and computationally efficiency, is demonstrated.

36 MATERIALS SCIENCE↗