Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “MACHINE LEARNING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Characterizing Mesoscale Cellular Convection in Marine Cold Air Outbreaks With a Machine Learning Approach

Abstract During marine cold‐air outbreaks (MCAOs), when cold polar air moves over warmer ocean, a well‐recognized cloud pattern develops, with open or closed mesoscale cellular convection (MCC) at larger fetch over open water. The Cold‐Air Outbreaks in the Marine Boundary Layer Experiment provided a comprehensive set of ground‐based in situ and remote sensing observations of MCAOs at a coastal location in northern Norway. MCAO periods that unambiguously exhibit open or closed MCC are determined. Individual cells observed with a profiling Ka‐band radar are identified using a watershed segmentation method. Using self‐organizing maps (SOMs), these cells are then objectively classified based on the variability in their vertical structure. The SOM nodes contain some information about the location of the cell transect relative to the center of the MCC. This adds classification noise, requiring numerous cell transects to isolate cell dynamical information. The SOM‐based classification shows that comparatively intense convection occurs only in open MCC. This convection undergoes an apparent lifecycle. Developing cells are associated with stronger updrafts, large spectrum width, larger amounts of liquid water, lower surface precipitation rates, and lower cloud tops than mature and weakening cells. The weakening of these cells is associated with the development of precipitation‐induced cold pools. The SOM classification also reveals less intense convection, with a similar lifecycle. More stratiform vertical cloud structures with weak vertical motions are common during closed MCC periods and are separated into precipitating and non‐precipitating stratiform cores. Convection is observed only occasionally in the closed MCC environment.

Meteorology & Atmospheric Sciences↗

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

Accurate Thermochemistry of Complex Lignin Structures via Density Functional Theory, Group Additivity, and Machine Learning

A molecular-level understanding of lignin structures and bond dissociation energies could facilitate depolymerization technologies. Still, this information is currently limited due to the lack of databases and the simplification of surrogate models. Here, substitution effects on seven common linkages in lignin polymers are systematically investigated. An automated reaction network generator is employed to create a database of structures. A new group additivity (GA) model based on principal component analysis (PCA) descriptors is introduced and trained on gas-phase density functional theory data of 4100 species at the M06-2X/6-311++G(d,p) level. Hydrogen bonds, local steric, and nonaromatic ring contributions are also incorporated. Lastly, we improve the accuracy of the group additivity model to reach the G4 theory by computing a data set of 770 species at this level and using a data fusion approach.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Opportunities for human factors in machine learning

Introduction The field of machine learning and its subfield of deep learning have grown rapidly in recent years. With the speed of advancement, it is nearly impossible for data scientists to maintain expert knowledge of cutting-edge techniques. This study applies human factors methods to the field of machine learning to address these difficulties. Methods Using semi-structured interviews with data scientists at a National Laboratory, we sought to understand the process used when working with machine learning models, the challenges encountered, and the ways that human factors might contribute to addressing those challenges. Results Results of the interviews were analyzed to create a generalization of the process of working with machine learning models. Issues encountered during each process step are described. Discussion Recommendations and areas for collaboration between data scientists and human factors experts are provided, with the goal of creating better tools, knowledge, and guidance for machine learning scientists.

97 MATHEMATICS AND COMPUTING↗

Learning to Branch with Interpretable Machine Learning Models

This presentation describes an algorithm for applying machine learning to branching to speed up the solution of integer optimization problems. These problems are challenging and solved multiple times a day by power systems operators. We show that our approach speeds up a widely used open-source optimization solver.

Bayramoglu, Selin↗

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics↗

Impact of classical statistics on thermal conductivity predictions of BAs and diamond using machine learning molecular dynamics

Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.

36 MATERIALS SCIENCE↗

AENET–LAMMPS and AENET–TINKER : Interfaces for accurate and efficient molecular dynamics simulations with machine learning potentials

Machine-learning potentials (MLPs) trained on data from quantum-mechanics based first-principles methods can approach the accuracy of the reference method at a fraction of the computational cost. To facilitate efficient MLP-based molecular dynamics and Monte Carlo simulations, an integration of the MLPs with sampling software is needed. Here, we develop two interfaces that link the atomic energy network (ænet) MLP package with the popular sampling packages TINKER and LAMMPS. The three packages, ænet, TINKER, and LAMMPS, are free and open-source software that enable, in combination, accurate simulations of large and complex systems with low computational cost that scales linearly with the number of atoms. Scaling tests show that the parallel efficiency of the ænet–TINKER interface is nearly optimal but is limited to shared-memory systems. The ænet–LAMMPS interface achieves excellent parallel efficiency on highly parallel distributed memory systems and benefits from the highly optimized neighbor list implemented in LAMMPS. We demonstrate the utility of the two MLP interfaces for two relevant example applications: the investigation of diffusion phenomena in liquid water and the equilibration of nanostructured amorphous battery materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterizing Seasonal Variation of the Atmospheric Mixing Layer Height Using Machine Learning Approaches

As machine learning becomes more integrated into atmospheric science, XGBoost has gained popularity for its ability to assess the relative contributions of influencing factors in the atmospheric boundary layer height. To examine how these factors vary across seasons, a seasonal analysis is necessary. However, dividing data by season reduces the sample size, which can affect result reliability and complicate factor comparisons. To address these challenges, this study replaces default parameters with grid search optimization and incorporates cross-validation to mitigate dataset limitations. Using XGBoost with four years of data from the atmospheric radiation measurement (ARM) (Southern Great Plains (SGP) C1 site, cross-validation stabilizes correlation coefficient fluctuations from 0.3 to within 0.1. With optimized parameters, the R value can reach 0.81. Analysis of the C1 site reveals that the relative importance of different factors changes across seasons. Lower tropospheric stability (LTS, ~0.53) is the dominant factor at C1 throughout the year. However, during DJF, latent heat flux (LHF, 0.44) surpasses LTS (0.22). In SON, LTS (0.58) becomes more influential than LHF (0.18). Further comparisons among the four long-term SGP sites (C1, E32, E37, and E39) show seasonal variations in relative importance. Notably, during JJA, the differences in the relative importance of the three factors across all sites are lower than in other seasons. This suggests that boundary layer development in the summer is not dominated by a single factor, reflecting a more intricate process likely influenced by seasonal conditions such as enhanced convective activity, higher temperatures, and humidity, which collectively contribute to a balanced distribution of parameter impacts. Furthermore, the relative importance of LTS gradually increases from morning to noon, indicating that LTS becomes more significant as the boundary layer approaches its maximum height. Consequently, the LTS in the early morning in autumn exhibits greater relative importance compared to other seasons. This reflects a faster development of the mixing layer height (MLH) in autumn, suggesting that it is easier to retrieve the MLH from the previous day during this period. The findings enhance understanding of boundary layer evolution and contribute to improved boundary layer parameterization.

54 ENVIRONMENTAL SCIENCES↗

Machine-learning identification of the variability of mean velocity and turbulence intensity for wakes generated by onshore wind turbines: Cluster analysis of wind LiDAR measurements

Light detection and ranging (LiDAR) measurements of isolated wakes generated by wind turbines installed at an onshore wind farm are leveraged to characterize the variability of the wake mean velocity and turbulence intensity during typical operations, which encompass a breadth of atmospheric stability regimes and rotor thrust coefficients. The LiDAR measurements are clustered through the k-means algorithm, which enables identifying the most representative realizations of wind turbine wakes while avoiding the imposition of thresholds for the various wind and turbine parameters. Considering the large number of LiDAR samples collected to probe the wake velocity field, the dimensionality of the experimental dataset is reduced by projecting the LiDAR data on an intelligently truncated basis obtained with the proper orthogonal decomposition (POD). The coefficients of only five physics-informed POD modes are then injected in the k-means algorithm for clustering the LiDAR dataset. The analysis of the clustered LiDAR data and the associated supervisory control and data acquisition and meteorological data enables the study of the variability of the wake velocity deficit, wake extent, and wake-added turbulence intensity for different thrust coefficients of the turbine rotor and regimes of atmospheric stability. Furthermore, the cluster analysis of the LiDAR data allows for the identification of systematic off-design operations with a certain yaw misalignment of the turbine rotor with the mean wind direction.

17 WIND ENERGY↗

Understanding Strain and Failure of a Knot in Polyethylene Using Molecular Dynamics with Machine-Learned Potentials

A neural network potential (NNP) has been developed by fitting to ab initio electronic structure data on hydrocarbons and is used to study failure of linear and knotted polyethylene (PE) chains. A linear PE chain must be highly strained before breaking as the stress is equally distributed across the chain. In contrast, the stress in a PE chain with a 31 or overhand knot, accumulates at the knot’s entrance/exit. We find the strain energy is greatest when the bond length and angle are strained simultaneously, and that the knot weakens the chain by increasing the variance of the C–C–C angle, thereby allowing rupture at lower bond strains. Here, we extend our analysis to both 51 and 52 knots and find that both break at the entrance/exit of a loop. Notably, molecular scale PE knots exhibit many of the same characteristics as knots in a macroscopic rope, with stick–slip phenomena upon tightening and similar points of failure.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Application of Machine Learning to Rotorcraft Health Monitoring

Machine learning is a powerful tool for data exploration and model building with large data sets. This project aimed to use machine learning techniques to explore the inherent structure of data from rotorcraft gear tests, relationships between features and damage states, and to build a system for predicting gear health for future rotorcraft transmission applications. Classical machine learning techniques are difficult, if not irresponsible to apply to time series data because many make the assumption of independence between samples. To overcome this, Hidden Markov Models were used to create a binary classifier for identifying scuffing transitions and Recurrent Neural Networks were used to leverage long distance relationships in predicting discrete damage states. When combined in a workflow, where the binary classifier acted as a filter for the fatigue monitor, the system was able to demonstrate accuracy in damage state prediction and scuffing identification. The time dependent nature of the data restricted data exploration to collecting and analyzing data from the model selection process. The limited amount of available data was unable to give useful information, and the division of training and testing sets tended to heavily influence the scores of the models across combinations of features and hyper-parameters. This work built a framework for tracking scuffing and fatigue on streaming data and demonstrates that machine learning has much to offer rotorcraft health monitoring by using Bayesian learning and deep learning methods to capture the time dependent nature of the data. Suggested future work is to implement the framework developed in this project using a larger variety of data sets to test the generalization capabilities of the models and allow for data exploration.

machine learning↗

Decision Science for Machine Learning (DeSciML)

The increasing use of machine learning (ML) models to support high-consequence decision making drives a need to increase the rigor of ML-based decision making. Critical problems ranging from climate change to nonproliferation monitoring rely on machine learning for aspects of their analyses. Likewise, future technologies, such as incorporation of data-driven methods into the stockpile surveillance and predictive failure analysis for weapons components, will all rely on decision-making that incorporates the output of machine learning models. In this project, our main focus was the development of decision scientific methods that combine uncertainty estimates for machine learning predictions, with a domain-specific model of error costs. Other focus areas include uncertainty measurement in ML predictions, designing decision rules using multiobjecive optimization, the value of uncertainty reduction, and decision-tailored uncertainty quantification for probability estimates. By laying foundations for rigorous decision making based on the predictions of machine learning models, these approaches are directly relevant to every national security mission that applies, or will apply, machine learning to data, most of which entail some decision context.

97 MATHEMATICS AND COMPUTING↗

Continual Learning for Production-Level Machine Learning in Particle Accelerators

Particle accelerators operate in complex environments where data distribution can change dynamically, leading to data drifts that significantly challenge Machine Learning (ML) models. These non-stationary conditions often cause ML models to deteriorate in performance, making it difficult to maintain reliable predictions in operation. The primary sources of data drifts are changes in accelerator settings and changes in equipment performance which cannot be measured directly. To bridge this gap between ML development and long-term deployment in operational settings, we identify key areas within particle accelerators where continual learning can help mitigate drift-induced performance degradation. We will provide a practical guide on selecting the appropriate method given resource constraints and desired stability plasticity trade offs. As a concrete example, we will present a real-world use case for anomaly detection to predict errant beams at the Spallation Neutron Source accelerator, where continual learning has been employed to demonstrate stable performance on drifting data streams. We will present practical challenges, lessons learned, and the results from the deployed ML model.

Rajput, Kishansingh [Thomas Jefferson National Acc↗