Developing statistical and machine learning models for predicting CO2 solubility in live crude oils
Not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
Not Available
Not Available
This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.
Big Earth Data are too big to be tractable to simple data inspection and require models to make sense of all the data. Useful models for Big Earth Data may be physical, statistical, or machine learning based. In many cases, hybrid models combine attributes of two or more of these types.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
It is known that an effective control system is the key condition for successful implementation of high-performance magnetic servo systems. Major issues to design such control systems are nonlinearity; unmodeled dynamics, such as secondary effects for copper resistance, stray fields, and saturation; and that disturbance rejection for the load effect reacts directly on the servo system without transmission elements. One typical approach to design control systems under these conditions is a special type of nonlinear feedback called gain scheduling. It accommodates linear regulators whose parameters are changed as a function of operating conditions in a preprogrammed way. In this paper, an on-line learning fuzzy control strategy is proposed. To inherit the wealth of linear control design, the relations between linear feedback and fuzzy logic controllers have been established. The exercise of engineering axioms of linear control design is thus transformed into tuning of appropriate fuzzy parameters. Furthermore, fuzzy logic control brings the domain of candidate control laws from linear into nonlinear, and brings new prospects into design of the local controllers. On the other hand, a self-learning scheme is utilized to automatically tune the fuzzy rule base. It is based on network learning infrastructure; statistical approximation to assign credit; animal learning method to update the reinforcement map with a fast learning rate; and temporal difference predictive scheme to optimize the control laws. Different from supervised and statistical unsupervised learning schemes, the proposed method learns on-line from past experience and information from the process and forms a rule base of an FLC system from randomly assigned initial control rules.
We introduce a machine-learning framework named statistics-informed neural network (SINN) for learning stochastic dynamics from data. This new architecture was theoretically inspired by a universal approximation theorem for stochastic systems, which we introduce in this paper, and the projection-operator formalism for stochastic modeling. Here, we devise mechanisms for training the neural network model to reproduce the correct statistical behavior of a target stochastic process. Numerical simulation results demonstrate that a well-trained SINN can reliably approximate both Markovian and non-Markovian stochastic dynamics. We demonstrate the applicability of SINN to coarse-graining problems and the modeling of transition dynamics. Furthermore, we show that the obtained reduced-order model can be trained on temporally coarse-grained data and hence is well suited for rare-event simulations.
The high complexity of modern aircraft and spacecraft requires elaborate Verification and Validation (V&V) approaches to make sure that such complex systems work properly and reliably. MARGInS is a framework for the analysis, understanding, and prediction of the behavior of a complex, hybrid system. MARGInS contains a set of machine learning and statistical algorithms for multivariate clustering, treatment learning, critical factor determination, time-series analysis, event prediction, and safety-boundary detection and characterization. The framework supports system testing and can be configured to find novel features in test suites, determine classes of behavior, propose new experiments that can efficiently explore and characterize the boundaries between classes of system behavior, and to create visualizations and reports.
High temperature oxidation and corrosion degradation mechanisms dictate the lifetime of materials critical to energy production. The combination of modeling and experimental approaches such as machine learning (ML) and data analytics, with sufficient experimental data, can accelerate the development of new materials while limiting its cost. In the present work, ML will be applied to two high temperature oxidation data libraries (Oak Ridge National Laboratory and National Air and Space Administration) that comprised of about 5000 mass change sample datasheets for a variety of materials and temperatures in dry air and air + 10 % H2O. A python code was developed to prepare the data for machine learning by collecting and formatting oxidation rate constants, alloy compositions and environment of exposure into a single data frame. Scikit-learn library and Statistics and Machine Learning Toolbox within MathWorks were then used to perform unsupervised clustering and supervised regression learning. The impact of dataset distribution on the performance of the developed ML models was evaluated. Potential strategies to improve the predictions and enhance extrapolative capability of the previously trained model were investigated.
Robust in situ magnetic field measurements are critical to understanding the various mechanisms that couple mass, momentum, and energy throughout our solar system. However, the spacecraft on which magnetometers are often deployed contaminate the magnetic field measurements via onboard subsystems including reaction wheels and magnetorquers. Two magnetometers can be deployed at different distances from the spacecraft to determine an approximation of the interfering field for subsequent removal, but constant data streams from both magnetometers can be impractical due to power and telemetry limitations. Here we propose a method to identify and remove time-varying magnetic interference from sources such as reaction wheels using statistical decomposition and convolutional neural networks, providing high-fidelity magnetic field data even in cases where dual-sensor measurements are not constantly available. For example, a measurement interval from the Parker Solar Probe outboard magnetometer experienced a 95.1% reduction in reaction wheel interference following application of the proposed technique.
An accurate representation of hydrodynamic force and torque experienced by every particle in a distribution can be obtained from particle resolved (PR) simulations. These unique quantities are influenced by the deterministic position of surrounding particles. However, systems simulated with this methodology are typically limited to particles due to the involved computational cost. This resource requirement is a major bottleneck in analyzing the effect of variations in particle distribution. Here, this article attempts to address this bottleneck by availing relatively inexpensive deep learning models. The surrogate models that we employ in this article use a physics‐based hierarchical framework and symmetry‐preserving neural networks to achieve robustness with limited training data. This article first performs additional generalizability tests on PR data of distinct distributions that are not involved in the training process. The models are then deployed on several different particle distributions. Impact of clustering and structure on the observed statistics are investigated.
Abstract Diffraction techniques can powerfully and nondestructively probe materials while maintaining high resolution in both space and time. Unfortunately, these characterizations have been limited and sometimes even erroneous due to the difficulty of decoding the desired material information from features of the diffractograms. Currently, these features are identified non-comprehensively via human intuition, so the resulting models can only predict a subset of the available structural information. In the present work we show (i) how to compute machine-identified features that fully summarize a diffractogram and (ii) how to employ machine learning to reliably connect these features to an expanded set of structural statistics. To exemplify this framework, we assessed virtual electron diffractograms generated from atomistic simulations of irradiated copper. When based on machine-identified features rather than human-identified features, our machine-learning model not only predicted one-point statistics (i.e. density) but also a two-point statistic (i.e. spatial distribution) of the defect population. Hence, this work demonstrates that machine-learning models that input machine-identified features significantly advance the state of the art for accurately and robustly decoding diffractograms.
Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.
Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.
Not Available
The industrial sector consumes about one-third of global energy, making them a frequent target for energy use reduction. Variation in energy usage is observed with weather conditions, as space conditioning needs to change seasonally, and with production, energy-using equipment is directly tied to production rate. Previous models were based on engineering analyses of equipment and relied on site-specific details. Others consisted of single-variable regressors that did not capture all contributions to energy consumption. Further, new modeling techniques could be applied to rectify these weaknesses. Applying data from 45 different manufacturing plants obtained from industrial energy audits, a supervised machine-learning model is developed to create a general predictor for industrial building energy consumption. The model uses features of air enthalpy, solar radiation, and wind speed to predict weather-dependency; motor, steam, and compressed air system parameters to capture support equipment contributions; and operating schedule, production rate, number of employees, and floor area to determine production-dependency. Results showed that a model that used a linear regressor over a transformed feature space could outperform a support vector machine and utilize features more representative of physical systems. Using informed parameters to build a reliable predictor will more accurately characterize a manufacturing facility's energy savings opportunities.