Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “support vector regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Verification and validation of developed short-term forecasting models

Recent advancements in machine learning (ML) and artificial intelligence (AI) technologies provide an opportunity for leveraging data-driven algorithms to predict future nuclear power plant (NPP) operating conditions by using recorded plant process data. Successfully implementing these models can lead to cost-reducing, conditioned-based predictive maintenance through optimized maintenance schedules and a reduction of unnecessary maintenance activities. This report discusses the verification and validation of short-term forecasting processes (i.e., data cleaning, feature selection, model optimization, and forecasting) developed in previous reports. The verification and validation (V&V) process demonstrates the expected precision and accuracy when the ML model encounters new datasets from different systems. Shapley additive explanations were used as the primary means of feature selection across these different data set. Individual models were trained for each data set, then validated through a cross-validation procedure. In this report, two different ML models were tasked to predict variables from three different plant process data sets with varying prediction horizons. The results indicate that support vector regression (SVR) outperformed long short-term memory (LSTM) neural networks in regard to each data set and each prediction horizon in this study, but further tuning and optimization could improve long short-term memory results. However, each forecasting model showed reduced performance as the prediction horizon was extended from 1 hour to 1 day ahead. Research is ongoing to evaluate the optimal input variable space, which is based on a given set of process parameters, to further improve forecasting accuracy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Comparative Study of Wind Energy Potential Estimation Methods for Wind Sites in Togo and Benin (West Sub-Saharan Africa)

The characterization of wind speed distribution and the optimal assessment of wind energy potential are critical factors in selecting a suitable site for wind power plants (WPP). The Weibull distribution law has been used extensively to analyze the wind characteristics of candidate WPP sites, and to estimate the available and deliverable energy. This paper presents a comparative study of five wind energy resource assessment methods as they applied to the context of wind sites in West Sub-Saharan Africa. We investigated three numerical approaches, namely, the adaptive neuro-fuzzy inference system (ANFIS), the multilayer perceptron method (MLP), and support vector regression (SVR), to derive the distribution law of wind speeds and to optimally quantify the corresponding wind energy potential. Next, we compared these three approaches to two well-known Weibull distribution law-based methods: the empirical method of Justus (EMJ) and the maximum likelihood method (MLM). Case study results indicated that the neural network-based methods, ANFIS and MLP, yielded the most accurate distribution fits and wind energy potential estimates, and consequently, are the most recommended methods for the wind sites in Togo and Benin. The orders of magnitude of the root mean squared error (RMSE) in estimating the recoverable energy using ANFIS were, respectively, 10-4 and 10-5 for Lomé and Cotonou, while MLP achieved an RMSE order of magnitude of 10-3 for both sites.

17 WIND ENERGY↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine learning-based ethylene concentration estimation, real-time optimization and feedback control of an experimental electrochemical reactor

With the increase in electricity supply from clean energy sources, electrochemical reduction of carbon dioxide (CO 2 ) has received increasing attention as an alternative source of carbon-based fuels. As CO 2 reduction is becoming a stronger alternative for the clean production of chemicals, the need to model, optimize and control the electrochemical reduction of the CO 2 process becomes inevitable. However, on one hand, a first-principles model to represent the electrochemical CO 2 reduction has not been fully developed yet because of the complexity of its reaction mechanism, which makes it challenging to define a precise state-space model for the control system. On the other hand, the unavailability of efficient concentration measurement sensors continues to challenge our ability to develop feedback control systems. Gas chromatography (GC) is the most common equipment for monitoring the gas product composition, but it requires a period of time to analyze the sample, which means that GC can provide only delayed measurements during the operation. Moreover, the electrochemical CO 2 reduction process is catalyzed by a fast-deactivating copper catalyst and undergoes a selectivity shift from the product-of-interest at the later stages of experiments, which can pose a challenge for conventional control methods. To this end, machine learning (ML) techniques provide a potential approach to overcome those difficulties due to their demonstrated ability to capture the dynamic behavior of a chemical process from data. Motivated by the above considerations, we propose a machine learning-based modeling methodology that integrates support vector regression and first-principles modeling to capture the dynamic behavior of an experimental electrochemical reactor; this model, together with limited gas chromatography measurements, is employed to predict the evolution of gas-phase ethylene concentration. The model prediction is directly used in a proportional-integral (PI) controller that manipulates the applied potential to regulate the gas-phase ethylene concentration at energy-optimal set-point values computed by a real-time process optimizer (RTO). Specifically, the RTO calculates the operation set-point by solving an optimization problem to maximize the economic benefit of the reactor. Finally, suitable compensation methods are introduced to further account for the experimental uncertainties and handle catalyst deactivation. The proposed modeling, optimization, and control approaches are the first demonstration of active control for a CO 2 electrolyzer and contribute to the automation and scale-up efforts for electrified manufacturing of fuels and chemicals starting from CO 2 .

42 ENGINEERING↗

Analytics-at-scale of Sensor Data for Digital Monitoring in Nuclear Plants (3 rd Annual Report)

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This report primarily focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 ± 0.014, 0.0026 ± 0, and 0.063 ± 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure. This report summarizes the Fiscal Year 2021 research progress encompassing the (1) data cleaning and feature selection necessary for ML applications; (2) development of short-term forecasting models to predict future plant process parameters for both single and multiple time steps ahead; and (3) validation of the feature selection methods and short-term forecasting models given new data from different systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Online LIBS–ML Framework for Dynamic Characterization of Heterogeneous Waste-Derived Gasification Feedstocks

LIBS−ML framework for real time feedstock characterization during continuous conveyor transport Heterogeneous waste derived feedstocks (e.g., waste coal, biomass and blends) introduce rapid variability in heating value and ash chemistry that affect gasifier operation, yet conventional laboratory characterization techniques are too slow to support proactive control. To address this gap, this study reports on an online, in situ, dynamic characterization framework that couple’s laser-induced breakdown spectroscopy (LIBS) with leakage safe machine learning (ML) regression to deliver real time, decision quality predictions of gasifier relevant properties. A controlled sample matrix spanning two different waste coals, two different biomasses, and engineered blends under two particle size conditions were constructed and benchmarked using standardized laboratory analyses for proximate/ultimate properties and ash composition. LIBS spectra were acquired dynamically as material flowed on a conveyor belt, using high energy 1064 nm laser ablation and shot averaging to improve repeatability and precision. Supervised regression models (multi layer perceptron (MLP) /artificial neural network (ANN), random forest (RF), and support vector regression (SVR)) and an optimized weighted ensemble were trained on emission line feature sets using nested cross validation with Bayesian hyperparameter tuning and validated against an independent hold out set. The proposed LIBS−ML workflow achieves near laboratory predictive fidelity across parametric targets (including higher heating value (HHV), ash content, fixed carbon, sulfur, major ash forming oxides, and initial deformation temperature (IDT)), with the weighted ensemble providing a robust default predictor under dynamic measurement conditions. These results demonstrate a practical pathway for real time feedstock characterization that can enable feedforward adjustments and more resilient gasifier operation for variable quality waste derived fuels.

Biomass↗

Classification and regression models of audio and vibration signals for machine state monitoring in precision machining systems

Here we present a data-driven method for monitoring machine status in manufacturing processes. Audio and vibration data from precision machining are used for inference in two operating scenarios: (a) variable machine health states (anomaly detection); and (b) settings of machine operation (state estimation). Audio and vibration signals are first processed through Fast Fourier Transform and Principal Component Analysis to extract transformed and informative features. These features are then used in the training of classification and regression models for machine state monitoring. Specifically, three classifiers (K-nearest neighbors, convolutional neural networks and support vector machines) and two regressors (support vector regression and neural network regression) were explored, in terms of their accuracy in machine state prediction. It is shown that the audio and vibration signals are sufficiently rich in information about the machine that 100% state classification accuracy could be accomplished. Data fusion was also explored, showing overall superior accuracy of data-driven regression models.

42 ENGINEERING↗

Gaussian Process Regression for Aggregate Baseline Load Forecasting

Demand response (DR) is one of the most effective ways to maintain the reliability and improve the flexibility of power systems. Accurate forecasts of baseline loads are essential for DR programs. In the era of big data, machine learning-based approaches present a unique opportunity for baseline load forecasting. Thus, this paper presents a machine learning-based approach using a relatively less explored algorithm, Gaussian process regression (GPR), to forecast aggregate baseline loads. As such, a dataset was generated using a set of EnergyPlus simulations. Using the generated dataset, a GPR-based forecasting model was developed. In addition, support vector regression (SVR)-, artificial neural network (ANN)-, and averaging-based models were developed as baseline models for comparison. These models were compared in terms of accuracy, simplicity, and integrity. The prediction performance of the models showed that the GPR-based model is more accurate and reliable than the others. Such high performance shows the potential of the GPR in baseline load forecasting. GPR, therefore, can be used for DR applications.

Amasyali, Kadir↗

Taylor-Expansion-Based Robust Power Flow in Unbalanced Distribution Systems: A Hybrid Data-Aided Method

Traditional power flow methods often adopt certain assumptions designed for passive balanced distribution systems, thus lacking practicality for unbalanced operation. moreover, their computation accuracy and efficiency are heavily subject to unknown errors and bad data in measurements or prediction data of distributed energy resources (ders). to address these issues, this paper proposes a hybrid data-aided robust power flow algorithm in unbalanced distribution systems, which combines taylor series expansion knowledge with a data-driven regression technique. the proposed method initiates a linearization power flow model to derive an explicitly analytical solution by modified taylor expansion. to mitigate the approximation loss that surges due to the der integration and bad data, we further develop a data-aided robust support vector regression approach to estimate the errors efficiently. comparative analysis in the 13-bus and 123-bus ieee unbalanced feeders shows that the proposed hybrid algorithm achieves superior computational efficiency, with guaranteed accuracy and robustness against outliers.

data-driven↗

Evapotranspiration partitioning estimates from 8 methods from 47 NEON sites, 2019-2021

This dataset provides daily estimates of evapotranspiration (ET) and the transpiration-to-evapotranspiration ratio (T/ET) across 47 terrestrial National Ecological Observatory Network (NEON) sites spanning diverse environmental and biome conditions in the United States across three years of data (2019-2021). Daily ET is reported in both energy units (MJ m⁻² day⁻¹) and equivalent water depth (mm day⁻¹), assuming a constant latent heat of vaporization of 2.45 MJ/kg. The primary method uses a hybrid recurrent neural network–Penman–Monteith framework (RNN-PM), which integrates physically based surface energy balance constraints with data-driven learning to partition ET into transpiration and evaporation components. Model inputs include in situ meteorological observations (air temperature, vapor pressure deficit, wind speed, and radiation) combined with satellite-derived land surface temperature, leaf area index, and soil moisture. For benchmarking and uncertainty assessment, T/ET estimates from seven additional models are included: Priestley-Taylor Jet Propulsion Laboratory (PT-JPL), Penman-Monteith (P-M), Two-Source Energy Balance (TSEB), Support Vector Regression (SVR), and Categorical Boosting (CatBoost), among others—spanning empirical, machine-learning, and process-based approaches (see methods section or linked publication for detailed descriptions). Data Package Contents: The dataset a csv files containing daily ET and T/ET estimates for each site and model, along with associated metadata files these variables. Data can be accessed using common spreadsheet software (e.g., Microsoft Excel, LibreOffice) or programming environments such as R or Python. Together, these data support cross-site comparisons of ecosystem water use, evaluation of ET partitioning methods, and development of improved land–atmosphere exchange models.

EARTH SCIENCE > ATMOSPHERE↗

In Silico Prediction of the Toxicity of Nitroaromatic Compounds: Application of Ensemble Learning QSAR Approach

In this work, a dataset of more than 200 nitroaromatic compounds is used to develop Quantitative Structure–Activity Relationship (QSAR) models for the estimation of in vivo toxicity based on 50% lethal dose to rats (LD 50 ). An initial set of 4885 molecular descriptors was generated and applied to build Support Vector Regression (SVR) models. The best two SVR models, SVR_A and SVR_B, were selected to build an Ensemble Model by means of Multiple Linear Regression (MLR). The obtained Ensemble Model showed improved performance over the base SVR models in the training set (R 2 = 0.88), validation set (R 2 = 0.95), and true external test set (R 2 = 0.92). The models were also internally validated by 5-fold cross-validation and Y-scrambling experiments, showing that the models have high levels of goodness-of-fit, robustness and predictivity. The contribution of descriptors to the toxicity in the models was assessed using the Accumulated Local Effect (ALE) technique. The proposed approach provides an important tool to assess toxicity of nitroaromatic compounds, based on the ensemble QSAR model and the structural relationship to toxicity by analyzed contribution of the involved descriptors.

54 ENVIRONMENTAL SCIENCES↗

Machine learning assisted phase and size-controlled synthesis of iron oxide particles

Synthesis of iron oxides with specific phases and particle sizes is a crucial challenge in various fields, including materials science, energy storage, biomedical applications, environmental science, and earth science. However, despite significant advances in this area, much of the current palette of particle outcomes has been based on time-consuming trial-and-error exploration of synthesis conditions. The present study was designed to explore a very different approach to 1) predict the outcome of synthesis from specified reaction parameters based on using machine learning (ML) techniques, and 2) correlate sets of parameters to obtain products with desired outcomes by a newly designed recommendation algorithm. To achieve this, four ML algorithms were tested, namely random forest, logistic regression, support vector machine, and k-nearest neighbor. Among the models, random forest outperformed the others, attaining 96% and 81% accuracy when predicting the phase and size of iron oxide particles in the test dataset. Surprisingly, the permutation feature importance analysis revealed that volume, which may strongly relate to pressure, was one of the important features, along with precursor concentration, pH, temperature, and time, influencing the phase and size of iron oxide particles during synthesis. To verify the robustness of the random forest models, prediction and experimental results were compared based on 24 randomly generated methods in additive and non-additive systems not included in the datasets. The predictions of product phase and particle size from the models agreed well with the experimental results. Furthermore, a searching and ranking algorithm was developed to recommend potential synthesis parameters for obtaining iron oxide products with the desired phase and particle size from previous studies in the dataset. Furthermore, this study lays the foundation for a closed-loop approach in materials synthesis and preparation, beginning with suggesting potential reaction parameters from the dataset and predicting potential outcomes, followed by conducting experiments and analyses, and ultimately enriching the dataset.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Integration and Optimization of a Waste Heat Driven Organic Rankine Cycle for Power Generation in Wastewater Treatment Plants

The study focuses on achieving energy self-sufficiency in Wastewater Treatment Plants by proposing a comprehensive model for integrating, sizing, and optimizing an Organic Rankine Cycle system. The Organic Rankine Cycle system is designed to utilize waste heat from the gensets at As Samra Wastewater Treatment Plant in Jordan, where it will contribute to the overall electrical energy supply of the plant. Real data from As Samra Wastewater Treatment Plant is used to model and calculate the available waste heat using TRNSYS® software. The Organic Rankine Cycle model is then developed using ASPEN PLUS® software to explore the impact of operational parameters and determine their optimal values for maximizing the plant's energy profile. An economic analysis is conducted to assess the feasibility of the proposed model, considering system components, installation, operation, and maintenance costs. To optimize the Organic Rankine Cycle system, the study employs the Multi-Output Support Vector Regression technique to capture nonlinear relationships between independent variables (fluid type, turbine inlet pressure, turbine inlet temperature, turbine outlet pressure, and mass flow rate) and dependent variables (pump power input, waste heat input, and turbine specific work). The Osprey optimization algorithm is used to address the multi-objective optimization problem, with the proposed Pareto-based Osprey Optimization Algorithm and the Multi-Objective Particle Swarm Optimization technique being employed to evaluate critical performance and economic parameters such as system thermal efficiency, net power output, and the levelized cost of electricity. The results of the optimization strategies indicate that the M-SVR model's prediction accuracy is significantly improved after parameter optimization, with the model returning high R 2 and low Mean Square Error values of 0.991 and 0.00216, respectively. The Pareto-Based Osprey Optimization Algorithm optimizer identifies the best working fluid as Isobutane/Isopentane in a ratio of 66:34, with optimal turbine inlet pressure and temperature of 15 bars and 218 °C, respectively. In conclusion, the Organic Rankine Cycle model at these optimal conditions achieves a cycle efficiency of 19.93% and an Levelized Cost of Electricity value of 0.0353 USD/kWh.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A comprehensive techno-eco-assessment of CO 2 enhanced oil recovery projects using a machine-learning assisted workflow

Carbon dioxide enhanced oil recovery (CO 2 -EOR) projects not only extract residual oil but also sequestrate CO 2 in the depleted reservoirs. Here, this study develops a machine-learning-based workflow to co-optimize the hydrocarbon recovery, CO 2 sequestration volume and project net present value (NPV) simultaneously. Considering the trade-off relationships among the objective functions, support vector regression with Gaussian kernel (Gaussian- SVR) proxies are coupled with multi-objective particle swarm optimization (PSO) protocol and generate Pareto optimal solutions. Taking advantage of the high computational efficacy of the proxy model, economic uncertainties introduced by tax credits, capital costs and oil price are investigated by this study. The results indicate that the tax incentive policy (Section 45Q) plays a vital role in enhancing the economic returns of CO 2 -EOR projects, especially under the depression of crude oil market. The proposed workflow has been successfully implemented to optimize a water alternative CO 2 (CO 2 -WAG) injection project in a depleted oil sand in the US. The optimization results yield an incremental oil production of 15.8 MM STB and 1.37 MM metric tons of CO 2 storage in a 20-year development strategy, with the highest project NPV to be 205.6 MM US dollars.

03 NATURAL GAS↗

Prediction and uncertainty quantification of shale well performance using multifidelity Monte Carlo

Uncertainty quantification is an integral component of reservoir management, especially considering the inherent uncertainty in subsurface systems. While a standard practice to estimate the uncertainty, Monte Carlo (MC) simulation is computationally intense when the sampling population comprises high-fidelity simulations. Alternatively, the Multi-fidelity Monte Carlo (MFMC) simulation overcomes this computational intensity by integrating low- and high-fidelity simulations. Our goal is to minimize the number of expensive high-fidelity simulations while maintaining accuracy and using numerous fast and cheap low-fidelity simulations to efficiently sample to input parameter space of interest. We selected gas production from unconventional wells to demonstrate the potential speedups and accuracy of the MFMC approach. The model fidelity usually determines the trade-off between accuracy and efficiency. While the high-fidelity model is more accurate, the low-fidelity model is more efficient. Our high-fidelity simulation includes reservoir simulations of a hydraulically fractured well. On the other hand, our low-fidelity model comprises the parallel-plate flow model. We used differential programming to efficiently solve the 1D flow model, where automatic differentiation is used to efficiently compute the gradients. We matched the production profile of high-fidelity simulations with our low-fidelity simulations. Then, we used a support vector regression to map the high- and low-fidelity input parameters. The mapping function is essential to tune the low-dimensional parameter space of the low-fidelity model to the high-dimensional parameter space of the high-fidelity model. We found that we can use a combination of 9 high fidelity and 10,000 low fidelity simulations to efficiently and accurately simulate pressure management. This method is at least two orders of magnitude faster than only using high-fidelity simulations. Finally, from a broader perspective, MFMC could efficiently estimate the uncertainty of various systems and models, integrating low- and high-fidelity models.

04 OIL SHALES AND TAR SANDS↗

Using machine learning to predict the correlation of spectra using SDSS magnitudes as an improvement to the Locus Algorithm

The Locus Algorithm is a new technique to improve the quality of differential photometry by optimising the choices of reference stars. At the heart of this algorithm is a routine to assess how good each potential reference star is by comparing its SDSS magnitude values to those of the target star. In this way, the difference in wavelength-dependent effects of the Earth’s atmospheric scattering between target and reference can be minimised. This paper sets out a new way to estimate the quality of each reference star using machine learning. A random subset of stars from SDSS with spectra was chosen. For each one, a suitable reference star, also with a spectrum, was chosen. The correlation between the two spectra in the SDSS r band (between 550 nm and 700 nm) was taken to be the gold-standard measure of how well they match up for differential photometry. The five SDSS magnitude values for each of these stars were used as predictors. A number of supervised machine learning models were constructed on a training set of the stars and were each evaluated on a testing set. The model using Support Vector Regression had the best performance of these models. It was then tested on a final, hold-out, validation set of stars to get an unbiased measure of its performance. With an R 2 of 0.62, the SVR model presents enhanced performance for the Locus Algorithm technique.

79 ASTRONOMY AND ASTROPHYSICS↗