Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Support Vector Machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

J Lemery↗

Earth Independent Medical Operations (EIMO) Datascope: Challenges and Potential Solutions

Data flows and storage/retrieval capacity are severely constrained during missions in space and challenges will become even greater during exploration class missions. There is a need for an artificial intelligence (AI)-based clinical decision support system (CDSS) to monitor and analyze data to provide real-time consultative support for crew medical officer (CMO) decision-making. EIMO is defined as the gradual transition of medical care and decision making from terrestrial to space-based assets, enabling support of astronaut health and performance and reducing overall mission risk. While a hallmark of this paradigm shift from low-earth orbit is that on-board care will increasingly become the responsibility of the astronauts for primary management and decision making, terrestrial assets will continue to be paramount in pre-mission screening and planning, as well as prevention, health maintenance and long-term care contingencies. New capabilities and systems that enable progressively more robust and resilient systems and crews will be necessary to reduce risk and increase probability of deep space exploration mission success. An aspiration for EIMO is to develop AI-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A “system of systems” approach is envisioned whereby EIMO will deploy AI-supported natural language processing and machine learning (ML) techniques to utilize embedded reference databases and real-time data streams [input vectors] from multiple data sources. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will feature mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats. Large amounts and variable sources of data can be leveraged to diagnose, inform treatment strategies, and potentially predict medical events and performance decrements. Inclusion of advanced training tools using extended reality will enable increasingly autonomous medical care to aid a CMO when ground support is unavailable or time-delayed beyond required action window, e.g., emergent medical situations. EIMO CDSS would require very large datasets to train pre-flight and significant amounts of data are needed to support ML via in-flight CDSS operations. An additional challenge will be to find sufficient data to train a model relevant to astronaut demographics. The rapid, accelerating evolution of this field creates a propitious solution space to leverage multi-modal AI through public-private partnership(s). The status of multi-modal AI systems today would preclude their use for long duration missions as they remain unreliable and are subject to “digital hallucinations” and other errors that could pose operational risk. A federated labs structure is being considered to test and optimize data flow from the multiple input vectors leading to field testing in suitable ground/flight analogs. Critical to the success of an EIMO CDSS will be integration and interoperability and success will be defined by a system that can serve as an in-flight medical consult for the CMO providing critical support during medical contingencies. Benefits to terrestrial medicine may be significant as an outflow of the EIMO medical system, particularly for remote areas and communities lacking significant infrastructure, personnel and resources.

Medical Operations↗

Large-Scale Numerical Simulations of Human Motion

This paper examines the feasibility of using massively-parallel and vector-processing supercomputers to solve large-scale optimal control problems for human movement. Specifically, we compare the computational expense of determining the optimal controls for the single support phase of walking using a conventional serial machine (a Silicon Graphics Personal Iris 4D25 workstation), a MIMD parallel machine (an Intel iPSC/860 comprising 128 processors), and a parallel-vector-processing machine (a Cray Y-MP 8/864). With the human body modeled as a 14 degree-of-freedom linkage actuated by 46 musculotendinous units, computation of the optimal controls for walking could take up to 3 months of CPU time on the Iris. Both the Cray Y-MP and the Intel iPSC/860 are able to reduce this time to practical levels. The optimal control solution for walking can be found with about 77 hours of CPU time on the Cray, and with about 88 hours of CPU time on the Intel. Although the overall speeds of the Cray and the Intel were found to be similar, the unique capabilities of each machine are best suited to different parts of the optimal control algorithm used. The Intel performed best in the calculation of the derivatives of the performance criterion and the constraints. In contrast, the Cray performed best during parameter optimization of the controls. These results suggest that the ideal computer architecture for solving very large-scale optimal control problems is a hybrid system in which a vector-processing machine is integrated into the communication network of a MIMD parallel machine.

Anderson, Frank C.↗

J-PLUS: Support vector regression to measure stellar parameters

Stellar parameters are among the most important characteristics in studies of stars which, in traditional methods, are based on atmosphere models. However, time, cost, and brightness limits restrain the efficiency of spectral observations. The Javalambre Photometric Local Universe Survey (J-PLUS) is an observational campaign that aims to obtain photometry in 12 bands. Owing to its characteristics, J-PLUS data have become a valuable resource for studies of stars. Machine learning provides powerful tools for efficiently analyzing large data sets, such as the one from J-PLUS, and enables us to expand the research domain to stellar parameters. The main goal of this study is to construct a support vector regression (SVR) algorithm to estimate stellar parameters of the stars in the first data release of the J-PLUS observational campaign. The training data for the parameter's regressions are featured with 12-waveband photometry from J-PLUS and are crossidentified with spectrum-based catalogs. These catalogs are from the Large Sky Area Multi-Object Fiber Spectroscopic Telescope, the Apache Point Observatory Galactic Evolution Experiment, and the Sloan Extension for Galactic Understanding and Exploration. We then label them with the stellar effective temperature, the surface gravity, and the metallicity. Ten percent of the sample is held out to apply a blind test. We develop a new method, a multi-model approach, in order to fully take into account, the uncertainties of both the magnitudes and the stellar parameters. The method utilizes more than 200 models to apply the uncertainty analysis. We present a catalog of 2 493 424 stars with the root mean square error of 160 K in the effective temperature regression, 0.35 in the surface gravity regression, and 0.25 in the metallicity regression. We also discuss the advantages of this multi-model approach and compare it to other machine-learning methods.

79 ASTRONOMY AND ASTROPHYSICS↗

Vehicle Position Detection Based on Machine Learning Algorithms in Dynamic Wireless Charging

Dynamic wireless charging (DWC) has emerged as a viable approach to mitigate range anxiety by ensuring continuous and uninterrupted charging for electric vehicles in motion. DWC systems rely on the length of the transmitter, which can be categorized into long-track transmitters and segmented coil arrays. The segmented coil array, favored for its heightened efficiency and reduced electromagnetic interference, stands out as the preferred option. However, in such DWC systems, the need arises to detect the vehicle’s position, specifically to activate the transmitter coils aligned with the receiver pad and de-energize uncoupled transmitter coils. This paper introduces various machine learning algorithms for precise vehicle position determination, accommodating diverse ground clearances of electric vehicles and various speeds. Through testing eight different machine learning algorithms and comparing the results, the random forest algorithm emerged as superior, displaying the lowest error in predicting the actual position.

47 OTHER INSTRUMENTATION↗

Using machine learning to predict the correlation of spectra using SDSS magnitudes as an improvement to the Locus Algorithm

The Locus Algorithm is a new technique to improve the quality of differential photometry by optimising the choices of reference stars. At the heart of this algorithm is a routine to assess how good each potential reference star is by comparing its SDSS magnitude values to those of the target star. In this way, the difference in wavelength-dependent effects of the Earth’s atmospheric scattering between target and reference can be minimised. This paper sets out a new way to estimate the quality of each reference star using machine learning. A random subset of stars from SDSS with spectra was chosen. For each one, a suitable reference star, also with a spectrum, was chosen. The correlation between the two spectra in the SDSS r band (between 550 nm and 700 nm) was taken to be the gold-standard measure of how well they match up for differential photometry. The five SDSS magnitude values for each of these stars were used as predictors. A number of supervised machine learning models were constructed on a training set of the stars and were each evaluated on a testing set. The model using Support Vector Regression had the best performance of these models. It was then tested on a final, hold-out, validation set of stars to get an unbiased measure of its performance. With an R 2 of 0.62, the SVR model presents enhanced performance for the Locus Algorithm technique.

79 ASTRONOMY AND ASTROPHYSICS↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

NEML2: An efficient and modular multiphysics constitutive modeling library for hybrid computing environments

This paper presents NEML2, an open-source, high-performance library developed for constitutive material modeling, designed to support the flexible and modular development of models for complex material behavior. Building on the foundational structure of its predecessor, NEML, the NEML2 library introduces significant improvements, including enhanced vectorization, automatic differentiation, and seamless integration with PyTorch, facilitating the application of machine learning techniques in material simulations. NEML2 provides a C++ backend with Python bindings, enabling users to create custom material models that can be executed efficiently on both CPU and GPU platforms. The library also supports coupling with Multiphysics simulation frameworks like MOOSE, making it suitable for realistic simulations involving coupled physical processes. Rigorous quality assurance through unit and regression testing ensures the reliability of results, while the extensible, user-friendly design encourages collaboration and reproducibility across the scientific community. This paper provides an overview of NEML2’s architecture, core features, and applications, highlighting its impact on accelerating material qualification and advancing computational methods in materials science.

GPU↗

Shorter function summaries for finite state machine-based high consequence systems using logic synthesis and tautologies (Final Report LDRD 24-1302)

Computer programs are often viewed as collections of functions – each function has parameters (inputs) and computes a return value, and each has potential side effects that modify program state (outputs). In this research, a Sandia symbolic execution tool designed to support “human-in-the-loop” analysis was modified to automatically create “function summaries,” and a new tool, “diaboolical,” was created to support enhancing readability of the summary using a novel approach to bit-vector simplification that leverages logic synthesis and tautologies. For this effort, students at Auburn University created several finite state machines (FSMs) to serve as exemplars for high-consequence systems. Function summaries for each of the machines were obtained, and then portions of the summaries were simplified using both diaboolical and the simplification procedure of a popular SMT solver. A comparison of the results shows that diaboolical can often produce smaller function summaries, with expression length improvements over the unsimplified function summaries ranging from 0% to 90% for diaboolical and 0% to 65% for the SMT solver, though diaboolical had a significantly greater cost in time. Diaboolical was evaluated against a collection of “arbitrary” C-code as well as FSM exemplars, and for both datasets it achieved an approximately 10% improvement in expression length compared to simplifications that could be obtained using existing techniques. Function summaries can assist assurance efforts that evaluate existing systems and their executable code. A smaller function summary is likely easier for humans to understand and could thus increase the ability and efficacy of assurance practices centered around the analysis of executable artifacts.

97 MATHEMATICS AND COMPUTING↗

Coupling High-Throughput Experiments and Regression Algorithms to Optimize PGM-Free ORR Electrocatalyst Synthesis

Over the past decades, significant improvement has been achieved in the performance of platinum group metal-free (PGM-free) materials as an alternative to Pt-based electrocatalysts for oxygen reduction reaction (ORR). However, further progress in ORR activity requires evaluation of precursors and synthesis approaches. In response to this challenge, we generated a first of its kind experimental data set of 36 samples using high-throughput synthesis and activity measurements. Several control parameters (e.g., Fe precursor identity, the precursor content, and pyrolysis temperature) were varied. We then developed several state-of-the-art machine learning (ML) based regression models to predict ORR activity, dependent on selected synthesis variables. Through an iterative algorithm, higher prediction accuracy (smaller root-mean-square error) was achieved. We identified that gradient boosting regression (GBR) and support vector regression (SVR), among several methods, work best for this data set. Aided by our ML-based surrogate models, we decided to alter catalyst synthesis conditions, which resulted in a 36% increase in measured ORR activity in comparison to the maximum ORR mass activity value of 21.9 A/g catalyst in the original data set. Overall, this combined experiment and machine learning approach represents a promising path forward toward developing highly efficient next-generation ORR electrocatalysts and, more generally, functional materials.

25 ENERGY STORAGE↗

Learning and discovering multiple solutions using physics-informed neural networks with random initialization and deep ensemble

In this work we explore the capability of physics-informed neural networks (PINNs) to discover multiple solutions. Many real-world phenomena governed by nonlinear differential equations (DEs), such as fluid flow, exhibit multiple solutions under the same conditions, yet capturing this solution multiplicity remains a significant challenge. A key difficulty lies in providing appropriate initial conditions or guesses, as widely used time-marching schemes and Newton’s method are highly sensitive to these choices when solving complex computational problems. While machine learning models, particularly PINNs, have shown promise in solving DEs, their ability to capture multiple solutions remains underexplored. In this work, we propose a simple and practical approach using PINNs to learn and discover multiple solutions. We first demonstrate that PINNs, when combined with random initialization and deep ensemble method—originally developed for uncertainty quantification—can effectively uncover multiple solutions to nonlinear ordinary and partial DEs. Although training large ensembles of PINNs may appear computationally demanding, this can be done efficiently using vectorization techniques supported by modern deep learning frameworks, allowing many networks to be trained simultaneously. Our approach highlights the critical role of initialization in shaping solution diversity, addressing an often-overlooked aspect of machine learning for scientific computing. Furthermore, we propose utilizing PINN-generated solutions as initial conditions or initial guesses for conventional numerical solvers to enhance accuracy and efficiency in capturing multiple solutions. Extensive numerical experiments, including the Allen–Cahn equation and cavity flow, where our approach successfully identifies both stable and unstable solutions, validate the effectiveness of our method. These findings establish a general and efficient framework for addressing solution multiplicity in nonlinear DEs.

97 MATHEMATICS AND COMPUTING↗

Research on computer systems benchmarking

This grant addresses the topic of research on computer systems benchmarking and is more generally concerned with performance issues in computer systems. This report reviews work in those areas during the period of NASA support under this grant. The bulk of the work performed concerned benchmarking and analysis of CPUs, compilers, caches, and benchmark programs. The first part of this work concerned the issue of benchmark performance prediction. A new approach to benchmarking and machine characterization was reported, using a machine characterizer that measures the performance of a given system in terms of a Fortran abstract machine. Another report focused on analyzing compiler performance. The performance impact of optimization in the context of our methodology for CPU performance characterization was based on the abstract machine model. Benchmark programs are analyzed in another paper. A machine-independent model of program execution was developed to characterize both machine performance and program execution. By merging these machine and program characterizations, execution time can be estimated for arbitrary machine/program combinations. The work was continued into the domain of parallel and vector machines, including the issue of caches in vector processors and multiprocessors. All of the afore-mentioned accomplishments are more specifically summarized in this report, as well as those smaller in magnitude supported by this grant.

Smith, Alan Jay↗

Projecting the Thermal Response in a HTGR-Type System during Conduction Cooldown Using Graph-Laplacian Based Machine Learning

Accurate prediction of an off-normal event in a nuclear reactor is dependent upon the availability of sensory data, reactor core physical condition, and understanding of the underlying phenomenon. This work presents a method to project the data from some discrete sensory locations to the overall reactor domain during conduction cooldown scenarios similar to High Temperature Gas-cooled Reactors (HTGRs). The existing models for conductive cooldown in a heterogeneous multi-body system, such as an assembly of prismatic blocks or pebble beds relies on knowledge of the thermal contact conductance, requiring significant knowledge of local thermal contacts and heat transport possibilities across those contacts. With a priori knowledge of bulk geometry features and some discrete sensors, a machine learning approach was devised. The presented work uses an experimental facility to mimic conduction cooldown with an assembly of 68 cylindrical rods initially heated to 1200 K. High-fidelity temperature data were collected using an infrared (IR) camera to provide training data to the model and validate the predicted temperature data. The machine learning approach used here first converts the macroscopic bulk geometry information into Graph-Laplacian, and then uses the eigenvectors of the Graph-Laplacian to develop Kernel functions. Support vector regression (SVR) was implemented on the obtained Kernels and used to predict the thermal response in a packed rod assembly during a conduction cooldown experiment. The usage of SVR modeling differs from most models today because of its representation of thermal coupling between rods in the core. When trained with thermographic data, the average normalized error is less than 2% over 400 s, during which temperatures of the assembly have dropped by more than 500 K. The rod temperature prediction performance was significantly better for rods in the interior of the assembly compared to those near the exterior, likely due to the model simplification of the surroundings.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Dataset for 'Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning', Water 2022

This data package presents forcing data, model code, and model output for classical machine learning models that predict monthly stream water temperature as presented in the manuscript ‘Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning’, Water (Weierbach et al., 2022). Specifically, for input forcing datasets we include two files each generated using the BASIN-3D data integration tool (Varadharajan et al., 2022) for stations in the Pacific Northwest and Mid Atlantic Hydrologic regions. Model code (written in python with the use of jupyter notebooks) includes codes for data preprocessing, training Multiple Linear Regression, Support Vector Regression, and Extreme Gradient Boosted Tree models, and additional notebooks for analysis of model output. We include specific model output files which represent modeling configurations presented in the manuscript also presented in an hdf5 format. Together, these data make up the workflow for predictions across three scenarios (single station, regional, and predictions in unmonitored basins) presented in the manuscript and allow for reproducibility of modeling procedures.

54 ENVIRONMENTAL SCIENCES↗

Programming the Navier-Stokes computer: An abstract machine model and a visual editor

The Navier-Stokes computer is a parallel computer designed to solve Computational Fluid Dynamics problems. Each processor contains several floating point units which can be configured under program control to implement a vector pipeline with several inputs and outputs. Since the development of an effective compiler for this computer appears to be very difficult, machine level programming seems necessary and support tools for this process have been studied. These support tools are organized into a graphical program editor. A programming process is described by which appropriate computations may be efficiently implemented on the Navier-Stokes computer. The graphical editor would support this programming process, verifying various programmer choices for correctness and deducing values such as pipeline delays and network configurations. Step by step details are provided and demonstrated with two example programs.

Middleton, David↗

A comprehensive techno-eco-assessment of CO 2 enhanced oil recovery projects using a machine-learning assisted workflow

Carbon dioxide enhanced oil recovery (CO 2 -EOR) projects not only extract residual oil but also sequestrate CO 2 in the depleted reservoirs. Here, this study develops a machine-learning-based workflow to co-optimize the hydrocarbon recovery, CO 2 sequestration volume and project net present value (NPV) simultaneously. Considering the trade-off relationships among the objective functions, support vector regression with Gaussian kernel (Gaussian- SVR) proxies are coupled with multi-objective particle swarm optimization (PSO) protocol and generate Pareto optimal solutions. Taking advantage of the high computational efficacy of the proxy model, economic uncertainties introduced by tax credits, capital costs and oil price are investigated by this study. The results indicate that the tax incentive policy (Section 45Q) plays a vital role in enhancing the economic returns of CO 2 -EOR projects, especially under the depression of crude oil market. The proposed workflow has been successfully implemented to optimize a water alternative CO 2 (CO 2 -WAG) injection project in a depleted oil sand in the US. The optimization results yield an incremental oil production of 15.8 MM STB and 1.37 MM metric tons of CO 2 storage in a 20-year development strategy, with the highest project NPV to be 205.6 MM US dollars.

03 NATURAL GAS↗

A Comparison of Machine Learning Methods to Forecast Tropospheric Ozone Levels in Delhi

Ground-level ozone is a pollutant that is harmful to urban populations, particularly in developing countries where it is present in significant quantities. It greatly increases the risk of heart and lung diseases and harms agricultural crops. This study hypothesized that, as a secondary pollutant, ground-level ozone is amenable to 24 h forecasting based on measurements of weather conditions and primary pollutants such as nitrogen oxides and volatile organic compounds. We developed software to analyze hourly records of 12 air pollutants and 5 weather variables over the course of one year in Delhi, India. To determine the best predictive model, eight machine learning algorithms were tuned, trained, tested, and compared using cross-validation with hourly data for a full year. The algorithms, ranked by R2 values, were XGBoost (0.61), Random Forest (0.61), K-Nearest Neighbor Regression (0.55), Support Vector Regression (0.48), Decision Trees (0.43), AdaBoost (0.39), and linear regression (0.39). When trained by separate seasons across five years, the predictive capabilities of all models increased, with a maximum R 2 of 0.75 during winter. Bidirectional Long Short-Term Memory was the least accurate model for annual training, but had some of the best predictions for seasonal training. Out of five air quality index categories, the XGBoost model was able to predict the correct category 24 h in advance 90% of the time when trained with full-year data. Separated by season, winter is considerably more predictable (97.3%), followed by post-monsoon (92.8%), monsoon (90.3%), and summer (88.9%). These results show the importance of training machine learning methods with season-specific data sets and comparing a large number of methods for specific applications.

54 ENVIRONMENTAL SCIENCES↗