Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “variable importance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modeling freight mode choice using machine learning classifiers: a comparative study using Commodity Flow Survey (CFS) data

This study explores the usefulness of machine learning classifiers for modeling freight mode choice. We investigate eight commonly used machine learning classifiers, namely Naïve Bayes, Support Vector Machine, Artificial Neural Network, K-Nearest Neighbors, Classification and Regression Tree, Random Forest, Boosting and Bagging, along with the classical Multinomial Logit model. US 2012 Commodity Flow Survey data are used as the primary data source; we augment it with spatial attributes from secondary data sources. The performance of the classifiers is compared based on prediction accuracy results. The current research also examines the role of sample size and training-testing data split ratios on the predictive ability of the various approaches. In addition, the importance of variables is estimated to determine how the variables influence freight mode choice. The results show that the tree-based ensemble classifiers perform the best. Specifically, Random Forest produces the most accurate predictions, closely followed by Boosting and Bagging. With regard to variable importance, shipment characteristics, such as shipment distance, industry classification of the shipper and shipment size, are the most significant factors for freight mode choice decisions.

42 ENGINEERING↗

An interpretable machine learning model for advancing terrestrial ecosystem predictions

We apply an interpretable Long Short-Term Memory (iLSTM) network for land-atmosphere carbon flux predictions based on time series observations of seven environmental variables. iLSTM enables interpretability of variable importance and variable-wise temporal importance to the prediction of targets by exploring internal network structures. The application results indicate that iLSTM not only improves prediction performance by capturing different dynamics of individual variables, but also reasonably interprets the different contribution of each variable to the target and its different temporal relevance to the target. This variable and temporal importance interpretation of iLSTM advances terrestrial ecosystem model development as well as our predictive understanding of the system.

Lu, Dan↗

Investigation of hydrometeorological influences on reservoir releases using explainable machine learning methods

Long short-term memory (LSTM) networks have demonstrated successful applications in accurately and efficiently predicting reservoir releases from hydrometeorological drivers including reservoir storage, inflow, precipitation, and temperature. However, due to its black-box nature and lack of process-based implementation, we are unsure whether LSTM makes good predictions for the right reasons. In this work, we use an explainable machine learning (ML) method, called SHapley Additive exPlanations (SHAP), to evaluate the variable importance and variable-wise temporal importance in the LSTM model prediction. In application to 30 reservoirs over the Upper Colorado River Basin, United States, we show that LSTM can accurately predict the reservoir releases with NSE ≥ 0.69 for all the considered reservoirs despite of their diverse storage sizes, functionality, elevations, etc. Additionally, SHAP indicates that storage and inflow are more influential than precipitation and temperature. Moreover, the storage and inflow show a relatively long-term influence on the release up to 7 days and this influence decreases as the lag time increases for most reservoirs. These findings from SHAP are consistent with our physical understanding. However, in a few reservoirs, SHAP gives some temporal importances that are difficult to interpret from a hydrological point of view, probably because of its ignorance of the variable interactions. SHAP is a useful tool for black-box ML model explanations, but the hydrological processes inferred from its results should be interpreted cautiously. More investigations of SHAP and its applications in hydrological modeling is needed and will be pursued in our future study.

54 ENVIRONMENTAL SCIENCES↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗

Implementation of a feature selection algorithm in FARM to identify important state variables and time-invariant matrices

The FARM (Feasible Actuator Range Modifier) software module is a component of the RAVEN-based FORCE framework for analysis of Integrated Energy Systems (IES). FARM aids the HERON software module in the evaluation of the optimal dispatch for the different IES components. Set-point trajectories are required to meet limits on both production variables (i.e., the variables to be optimized such as the electrical power, the hydrogen production rate, etc.) and process variables tied to the service life of equipment (e.g., steam flowrate, vessel pressure, turbine firing temperature, etc.). To evaluate the feasibility of HERON generated set-points and to do so in an acceptable time, FARM employs reduced order models to represent the dynamic behavior of the systems to be dispatched. These surrogate models take the form of a linear dynamic system with sets of Linear Parameter Varying (LPV) matrices that are mapped to the system operating space. These matrices are derived from the trajectories of system state variables and system output variables during transients. The accuracy of LPV matrices depends on the selection of state variables. In previous reports, state variables were selected by adopting a complicated workflow requiring multiple software licenses and an advanced level of user expertise. In this report, a new workflow that automates the state variable selection process is presented. It significantly reduces the frequency of user interventions and does not require multiple software licenses. Each module in the new workflow is described in detail, and the input / output examples in each step of the workflow are provided. It was demonstrated that this workflow can greatly reduce the complexity of the state variable selection process, and that the updated FARM-Gamma and FARM-Delta validators can benefit from this workflow when solving the power dispatch problem of a representative IES test case. Finally, some code improvements that can further enhance the efficiency are suggested.

42 ENGINEERING↗

Machine Learning in Infectious Disease for Risk Factor Identification and Hypothesis Generation: Proof of Concept Using Invasive Candidiasis

Machine learning (ML) models can handle large data sets without assuming underlying relationships and can be useful for evaluating disease characteristics, yet they are more commonly used for predicting individual disease risk than for identifying factors at the population level. We offer a proof of concept applying random forest (RF) algorithms to Candida-positive hospital encounters in an electronic health record database of patients in the United States. Candida-positive encounters were extracted from the Cerner HealthFacts database; invasive infections were laboratory-positive sterile site Candida infections. Features included demographics, admission source, care setting, physician specialty, diagnostic and procedure codes, and medications received before the first positive Candida culture. We used RF to assess risk factors for 3 outcomes: any invasive candidiasis (IC) vs non-IC, within-species IC vs non-IC (eg, invasive C. glabrata vs noninvasive C. glabrata), and between-species IC (eg, invasive C. glabrata vs all other IC). Fourteen of 169 (8%) variables were consistently identified as important features in the ML models. When evaluating within-species IC, for example, invasive C. glabrata vs non-invasive C. glabrata, we identified known features like central venous catheters, intensive care unit stay, and gastrointestinal operations. In contrast, important variables for invasive C. glabrata vs all other IC included renal disease and medications like diabetes therapeutics, cholesterol medications, and antiarrhythmics. Known and novel risk factors for IC were identified using ML, demonstrating the hypothesis-generating utility of this approach for infectious disease conditions about which less is known, specifically at the species level or for rarer diseases.

60 APPLIED LIFE SCIENCES↗

Identifying Entangled Physics Relationships Through Sparse Matrix Decomposition to Inform Plasma Fusion Design

We report a sustainable burn platform through inertial confinement fusion (ICF) has been an ongoing challenge for over 50 years. Mitigating engineering limitations and improving the current design involves an understanding of the complex coupling of physical processes. While sophisticated simulation codes are used to model ICF implosions, these tools contain necessary numerical approximation but miss physical processes that limit predictive capability. Identification of relationships between controllable design inputs to ICF experiments and measurable outcomes (e.g., neutron yield, neutron velocity, areal density) from performed experiments can help guide the future design of experiments and development of simulation codes, to potentially improve the accuracy of the computational models used to simulate ICF experiments. We use sparse matrix decomposition methods to identify clusters of a few related design variables. Sparse principal component analysis (SPCA) identifies groupings that are related to the physical origin of the variables (laser, hohlraum, and capsule). A variable importance analysis finds that in addition to variables highly correlated with neutron yield, such as picket power and laser energy, variables that represent a dramatic change of the ICF design, such as number of pulse steps, are also very important. The obtained sparse components are then used to train a random forest (RF) regression surrogate for predicting total yield. The RF performance on the training and testing data compares with the performance of the RF trained using all the design variables considered. This work is intended to inform design changes in future ICF experiments by augmenting the expert intuition and simulation results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data‐driven identification of environmental variables influencing phenotypic plasticity to facilitate breeding for future climates

Summary Phenotypic plasticity describes a genotype's ability to produce different phenotypes in response to different environments. Breeding crops that exhibit appropriate levels of plasticity for future climates will be crucial to meeting global demand, but knowledge of the critical environmental factors is limited to a handful of well‐studied major crops. Using 727 maize ( Zea mays L.) hybrids phenotyped for grain yield in 45 environments, we investigated the ability of a genetic algorithm and two other methods to identify environmental determinants of grain yield from a large set of candidate environmental variables constructed using minimal assumptions. The genetic algorithm identified pre‐ and postanthesis maximum temperature, mid‐season solar radiation, and whole season net evapotranspiration as the four most important variables from a candidate set of 9150. Importantly, these four variables are supported by previous literature. After calculating reaction norms for each environmental variable, candidate genes were identified and gene annotations investigated to demonstrate how this method can generate insights into phenotypic plasticity. The genetic algorithm successfully identified known environmental determinants of hybrid maize grain yield. This demonstrates that the methodology could be applied to other less well‐studied phenotypes and crops to improve understanding of phenotypic plasticity and facilitate breeding crops for future climates.

Kusmec, Aaron↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗

Spatial Microbial Respiration Variations in the Hyporheic Zones Within the Columbia River Basin

Abstract While the hyporheic zone (HZ) accounts for a significant portion of whole stream CO 2 concentrations, HZ respiration modeling studies are lacking in quantifying their contributions to the total CO 2 at large watershed/basin scales. Quantifying the contribution of anaerobic respiration is also underappreciated. This study used a carbon‐nitrogen‐coupled river corridor model to quantify HZ aerobic and anaerobic respiration and determined the key factors controlling their spatial variability within the Columbia River Basin (CRB). The modeled respiration patterns showed high spatial variability. Among the nine sub‐basins composing the CRB, the Lower Columbia and the Willamette, which receive higher precipitation, had higher respiration. Medium‐sized rivers (fourth to sixth orders) produced the highest aerobic and anaerobic respiration among reaches of different sizes. At the basin scale, aerobic respiration is dominant, representing approximately 98.7% of the total respiration across the CRB. While most of the reaches were dominant with aerobic respiration, reaches in agricultural land showed a relatively higher anaerobic respiration (18%) ratio. A variable importance analysis showed that hyporheic exchange flux controlled most of the spatial variability of HZ respiration, dominating over other physical variables such as residence time, stream dissolved organic carbon (DOC), nitrate, and dissolved oxygen (DO). The influence of substrate concentration (DOC and DO) is larger in modeling anaerobic respiration than aerobic respiration. Future efforts will focus on improving the estimation of the HZ exchange flux and the implementation of spatially explicit parameterizations for the reactions of interest to reduce model uncertainty.

54 ENVIRONMENTAL SCIENCES↗

Semi-supervised Bayesian Low-shot Learning

Deep neural networks (NNs) typically outperform traditional machine learning (ML) approaches for complicated, non-linear tasks. It is expected that deep learning (DL) should offer superior performance for the important non-proliferation task of predicting explosive device configuration based upon observed optical signature, a task which human experts struggle with. However, supervised machine learning is difficult to apply in this mission space because most recorded signatures are not associated with the corresponding device description, or “truth labels.” This is challenging for NNs, which traditionally require many samples for strong performance. Semi-supervised learning (SSL), low-shot learning (LSL), and uncertainty quantification (UQ) for NNs are emerging approaches that could bridge the mission gaps of few labels and rare samples of importance. NN explainability techniques are important in gaining insight into the inferential feature importance of such a complex model. In this work, SSL, LSL, and UQ are merged into a single framework, a significant technical hurdle not previously demonstrated. Exponential Average Adversarial Training (EAAT) and Pairwise Neural Networks (PNNs) are chosen as the SSL and LSL methods of choice. Permutation feature importance (PFI) for functional data is used to provide explainability via the Variable importance Explainable Elastic Shape Analysis (VEESA) pipeline. A variety of uncertainty quantification approaches are explored: Bayesian Neural Networks (BNNs), ensemble methods, concrete dropout, and evidential deep learning. Two final approaches, one utilizing ensemble methods and one utilizing evidential learning, are constructed and compared using a well-quantified synthetic 2D dataset along with the DIRSIG Megascene.

97 MATHEMATICS AND COMPUTING↗

Modeling the distribution of the endangered Jemez Mountains salamander (Plethodon neomexicanus) in relation to geology, topography, and climate

The Jemez Mountains salamander (Plethodon neomexicanus; hereafter JMS) is an endangered salamander restricted to the Jemez Mountains in north-central New Mexico, United States. This strictly terrestrial and lungless species requires moist surface conditions for activities such as mating and foraging. Threats to its current habitat include fire suppression and ensuing severe fires, changes in forest composition, habitat fragmentation, and climate change. Forest composition changes resulting from reduced fire frequency and increased tree density suggest that its current aboveground habitat does not mirror its historically successful habitat regime. However, because of its limited habitat area and underground behavior, we hypothesized that geology and topography might play a significant role in the current distribution of the salamander. We modeled the distribution of the JMS using a machine learning algorithm to assess how geology, topography, and climate variables influence its distribution. The best habitat suitability model indicates that geology type and maximum winter temperature (November to March) were most important in predicting the distribution of the salamander (23.5% and 50.3% permutation importance, respectively). Minimum winter temperature was also an important variable (21.4%), suggesting this also plays a role in salamander habitat. Our habitat suitability map reveals low uncertainty in model predictions, and we found slight discrepancies between the designated critical habitat and the most suitable areas for the JMS. Because geological features are important to its distribution, we recommend that geological and topographical data are considered, both during survey design and in the description of localities of JMS records once detected.

59 BASIC BIOLOGICAL SCIENCES↗

How Well do Earth System Models Capture Apparent Relationships Between Phytoplankton Biomass and Environmental Variables?

Abstract As phytoplankton form the base of the marine food web, understanding the controls on their abundance is fundamental to understanding marine ecology and its sensitivity to global climate change. While many Earth System Models (ESMs) predict phytoplankton biomass, it is unclear whether they properly capture the mechanistic relationships that control this quantity in the real ocean. We used Random Forest analysis to analyze the output of 13 ESMs as well as two observational data sets. The target variable was phytoplankton carbon and the predictors included environmental parameters known to influence phytoplankton, including nutrients, light, mixed layer depth, salinity, temperature, and upwelling. We examined the following: (a) What fractions of variability in ESMs and observations can be linked to the large‐scale environmental variables simulated by ESMs? (b) What are the dominant predictors and relationships affecting phytoplankton biomass? (c) How well do ESMs simulate phytoplankton carbon and do they simulate the relationships we see in observations? About 88%–96% of the variability in observational data sets and greater than 98% in the ESMs was accounted for by environmental variables known to influence phytoplankton biomass. The dominant predictors in the observational data sets were shortwave radiation and dissolved iron, with temperature and ammonium also relatively important. All the ESMs show that shortwave radiation is the most important variable and most of them predict the right sign of sensitivity to most variables. However, the models predict that biomass reaches maximum levels at unrealistically low levels of iron and unrealistically high levels of light.

Environmental Sciences & Ecology↗

Exploring the Effects of Population and Employment Characteristics on Truck Flows: An Analysis of NextGen NHTS Origin-Destination Data

Truck transportation remains the dominant mode of US freight transportation because of its advantages, such as the flexibility of accessing pickup and drop-off points and faster delivery. Because of the massive freight volume transported by trucks, understanding the effects of population and employment characteristics on truck flows is critical for better transportation planning and investment decisions. The US Federal Highway Administration published a truck travel origin-destination data set as part of the Next Generation National Household Travel Survey program. This data set contains the total number of truck trips in 2020 within and between 583 predefined zones encompassing metropolitan and nonmetropolitan statistical areas within each state and Washington, DC. In this study, origin-destination-level truck trip flow data was augmented to include zone-level population and employment characteristics from the US Census Bureau. Census population and County Business Patterns data were included. The final data set was used to train a machine learning algorithm-based model, Extreme Gradient Boosting (XGBoost), where the target variable is the number of total truck trips. Shapley Additive ExPlanation (SHAP) was adopted to explain the model results. Results showed that the distance between the zones was the most important variable and had a nonlinear relationship with truck flows.

Uddin, Majbah↗

Machine-learning-based investigation of the variables affecting summertime lightning occurrence over the Southern Great Plains

Lightning is affected by many factors, many of which are not routinely measured, well understood, or accounted for in physical models. Several commonly used machine learning (ML) models have been applied to analyze the relationship between Atmospheric Radiation Measurement (ARM) data and lightning data from the Earth Networks Total Lightning Network (ENTLN) in order to identify important variables affecting lightning occurrence in the vicinity of the Southern Great Plains (SGP) ARM site during the summer months (June, July, August and September) of 2012 to 2020. Testing various ML models, we found that the random forest model is the best predictor among common classifiers. When convective clouds were detected, it predicts lightning occurrence with an accuracy of 76.9 % and an area under the curve (AUC) of 0.850. Using this model, we further ranked the variables in terms of their effectiveness in nowcasting lightning and identified geometric cloud thickness, rain rate and convective available potential energy (CAPE) as the most effective predictors. The contrast in meteorological variables between no-lightning and frequent-lightning periods was examined for hours with CAPE values conducive to thunderstorm formation. Besides the variables considered for the ML models, surface variables and mid-altitude variables (e.g., equivalent potential temperature and minimum equivalent potential temperature, respectively) have statistically significant contrasts between no-lightning and frequent-lightning hours. For example, the minimum equivalent potential temperature from 700 to 500 hPa is significantly lower during frequent-lightning hours compared with no-lightning hours. Finally, a notable positive relationship between the intracloud (IC) flash fraction and the square root of CAPE ($\sqrt{CAPE}$) was found, suggesting that stronger updrafts increase the height of the electrification zone, resulting in fewer flashes reaching the surface and consequently a greater IC flash fraction.

54 ENVIRONMENTAL SCIENCES↗

A machine learning approach for identifying variables associated with risk of developing neutralizing antidrug antibodies to factor VIII

A key unmet need in the management of hemophilia A (HA) is the lack of clinically validated markers that are associated with the development of neutralizing antibodies to Factor VIII (FVIII) (commonly referred to as inhibitors). This study aimed to identify relevant biomarkers for FVIII inhibition using Machine Learning (ML) and Explainable AI (XAI) using the My Life Our Future (MLOF) research repository. The dataset includes biologically relevant variables such as age, race, sex, ethnicity, and the variants in the F8 gene. In addition, we previously carried out Human Leukocyte Antigen Class II (HLA-II) typing on samples obtained from the MLOF repository. Using this information, we derived other patient-specific biologically and genetically important variables. These included identifying the number of foreign FVIII derived peptides, based on the alignment of the endogenous FVIII and infused drug sequences, and the foreign-peptide HLA-II molecule binding affinity calculated using NetMHCIIpan. The data were processed and trained with multiple ML classification models to identify the top performing models. The top performing model was then chosen to apply XAI via SHAP, (SHapley Additive exPlanations) to identify the variables critical for the prediction of FVIII inhibitor development in a hemophilia A patient. Using XAI we provide a robust and ranked identification of variables that could be predictive for developing inhibitors to FVIII drugs in hemophilia A patients. These variables could be validated as biomarkers and used in making clinical decisions and during drug development. The top five variables for predicting inhibitor development based on SHAP values are: (i) the baseline activity of the FVIII protein, (ii) mean affinity of all foreign peptides for HLA DRB 3, 4, & 5 alleles, (iii) mean affinity of all foreign peptides for HLA DRB1 alleles), (iv) the minimum affinity among all foreign peptides for HLA DRB1 alleles, and (v) F8 mutation type.

60 APPLIED LIFE SCIENCES↗

Simple Statistical Models for Predicting Overpressure Due to CO2 and Low-Salinity Waste-Fluid Injection into Deep Saline Formations

Deep saline aquifers have been used for waste-fluid disposal for decades and are the proposed targets for large-scale CO2 storage to mitigate CO2 concentration in the atmosphere. Due to relatively limited experience with CO2 injection in deep saline formations and given that the injection targets for CO2 sometimes are the same as waste-fluid disposal formations, it could be beneficial to model and compare both practices and learn from the waste-fluid disposal industry. In this paper, we model CO2 injection in the Patterson Field, which has been proposed as a site for storage of 50 Mt of industrial CO2 over 25 years. We propose general models that quickly screen the reservoir properties and calculate pressure changes near and far from the injection wellbore, accounting for variable reservoir properties. The reservoir properties we investigated were rock compressibility, injection rate, vertical-to-horizontal permeability ratio, average reservoir permeability and porosity, reservoir temperature and pressure, and the injectant total dissolved solids (TDS) in cases of waste-fluid injection. We used experimental design to select and perform simulation runs, performed a sensitivity analysis to identify the important variables on pressure build-up, and then fit a regression model to the simulation runs to obtain simple proxy models for changes in average reservoir pressure and bottomhole pressure. The CO2 injection created more pressure compared to saline waste-fluids, when similar mass was injected. However, we found a more significant pressure buildup at the caprock-reservoir interface and lower pressure buildup at the bottom of the reservoir when injecting CO2 compared with waste-fluid injection.

Ansari, Esmail↗