Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “variable importance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modeling freight mode choice using machine learning classifiers: a comparative study using Commodity Flow Survey (CFS) data

This study explores the usefulness of machine learning classifiers for modeling freight mode choice. We investigate eight commonly used machine learning classifiers, namely Naïve Bayes, Support Vector Machine, Artificial Neural Network, K-Nearest Neighbors, Classification and Regression Tree, Random Forest, Boosting and Bagging, along with the classical Multinomial Logit model. US 2012 Commodity Flow Survey data are used as the primary data source; we augment it with spatial attributes from secondary data sources. The performance of the classifiers is compared based on prediction accuracy results. The current research also examines the role of sample size and training-testing data split ratios on the predictive ability of the various approaches. In addition, the importance of variables is estimated to determine how the variables influence freight mode choice. The results show that the tree-based ensemble classifiers perform the best. Specifically, Random Forest produces the most accurate predictions, closely followed by Boosting and Bagging. With regard to variable importance, shipment characteristics, such as shipment distance, industry classification of the shipper and shipment size, are the most significant factors for freight mode choice decisions.

42 ENGINEERING↗

An interpretable machine learning model for advancing terrestrial ecosystem predictions

We apply an interpretable Long Short-Term Memory (iLSTM) network for land-atmosphere carbon flux predictions based on time series observations of seven environmental variables. iLSTM enables interpretability of variable importance and variable-wise temporal importance to the prediction of targets by exploring internal network structures. The application results indicate that iLSTM not only improves prediction performance by capturing different dynamics of individual variables, but also reasonably interprets the different contribution of each variable to the target and its different temporal relevance to the target. This variable and temporal importance interpretation of iLSTM advances terrestrial ecosystem model development as well as our predictive understanding of the system.

Lu, Dan↗

Investigation of hydrometeorological influences on reservoir releases using explainable machine learning methods

Long short-term memory (LSTM) networks have demonstrated successful applications in accurately and efficiently predicting reservoir releases from hydrometeorological drivers including reservoir storage, inflow, precipitation, and temperature. However, due to its black-box nature and lack of process-based implementation, we are unsure whether LSTM makes good predictions for the right reasons. In this work, we use an explainable machine learning (ML) method, called SHapley Additive exPlanations (SHAP), to evaluate the variable importance and variable-wise temporal importance in the LSTM model prediction. In application to 30 reservoirs over the Upper Colorado River Basin, United States, we show that LSTM can accurately predict the reservoir releases with NSE ≥ 0.69 for all the considered reservoirs despite of their diverse storage sizes, functionality, elevations, etc. Additionally, SHAP indicates that storage and inflow are more influential than precipitation and temperature. Moreover, the storage and inflow show a relatively long-term influence on the release up to 7 days and this influence decreases as the lag time increases for most reservoirs. These findings from SHAP are consistent with our physical understanding. However, in a few reservoirs, SHAP gives some temporal importances that are difficult to interpret from a hydrological point of view, probably because of its ignorance of the variable interactions. SHAP is a useful tool for black-box ML model explanations, but the hydrological processes inferred from its results should be interpreted cautiously. More investigations of SHAP and its applications in hydrological modeling is needed and will be pursued in our future study.

54 ENVIRONMENTAL SCIENCES↗

Developing a Customizable Composite Drought Index for Pakistan

Pakistan has experienced intense agricultural droughts in recent years, with the southern provinces experiencing the most severe drought conditions. These prolonged dry conditions often result in failed crop production, impacting families and communities. To reduce the impacts a drought may have on a community by identifying dry conditions, drought indices are used. A customizable composite drought index (CDI) is developed to improve the spatial and temporal understanding of historical agricultural droughts that have affected Pakistan. An in depth analysis using 9 input variables will provide information regarding the most important variables for differing locations. This framework can enhance drought monitoring and forecasting systems. The performance of the CDI will be evaluated by using production data from different crops that have significant economic value in Pakistan. These crops include wheat, rice, maize, cotton and barley. The expected relationship between the CDI and crop production would be to see production decrease when the CDI decreases, an indication that crops failed due to drought. The CDI is made more precise by evaluating only agricultural areas using the months of the growing season for each specific crop. This will allow for a custom drought index to be developed based on crop type and geographic location that only uses the most important variables, while still capturing the full extent of historical drought.

Caily Schwartz↗

Vegetation Greening Mitigates the Impacts of Increasing Extreme Rainfall on Runoff Events

Future flood risk assessment has primarily focused on heavy rainfall as the main driver, with the assumption that projected increases in extreme rain events will lead to subsequent flooding. However, the presence of and changes in vegetation have long been known to influence the relationship between rainfall and runoff. Here, we extract historical (1850–1880) and projected (2070–2100) daily extreme rainfall events, the corresponding runoff, and antecedent conditions simulated in a prominent large Earth system model ensemble to examine the shifting extreme rainfall and runoff relationship. Even with widespread projected increases in the magnitude (78% of the land surface) and number (72%) of extreme rainfall events, we find projected declines in event-based runoff ratio (runoff/rainfall) for a majority (57%) of the Earth surface. Runoff ratio declines are linked with decreases in antecedent soil water driven by greater transpiration and canopy evaporation (both linked to vegetation greening) compared to areas with runoff ratio increases. Using a machine learning regression tree approach, we find that changes in canopy evaporation is the most important variable related to changes in antecedent soil water content in areas of decreased runoff ratios (with minimal changes in antecedent rainfall) while antecedent ground evaporation is the most important variable in areas of increased runoff ratios. Our results suggest that simulated interactions between vegetation greening, increasing evaporative demand, and antecedent soil drying are projected to diminish runoff associated with extreme rainfall events, with important implications for society.

soil moisture↗

Rapid solidification of metallic particulates

In order to maximize the heat transfer coefficient the most important variable in rapid solidification is the powder particle size. The finer the particle size, the higher the solidification rate. Efforts to decrease the particle size diameter offer the greatest payoff in attained quench rate. The velocity of the liquid droplet in the atmosphere is the second most important variable. Unfortunately the choices of gas atmospheres are sharply limited both because of conductivity and cost. Nitrogen and argon stand out as the preferred gases, nitrogen where reactions are unimportant and argon where reaction with nitrogen may be important. In gas atomization, helium offers up to an order of magnitude increase in solidification rate over argon and nitrogen. By contrast, atomization in vacuum drops the quench rate several orders of magnitude.

Grant, N. J.↗

Rapid solidification of metallic particulates

In order to maximize the heat transfer coefficient the most important variable in rapid solidification is the powder particle size. The finer the particle size, the higher the solidification rate. Efforts to decrease the particle size diameter offer the greatest payoff in attained quench rate. The velocity of the liquid droplet in the atmosphere is the second most important variable. Unfortunately the choices of gas atmospheres are sharply limited both because of conductivity and cost. Nitrogen and argon stand out as the preferred gases, nitrogen where reactions are unimportant and argon where reaction with nitrogen may be important. In gas atomization, helium offers up to an order of magnitude increase in solidification rate over argon and nitrogen. By contrast, atomization in vacuum drops the quench rate several orders of magnitude. Previously announced in STAR as N82-26433

Grant, N. J.↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗

Implementation of a feature selection algorithm in FARM to identify important state variables and time-invariant matrices

The FARM (Feasible Actuator Range Modifier) software module is a component of the RAVEN-based FORCE framework for analysis of Integrated Energy Systems (IES). FARM aids the HERON software module in the evaluation of the optimal dispatch for the different IES components. Set-point trajectories are required to meet limits on both production variables (i.e., the variables to be optimized such as the electrical power, the hydrogen production rate, etc.) and process variables tied to the service life of equipment (e.g., steam flowrate, vessel pressure, turbine firing temperature, etc.). To evaluate the feasibility of HERON generated set-points and to do so in an acceptable time, FARM employs reduced order models to represent the dynamic behavior of the systems to be dispatched. These surrogate models take the form of a linear dynamic system with sets of Linear Parameter Varying (LPV) matrices that are mapped to the system operating space. These matrices are derived from the trajectories of system state variables and system output variables during transients. The accuracy of LPV matrices depends on the selection of state variables. In previous reports, state variables were selected by adopting a complicated workflow requiring multiple software licenses and an advanced level of user expertise. In this report, a new workflow that automates the state variable selection process is presented. It significantly reduces the frequency of user interventions and does not require multiple software licenses. Each module in the new workflow is described in detail, and the input / output examples in each step of the workflow are provided. It was demonstrated that this workflow can greatly reduce the complexity of the state variable selection process, and that the updated FARM-Gamma and FARM-Delta validators can benefit from this workflow when solving the power dispatch problem of a representative IES test case. Finally, some code improvements that can further enhance the efficiency are suggested.

42 ENGINEERING↗

Predicting Ecologically Important Vegetation Variables from Remotely Sensed Optical/Radar Data Using Neural Networks

A number of satellite sensor systems will collect large data sets of the Earth's surface during NASA's Earth Observing System (EOS) era. Efforts are being made to develop efficient algorithms that can incorporate a wide variety of spectral data and ancillary data in order to extract vegetation variables required for global and regional studies of ecosystem processes, biosphere-atmosphere interactions, and carbon dynamics. These variables are, for the most part, continuous (e.g. biomass, leaf area index, fraction of vegetation cover, vegetation height, vegetation age, spectral albedo, absorbed photosynthetic active radiation, photosynthetic efficiency, etc.) and estimates may be made using remotely sensed data (e.g. nadir and directional optical wavelengths, multifrequency radar backscatter) and any other readily available ancillary data (e.g., topography, sun angle, ground data, etc.). Using these types of data, neural networks can: 1) provide accurate initial models for extracting vegetation variables when an adequate amount of data is available; 2) provide a performance standard for evaluating existing physically-based models; 3) invert multivariate, physically based models; 4) in a variable selection process, identify those independent variables which best infer the vegetation variable(s) of interest; and 5) incorporate new data sources that would be difficult or impossible to use with conventional techniques. In addition, neural networks employ a more powerful and adaptive nonlinear equation form as compared to traditional linear, index transformations, and simple nonlinear analyses. These neural networks attributes are discussed in the context of the authors' investigations of extracting vegetation variables of ecological interest.

Kimes, Daniel S.↗

A Framework for Categorizing Important Project Variables

While substantial research has led to theories concerning the variables that affect project success, no universal set of such variables has been acknowledged as the standard. The identification of a specific set of controllable variables is needed to minimize project failure. Much has been hypothesized about the need to match project controls and management processes to individual projects in order to increase the chance for success. However, an accepted taxonomy for facilitating this matching process does not exist. This paper surveyed existing literature on classification of project variables. After an analysis of those proposals, a simplified categorization is offered to encourage further research.

Parsons, Vickie S.↗

Machine Learning in Infectious Disease for Risk Factor Identification and Hypothesis Generation: Proof of Concept Using Invasive Candidiasis

Machine learning (ML) models can handle large data sets without assuming underlying relationships and can be useful for evaluating disease characteristics, yet they are more commonly used for predicting individual disease risk than for identifying factors at the population level. We offer a proof of concept applying random forest (RF) algorithms to Candida-positive hospital encounters in an electronic health record database of patients in the United States. Candida-positive encounters were extracted from the Cerner HealthFacts database; invasive infections were laboratory-positive sterile site Candida infections. Features included demographics, admission source, care setting, physician specialty, diagnostic and procedure codes, and medications received before the first positive Candida culture. We used RF to assess risk factors for 3 outcomes: any invasive candidiasis (IC) vs non-IC, within-species IC vs non-IC (eg, invasive C. glabrata vs noninvasive C. glabrata), and between-species IC (eg, invasive C. glabrata vs all other IC). Fourteen of 169 (8%) variables were consistently identified as important features in the ML models. When evaluating within-species IC, for example, invasive C. glabrata vs non-invasive C. glabrata, we identified known features like central venous catheters, intensive care unit stay, and gastrointestinal operations. In contrast, important variables for invasive C. glabrata vs all other IC included renal disease and medications like diabetes therapeutics, cholesterol medications, and antiarrhythmics. Known and novel risk factors for IC were identified using ML, demonstrating the hypothesis-generating utility of this approach for infectious disease conditions about which less is known, specifically at the species level or for rarer diseases.

60 APPLIED LIFE SCIENCES↗

Identifying Entangled Physics Relationships Through Sparse Matrix Decomposition to Inform Plasma Fusion Design

We report a sustainable burn platform through inertial confinement fusion (ICF) has been an ongoing challenge for over 50 years. Mitigating engineering limitations and improving the current design involves an understanding of the complex coupling of physical processes. While sophisticated simulation codes are used to model ICF implosions, these tools contain necessary numerical approximation but miss physical processes that limit predictive capability. Identification of relationships between controllable design inputs to ICF experiments and measurable outcomes (e.g., neutron yield, neutron velocity, areal density) from performed experiments can help guide the future design of experiments and development of simulation codes, to potentially improve the accuracy of the computational models used to simulate ICF experiments. We use sparse matrix decomposition methods to identify clusters of a few related design variables. Sparse principal component analysis (SPCA) identifies groupings that are related to the physical origin of the variables (laser, hohlraum, and capsule). A variable importance analysis finds that in addition to variables highly correlated with neutron yield, such as picket power and laser energy, variables that represent a dramatic change of the ICF design, such as number of pulse steps, are also very important. The obtained sparse components are then used to train a random forest (RF) regression surrogate for predicting total yield. The RF performance on the training and testing data compares with the performance of the RF trained using all the design variables considered. This work is intended to inform design changes in future ICF experiments by augmenting the expert intuition and simulation results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data‐driven identification of environmental variables influencing phenotypic plasticity to facilitate breeding for future climates

Summary Phenotypic plasticity describes a genotype's ability to produce different phenotypes in response to different environments. Breeding crops that exhibit appropriate levels of plasticity for future climates will be crucial to meeting global demand, but knowledge of the critical environmental factors is limited to a handful of well‐studied major crops. Using 727 maize ( Zea mays L.) hybrids phenotyped for grain yield in 45 environments, we investigated the ability of a genetic algorithm and two other methods to identify environmental determinants of grain yield from a large set of candidate environmental variables constructed using minimal assumptions. The genetic algorithm identified pre‐ and postanthesis maximum temperature, mid‐season solar radiation, and whole season net evapotranspiration as the four most important variables from a candidate set of 9150. Importantly, these four variables are supported by previous literature. After calculating reaction norms for each environmental variable, candidate genes were identified and gene annotations investigated to demonstrate how this method can generate insights into phenotypic plasticity. The genetic algorithm successfully identified known environmental determinants of hybrid maize grain yield. This demonstrates that the methodology could be applied to other less well‐studied phenotypes and crops to improve understanding of phenotypic plasticity and facilitate breeding crops for future climates.

Kusmec, Aaron↗

A Researcher's Perspective on Function Allocation and Its Application to Air Traffic Management

Functional Allocation is an important research area for NASA going forward, and aims at exploring a range of potential options for future ATM (Air Traffic Management) operation concepts. The goal of this work is to gain some insight into the relative strengths and limitations of a range of different concepts, and to identify the important variables as well as any critical or threshold values for those variables for each of those concepts. This information will be captured in a form that can be easily referenced, and can then be used by decision makers to chose which concept or concepts best suit the needs of the US NAS (National Airspace System). The work presented here is one researcher's perspective on this process. A brief overview of Functional Allocation will be followed by more detailed discussion about two of the topics being currently investigated, as well as why and how this work differs from the way we have done ATM research in the past. More specifically, the goal of the analysis work is to understand the reasons for any differences in our metrics (such as delay) when adjusting input variables rather than trying to analyze the raw numbers themselves. In other words, finding which variables different concepts are sensitive to is important for this work, while finding a best or recommended operational concept is not. This will be followed by some examples of this process using results from our current studies. Finally, there will be a discussion on where we are in the process as well as our next steps.

Air-Ground Studies↗

Mapping National Forest Aboveground Biomass in Mexico by Integrating GEDI, Sentinel‐1 and Sentinel‐2 Data

Accurate mapping of forest aboveground biomass density (AGBD) is required to better understand the role of forests in the global carbon cycle and to support international policies for climate change mitigation and adaptation. Mexico is one of the countries having great potential for the United Nations Programme on Reducing Emissions from Deforestation and Forest Degradation (or UN-REDD program) and there is a growing demand for unbiased Monitoring Reporting Verification systems at a national level. As an effort under NASA’s Carbon Monitoring System (CMS) program, we developed a machine learning model using multi-stream remote sensing measurements as well as topographic data to create a high spatial resolution AGBD map (~100 m) over Mexico (circa 2020). The remote sensing data includes Global Ecosystem Dynamic Investigation (GEDI) lidar, Sentinel 1 Synthetic-Aperture Radar (SAR), and Sentinel-2 multispectral imagery (MSI). GEDI onboard the International Space Station provides unprecedented forest structure and AGBD sampling datasets for model training and validation practices. Our analysis indicates that the developed random forest model can capture 63 % of the spatial variation (RMSE = 33.7 Mg/ha) of AGBD of Mexican forests. We find that shortwave infrared bands of Sentinel-2 MSI and topographical variables from elevation data are the most important variables in the developed AGBD model. Our study highlights methodological opportunities in synergistic uses of multiple sensors for large-scale forest AGBD mapping and shows potential for retrospective analysis and operational monitoring of forest AGBD and its dynamics.

Taejin Park↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗