Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “variable importance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Constrained or unconstrained? Neural-network-based equation discovery from data

Throughout many fields, practitioners often rely on differential equations to model systems. Yet, for many applications, the theoretical derivation of such equations and/or the accurate resolution of their solutions may be intractable. Instead, recently developed methods, including those based on parameter estimation, operator subset selection, and neural networks, allow for the data-driven discovery of both ordinary and partial differential equations (PDEs), on a spectrum of interpretability. The success of these strategies is often contingent upon the correct identification of representative equations from noisy observations of state variables and, as importantly and intertwined with that, the mathematical strategies utilized to enforce those equations. Specifically, the latter has been commonly addressed via unconstrained optimization strategies. Representing the PDE as a neural network, we propose to discover the PDE (or the associated operator) by solving a constrained optimization problem and using an intermediate state representation similar to a physics-informed neural network (PINN). The objective function of this constrained optimization problem promotes matching the data, while the constraints require that the discovered PDE is satisfied at a number of spatial collocation points. We present a penalty method and a widely used trust-region barrier method to solve this constrained optimization problem, and we compare these methods on numerical examples. Our results on several example problems demonstrate that the latter constrained method outperforms the penalty method, particularly for higher noise levels or fewer collocation points. This work motivates further exploration into using sophisticated constrained optimization methods in scientific machine learning, as opposed to their commonly used, penalty-method or unconstrained counterparts. For both of these methods, we solve these discovered neural network PDEs with classical methods, such as finite difference methods, as opposed to PINNs-type methods relying on automatic differentiation. Here, we briefly highlight how simultaneously fitting the data while discovering the PDE improves the robustness to noise and other small, yet crucial, implementation details.

Data-driven discovery↗

Evaluation of average leaf inclination angle quantified by indirect optical instruments in crop fields

Average leaf inclination angle ($\overline{θ}$ L ) is an important canopy structure variable that influences light regime, photosynthesis, and evapotranspiration of plants. $\overline{θ}$ L can be measured through direct methods (e.g., protractor), which are labor-intensive and time-consuming, or through indirect optical instruments, which are more efficient than the direct methods. However, uncertainties of different indirect optical instruments for quantifying $\overline{θ}$ L remain largely unquantified. In this study, we evaluated and compared the performances of three major indirect optical instruments: (1) LAI-2200, (2) 30°-tilted camera, and (3) digital hemispherical photography (DHP), in different crop fields over a growing season, benchmarked with direct measurements. LAI-2200 and 30°-tilted camera showed higher agreement with direct $\overline{θ}$ measurements (R 2 = 0.54, RMSE = 7.37°; R 2 = 0.58, RMSE = 8.08°) than DHP (R 2 = 0.14, RMSE = 13.96°). Different performances of indirect optical instruments could be attributed to the accuracy of gap fraction measurement and the performance of the $\overline{θ}$ L quantification algorithms. When using the LAI-2200 algorithm, larger gap fraction gradients over view zenith angles led to larger $\overline{θ}$ L values, and smaller gap fraction gradients led to smaller $\overline{θ}$ L values. Such error propagation was larger in sparse canopy than in dense canopy. The Wilson G function of the LAI-2200 algorithm performed better in estimating $\overline{θ}$ L than the G function based on the ellipsoidal LAD function used by the CAN_EYE algorithm. We also proposed a modification of the LAI-2200 algorithm, which further improved the performance of LAI-2200 and 30°-tilted cameras in estimating $\overline{θ}$ L . We envision that the low-cost 30°-tilted cameras provide a promising sensor solution to continuously monitor canopy structure for various ecosystems.

30°-tilted camera↗

Stromataxic Stabilization of a Metastable Layered ScFeO3 Polymorph

Metastable polymorphs--materials with the same stoichiometry as the ground state but a different crystal structure--enable many critical technologies. This work describes the development of a stabilization approach for metastable polymorphs that are difficult to achieve through other stabilization techniques (such as epitaxy or quenching) called stromataxy. Stromataxy is a method based on controlling the precursor structure during the initial stages of material growth to dictate phase formation. To illustrate this approach, we controlled the atomic layering of the precursors of ScFeO3 and stabilized the metastable P63cm phase, under conditions that previously led to the ground-state Ia3¯ bixbyite phase. Ab initio mechanistic calculations highlight the importance of the variable oxidation state of Fe and the layer stability during layer-by-layer growth. The broad applicability of a stromataxy approach was demonstrated by stabilizing this metastable phase on substrates that have previously been shown to stabilize other polymorphs under continuous growth. Stromataxy is shown as a viable option for accessing polymorphs that are close in energy, difficult to differentiate by strain, or that lack a well epitaxially matched substrate.

calculations↗

Data for Root Exudation Links Root Traits to Soil Functioning in Agroecosystems

Root exudation is a key process for plant nutrient acquisition, but the controls on root exudation and its relationship to soil C and N processes in agroecosystems are unclear. We hypothesized that root exudation rates would be related to root morphological traits, N fertilization, and soil moisture. We also anticipated that root exudation would be correlated with bulk soil enzyme activity. Root exudation, root traits, and bulk soil extracellular enzyme activity were assessed in maize (Zea mays L.), soybean (Glycine max (L.) Merr.), biomass sorghum (Sorghum bicolor (L.) Moench), giant miscanthus (Miscanthus × giganteus), and switchgrass (Panicum virgatum L.). Measurements were taken in situ during two growing seasons with contrasting precipitation regimes, and N fertilization rate was varied in sorghum during one year. Specific root exudation (per unit root surface area) was negatively related to root diameter and was generally higher in annuals than perennials. Sorghum N fertilization did not affect root exudation rates, and soil moisture regime had no effect on annual root exudation rates within maize, sorghum, and miscanthus. Specific root exudation was negatively related to bulk soil C- and N-degrading soil enzyme activities. Intrinsic plant characteristics appeared more important than environmental variables in controlling in situ root exudation rates. The relationships between root diameter, root exudation, and soil C and N processes link root morphological traits to soil functions and demonstrate the potential tradeoffs among plant nutrient acquisition strategies in agroecosystems.

Biomass Analytics↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Adoption of Plug-in Electric Vehicles: Local Fuel Use and Greenhouse Gas Emissions Reductions Across the U.S.

The dependence on gasoline-powered light-duty automobiles has made U.S. households vulnerable to the burden of fuel costs. Tailpipe emissions from these vehicles constitute 58% of greenhouse gas (GHG) emissions in the U.S., which are damaging to the environment (EPA, 2023). The adoption of plug-in electric vehicles (PEVs) has been shown to effectively reduce fuel costs and GHG emissions. However, local effects on these benefits are not well understood by American consumers, potentially limiting adoption and therefore the realization of PEV benefits at scale (MacInnis & Krosnick, 2020; EY Americas, 2023). To fill this research gap, this study estimates the fuel cost savings and GHG emission reductions at the state and ZIP code levels by considering local fuel prices, vehicle class preference, average vehicle model year, fuel efficiencies, and driving intensities. The study's findings reveal that the adoption of PEVs can yield substantial benefits in terms of fuel cost savings and GHG emission reductions nationwide. Specifically, driving a battery electric vehicle (BEV) is estimated to result in annual savings of up to $\$2,200$, while driving a plug-in hybrid electric vehicle (PHEV) can lead to savings up to $\$1,500$, when compared to an internal combustion engine vehicle (ICEV) of equivalent size. Moreover, using population-weighted averages by ZIP code, BEVs and PHEVs show the potential to save 400 and 200 grams of carbon dioxide equivalent per mile, respectively, compared to a representative ICEV of the same class. The magnitude of fuel cost savings and emissions reduction vary by region due to various factors. Generally, regions with high gasoline prices, low electricity prices, preferences for larger vehicles, and high driving intensities tend to see relatively large fuel savings. The emissions reductions are more pronounced in areas with clean grids where consumer preferences lie with large vehicles. This regional variability underscores the importance of considering local contextual factors when assessing the potential benefits of PEV adoption. In more than 99% of U.S. ZIP codes, PEVs result in overall savings in fuel use (and subsequent costs) and GHG emissions. While not a central focus of this analysis, reductions in GHG tailpipe emissions from PEV adoption would also come with reductions in criteria pollutant emissions, contributing to improved local air quality depending on the PEV penetration, population density, and electricity generation infrastructure in the locality.

33 ADVANCED PROPULSION SYSTEMS↗

Modeling Oxygen Partial Pressure in Solid Oxide Electrolysis Cells: The Microstructure Effect

Oxygen partial pressure is an important thermodynamic state variable that affects both the performance and degradation of solid oxide electrolysis cells. In this work, a 3D model developed from Virkar’s 1D model has been applied to reconstructed and synthetic microstructures of Ni-YSZ-GDC-LSCF cell. The effect of the microstructures, including the thickness of the YSZ and GDC layer, and the compositions of the hydrogen and oxygen electrode, on the distribution of oxygen partial pressure has been investigated. The results show that the maximum oxygen partial pressure may occur on the interface between oxygen electrode and GDC layer or between GDC layer and YSZ layer depending on the rate of oxygen ion exchange between GDC and YSZ. A thicker GDC layer lowers the maximum oxygen partial pressure in the cell, while a thicker YSZ layer lowers the maximum oxygen partial pressure in the hydrogen electrode. In addition, the Ni:YSZ ratio and porosity also affects the maximum partial pressure. These findings provide insights on mitigating degradation in solid oxide electrolysis cell by tuning the microstructures.

Lei, Yinkai↗

Key Environmental and Ecological Variables of Wetland CH 4 and CO 2 Fluxes Change With Warming

Wetlands are important ecosystems for the global carbon cycle, impacting regional and global methane (CH 4 ) and carbon dioxide (CO 2 ) budgets. This study examines how environmental and ecological variables impact wetland CH 4 flux and net ecosystem exchange of CO 2 (NEE) across 17 sites globally. We also quantified the importance of variables for each wetland type and site at monthly scale under normal and warm temperatures using dominance analysis. We identified soil and air temperature (TS, TA, respectively) as key variables influencing wetland CH4, and latent heat (LE) and shortwave radiation (SW) for NEE under normal and warm conditions. However, the importance of some variables shifted with warming. For predicting the variability of wetland CH4 flux under warming, gross primary productivity (GPP) and LE, replacing wind direction (WD), were dominant variables for tropical swamps, while NEE was important for high-latitude fens and bogs under warm temperatures. For wetland NEE, the role of TA and TS decreased across all wetland types with warming, while vapor pressure deficit (VPD) became more important for mid and high-latitude wetlands. Our results reveal the complex responses of wetland carbon flux to environmental and ecological variables with warming and provide new insights into improving wetland models by incorporating additional variables and accounting for the changing roles of variables in carbon flux under warming.

54 ENVIRONMENTAL SCIENCES↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Limited Role of Absolute Humidity in Intraurban Heat Variability

Abstract Monitoring and understanding the variability of heat within cities is important for urban planning and public health, and the number of studies measuring intraurban temperature variability is growing. Recognizing that the physiological effects of heat depend on humidity as well as temperature, measurement campaigns have included measurements of relative humidity alongside temperature. However, the role the spatial structure in humidity, independent from temperature, plays in intraurban heat variability is unknown. Here we use summer temperature and humidity from networks of stationary sensors in multiple cities in the United States to show spatial variations in the absolute humidity within these cities are weak. This variability in absolute humidity plays an insignificant role in the spatial variability of the heat index and humidity index (humidex), and the spatial variability of the heat metrics is dominated by temperature variability. Thus, results from previous studies that considered only intraurban variability in temperature will carry over to intraurban heat variability. Also, this suggests increases in humidity from green infrastructure interventions designed to reduce temperature will be minimal. In addition, a network of sensors that only measures temperature is sufficient to quantify the spatial variability of heat across these cities when combined with humidity measured at a single location, allowing for lower-cost heat monitoring networks. Significance Statement Monitoring the variability of heat within cities is important for urban planning and public health. While the physiological effects of heat depend on temperature and humidity, it is shown that there are only weak spatial variations in the absolute humidity within nine U.S. cities, and the spatial variability of heat metrics is dominated by temperature variability. This suggests increases in humidity will be minimal resulting from green infrastructure interventions designed to reduce temperature. It also means a network of sensors that only measure temperature is sufficient to quantify the spatial variability of heat across these cities when combined with humidity measured at a single location.

Meteorology & Atmospheric Sciences↗

Investigating the Impact of Power-Take-Off System Parameters and Control Law on a Rotational Wave Energy Converter’s Peak-to-Average Power Ratio Reduction

Due to the irregular nature of real waves, the power captured in a wave energy converter (WEC) system is highly variable. This is an important barrier to the effective use of WECs. To address this challenge, this study focuses on a rotational WEC power-take-off system in which high-speed and high-efficiency generators along with a torque/power smoothing inertia element can be effectively utilized. In the first phase of this study, the U.S. Department of Energy’s reference model 3 (WEC-Sim RM3; two-body point absorber), along with a slider-crank WEC, were integrated for linear to rotational conversion. Relative motion between the float and spar in RM3 was the driving force for this slider-crank WEC, which is connected to a motor/generator set through a gearbox. RM3 geometry was scaled down by 25 times to work within the limits of the physical motor/generator set used in the experimentation. Once the integration in a hardware-in-the-loop simulation environment was successfully completed, data on the peak-to-average power ratio was collected for various wave conditions including regular and irregular waves. The control algorithm designed to keep the system in resonance with waves was able to maintain relatively high speed depending on the specific gear ratio and wave period. Initial results with hardware-in-the-loop simulations reveal that gear ratio and crank radius have a strong impact on the peak-to-average power ratio. In addition, it was found that output power from the generator was maximized at a larger gear ratio, as the crank radius was increased.

50 EE - Wind and Water Power Program - Water (EE-4↗

The relative importance of wind and hydroclimate drivers in modulating the interannual variability of dust emissions in Earth system models

Windblown dust emissions are controlled by near-surface wind speed and sediment erodibility, the latter modulated by hydroclimate and land-use conditions. Accurate representations of these drivers are critical for reproducing historical dust variability and projecting future dust changes in Earth system models (ESMs). This study examines the discrepancies among 21 ESMs in the relative importance of wind speed versus five hydroclimate drivers in explaining the historical (1980–2014) variability of dust emissions from global drylands. In hyperarid areas, models show poor agreement in the simulated dust variability, with only 9 % out of 210 inter-model comparisons exhibiting significant positive correlations. In contrast, arid and semiarid areas exhibit a dual pattern driven by a “double-edged sword” effect of land surface memory: models with coherent hydroclimate variability show better agreement, whereas those with divergent hydroclimate representations show larger disagreement. While the ESMs capture the dominant role of wind speed in hyperarid areas, they diverge markedly in the relative contributions of wind and hydroclimate drivers in arid and semiarid areas. Replacing the Zender et al. (2003) dust scheme with the Kok et al. (2014) scheme in CESM and E3SM generally strengthens hydroclimate influences while reducing wind speed contributions to simulated dust variability. MERRA-2 reanalysis produces stronger wind influences than most ESMs across all dryland regions. These results underscore the need for improved near-surface wind simulations in hyperarid areas and more realistic land surface and hydroclimate representations in arid and semiarid areas to reduce uncertainties in global dust emission simulations.

Li, Xinzhu [Michigan Technological University, Hou↗

Dispatching Long Duration Storage on High PV Systems

Long duration storage is expected to become increasingly important as shares of variable renewable energy increase. This presentation details some of the challenges of representing long duration storage, the importance of accurately modeling its dispatch, and a few methods for improving dispatch models.

dispatch↗

Disentangling the hydrological and hydraulic controls on streamflow variability in Energy Exascale Earth System Model (E3SM) V2 – a case study in the Pantanal region

Abstract. Streamflow variability plays a crucial role in shaping the dynamics and sustainability of Earth's ecosystems, which can be simulated and projected by a river routing model coupled with a land surface model. However, the simulation of streamflow at large scales is subject to considerable uncertainties, primarily arising from two related processes: runoff generation (hydrological process) and river routing (hydraulic process). While both processes have impacts on streamflow variability, previous studies only calibrated one of the two processes to reduce biases in the simulated streamflow. Calibration focusing only on one process can result in unrealistic parameter values to compensate for the bias resulting from the other process; thus other water-related variables remain poorly simulated. In this study, we performed several experiments with the land and river components of the Energy Exascale Earth System Model (E3SM) over the Pantanal region to disentangle the hydrological and hydraulic controls on streamflow variability in coupled land–river simulations. Our results show that the generation of subsurface runoff is the most important factor for streamflow variability contributed by the runoff generation process, while floodplain storage effect and main-channel roughness have significant impacts on streamflow variability through the river routing process. We further propose a two-step procedure to robustly calibrate the two processes together. The impacts of runoff generation and river routing on streamflow are appropriately addressed with the two-step calibration, which may be adopted by developers of land surface and earth system models to improve the modeling of streamflow.

54 ENVIRONMENTAL SCIENCES↗

Automatic detection of cataclysmic variables from SDSS images

Abstract Investigating rare and new objects have always been an important direction in astronomy. Cataclysmic variables (CVs) are ideal and natural celestial bodies for studying the accretion process of semi-detached binaries with accretion processes. However, the sample size of CVs must increase because a lager gap exists between the observational and the theoretical expanding CVs. Astronomy has entered the big data era and can provide massive images containing CV candidates. CVs as a type of faint celestial objects, are highly challenging to be identified directly from images using automatic manners. Deep learning has rapidly developed in intelligent image processing and has been widely applied in some astronomical fields with excellent detection results. YOLOX, as the latest YOLO framework, is advantageous in detecting small and dark targets. This work proposes an improved YOLOX-based framework according to the characteristics of CVs and Sloan Digital Sky Survey (SDSS) photometric images to train and verify the model to realise CV detection. We use the Convolutional Block Attention Module to increase the number of output features with the feature extraction network and adjust the feature fusion network to obtain fused features. Accordingly, the loss function is modified. Experimental results demonstrate that the improved model produces satisfactory results, with average accuracy (mean average Precision at 0.5) of 92.0%, Precision of 92.9%, Recall of 94.3%, and $F1-score$ of 93.6% on the test set. The proposed method can efficiently achieve the identification of CVs in test samples and search for CV candidates in unlabeled images. The image data vastly outnumber the spectra in the SDSS-released data. With supplementary follow-up observations or spectra, the proposed model can help astronomers in seeking and detecting CVs in a new manner to ensure that a more extensive CV catalog can be built. The proposed model may also be applied to the detection of other kinds of celestial objects.

Astronomy & Astrophysics↗

Evapotranspiration Partitioning Using Flux Tower Data in a Semi-Arid Ecosystem

Information about evapotranspiration (ET) and its components, that is, evaporation and transpiration, is crucial for a wide range of water and ecosystem management applications. However, partitioning ET into its two components is often challenging because of their spatiotemporal variabilities and lack of process understanding. This study developed a machine learning (ML) framework to shed light on ET processes and assess the relative importance of different drivers by incorporating hydrometeorology and biomass productivity variables. The Shapley Additive Explanations (SHAP) approach was applied to enhance explainability and rank the importance of ET drivers and their components. A total of 62 variables covering hydrometeorological and biomass productivity dimensions were considered from the Reynolds Creek Critical Zone Observatory (CZO) station in Idaho. The variable importance assessment identified the leading drivers individually for evaporation, transpiration and ET (soil water content for evaporation, vapour pressure deficit for transpiration and soil water content for ET). The results further highlighted the value of combining hydrometeorological and biomass productivity variables to achieve better predictability of ET processes.

54 ENVIRONMENTAL SCIENCES↗

Modeling household online shopping demand in the U.S.: a machine learning approach and comparative investigation between 2009 and 2017

Despite the rapid growth of online shopping and research interest in the relationship between online and in-store shopping, national-level modeling and investigation of the demand for online shopping with a prediction focus remain limited in the literature. Here, this paper differs from prior work and leverages two recent releases of the U.S. National Household Travel Survey (NHTS) data for 2009 and 2017 to develop machine learning (ML) models, specifically gradient boosting machine (GBM), for predicting household-level online shopping purchases. The NHTS data allow for not only conducting nationwide investigation but also at the level of households, which is more appropriate than at the individual level given the connected consumption and shopping needs of members in a household. We follow a systematic procedure for model development including employing Recursive Feature Elimination algorithm to select input variables (features) in order to reduce the risk of model overfitting and increase model explainability. Among several ML models, GBM is found to yield the best prediction accuracy. Extensive post-modeling investigation is conducted in a comparative manner between 2009 and 2017, including quantifying the importance of each input variable in predicting online shopping demand, and characterizing value-dependent relationships between demand and the input variables. In doing so, two latest advances in machine learning techniques, namely Shapley value-based feature importance and Accumulated Local Effects plots, are adopted to overcome inherent drawbacks of the popular techniques in current ML modeling. The modeling and investigation are performed at the national level, with a number of findings obtained. The models developed and insights gained can be used for online shopping-related freight demand generation and may also be considered for evaluating the potential impact of relevant policies on online shopping demand.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗