Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Machine learning predictions for local electronic properties of disordered correlated electron systems

We present a scalable machine learning (ML) model to predict local electronic properties such as on-site electron number and double occupation for disordered correlated electron systems. Our approach is based on the locality principle, or the nearsightedness nature, of many-electron systems, which means local electronic properties depend mainly on the immediate environment. A ML model is developed to encode this complex dependence of local quantities on the neighborhood. We demonstrate our approach using the square-lattice Anderson-Hubbard model, which is a paradigmatic system for studying the interplay between Mott transition and Anderson localization. We develop a lattice descriptor based on the group-theoretical method to represent the on-site random potentials within a finite region. The resultant feature variables are used as input to a multilayer fully connected neural network, which is trained from data sets of variational Monte Carlo (VMC) simulations on small systems. We show that the ML predictions agree reasonably well with the VMC data. Our work underscores the promising potential of ML methods for multiscale modeling of correlated electron systems.

36 MATERIALS SCIENCE↗

Calibration and Rapid-Adoption Forecasting Techniques

CRAFT (Calibration and Rapid-Adoption Forecasting Techniques) CRAFT is a Python-based project for processing, analyzing, and modeling atmospheric or environmental data. It uses machine learning techniques, specifically Random Forest Regression, to create emulators for various environmental variables such as gross primary production and soil water content. It then uses these emulators to robustly test the parameter space of mechanistic models to provide posterior estimations of the free parameters.

Robins, Zachary↗

A Latent-Variable Formulation of the Poisson Canonical Polyadic Tensor Model: Maximum Likelihood Estimation and Fisher Information

We establish parameter inference for the Poisson canonical polyadic (PCP) tensor model through a latent-variable formulation. Our approach exploits the observation that any random PCP tensor can be derived by marginalizing an unobservable random tensor of one dimension larger. The loglikelihood of this larger dimensional tensor, referred to as the “complete” loglikelihood, is comprised of multiple rank one PCP loglikelihoods. Using this methodology, we first derive maximum likelihood estimators for the PCP model and demonstrate that several existing algorithms for fitting non-negative matrix and tensor factorizations are Expectation-Maximization algorithms. Next, we derive the observed and expected Fisher information matrices for the PCP model. The Fisher information provides us crucial insights into the well-posedness of the tensor model, such as the role that tensor rank plays in identifiability and indeterminacy. For the special case of rank one PCP models, we demonstrate that these results are greatly simplified.

97 MATHEMATICS AND COMPUTING↗

COBRA:COMPUTED-TOMOGRAPHY BASED RANDOM-FIELD APPROXIMATION

SF-25-115 COBRA (COmputed-tomography Based Random-field Approximation) is a Python application for generating statistically equivalent random fields from CT-scan imagery. It leverages Karhunen–Loève expansions to model microstructural variability, enabling users to: Preprocess CT scans (filtering and Gaussian transformation); Fit covariance kernels fromempirical data; Solve eigenproblems to obtain KL modes; Sample random fields onsistent with fitted statistics; Postprocess samples back into the physical domain.

Hu, Tianchen↗

Climate, Hydrology, and Nutrients Control the Seasonality of Si Concentrations in Rivers

Abstract The seasonal behavior of fluvial dissolved silica (DSi) concentrations, termedDSi regime, mediates the timing of DSi delivery to downstream waters and thus governs river biogeochemical function and aquatic community condition. Previous work identified five distinct DSi regimes across rivers spanning the Northern Hemisphere, with many rivers exhibiting multiple DSi regimes over time. Several potential drivers of DSi regime behavior have been identified at small scales, including climate, land cover, and lithology, and yet the large‐scale spatiotemporal controls on DSi regimes have not been identified. We evaluate the role of environmental variables on the behavior of DSi regimes in nearly 200 rivers across the Northern Hemisphere using random forest models. Our models aim to elucidate the controls that give rise to (a) average DSi regime behavior, (b) interannual variability in DSi regime behavior (i.e., Annual DSi regime), and (c) controls on DSi regime shape (i.e., minimum and maximum DSi concentrations). Average DSi regime behavior across the period of record was classified accurately 59% of the time, whereas Annual DSi regime behavior was classified accurately 80% of the time. Climate and primary productivity variables were important in predicting Average DSi regime behavior, whereas climate and hydrologic variables were important in predicting Annual DSi regime behavior. Median nitrogen and phosphorus concentrations were important drivers of minimum and maximum DSi concentrations, indicating that these macronutrients may be important for seasonal DSi drawdown and rebound. Our findings demonstrate that fluctuations in climate, hydrology, and nutrient availability of rivers shape the temporal availability of fluvial DSi.

Environmental Sciences & Ecology↗

Code Description for "Brief Communication: Monitoring snow depth using small, cheap, and easy-to-deploy ground surface temperature sensors"

Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We train a random forest machine learning model to predict snow depth from variability in ground surface temperature. To our knowledge, this is the first time that small ground surface temperature sensors have been used to estimate snow depth. The model performs well at sites where the model was trained and at pan-arctic evaluation sites (RMSE <= 0.15 m). Small temperature sensors are cheap and easy-to-deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring to an extent previously infeasible. The model is flexible and can be applied to datasets retroactively to retrieve snow depth estimates at additional sites. This code package includes a *.joblib file of the trained random forest model and a *.ipynb file showing how to clean input data, train the random forest model, and apply the model.

Bachand, Claire↗

Estimating carrying capacity for juvenile salmon using quantile random forest models

Abstract Establishing robust methods and metrics to evaluate habitat quality is critical for the recovery of endangered Pacific salmonids ( Oncorhynchus spp.). A variety of modeling approaches are used for status and trend monitoring of anadromous species throughout the Pacific Northwest, USA, but current methods may fail to capture the complex relationship between fish and habitat and are often limited in predictive power beyond specific watersheds. Further, the focus on species distribution and abundance is not easily manipulated to predict carrying capacity and traditional stock‐recruitment analyses are reliant on long‐term data which are not always available. In this study, we developed a quantile random forest model to provide estimates of habitat carrying capacity for Chinook salmon ( O. tshawytscha ) parr during the summer months, at both the site and watershed scale. Quantile random forest models allow for the consideration of noisy data, correlated variables, and non‐linear relationships: common features in fish–habitat datasets. We leveraged Columbia Habitat Monitoring Program data to select habitat co‐variates and predict capacity at those sites. We also identified a set of globally available attributes to extrapolate capacity estimate predictions throughout wadeable streams within the Columbia River basin. Total capacity estimates for watersheds closely matched estimates from alternative fish productivity models. Carrying capacity estimates based on quantile random forest models, like those presented here, provide managers a framework to guide the identification, prioritization, and development of habitat rehabilitation actions to recover salmon populations.

See, Kevin E.↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗

Spatial patterns of snow distribution in the sub-Arctic

Abstract. The spatial distribution of snow plays a vital role in sub-Arctic and Arctic climate, hydrology, and ecology due to its fundamental influence on the water balance, thermal regimes, vegetation, and carbon flux. However, the spatial distribution of snow is not well understood, and therefore, it is not well modeled, which can lead to substantial uncertainties in snow cover representations. To capture key hydro-ecological controls on snow spatial distribution, we carried out intensive field studies over multiple years for two small (2017–2019; ∼ 2.5 km2) sub-Arctic study sites located on the Seward Peninsula of Alaska. Using an intensive suite of field observations (> 22 000 data points), we developed simple models of the spatial distribution of snow water equivalent (SWE) using factors such as topographic characteristics, vegetation characteristics based on greenness (normalized different vegetation index, NDVI), and a simple metric for approximating winds. The most successful model was random forest, using both study sites and all years, which was able to accurately capture the complexity and variability of snow characteristics across the sites. Approximately 86 % of the SWE distribution could be accounted for, on average, by the random forest model at the study sites. Factors that impacted year-to-year snow distribution included NDVI, elevation, and a metric to represent coarse microtopography (topographic position index, TPI), while slope, wind, and fine microtopography factors were less important. The characterization of the SWE spatial distribution patterns will be used to validate and improve snow distribution modeling in the Department of Energy's Earth system model and for improved understanding of hydrology, topography, and vegetation dynamics in the sub-Arctic and Arctic regions of the globe.

54 ENVIRONMENTAL SCIENCES↗

Alert Classification for the ALeRCE Broker System: The Light Curve Classifier

We present the first version of the Automatic Learning for the Rapid Classification of Events (ALeRCE) broker light curve classifier. ALeRCE is currently processing the Zwicky Transient Facility (ZTF) alert stream, in preparation for the Vera C. Rubin Observatory. The ALeRCE light curve classifier uses variability features computed from the ZTF alert stream and colors obtained from AllWISE and ZTF photometry. We apply a balanced random forest algorithm with a two-level scheme where the top level classifies each source as periodic, stochastic, or transient, and the bottom level further resolves each of these hierarchical classes among 15 total classes. This classifier corresponds to the first attempt to classify multiple classes of stochastic variables (including core- and host-dominated active galactic nuclei, blazars, young stellar objects, and cataclysmic variables) in addition to different classes of periodic and transient sources, using real data. We created a labeled set using various public catalogs (such as the Catalina Surveys and Gaia DR2 variable stars catalogs, and the Million Quasars catalog), and we classify all objects with ≥6 g-band or ≥6 r-band detections in ZTF (868,371 sources as of 2020 June 9), providing updated classifications for sources with new alerts every day. For the top level we obtain macro-averaged precision and recall scores of 0.96 and 0.99, respectively, and for the bottom level we obtain macro-averaged precision and recall scores of 0.57 and 0.76, respectively. Updated classifications from the light curve classifier can be found at the ALeRCE Explorer website (http://alerce.online).

47 OTHER INSTRUMENTATION↗

Role of crystallographic orientation on intragranular void growth in polycrystalline FCC materials

In this study, we study the effect of crystallographic orientation and applied triaxiality on the growth of intragranular voids. Two 3D full-field micromechanics methods are used, the dilatational visco-plastic fast-Fourier transform (DVP-FFT) and the crystal plasticity Finite Elements (CP-FE), both of which incorporate a combination of crystalline plasticity and dilatational plasticity. We demonstrate with several select cases that predictions of void growth from both formulations agree qualitatively. With the more computationally efficient DVP-FFT, additional effects of polycrystalline microstructure and the influence of nearest neighborhood are investigated. Crystals bearing a single intracrystalline void are studied in three types of 3D microstructural environments: isolated single crystals, individual equal-sized grains within a regular polycrystal, and individual variable sized grains within a polycrystal with grains and voids randomly located. We show that loading type plays a significant role. In strain-rate controlled conditions, voids in the hardest [111]-crystals grow the fastest in time, whereas in stress-controlled conditions, voids in the softest [100]-crystal grow the fastest in time. The analysis reveals that on average void growth is slower for the same starting orientation in the polycrystal than in the single crystal. We find that at the highest triaxiality tested that the correlation between crystal orientation and void growth rate in the polycrystal strengthens, drawing closer to that seen in the isolated single crystals. These results and model can help guide the microstructural design of polycrystalline materials with high strength and damage-tolerance in high-rate deformation.

36 MATERIALS SCIENCE↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

An Empirical Study on the Use of the Rancor Microworld Simulator to Support Full-scope Data Collection

A lack of data has been identified as a major challenge in human reliability analysis (HRA). Accordingly, several institutes and researchers have tried to collect HRA data from different data sources such as actual historical measurements, expert judgements, or simulator studies. While the most recent studies predominantly focus on collecting data using full-scope simulators with actual operators, Idaho National Laboratory (INL) has begun to collect HRA data using a simplified simulator, i.e., the Rancor Microworld simulator, with student participants. Full-scope studies have been known to have several intrinsic challenges to securing enough quantity of the data due to many reasons like the high cost for performing experiments or requiring actual operators’ cooperation. The ultimate goal of INL’s effort aims to infer actual operators’ data collected from a full-scope simulator on the basis of microworld data with student subjects as well as collect additional data that could be missed in the full-scope research. As a first step to achieve this goal, this paper projects an experimental plan for investigating the differences in human performance between individuals in two groups: 1) an actual operator and 2) a student when using the Rancor Microworld simulator. A randomized factorial experiment design has been developed with two independent variables, i.e., type of scenario and type of subject. Six human performance measures, i.e., 1) time, 2) error, 3) workload, 4) situation awareness, 5) patterns of attention and 6) the number of manipulations were selected. A couple of scenarios and their procedures available to the Rancor Microworld simulator have been developed.

99 GENERAL AND MISCELLANEOUS↗

Brief communication: Monitoring snow depth using small, cheap, and easy-to-deploy snow–ground interface temperature sensors

Abstract. Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We trained a random forest machine learning model to predict snow depth from variability in snow–ground interface temperature. The model performed well on Alaska's Seward Peninsula where it was trained and at Arctic evaluation sites (RMSE ≤ 0.15 m). It performed poorly at temperate sites with deeper snowpacks, partially due to training data limitations. Small temperature sensors are cheap and easy to deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring at high latitudes to an extent previously infeasible.

54 ENVIRONMENTAL SCIENCES↗

A semi-agnostic ansatz with variable structure for variational quantum algorithms

Quantum machine learning—and specifically Variational Quantum Algorithms (VQAs)—offers a powerful, flexible paradigm for programming near-term quantum computers, with applications in chemistry, metrology, materials science, data science, and mathematics. Here, one trains an ansatz, in the form of a parameterized quantum circuit, to accomplish a task of interest. However, challenges have recently emerged suggesting that deep ansatzes are difficult to train, due to flat training landscapes caused by randomness or by hardware noise. This motivates our work, where we present a variable structure approach to build ansatzes for VQAs. Our approach, called VAns (Variable Ansatz), applies a set of rules to both grow and (crucially) remove quantum gates in an informed manner during the optimization. Consequently, VAns is ideally suited to mitigate trainability and noise-related issues by keeping the ansatz shallow. We employ VAns in the variational quantum eigensolver for condensed matter and quantum chemistry applications, in the quantum autoencoder for data compression and in unitary compilation problems showing successful results in all cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Effects of forest structural and compositional change on forest microclimates across a gradient of disturbance severity

Forest structural diversity and community composition are key in regulating forest microclimates. When disturbance affects structural diversity or composition, forest microclimates may be altered due to changes in soil temperature, soil water content, and light availability. It is unclear however which structural or compositional components, when changed or to what extent, result in microclimatic change. To address this question, we used data from a large scale, manipulative stem-girdling experiment in northern, lower Michigan—the Forest Resilience and Threshold Experiment (FoRTE). FoRTE follows a factorial design with multiple levels of disturbance severity (0, 45, 65, 85%) based on targeted reductions in gross leaf area index via stem-girdling induced mortality. These disturbance severity treatments are applied in two ways: either as top-down (largest trees are killed) or bottom-up (small to medium trees killed) treatments. We examined how multiple components of structural diversity and community composition changed as a product of disturbance severity and type, and then tested for resulting effects on forest microclimates (light availability, soil temperature, and soil water), using a multivariate, Random Forest framework. We found that measures of community composition (species richness, species evenness, and Shannon-Wiener Diversity Index) and stand structure (basal area, standard deviation of DBH, tree size diversity) declined more following disturbance than did measures of canopy cover, heterogeneity, arrangement, or height. However, when changes in each variable from pre- to post-disturbance, measured as log change, were employed in a multivariate, Random Forest regression framework, structural diversity measures of heterogeneity (rugosity, top rugosity), cover (canopy cover), and arrangement (porosity) were the most influential variables, but with differences among bottom-up and top-down treatments We found that the death of large trees from disturbance impacts soil temperature, water, and light environments more substantially and uniformly across disturbance gradients than does the death of smaller trees. Furthermore, our results have implications for both statistical and process-based modeling of forest disturbance.

54 ENVIRONMENTAL SCIENCES↗

Ensemble Estimation of Historical Evapotranspiration for the Conterminous U.S.

Abstract Evapotranspiration (ET) is the largest component of the water budget, accounting for the majority of the water available from precipitation. ET is challenging to quantify because of the uncertainties associated with the many ET equations currently in use, and because observations of ET are uncertain and sparse. In this study, we combine information provided by available ET data and equations to produce a new monthly data set for ET for the conterminous U.S. (CONUS). These maps are produced from 1895 to 2018 at an 800 m spatial scale, marking a finer resolution than currently available products over this time period. In our approach, the relative performance of a suite of ET equations is assessed using water balance, flux tower, and remotely sensed ET estimates. At the observation locations, we use error distributions to quantify relative weights for the equations and use these in a modified Bayesian model averaging weighted ensemble approach. The relative weights are spatially generalized using a random forest regression, which is applied to wall‐to‐wall explanatory variable maps to generate CONUS‐wide relative weight maps and ensemble estimates. We assess the performance of the ensemble using a reserved subset of the observations and compare this performance against other national‐scale map products for historical to modern ET. The ensemble ET maps are shown to provide an improved accuracy over the alternative comparison products. These ET maps could be useful for a variety of hydrologic modeling and assessment applications that benefit from a long record, such as the study of periods of water scarcity through time.

Environmental Sciences & Ecology↗