Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

How Should Machine Learning Be Successfully Used for Wind Speed Vertical Extrapolation?

An accurate characterization of the wind resource available at hub-height is required for an efficient and bankable wind farm project. However, direct measurement of wind speed at the constantly increasing height of the hub of commercial wind turbines is oftentimes challenging and expensive, so that it is common practice to vertically extrapolate the wind resource from lower and more easily accessible levels. Conventional techniques for wind speed vertical extrapolation include the use of a power law and a logarithmic profile. While simple, the limits in accuracy of these methods have been shown in various studies. Recently, machine learning has been proposed as a new method to vertically extrapolate winds. All the published studies on the topic assess the performance of machine learning techniques in vertically extrapolating the wind resource at the same location where the algorithm has been trained. However, in real-world applications, the wind resource is measured at the instrument location, but it then needs to be extrapolated at hub height at the location of the wind turbines within the find farm. To be able to fully recommend the use of machine learning techniques over the simple power law and logarithmic law, the spatial variability of the performance improvements of the machine learning approaches needs to be assessed. Here, we propose a round-robin validation of a machine learning-based method for wind speed extrapolation. We use 20 months of observations at four locations spanning a 100 km wide region at the Southern Great Plains (SGP) atmospheric observatory, in north-central Oklahoma. At each location, we train a random forest to predict 30-min average wind speed at 143 m AGL. We use as input features lidar wind speed at 65 m AGL, time of day, sonic anemometer wind speed at 4 m AGL, turbulent kinetic energy, and Obukhov length. First, we perform a same-site comparison of the performance of the proposed random forest against the conventional techniques for wind speed extrapolation (namely power law and logarithmic profile, with widely accepted stability corrections). We find that the random forest outperforms the power law in vertically extrapolating wind speed in all the considered stability regimes, with a 33% reduction in MAE for stable conditions, and a 31% reduction in unstable conditions. Similar results are found when comparing predictions of extrapolated winds from the logarithmic profile and the random forest with the observed values. Next, we propose a round-robin validation, to use the random forest trained at each site to extrapolate wind speed at the remaining three sites. We find that the performance of the random forest approach degrades when the algorithm is tested at a site different than the training one. However, even under those circumstances, the machine learning-based approach still outperforms the conventional techniques for wind speed extrapolation, with, on average, a reduction in mean absolute error between 15 and 20% over the conventional methods, with the largest benefits obtained under stable conditions.

Monte Carlo↗

Surface temperatures reveal the patterns of vegetation water stress and their environmental drivers across the tropical Americas

Vegetation is a key component in the global carbon cycle as it stores ~450 GtC as biomass, and removes about a third of anthropogenic CO2 emissions. However, in some regions, the rate of plant carbon uptake is beginning to slow, largely because of water stress. Here, we develop a new observation-based methodology to diagnose vegetation water stress and link it to environmental drivers. We used the ratio of remotely sensed land surface to near surface atmospheric temperatures (LST/Tair) to represent vegetation water stress, and built regression tree models (random forests) to assess the relationship between LST/Tair and the main environmental drivers of surface energy fluxes in the tropical Americas. We further determined ecosystem traits associated with water stress and surface energy partitioning, pinpointed critical thresholds for water stress, and quantified changes in ecosystem carbon uptake associated with crossing these critical thresholds. Additionally, we found that the top drivers of LST/Tair, explaining over a quarter of its local variability in the study region, are (1) radiation, in 58% of the study region; (2) water supply from precipitation, in 30% of the study region; and (3) atmospheric water demand from vapor pressure deficits (VPD), in 22% of the study region. Regions in which LST/Tair variation is driven by radiation are located in regions of high aboveground biomass or at high elevations, while regions in which LST/Tair is driven by water supply from precipitation or atmospheric demand tend to have low species richness. Carbon uptake by photosynthesis can be reduced by up to 80% in water-limited regions when critical thresholds for precipitation and air dryness are exceeded simultaneously, that is, as compound events. Our results demonstrate that vegetation structure and diversity can be important for regulating surface energy and carbon fluxes over tropical regions.

59 BASIC BIOLOGICAL SCIENCES↗

Use of Satellite, Surface Observations and Numerical Weather Prediction Model Data to Improve Cloud Base Height and Cloud Base Vertical Velocity Estimation

Cloud base height (CBH) and cloud base vertical velocity (CBVV) are important variables that impact the overall climate in a region as they influence the formulation, longevity, and evolution of clouds. Retrieval of both parameters have long used ground instrumentation (e.g., Doppler lidar (DL), ground base radar); however, retrieving CBH from satellites is particularly challenging given that space-based instruments only observe cloud tops. In this manuscript, CBH is retrieved using a multi-linear regression equation, while CBVV used a random forests model. Both retrievals combine satellite and numerical weather prediction data. The satellite data used are the Visible Infrared Imaging Radiometer Suite imagery, while measurements of CBH and CBVV include DL and radiosonde data at the Southern Great Plains (SGP) Atmospheric Radiation Measurement observatory. Data from 83 summer days (May-August) in 2018–2021 featuring cumulus clouds forced by solar heating were examined and used to train the models, with years 2022–2023 used for validation. Various spatial domains were defined with one large (2.4° longitude by 2.0° latitude) SGP domain being split into smaller sections (smallest being 0.99° and 0.61° longitude and latitude respectably). CBH and CBVV values obtained from the DL as compared to the models show root mean square errors between 150 and 200 m, with CBVV values between 0.45 and 1 ms -1 . Finally, it was found that the CBH formulation performs well over all domains, while the CBVV retrievals become less accurate due to more turbulence being introduced into the observations as the number of DL stations decreases in the smaller domains.

54 ENVIRONMENTAL SCIENCES↗

The hyperplane of early-type galaxies: using stellar population properties to increase the precision and accuracy of the fundamental plane as a distance indicator

ABSTRACT We use deep spectroscopy from the SAMI (Sydney-AAO Multi-object Integral) Galaxy Survey to explore the precision of the fundamental plane (FP) of early-type galaxies as a distance indicator for future single-fibre spectroscopy surveys. We study the optimal trade-off between sample size and signal-to-noise ratio (SNR), and investigate which additional observables can be used to construct hyperplanes with smaller intrinsic scatter than the FP. We add increasing levels of random noise (parametrized as effective exposure time) to the SAMI spectra to study the effect of increasing measurement uncertainties on the FP- and hyperplane-inferred distances. We find that, using direct-fit methods, the values of the FP and hyperplane best-fitting coefficients depend on the spectral SNR, and reach asymptotic values for a mean $\langle \mathrm{ SNR} \rangle =40\, \mathrm{\mathring{\rm A}}^{-1}$. As additional variables for the FP we consider three stellar-population observables: light-weighted age, stellar mass-to-light ratio, and a novel combination of Lick indices ($I_\mathrm{age}$). For an $\langle \mathrm{ SNR} \rangle =45~\mathrm{\mathring{\rm A}}^{-1}$ (equivalent to 1-h exposure on a 4-m telescope), all three hyperplanes outperform the FP as distance indicators. Being an empirical spectral index, $I_\mathrm{age}$ avoids the model-dependent uncertainties and bias underlying age and mass-to-light ratio measurements, yet yields a 10 per cent reduction of the median distance uncertainty compared to the FP. We also find that, as a by-product, the $I_\mathrm{age}$ hyperplane removes most of the reported environment bias of the FP. After accounting for the different SNR, these conclusions also apply to a 50 times larger sample from SDSS-III (Sloan Digital Sky Survey). However, in this case, only $\mathrm{ age}$ removes the environment bias.

D’Eugenio, Francesco (ORCID:0000000323888172)↗

Computational materials reliability assessment of hydrogen fueled gas turbine power generation engines

The use of blended fuel sources in land based gas turbine engines drives variations in the resulting operational profile (temperatures and pressures) which can impact engine reliability. Furthermore, variability in the manufacture of components affects the resulting microstructure which directly impacts material performance and reliability. Currently, data-driven models are typically used for maintaining and inspecting fleets of engines. Without explicitly capturing material and operational sources of variability conservatism must be used in developing component-level reliability models. Therefore, there exists an opportunity to use information from materials-scale physics models to better inform reliability modeling and reduce conservatism; the impact is more cost-efficient operation and maintenance of current and future fleets. Specifically, this work establishes a computational framework for evaluating the probabilistic high temperature creep performance of hot-section Ni-based superalloys where uncertainty comes from both microstructural and operational variability. A novel high-fidelity physics model which phenomenologically captures grain-boundary sensitive phenomena has been established. A probabilistic calibration procedure was used to calibrate the model and capture uncertainty in the parameterized model coefficients. A design of experiments methodology was established for identifying informative microstructural digital representations for suitable for forward model evaluation. Results show that training a machine-learning surrogate using this design criteria outperforms random selection of microstructural representations. Finally, two surrogate models were developed: (1) a deterministic surrogate model which predicts the local field response given microstructure, constitutive model parameters, and operating conditions (stress, temperature) and (2) a probabilistic model, where uncertainty comes from constitutive law uncertainty, built using denoising diffusion probabilistic models which samples responses given (1) microstructure and (2) operating conditions. These surrogate models enable partner Siemens Energy to rapidly perform UQ analysis specific to creep deformation across a range of microstructures and operating conditions. The impact is that these ML and physics codes can be used to establish more advanced reliability models for the inspection, servicing, and maintenance of land based gas turbine engines.

36 MATERIALS SCIENCE↗

Range Hood Use and Effectiveness in Reducing Indoor Air Pollution During Gas and Induction Cooking

The Cooking Energy and Ventilation Impacts on Children's Asthma (CEVICA) study measured cooking frequency, range hood use, indoor air quality and respiratory health indicators of children with asthma living in homes with gas stoves in California's San Joaquin Valley. The study installed electric induction stoves and repeated measurements over three 2-week intensive periods, at baseline and at the end of two consecutive 3-month study phases. Stove replacements occurred at the start of Phase 1 or Phase 2 by random assignment. There were 4184 cooking events identified by automated analysis of time-series data from temperature sensors mounted above the cooktops and 1038 related range hood usage events detected from data recorded by anemometers, smart plugs, or motor loggers. Analysis of 1-minute resolved PM2.5 and NO 2 data identified and quantified 2685 PM 2.5 events and 2606 NO 2 events. Range hood use was characterized as a binary variable (>3 min vs. <3 min use). Range hood use was more common during cooking events associated with particle emissions and longer cooking durations. PM 2.5 concentrations during events with range hood use were comparable to those without use, which could result from limited effectiveness or if range hoods were preferentially used during higher-emission cooking scenarios. In homes with gas cooking, integrated NO 2 concentrations were about 45 percent higher during cooking events with no range hood use compared to those range hood use. The lowest pollutant levels were observed when the range hood operated for more than half of the cooking duration. These findings show that operation of venting range hood during cooking can substantially reduce short-term indoor exposure NO 2 in homes with gas cooking.

Fang, Yi↗

A Parameterization of the Cloud Scattering Polarization Signal Derived From GPM Observations for Microwave Fast Radative Transfer Models

Microwave cloud polarized observations have shown the potential to improve precipitation retrievals since they are linked to the orientation and shape of ice habits. Stratiform clouds show larger brightness temperature (TB) polarization differences (PDs), defined as the vertically polarized TB (TBV) minus the horizontally polarized TB (TBH), with ~10 K PD values at 89 GHz due to the presence of horizontally aligned snowflakes, while convective regions show smaller PD signals, as graupel and/or hail in the updraft tend to become randomly oriented. The launch of the global precipitation measurement (GPM) microwave imager (GMI) has extended the availability of microwave polarized observations to higher frequencies (166 GHz) in the tropics and midlatitudes, previously only available up to 89 GHz. This study analyzes one year of GMI observations to explore further the previously reported stable relationship between the PD and the observed TBs at 89 and 166 GHz, respectively. The latitudinal and seasonal variability is analyzed to propose a cloud scattering polarization parameterization of the PD-TB relationship, capable of reconstructing the PD signal from simulated TBs. Given that operational radiative transfer (RT) models do not currently simulate the cloud polarized signals, this is an alternative and simple solution to exploit the large number of cloud polarized observations available. Finally, the atmospheric radiative transfer simulator (ARTS) is coupled with the weather research and forecasting (WRF) model, in order to apply the proposed parameterization to the RT simulated TBs and hence infer the corresponding PD values, which show to reproduce the observed GMI PDs well.

54 ENVIRONMENTAL SCIENCES↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

Causal relationship between mitochondrial-associated proteins and cerebral aneurysms: a Mendelian randomization study

Background Cerebral aneurysm is a high-risk cerebrovascular disease with a poor prognosis, potentially linked to multiple factors. This study aims to explore the association between mitochondrial-associated proteins and the risk of cerebral aneurysms using Mendelian randomization (MR) methods. Methods We used GWAS summary statistics from the IEU Open GWAS project for mitochondrial-associated proteins and from the Finnish database for cerebral aneurysms (uIA, aSAH). The association between mitochondrial-associated exposures and cerebral aneurysms was evaluated using MR-Egger, weighted mode, IVW, simple mode and weighted median methods. Reverse MR assessed reverse causal relationship, while sensitivity analyses examined heterogeneity and pleiotropy in the instrumental variables. Significant causal relationship with cerebral aneurysms were confirmed using FDR correction. Results Through MR analysis, we identified six mitochondrial proteins associated with an increased risk of aSAH: AIF1 (OR: 1.394, 95% CI: 1.109–1.752, p = 0.0044), CCDC90B (OR: 1.318, 95% CI: 1.132–1.535, p = 0.0004), TIM14 (OR: 1.272, 95% CI: 1.041–1.553, p = 0.0186), NAGS (OR: 1.219, 95% CI: 1.008–1.475, p = 0.041), tRNA PusA (OR: 1.311, 95% CI: 1.096–1.569, p = 0.003), and MRM3 (OR: 1.097, 95% CI: 1.016–1.185, p = 0.0175). Among these, CCDC90B, tRNA PusA, and AIF1 demonstrated a significant causal relationship with an increased risk of aSAH (FDR q < 0.1). Three mitochondrial proteins were associated with an increased risk of uIA: CCDC90B (OR: 1.309, 95% CI: 1.05–1.632, p = 0.0165), tRNA PusA (OR: 1.306, 95% CI: 1.007–1.694, p = 0.0438), and MRM3 (OR: 1.13, 95% CI: 1.012–1.263, p = 0.0303). In the reverse MR study, only one mitochondrial protein, TIM14 (OR: 1.087, 95% CI: 1.004–1.177, p = 0.04), showed a causal relationship with aSAH. Sensitivity analysis did not reveal heterogeneity or pleiotropy. The results suggest that CCDC90B, tRNA PusA, and MRM3 may be common risk factors for cerebral aneurysms (ruptured and unruptured), while AIF1 and NAGS are specifically associated with an increased risk of aSAH, unrelated to uIA. TIM14 may interact with aSAH. Conclusion Our findings confirm a causal relationship between mitochondrial-associated proteins and cerebral aneurysms, offering new insights for future research into the pathogenesis and treatment of this condition.

Wang, Shuai↗

Significance of Low‐Velocity Zones on Solute Retention in Rough Fractures

Natural fractures are characterized by high internal heterogeneity. This internal variability is the cause of flow channeling, which in turn leads to contaminant transport taking place primarily along the high-velocity channels. Mass exchange between the high-velocity channels and the low-velocity zones has the potential to enhance contaminant retention, due to solute diffusion into the low-velocity zones and subsequent exposure to additional surface area for diffusion into the bordering rock matrix. Here, we derive a random walk particle tracking method for heterogeneous fractures, which includes an additional term to account for the aperture gradient. The method takes into account advection, diffusion in the fracture and matrix diffusion. The developed numerical framework is applied to assess the effect of low-velocity zones in rough self-affine fractures. The results show that diffusion into low-velocity zones has a visible but modest impact on contaminant retention. The magnitude of this impact does not change considerably, regardless of whether diffusion into the rock matrix is considered in the model, and increases for a decreasing average Péclet number of the fracture.

58 GEOSCIENCES↗

Global variation in the fraction of leaf nitrogen allocated to photosynthesis

Plants invest a considerable amount of leaf nitrogen in the photosynthetic enzyme ribulose-1,5-bisphosphate carboxylase-oxygenase (RuBisCO), forming a strong coupling of nitrogen and photosynthetic capacity. Variability in the nitrogen-photosynthesis relationship indicates different nitrogen use strategies of plants (i.e., the fraction nitrogen allocated to RuBisCO; fLNR), however, the reason for this remains unclear as widely different nitrogen use strategies are adopted in photosynthesis models. Here, we use a comprehensive database of in situ observations, a remote sensing product of leaf chlorophyll and ancillary climate and soil data, to examine the global distribution in fLNR using a random forest model. We find global fLNR is 18.2 ± 6.2%, with its variation largely driven by negative dependence on leaf mass per area and positive dependence on leaf phosphorus. Some climate and soil factors (i.e., light, atmospheric dryness, soil pH, and sand) have considerable positive influences on fLNR regionally. This study provides insight into the nitrogen-photosynthesis relationship of plants globally and an improved understanding of the global distribution of photosynthetic potential.

54 ENVIRONMENTAL SCIENCES↗

An Open‐Source, Physics‐Based, Tropical Cyclone Downscaling Model With Intensity‐Dependent Steering

Abstract An open‐source, physics‐based tropical cyclone (TC) downscaling model is developed, in order to generate a large climatology of TCs. The model is composed of three primary components: (a) a random seeding process that determines genesis, (b) an intensity‐dependent beta‐advection model that determines the track, and (c) a non‐linear differential equation set that determines the intensification rate. The model is entirely forced by the large‐scale environment. Downscaling ERA5 reanalysis data shows that the model is generally able to reproduce observed TC climatology, such as the global seasonal cycle, genesis locations, track density, and lifetime maximum intensity distributions. Inter‐annual variability in TC count and power‐dissipation is also well captured, on both basin‐wide and global scales. Regional TC hazard estimated by this model is also analyzed using return period maps and curves. In particular, the model is able to reasonably capture the observed return period curves of landfall intensity in various sub‐basins around the globe. The incorporation of an intensity‐dependent steering flow is shown to lead to regionally dependent changes in power dissipation and return periods. Advantages and disadvantages of this model, compared to other downscaling models, are also discussed.

Meteorology & Atmospheric Sciences↗

Knowledge of lactation amenorrhea method among postpartum women in Ethiopia: a facility-based cross-sectional study

While the importance of knowledge about contraceptives in improving their utilization and thereby reducing the risk of unintended pregnancies is well documented, there are limited studies documented about the Lactational Amenorrhea Method (LAM). Thus, understanding the knowledge of postpartum mothers about LAM is essential for designing tailored interventions. This study assessed the level of knowledge about LAM and its associated factors among postpartum mothers in Ethiopia. A facility-based cross-sectional study was conducted among 3148 randomly selected postpartum participants. The study utilized multistage sampling approach in hospitals located across five regions and one city administration in Ethiopia. Data were collected using face-to-face interviews at discharge. A participant was categorized as having knowledge of LAM if she correctly answered the three LAM criteria: amenorrhea, the first 6 months, and exclusive breast feeding. A binary logistic regression model was used to identify factors associated with knowledge of LAM. Variables with p < 0.25 in the binary logistic regression were included in the multiple logistic regression. Then, associations were described using the adjusted odds ratio (AOR) along with the 95% confidence interval (CI), and statistical significance was declared at p < 0.05. Only four in 10 participants (40.6%; 95% CI 38.9–42.3) had knowledge of LAM. Participants who attended college or above educational level (AOR = 2.1, 95% CI 1.5–2.8), those with parity of two (AOR = 2.3; 95% CI 1.6–3.6) or more than two (AOR = 2.4; 95% CI 1.5–4.0), those who expressed a desire for further fertility (AOR = 1.3; 95% CI 1.1–1.5), individuals who received counselling on LAM (AOR = 3.0; 95% CI 2.6–3.7), and those who gave birth in hospital (AOR = 2.6; 95% CI 1.4–2.6) had higher odds of knowledge about LAM, compared to their counter parts. In contrary, participants resided far away from health facilities had 30% lower odd of knowledge about LAM compared to those resided near the health facilities (AOR = 0.70; 95% CI 0.6–0.8). The proportion of participants who had knowledge of LAM was low. Strengthening counseling about LAM during antenatal care and delivery with due attention to women with limited access to health facilities should be considered for increasing their level of knowledge on LAM.

60 APPLIED LIFE SCIENCES↗

Analysis of Random Forest Modeling Strategies for Multi-Step Wind Speed Forecasting

Although the random forest (RF) model is a powerful machine learning tool that has been utilized in many wind speed/power forecasting studies, there has been no consensus on optimal RF modeling strategies. This study investigates three basic questions which aim to assist in the discernment and quantification of the effects of individual model properties, namely: (1) using a standalone RF model versus using RF as a correction mechanism for the persistence approach, (2) utilizing a recursive versus direct multi-step forecasting strategy, and (3) training data availability on model forecasting accuracy from one to six hours ahead. These questions are investigated utilizing data from the FINO1 offshore platform and Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) C1 site, and testing results are compared to the persistence method. At FINO1, due to the presence of multiple wind farms and high inter-annual variability, RF is more effective as an error-correction mechanism for the persistence approach. The direct forecasting strategy is seen to slightly outperform the recursive strategy, specifically for forecasts three or more steps ahead. Finally, increased data availability (up to ~8 equivalent years of hourly training data) appears to continually improve forecasting accuracy, although changing environmental flow patterns have the potential to negate such improvement. We hope that the findings of this study will assist future researchers and industry professionals to construct accurate, reliable RF models for wind speed forecasting.

54 ENVIRONMENTAL SCIENCES↗

Projected U.S. drought extremes through the twenty-first century with vapor pressure deficit

Global warming is expected to enhance drought extremes in the United States throughout the twenty-first century. Projecting these changes can be complex in regions with large variability in atmospheric and soil moisture on small spatial scales. Vapor Pressure Deficit (VPD) is a valuable measure of evaporative demand as moisture moves from the surface into the atmosphere and a dynamic measure of drought. Here, VPD is used to identify short-term drought with the Standardized VPD Drought Index (SVDI); and used to characterize future extreme droughts using grid dependent stationary and non-stationary generalized extreme value (GEV) models, and a random sampling technique is developed to quantify multimodel uncertainties. The GEV analysis was performed with projections using the Weather Research and Forecasting model, downscaled from three Global Climate Models based on the Representative Concentration Pathway 8.5 for present, mid-century and late-century. Results show the VPD based index (SVDI) accurately identifies the timing and magnitude short-term droughts, and extreme VPD is increasing across the United States and by the end of the twenty-first century. The number of days VPD is above 9 kPa increases by 10 days along California’s coastline, 30–40 days in the northwest and Midwest, and 100 days in California’s Central Valley.

54 ENVIRONMENTAL SCIENCES↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Marine Hydrokinetic Tidal Turbine

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset is part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with other energy technologies. This dataset contains inputs and outputs from simulations of a floating marine hydrokinetic turbine over approximately half a tidal cycle (~6.6 hours). Inflow conditions were derived from field measurements in Alaska’s Cook Inlet and represent a tidal environment in which the current speed ramps from near 0 m/s to a peak of 3 m/s and back. The original acoustic doppler current profiler dataset is publicly available on the Marine and Hydrokinetic Data Repository. In a full tidal cycle, the flow reverses and the rotor would reorient; this reversal was not modeled. In the Cook Inlet campaign , turbulence intensity was similar in both directions. Two inflow cases are included. In the first case, labeled “raw” in the files, the measured current time series was used directly in the InflowWind module of OpenFAST. Speed and direction were applied as a function of time and elevation, uniformly in the horizontal direction. With full spatial coherence, this approach captures high turbulent variability and results in pronounced power fluctuations, so it is considered a conservative, near-worst-case representation of loading. In the second case, labeled “average” in the files, a 30-minute moving average was applied to extract the slowly varying mean speed. The residual fluctuations about this mean were used to generate spatially varying, full-field turbulence inputs with TurbSim, giving a more physically realistic representation of the inflow across the rotor disk. Two random realizations were used to produce distinct inflow conditions for two OpenFAST simulations representing a two-turbine array. The same turbulence intensity is applied across the full time series, producing larger fluctuations at the start and end, where the mean speed is low. The second case is the more appropriate framework for performance and power assessment but overpredicts turbulence at lower flow speeds and underpredicts it at higher speeds. As the floating platform moves and the rotor changes its x-position, Taylor’s frozen turbulence hypothesis used by InflowWind assumes a constant rather than a time-varying mean velocity, introducing some inaccuracy in the velocity plane sampling. The turbine modeled is the 500-kW Reference Model 1, a horizontal-axis two-bladed hydrokinetic turbine on a four-column floating semisubmersible substructure . Simulations were performed using OpenFAST v4.1 with the Reference Open Source Controller (ROSCO) v2.10. All input files required to reproduce the simulations are included. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel . This unit supports up to 2.5 MW, but NLR has only a single 1.25-MW stack. The datasets report hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. The system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operating current of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The simulated tidal turbine time series data was translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz. Each zip file represents a single tidal electrolysis experiment and is named: {technology}_{inflow method}_{number of 500 kW tidal turbines connected} For instance, “tidal-500kW-RM1_average_2.zip” is a 6-hour experiment using the 500-kW tidal reference model, scaled by 2x (1-MW) to better match the electrolyzer maximum of 1.25MW, fed with the 30-minute moving average current case. Each zip folder contains the following files: A .csv file of raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. A .csv file combines all tidal profiles as "combined_tidal_experiments.csv." A separate experiment, “characterization_200.zip,” shows the MC250 electrolyzer steady-state response with 30-minute load steps over 5 hours and is accessible with this entry.

08 HYDROGEN↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Reconstruction of Six-Dimensional Phase Space

A phase space is a mathematical representation of all possible physical states of a system. Particle beams at Fermilab exist within a six-dimensional (6D) phase space defined by three positional components, (x, y, z) and three momentum components, (px, py, pz). To reconstruct this space implies taking measurement data from detectors and mapping out particle behavior using computational methods. The beam detectors, however, are only able to detect spatial distribution among the events of the beam, therefore being limited to positional data. Also, due to the vast number of events in a particle beam, it is extremely difficult to analyze and differentiate every single one’s behavior. However, with Machine Learning (ML), which can distinguish between patterns and map out particle behavior more efficiently. We first used the particle beam software, G4beamline, to simulate a 10,000-event muon beam, adjusting parameters such as initial momentum magnitude (p¬0) and virtual detector position. Using ten virtual detectors, we analyzed p0 values such that minimum 9,990 events were analyzed by every detector. We then input the data from these beam simulations to a C++ program, that randomly selects 100 events, and creates a 2D histogram based on spatial distribution, detector position, and event intensity. This process is repeated 100 times to create 100 histograms per p0 value. These images were then input to a modified ResNet18 Convolutional Neural Network (CNN) for training, and to predict p0 from some unseen set of histograms. The model was accurate when trained on momentum increments of 5 MeV/c and provided with denser training samples around highly variable test values. These results displayed machine learning being able to accurately predict p0 from being trained on different particle behaviors.

Shirlee, Jermain [Fermilab]↗