Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Dynamic Modeling of a Kaplan Hydroturbine Using Optimal Parametric Tuning and Real Plant Operational Data

To address grid variability caused by renewable energy integration and to maintain grid reliability and resilience, hydropower must quickly adjust its power generation over short time periods. This changing energy generation landscape requires advance technology integration and adaptive parameter optimization for hydropower systems via digital twin effort. However, this is difficult owing to the lack of characterization and modeling for the nonlinear nature of hydroturbines. To solve this issue, this paper first formulates a six-coefficient Kaplan hydroturbine model and then proposes a parametric optimization tuning framework based on the Nelder–Mead algorithm for adaptive dynamic learning of the six-coefficients so as to build models that describe the turbine. To assess the performance of the proposed optimal parametric tuning technique, operational data from a real-world Kaplan hydroturbine unit are collected and used to model the relationship between the gate opening and the generated power production. The findings show that the proposed technique can effectively and adaptively learn the unknown dynamics of the Kaplan hydroturbine while optimally tune the unknown coefficients to match the generated power output from the real hydroturbine unit with an inaccuracy of less than 5%. The method can be used to provides optimal tuning of parameters critical for controller design, operational optimization and daily maintenance for hydroturbines in general.

13 HYDRO ENERGY↗

Consistency-Enhanced Evolution for Variable Selection Can Identify Key Chemical Information from Spectroscopic Data

In the last few decades, spectroscopic techniques such as near-infrared (NIR) spectroscopy have gained wide applications in several industries, such as the pharmaceutical, agricultural, oil, and gas industries. As a result, various soft sensors have been developed to predict sample properties from spectroscopic readings. Because the spectroscopic readings at different wavelengths, especially at the adjacent wavelengths, are highly correlated, it has been shown that variable selection could significantly improve a soft sensor’s prediction performance while reducing the model complexity. To improve the prediction performance, most variable selection methods focus on identifying the variables (i.e., wavelengths or wavelength segments) that are strongly correlated with the dependent variable. Although many successful applications have been reported, these variable selection methods do have their limitations. Specifically, the selected wavelengths sometimes show little connection to the chemical bounds or functional groups presenting in the sample. In addition, the selected variables can be quite sensitive to the choice of the training samples. In this work, we address these limitations from a different perspective: if a variable selection algorithm can identify the truly relevant input variables, it should consistently identify the same subset of variables regardless of the choice of the training samples. Therefore, we propose a variable selection method that aims to improve the consistency of variable selection resulting from different training samples. Furthermore, the new algorithm is termed consistency-enhanced evolution for variable selection (CEEVS). To demonstrate the performance and robustness of CEEVS, we compare the proposed method with three representative variable selection methods using five published NIR data sets. These case studies clearly demonstrate that by improving the variable selection consistency, we can not only achieve improved prediction performance, but also identify key chemical information from spectroscopic data.

42 ENGINEERING↗

Review of Intrusion Detection Methods and Tools for Distributed Energy Resources

Recent trends in the growth of distributed energy resources (DER) in the electric grid and newfound malware frameworks that target internet of things (IoT) devices is driving an urgent need for more reliable and effective methods for intrusion detection and prevention. Cybersecurity intrusion detection systems (IDSs) are responsible for detecting threats by monitoring and analyzing network data, which can originate either from networking equipment or end-devices. Creating intrusion detection systems for PV/DER networks is a challenging undertaking because of the diversity of the attack types and intermittency and variability in the data. Distinguishing malicious events from other sources of anomalies or system faults is particularly difficult. New approaches are needed that not only sense anomalies in the power system but also determine causational factors for the detected events. In this report, a range of IDS approaches were summarized along with their pros and cons. Using the review of IDS approaches and subsequent gap analysis for application to DER systems, a preliminary hybrid IDS approach to protect PV/DER communications is formed in the conclusion of this report to inform ongoing and future research regarding the cybersecurity and resilience enhancement of DER systems.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Physical Insights From the Multidecadal Prediction of North Atlantic Sea Surface Temperature Variability Using Explainable Neural Networks

Abstract North Atlantic sea surface temperatures (NASST), particularly in the subpolar region, are among the most predictable in the world's oceans. However, the relative importance of atmospheric and oceanic controls on their variability at multidecadal timescales remain uncertain. Neural networks (NNs) are trained to examine the relative importance of oceanic and atmospheric predictors in predicting the NASST state in the Community Earth System Model 1 (CESM1). In the presence of external forcings, oceanic predictors outperform atmospheric predictors, persistence, and random chance baselines out to 25‐year leadtimes. Layer‐wise relevance propagation is used to unveil the sources of predictability, and reveal that NNs consistently rely upon the Gulf Stream‐North Atlantic Current region for accurate predictions. Additionally, CESM1‐trained NNs successfully predict the phasing of multidecadal variability in an observational data set, suggesting consistency in physical processes driving NASST variability between CESM1 and observations.

Geology↗

Learning Planar Ising Models Software

Learning Planar Ising Models is a software package written in Matlab for learning relationships among variable in a dataset using graphical models. The software package implements a generally-applicable algorithm for learning planar Ising models from any multivariate dataset. The code provides an algorithm for learning the best planar Ising model to approximate an arbitrary collection of binary random variables (possibly from sample data). Given the set of all pairwise correlations among variables, we select a planar graph and optimal planar Ising model defined on this graph to best approximate that set of correlations. The software includes demonstrations of the algorithm in simulations and for applications on publicly available datasets. Details of the algorithm, demonstration simulations, and applications are given in Johnson, et al; 2016. Reference: Johnson, J. K., Oyen, D., Chertkov, M., and Netrapalli, P. (2016). Learning planar Ising models. Journal of Machine Learning Research.

Oyen, Diane↗

Forest aboveground biomass estimation through integration of sentinel-2 and PALSAR-2 time series: assessing models trained on GEDI and field inventory benchmarks

Accurate and spatially explicit forest Aboveground Biomass (AGB) mapping through remote sensing is critical for quantifying terrestrial carbon stocks and informing effective forest management strategies. However, AGB estimation in dense forests with complex terrain remains challenging due to satellite sensor signal saturation problem (saturation issue occurs in high biomass forests), structural complexity, and limited ground truth for calibration. This study presents a novel framework that integrates multi-temporal Sentinel-2 optical imagery, ALOS PALSAR-2 Synthetic Aperture Radar (SAR) data, and topographic variables with explainable Machine Learning to map AGB across mountainous forests within subtropical and temperate oceanic climate zones of Mexico. We evaluate the effects of temporal granularity and sensor synergy by comparing multiple temporal inputs and sensor configurations (Sentinel-2, PALSAR-2, and their fusion), and assess model performance using two reference datasets: NASA GEDI LiDAR-derived biomass and Mexico’s National Forest and Soil Inventory (INFyS). Our results showed that models trained on INFyS consistently outperformed those trained on GEDI, highlighting limitations in GEDI’s reliability in biomass estimates within this study region. Furthermore, the integration of Sentinel-2 and PALSAR-2 provided improved predictions compared to single-sensor models, particularly when combined with temporally explicit yearly statistics. The best-performing model, which was trained on INFyS data, and considered both Sentinel-2 and PALSAR-2 yearly statistics, as well as topographic variables, achieved an R2 of 0.64, RMSE of 51.10 Mg/ha, and relative RMSE (rRMSE) of 58.69%. Explainable ML analysis identified Sentinel-2 spectral indices and topographic features as key predictors, while PALSAR-2 metrics provided complementary information, partially mitigating saturation effects in high-biomass areas. Specifically, integrating both sensors substantially improved AGB estimation in high biomass forest (≥200 Mg/ha), yielding 98% gains over optical-only model, with resulting estimates exceeding GEDI L4B by 29% and ESA-CCI-BIOMASS by 174%. Terrain-stratified analysis indicated close agreement with GEDI in low-slope areas, with increasing divergence as slope steepness increased, while estimates remained consistently higher than ESA-CCI-BIOMASS across all slope classes. The proposed approach advances multi-sensor fusion and temporal feature engineering for AGB mapping using open-access satellite datasets, providing a scalable and reproducible framework for annual biomass monitoring in topographically complex mountainous forests. The resulting 25 m resolution biomass product has the potential to provide spatially detailed information for forest monitoring and may support applications in carbon accounting and forest management.

54 ENVIRONMENTAL SCIENCES↗

Disentangling the Impact of the COVID‐19 Lockdowns on Urban NO 2 From Natural Variability

Abstract TROPOMI satellite data show substantial drops in nitrogen dioxide (NO 2 ) during COVID‐19 physical distancing. To attribute NO 2 changes to NO x emissions changes over short timescales, one must account for meteorology. We find that meteorological patterns were especially favorable for low NO 2 in much of the United States in spring 2020, complicating comparisons with spring 2019. Meteorological variations between years can cause column NO 2 differences of ~15% over monthly timescales. After accounting for solar angle and meteorological considerations, we calculate that NO 2 drops ranged between 9.2% and 43.4% among 20 cities in North America, with a median of 21.6%. Of the studied cities, largest NO 2 drops (>30%) were in San Jose, Los Angeles, and Toronto, and smallest drops (<12%) were in Miami, Minneapolis, and Dallas. These normalized NO 2 changes can be used to highlight locations with greater activity changes and better understand the sources contributing to adverse air quality in each city.

54 ENVIRONMENTAL SCIENCES↗

How Well Can CMIP6 Models Represent the Observed Influence of the Pacific and Indian Oceans on the Indian Summer Monsoon Rainfall?

This study evaluates the ability of CMIP6 climate models to simulate the observed effects of tropical Pacific and Indian Ocean sea surface temperature anomalies (SSTAs) on Indian summer monsoon rainfall (ISMR) variability. Using observational data and the large ensemble historical simulations of seven CMIP6 models from 1950 to 2014, we applied a cyclostationary linear inverse model (CS-LIM) to isolate the impacts of tropical Pacific SSTAs, Indian Ocean SSTAs and their interaction on the interannual variability of ISMR. Overall, CMIP6 models well reproduced the observed enhanced (reduced) ISMR variability from Pacific SSTAs (Indian Ocean SSTAs and the Indo-Pacific interaction), but with varying spatial patterns and magnitudes. While CESM2 and E3SM-2-0 showed the best agreement with observations for the effects of Pacific SSTAs and the Indo-Pacific interaction, respectively, CMIP6 models showed mixed results for the impacts from Indian Ocean SSTAs. Composite analysis of ISMR anomalies during the developing phases of pure and co-occurring El Niño-Southern Oscillation (ENSO) and Indian Ocean dipole (IOD) events revealed that the impacts from Pacific SSTAs were captured reasonably well by E3SM-2-0, CESM2, MIROC6, and MPI-ESM1-2-LR, while E3SM-2-0 also showed the best agreement with observations for the effects from the Indo-Pacific interaction. However, all models showed substantial biases in simulating the Indian Ocean SSTA impacts on ISMR, especially for pure El Niño events. Overall, this study provides new insights into how individual CMIP6 models simulate the isolated impacts from the tropical Pacific and Indian Oceans, which has important applications for improving ISMR predictions and interpreting ISMR future projections.

monsoon↗

Machine Learning Prediction of Tritium‐Helium Groundwater Ages in the Central Valley, California, USA

Abstract Groundwater ages provides insight into recharge rates, flow velocities, and vulnerability to contaminants. The ability to predict groundwater ages based on more accessible parameters via Machine Learning (ML) would advance our ability to guide sustainable management of groundwater resources. In this study, ML models were trained and tested on a large data set of tritium concentrations and tritium‐helium groundwater ages from the California Central Valley, a large groundwater basin with complex land use, irrigation, and water management practices. The ML models were trained on 63 features, including location, well construction information, landscape characteristics, and climate variables, water chemistry, and stable isotopes. The Bagging regressor method can accurately classify (F1‐score = 0.91) groundwater samples as either modern or pre‐modern whereas the accuracy of the ML prediction of continuous tritium‐helium groundwater ages is limited and explains only of the variability in this data set. In general, ML groundwater age prediction relies mostly on features related to (a) the source of groundwater recharge, (b) contaminant history, (c) aquifer materials, (d) well construction, and (e) geochemical reactions along flow paths.

54 ENVIRONMENTAL SCIENCES↗

Correlating processing variables to material properties in recycled polypropylene: A data‐driven approach

Abstract Polypropylene (PP) is one of the most widely used plastics, yet its recycling remains limited, with less than 1% of solid waste PP being reprocessed. Mechanical recycling through extrusion is the most practical method, but inconsistent reprocessing conditions introduce variability in material properties. While temperature, screw speed, and residence time influence the thermomechanical stress applied during reprocessing, there are no standardized guidelines for optimizing these parameters. This study examines how these factors shape the properties of recycled PP, using conditions designed to mimic post‐industrial recycled (PIR) scrap. Residence time was measured using colorimetric tracking and correlated with molecular weight, viscosity, and mechanical properties over multiple extrusion cycles. Data‐driven modeling, including response surface methodology, support vector machines, and artificial neural networks, identified processing temperature as the dominant factor in material degradation, followed by residence time. Mechanical properties remained stable, while viscosity decreased predictably with increasing residence time. By linking reprocessing conditions to property evolution, this study provides a method to optimize processing parameters and reduce variability in recycled PP. These findings help manufacturers improve process control, making recycled PP more predictable for reuse in manufacturing. Highlights Study of PIR‐quality PP without additives or compatibilizers. Residence time analysis shows processing temperature drives PP property changes. Mark‐Houwink enables quick molecular weight checks for quality control. Models predict mechanical and rheological shifts in reprocessing. Optimized processing parameters minimize property degradation in recycling.

Estela‐García, John E. [Polymer Engineering Center↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

Combining variational autoencoders and physical bias for improved microscopy data analysis *

Electron and scanning probe microscopy produce vast amounts of data in the form of images or hyperspectral data, such as electron energy loss spectroscopy or 4D scanning transmission electron microscope, that contain information on a wide range of structural, physical, and chemical properties of materials. To extract valuable insights from these data, it is crucial to identify physically separate regions in the data, such as phases, ferroic variants, and boundaries between them. In order to derive an easily interpretable feature analysis, combining with well-defined boundaries in a principled and unsupervised manner, here we present a physics augmented machine learning method which combines the capability of variational autoencoders to disentangle factors of variability within the data and the physics driven loss function that seeks to minimize the total length of the discontinuities in images corresponding to latent representations. Our method is applied to various materials, including NiO-LSMO, BiFeO 3 , and graphene. The results demonstrate the effectiveness of our approach in extracting meaningful information from large volumes of imaging data. The customized codes of the required functions and classes to develop phyVAE is available at https://github.com/arpanbiswas52/phy-VAE.

97 MATHEMATICS AND COMPUTING↗

Cost of Fish Exclusion and Passage Technologies for Hydropower

Hydropower represents a reliable source of renewable energy and accounts for approximately 7% of the total electrical generation in the United States. Future expansion of hydropower is likely to be in the form of either smaller new stream development projects or powering existing non-powered dams. For these new projects to be successful, careful analysis of risks, costs, and uncertainty to offset reduced power production as well as ensuring the protection and safe passage of migratory fish to gain public support, will be required. Exclusion and passage are two common approaches to protect fish from entrainment and impingement at hydropower facilities. The thresholds for entrainment risk and requirements for exclusion and passage often differ depending on the species involved, the characteristics of the facility, and the goals of stakeholders. While the costs associated with environmental mitigations represent a large proportion of the total costs required for the licensing of hydropower facilities, little quantitative information is present within the literature regarding the specific costs of fish exclusion and passage. Working with FOA awardee Natel Energy, scientists at Oak Ridge National Laboratory were tasked with assessing the capital construction costs for downstream fish exclusion and passage infrastructure. This report used keyword searches of an existing environmental mitigation cost data set and manual extraction of additional cost data associated with protection, mitigation, and enhancement (PM&E) measures related to positive barrier screening and passage from regulatory licensing documents available in the Federal Energy Regulatory Commission (FERC) eLibrary. This approach yielded a total of 50 PM&E mitigation measures with estimated capital construction costs pertaining to positive barrier screens, 142 pertaining to passage studies, and 26 pertaining to passage-related studies. PM&E measures associated with positive barrier screens represented <10% of the 171 total FERC project dockets available in the data set. These data were highly skewed toward conventional relicensing projects, as <7% were associated with new stream development (NSD) projects. Results from these data indicate highly variable costs associated with fish screening, with flow-normalized costs one to two orders of magnitude higher for screening with the highest exclusion capability (≤0.09 in. spacing) compared with coarser screening (1 to 2 in.). Furthermore, estimated capital costs of passage infrastructure were positively related to the scale of the project based on installed capacity for some, but not all, types of passage. These data provide an initial baseline for estimating exclusion and passage costs for hydropower development and may help developers consider options for more fish-friendly generation technologies, though gaps remain relating to a lack of data, particularly for NSD projects. More data may still be available within the FERC eLibrary, but significant effort will be required to manually identify and extract the data for future analyses.

13 HYDRO ENERGY↗

Merged Aerosol Value-Added Product Report

The Merged Aerosol Value-Added Product (VAP) simplifies scientists’ use of Atmospheric Radiation Measurement (ARM) User Facility aerosol data by performing several tedious, time-consuming tasks for the users. First, the VAP identifies the best data available when multiple datastreams exist for a single geophysical quantity so that ARM users do not have to research this for themselves. Second, the VAP consolidates multiple ARM aerosol datastreams into a single file for ARM data users so that they do not have to download, open, and read multiple files for their analysis. Next, the VAP transforms all measurements onto a common one-hour timestamp. The one-hour resolution matches the time resolution of the slowest instrument. Instruments with faster sampling rates than one measurement per hour are averaged over the time interval. Finally, the VAP reads the QA/QC variables and marks data with known issues as missing, so that users do not have to spend excessive time cleaning data. This includes incorporating Data Quality Reports (DQRs) that exist at the time when the VAP data is generated. DQRs are reports filed by instrument mentors or data users that indicate a problem with the output data of individual instruments.

54 ENVIRONMENTAL SCIENCES↗

Cross-national analysis of food security drivers: comparing results based on the Food Insecurity Experience Scale and Global Food Security Index

Abstract The second UN Sustainable Development Goal establishes food security as a priority for governments, multilateral organizations, and NGOs. These institutions track national-level food security performance with an array of metrics and weigh intervention options considering the leverage of many possible drivers. We studied the relationships between several candidate drivers and two response variables based on prominent measures of national food security: the 2019 Global Food Security Index (GFSI) and the Food Insecurity Experience Scale’s (FIES) estimate of the percentage of a nation’s population experiencing food security or mild food insecurity (FI ). We compared the contributions of explanatory variables in regressions predicting both response variables, and we further tested the stability of our results to changes in explanatory variable selection and in the countries included in regression model training and testing. At the cross-national level, the quantity and quality of a nation’s agricultural land were not predictive of either food security metric. We found mixed evidence that per-capita cereal production, per-hectare cereal yield, an aggregate governance metric, logistics performance, and extent of paid employment work were predictive of national food security. Household spending as measured by per-capita final consumption expenditure (HFCE) was consistently the strongest driver among those studied, alone explaining a median of 92% and 70% of variation (based on out-of-sample R 2 ) in GFSI and FI , respectively. The relative strength of HFCE as a predictor was observed for both response variables and was independent of the countries used for model training, the transformations applied to the explanatory variables prior to model training, and the variable selection technique used to specify multivariate regressions. The results of this cross-national analysis reinforce previous research supportive of a causal mechanism where, in the absence of exceptional local factors, an increase in income drives increase in food security. However, the strength of this effect varies depending on the countries included in regression model fitting. We demonstrate that using multiple response metrics, repeated random sampling of input data, and iterative variable selection facilitates a convergence of evidence approach to analyzing food security drivers.

42 ENGINEERING↗

Dynamically Downscaled (WRF) 1km, Hourly Meteorological Conditions 1987-2020. East/Taylor Watersheds

This dataset contains meteorological output from the Weather Research and Forecasting (WRF) version 3.8.1. This dataset has been created to 1) investigate hydrometeorological processes impacting the East River and water-delivery to the Critical Zone, and 2) provide meteorological forcing data for distributed Earth-science modeling applications in the East River watershed. Variables have a 1 kilometer spatial resolution and hourly temporal resolution and encompass a rectangular region encompassing the East and Taylor River watersheds, Colorado, near the town of Crested Butte. WRF was forced using Climate Forecast System Reanalysis (CFSR) lateral boundary conditions. Each .zip file contains one "water year" of data (October 1 -- September 30; i.e. water year 2017 starts October 1, 2016 and ends September 30, 2017). Each zip folder contains 12 netcdf (.nc) files containing one month of hourly data each and are approximately 250mb. Model timestamps are in UTC time.The files contain the following data variables:EAST_MASK: binary mask of the watershed regionTAYLOR_MASK: binary mask of the watershed regionGLW downwelling longwave radiation (w/m2)HR_PRCP: Hourly Precipitation Rate (mm/hr). Includes all hydrometeors (solid+liquid). HFX NoahMP LSM total grid-cell modelled sensible heat flux (w/m2) [positive towards atmosphere]LH NoahMP LSM total grid-cell modelled latent heat flux (w/m2) [positive towards atmosphere; can be converted to ET]PSFC Surface Barometric Pressure (hPa)Q2 Two-meter specific humidity (kg/kg)SWDOWN Downwelling shortwave solar radiation (w/m2)SWNORM Terrain-normal downwelling shortwave radiation (w/m2)T2 Two-meter air temperature (deg K)U10 10-m U-component of wind velocity (m/s)V10 10m V-component of wind velocity (m/s)XLAT Latitude of grid-center point XLONG Longitude of grid-center point XTIME Model timestamp, **in UTC**

54 ENVIRONMENTAL SCIENCES↗

Panta Rhei benchmark dataset: socio-hydrological data of paired events of floods and droughts

As the adverse impacts of hydrological extremes increase in many regions of the world, a better understanding of the drivers of changes in risk and impacts is essential for effective flood and drought risk management and climate adaptation. However, there is currently a lack of comprehensive, empirical data about the processes, interactions, and feedbacks in complex human–water systems leading to flood and drought impacts. Here we present a benchmark dataset containing socio-hydrological data of paired events, i.e. two floods or two droughts that occurred in the same area. The 45 paired events occurred in 42 different study areas and cover a wide range of socio-economic and hydro-climatic conditions. The dataset is unique in covering both floods and droughts, in the number of cases assessed and in the quantity of socio-hydrological data. The benchmark dataset comprises (1) detailed review-style reports about the events and key processes between the two events of a pair; (2) the key data table containing variables that assess the indicators which characterize management shortcomings, hazard, exposure, vulnerability, and impacts of all events; and (3) a table of the indicators of change that indicate the differences between the first and second event of a pair. The advantages of the dataset are that it enables comparative analyses across all the paired events based on the indicators of change and allows for detailed context- and location-specific assessments based on the extensive data and reports of the individual study areas. The dataset can be used by the scientific community for exploratory data analyses, e.g. focused on causal links between risk management; changes in hazard, exposure and vulnerability; and flood or drought impacts. The data can also be used for the development, calibration, and validation of socio-hydrological models. The dataset is available to the public through the GFZ Data Services (Kreibich et al., 2023, https://doi.org/10.5880/GFZ.4.4.2023.001).

54 ENVIRONMENTAL SCIENCES↗