Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data needs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

U.S. Manufacturing Water Use Data and Estimates: Current State, Limitations, and Future Needs for Supporting Manufacturing Research and Development

Water is essential to manufacturing operations; without it, many facilities could not operate or meet production demands. Physical, reputational, and regulatory risks to water supplies compounded by climate change-induced impacts on hydrological conditions threaten the adequacy of water supplies for manufacturing. Manufacturing water use has not been a major focus of either water or manufacturing-related research. Research and development (R&D) aimed at helping manufacturers use water more sustainably and adapt to changing water conditions is needed to ensure a thriving sector and economy. However, the ability to identify R&D needs is severely limited due to a lack of current, statistically representative data on manufacturing water use and its environmental implications. In this Perspective, we outline four key questions to inform R&D on manufacturing use and highlight how the current state of water data in the United States does not support the adequate investigation of these questions. We make recommendations for the water data characteristics needed to explore the research questions and knowledgeably inform R&D on manufacturing water use.

McCall, James↗

State-of-the-art of data collection, analytics, and future needs of transmission utilities worldwide to account for the continuous growth of sensing data

Nowadays, transmission system operators require higher degree of observability in real-time to gain situational awareness and improve the decision-making process to guarantee a safe and reliable operation. Digitalization of energy systems allows utilities to monitor the system dynamic performance in real-time at fast time scales. The use of such technologies has unlocked new opportunities to introduce new data driven algorithms for improving the stability assessment and control of the system. Motivated by these challenges, a group of experts have worked together to highlight and establish a baseline set of these common concerns, which can be used as motivation to propose innovative analytics and data-driven solutions. In this document, the results of a survey on 10 transmission system operators around the world are presented and it aims to understand the current practices of the participating companies, in terms of data acquisition, handling, storage, modelling and analytics. The overall objective of this document is to capture the actual needs from the interviewed utilities, thereby laying the groundwork for setting valid assumptions for the development of advanced algorithms in this field.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Proceedings for the Workshop on Applied Nuclear Data Activities 2024

The Workshop for Applied Nuclear Data Activities (WANDA) is designed to increase communication among nuclear data (ND) users in multidisciplinary federal programs, ND producers, ND funders, and other ND experts. It also presents an opportunity to cross-pollinate ideas as well as introduce ND gaps identified by federal programs to ND experts and ND capabilities to the various federal ND users. WANDA 2024 included five technical sessions, three of which focused on Fusion Energy Sciences (FES)—FES Fusion Neutronics, FES Tritium Production, and FES Material Damage—and two stand-alone sessions—Isotopes and Targetry for Nuclear Data and Uncertainty Quantification. The FES sessions successfully brought new voices to the WANDA discussions, expanding the application space in which nuclear data are critical. FES programs need accurate nuclear data with realistic uncertainty quantification to properly estimate, for example, shielding, activation, tritium production, helium production, structural material integrity, and superconducting magnet operation. This includes a variety of projectile (neutrons, photons, charged particles) and target atoms. One of the action items common to all the FES sessions was a need to perform sensitivity studies to identify the prioritization of nuclear data needs. The Isotopes and Targetry session highlighted the many capabilities available to produce high-quality targets for nuclear data measurements, including 3D printing with spherical powders, combustion synthesis coupled with spin coating & electrospraying, inkjet printing, and isotopic doping. These new methods open doors for more accurate measurement, but it was also stressed that sample characterization following any method of fabrication is of the highest importance to accurately interpret nuclear data measurement results that used that sample. The Uncertainty Quantification (UQ) session was broken into two categories: nuclear data uncertainty quantification and the use of that uncertainty quantification. Thematic to the UQ session was the loss of information when going from nuclear data measurement, to evaluation, to evaluated file, and finally to neutron transport calculations. Current evaluated ND libraries typically only contain covariances, which assume that the probability distributions are Gaussian. Beyond being a simplified assumption for many evaluations, this can lead to negative values on many observables when attempting to sample the covariance. The covariance format, however, is very efficient in that a simple set of linear equations can transform uncertainty from parameters or cross sections to the application of interest. Focused collaboration is needed between nuclear data evaluators and nuclear data users to ensure that needs are being met.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Lessons Learned on the Interface Between Quantum and Conventional Networking

The future Quantum Internet is expected to be based on a hybrid architecture with core quantum transport capabilities complemented by conventional networking. Practical and foundational considerations indicate the need for conventional control and data planes that (i) utilize extensive existing telecommunications fiber infrastructure, and (ii) provide parallel conventional data channels needed for quantum networking protocols. We propose a quantum-conventional network (QCN) harness to implement a new architecture to meet these requirements. The QCN control plane carries the control and management traffic, whereas its data plane handles the conventional and quantum data communications. We established a local area QCN connecting three quantum laboratories over dedicated fiber and conventional network connections. We describe considerations and tradeoffs for layering QCN functionalities, informed by our recent quantum entanglement distribution experiments conducted over this network.

Alshowkan, Muneer↗

Anonymization of Network Traces Data through Condensation-based Differential Privacy

Network traces are considered a primary source of information to researchers, who use them to investigate research problems such as identifying user behavior, analyzing network hierarchy, maintaining network security, classifying packet flows, and much more. However, most organizations are reluctant to share their data with a third party or the public due to privacy concerns. Therefore, data anonymization prior to sharing becomes a convenient solution to both organizations and researchers. Although several anonymization algorithms are available, few of them allow sufficient privacy (organization need), acceptable data utility (researcher need), and efficient data analysis at the same time. This article introduces a condensation-based differential privacy anonymization approach that achieves an improved tradeoff between privacy and utility compared to existing techniques and produces anonymized network trace data that can be shared publicly without lowering its utility value. Our solution also does not incur extra computation overhead for the data analyzer. A prototype system has been implemented, and experiments have shown that the proposed approach preserves privacy and allows data analysis without revealing the original data even when injection attacks are launched against it. When anonymized datasets are given as input to graph-based intrusion detection techniques, they yield almost identical intrusion detection rates as the original datasets with only a negligible impact.

97 MATHEMATICS AND COMPUTING↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Pump House in East River Watershed, Colorado 2019-2024

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors near Pump House at Mount Crested Butte in the East River Watershed. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format ER-X-Y, where ER refers to East River, X is the location identifier, and Y is the depth identifier at the same X (shallow Y=1). Specifically, ER-PHS, ER-LMC, ER-LMF, and ER-SMN are associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and ER-RBTn (upslope n=1) are sampling transects during the 2019 Rootball Campaign. The sample and location information can be found in metadata.csv. Sampling and Measurements Each sample falls into one of the three sampling methods – (1) intact cores, (2) repacked samples, or (3) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. Both intact cores and repacked samples were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores and in-situ soil sensors could be prioritized because these sampling methods are less destructive. While the repacked samples were packed to the target bulk density (estimated post-sampling, when sample volume was known), these samples had altered pore structures. Nevertheless, intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method also has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Trail Creek in Taylor River Watershed, Colorado 2024-2025

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors near Trail Creek. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format TR-X-Y, where TR refers to Trail Creek, X is the treatment block identifier, and Y is the location identifier. Specifically, TR-ASCC1 is the control treatment block under the Adaptive Silviculture for Climate Change (ASCC) project, and TR-ASCC2 is the clear-cut treatment block. TR-ASCC-EHSn is associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and TR-ASCC-ERTn (upslope n=1) are ecohydrology sites along the electrical resistivity tomography transects. The sample and location information can be found in metadata.csv, and the data from the soil sensors will be included in a future data version when the observation period becomes sufficiently long for data analysis. Sampling and Measurements Each sample falls into one of the two sampling methods – (1) intact cores or (2) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. The intact cores were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Snodgrass Mountain in East River Watershed, Colorado 2020-2025

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors at Snodgrass Mountain. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format SG-X-Y, where SG refers to Snodgrass Mountain, X is the location identifier, and Y is the depth identifier at the same X (shallow Y=1). Specifically, SG-EHS is associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and SG-ERTn (upslope n=1) are points along the Snodgrass electrical resistivity tomography transect not associated with the existing site names in the directory. The sample and location information can be found in metadata.csv. Sampling and Measurements Each sample falls into one of the three sampling methods – (1) intact cores, (2) repacked samples, or (3) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. Both intact cores and repacked samples were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores and in-situ soil sensors could be prioritized because these sampling methods are less destructive. While the repacked samples were packed to the target bulk density (estimated post-sampling, when sample volume was known), these samples had altered pore structures. Nevertheless, intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method also has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Adaptive Generation of Training Data for ML Reduced Model Creation

Machine learning proxy models are often used to speed up or completely replace complex computational models. The greatly reduced and deterministic computational costs enable new use cases such as digital twin control systems and global optimization. The challenge of building these proxy models is generating the training data. A naive uniform sampling of the input space can result in a non-uniform sampling of the output space of a model. This can cause gaps in the training data coverage that can miss finer scale details resulting in poor accuracy. While larger and larger data sets could eventually fill in these gaps, the computational burden of full-scale simulation codes can make this prohibitive. In this paper, we present an adaptive data generation method that utilizes uncertainty estimation to identify regions where training data should be augmented. By targeting data generation to areas of need, representative data sets can be generated efficiently. The effectiveness of this method will be demonstrated on a simple one-dimensional function and a complex multidimensional physics model.

Cianciosa, Mark↗

Intelligent Prediction of States in Multi-port Autonomous Reconfigurable Solar power plant (MARS)

In power electronics, prediction of states may be used for identification of faults, determination of aging of components, identification of bad data measurements, among others. Prediction of states in power electronics have broadly been based on: (a) physics-based models, (b) data-driven models, and (c) hybrid models. In this paper, data-driven approaches are presented for intelligent prediction of states in multi-port autonomous reconfigurable solar power plant (MARS) and compared. The data-set needed to train the data-driven models based on artificial intelligence (AI) algorithms has been identified and the trained models are evaluated under different extrapolated normal and abnormal operating conditions. The AI algorithms include nonlinear auto-regressive exogenous model (NARX), spiking neural networks (SNN), and decision tree. The models are compared and contrasted. The best model (NARX) is evaluated under different normal and abnormal operating conditions that have indicated accurate prediction.

Debnath, Suman↗

Data Movement Categories

We have endeavored to classify the commonly seen data movement needs, as observed in data-intensive institutions (both commercial and non-profit), into four categories. Knowing how to map a data movement task into one of the four categories helps select proper data mover tools. For each category, how the data storage is involved, high-level examples and the nature of typical solutions are described. Finally, some general remarks are provided to help further orient readers new to this field - the 4th IT pillar.

99 GENERAL AND MISCELLANEOUS↗

Improving the Quality of Geothermal Data Through Data Standards and Pipelines Within the Geothermal Data Repository: Preprint

For machine learning outputs to be applicable to real world problems, high quality data are needed to ensure high quality results. With the more recent emphasis on machine learning in geothermal, there is an increasing need for greater focus on the quality of the data available for use in these projects. For example, Geothermal Operational Optimization Using Machine Learning (GOOML) utilized large quantities of geothermal power plant operational data to inform power plant operational configurations to maximize power generation. High quality datasets result from dependable sensors or devices collecting data, high frequency of measurements, sufficient data points, adequate metadata, reliable storage of data, and sufficient data curation. Another component that contributes to high quality data is reusability, which can be enhanced through data standardization. Data Standardization creates consistency in formatting and contents of like datasets, lessening preprocessing requirements and ensuring adequate information provided by a given dataset. The Geothermal Data Repository (GDR) aims to help improve data quality through automated data standardization for high-value datasets through the implementation of data pipelines alongside reliable and accessible long-term storage for datasets. As such, the GDR has decided to shift away from recommending the use of Excel-based content models and towards the implementation of automated data pipelines. This takes the burden of data standardization off the user and project team and will increase the availability of standardized geothermal data available through the GDR. A set of recommendations, or a data standard for each data type will exist with each data pipeline in order to advise data collection for maximum usability for future research. This paper serves to describe the GDR's proposed transition towards data standardization through automated data pipelines, to discuss the need for and value of such a shift, and to call for suggestions from the community regarding the most useful data standards and pipelines.

data↗

Informing forest carbon inventories under the Paris Agreement using ground-based forest monitoring data

Human interactions with forests have shaped Earth's climate for millennia and will continue to do so as we target net-zero emission goals. Accurately characterizing these climate impacts requires making reliable forest carbon data available for forest monitoring and planning. Here, we develop a semi-automated process for submitting forest carbon measurements from the largest relevant scientific database to the International Panel on Climate Change's Emission Factor Database, which currently has sparse forest carbon data. Building this bridge from scientific research to international policy is an important step towards managing forests in a net-zero motivated future. Humans have been influencing Earth's climate via transformative impacts on forests for millennia, and forests are now recognized as critical to climate change mitigation under the Paris Agreement. The efficacy of climate change mitigation planning and reporting depends on quality data on forest carbon (C) stocks and changes. The Emission Factor Database (EFDB) of the International Panel on Climate Change (IPCC) is intended to be a definitive source for such data, but needs comprehensive and well-documented data to be so. To facilitate submission of forest C estimates from scientific studies to EFDB, we develop and document a process for semi-automated data submission from the Global Forest C database (ForC v4.0), which is the largest compilation of ground-based forest C estimates. We then assess the data currently available through ForC and provide recommendations for improving forest data collection, analysis, and reporting. As of September 2024, ForC contained ~19,286 records potentially relevant to EFDB, 1068 of which had been submitted and posted to EFDB. These represented 19% of the total EFDB records for forest land. Records were unevenly distributed across variables and geographic regions. ForC records (37%) reviewed could not be submitted because the original publication lacked required information. In the future, ground-based forest C estimates should target gaps in the record, and studies should ensure that they report all information necessary for inclusion in EFDB. Given that climate change is rapidly impacting the world's forests, timely reporting of recent estimates will be critical to accurate forest C inventories.

54 ENVIRONMENTAL SCIENCES↗

Advances in building data management for building performance standards using the SEED platform

Reducing energy consumption and greenhouse gas emissions in the built environment is a critical step in achieving emission goals to mitigate climate change impacts. Local, federal, and international jurisdictions are deploying several methods to reduce energy and emissions such as voluntary and mandatory benchmarking and building performance standards, requiring building owners to reach energy and emission targets. Jurisdictions leveraging benchmarking and building performance standards require knowledge of the buildings covered; which is a large task due to staffing constraints, limited information on building characteristics and tax parcel data, and the need for advanced data management techniques to align datasets. This paper describes an open-source platform's recent advances to create consistent taxonomies, identify erroneous data, enable auditability, and track building performance. The paper concludes with two use cases on how the platform has been used by jurisdictions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Entropy and Boundary Based Adversarial Learning for Large Scale Unsupervised Domain Adaptation

Supervised semantic segmentation methods provide state-of-the-art performance, but their performance is limited by the amount of quality labeled data they need for training. Scarcity of labeled data and non-transferablity of models, due to cross-domain discrepancy makes it a bigger challenge for remote sensing imagery analysis. In this work, we approach this problem through adversarial learning, driven by entropy and boundary of region-of-interest for unsupervised domain adaptation. This concept helps with better boundary prediction and encourages target domain entropy maps (probability/uncertainty maps) to be similar to source domains. In particular, we showed that deriving informative entropy through the adversarial learning is essential to enable the adaptation. We used a large scale cross country building extraction dataset to validate the framework. The experimental results show the usefulness of considering boundary and entropy driven adversarial learning for adaptation.

Makkar, Nikhil↗

The critical importance of software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data storage needs by the foreseen increase in data rates in the next decade for HL-LHC is not sustainable within the current budgets. As a result, a more efficient usage of computing resources is required in order to realise the physics potential of future experiments. Software and computing are an integral part of experimental design, trigger and data acquisition, simulation, reconstruction, and analysis, as well as related theoretical predictions. A significant investment in computing and software is therefore critical. Advances in software and computing, including artificial intelligence (AI) and machine learning (ML), will be key for solving these challenges. Making better use of new processing hardware such as graphical processing units (GPUs) or ARM chips is a growing trend. This forms part of a computing solution that makes efficient use of facilities and contributes to the reduction of the environmental footprint of HEP computing. The HEP community already provided a roadmap for software and computing for the last EPPSU, and this paper updates that, with a focus on the most resource critical parts of our data processing chain.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification in High Explosives Equations of State [Slides]

In order to fit models to data, we need a way to simulate the data given a model. Our projects involved simulators for the Manganin gauge experiment and for cylinder tests, which we use to calibrate models to experiments. We call the parameters we seek to optimize the degrees of freedom (DOF). We use a Bayesian approach with its associated UQ benefits.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Object storage model for CMS data

In CMS, data access and management is organized around the data-tier model: a static definition of what subset of event information is available in a particular dataset, realized as a collection of files. In previous work, we have proposed a novel data management model that obviates the need for data tiers by exploding files into individual event data product objects. In this work, we estimate the potential savings in data volume based on user analysis patterns.

Smith, Nick↗