Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data guidelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Refined Method to Translate Solar Data Quality Assessment Flags to Estimated Measurement Uncertainty

Integrating solar resource uncertainties due to radiometer measurement performance and operational data quality assessment can provide improved estimates of economic bankability, system design performance, and compliance of solar energy conversion systems. Estimating radiometer measurement uncertainty is an established procedure consistent with recognized best practices and international guidelines. SERI QC is a robust solar data quality assessment software tool that has been in continuous use for more than three decades. This report, the fourth of six for the Data Quality and Uncertainty Integration Project, presents a refined algorithm description for software to translate solar resource data quality assessment results into estimated uncertainty values in a Solar Resource Operational Uncertainty Integrator (SROUI) application. This algorithm requires three-component solar irradiance measurements - global horizontal irradiance, direct normal irradiance, and diffuse horizontal irradiance - collected at 1- to 60-minute intervals, as described in the previous deliverables. The development of this report as Deliverable 6.4 was an iterative process that included reviews and feedback from the project team on initial drafts designed to refine how the new software could best support determining solar resource data uncertainty. The results of this effort will contribute to the final software system development by National Renewable Energy Laboratory staff.

14 SOLAR ENERGY↗

Aerobic respiration controls on shale weathering, Geochimica et Cosmochimica Acta, 2023: Dataset

This data package was generated in order to support the development of a deep-time weathering model and to assess the coupling between shale weathering and aerobic respiration in the paper “Aerobic respiration controls on shale weathering” by Stolze et al., Geochimica et Cosmochimica Acta (2023). The package contains two csv files providing the average CO2(g) concentration profiles [ppm] and mineral concentration profiles [wt%], respectively. The CO2(g) concentration profiles were measured in the vicinity of the monitoring well PLM2 between January 2018 and April 2019. The gas samples were collected in the unsaturated zone to a depth of 1.52 m. The mineral concentration profiles were determined by X-Ray diffraction (XRD). The XRD measurements were performed on sub-core samples collected in the monitoring well PLM3 down to a depth of 7.01 m. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Update on 2024-05-28: Revised versions of the CSV data files (CO2_data_GCA_Stolze_et_al_2023.csv and XRD_data_GCA_Stolze_et_al_2023.csv) were made to apply ESS-DIVE's CSV reporting format guidelines. Updated versions of the File Level Metadata (v2_20240528_flmd.csv) and Data Dictionary (v2_20240528_dd.csv) files were updated to reflect the changes made to the CSV files.

54 ENVIRONMENTAL SCIENCES↗

ESS-DIVE Reporting Format for Sample-based Water and Soil Chemistry Measurements

The ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) reporting format for sample-based water and soil chemistry measurements are written guidelines and spreadsheet templates, that facilitate archiving water/soil/sediment chemistry data. The templates and terminology lists provided as part of the reporting format can help organize data and potentially enable data reuse.The reporting format consists of a general instructions file (instructions.md) as well as more detailed instructions files for each template (files beginning “Detailed_Instructions_*.md). The “examples” folder includes an example of each data template with just the limited set of required fields filled out as well as other examples with both required and optional fields complete.The 'templates' folder contains CSV templates for data, methods, and terminology file templates. Similar to the examples folder, the blank templates are provided with versions for required fields only as well as required and optional fields. Lastly, the term list folder provides common terminology used in each of the templates as well as a definition and any constraints for the fields.

54 ENVIRONMENTAL SCIENCES↗

Current and future federal and state sampling guidance for per- and polyfluoroalkyl substances in environmental matrices

Per- and polyfluoroalkyl substances (PFAS) are a class of emerging contaminants composed of an estimated 5000 to 10,000 human-made, fluorinated, organic chemicals. Due to the complexity of PFAS, the need for multiple environmental matrix considerations and the absence of a promulgated federal standard for environmental sampling and analysis, U.S. states have begun developing health-based regulatory and/or guidance values for a limited number of PFAS in environmental matrices. As there is a growing body of science to inform PFAS sampling guidance standard development, it is important to understand which U.S. states are implementing sampling guidelines and how they plan to handle emerging PFAS. This critical review discusses the current and impending federal and state sampling guidelines for PFAS in environmental matrices, the data gaps surrounding PFAS sampling guidance in U.S. states, and the future impacts of impending guidance documents and regulations. Ten federal guidance documents are available for PFAS sampling guidance and analysis. The maximum number of PFAS covered in these guidance documents is 25 analytes spanning across 8 unique media. While the EPA has developed several different sampling and analytical guidelines for PFAS, there is no formal regulation of PFAS or requirements of states to enforce these guidelines. Consequently, only 31 states have informally adopted sampling guidelines, while the other 19 states have no guidance documentation in place for PFAS. The introduction of new PFAS sampling guidelines by the EPA, as well as updated analytical guidelines that target more PFAS or total organofluoride, is expected to continuously shift the landscape of federal and state guidance for PFAS sampling moving forward.

54 ENVIRONMENTAL SCIENCES↗

National and Regional Initiatives to Promote Energy Efficiency and Renewable Energy Through State Energy Offices

The National Association of State Energy Officials (NASEO) worked with the U.S. Department of Energy’s (DOE) Office of Energy Efficiency and Renewable Energy (EERE), Weatherization and Intergovernmental Programs Office (WIP) over a ten year period to provide technical assistance, research and analyses, and enhanced coordination between DOE and the State and Territory Energy Offices. Over the life of the agreement, NASEO worked with WIP and the states to provide statewide strategic energy plan analyses and recommendations, including customized technical assistance to the states; peer to peer financing assistance via NASEO’s Financing Committee and supporting activities; training on core energy policies and programs, including dialogues and resources to support energy-air coordination, a training for new State Energy Office Directors, peer to peer exchange via a rural energy taskforce, and support to states to enhance buildings efficiency, home energy labeling, and technology deployment opportunities; support for State Energy Program metrics; and regional coordination via regional coordinators and peer exchange opportunities both in-person and online. Over the life of the project, the position and role of State Energy Offices within state government has changed, with now 80 percent of State Energy Office Directors serving as their Governor’s energy advisor, or reporting directly to their Governor’s energy advisor. This shift in stature has made it ever more crucial for State Energy Offices to receive timely and relevant technical assistance across a range of energy issue areas. NASEO, in collaboration with WIP, was able to deliver this technical assistance, and energy efficiency and renewable energy deployment has accelerated across the country. The State Energy Office Directors and their staff continue to engage in NASEO’s Committees – many of which were supported through this agreement – and have provided formal and informal feedback on the value of the programs supported through this award (e.g., reporting via survey increased understanding of their roles and technical assistance offerings provided by WIP and NASEO following the New Director Trainings). Moreover, resources developed through this award (e.g., Comprehensive State Energy Planning Guidelines, State Energy Loan Fund Database, Rural Data Resources for State Energy Planning and Programs, The Value of Adding home Energy Score to Low-Income Energy Efficiency Programs, etc.) have been cited by states as instrumental to their understanding of specific energy issue areas and in many cases led directly to enhanced program design within a state (e.g. Iowa citing their review of NASEO’s Comprehensive State Energy Planning Guidelines as a necessary first step in their planning process, later following may of the steps outlined in the guidelines). The priorities of the states and federal government have evolved over the last decade, with an increasing focus on climate mitigation and adaptation, equity impacts and considerations, energy security and resilience, and enhanced energy efficiency and renewable energy technology deployment, but the roots of these new priorities are based in the state and federal policies and programs, and research and analyses, that were supported in part through this award and other complementary initiatives. NASEO looks forward to continuing to support the states, and collaborate and coordinate with DOE, as we build on this important foundation in the years ahead.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Energy performance evaluation of the ASHRAE Guideline 36 control and reinforcement learning–based control using field measurements

This study evaluates the energy performance of ASHRAE Guideline 36–compliant control (ASHRAE 36 control) and reinforcement learning (RL)–based control through experimental field tests and a simulation study. Three field tests were conducted at Oak Ridge National Laboratory’s commercial building test facility in Oak Ridge, Tennessee: a baseline with a baseline conventional control, a test with ASHRAE 36 control, and a test with RL-based control. The selected ASHRAE 36 controls were trim and respond control, as well as variable air volume (VAV) box control. We compared the measured supply air temperature of the rooftop unit, VAV box supply air temperature, and VAV box supply airflow rate across the three test cases. The field data indicated that ASHRAE 36 controls operated as specified by ASHRAE Guideline 36. Based on these data, ASHRAE 36 control achieved a 45 % reduction in hourly averaged HVAC energy consumption compared with the baseline, and RL-based control achieved a 66 % reduction. These potential annual energy savings were confirmed using a calibrated whole-building energy model. Compared with the baseline, ASHRAE 36 control reduced HVAC energy consumption by 42 %, and RL-based control achieved a 54 % reduction. Furthermore, RL-based control reduced total HVAC energy consumption by 21 % more than ASHRAE 36 control.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

LiDAR Measurements of Wind Shear Exponents and Turbulence Intensity Offshore the Northeast United States

This paper presents wind speed shear exponents and turbulence intensity measurements collected from LiDARs measuring wind speeds from 40m to 200m above sea-level and provides comparisons to industry design guidelines. The high-altitude wind speed data are unique and represent some of the first measurements made offshore in this part of the country, which is actively being developed for offshore wind. The data is used to support the New England Aqua Ventus I Floating Offshore Wind Farm to be located 17km offshore the Northeast United States. Multiple LiDAR measurements were made using a DeepCLiDAR floating buoy and LiDARs located on a nearby island. Here, the measured wind speed shear exponents are compared against industry standard mesoscale model outputs and offshore design codes including the American Bureau of Shipping, American Petroleum Institute, and DNV-GL guides. Significant variation in the vertical wind speed profile occurs throughout the year which is not addressed in design standards. Additionally, turbulence intensity measurements made from the LiDAR, although not widely accepted in the scientific community, are presented and compared against industry guidelines.

42 ENGINEERING↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Multicriteria Measures to Assess the Sustainability of Diets: A Systematic Review

Abstract Context Assessing the overall sustainability of a diet is a challenging undertaking requiring a holistic approach capable of addressing the multicriteria nature of this concept. Objective The aim was to identify and summarize the multicriteria measures used to assess the sustainability characteristics of diets reported at the individual level by healthy adults. Data Sources Articles were identified via PubMed, Scopus, and Web of Science. The search strategy consisted of key words and MeSH terms, and was concluded in September 2022, covering references in English, Spanish, and Portuguese. Data Extraction This systematic review followed the PRISMA guidelines. The search identified 5663 references, from which 1794 were duplicates. Two reviewers independently screened the titles and abstracts of each of the 3869 records and the full-text of the 144 references selected. Of these, 7 studies met the inclusion criteria. Data Analysis A total of 6 multicriteria measures were identified: 3 different Sustainable Diet Indices, the Quality Environmental Costs of Diet, the Quality Financial Costs of Diet, and the Environmental Impact of Diet. All of these incorporated a health/nutrition dimension, while the environmental and economic dimensions were the second and the third most integrated, respectively. A sociocultural sustainability dimension was included in only 1 of the measures. Conclusion Despite some methodological concerns in the development and validation process of the identified measures, their inclusion is considered indispensable in assessing the transition towards sustainable diets in future studies. Systematic Review Registration PROSPERO registration no. CRD42022358824.

Rei, Mariana (ORCID:0000000189453708)↗

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

Terrestrial laser scanning data (Levels 0 and 1) from Urban Biogeochemistry Pilot Project sites, Knoxville, Tennessee, Jul 2024 - Jul 2025

This data package contains data from terrestrial laser scanning (TLS) at five urban park sites in Knoxville, Tennessee, USA. All parks include open-grown and/or closed-canopy trees and mixed nearby land use. These study sites were established as part of the Urban Biogeochemistry Pilot Project, which has an overall goal of better understanding how hydrobiogeochemical cycling is altered within the human environment. These five sites represent a gradient of urbanization, and were instrumented to understand hydrological and biogeochemical cycling (e.g., soil moisture, soil physical properties and biogeochemistry, tree transpiration, species type). The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree- and stand-level characterization of woody structure and leaf area. TLS scans were placed to capture the area around trees with sap flow sensors, and as much of a 50 m radius area around the meteorological station as possible given site property limits. Derived products will allow upscaling of water content and transpiration data. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

Exploring physics of ferroelectric domain walls via Bayesian analysis of atomically resolved STEM data

The physics of ferroelectric domain walls is explored using the Bayesian inference analysis of atomically resolved STEM data. We demonstrate that domain wall profile shapes are ultimately sensitive to the nature of the order parameter in the material, including the functional form of Ginzburg-Landau-Devonshire expansion, and numerical value of the corresponding parameters. The preexisting materials knowledge naturally folds in the Bayesian framework in the form of prior distributions, with the different order parameters forming competing (or hierarchical) models. Here, we explore the physics of the ferroelectric domain walls in BiFeO 3 using this method, and derive the posterior estimates of relevant parameters. More generally, this inference approach both allows learning materials physics from experimental data with associated uncertainty quantification, and establishing guidelines for instrumental development answering questions on what resolution and information limits are necessary for reliable observation of specific physical mechanisms of interest.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

High-throughput spin-bath characterization of spin defects in semiconductors

Detailed knowledge of the local environments of spin defects in semiconductors, such as nitrogenvacancy (NV) centers in diamond or divacancies in silicon carbide, is crucial for optimizing control and entanglement protocols in quantum sensing and information applications. However, at present a direct experimental characterization of individual defect environments is not scalable, as conventional spin-bath measurements are time consuming and difficult to automate. Achieving high-throughput characterization requires short experiments to probe the spin bath. However, with fewer and noisier measurements, the inverse problem of recovering spin-bath properties from measured data becomes ill posed, with multiple spin baths having a high likelihood of yielding the same data. In this work, we present a set of computational tools to resolve the ill-posed inverse problem of recovering the atomic positions and hyperfine couplings of random nuclei surrounding spin defects from sparse, noisy experimental coherence data, which can be obtained in hours. Here, we use a trans-dimensional Bayesian approach that incorporates ab initio data to yield full posterior distributions over nuclear spin environments, enabling robust recovery from limited data. We also provide practical tools and guidelines to determine the limits of detectability for hyperfine couplings under specific dynamical decoupling sequences and sampling conditions. In addition, we demonstrate how the tools developed here, in combination with ab initio simulations of spin baths, can guide the design of efficient experimental protocols for application-specific high-throughput screening. To showcase the utility of our approach, we apply it to design fast dynamical decoupling experiments to characterize the spin baths often individual NV centers in diamond. While the primary focus is on accelerating spin-bath characterization of spin defects, this Bayesian approach also lays the foundation for digital-twin studies of spin defects, where a virtual model of the spin-defect system evolves in real time with ongoing experimental measurements. Together, the set of tools we designed and applied paves the way for scalable deployment of spin defects in semiconductors for quantum sensing and information applications.

Bayesian methods↗

CheKiPEUQ

Parameter estimation for complex physical problems often suffers from finding 'solutions' that are not physically realistic. The CheKiPEUQ software provides tools for finding physically realistic parameter estimates.and CheKiPEUQ provide tools for making graphs of the parameter positions within parameter space as well as plots of the final simulation results. The primary purpose of the CheKiPEUQ software is to enable more physically realistic parameter estimation from comparing simulations to experiments. Specifically, when prior knowledge is available about the region of parameter space which is physically realistic and when the level of uncertainty from experiment can be estimated. The software is intentionally general and can be used for almost any type of simulation. The software is made in a user friendly manner so that users can simply enter the required data and then run the program, following guidelines provided by the authors along with some trial and error. Examples are provided. Users do not need to understand the methodology that will be namedi in the following sentences. While CheKiPEUQ can be used for conventional parameter estimation and other uses, the primary uses for CheKiPEUQ are: 1) Bayesian Parameter Estimation, 2) Bayesian Model Discrimination, 3) Bayesian Design of Experiments. For more information see the project website, documentation, examples, and related publications.

Savara, Aditya [Oak Ridge National Lab. (ORNL), Oa↗

Videos, photos, and AI-derived grain size data associated with “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization” under review. This data package includes five data types: 1) raw photos and videos from drone survey and walking smartphone surveys; 2) images derived from raw videos; 3) manual labeling of reference scales; 4) metadata for all images and photo resolution derived from artificial intelligence (AI) models or manual labels, 5) grain size data obtained from AI models for all photos, 6) metadata and grain size data after quality control, 7) summaries of sample efficiency for all data, and 8) computational fluid dynamics (CFD) data used to support hydro-biogeochemical (HBGC) parameter estimation. Such data is used to 1) demonstrate significant improvements in accuracy, efficiency, and quality control for grain size data collection with the help of AI models, 2) study the spatial heterogeneity of grain size and observation reproducibility based on tens of thousands of data points generated by the AI models, and 3) evaluate the impacts of grain size heterogeneity on key HBGC parameters across sediment-to-reach and hourly-to-yearly scales. In particular, the data package contains 116 folders and 179696 files. The files include 41 videos in .mov format, 64047 photos in .jpg format, 13541 video-derived photos in .png format, 12747 segmentation mask data in .tif format, 12747 segmentation data in .json format, 24771 .csv files that with metadata and grain size for each individual photo as well as water depth and velocity data from CFD and observation, 51791 .txt files of raw AI predicted labels, and 11 flight record data in .srt format. The summary for all metadata and grain size statistics information is included in “Scales_V3_NG.csv” and “Statistics_V3_NG.csv”. The summary for data that pass data quality control (QC) level 0-2 is included in “QCStatistics_V3_NG.csv”. The QC level 0 represents photos whose photo resolution is positive, excluding photos that miss reference scale. The QC level 1 means reference scale circularity uncertainty is less than 5% for smartphone images while representing photo resolution is larger than 0.44 mm/pixel for drone images. The QC level 2 means excluding photos whose grain number is less than 100, a minimum number of grains recommended by classic literature. The summary for each video’s name, length, frame rates, survey area, grain number, survey efficiency, etc. can be found in “QCSummary_V3_NG.csv”. The summary for site name, GPS coordinates, and number of images at each site can be found in “SitesSummary_V3_*.csv” files. Overall computational efficiency summary is reported in Table 4 of accompanying manuscript. Additionally, the nitrate concentration data used in this work was downloaded from an existing dataset published on ESS-DIVE (Boat-Dragged Sensor Hanford Reach.csv; Conner A. et al., 2020). We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Port of Benton, and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the data were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate data collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Performance Results for Sensor Assignment Problem as Solved on a Multi-Node Cluster

An earlier report described a procedure for optimal sensor set selection and its implementation on a computational cluster. This new and innovative capability was developed to facilitate a reduction in operations staffing levels to improve plant economics. By automating surveillance and maintenance tasks through early detection of degrading sensors and equipment, staff can be more efficiently deployed. The method uses automated reasoning and domain knowledge in the form of the conservation equations to infer from plant measurements the state of equipment health. Inclusion of domain knowledge addresses the problem that exists with pure data-driven methods that there are no rigorous guidelines for determining what constitutes an adequate sensor set. Formalizing the procedure for sensor set selection as we have done results in a more reliable and explainable diagnosis of plant equipment health. Importantly, from the standpoint of the plant owner, personnel are provided with an early and explicit diagnosis of an equipment problem. That in principle automates the process and eliminates having to send personnel into the plant to find the cause as typically occurs when a data-driven method detects an anomaly. In this report we describe first results obtained using a computational cluster to solve the sensor set selection problem as framed above. The case described addresses the problem of equipment health monitoring in the high-pressure (HP) feedwater system of a pressurized light water reactor as seen through the eyes of our collaborating utility partner. Maintenance of this system can amount to millions of dollars per year if equipment health issues go undiagnosed and lead to loss of function. On examining the potential that is inherent in the installed sensor set for diagnosing equipment health degradation, it was found that greater fault resolution capability can be achieved using a sensor set that is 20 percent fewer in number. The take-away is that compared to the installed sensor set there exists a more strategic assignment of sensors that will furnish better health monitoring capability and with fewer sensors. Where the problem defies solution by manual inspection, as is the case here, one can be found by an algorithm. The solution was obtained in four hours using 30 computational cores. The HP feedwater problem as posed above illustrates the added value of approaching the sensor selection problem as one amenable to algorithmic solution. This problem is of interest to advanced reactor designers and to utilities that are setting up remote monitoring and diagnostic centers.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications

Cybersecurity breaches targeting electrical substations constitute a significant threat to the integrity of the power grid, necessitating comprehensive defense and mitigation strategies. Any anomaly in information and communication technology (ICT) should be detected for secure communications between devices in digital substations. This paper proposes large language models (LLMs), e.g., ChatGPT, for the cybersecurity of IEC 61850-based communications. Multi-cast messages such as generic object oriented system events (GOOSE) and sampled values (SV) are used for case studies. The proposed LLM-based cybersecurity framework includes, for the first time, data pre-processing of communication systems and human-in-the-loop (HITL) training (considering the cybersecurity guidelines recommended by humans). The results show a comparative analysis of detected anomaly data carried out based on the performance evaluation metrics for different LLMs. A hardware-in-the-loop (HIL) testbed is used to generate and extract a dataset of IEC 61850 communications.

ChatGPT↗

Materials Data Science Ontology(MDS-Onto): Unifying Domain Knowledge in Materials and Applied Data Science

Ontologies have gained popularity in the scientific community as a way to standardize terminologies in organizations’ data. Although certain cohorts have created frameworks with rules and guidelines on creating ontologies, there exist significant variations in how Materials Science ontologies are currently developed. We seek to provide guidance in the form of a unified automated framework for developing interoperable and modular ontologies for Materials Data Science that simplifies the ontology terms matching by establishing a semantic bridge up to the Basic Formal Ontology(BFO). This framework provides key recommendations on how ontologies should be positioned within the semantic web, what knowledge representation language is recommended, and where ontologies should be published online to boost their findability and interoperability. Two fundamental components of the MDS-Onto framework are the bilingual package called FAIRmaterials for ontology creation and FAIRLinked, for FAIR data creation. To showcase the practical capabilities of FAIRmaterials, we present two exemplar domain ontologies of MDS-Onto: Synchrotron X-Ray Diffraction and Photovoltaics.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗