Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data guidelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

Analytical methods for online data quality assessment

This chapter provides a comprehensive overview of the main steps for algorithmic sensor signal quality assessment, which can enhance the decision-making process for water resource recovery facility (WRRF) operation and optimization. It introduces the concept of redundancy as the basis for data quality assessment. It also explains the typical data processing pipeline, which consists of preliminary analysis, data pre-processing, and specific algorithmic approaches. Each of these processes is presented and discussed in three separate sections. Importantly, this chapter introduces the main approaches for data quality assessment, provides guidelines for selecting the most suitable one and the key performance indicators to evaluate them and explains how to collect metadata through such an algorithmic approach.

Aguado, Daniel↗

A Parameterization Study of Sew-EZ Materials: Types #6 and #8

Two material types identified by Sew-EZ were tested in various configurations, and under various conditions, by Sandia National Laboratories (SNL). The primary focus of this study was to assess the filtration performance of these two materials and identify if they perform similarly to certified N95 respirators. Testing was conducted on two systems which use distinctly different techniques to characterize the aerosol penetration characteristics of materials: a) R&D Filtration System: A large-scale R&D filtration system was used with testing parameters that mimicked NIOSH guidelines, where possible. Efficiency data as a function of particle size was attained using NaC1 as the test aerosol and a Scanning Mobility Particle Sizer (SMPS) for measurements. A more detailed system description can be found in Omana et al. 2020. b) Automated Tester: A commercial, automated filter tester (100Xs, Air Techniques International) was used to provide penetration/efficiency data for Sew EZ materials. The 100Xs aerosolizes a polydisperse NaC1 aerosol with a consistent concentration and size profile. The 100Xs manual (Air Techniques International 2018) states, "The aerosol particle size and distribution are designed to meet all requirements as defined in the relevant sections of NIOSH 42 CFR, Part 84 (pg. 32)."

36 MATERIALS SCIENCE↗

Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science

Abstract Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph‐based data models elucidate the interconnectedness among core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these “knowledge graphs” (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally accepted, open‐access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open‐source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object‐oriented classification and graph‐oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates) representing biomedical entities such as gene, disease, chemical, anatomic structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

60 APPLIED LIFE SCIENCES↗

Navigating the Expansive Landscapes of Soft Materials: A User Guide for High-Throughput Workflows

Synthetic polymers are highly customizable with tailored structures and functionality, yet this versatility generates challenges in the design of advanced materials due to the size and complexity of the design space. Thus, exploration and optimization of polymer properties using combinatorial libraries has become increasingly common, which requires careful selection of synthetic strategies, characterization techniques, and rapid processing workflows to obtain fundamental principles from these large data sets. Herein, we provide guidelines for strategic design of macromolecule libraries and workflows to efficiently navigate these high-dimensional design spaces. We describe synthetic methods for multiple library sizes and structures as well as characterization methods to rapidly generate data sets, including tools that can be adapted from biological workflows. We further highlight relevant insights from statistics and machine learning to aid in data featurization, representation, and analysis. This Perspective acts as a “user guide” for researchers interested in leveraging high-throughput screening toward the design of multifunctional polymers and predictive modeling of structure–property relationships in soft materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats

Abstract Research can be more transparent and collaborative by using Findable, Accessible, Interoperable, and Reusable (FAIR) principles to publish Earth and environmental science data. Reporting formats—instructions, templates, and tools for consistently formatting data within a discipline—can help make data more accessible and reusable. However, the immense diversity of data types across Earth science disciplines makes development and adoption challenging. Here, we describe 11 community reporting formats for a diverse set of Earth science (meta)data including cross-domain metadata (dataset metadata, location metadata, sample metadata), file-formatting guidelines (file-level metadata, CSV files, terrestrial model data archiving), and domain-specific reporting formats for some biological, geochemical, and hydrological data (amplicon abundance tables, leaf-level gas exchange, soil respiration, water and sediment chemistry, sensor-based hydrologic measurements). More broadly, we provide guidelines that communities can use to create new (meta)data formats that integrate with their scientific workflows. Such reporting formats have the potential to accelerate scientific discovery and predictions by making it easier for data contributors to provide (meta)data that are more interoperable and reusable.

54 ENVIRONMENTAL SCIENCES↗

Common-Cause Component Group Modeling Issues in Probabilistic Risk Assess

Common cause failures (CCFs) have been recognized as significant risk contributors in probabilistic risk assessments (PRAs) for commercial nuclear power plants (NPPs) since 1980s. A series of reports including those of United States Nuclear Regulatory Commission (NRC) regulations (NUREGs) have been published from late 1980s to early 2000s to provide guidelines for performing CCF event data analysis and modeling CCF in PRA. However, there are still some issues existed in CCF modeling and CCF parameter estimations. One of such issues is the modeling of multiple common-cause component group (CCCG) for the same component in a PRA model. This paper exams the CCCG definition and the requirements from the PRA standard, investigates the CCCG modeling practices and issues existing in PRAs, and presents systematic approaches and guidelines to model CCCG in PRA properly and consistently.

99 GENERAL AND MISCELLANEOUS↗

Photovoltaic System Health-State Architecture for Data-Driven Failure Detection

The timely detection of photovoltaic (PV) system failures is important for maintaining optimal performance and lifetime reliability. A main challenge remains the lack of a unified health-state architecture for the uninterrupted monitoring and predictive performance of PV systems. To this end, existing failure detection models are strongly dependent on the availability and quality of site-specific historic data. The scope of this work is to address these fundamental challenges by presenting a health-state architecture for advanced PV system monitoring. The proposed architecture comprises of a machine learning model for PV performance modeling and accurate failure diagnosis. The predictive model is optimally trained on low amounts of on-site data using minimal features and coupled to functional routines for data quality verification, whereas the classifier is trained under an enhanced supervised learning regime. The results demonstrated high accuracies for the implemented predictive model, exhibiting normalized root mean square errors lower than 3.40% even when trained with low data shares. The classification results provided evidence that fault conditions can be detected with a sensitivity of 83.91% for synthetic power-loss events (power reduction of 5%) and of 97.99% for field-emulated failures in the test-bench PV system. Finally, this work provides insights on how to construct an accurate PV system with predictive and classification models for the timely detection of faults and uninterrupted monitoring of PV systems, regardless of historic data availability and quality. Such guidelines and insights on the development of accurate health-state architectures for PV plants can have positive implications in operation and maintenance and monitoring strategies, thus improving the system’s performance.

photovoltaics↗

Characterization of the In Vivo Deuteration of Native Phospholipids by Mass Spectrometry Yields Guidelines for Their Regiospecific Customization

Customization of deuterated biomolecules is vital for many advanced biological experiments including neutron scattering. However, because it is challenging to control the proportion and regiospecificity of deuterium incorporation in live systems, often only two or three synthetic lipids are mixed together to form simplistic model membranes. This limits the applicability and biological accuracy of the results generated with these synthetic membranes. Despite some limited prior examination of deuterating Escherichia coli lipids in vivo, this approach has not been widely implemented. In this report an extensive mass spectrometry-based profiling of E. coli phospholipid deuteration states with several different growth media was performed, and a computational method to describe deuterium distributions with a one-number summary is introduced. The deuteration states of 36 lipid species were quantitatively profiled in 15 different growth conditions, and tandem mass spectrometry was used to reveal deuterium localization. Regressions were employed to enable the prediction of lipid deuteration for untested conditions. Small-angle neutron scattering was performed on select deuterated lipid samples, which validated the deuteration states calculated from the mass spectral data. Based on these experiments, guidelines for the design of specifically deuterated phospholipids are described. This unlocks even greater capabilities from neutron-based techniques, enabling experiments that were formerly impossible.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Causes and consequences of experimental variation in Nicotiana benthamiana transient expression

Infiltration of Agrobacterium tumefaciens into Nicotiana benthamiana has become a foundational technique in plant biology, enabling efficient delivery of transgenes in planta with technical ease, robust signal, and relatively high throughput. Despite transient expression’s prevalence in disciplines such as synthetic biology, little work has been done to describe and address the variability inherent in this system, a concern for experiments that rely on highly quantitative readouts. In a comprehensive analysis of N. benthamiana agroinfiltration experiments, we model sources of variability that affect transient expression. Our findings emphasize the need to validate normalization methods under the specific conditions of each study, as distinct normalization schemes do not always reduce variation either within or between experiments. Using a dataset of 1915 plants collected over three years, we develop a model of variation in N. benthamiana transient expression, using power analysis to determine the number of individual plants required for a given effect size. Drawing on our longitudinal data, these findings inform practical guidelines for minimizing variability through strategic experimental design and power analysis, providing a foundation for more robust and reproducible use of N. benthamiana in quantitative plant biology and synthetic biology applications.

Tang, Sophia N. [Joint BioEnergy Institute (JBEI),↗

Data from: “Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats”

This dataset contains supplementary information for a manuscript describing the ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) data repository's community data and metadata reporting formats. The purpose of creating the ESS-DIVE reporting formats was to provide guidelines for formatting some of the diverse data types that can be found in the ESS-DIVE repository. The 6 teams of community partners who developed the reporting formats included scientists and engineers from across the Department of Energy National Lab network. Additionally, during the development process, 247 individuals representing 128 institutions provided input on the formats. The primary files in this dataset are 10 data and metadata crosswalk for ESS-DIVE’s reporting formats (all files ending in _crosswalk.csv). The crosswalks compare elements used in each of the reporting formats to other related standards and data resources (e.g., repositories, datasets, data systems). This dataset also contains additional files recommended by ESS-DIVE’s file-level metadata reporting format. Each data file has an associated dictionary (files ending in _dd.csv) which provide a brief description of each standard or data resource consulted in the data reporting format development process. The flmd.csv file describes each file contained within the dataset.

54 ENVIRONMENTAL SCIENCES↗

KS4A-Omics1.0_FspDS682

Soil fungi facilitate the translocation of inorganic nutrients from soil minerals to other microorganisms and plants. This ability is particularly advantageous in impoverished soils, because fungal mycelial networks can bridge otherwise spatially disconnected and inaccessible nutrient hotspots. However, the molecular mechanisms underlying fungal mineral weathering and transport through soil remains poorly understood. Here, we addressed this knowledge gap by directly visualizing nutrient acquisition and transport through fungal hyphae in a mineral doped soil micromodel using a multimodal imaging approach. Here, we observed how a representative of common saprotrophic soil fungi, Fusarium sp. DS 682, exhibited a mechanosensory response (thigmotropism) around obstacles and through pore spaces (~12 μm) in the presence of minerals.This study establishes the significance of fungal biology and nutrient translocation mechanisms in maintaining fungal growth under water and nutrient limitations in a soil-like microenvironment, using a high-throughput multi-omic analysis approach. Data package KS4A-Omics.1.0_FspDS682 (Publication: Fungal Mineral Weathering Mechanisms Revealed Through Direct Molecular Visualization) contents reported here are the first version (1.0) and contain pre- and post-processed data using high throughput data capture technologies for multi-omic analysis and integration, this data package contains raw and post-processed experimental data for X-Ray Absorption Near Edge Structure Spectroscopy (XANES/XRF), Optical Microscopy, Proteomics, Scanning Electron Microscope (SEM), Time-of-Flight Secondary Ion Mass Spectroscopy (ToF-SIMS), X-Ray Diffraction Spectroscopy (XRD) files, and X- Ray Photoelectron Spectroscopy (XPS) using EMSL capabilities. This data package DOI contains a comprehensive collection of high-throughput multi-omics data and process metadata catalog. Support files include additional data download contents “Read Me” with dataset descriptor information and data source method application ontologies (see data dictionary section). Reported data download content is structured for compliance with reported guidelines provided by community standard initiatives and publisher stakeholder policies supporting FAIR data principles.

47 OTHER INSTRUMENTATION↗

A transfer learning approach to energy-efficient control of small and medium-sized commercial buildings

Model-free reinforcement learning (RL) provides a data-driven and adaptive approach to optimize building energy use while satisfying occupant comfort. This powerful tool does not need any prior knowledge about the environment and system it is optimizing and can adapt its policy based on the changes in captures. Like any other data-driven tool, it faces high training costs due to the extensive agent-environment interactions required to capture long-term building dynamics and user comfort. Transfer learning, particularly policy distillation, offers a promising way to accelerate training by leveraging pretrained RL agents in different building and system types. Here, this study investigates online student distillation, in which the student model updates its neural network weights using outputs from teacher models. The work introduces a student distillation strategy designed for efficient knowledge transfer, along with a teacher selection method that ensures high-quality guidance. The approach is validated using a highly calibrated whole building energy model for a small/medium commercial building test facility. Results show substantial reductions in training time and data requirements while surpassing the performance of ASHRAE Guideline 36, an advanced rule-based control strategy. The distilled RL model required 45% less data and achieved 20% higher cumulative rewards than a state-of-the-art RL model, with faster convergence and lower energy consumption. These outcomes demonstrate that effective transfer learning enables a scalable and data-efficient energy management solution for commercial buildings.

ASHRAE guideline 36↗

Update on Activities Related to the Library of Graphite Microstructures

This report provides an overview and update of the ongoing efforts to create a comprehensive library of microstructures for nuclear graphite and carbon-based materials under consideration for nuclear applications. The library includes data on microstructural characterization of unirradiated graphite materials, a guide to the techniques used to analyze graphite (which complements the ASME guidelines and ASTM standards), a summary of characterization data for neutron-irradiated or oxidized material, and a compendium of microstructural information for carbon-based materials. These efforts are being conducted at various length scales for the filler and binder phases in graphite to better understand graphite’s local structure and property relationships. This report is a follow-up to the previous milestone report titled Report on initial development of a database of nuclear graphite characteristics based on microstructural characterization, ORNL/TM/-2023/2992, published in July 2023.The effort to develop the library of microstructures supports the US Department of Energy Office of Advanced Reactor Technologies program objectives of aiding the material selection, licensing, management, and core assessments of a graphite core by documenting the unirradiated microstructure of relevant grades or characterizing the microstructure’s evolution under the reactor environment. Additionally, this project aims to provide (1) information and guidelines for the characterizing of graphite and (2) a protocol to assess a nuclear graphite grade.

36 MATERIALS SCIENCE↗

Tritium Fires: Simulation and Safety Assessment

This is the Sandia report from a joint NSRD project between Sandia National Labs and Savannah River National Labs. The project involved development of simulation tools and data intended to be useful for tritium operations safety assessment. Tritium is a synthetic isotope of hydrogen that has a limited lifetime, and it is found at many tritium facilities in the form of elemental gas (T 2 ). The most serious risk of reasonable probability in an accident scenario is when the tritium is released and reacts with oxygen to form a water molecule, which is subsequently absorbed into the human body. This tritium oxide is more readily absorbed by the body and therefore represents a limiting factor for safety analysis. The abnormal condition of a fire may result in conversion of the safer T 2 inventory to the more hazardous oxidized form. It is this risk that tends to govern the safety protocols. Tritium fire datasets do not exist, so prescriptive safety guidance is largely conservative and reliant on means other than testing to formulate guidelines. This can have a consequence in terms of expensive and/or unnecessary mitigation design, handling protocols, and operational activities. This issue can be addressed through added studies on the behavior of tritium under representative conditions. Due to the hazards associated with the tests, this is being approached mainly from a modeling and simulation standpoint and surrogate testing. This study largely establishes the capability to generate simulation predictions with sufficiently credible characteristics to be accepted for safety guidelines as a surrogate for actual data through a variety of testing and modeling activities.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗