Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Lidar / Processed Data

This dataset contains processed, standardized data from the UND scanning Doppler lidar at WFIP3's BARG site, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.

17 WIND ENERGY↗

NANT Site - Lidar / Processed Data Reformatted

This dataset contains standardized data from the PNNL scanning Doppler lidar (S/N 184), consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.

17 WIND ENERGY↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Data Privacy for the Grid: Toward a Data Privacy Standard for Inverter-Based and Distributed Energy Resources

The traditional approach to planning the distribution grid has focused on reliability in the context of gradual and reasonably predictable load growth. Forecasts of load growth, combined with asset management practices, were used by system planners to identify upgrades to the system to maintain or improve reliability. The decisions, typically based within load flow analysis tools, included considerations about contingency scenarios and corporate forecasts (i.e., top-down predictions at a summary level of what would happen in a particular area that could impact load growth and behavior). As a result, today, this traditional approach no longer fits all purposes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

HPC and Cloud Convergence Beyond Technical Boundaries: Strategies for Economic Sustainability, Standardization, and Data Accessibility

At the IEEE/ACM International Conference for High-Performance Computing, Networking, Storage, and Analysis (SC23), held in Denver, experts discussed the convergence of high-performance computing and cloud computing. Experts explored how this integration could address current scientific computing limitations, enhance computational capabilities, and foster global collaboration while focusing on economic, security, technical, and community challenges and opportunities.

97 MATHEMATICS AND COMPUTING↗

CityBES v2021

City Buildings, Energy, and Sustainability (CityBES) is a web-based data and computing platform, focusing on energy modeling and analysis of a city's building stock to support district or city-scale building energy efficiency programs. CityBES uses an international open data standard, CityGML, to represent and exchange 3D city models. CityBES employs EnergyPlus to simulate building energy use and savings from energy efficient retrofits. Other CityBES features include energy benchmarking, district heating and cooling system modeling, rooftop PV analysis, building performance visualization, heat resilience modeling, as well as urban scale mapping of microclimate and heat vulnerability at census tract level. Different from other tools, CityBES uses integrated open and standard 3D city building data and models each individual building using EnergyPlus. CityBES can be used by urban planners, city energy managers, building owners, utilities, energy consultants and researchers.

Hong, Tianzhen↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

2025 TEM Workshop

The TEM Data Management Workshop will take place on August 26 from 9 a.m. to 12 p.m. MT, and will be held virtually on TEAMS. The primary goal of this workshop is to engage NSUF users and stakeholders in discussions about the data needs for the utilization of AI and ML in the analysis of TEM data. Key topics to be covered include data storage, data sharing, data tagging, metadata inclusion, standardized data formats, data augmentation, and annotated training datasets. Additionally, the workshop will provide valuable insights into resources such as the Nuclear Research Data System (NRDS) for data storage and sharing, as well as open-source codes for data analysis.

Bachhav, Mukesh↗

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Comparability of Liquid Chromatography Tandem Mass Spectrometry Analysis of Dissolved Organic Matter across Laboratories

Non-targeted liquid chromatography tandem highresolution mass spectrometry (LC−MS/MS) is increasingly applied for the structure-resolved chemical analysis of dissolved organic matter (DOM). With new developments in MS instrumentation and analysis software, the approach has gained substantial momentum over the past decade. However, achieving high-quality analytical data that is reproducible and comparable across laboratories can be a bottleneck in non-targeted metabolomics and organic matter chemical analysis, especially for data reuse in repository-scale analyses. Understanding the capabilities as well as challenges of comparing LC−MS/MS data from different laboratories is necessary for inferring global trends from public data sets. To illuminate instrumentation factors that drive differences and variability, we used a standardized data analysis pipeline, including classical (CMN) and featurebased molecular networking (FBMN), to analyze data from a ring trial by 24 laboratories on identical sample sets of algal and DOM extracts that were mixed in predefined concentrations and spiked with standards. Our results showed that data sets from similar mass spectrometer types with unified instrument parameters were qualitatively comparable, resolving the same general trends and shared mass spectral features. Interlaboratory comparability was best for high-intensity features, while low-intensity features showed greater detection variability. Our analysis also highlights challenges when comparing data from instruments with different acquisition rates or operating with less standardized methods. Lastly, we provide recommendations for data integration, public data sharing, standardization, and best practices for standardized LC−MS/MS data acquisition, which will be critical for long-term time series and intercomparability of DOM chemical analyses.

DOM↗

Best Practices for Nuclear Experiment Data Preservation at Idaho National Laboratory: A Guide for Researchers and Reactor Operators

Preserving experimental data is essential for supporting advancements in nuclear science and ensuring the longevity of Idaho National Laboratory's contributions to reactor technology and safety. This report provides a comprehensive guide to best practices for experimental data management and preservation, focusing on standardized data formats, redundancy in storage, metadata documentation, and alignment with international standards. By following these recommendations, experimentalists and reactor operators can enhance the accessibility, reproducibility, and utility of critical datasets for regulatory review, validation computational methods, and future research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

RC-SFA Data Management Templates and Guidance for Standardized, Reusable AI-Ready Data Packages

This data package provides templates and supporting documentation developed by the River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) to communicate its approach to managing and publishing AI-ready data. The package is intended to help data users and data producers understand the structures, metadata practices, and quality-control approaches that support consistent, reusable, and machine-actionable data products across RC-SFA studies. Rather than focusing on a single experimental dataset, this package documents the data management framework used to make RC-SFA data easier to find, ingest, navigate, and interpret. The materials in this package reflect RC-SFA practices for standardized data package organization, including the use of a human- and machine-readable README, file-level metadata, data dictionaries, descriptive file naming, method identifiers, and automated and review-based quality assurance procedures. Together, these components illustrate how RC-SFA extends FAIR data principles toward AI-readiness by prioritizing deep metadata, consistency across data packages, and support for informed downstream reuse by both humans and computational tools. This dataset is comprised of (1) readme; (2) presentation slides with an overview of RC-SFA approach and guidance; (3) document of RC-SFA best practices; (4) data dictionary (dd); (5) file level metadata (flmd); and a subfolder containing templates for dd and flmd. All files are .csv and .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

AI-readiness↗

Improving and Automating Building Model Data Exchange

There are many instances throughout a project’s lifecycle where there arises a need for quick and accurate risk assessment of building designs. For example, an unexpected design change during construction may necessitate structural engineers to perform a seismic risk assessment on analytical models of the updated building design using high fidelity structural analysis software, such as ANSYS or Abaqus. However, the efficiency of such workflows often depends upon the interoperability of architectural design software and structural analysis software. When the quality of this interoperability is lacking or even non-existent, the efficiency of virtual engineering workflows is hampered, which increases project costs. A McGraw Hill industry survey of professional users of Building Information Modeling (BIM) technologies found that there is high demand for BIM interoperability for structural analysis, but that the value/difficulty ratio is currently too low for practical use. There have been efforts by the academic community to facilitate model data exchange between the architectural design and structural analysis domains, but such solutions have not been widely adopted by industry, face technical challenges, and oftentimes are limited in applicability for users of various BIM software. Therefore, INL is developing capabilities to improve, automate, and generalize model data exchange between architectural BIM software (e.g., Revit) and structural analysis software (e.g., SAP2000, ANSYS). The goal is to help expedite and automate as much of the pre-processing step for creating analytical models in finite element analysis software as reasonably as possible. Such a "BIM-to-FEA" conversion tool should provide direct benefit to end-users through accuracy, automation, quick turn-around, and wide applicability. To generalize the application of this BIM-to-FEA conversion tool and increase its useability among the many different commercial BIM software currently used by industry, the program is being developed with the concept of openBIM. OpenBIM is the application of non-proprietary, open data standards that allow for BIM model data exchange in a format that is accessible, retainable, and useable for all users. The most widely used open, non-proprietary data exchange format for BIM is the Industry Foundation Classes (IFC) schema. IFC is developed by buildingSMART international and is ISO certified (ISO 16739-1:2018). The BIM-to-FEA conversion tool is being developed for compatibility with typical commercial building designs of steel framed structures. The tool is currently capable of importing architectural BIM data of framed building structures, recognizing and extracting the aspects of the model that are required for structural analysis, adjusting the connectivity of frame members, and finally exporting to an analytical model stored in the IFC format. The exported IFC analytical model can then be imported into various openBIM compliant software, such as SAP2000. Such capabilities have already been tested on commercial software, as shown above, and continue to be improved. Work is underway to test the conversion on various commercial BIM software, develop a user-friendly interface, incorporate the program into the broader DeepLynx data warehouse project being developed by INL, and to eventually open-source the tool for the benefit of the community. Future development of the tool envisions the ability for efficient iterative risk assessment of generative building designs, all within a workflow utilizing open-source tools. One such open-source tool will be MOOSE, an advanced finite element analysis tool developed at INL. The conversion tool will also branch out from typical commercial building designs and will aim to incorporate nuclear construction. The aim will be to convert both structural and non-structural components of nuclear facilities, such as curved concrete containment structures and piping systems, respectively.

97 MATHEMATICS AND COMPUTING↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

AmeriFlux BASE data pipeline to support network growth and data sharing

Abstract AmeriFlux is a network of research sites that measure carbon, water, and energy fluxes between ecosystems and the atmosphere using the eddy covariance technique to study a variety of Earth science questions. AmeriFlux’s diversity of ecosystems, instruments, and data-processing routines create challenges for data standardization, quality assurance, and sharing across the network. To address these challenges, the AmeriFlux Management Project (AMP) designed and implemented the BASE data-processing pipeline. The pipeline begins with data uploaded by the site teams, followed by the AMP team’s quality assurance and quality control (QA/QC), ingestion of site metadata, and publication of the BASE data product. The semi-automated pipeline enables us to keep pace with the rapid growth of the network. As of 2022, the AmeriFlux BASE data product contains 3,130 site years of data from 444 sites, with standardized units and variable names of more than 60 common variables, representing the largest long-term data repository for flux-met data in the world. The standardized, quality-ensured data product facilitates multisite comparisons, model evaluations, and data syntheses.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Multi-Site Observational Study to Assess Biomarkers for Susceptibility or Resilience to Chronic Pain: The Acute to Chronic Pain Signatures (A2CPS) Study Protocol

Chronic pain has become a global health problem contributing to years lived with disability and reduced quality of life. Advances in the clinical management of chronic pain have been limited due to incomplete understanding of the multiple risk factors and molecular mechanisms that contribute to the development of chronic pain. The Acute to Chronic Pain Signatures (A2CPS) Program aims to characterize the predictive nature of biomarkers (brain imaging, high-throughput molecular screening techniques, or “omics,” quantitative sensory testing, patient-reported outcome assessments and functional assessments) to identify individuals who will develop chronic pain following surgical intervention. The A2CPS is a multisite observational study investigating biomarkers and collective biosignatures (a combination of several individual biomarkers) that predict susceptibility or resilience to the development of chronic pain following knee arthroplasty and thoracic surgery. This manuscript provides an overview of data collection methods and procedures designed to standardize data collection across multiple clinical sites and institutions. Pain-related biomarkers are evaluated before surgery and up to 3 months after surgery for use as predictors of patient reported outcomes 6 months after surgery. The dataset from this prospective observational study will be available for researchers internal and external to the A2CPS Consortium to advance understanding of the transition from acute to chronic postsurgical pain.

60 APPLIED LIFE SCIENCES↗