Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Towards Integrating Data Quality Assessments and Radiometer Uncertainty for Determining the Expanded Uncertainty of Three-Component Solar Radiation Measurements

Accurate solar irradiance data are fundamental for determining the design and performance characteristics of photovoltaic systems. The uncertainty of solar irradiance measurements depends on many factors including radiometer design, calibration, installation, maintenance, and operational environment. The key contributors to this uncertainty can be classified as the measurement uncertainty of a particular radiometer and the operational uncertainty determined for the time of measurement. Radiometer measurement uncertainty estimates (U R ) can be based on well-established methods used as part of the radiometer calibration process. Estimates of operational uncertainties (U o ) require consideration of additional site-specific factors that affect data quality. A method is needed for establishing the accuracy of solar irradiance data by integrating an existing data quality process and measurement uncertainty estimates for specific radiometers. An algorithm has been developed to integrate data quality analyses and measurement uncertainty estimates for three-component solar irradiance data: global horizontal (total hemispheric) irradiance, direct normal (beam) irradiance, and diffuse horizontal (sky) irradiance collected at one- to 60-minute intervals. The algorithm has been tested using one-minute irradiance measurements. The goal of the project is to distribute a user-friendly software package based on the new algorithm.

data integrity↗

An integrated data management and informatics framework for continuous drug product manufacturing processes: A case study on two pilot plants

The pharmaceutical industry continuously looks for ways to improve its development and manufacturing efficiency. In recent years, such efforts have been driven by the transition from batch to continuous manufacturing and digitalization in process development. To facilitate this transition, integrated data management and informatics tools need to be developed and implemented within the framework of Industry 4.0 technology. Here, in this regard, the work aims to guide the data integration development of continuous pharmaceutical manufacturing processes under the Industry 4.0 framework, improving digital maturity and enabling the development of digital twins. This paper demonstrates two instances where a data integration framework has been successfully employed in academic continuous pharmaceutical manufacturing pilot plants. Details of the integration structure and information flows are comprehensively showcased. Approaches to mitigate concerns in incorporating complex data streams, including integrating multiple process analytical technology tools and legacy equipment, connecting cloud data and simulation models, and safeguarding cyber-physical security, are discussed. Critical challenges and opportunities for practical considerations are highlighted.

59 BASIC BIOLOGICAL SCIENCES↗

Towards physics-inspired data-driven weather forecasting: integrating data assimilation with a deep spatial-transformer-based U-NET in a case study with ERA5

Abstract. There is growing interest in data-driven weather prediction (DDWP), e.g., using convolutional neural networks such as U-NET that are trained on data from models or reanalysis. Here, we propose three components, inspired by physics, to integrate with commonly used DDWP models in order to improve their forecast accuracy. These components are (1) a deep spatial transformer added to the latent space of U-NET to capture rotation and scaling transformation in the latent space for spatiotemporal data, (2) a data-assimilation (DA) algorithm to ingest noisy observations and improve the initial conditions for next forecasts, and (3) a multi-time-step algorithm, which combines forecasts from DDWP models with different time steps through DA, improving the accuracy of forecasts at short intervals. To show the benefit and feasibility of each component, we use geopotential height at 500 hPa (Z500) from ERA5 reanalysis and examine the short-term forecast accuracy of specific setups of the DDWP framework. Results show that the spatial-transformer-based U-NET (U-STN) clearly outperforms the U-NET, e.g., improving the forecast skill by 45 %. Using a sigma-point ensemble Kalman (SPEnKF) algorithm for DA and U-STN as the forward model, we show that stable, accurate DA cycles are achieved even with high observation noise. This DDWP+DA framework substantially benefits from large (O(1000)) ensembles that are inexpensively generated with the data-driven forward model in each DA cycle. The multi-time-step DDWP+DA framework also shows promise; for example, it reduces the average error by factors of 2–3. These results show the benefits and feasibility of these three components, which are flexible and can be used in a variety of DDWP setups. Furthermore, while here we focus on weather forecasting, the three components can be readily adopted for other parts of the Earth system, such as ocean and land, for which there is a rapid growth of data and need for forecast and assimilation.

54 ENVIRONMENTAL SCIENCES↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Enabling Data Exchange and Data Integration with the Common Information Model: An Introduction for Power Systems Engineers and Application Developers

The Common Information Model (CIM) is an open-source information model that is used to model an electrical network and the various equipment used on the network. CIM is widely used for data exchange of bulk transmission power systems and is finding increasing use for distribution systems. Use of a non-proprietary information model (such as CIM) that has been agreed upon and adopted by numerous utilities, vendors, and researchers allows significant reduction in the effort and cost of data integration. Likewise, adoption of open data platforms built around the CIM increases available functionalities for managing and optimizing the smart grid of the future. This report is intended as an introduction to CIM for utility engineers, power systems researchers, and application developers, providing a broad view of the CIM and how particular profiles can be adapted for various use cases. Unlike most other CIM introduction documents and the International Electrotechnical Commission (IEC) standards (which are mostly targeted to an audience of data scientists, enterprise database managers, and platform developers), this report is intended for users of traditional power systems analysis software and other readers without any prior experience with canonical information models, data profiles, or UML modeling.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Low Yield Nuclear Monitoring Physics Experiment 1 – Integrated Data Acquisition System Design and Initial Observations

The report documents the design of the Integrated Data AcQuisition (IDAQ) system and observations recorded during the first in a series of underground chemical explosions conducted on the Nevada National Security Site (NNSS) in southern Nevada. Experiments are funded as part of Low Yield Nuclear Monitoring (LYNM) research and development within the United States National Nuclear Security Administration NA-22 nuclear non-proliferation program. The series is part of the broader Physical Experiment 1 (PE1) being conducted in and around the P-tunnel facility on the NNSS. Each explosive experiment utilizes several tons of comp-B to generate signals recorded by a broad suite of instrumentation. The IDAQ serves as the backbone for all subsurface instrumentation providing precise time synchronization, remote control, data exfiltration and backup, along with recording several sensing modalities throughout the underground complex that includes ground motion, environmental conditions, and electromagnetic signals.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A data integration framework of additive manufacturing based on FAIR principles

Abstract Laser-powder bed fusion (L-PBF) is a popular additive manufacturing (AM) process with rich data sets coming from both in situ and ex situ sources. Data derived from multiple measurement modalities in an AM process capture unique features but often have different encoding methods; the challenge of data registration is not directly intuitive. In this work, we address the challenge of data registration between multiple modalities. Large data spaces must be organized in a machine-compatible method to maximize scientific output. FAIR (findable, accessible, interoperable, and reusable) principles are required to overcome challenges associated with data at various scales. FAIRified data enables a standardized format allowing for opportunities to generate automated extraction methods and scalability. We establish a framework that captures and integrates data from a L-PBF study such as radiography and high-speed camera video, linking these data sets cohesively allowing for future exploration. Graphical abstract

36 MATERIALS SCIENCE↗

LSTM-Based Data Integration to Improve Snow Water Equivalent Prediction and Diagnose Error Sources

Accurate prediction of snow water equivalent (SWE) can be valuable for water resource managers. Recently, deep learning methods such as long short-term memory (LSTM) have exhibited high accuracy in simulating hydrologic variables and can integrate lagged observations to improve prediction, but their benefits were not clear for SWE simulations. Here we tested an LSTM network with data integration (DI) for SWE in the western United States to integrate 30-day-lagged or 7-day-lagged observations of either SWE or satellite-observed snow cover fraction (SCF) to improve future predictions. SCF proved beneficial only for shallow-snow sites during snowmelt, while lagged SWE integration significantly improved prediction accuracy for both shallow- and deep-snow sites. The median Nash–Sutcliffe model efficiency coefficient (NSE) in temporal testing improved from 0.92 to 0.97 with 30-day-lagged SWE integration, and root-mean-square error (RMSE) and the difference between estimated and observed peak SWE values d max were reduced by 41% and 57%, respectively. DI effectively mitigated accumulated model and forcing errors that would otherwise be persistent. Moreover, by applying DI to different observations (30-day-lagged, 7-day-lagged), we revealed the spatial distribution of errors with different persistent lengths. For example, integrating 30-day-lagged SWE was ineffective for ephemeral snow sites in the southwestern United States, but significantly reduced monthly-scale biases for regions with stable seasonal snowpack such as high-elevation sites in California. These biases are likely attributable to large interannual variability in snowfall or site-specific snow redistribution patterns that can accumulate to impactful levels over time for nonephemeral sites. These results set up benchmark levels and provide guidance for future model improvement strategies.

54 ENVIRONMENTAL SCIENCES↗