Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Bias Correcting NOAA's High-Resolution Rapid Refresh (HRRR) Wind Resource Data for Grid Integration Applications [Slides]

Many weather years of high-quality wind data are widely accepted in the grid integration community to be important for studying wind energy technical potential, energy system operations, and grid resilience. NREL makes high-quality wind and solar resource data available. NREL's Grid-Atmosphere workshop (March 2024) identified NREL National Solar Radiation Database as widely used in grid integration modeling, but there is less agreement on commonly used wind datasets. One important factor identified by ESIG's 2023 report 'Weather Dataset Needs for Planning and Analyzing Modern Power Systems' for gold standard wind data is regular updates. To address the need for regular updates, NREL's team can now process all currently available and regularly updated High-Resolution Rapid Refresh (HRRR) outputs. HRRR is an hourly-updated operational forecast product produced by the National Oceanic and Atmospheric Administration (NOAA) (Dowell et al., 2022). One barrier to NREL using HRRR is systematic bias and consistency with NREL's existing wind datasets (e.g. WIND Toolkit, 'WTK') across weather years. To address this barrier, we show that the HRRR can be interpolated and bias-corrected to be consistent with NRE's existing datasets. We call the new dataset BC-HRRR (bias-corrected HRRR). As with historical datasets like the WTK, BC-HRRR is intended for use in grid integration modeling (e.g., capacity expansion, production cost, and resource adequacy modeling). BC-HRRR's (2015-present) consistency with WTK (2007-2013) allows NREL to extend internal grid integration tooling with 15+ weather years of wind data with low-overhead extensibility to future years as they are made available by NOAA. The rest of this slide deck documents the BC-HRRR processing methods, validation, and its implications for intended use.

17 WIND ENERGY↗

Model-based interface design for smart field-device integration

Operational complexity is ever-increasing for electric utilities that face challenges including integration of DERs, customer expectation of energy choices, the proliferation of non-utility-owned resources, new business models with energy service providers, and new technology with IT/OT convergence. To maintain and improve the quality of operations, planning, and decision-making in general, utilities need to manage and navigate the complexity. Managing complexity requires a modular, scalable, and flexible solution. Connecting large amounts of DERs and introducing new services requires increased grid control and evolving applications. In this paper, we show an approach utilizing model-based standardized interfaces that simplifies integration and deployment of new algorithms and smart field devices for interoperability across legacy or new systems. A modular design is presented for a reference implementation of widely used DNP3 and IEEE 2030.5 interfaces within an open-source, standards-based data integration platform for integration of smart field devices with independently developed, best-of-breed applications.

Model-driven development, system integration, smar↗

Ten questions concerning Large Language Models (LLMs) for building applications

Large Language Models (LLMs) are emerging as powerful AI tools capable of transforming how building information is collected, processed, analyzed, and applied across diverse research areas. Their capabilities can help building operators, facility managers and other stakeholders such as designers, architects and engineers by providing actionable insights for decision-making across planning, construction, operations, and maintenance of buildings and facilities. This paper explores ten key questions concerning the role of LLMs in shaping sustainable, intelligent, and human-centric buildings. From fundamental definitions to advanced applications, we examine how LLMs facilitate decision-making across the life cycle of buildings and energy systems. LLMs can enhance life cycle assessments (LCA), building energy simulations, and real-time data integration, empowering more efficient and adaptive human-AI environments. They can also contribute to streamlining regulatory compliance, improving post-occupancy evaluations, and fostering more inclusive and participatory design processes. Additionally, this paper addresses the ethical challenges posed by LLMs, such as bias, data privacy, and environmental impacts, and explores their potentials in advancing intelligent digital twins (DT) for ongoing building operations and maintenance. Built upon our applied research using LLMs and the review of tools, datasets, and research gaps, we provide a forward-looking perspective on how LLMs can drive innovation, collaboration, and productivity in the built environment while supporting ethical and effective implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Security assessment and impact analysis of cyberattacks in integrated T&D power systems

In this paper, we examine the impact of cyberattacks in an integrated transmission and distribution (T&D) power grid model with distributed energy resource (DER) integration. We adopt the OCTAVE Allegro methodology to identify critical system assets, enumerate potential threats, analyze and prioritize risks for threat scenarios. Based on the analysis, attack strategies and exploitation scenarios are identified which could lead to system compromise. Specifically, we investigate the impact of data integrity attacks in inverted-based solar PV controllers, control signal blocking attacks in protective switches and breakers, and coordinated monitoring and switching time-delay attacks. Index Terms—Cyberattacks, security assessment, impact analysis, case studies, integrated power systems.

14 SOLAR ENERGY↗

KG-Hub—building and exchanging biological knowledge graphs

Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of KGs is lacking. Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of KGs. Features include a simple, modular extract–transform–load pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial–environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate KGs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph ML, including node embeddings and training of models for link prediction and node classification.

59 BASIC BIOLOGICAL SCIENCES↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Quantifying Groundwater Response and Uncertainty in Beaver‐Influenced Mountainous Floodplains Using Machine Learning‐Based Model Calibration

Abstract Beavers ( Castor canadensis ) alter river corridor hydrology by creating ponds and inundating floodplains, and thereby improving surface water storage. However, the impact of inundation on groundwater, particularly in mountainous alluvial floodplains with permeable gravel/cobble layers overlain by a soil layer, remains uncertain. Numerical modeling across various floodplain structures considers topographic and sediment complexity and multidirectional flow, linking inundation to groundwater response. This study develops a model‐data integration workflow to address uncertainty in groundwater response to beaver‐induced inundations in a mountainous alluvial floodplain in the Upper Colorado River Basin. Uncertain factors include seasonal hydrologic dynamics, hydraulic conductivities, floodplain structures, and meteorological forcings. We employed an ensemble of groundwater models, based on geophysical and hydrologic data, with machine learning‐based calibration using a neural density estimator. This allowed us to quantify the vertical flux from the soil layer to the permeable gravel bed, the down‐valley underflow within the gravel bed, and their ratios. Results show a significant increase in the vertical flux relative to down‐valley underflow, from 2 during dry pond periods to 20 during wet periods, serving as an analogy for conditions without and with beaver ponds. The study highlights the influence of floodplain structure on groundwater storage, water balance, and water quality impacted by beaver ponds. A thick gravel bed layer, with a large down‐valley underflow, minimizes the effect of beaver‐induced inundation on water quality. We emphasize the need for field‐scale measurements of floodplain structure and improved characterization of evapotranspiration changes to reduce uncertainty in groundwater response. Plain Language Summary Beavers change the flow of water in river corridors by creating ponds, expanding wetlands, and flooding floodplains. This increases surface water area, promotes plant growth, and enhances biodiversity. However, the impact of this flooding on groundwater flow is not well understood, especially in mountainous areas with gravel layers where water moves easily beneath soil. In this study, we used numerical modeling to investigate how beaver ponds influence groundwater in a mountainous floodplain of the Upper Colorado River Basin. We adapted a machine learning method to validate our numerical models using multiple field data sets. Our findings show that beaver ponds significantly increase vertical water flow from the soil to the gravel during wet periods, compared to when the ponds are fully drained. The study also highlights the importance of floodplain structure in controlling both water flow in gravel layers along the river direction and vertical flow from the soil to the gravel with the presence of beavers. To reduce uncertainty in groundwater response, we emphasize the need for more field‐scale measurements of floodplain structure, hydraulic properties, and evapotranspiration changes. Key Points Floodplain structures and hydraulic conductivities are important for groundwater response with beaver ponds in mountainous floodplains Large down‐valley underflow in permeability‐stratified floodplains reduces beaver‐induced impacts on groundwater storage and water quality Machine learning‐based model calibration methods are effective for estimating posterior distributions of groundwater model parameters

Wang, Lijing↗

ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability

Earth system predictability is challenged by the complexity of environmental dynamics and the multitude of variables involved. Current AI foundation models, although advanced by leveraging large and heterogeneous data, are often constrained by their size and data integration, limiting their effectiveness in addressing the full range of Earth system prediction challenges. To overcome these limitations, we introduce the Oak Ridge Base Foundation Model for Earth System Predictability (ORBIT), an advanced vision transformer model that scales up to 113 billion parameters using a novel hybrid tensor-data orthogonal parallelism technique. As the largest model of its kind, ORBIT surpasses the current climate AI foundation model size by a thousandfold. Performance scaling tests conducted on the Frontier supercomputer have demonstrated that ORBIT achieves 684 petaFLOPS to 1.6 exaFLOPS sustained throughput, with scaling efficiency maintained at 41% to 85% across 49,152 AMD GPUs. These breakthroughs establish new advances in AI-driven climate modeling and demonstrate promise to significantly improve the Earth system predictability.

Wang, Xiao↗

Blockchain-Enabled Secure Device-to-Device Communication in Software-Defined Networking

The Internet of Things (IoT) continues to increase the demand for seamless communication among IoT devices. The rapid growth of IoT devices has led to an exponential increase in device-to-device (D2D) communication within the Software-Defined Networking (SDN), though it enables a flexible archi-tecture for managing network resources. However, traditional security models face challenges (e.g., Security, privacy, and trust) in addressing the dynamic and decentralized nature of these communications. Despite of these challenges, this paper proposes a novel approach that leverages blockchain technology to enhance the security, privacy, and trustworthiness of D2D communication within an SDN environment. The proposed approach integrates blockchain nodes in sDN components to establish a decentralized ledger for transparent and verifiable records. Smart contracts enforce authentication rules to ensure that only authenticated devices can access the network and engage in transactions securely. It also automates the security policies to ensure temper resistance execution using the cryptographic mechanism for data integrity and authentic communication. The Implementation of the proposed algorithms validates the resilience of the proposed approach against cyberattacks. Overall, the proposed approach enables efficient and secure D2D communication for resilient SDN infrastructure in IoT ecosystems.

Das, Debashis↗

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learned Empirical Numerical Integrator from Simulated Data

Recently, a number of state-of-the-art surrogate machine learning (ML) models have been designed for global weather and climate prediction, which have been trained using reanalysis data products. Reanalysis data products are constructed using numerical model simulations that combine numerical integration of partial differential equations and parameterization schemes. These products are typically only archived and made available using coarsened spatial and temporal resolutions. This study explores the impact of the numerical generation methods used to produce the training datasets and the temporal resolution of those datasets on machine learning surrogate models. Using the nonlinear vector autoregression (NVAR) machine as an explainable ML technique, simple dynamical systems are emulated with ML models trained on data produced by three classical numerical integration schemes. NVAR is validated as a skillful ML method, capable of producing accurate predictions and, more importantly, reconstructing both the underlying dynamics and the numerical integration scheme used to generate the training data. However, the machine fails to generalize predictions on unseen test data generated by different numerical integration schemes, despite the underlying dynamical system being the same. This result provides a word of caution for the growing field of machine learning emulation of weather and climate dynamics. Furthermore, we illustrate using NVAR that training on temporally coarsened data may increase the required complexity of ML models and potentially introduce new numerical challenges. Finally, we discover that empirical integration schemes with arbitrary time-stepping sizes can be constructed directly from the data, which implies a potential for the development of empirical numerical integration schemes.

54 ENVIRONMENTAL SCIENCES↗

Neutron Absorber Plate Characterization Plan for Criticality Experiments Design [Slides]

The goal of this presentation is to show the feasibility of performing a critical experiment with a B 4 C/Aluminum alloy neutron absorber plate. The motivation of this experiment is to provide integral data on modern neutron absorber product “Boralcan” by Rio Tinto. In 2020, used in about 20% of spent fuel pools in the US, and in casks for storage and transportation in the US and in Europe. The added value of this experiment: Neutron absorber plate modeling method verification: homogeneous or not, a lot of potential savings by relaxing boron loading credit limits. No available critical benchmark involving Boralcan, other boron-related experiments are relatively old. Nuclear data validation for B, C, and Al in the form of B 4 C in Al matrix.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking (FY2020 Progress Report)

The development of algorithms for machine learning and data analysis for the 3013 Surveillance Program is a collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). For corrosion detection, Laser Confocal Microscope (LCM) or Wide Area 3D Measurement System (WAMS) data is extracted from large binary files, with software written to convert the data to physical attributes (e.g. height, color and grayscale values; all as functions of a location in a plane projection). It is the objective of this project to produce a user-friendly interface that incorporated all operations needed to perform surface examination. For this reason, a Matlab-based Graphical User Interface (GUI) was created to integrate data input with software developed for processing and evaluation. In summary, the GUI permits selective downloading of binary data, interrogation of attributes, data labeling, flagging of significant features, execution of Machine Learning (ML) algorithms, output of parameters for trained ML algorithms, reports of ML model accuracy with respect to labeled data, and generation of graphical representations of various analyses.

3013 corrosion↗

Stratigraphically Controlled Stress Variations at the Hydraulic Fracture Test Site-1 in the Midland Basin, TX

We investigated the relationship between stratigraphy, stress, and microseismicity at the Hydraulic Fracture Test Site-1. The site comprises two sets of horizontal wells in the Wolfcamp shale and a deviated well drilled after hydraulic fracturing. Regional stresses indicate normal/strike-slip faulting with E-W compression. Stress measurements in vertical and horizontal wells show that the minimum principal stress varies with depth. Strata with high clay and organic content show high values of the least compressive stress, consistent with the theory of viscous stress relaxation. By integrating data from core, logs, and the hydraulic fracturing stages, we constructed a stress profile for the Wolfcamp sequence, which predicts how much pressure is required for hydraulic fracture growth. We applied the results to fracture orientation data from image logs to determine the population of pre-existing faults that are expected to slip during stimulation. We also determined microseismic focal plane mechanisms and found slip on steeply dipping planes striking NW, consistent with the orientations of potentially active faults predicted by the stress model. This case study represents a general approach for integrating stress measurements and rock properties to predict hydraulic fracture growth and the characteristics of injection-induced microseismicity.

microseismicity↗

Systematic characterization of human gut microbiome-secreted molecules by integrated multi-omics

The human gut microbiome produces a complex mixture of biomolecules that interact with human physiology and play essential roles in health and disease. Crosstalk between micro-organisms and host cells is enabled by different direct contacts, but also by the export of molecules through secretion systems and extracellular vesicles. The resulting molecular network, comprised of various biomolecular moieties, has so far eluded systematic study. Here we present a methodological framework, optimized for the extraction of the microbiome-derived, extracellular biomolecular complement, including nucleic acids, (poly)peptides, and metabolites, from flash-frozen stool samples of healthy human individuals. Our method allows simultaneous isolation of individual biomolecular fractions from the same original stool sample, followed by specialized omic analyses. The resulting multi-omics data enable coherent data integration for the systematic characterization of this molecular complex. Our results demonstrate the distinctiveness of the different extracellular biomolecular fractions, both in terms of their taxonomic and functional composition. This highlights the challenge of inferring the extracellular biomolecular complement of the gut microbiome based on single-omic data. The developed methodological framework provides the foundation for systematically investigating mechanistic links between microbiome-secreted molecules, including those that are typically vesicle-associated, and their impact on host physiology in health and disease.

59 BASIC BIOLOGICAL SCIENCES↗

Advanced Long-Term Environmental Monitoring Systems (ALTEMIS) Artificial Intelligence Data Management Plan

Across the Department of Energy’s Environmental and Legacy Management sites, complex groundwater plumes exist that will require long-term monitoring to ensure remedial actions that have been put in place remain effective decades into the future. The current monitoring paradigm predominantly consists of groundwater well sampling, whereby samples are collected, concentrations analyzed, and plume anomalies are detected after they have occurred. The Advanced Long Term Environmental Monitoring Systems (ALTEMIS) program is a multi-lab, multi-institution team of researchers that is deploying spatially integrative technologies (i.e., real-time in situ sensor networks), coupled with artificial intelligence and machine learning, to establish a more proactive monitoring paradigm. Within this approach, plume anomalies can be predicted, and corrective actions can be established prior to the occurrence, offering a more cost-effective and robust approach to long-term monitoring. The team has deployed a variety of different in situ sensing technologies at the Savannah River Site’s F-Area Hazardous Waste Management Facility around the F-Area Seepage Basins, which are unlined basins that received 7 billion liters of acidic low-level radioactive waste from the 1950s until the late 1980s. The technologies and techniques that the team is deploying are intended to ensure that the remedial actions that have been taken by the site remain effective decades into the future. Foundational to this approach is a robust, integrated data management and analysis plan to ensure accurate and timely reporting from the variety of sensor systems that are in place. This report will outline the data management plan that has been implemented by the ALTEMIS team at the Savannah River Site and will serve as a blueprint as the technology is translated to new sites across the DOE Complex.

54 ENVIRONMENTAL SCIENCES↗

NETL Coal Energy Atlas: A Collection of Coal/Energy Related Maps

The NETL Coal Energy Atlas contains a comprehensive collection of coal and energy-related maps and graphics curated by the National Energy Technology Laboratory (NETL) Systems Analysis group. It serves as a living document providing an overview of the U.S. coal and energy sectors. The volume is structurally organized into six key thematic areas. Ultimately, the atlas functions as a modular baseline for data integration, allowing researchers to drill down into specific regional locations or customize geographic base layers for advanced systems analysis.

bituminous coal↗

Phase‐Change‐Memory Process at the Limit: A Proposal for Utilizing Monolayer Sb 2 Te 3

Abstract One central task of developing nonvolatile phase change memory (PCM) is to improve its scalability for high‐density data integration. In this work, by first‐principles molecular dynamics, to date the thinnest PCM material possible (0.8 nm), namely, a monolayer Sb 2 Te 3 , is proposed. Importantly, its SET (crystallization) process is a fast one‐step transition from amorphous to hexagonal phase without the usual intermediate cubic phase. An increased spatial localization of electrons due to geometrical confinement is found to be beneficial for keeping the data nonvolatile in the amorphous phase at the 2D limit. The substrate and superstrate can be utilized to control the phase change behavior: e.g., with passivated SiO 2 (001) surfaces or hexagonal Boron Nitride, the monolayer Sb 2 Te 3 can reach SET recrystallization in 0.54 ns or even as fast as 0.12 ns, but with unpassivated SiO 2 (001), this would not be possible. Besides, working with small volume PCM materials is also a natural way to lower power consumption. Therefore, the proposed PCM working process at the 2D limit will be an important potential strategy of scaling the current PCM materials for ultrahigh‐density data storage.

2D limit↗