Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

LSTM-Based Data Integration to Improve Snow Water Equivalent Prediction and Diagnose Error Sources

Accurate prediction of snow water equivalent (SWE) can be valuable for water resource managers. Recently, deep learning methods such as long short-term memory (LSTM) have exhibited high accuracy in simulating hydrologic variables and can integrate lagged observations to improve prediction, but their benefits were not clear for SWE simulations. Here we tested an LSTM network with data integration (DI) for SWE in the western United States to integrate 30-day-lagged or 7-day-lagged observations of either SWE or satellite-observed snow cover fraction (SCF) to improve future predictions. SCF proved beneficial only for shallow-snow sites during snowmelt, while lagged SWE integration significantly improved prediction accuracy for both shallow- and deep-snow sites. The median Nash–Sutcliffe model efficiency coefficient (NSE) in temporal testing improved from 0.92 to 0.97 with 30-day-lagged SWE integration, and root-mean-square error (RMSE) and the difference between estimated and observed peak SWE values d max were reduced by 41% and 57%, respectively. DI effectively mitigated accumulated model and forcing errors that would otherwise be persistent. Moreover, by applying DI to different observations (30-day-lagged, 7-day-lagged), we revealed the spatial distribution of errors with different persistent lengths. For example, integrating 30-day-lagged SWE was ineffective for ephemeral snow sites in the southwestern United States, but significantly reduced monthly-scale biases for regions with stable seasonal snowpack such as high-elevation sites in California. These biases are likely attributable to large interannual variability in snowfall or site-specific snow redistribution patterns that can accumulate to impactful levels over time for nonephemeral sites. These results set up benchmark levels and provide guidance for future model improvement strategies.

54 ENVIRONMENTAL SCIENCES↗

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES↗

The Role of Snowmelt and Subsurface Heterogeneity in Headwater Hydrology of a Mountainous Catchment in Colorado: A Model‐Data Integration Approach

Mountainous headwater streams are sustained by both snowmelt‐driven streamflow and groundwater discharge in the Upper Colorado River Basin. However, predicting headwater stream discharge magnitude and peak flow timing is challenging in mountainous terrains, where snowmelt rates vary with vegetation type and elevation, and heterogeneous subsurface physical properties influence groundwater storage and its release. We used a model‐data integration approach to investigate the roles of snowmelt and subsurface structure in stream discharge and groundwater level. We ran an ensemble of 100 integrated surface‐subsurface hydrologic models for a mountainous headwater catchment near Crested Butte, Colorado, USA. We also evaluated and calibrated these models against observed data sets, including snow depth measurements using distributed temperature probes, stream discharge, and groundwater levels. Calibration with multiple data sources using neural density estimators has further constrained uncertainty in subsurface properties and snowmelt rates. Results indicated that observed slower snowmelt rates in evergreen forests delayed the peak flow and baseflow onset. In upstream areas with lower subsurface permeability, water was stored within the subsurface but was not released as interflow or shallow groundwater flow, and thereby not contributing to downstream streamflow during recession limb periods. Double peaks in groundwater occurred in areas with spatial subsurface heterogeneity, in our case due to the contrast between granodiorite and Mancos shale. These process‐based insights into groundwater and snowmelt dynamics in mountainous headwaters will help improve predictions of headwater hydrology.

Wang, Lijing [University of Connecticut, Storrs, C↗

Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower (DIVERS-H)

U.S. hydropower plants face potential threats from shrinking water supply, rising demands, and warmer stream temperatures from various causes. Power plant owners, operators, and regulators require new tools to take advantage of and interpret the diverse range of scientific data being produced by both observational methods (for example, satellite, radar, stream gauges) and computer modeling methods that evaluate and predict how earth's dynamic systems (atmosphere, oceans, land surface, and sea ice) are changing and interacting. Combining datasets such as these with AI-based analyses introduces a novel decision support system to help users anticipate and address potential impacts on power generation stations. This new technology has been named DIVERS-H for "Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower." In Phase I, technical feasibility was established with the development and demonstration of all the new technologies that are required. Most notably, DIVERS-H will use new artificial intelligence (AI) methods to capture the complex dynamics of water availability, demand, and environmental changes. In addition, new data management software was developed, and a prototype user interface was implemented as the precursor to a full scale decision support system. With technical research complete, the project focus now shifts to development of a commercial software product to provide users with actionable insight into water availability and the risk/resilience of critical systems at their locations of interest. Although DIVER-H was originally conceived as a tool for hydroelectric power applications, the same underlying technology can be readily applied to other water-consuming systems including coal, natural gas, oil, and nuclear power plants.

Chaudhary, Aashish [Kitware, Inc., Clifton Park, N↗

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Machine learning-enabled model-data integration for predicting subsurface water storage

Subsurface water storage (SWS) is a key variable of the climate system and a storage component for precipitation and radiation anomalies, inducing persistence in the climate system. It plays a critical role in climate-change projections and can mitigate the impacts of climate change on ecosystems. However, because of the difficult accessibility of the underground, hydrologic properties and dynamics of SWS are poorly known. Direct observations of SWS are limited, and accurate incorporation of SWS dynamics into Earth system land models remains challenging. We propose a machine learning-enabled model-data integration framework to improve the SWS prediction at local to conus scales in a changing climate by leveraging all the available observation and simulation resources, as well as to inform the model development and guide the observation collection. The accurate prediction will enable an optimal decision of water management and land use and improve the ecosystem's resilience to the climate change.

Lu, Dan↗

Battery data integrity and usability: Navigating datasets and equipment limitations for efficient and accurate research into battery aging

A tremendous commitment of resources is needed to acquire, understand and apply battery data in terms of performance and aging behavior. There are many state of performance (SOP) and state of health (SOH) metrics that are useful to guide alignment of batteries to end-use, yet how these metrics are measured or extracted can make the difference between usable, valuable datasets versus data that lacks the necessary integrity to meet baseline confidence levels for SOP/SOH quantification. This work will speak to 1) types of data that support SOP and SOH evaluations on mechanistic terms, 2) measurement conditions needed to assure high data integrity, 3) equipment limitations that can compromise data high fidelity, and 4) the impact of cell polarization on data quality. A common goal in battery research and field use is to work from a data platform that supports economical paths of data capture while minimizing down-time for battery diagnostics. An ideal situation would be to utilize data obtained during normal daily use (“pulses or cycles of convenience”) without stopping the daily duty cycles to perform dedicated SOP/SOH diagnostic routines. However, difficulties arise in trying to make use of daily duty cycle data (denoted as cycle-by-cycle, CBC) that underscores the need for standardization of conditions: temperature and duty cycles can vary over the course of a day and throughout a week, month and year; polarization can develop within an immediate cycle and throughout successive cycles as a hysteresis. If CBC data is envisioned as a data source to determine performance and aging trends, it should be recognized that polarization is a frequent consequence of CBC and thus makes it difficult to separate reversible and irreversible components to metrics such as capacity loss and resistance increase over aging. Since CBC conditions can have a major impact on data usability, we will devote part of this paper to CBC data conditioning and management. Differential analyses will also be discussed as a means to detect changing trends in data quality. Our target cell chemistries will be lithium-ion types NMC/graphite and LMO/LTO.

25 ENERGY STORAGE↗

sciCAN: single-cell chromatin accessibility and gene expression data integration via cycle-consistent adversarial network

The boom in single-cell technologies has brought a surge of high dimensional data that come from different sources and represent cellular systems from different views. With advances in these single-cell technologies, integrating single-cell data across modalities arises as a new computational challenge. Here, we present an adversarial approach, sciCAN, to integrate single-cell chromatin accessibility and gene expression data in an unsupervised manner. We benchmarked sciCAN with 5 existing methods in 5 scATAC-seq/scRNA-seq datasets, and we demonstrated that our method dealt with data integration with consistent performance across datasets and better balance of mutual transferring between modalities than the other 5 existing methods. We further applied sciCAN to 10X Multiome data and confirmed that the integrated representation preserves biological relationships within the hematopoietic hierarchy. Finally, we investigated CRISPR-perturbed single-cell K562 ATAC-seq and RNA-seq data to identify cells with related responses to different perturbations in these different modalities.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating data types to estimate spatial patterns of avian migration across the Western Hemisphere

For many avian species, spatial migration patterns remain largely undescribed, especially across hemispheric extents. Recent advancements in tracking technologies and high-resolution species distribution models (i.e., eBird Status and Trends products) provide new insights into migratory bird movements and offer a promising opportunity for integrating independent data sources to describe avian migration. Here, we present a three-stage modeling framework for estimating spatial patterns of avian migration. First, we integrate tracking and band re-encounter data to quantify migratory connectivity, defined as the relative proportions of individuals migrating between breeding and nonbreeding regions. Next, we use estimated connectivity proportions along with eBird occurrence probabilities to produce probabilistic least-cost path (LCP) indices. In a final step, we use generalized additive mixed models (GAMMs) both to evaluate the ability of LCP indices to accurately predict (i.e., as a covariate) observed locations derived from tracking and band re-encounter data sets versus pseudo-absence locations during migratory periods and to create a fully integrated (i.e., eBird occurrence, LCP, and tracking/band re-encounter data) spatial prediction index for mapping species-specific seasonal migrations. To illustrate this approach, we apply this framework to describe seasonal migrations of 12 bird species across the Western Hemisphere during pre- and postbreeding migratory periods (i.e., spring and fall, respectively). We found that including LCP indices with eBird occurrence in GAMMs generally improved the ability to accurately predict observed migratory locations compared to models with eBird occurrence alone. Using three performance metrics, the eBird + LCP model demonstrated equivalent or superior fit relative to the eBird-only model for 22 of 24 species–season GAMMs. In particular, the integrated index filled in spatial gaps for species with over-water movements and those that migrated over land where there were few eBird sightings and, thus, low predictive ability of eBird occurrence probabilities (e.g., Amazonian rainforest in South America). This methodology of combining individual-based seasonal movement data with temporally dynamic species distribution models provides a comprehensive approach to integrating multiple data types to describe broad-scale spatial patterns of animal movement. Further development and customization of this approach will continue to advance knowledge about the full annual cycle and conservation of migratory birds.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating Data Centers and Grid Technologies at Scale

This presentation focuses on the challenge of integrating AI-driven data centers with the power grid at scale. It examines the AI data center capacity challenge and the role of new Medium Voltage Direct Current (MVDC) and other grid-enhancing technologies in enabling efficient and reliable power delivery. The session will highlight the National Laboratory of the Rockies' ARIES capabilities and planning tools, along with collaborative examples involving Verrus, Compass, and Schneider through the Agora test bed for grid-friendly data center evaluations, and ON. Energy for UPS evaluation. It will showcase the NLR Stable Grid Platform for studying oscillations caused by large-scale data centers, along with planning tools to assess grid security and reliability. Additionally, the presentation covers reconductoring strategies to increase grid capacity and explores innovative data center architectures, including the Advanced DC Architectures with Power-electronic Transformers (ADAPT) platform, which enables testing of complete DC architectures for data centers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Framework to Demonstrate a DNP3 Interface With a CIM-Based Data Integration Platform: Preprint

The contemporary electrical grid is characterized by its complexity and abundance of data. A control-rich environment supported by information and communication technologies within an Advanced Distribution Management System (ADMS) presents a viable and cost-effective option for utility companies aiming to implement advanced real-time analytical schemes for monitoring and remotely controlling distribution feeders. Modular platform-based approaches to distribution operations require a structured framework for acquiring field device measurements, performing analytics, converting the setpoint to the correct protocol, and sending it on the appropriate communications network to the field devices. We present the development and deployment of an application service to integrate an open-source standardsbased platform with an ADMS test bed with field devices using the Distributed Network Protocol (DNP3) for data exchange. The step-by-step procedure for establishing the DNP3-Master service on an open-source distribution platform is outlined, comprehensively explaining the Master setup process. Moreover, sample use case results highlight the capabilities of the DNP3- Master service setup. Results demonstrate the scalability and configurability of the DNP3-Master service, making it adaptable for integration with other relevant applications, thus providing potential opportunities for real-world field trials and real-time assessments.

ADMS↗

Abridged spectral matrix inversion: parametric fitting of X-ray fluorescence spectra following integrative data reduction

Recent improvements in both X-ray detectors and readout speeds have led to a substantial increase in the volume of X-ray fluorescence data being produced at synchrotron facilities. This in turn results in increased challenges associated with processing and fitting such data, both temporally and computationally. Herein an abridging approach is described that both reduces and partially integrates X-ray fluorescence (XRF) data sets to obtain a fivefold total improvement in processing time with negligible decrease in quality of fitting. The approach is demonstrated using linear least-squares matrix inversion on XRF data with strongly overlapping fluorescent peaks. This approach is applicable to any type of linear algebra based fitting algorithm to fit spectra containing overlapping signals wherein the spectra also contain unimportant (non-characteristic) regions which add little (or no) weight to fitted values, e.g. energy regions in XRF spectra that contain little or no peak information.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Accelerating Nuclear-Integrated Data Centers in the USA: SWOT Analysis, Power-Thermal Management Strategies, and Industrial-Scale Demonstration and Potential Deployment

Driven by the growth in digital services, cloud computing, AI, and manufacturing, data centers face rising energy demands that challenge traditional power sources and cooling efficiency. This study explores using nuclear power to meet these demands, focusing on accelerated reactor technology deployment and highlighting needs such as N+1/N+2 power supplies and integrated power-thermal management. A SWOT analysis addresses grid connectivity, reactors, and site selection, particularly DOE sites. Reactor technology demonstration and deployment could be accelerated by leveraging test facilities such as MARVEL, MAGNET, TED, FAS, DOME, LOTUS, ATR, Energy System Proving Grounds, and upcoming Energy Launch Pads, along with modeling and simulation tools such as RELAP5, MOOSE, VERA, RAVEN, and FORCE. The potential power and thermal management options, including various cooling technologies, waste-heat utilization, and an industrial-scale demonstration plan, aim to accelerate the integration of nuclear power and data centers in the USA, while emphasizing community and stakeholder engagement and synergistic efforts.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Risk-informed Hierarchical Control of Behind-the-Meter DERs with AMI Data Integration (Final Technical Report)

This project addresses several key barriers to implement the next generation demand response applications and provides a clear understanding of implementing hierarchical and standalone control using AMI data. Through this program, Eaton has developed and tested a meter-as-a-controller prototype with the help of other partners--- National Renewable Energy Laboratory (NREL), Electric Power Research Institute (EPRI), Pecan St Inc. (PSI), and Delaware Electric Cooperative (DEC). The controller can utilize residential controllable loads such as heating, ventilation, and air conditioner (HVAC), electric water heater and distributed energy resources like solar PV and battery energy storage systems for off-setting the demand that is required from the grid, thus providing reliable grid-services for demand reduction or peak shaving. The controller is also capable of coordinating the resources of the premises for better management and energy efficiency while meeting the comfort bound of the premises owner as quality-of-service. The development has been demonstrated in a three virtual-home setup at system performance lab of NREL with real appliances (HVAC, electric water heater, solar PV, and battery). The technology has also been proved through laboratory and field demonstration with successful interconnectivity (e.g., end-to-end communication and data exchange) between the residential appliances and utility through the RF network at Delaware Electric Co-op (DEC) in Delaware.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

High-frequency Data Integration for Landscape Model Calibration of Carbon Fluxes Across Diverse Tidal Marshes

Terrestrial Aquatic Interfaces (TAIs), and tidal wetlands in particular, store large amounts of carbon yet are not well represented in Earth System Models (ESMs). Predictions of carbon cycling and greenhouse gas (GHG) emissions in tidal wetlands are highly uncertain. Eddy covariance (EC) towers provide ecosystem-scale GHG flux data at a temporal resolution (every 30min) that is helpful for parameterizing and improving mechanistic realism in ESMs. We propose to use a network of eddy covariance towers and standardized ancillary data streams, along with mesocosm experiments and statistical analyses, across diverse tidal wetlands of North America to develop and improve biogeochemical modeling at the TAI. Our overarching objective is to improve understanding and process-based modeling of gross primary productivity (GPP) and CH4 emission responses, both non-linear and asynchronous, to stressors including plant inundation, disturbance, salinity and nitrogen loading.

54 ENVIRONMENTAL SCIENCES↗

Summary of Thermodynamic Data Integration (M4SF-23LL010301052)

This progress report (Level 4 Milestone Number M4SF-23LL010301052) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite Activity Number M4SF-23LL010301052. SUPCRTNE is being developed as the primary engine for thermodynamic database development in support of geologic disposal of high-level nuclear waste. As used here, “SUPCRT” refers to both a computer program and its supporting database. Additional letters or numbers refer to specific variants (SUPCRT92, Johnson et al., 1992; SUPCRTBL, Zimmer et al., 2016; and SUPCRTNE). A SUPCRT database contains the data required to calculate the thermodynamic properties of solid, gas, and aqueous species over a wide range of temperature and pressure. Normally SUPCRT is used to create a higher-level data base that directly supports modeling and simulation codes such as EQ3/6, GWB, PHREEQC, and PFLOTRAN.

58 GEOSCIENCES↗

Million-scale data integrated deep neural network for phonon properties of heuslers spanning the periodic table

Existing machine learning potentials for predicting phonon properties of crystals are typically limited on a material-to-material basis, primarily due to the exponential scaling of model complexity with the number of atomic species. We address this bottleneck with the developed Elemental Spatial Density Neural Network Force Field, namely Elemental-SDNNFF. The effectiveness and precision of our Elemental-SDNNFF approach are demonstrated on 11,866 full, half, and quaternary Heusler structures spanning 55 elements in the periodic table by prediction of complete phonon properties. Self-improvement schemes including active learning and data augmentation techniques provide an abundant 9.4 million atomic data for training. Deep insight into predicted ultralow lattice thermal conductivity (<1 Wm –1 K –1 ) of 774 Heusler structures is gained by p–d orbital hybridization analysis. Additionally, a class of two-band charge-2 Weyl points, referred to as “double Weyl points”, are found in 68% and 87% of 1662 half and 1550 quaternary Heuslers, respectively.

36 MATERIALS SCIENCE↗