Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Integration of Data Analytics with Plant System Health Program

This report summarizes the R&D activities of the Plant Health Management (PHM) project during fiscal year 2020 (FY20). This project focuses on the development of methods that integrate component health data and propagate such information at the system level to evaluate system sources of risk. This project development lives in cooperation with the Risk Informed Asset Management (RIAM) project. The RIAM project, in fact, uses the reliability models and data generated by the PHM project to optimize plant operations (e.g., maintenance/replacement schedule, optimal maintenance posture). This year’s activities for the PHM project focused mainly on the development of two classes of models. The first one includes a series of component reliability models which include aging, testing and maintenance. The second class includes a series of system reliability models which focus on the secondary side of exiting U.S. reactors. These models are based on fault tree logic structures such that they can be used by existing plant PRA software. We also started to focus on the management of health data and we tackled this issue in two directions. The first one focuses on the integration of monitor data with simulation models to assess component health. This approach is moving from a classical data based to a model+data based approach with the goal of improving higher component health information. The second direction focuses on linking equipment reliability data (e.g., maintenance/failure reports, component monitoring data) directly to system reliability models using a safety margin based language rather than a probability based language. The main advantage of a safety margin based language is that it can provide to system engineers more tangible information on system/component health and how it propagates to the system level.

97 MATHEMATICS AND COMPUTING↗

A Simple Standard for Sharing Ontological Mappings (SSSOM)

Abstract Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Or are they associated in some other way? Such relationships between the mapped terms are often not documented, which leads to incorrect assumptions and makes them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Furthermore, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. We have developed the Simple Standard for Sharing Ontological Mappings (SSSOM) which addresses these problems by: (i) Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. (ii) Defining an easy-to-use simple table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data principles. (iii) Implementing open and community-driven collaborative workflows that are designed to evolve the standard continuously to address changing requirements and mapping practices. (iv) Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases in detail and survey some of the existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable and Reusable (FAIR). The SSSOM specification can be found at http://w3id.org/sssom/spec. Database URL: http://w3id.org/sssom/spec

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Radiation-induced bowing of SiC/SiC composites under neutron flux gradients—integral experimental data for model validation

Here, the radiation-induced swelling of SiC and its composites, including strong dependencies on temperature and dose, can drive significant lateral bowing in the presence of temperature and/or dose gradients. In recent years, simulations have been performed to assess the extent of bowing in SiC composite light-water reactor (LWR) fuel cladding and boiling water reactor (BWR) channel boxes. However, to date, no integral experimental data exist to validate these models. This work provides the first experimental bowing evaluation of three ∼380 mm long SiC composite specimens irradiated under varying neutron dose gradients (∼50°C–60°C, 0.03–0.06 dpa): two tubes (∼9.8 mm diameter) and a miniature BWR channel box (∼30 mm square). The measured radiation-induced length swelling (∼0.3%–0.7% linear) was consistently 10%–21% higher than values obtained from 3D finite element structural analyses with inputs from 3D radiation transport calculations. This discrepancy could be at least partially explained by differences in dose rate (∼10 -8 dpa/s) compared to the literature data (∼10-6 dpa/s) used to establish the dose-to-swelling correlations in the model. Nevertheless, the modeled bowing magnitudes (<2 mm) obtained from finite element analyses and simple analytical equations were within the bounds of the experimental measurements for all specimens. With improved confidence in the ability to predict the structural response and measure the macroscopic deformations, future experiments will target transient bowing under neutron flux gradients at representative LWR temperatures and assess whether grid spacers can mitigate the tens of millimeters of bowing that would otherwise be expected in ∼4 m long LWR components.

bowing↗

Cyber physical attack detection

A cyber-security threat detection system and method stores physical data measurements from a cyber-physical system and extracts synchronized measurement vectors synchronized to one or more timing pulses. The system and method synthesize data integrity attacks in response to the physical data measurements and applies alternating parameterized linear and non-linear operations in response to the synthesized data integrity attacks. The synthesis renders optimized model parameters used to detect multiple cyber-attacks.

Ferragut, Erik M.↗

Assessing carbon storage capacity and saturation across six central US grasslands using data–model integration

Abstract. Future global changes will impact carbon (C) fluxes and pools in most terrestrial ecosystems and the feedback of terrestrial carbon cycling to atmospheric CO2. Determining the vulnerability of C in ecosystems to future environmental change is thus vital for targeted land management and policy. The C capacity of an ecosystem is a function of its C inputs (e.g., net primary productivity – NPP) and how long C remains in the system before being respired back to the atmosphere. The proportion of C capacity currently stored by an ecosystem (i.e., its C saturation) provides information about the potential for long-term C pools to be altered by environmental and land management regimes. We estimated C capacity, C saturation, NPP, and ecosystem C residence time in six US grasslands spanning temperature and precipitation gradients by integrating high temporal resolution C pool and flux data with a process-based C model. As expected, NPP across grasslands was strongly correlated with mean annual precipitation (MAP), yet C residence time was not related to MAP or mean annual temperature (MAT). We link soil temperature, soil moisture, and inherent C turnover rates (potentially due to microbial function and tissue quality) as determinants of carbon residence time. Overall, we found that intermediates between extremes in moisture and temperature had low C saturation, indicating that C in these grasslands may trend upwards and be buffered against global change impacts. Hot and dry grasslands had greatest C saturation due to both small C inputs through NPP and high C turnover rates during soil moisture conditions favorable for microbial activity. Additionally, leaching of soil C during monsoon events may lead to C loss. C saturation was also high in tallgrass prairie due to frequent fire that reduced inputs of aboveground plant material. Accordingly, we suggest that both hot, dry ecosystems and those frequently disturbed should be subject to careful land management and policy decisions to prevent losses of C stored in these systems.

58 GEOSCIENCES↗

Automated Cloud Based Long Short-Term Memory Neural Network Based SWE Prediction

Snow derived water is a critical component of the US water supply. Measurements of the Snow Water Equivalent (SWE) and associated predictions of peak SWE and snowmelt onset are essential inputs for water management efforts. This paper aims to develop an integrated framework for real-time data ingestion, estimation, prediction and visualization of SWE based on daily snow datasets. In particular, we develop a data-driven approach for estimating and predicting SWE dynamics using the Long Short-Term Memory neural network (LSTM) method. Our approach uses historical datasets (precipitation, air temperature, SWE, and snow thickness) collected at NRCS Snow Telemetry (SNOTEL) stations to train the LSTM network and current year data to predict SWE behavior. The performance of our prediction was compared for different prediction dates and prediction training datasets. Our results suggest that the proposed LSTM network can be an efficient tool for forecasting the SWE timeseries, as well as Peak SWE and snowmelt timing. Results showed that the window size impacts the model performance (where the Nash Sutcliffe efficiency (NSE) ranged from 0.96 to 0.85 and the Rooted Mean Square Error (RMSE) ranged from 0.038 to 0.07) with an optimum number that should be calibrated for different stations and climate conditions. In addition, by implementing the LSTM prediction capability in a cloud based site-monitoring platform, we automate model-data integration. By making the data accessible through a graphical web interface and an underlying API which exposes both training and prediction capabilities. The associated results can be made easily accessible to a broad range of stakeholders.

54 ENVIRONMENTAL SCIENCES↗

An overview of data tools for representing and managing building information and performance data

Building information modeling (BIM) has been widely adopted for representing and exchanging building data across disciplines during building design and construction. However, BIM's use in the building operation phase is limited. With the increasing deployment of low-cost sensors and meters, as well as affordable digital storage and computing technologies, growing volumes of data have been collected from buildings, their energy services systems, and occupants. Such data are crucial to help decision makers understand what, how, and when energy is consumed in buildings—a critical step to improving building performance for energy efficiency, demand flexibility, and resilience. However, practical analyses and use of the collected data are very limited due to various reasons, including poor data quality, ad-hoc representation of data, and lack of data science skills. To unlock value from building data, there is a strong need for a toolchain to curate and represent building information and performance data in common standardized terminologies and schemas, to enable interoperability between tools and applications. This study selected and reviewed 24 data tools based on common use cases of data across the building life cycle, from design to construction, commissioning, operation, and retrofits. The selected data tools are grouped into three categories: (1) data dictionary or terminology, (2) data ontology and schemas, and (3) data platforms. The data are grouped into ten typologies covering most types of data collected in buildings. This study resulted in five main findings: (1) most data representation tools can represent their intended data typologies well, such as Green Button for smart meter data and Brick schema for metadata of sensors in buildings and HVAC systems, but none of the tools cover all ten types of data; (2) there is a need for data schemas to represent the basis of design data and metadata of occupant data; (3) standard terminologies such as those defined in BEDES are only adopted in a few data tools; (4) integrating data across various stages in the building life cycle remains a challenge; and (5) most data tools were developed and maintained by different parties for different purposes, their flexibility and interoperability can be improved to support broader use cases. Finally, recommendations for future research on building data tools are provided for the data and buildings community based on the FAIR principles to make data Findable, Accessible, Interoperable, and Reusable.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A scalable, open-source implementation of a large-scale mechanistic model for single cell proliferation and death signaling

Mechanistic models of how single cells respond to different perturbations can help integrate disparate big data sets or predict response to varied drug combinations. However, the construction and simulation of such models have proved challenging. Here, we developed a python-based model creation and simulation pipeline that converts a few structured text files into an SBML standard and is high-performance- and cloud-computing ready. We applied this pipeline to our large-scale, mechanistic pan-cancer signaling model (named SPARCED) and demonstrate it by adding an IFNγ pathway submodel. We then investigated whether a putative crosstalk mechanism could be consistent with experimental observations from the LINCS MCF10A Data Cube that IFNγ acts as an anti-proliferative factor. The analyses suggested this observation can be explained by IFNγ-induced SOCS1 sequestering activated EGF receptors. This work forms a foundational recipe for increased mechanistic model-based data integration on a single-cell level, an important building block for clinically-predictive mechanistic models.

60 APPLIED LIFE SCIENCES↗

Cyber Resilient Design and Controls for Energy Systems

As the grid of today is evolving into a highly distributed and autonomous cyber-physical system with higher penetration of Distributed Energy Resources (DERs) and autonomous controls that are dependent on the cyber infrastructure, there are novel and increasing challenges to cyber-resilient operation of the grid. While the traditional mechanisms take a "bolt on" approach to cybersecurity and resiliency, this talk presents two novel technologies developed at NREL that "bake in" security and resilience into the design and control of the grid. An Adaptive Resilience Metrics framework is a novel mechanism for system owners and operators to understand the impact of cyber and physical system design and topology on the resilient operation, while also helping determine optimal policies in response to adverse cyber events. Along with that a zero-knowledge proof-based technique helps incorporate security into the controls layer by focusing on computational integrity rather than data integrity. By leveraging design and control properties unique to Operational Technology systems these technologies will not only help enhance cyber-resilience for the grid of today, but a highly distributed and autonomous grid of tomorrow.

cyber resilience↗

DIRECT RF SAMPLING BASED LLRF CONTROL SYSTEM FOR C-BAND LINEAR ACCELERATOR

Low Level RF (LLRF) control systems of linear accel- erators (LINACs) are typically implemented with hetero- dyne based architectures, which have complex analog RF mixers for up and down conversion. The Gen 3 Radio Fre- quency System-on-Chip (RFSoC) device from AMD Xilinx integrates data converters with maximum RF frequency of 6 GHz. This enables direct RF sampling of C-band LLRF signal typically operated at 5.712 GHz without any analogue mixers, which can significantly simplify the system architec- ture. The data converters sample RF signals in higher order Nyquist zones and then up or down convert digitally by the integrated data path in RFSoC. The closed-loop feedback control firmware implemented in FPGA integrated in RF- SoC can process the base-band signal from the ADC data path and calculate the updated phase and amplitude to be up- mixed by the DAC data path. We have developed a C-band LLRF control RFSoC platform with direct RF sampling, which targets Cool Copper Collider (𝐶3) and other C or S band LINAC research and development projects. In this paper, the architecture of the platform will be described. We have optimized the configuration of the data converter and characterized performance of them with RF pulses. The test results for some of the key performance parameters for the LLRF platform with our custom solid-state amplifier, such as phase and amplitude stability, will be discussed in this paper.

Liu, C↗

Ranking Biological Features in Soil-Based Microbial Multi-Omics Data with Integration Modeling

Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.

54 ENVIRONMENTAL SCIENCES↗

3D Multiresolution Velocity Model Fusion with Probability Graphical Models

ABSTRACT The variability in spatial resolution of seismic velocity models obtained via tomographic methodologies is attributed to many factors, including inversion strategies, ray-path coverage, and data integrity. Integration of such models, with distinct resolutions, is crucial during the refinement of community models, thereby enhancing the precision of ground-motion simulations. Toward this goal, we introduce the probability graphical model (PGM), combining velocity models with heterogeneous resolutions and nonuniform data point distributions. The PGM integrates data relations across varying resolution subdomains, enhancing detail within low-resolution (LR) domains by utilizing information and prior knowledge from high-resolution (HR) subdomains through a maximum posterior problem. Assessment of efficacy, utilizing both 2D and 3D velocity models—consisting of synthetic checkerboard models and a fault-zone model from Ridgecrest, California—demonstrates noteworthy improvements in accuracy, compared to state-of-the-art fusion techniques. Specifically, we find reductions of 30% and 44% in computed travel-time residuals for 2D and 3D models, respectively, as compared to conventional smoothing techniques. Unlike conventional methods, the PGM’s adaptive weight selection facilitates preserving and learning details from complex, nonuniform HR models and applies the enhancements to the LR background domain.

Geochemistry & Geophysics↗

Integrating AI Data Centers with the Power Grid

The rapid expansion of artificial intelligence (AI) has triggered an unprecedented surge in electricity demand, with US data center energy use projected to double or triple 2023 levels by 2028. This exponential growth places strain on grid infrastructure, which can hinder timely construction of desired computing capacity. To bridge this supply-demand gap, utilities and AI developers are increasingly turning to demand flexibility, a strategy that incentivizes shifting or reducing power use during peak periods of grid stress. Data centers are uniquely equipped for flexible operations due to their digital workloads, built-in redundancy, and onsite energy assets. This article outlines four primary mechanisms to enable data center flexibility: computational load flexibility (shifting tasks temporally or geographically), flexible use of core facility infrastructure adjustments, energy storage utilization, and onsite electricity generation. To encourage adoption, utilities are deploying new tariff designs, including voluntary interruptible service riders, mandated flexibility requirements, and streamlined interconnection processes for flexible loads. For the highly capitalized and rapidly growing AI industry, the primary motivators for embracing these strategies are expediting facility interconnection, satisfying emerging regulatory mandates, and mitigating community resistance. While demand flexibility cannot substitute the long-term need for new bulk power generation, it serves as an essential, immediate solution for enabling near-term deployment. By transforming data centers from grid stressors into stabilizing assets, flexible operations can ensure reliable grid integration, ease market pressures, and support a resilient power system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

From Modular ADMS to Plug-and-Play Ops: Distribution Grid Operations with Platform-Level Orchestration to Enable Ambitious App Hosting

The core function of the distribution grid is to provide electricity to consumers affordably, reliably, and securely. In pursuing these core objectives, distribution utilities are accountable to customers, regulators, and in some cases, shareholders. Other third parties such as aggregators and microgrids can also have a stake in the smooth operation of the grid. Each of these stakeholders has economic, business, and/or governance objectives that inform their expectations of the distribution grid. This multi-objective, multi-stakeholder environment creates tension that must be reconciled to successfully design and operate the distribution grid. Innovative companies are competing to bring high-tech solutions to electric utilities and their customers that address each of these objectives. Many developers of advanced distribution management systems (ADMS) and distributed energy resource management systems (DERMS) have adopted a modular architecture that allows grid operators to select functions and features according to their individual system needs. A modular platform also allows the solution provider to develop and integrate specific new product modules; however, the need to pursue multiple objectives with a fixed set of controllable devices makes integration expensive whether it is done at the product development stage or the deployment stage. This cost creates a significant barrier to adoption and can lengthen the product to market time of new solutions. To fundamentally address the complexity of system integration for distribution grid operations, the U.S. Department of Energy Office of Electricity has funded the GridAPPS-D project at PNNL, which streamlines integration by contributing to standards development, defining system architecture, applying advanced mathematics, and developing open-source software to demonstrate the concept of an open data-integration platform for distribution operations. The open data-integration platform concept enables system operators and solution providers to deploy ambitious, best-of-breed applications (or apps) without continually reengineering for integration. Ambitious apps developed by different solution providers will inevitably attempt to achieve different control objectives with the same set of controllable devices. If the open platform itself can resolve these conflicts in a way that achieves the best available outcomes for all apps, doesn’t restrict the ambitious design of apps, and ensures safe and secure operations, apps will be able to plug-and-play with the platform at the same time as other ambitious apps. In this paper, we describe a framework called App Deconfliction that empowers a platform to assign setpoints to controllable devices based on the values preferred by different apps (and even external stakeholder entities like customers or aggregators). The App Deconfliction framework is compatible with several methods for determining setpoint values. We present two methods based on game theory that provide a subtle built-in incentive structure for developers to adapt their apps to the fact that they will be operating in a moderated multi-app environment and to favor device setpoints that have the most effect on their objectives over those that have the least effect. Our simulation-based demonstrations have shown that game-theory-based deconfliction can lead to a 7% improvement in control space utilization compared to design-based methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SAM-ML: Integrating data-driven closure with nuclear system code SAM for improved modeling capability

Advanced reactors often involve complicated thermal-fluid (T-F) phenomena. Modeling such phenomena with the traditional one-dimensional (1-D) system code is a challenging task. The System Analysis Module (SAM), a modern nuclear system code, has developed a coarse mesh multi-dimensional (multi-D) flow model to capture the spatial effect of T-F phenomena in advanced reactors. As a coarse mesh solver, constitutive relations are required for SAM's multi-D model for unresolved fine-scale physics, such as turbulence. Here this work presents a novel approach that integrates neural networks as data-driven closure for SAM's multi-D flow model. The data-driven closure is trained with fine-resolution data to ensure its accuracy while maintaining a coarse mesh setup to ensure its efficiency and consistency with SAM. We demonstrate the applicability of this SAM-ML capability in an open volume thermal stratification problem, where a neural network model serves as the eddy viscosity closure. A customized interface between the neural network and SAM is developed to ensure flexible and efficient data exchange. The SAM-ML results demonstrate superior performance compared to SAM's built-in zero-equation eddy viscosity closure. The case study shows that although the generalization capability of the data-driven closure still needs to be improved for different transient case or different geometric setup, SAM -ML demonstrates good potential for challenging simulation problems with improved accuracy and computational efficiency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Expert‐in‐the‐loop design of integral nuclear data experiments

Abstract Nuclear data are fundamental inputs to radiation transport codes used for reactor design and criticality safety. The design of experiments to reduce nuclear data uncertainty has been a challenge for many years, but advances in the sensitivity calculations of radiation transport codes within the last two decades have made optimal experimental design possible. The design of integral nuclear experiments poses numerous challenges not emphasized in classical optimal design, in particular, constrained design spaces (in both a statistical and engineering sense), severely under‐determined systems, and optimality uncertainty. We present a design pipeline to optimize critical experiments that uses constrained Bayesian optimization within an iterative expert‐in‐the‐loop framework. We show a successfully completed experiment campaign designed with this framework that involved two critical configurations and multiple measurements that targeted compensating errors in 239 Pu nuclear data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Bayesian Monte Carlo Evaluation Framework for Cross Sections Nuclear Data and Integral Benchmark Experiments

The new Bayesian Monte Carlo (MC) evaluation framework described in this abstract has been conceived as an attempt to improve nuclear data evaluations of differential crosssection data by removing the following two approximations conventionally employed for nuclear data evaluations: all probability density functions (PDFs) of all data and model parameters, both prior and posterior, are assumed to be normal (i.e., Gaussian) PDFs, and all uncertainties and covariances are propagated using a linear approximation. With these approximations removed, the Bayesian MC (BMC) framework could be used to account for nonlinear effects and would enable improved evaluations of differential cross sections and IBE data that are presently performed based on the assumptions itemized above. The BMC would also improve upon the uniform sampling of IBE parameters from within ranges defined by their evaluated uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗