Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Distributed Ledger Technology Preliminary Performance Assessment for Applications in Utilities Operational Technology

This paper provides descriptions of the key components of different distributed ledger technology platforms. Distributed ledger technology (DLT) allows for distribution of databases among different organizations and devices. The platforms use cryptographically linked “blocks” to store and verify transactional information between these organizations. DLT increases data security, data integrity, trust among its participants. Different organizations are looking to deploy this distributed and decentralized approach to avoid the single-point-of-failure vulnerabilities associated with centralized data repositories. In this study we examine twelve different DLT platforms. There is agreement within the community that of all the platforms considered here, Hyperledger and Ethereum are the most mature when it comes to privacy and permissions. These DLT platforms are being used for applications such as transactive energy, health care, and the food and goods supply chain. However, further development is required to realize the full promise of DLT. Our assessment includes a general description of each DLT and its key characteristics. Such characteristics include consensus protocol and cryptography used, public vs. private, and permissioned or permissionless. The selection and implementation of a DLT architecture depends heavily on the use case and performance requirements. During this research we found key parameters to measure performance and existing tools for assessment. Four different parameters were identified 1) consensus, 2) throughput, 3) latency, and 4) scalability. The architectures of Hyperledger Caliper and Blockbench are described as different performance assessment frameworks. From this preliminary study it is evident that there are dissimilarities on the performance assessments methods developers and users are characterizing DLT architectures. The purpose of this paper is to identify key parameters to test performance, tools that are being used and provide information on results from previous studies.

97 MATHEMATICS AND COMPUTING↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

LYNM-PE1 Seismic Parameters from Borehole Log, Laboratory, and Tabletop Measurements

The goal of this work is to provide a database of quality-checked seismic parameters that can be integrated with the Geologic Framework Model (GFM) for the LYNM-PE1 (Low Yield Nuclear Monitoring – Physical Experiment 1) testbed. We integrated data from geophysical borehole logs, tabletop measurements on collected core, and laboratory measurements. We reviewed for internal consistency among each measurement type, documented the caveats of measurement conditions, and integrated lithologic logs to check the validity of outlier values. The resulting consolidated parameter tables can be used as inputs for modeling and analysis codes and are designed to interface with the GFM, which is being actively developed.

58 GEOSCIENCES↗

Using Information Automation and Human Technology Integration to Implement Integrated Operations for Nuclear

The purpose of the research effort described in this report is to develop and demonstrate an approach to the design and implementation of advanced, automated systems intended to increase operational efficiencies at nuclear power plants. In particular, we describe methods for considering human-technology integration (HTI) issues throughout the various phases of system design, test, and implementation and how these considerations help promote effective design. This research project will develop planning tools and comprehensive guidance on how HTI principles and methods, in combination with information automation technologies, can enable effective data integration and coordination for full nuclear plant modernization. Specifically, this research project will develop an approach to automate the mapping of data from plant systems and processes to application needs, thereby significantly reducing the amount of human workload currently required for the execution of these tasks. In addition to developing an automated solution as a replacement for these tasks, this research will also analyze digitalization’s effectiveness in reducing operational costs of compliance related activities. Compliance activities are estimated to account for as much as 50% of operations and maintenance (i.e., non-fuel and non-capital) costs of plant operation.

97 MATHEMATICS AND COMPUTING↗

Dataset for 'Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning', Water 2022

This data package presents forcing data, model code, and model output for classical machine learning models that predict monthly stream water temperature as presented in the manuscript ‘Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning’, Water (Weierbach et al., 2022). Specifically, for input forcing datasets we include two files each generated using the BASIN-3D data integration tool (Varadharajan et al., 2022) for stations in the Pacific Northwest and Mid Atlantic Hydrologic regions. Model code (written in python with the use of jupyter notebooks) includes codes for data preprocessing, training Multiple Linear Regression, Support Vector Regression, and Extreme Gradient Boosted Tree models, and additional notebooks for analysis of model output. We include specific model output files which represent modeling configurations presented in the manuscript also presented in an hdf5 format. Together, these data make up the workflow for predictions across three scenarios (single station, regional, and predictions in unmonitored basins) presented in the manuscript and allow for reproducibility of modeling procedures.

54 ENVIRONMENTAL SCIENCES↗

A Causal Approach to Integrate Component Health Data into System Reliability Models

Two of the challenges of current plant reliability approaches are the ability to integrate plant health data, and to support decision making. Condition based data and diagnostic/prognostic information are in fact not considered into plant reliability models to inform system engineers on the most critical components. Currently, the propagation of quantitative health data from the component to the system level is a challenge given the diverse nature/structure of the data. On the other hand, plant reliability methods (which are typically based on fault-trees or reliability block diagrams) can effectively propagate data from the component to the system level, but values of failure rates or failure probabilities are an approximated integral representation of the past industry-wide operational experience, and it neglects the present component health status (e.g., diagnostic and condition-based data) and health projection (when available from prognostic data). Our first claim is that system reliability models should propagate health information from the component to the system/plant level in order to provide a quantitative snapshot of system/plant health and identify the most critical components. Our second claim is that component health should be informed solely by that specific component current and historical performance data and should not be an approximated integral representation of the past industry-wide operational experience. This paper is directly supporting these two claims by proposing a different approach to perform reliability modeling which relies on available component diagnostic, prognostic and condition-based data to measure component health, and it propagates this information through fault tree models. The propagation of health data from the component to the system level is performed not in terms of probability, but in terms of margins where margin is defined as the “distance” between the present actual status and an undesired event (e.g., failure or unacceptable performance). Through a cause-effect lens, while classical reliability models target the effect associated to a component performance, a margin-based approach focuses on the cause of an undesired component performance (i.e., component health). Hence, thinking of reliability in terms of margins implies decision making based on causal reasoning. We will show how fault tree models can be solved using a margin language and how this process can effectively assist system engineers to identify the most critical components.

97 MATHEMATICS AND COMPUTING↗

A Novel Architecture for Attack-Resilient Wide-Area Protection and Control System in Smart Grid

Wide-area protection and control (WAPAC) systems are widely applied in the energy management system (EMS) that rely on a wide-area communication network to maintain system stability, security, and reliability. As technology and grid infrastructure evolve to develop more advanced WAPAC applications, however, so do the attack surfaces in the grid infrastructure. This paper presents an attack-resilient system (ARS) for the WAPAC cybersecurity by seamlessly integrating the network intrusion detection system (NIDS) with intrusion mitigation and prevention system (IMPS). In particular, the proposed NIDS utilizes signature and behavior-based rules to detect attack reconnaissance, communication failure, and data integrity attacks. Further, the proposed IMPS applies state transition-based mitigation and prevention strategies to quickly restore the normal grid operation after cyberattacks. As a proof of concept, we validate the proposed generic architecture of ARS by performing experimental case study for wide-area protection scheme (WAPS), one of the critical WAPAC applications, and evaluate the proposed NIDS and IMPS components of ARS in a cyber-physical testbed environment. Our experimental results reveal a promising performance in detecting and mitigating different classes of cyberattacks while supporting an alert visualization dashboard to provide an accurate situational awareness in real-time.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The need for an integrated multi-OMICs approach in microbiome science in the food system

Microbiome science as an interdisciplinary research field has evolved rapidly over the past two decades, becoming a popular topic not only in the scientific community and among the general public, but also in the food industry due to the growing demand for microbiome-based technologies that provide added-value solutions. Microbiome research has expanded in the context of food systems, strongly driven by methodological advances in different -omics fields that leverage our understanding of microbial diversity and function. However, managing and integrating different complex -omics layers are still challenging. Within the Coordinated Support Action MicrobiomeSupport (https://www.microbiomesupport.eu/), a project supported by the European Commission, the workshop “Metagenomics, Metaproteomics and Metabolomics: the need for data integration in microbiome research” gathered 70 participants from different microbiome research fields relevant to food systems, to discuss challenges in microbiome research and to promote a switch from microbiome-based descriptive studies to functional studies, elucidating the biology and interactive roles of microbiomes in food systems. A combination of technologies is proposed. This will reduce the biases resulting from each individual technology and result in a more comprehensive view of the biological system as a whole. Although combinations of different datasets are still rare, advanced bioinformatics tools and artificial intelligence approaches can contribute to understanding, prediction, and management of the microbiome, thereby providing the basis for the improvement of food quality and safety.

60 APPLIED LIFE SCIENCES↗

ZFS Interface For Accelerators

As data volume increases in HPC centers, storing data in reasonable amounts of space and time while maintaining data integrity and recoverability becomes ever more difficult. ZFS is a filesystem that provides these features, among many others, that make it attractive for usage in HPC environments. It has the ability to compress data, compute checksums, and compute redundancy codes within the same runtime, rather than compute each separately, without knowledge of the underlying structure of the filesystem. However, measurements have shown that compression on CPUs can be incredibly inefficient and significantly reduce the performance of ZFS. The ZFS Interface for Accelerators (Z.I.A.) was developed to provide an interface to route data to accelerators while being processed by ZFS, so that the ZFS infrastructure and features are maintained, while allowing for ZFS administrators to provide faster implementations of compression, or other features, to their instance of ZFS.

Lee, Jason↗

MyCrunchGPT: A LLM Assisted Framework for Scientific Machine Learning

Scientific machine learning (SciML) has advanced recently across many different areas in computational science and engineering. Here, the objective is to integrate data and physics seamlessly without the need of employing elaborate and computationally taxing data assimilation schemes. However, preprocessing, problem formulation, code generation, postprocessing, and analysis are still time- consuming and may prevent SciML from wide applicability in industrial applications and in digital twin frameworks. Here, we integrate the various stages of SciML under the umbrella of ChatGPT, to formulate MyCrunchGPT, which plays the role of a conductor orchestrating the entire workflow of SciML based on simple prompts by the user. Specifically, we present two examples that demonstrate the potential use of MyCrunchGPT in optimizing airfoils in aerodynamics, and in obtaining flow fields in various geometries in interactive mode, with emphasis on the validation stage. To demonstrate the flow of the MyCrunchGPT, and create an infrastructure that can facilitate a broader vision, we built a web app based guided user interface, that includes options for a comprehensive summary report. The overall objective is to extend MyCrunchGPT to handle diverse problems in computational mechanics, design, optimization and controls, and general scientific computing tasks involved in SciML, hence using it as a research assistant tool but also as an educational tool. While here the examples focus on fluid mechanics, future versions will target solid mechanics and materials science, geophysics, systems biology, and bioinformatics.

97 MATHEMATICS AND COMPUTING↗

Rapid Detection of Anomalies in Battery Energy Storage System Data

Data analytics is pivotal in assessing the technical characteristics and performance of Battery Energy Storage Systems (BESS), underpinning BESS modeling, optimization, and control. However, raw datasets frequently harbor anomalies from measurement errors and equipment malfunctions, impacting BESS reliability and analysis accuracy To address the challenge, this paper presents a novel methodology for the rapid detection of anomalous charge or discharge cycles within BESS operational data, expediting the cleaning process while ensuring data integrity. We’ve collected diverse and comprehensive real-world BESS operational datasets in collaboration with the Electric Power Research Institute and multiple Washington State utilities. These datasets serve dual roles: enabling comprehensive data exploration and analysis for understanding underlying challenges and method development, while also acting as a vital validation resource, demonstrating practical effectiveness. The proposed method detects anomalies and aids in their resolution, improving system performance characterization precision. It also reveals recurring data anomaly sources, offering insights for data collection and handling enhancement. Practitioners can gain valuable insights from the identified anomalous cycles in the real-world datasets along with the investigative process for root cause analyses and essential data cleaning steps.

Crawford, Aladsair J.↗

FL‐ADS: Federated learning anomaly detection system for distributed energy resource networks

Abstract With the ongoing development of Distributed Energy Resources (DER) communication networks, the imperative for strong cybersecurity and data privacy safeguards is increasingly evident. DER networks, which rely on protocols such as Distributed Network Protocol 3 and Modbus, are susceptible to cyberattacks such as data integrity breaches and denial of service due to their inherent security vulnerabilities. This paper introduces an innovative Federated Learning (FL)‐based anomaly detection system designed to enhance the security of DER networks while preserving data privacy. Our models leverage Vertical and Horizontal Federated Learning to enable collaborative learning while preserving data privacy, exchanging only non‐sensitive information, such as model parameters, and maintaining the privacy of DER clients' raw data. The effectiveness of the models is demonstrated through its evaluation on datasets representative of real‐world DER scenarios, showcasing significant improvements in accuracy and F1‐score across all clients compared to the traditional baseline model. Additionally, this work demonstrates a consistent reduction in loss function over multiple FL rounds, further validating its efficacy and offering a robust solution that balances effective anomaly detection with stringent data privacy needs.

Purohit, Shaurya [Iowa State University Ames Iowa ↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Novel Proposals for FAIR, Automated, Recommendable, and Robust Workflows

Lightning talks of the Workflows in Support of Large-Scale Science (WORKS) workshop are a venue where the workflow community (researchers, developers, and users) can discuss work in progress, emerging technologies and frameworks, and training and education materials. This paper summarizes the WORKS 2022 lightning talks, which cover five broad topics: data integrity of scientific workflows; a machine learning-based recommendation system; a Python toolkit for running dynamic ensembles of simulations; a cross-platform, high-performance computing utility for processing shell commands; and a meta(data) framework for reproducing hybrid workflows.

Abhinit, Ishan↗

Intelliquench: An Adaptive Machine Learning System for Detection of Superconducting Magnet Quenches

In superconducting magnets, the irreversible transition of a portion of the conductor to resistive state is called a “quench.” Having large stored energy, magnets can be damaged by quenches due to localized heating, high voltage, or large force transients. Unfortunately, current quench protection systems can only detect a quench after it happens, and mitigating risks in Low Temperature Superconducting (LTS) accelerator magnets often requires fast response (down to ms). Additionally, protection of High Temperature Superconducting (HTS) magnets is still suffering from prohibitively slow quench detection. In this study, we lay the groundwork for a quench prediction system using an auto-encoder fully-connected deep neural network. After dynamically trained with data features extracted from acoustic sensors around the magnet, the system detects anomalous events seconds before the quench in most of our data. While the exact nature of the events is under investigation, we show that the system can “forecast” a quench before it happens under magnet training conditions through a randomized experiment. This opens up the way of integrated data processing, potentially leading to faster and better diagnostics and detection of magnet quenches

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

AmpSuite

Seismic amplitudes offer vital information about explosion source characteristics, including discrimination and yield estimation. To take advantage of this, we developed an interactive Python package to measure, control data quality, generate broad area propagation models and perform discrimination and estimate yield. Propagation models are essential in support of transportable yield and broad area discrimination. The key benefit of this package will be its ability to continuously integrate data and new techniques. The AmpSuite framework will provide standardized, repeatable, and accurate model generation and characterization routines. The capability is crucial for monitoring agencies tasked with rapid and high-quality seismic event characterization. The AmpSuite software includes a series of independent modules to perform: • Direct Phase Amplitude Measurement and Storage • Coda Envelope Measurement and Storage • Data Quality Control • New Propagation Model Developments • Seismic Discrimination and Analysis • Yield Estimation and supporting utility software. The AmpSuite software provides comprehensive solutions for monitoring agencies seeking to optimize model generation and event analysis within a contemporary Python framework. Stakeholders (AFTAC) have begun to move towards the Python language for scientific analysis as a new workforce emerges.

Alfaro, Richard↗