Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data protection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Fermilab PIP II machine protection system digitized data noise elimination scheme and its FPGA implementation

In Fermilab's PIP-II machine protection system, beam loss signals from various detectors are digitized at 125 MS/s. Noise from both high-frequency sources and low-frequency 60 Hz AC power equipment can contaminate the data. To suppress noise across these ranges especially 60 Hz and its harmonics, which overlap with beam loss signal frequencies advanced digital processing beyond standard filtering is required. Several real-time functional blocks were simulated and tested on an FPGA: (1) a dual time-constant discharging integrator filter, (2) a de-ripple baseline extraction and storage block, and (3) a fast-recovery discharging integrator. The nonlinear IIR integrator filter removes high-frequency noise and feeds into the baseline extractor. Upon detecting abrupt beam loss, it switches to a longer time constant to prevent baseline distortion. The de-ripple block calculates a valid baseline by averaging over multiple 60 Hz periods, storing results in a 4096-word FPGA RAM. This baseline is subtracted from raw data before integration by the fast-recovery block, which resets quickly after use. All blocks achieved expected performance.

Wu, J. [Fermilab] (ORCID:0000000344329521)↗

Fermilab PIP II Machine Protection System Digitized Data Noise Elimination Scheme and Its FPGA Implementation

In Fermilab's PIP-II machine protection system, beam loss signals from various detectors are digitized at 125 MS/s. Noise from both high-frequency sources and low-frequency 60 Hz AC power equipment can con-taminate the data. To suppress noise across these ranges especially 60 Hz and its harmonics, which overlap with beam loss signal frequencies advanced digital processing beyond standard filtering is re-quired. Several real-time functional blocks were simu-lated and tested on an FPGA: (1) a dual time-constant discharging integrator filter, (2) a de-ripple baseline extraction and storage block, and (3) a fast-recovery discharging integrator. The nonlinear IIR integrator filter removes high-frequency noise and feeds into the baseline extractor. Upon detecting abrupt beam loss, it switches to a longer time constant to prevent baseline distortion. The de-ripple block calculates a valid base-line by averaging over multiple 60 Hz periods, storing results in a 4096-word FPGA RAM. This baseline is subtracted from raw data before integration by the fast-recovery block, which resets quickly after use. All blocks achieved expected performance.

Wu, Jinyuan [Fermilab] (ORCID:0000000344329521)↗

Homomorphic Encryption for Electrical Metering Aggregation: Protecting the Privacy of Building Tenants

Electrical meters are devices that measure consumer electricity usage. The data collected by these meters is necessary for utility billing and electrical grid management but can also be used to assess the environmental impact of buildings. Prior research has found that unprotected metering data could potentially be used to infer some information about the behaviors of building tenants by detecting changes in electricity usage. For example, a period of low electricity usage could suggest that the tenants are not in the building. As smart metering becomes more common, there is a growing need for data privacy protections for metering data that do not negatively impact the quality and availability of data used for energy management and billing applications. To identify potential solutions, we developed a Python-based data aggregation platform to analyze the potential efficacy of privacy-enhancing technologies for energy metering applications. This platform aggregates groups of metering sites into virtual buildings, which could potentially detach changes in electrical activity from individual tenants, making it more difficult to track the activity of a specific tenant. To further protect data during analysis, this project utilizes homomorphic encryption as part of its initial approach. Homomorphic encryption offers a means of protecting energy consumption data while permitting mathematical operations to be performed without the need to know the data contents. This allows for data to be processed into usable statistics without revealing energy consumption information. A series of homomorphic encryption libraries were evaluated to determine their applicability and limitations in the context of metering data. The use of these techniques may help to reassure consumers and encourage further adoption of smart grid infrastructure.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A System for Standardizing and Combining U.S. Environmental Protection Agency Emissions and Waste Inventory Data

The U.S. Environmental Protection Agency (USEPA) provides databases that agglomerate data provided by companies or states reporting emissions, releases, wastes generated, and other activities to meet statutory requirements. These databases, often referred to as inventories, can be used for a wide variety of environmental reporting and modeling purposes to characterize conditions in the United States. Yet, users are often challenged to find, retrieve, and interpret these data due to the unique schemes employed for data management, which could result in erroneous estimations or double-counting of emissions. To address these challenges, a system called Standardized Emission and Waste Inventories (StEWI) has been created. The system consists of four python modules that provide rapid access to USEPA inventory data in standard formats and permit filtering and combination of these inventory data. When accessed through StEWI, reported emissions of carbon dioxide to air and ammonia to water are reduced approximately two- and four-fold, respectively, to avoid duplicate reporting. StEWI will greatly facilitate the use of USEPA inventory data in chemical release and exposure modeling and life cycle assessment tools, among other things. To date, StEWI has been used to build the recent USEEIO model and the baseline electricity life cycle inventory database for the Federal LCA Commons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

End-to-end microgrid protection using distributed data-driven methods

This paper introduces an end-to-end microgrid protection framework that offers real-time system monitoring, fault-related decision making, and circuit breaker control. This is achieved through the design of distributed data-driven techniques based on the support vector machine method, where each relay is responsible for distributed data collection, fault detection, fault localization, and fault isolation. Local communication is established among neighboring relays, fostering cooperative fault localization and isolation. This decentralized design not only reduces the computational and communication requirements but also enables the adaptability of each relay under varying operational dynamics. The proposed end-to-end protection framework was validated using MATLAB/Simulink simulations on a 100% renewable microgrid, achieving an accuracy of 93.1% with response time of 0.0523 s, in protecting against a range of fault scenarios that are characterized by various types, locations, impedances, load conditions, photovoltaic power levels, and microgrid operating modes.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Virtual Log-Structured Storage for High-Performance Streaming

Over the past decade, given the higher number of data sources (e.g., Cloud applications, Internet of things) and critical business demands, Big Data transitioned from batch-oriented to real-time analytics. Stream storage systems, such as Apache Kafka, are well known for their increasing role in real-time Big Data analytics. For scalable stream data ingestion and processing, they logically split a data stream topic into multiple partitions. Stream storage systems keep multiple data stream copies to protect against data loss while implementing a stream partition as a replicated log. This architectural choice enables simplified development while trading cluster size with performance and the number of streams optimally managed. This paper introduces a shared virtual log-structured storage approach for improving the cluster throughput when multiple producers and consumers write and consume in parallel data streams. Stream partitions are associated with shared replicated virtual logs transparently to the user, effectively separating the implementation of stream partitioning (and data ordering) from data replication (and durability). We implement the virtual log technique in the KerA stream storage system. When comparing with Apache Kafka, KerA improves the cluster ingestion throughput by up to 4x when multiple producers write over hundreds of data streams.

consistent stream ordering↗

Protection of Inverter-Dependent Transmission Systems (PROTECT-IT)

This presentation highlights the overall objectives of the SETO funded protection project. The main technical approaches are also highlighted. This high impact project produces multiple innovative outcomes, including comprehensive impact study of how IBR affects protection elements, simplified low-order IBR model for protection engineers, enhanced and data-driven protection design, co-design concept for coordination protection and IBRs.

14 SOLAR ENERGY↗

The Statistical Spread of Transmission Outages on a Fast Protection Time Scale Based on Utility Data

When there is a fault, the protection system automatically removes one or more transmission lines on a fast time scale of less than one minute. The outaged lines form a pattern in the transmission network. We extract these patterns from utility outage data, determine some key statistics of these patterns, and then show how to generate new patterns consistent with these statistics. The generated patterns provide a new and easily feasible way to model the overall effect of the protection system at the scale of a large transmission system. This new data-driven generative modeling of protection is expected to contribute to simulations of disturbances in large grids so that they can better quantify the risk of blackouts. Analysis of the pattern sizes suggests an index that describes how much outages spread in the transmission network at the fast timescale.

Transmission↗

Networked Microgrid Ownership, Data, and Control Implications: Challenges and Open Questions

Microgrid deployments increasingly favor the potential to form networks for greater benefits to resilience, reliability, and energy sovereignty. Both independent and networked micro-grids predominantly have a single-entity-ownership and control, where the associations from ownership to data requirements to control functions to microgrid objectives is linear. The emerging model, however, is cyclical, with bidirectional causal impacts between each of the 4 pillars: there are more complex mixed ownership models across the physical, electrical, data, communications, protection, and control boundaries that impact the data requirements for meeting control functions that help realize the use-cases or objectives. This paper is the first to delineate the pillars for effective ownership and controllability of both independent as well as networked microgrids through the cyclical model, and present barriers to the adoption of such a model.

Sundararajan, Aditya↗

Distributed fiber sensor and machine learning data analytics for pipeline protection against extrinsic intrusions and intrinsic corrosions

This paper presents an integrated technical framework to protect pipelines against both malicious intrusions and piping degradation using a distributed fiber sensing technology and artificial intelligence. A distributed acoustic sensing (DAS) system based on phase-sensitive optical time-domain reflectometry (φ-OTDR) was used to detect acoustic wave propagation and scattering along pipeline structures consisting of straight piping and sharp bend elbow. Signal to noise ratio of the DAS system was enhanced by femtosecond induced artificial Rayleigh scattering centers. Data harnessed by the DAS system were analyzed by neural network-based machine learning algorithms. The system identified with over 85% accuracy in various external impact events, and over 94% accuracy for defect identification through supervised learning and 71% accuracy through unsupervised learning.

Peng, Zhaoqiang↗

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗

Enabling end-to-end secure federated learning in biomedical research on heterogeneous computing environments with APPFLx

Facilitating large-scale, cross-institutional collaboration in biomedical machine learning (ML) projects requires a trustworthy and resilient federated learning (FL) environment to ensure that sensitive information such as protected health information is kept confidential. Specifically designed for this purpose, this work introduces APPFLx - a low-code, easy-to-use FL framework that enables easy setup, configuration, and running of FL experiments. APPFLx removes administrative boundaries of research organizations and healthcare systems while providing secure end-to-end communication, privacy-preserving functionality, and identity management. Furthermore, it is completely agnostic to the underlying computational infrastructure of participating clients, allowing an instantaneous deployment of this framework into existing computing infrastructures. Experimentally, the utility of APPFLx is demonstrated in two case studies: (1) predicting participant age from electrocardiogram (ECG) waveforms, and (2) detecting COVID-19 disease from chest radiographs. Here, ML models were securely trained across heterogeneous computing resources, including a combination of on-premise high-performance computing and cloud computing facilities. By securely unlocking data from multiple sources for training without directly sharing it, these FL models enhance generalizability and performance compared to centralized training models while ensuring data remains protected. In conclusion, APPFLx demonstrated itself as an easy-to-use framework for accelerating biomedical studies across organizations and healthcare systems on large datasets while maintaining the protection of private medical data.

Biomedical Research↗

New data-driven approach to bridging power system protection gaps with deep learning

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this paper, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A combined convolutional neural network and long short-term memory (CNN-LSTM) network is implemented to achieve data translation invariance and capture the temporal correlation of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN-LSTM--based method avoids the complicated, manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on two kinds of protection gaps: high-impedance faults and transformer inter-turn faults. Lastly, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

42 ENGINEERING↗

Nine Canyon Long-Duration Energy Storage: A Feasibility Study

The Nine Canyon Long Duration Energy Storage (LDES) Feasibility Study explores the technical and economic viability of deploying advanced energy storage technologies at Energy Northwest's (EN) Nine Canyon (9C) Wind Project site in Benton County, Washington. Supported by the Washington State Department of Commerce and the U.S. Department of Energy’s Office of Electricity under its LDES Voucher Program, the study represents a collaborative effort between EN, Pacific Northwest National Laboratory (PNNL), and ARES North America. At the core of this effort is the development of a generalized techno-economic modeling framework and evaluation tool designed to assess the value proposition of LDES projects across a variety of contexts. The modeling tool is technology-agnostic and accommodates user-defined parameters such as rated power, energy duration, round-trip efficiency, capital and operational costs, and dispatch constraints. It also integrates economic inputs, including market prices, energy revenue structures, and financing parameters to evaluate performance through key metrics. The tool provides utilities with a transparent, adaptable platform to support decision-making, investment prioritization, and portfolio planning for various storage technologies. To guide scenario design and interpretation, the study first surveyed the LDES technology landscape, including lithium-ion batteries, flow batteries, non-hydro gravity storage, and thermo-mechanical systems, comparing cost trajectories, technical performance, safety and hazards, materials sourcing and recyclability, and spatial/siting considerations. This literature-grounded review highlights technology trade-offs and reinforces the need to align technology choice with site characteristics, use cases, and project objectives. A companion chapter examines ownership structures (EN ownership, third-party ownership, shared models) and offtake options (energy marketing, capacity/energy PPAs, time-of-use PPAs, block-delivery PPAs, and tolling), where PPAs (power purchase agreements) represent contractual arrangements for buying and selling electricity. The chapter also highlights implications for risk allocation, capital access, operational control, and revenue certainty. The study also evaluates supervisory control and data acquisition (SCADA) and transmission interconnection pathways, options include upgrading the existing SCADA or deploying a dedicated LDES controller, with attention to protection schemes, data telemetry, cybersecurity, and regulatory coordination with BPA. In addition, an ARES-specific geotechnical and hydrology assessment presented in the appendix screens multiple corridors for slope stability, bearing capacity, cut-and-fill magnitude, and stormwater behavior.

25 ENERGY STORAGE↗

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics↗