Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data protection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Accessing protected data by a high-performance computing cluster

A data protection system is provided that allows applications to access protected data in a way that restricts applications from outputting to unauthorized targets any unprotected data derived from the protected data and that ensures that the applications do not have access to a key that allows access to the unprotected data. The data protection system provides a policy server that may execute on a service node of a high performance computing system and a data encryption process that may execute on each compute node that is allocated to an application or batch job. The policy server maintains policies of entities specifying access control for protected data. The data encryption process generates a secure execution environment for an application process and interfaces with the policy server to retrieve keys for decrypting protected data in accordance with a policy, and it decrypts and provides the decrypted data to the application process.

Barnes, Peter↗

An Approach to Data Management Planning for Protected Data Projects

This report describes what is required in a data management plan for data that needs to be protected in some fashion. Here we provide an overview of what a data manager should consider, including the data ingestion and various extract-transform-load processes, metadata considerations and documentation, to what might need to be accounted for in the event of data loss. In addition, this report includes two appendices: forms that, when filled out, make the user compliant with DOE data management requirements as well as additional requirements for handling protected data at ORNL.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Method and apparatus for data protection in memory devices

An apparatus and method for efficiently transmitting data are described. A transmitter sends data to a receiver. An encoder of the transmitter divides a received first block of data into multiple sub-blocks. The encoder selects a portion of each sub-block to compare to one another. A portion in a particular sub-block has a same offset and a same size as other portions of other sub-blocks. If the encoder determines the multiple portions match one another, the encoder sends, to the receiver, a second block of data corresponding to the first block of data. The second block of data has a same size as a size of the received first block of data, and the second block of data includes security data from one of multiple error correction schemes. Therefore, the second block of data provides security without increasing an amount of data to transmit.

SeyedzadehDelcheh, SeyedMohammad↗

Decentralized Collaborative Learning with Probabilistic Data Protection

We discuss future directions of Blockchain as a collaborative value co-creation platform, in which network participants can gain extra insights that cannot be accessed when disconnected from the others. As such, we propose a decentralized machine learning framework that is carefully designed to respect the values of democracy, diversity, and privacy. Specifically, we propose a federated multi-task learning framework that integrates a privacy-preserving dynamic consensus algorithm. We show that a specific network topology called the expander graph dramatically improves the scalability of global consensus building. We conclude the paper by making some remarks on open problems.

Ide, Tsuyoshi↗

Development of a Data Overflow Protection System for Super-Kamiokande to Maximize Data from Nearby Supernovae

Neutrinos from very nearby supernovae, such as Betelgeuse, are expected to generate more than ten million events over 10 s in Super-Kamokande (SK). At such large event rates, the buffers of the SK analog-to-digital conversion board (QBEE) will overflow, causing random loss of data that are critical for understanding the dynamics of the supernova explosion mechanism. In order to solve this problem, two new data-acquisition (DAQ) modules were developed to aid in the observation of very nearby supernovae. The first of these, the SN module, is designed to save only the number of hit photomultiplier tubes during a supernova burst and the second, the Veto module, prescales the high-rate neutrino events to prevent the QBEE from overflowing based on information from the SN module. In the event of a very nearby supernova, these modules allow SK to reconstruct the time evolution of the neutrino event rate from beginning to end using both QBEE and SN module data. This paper presents the development and testing of these modules together with an analysis of supernova-like data generated with a flashing laser diode. We demonstrate that the Veto module successfully prevents DAQ overflows for Betelgeuse-like supernovae as well as the long-term stability of the new modules. During normal running the Veto module is found to issue DAQ vetos a few times per month resulting in a total dead-time less than 1 ms, and does not influence ordinary operations. Additionally, using simulation data we find that supernovae closer than 800 pc will trigger the Veto module, resulting in a prescaling of the observed neutrino data.

F20 Instrumentation and technique↗

Hierarchical Data-Driven Protection for Microgrid with 100% Renewable Penetration: Preprint

The accurate detection and isolation of faults is critical for the reliable operation of microgrids (MGs). Traditional protection approaches are even more challenged for 100% renewable MGs because inverter-based resources (IBRs) are the only sources for fault current which are usually low and unpredictable/non-uniform. This calls for new protection scheme that can identify IBR fault responses and detect faults in MGs. Data-driven based protection can learn the pattern of IBR fault responses and make the correct decision to identify faults. Therefore, this paper presents a data-driven approach for fault localization in island MGs. The approach builds a training dataset of comprehensive fault scenarios that can be used to learn fault characteristics from processed measurements. The localization task is modeled as a binary classification problem at each relay, which simplifies the learning process. Then, a hierarchical decision mechanism is used to identify the fault location. The proposed approach is assessed using an exemplary MG with several grid-forming (GFM) and grid-following (GFL) inverters, where accurate estimation of fault location is achieved. The data-driven based protection approach developed in this paper provides a generic framework and useful guidance for power system protection engineers to achieve reliable protection for MGs with 100% renewables.

artificial intelligence↗

Data-Driven Protection Software to classify fault locations by protective zone in distribution systems with high PV penetration

The software contains (a) the source codes to generate Point-on-Wave (PoW) transient data for any feeder model in Alternative Transient Program (ATP) format. Codes provide options to change different steady state settings, including the loading condition and PV capacity and transient state setting like faults type, location and initiation time (b) data post-processing source code to converted data from native format to COMTRADE, csv, HDF5 (c) Docker container to train CNN to classify fault locations by protective zone. The container takes dataset and other training parameters (sampling rate, training epochs, batch size etc) as input to train CNN. The container writes back the trained CNN model, training and testing metrics and plots to the local workstation

Ramesh, Meghana↗

Y-12 Groundwater Protection Program Data Management Plan

This Data Management Plan (DMP) describes the processes in place to ensure the integrity of groundwater monitoring information collected by the U.S. Department of Energy (DOE), National Nuclear Security Administration (NNSA), Y-12 National Security Complex (Y-12), Groundwater Protection Program (GWPP). This information includes program plans, reports, and computer systems used to capture monitoring station information and analytical data. The primary computer system used by the GWPP is the Groundwater Information Management System (GIMS). Procedures used to ensure the integrity of the data are included in this document by reference.

54 ENVIRONMENTAL SCIENCES↗

Innovation in Radiological Security, Part 2 of 2 – Insights into Developing a Cloud-hosted Security Technology

A Sentry Remote Monitoring System (Sentry-RMS) is a stand-alone security system that provides detection, assessment, and communication of priority alarms as an additional means of thwarting internal and external threats to sites that maintain radiological material. The SEntry-RMS CommUnications and REsponse (Sentry-SECURE) platform is an optional feature of the Sentry-RMS that relays priority alarm information to the identified response stakeholders. Sentry-SECURE is hosted in a cloud environment that abstracts the data owner’s and data consumer’s platforms to allow for greater information sharing. This promotes situational awareness amongst authorized users and enables future innovation among modern response platforms. When securely architecting a cloud solution such as this, the use of design paradigms can be an effective tool to increase the accuracy and reliability of cyber- and information-security-related decisions made throughout the development process. This approach also supports the categorization of design considerations into three levels: industry concepts, project approaches, and data protections for digital processes. Industry concepts consist of the notional underpinnings that guide or motivate a security process, system, or design but often lack any tangible attributes. Project approaches represent decisions made during the design and development process to prioritize a solution, method, or practice above another that may provide a comparable functional output but lacks a desired security benefit. Data protections for digital processes represent the selection, integration, and implementation of specific controls for a given asset. This paper will explore specific examples of how Sentry-SECURE has been designed to account for considerations at each of these three levels, while balancing the operational intent of the platform with the security enhancements necessary to maintain data integrity, availability, and confidentiality.

assessment, RMS, security, physical, cloud, Cyber ↗

Fermilab PIP-II machine protection system digitized data noise elimination scheme and its FPGA implementation

In Fermilab's PIP-II machine protection system, beam loss signals from various detectors are digitized at 125 MS/s. Noise from both high-frequency sources and low-frequency 60 Hz AC power equipment can contaminate the data. To suppress noise across these ranges—especially 60 Hz and its harmonics, which overlap with beam loss signal frequencies—advanced digital processing beyond standard filtering is required. Several real-time functional blocks were simulated and tested on an FPGA: (1) a dual time-constant discharging integrator filter, (2) a de-ripple baseline extraction and storage block, and (3) a fast-recovery discharging integrator. The nonlinear IIR integrator filter removes high-frequency noise and feeds into the baseline extractor. Upon detecting abrupt beam loss, it switches to a longer time constant to prevent baseline distortion. The de-ripple block calculates a valid baseline by averaging over multiple 60 Hz periods, storing results in a 4096-word FPGA RAM. This baseline is subtracted from raw data before integration by the fast-recovery block, which resets quickly after use. All blocks achieved expected performance and were successfully implemented on a low-cost FPGA.

Wu, Jinyuan [Fermilab]↗

Privacy policy robustness to reverse engineering

Differential privacy policies allow one to preserve data privacy while sharing and analyzing data. However, these policies are susceptible to an array of attacks. In particular, often a portion of the data desired to be privacy protected is exposed online. Access to these pre-privacy protected data samples can then be used to reverse engineer the privacy policy. With knowledge of the generating privacy policy, an attacker can use machine learning to approximate the full set of originating data. Bayesian inference is one method for reverse engineering both model and model parameters. We present a methodology for evaluating and ranking privacy policy robustness to Bayesian inference-based reverse engineering, and demonstrated this method across data with a variety of temporal trends.

Kusne, Aaron Gilad↗

Bridging Power System Protection Gaps with Data-driven Approaches

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this research, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A convolutional neural network (CNN) based fault detection approach is implemented to achieve data translation invariance of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN–based method avoids the complicated manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on four kinds of protection gaps: high impedance faults, transformer/generator inter-turn faults, distribution system PV circuit faults, and the mis-operation situations of Zone 3 line protection relays operating under system stress. Finally, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

FY 2025 Multidimensional Data Correlation Platform: Unified Software Architecture for Advanced Materials and Manufacturing Technologies Data Management and Processing

The Advanced Materials and Manufacturing Technologies (AMMT) program continues to advance a data-driven approach to demonstrate the utility of additive manufacturing for fabricating components for nuclear applications. A key scientific goal is to leverage data to better understand manufacturing outcomes and thereby improve the performance, reliability, and lifespan of nuclear components. Ultimately, this effort supports the development of standards for certification and qualification of additively manufactured components, enabling broader industry adoption. In support of this objective, the AMMT program is building and deploying a data management platform to record, index, analyze, and make available the manufacturing data generated across the AMMT program. In FY 2023, the team conceptualized the architecture of the platform and, in FY 2024, deployed the first functional version at the Oak Ridge National Laboratory (ORNL) Manufacturing Demonstration Facility (MDF). In FY 2025, the platform was officially opened to all AMMT members. To enable this expansion, core modifications and enhancements were developed, including improvements to the user interface and workflows for data entry and retrieval. Most notably, robust security and access control mechanisms were implemented to protect data and manage information sharing. This effort featured a logging system, protected views, and controlled access mechanisms. This report documents these enhancements and the transition of the platform into program-wide use.

36 MATERIALS SCIENCE↗

Hybrid Interior-Point/Active-Set SCOPF Algorithms Exploiting Power System Characteristics

These protected data were produced under agreement no. with the U.S. Department of Energy and may not be published, disseminated, or disclosed to others outside the Government until 5 years after development of information under this agreement, unless express written authorization is obtained from the recipient. Upon expiration of the period of protection set forth in this Notice, the Government shall have unlimited rights in this data. This Notice shall be marked on any reproduction of this data, in whole or in part.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Utah FORGE: Sanvean Technologies Drilling Data from Well 16B(78)-32

Included here is Sanvean Technologies bit sensor data amalgamated with data from National Oilwell Varco's (NOV) BlackBox tool for Reed Hycalog bits used during drilling of Well 16B(78)-32. The dataset contains information collected at the bit while drilling including rate of penetration (ROP), top drive torque, and bit box temperature. The data was recorded at the bit box and top sub of the motor. RPM was measured by onboard gyro recording continuously in each sensor, and shock levels were also recorded on X, Y and Z axis. This data was merged with EDR in time format and saved in file sets (the zipped files) then output into CSV files. Please note: fields in the CSV files, such as the date field, may need to be formatted to display properly. There is an additional zipped folder in each dataset that is password protected. Sanvean GameChanger Viewer software must be used to view this password protected data. Information on how to use and download this free software is also included here.

15 GEOTHERMAL ENERGY↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Protection Against Graph-Based False Data Injection Attacks on Power Systems

Graph signal processing (GSP) has emerged as a powerful tool for practical network applications, including power system monitoring. By representing power system voltages as smooth graph signals, recent research has focused on developing GSP-based methods for state estimation, attack detection, and topology identification. Included, efficient methods have been developed for detecting false data injection (FDI) attacks, which until now were perceived as non-smooth with respect to the graph Laplacian matrix. Consequently, these methods may not be effective against smooth FDI attacks. In this paper, we propose a graph FDI (GFDI) attack that minimizes the Laplacian-based graph total variation (TV) under practical constraints. In addition, we develop a low-complexity algorithm that solves the non-convex GDFI attack optimization problem using ell_1-norm relaxation, the projected gradient descent (PGD) algorithm, and the alternating direction method of multipliers (ADMM). We then propose a protection scheme that identifies the minimal set of measurements necessary to constrain the GFDI output to high graph TV, thereby enabling its detection by existing GSP-based detectors. Our numerical simulations on the IEEE-57 bus test case reveal the potential threat posed by well-designed GSP-based FDI attacks. Moreover, we demonstrate that integrating the proposed protection design with GSP-based detection can lead to significant hardware cost savings compared to previous designs of protection methods against FDI attacks.

Morgenstern, Gal↗