Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data protection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Enabling end-to-end secure federated learning in biomedical research on heterogeneous computing environments with APPFLx

Facilitating large-scale, cross-institutional collaboration in biomedical machine learning (ML) projects requires a trustworthy and resilient federated learning (FL) environment to ensure that sensitive information such as protected health information is kept confidential. Specifically designed for this purpose, this work introduces APPFLx - a low-code, easy-to-use FL framework that enables easy setup, configuration, and running of FL experiments. APPFLx removes administrative boundaries of research organizations and healthcare systems while providing secure end-to-end communication, privacy-preserving functionality, and identity management. Furthermore, it is completely agnostic to the underlying computational infrastructure of participating clients, allowing an instantaneous deployment of this framework into existing computing infrastructures. Experimentally, the utility of APPFLx is demonstrated in two case studies: (1) predicting participant age from electrocardiogram (ECG) waveforms, and (2) detecting COVID-19 disease from chest radiographs. Here, ML models were securely trained across heterogeneous computing resources, including a combination of on-premise high-performance computing and cloud computing facilities. By securely unlocking data from multiple sources for training without directly sharing it, these FL models enhance generalizability and performance compared to centralized training models while ensuring data remains protected. In conclusion, APPFLx demonstrated itself as an easy-to-use framework for accelerating biomedical studies across organizations and healthcare systems on large datasets while maintaining the protection of private medical data.

Biomedical Research↗

New data-driven approach to bridging power system protection gaps with deep learning

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this paper, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A combined convolutional neural network and long short-term memory (CNN-LSTM) network is implemented to achieve data translation invariance and capture the temporal correlation of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN-LSTM--based method avoids the complicated, manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on two kinds of protection gaps: high-impedance faults and transformer inter-turn faults. Lastly, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

42 ENGINEERING↗

Nine Canyon Long-Duration Energy Storage: A Feasibility Study

The Nine Canyon Long Duration Energy Storage (LDES) Feasibility Study explores the technical and economic viability of deploying advanced energy storage technologies at Energy Northwest's (EN) Nine Canyon (9C) Wind Project site in Benton County, Washington. Supported by the Washington State Department of Commerce and the U.S. Department of Energy’s Office of Electricity under its LDES Voucher Program, the study represents a collaborative effort between EN, Pacific Northwest National Laboratory (PNNL), and ARES North America. At the core of this effort is the development of a generalized techno-economic modeling framework and evaluation tool designed to assess the value proposition of LDES projects across a variety of contexts. The modeling tool is technology-agnostic and accommodates user-defined parameters such as rated power, energy duration, round-trip efficiency, capital and operational costs, and dispatch constraints. It also integrates economic inputs, including market prices, energy revenue structures, and financing parameters to evaluate performance through key metrics. The tool provides utilities with a transparent, adaptable platform to support decision-making, investment prioritization, and portfolio planning for various storage technologies. To guide scenario design and interpretation, the study first surveyed the LDES technology landscape, including lithium-ion batteries, flow batteries, non-hydro gravity storage, and thermo-mechanical systems, comparing cost trajectories, technical performance, safety and hazards, materials sourcing and recyclability, and spatial/siting considerations. This literature-grounded review highlights technology trade-offs and reinforces the need to align technology choice with site characteristics, use cases, and project objectives. A companion chapter examines ownership structures (EN ownership, third-party ownership, shared models) and offtake options (energy marketing, capacity/energy PPAs, time-of-use PPAs, block-delivery PPAs, and tolling), where PPAs (power purchase agreements) represent contractual arrangements for buying and selling electricity. The chapter also highlights implications for risk allocation, capital access, operational control, and revenue certainty. The study also evaluates supervisory control and data acquisition (SCADA) and transmission interconnection pathways, options include upgrading the existing SCADA or deploying a dedicated LDES controller, with attention to protection schemes, data telemetry, cybersecurity, and regulatory coordination with BPA. In addition, an ARES-specific geotechnical and hydrology assessment presented in the appendix screens multiple corridors for slope stability, bearing capacity, cut-and-fill magnitude, and stormwater behavior.

25 ENERGY STORAGE↗

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics↗

Optimal Control of Differentially Private EV Charging: A Scalable Learning Approach Under Uncertainty

Internet of Things (IoT)-enabled electric vehicles (IoEVs) enable intelligent charging coordination that accounts for grid congestion. However, increased data exchange raises privacy concerns, as charging patterns can reveal sensitive driver behavior to grid operators. Here, we propose a differentially private (DP) EV charging framework that enables coordinated control while protecting driver data with theoretical privacy guarantees. Nevertheless, integrating DP inevitably introduces uncertainty into the control strategy for EVs, which can lead to infeasible solutions. To tackle this challenge, we develop a feasible and scalable control algorithm based on constrained reinforcement learning (CRL) and convex hulls. While our framework is designed to handle the uncertainty introduced by DP, it is general and also applicable to other sources of uncertainty in EV charging, such as the stochastic nature of driver behavior and renewable variability. This ensures feasible and privacy-preserving coordination of EV charging at scale. Our method constructs convex hulls within the action space to guarantee feasibility under stochastic constraints and incorporates constraint reduction techniques to improve scalability. Case studies based on IEEE benchmark systems demonstrate that the proposed approach effectively balances feasibility under uncertainty, scalability, and privacy in large-scale EV charging control.

Engineering - Power transmission and distribution↗

Complete and Correct Transfer of Information (CACTI)

Many distributed systems, file transfer mechanisms, and message passing systems offer reliability mechanisms such as acknowledgements, retries, and durability. While these tools may be “good enough” for their typical use cases, they may not offer sufficient coverage for the wide range of faults that impact data transfers and communication. A gap in the reliability measures may lead to some small amount of data loss. Some high-consequence systems cannot tolerate the loss or corruption of even a single record. We present seven principles that will counter a wide range of faults and protect against data loss and corruption. These principles bring together lessons learned from a wide range of technologies and can inform appropriate system design and application usage. These principles will help readers reason on how prevent data loss in a multi-hop pipeline and how to properly use tools that may have a deficiency in reliability.

97 MATHEMATICS AND COMPUTING↗

Mitigate: An Adaptive Network Data Anonymization Tool Using Condensation-Based Differential Privacy

Modern network devices collect a large amount of data that can be analyzed to identify bottlenecks, anomalies, cyber-attacks, etc. Therefore, there is often a need to analyze such collections of network data quite often by an external expert or by the research community. However, these collections of data contain sensitive, proprietary information. In order for the network data to be shared, it must first be anonymized. The overall objective of this project is to develop an innovative privacy management tool to anonymize network data and achieve sufficient privacy, acceptable data utility, and efficient data analysis at the same time. No existing anonymization methods can achieve all of these at the same time. The core of this technology is a differential private clustering algorithm that provides strong privacy protection, preserves data properties important for subsequent analysis, and allows the party receiving the anonymized data to conduct analysis directly on anonymized data without the need of decryption or any extra processing. The research carried out was to design, implement and verify a solution to this problem by completing the following tasks: 1) developing the core technology; 2) developing a context based method that automatically recommends fields that must be anonymized; 3) conducted experiments showing superior results using our approach compared to existing tools, and 4) developed an intuitive but basic user interface. The research that was conducted generated novel algorithmic techniques that utilize state-of-the-art methods such as condensation, differential privacy preservation, clustering, automated tuning based on contextual awareness, and recommendation techniques to specify columns to users for anonymization leading to optimal privacy that allows research analysis on the dataset. Experiments were conducted to evaluate the efficacy of these novel algorithmic techniques by performing analysis on original non-anonymized datasets, then conducting analysis on the same yet anonymized datasets and comparing the results of the analyses. Overall, the anonymized analysis results were within 1% of the original results, verifying that the generated technology not only guarantees a high level of privacy but also enables research analysis as if it were conducted on the original dataset. Potential applications of this technology include anonymization of any type of structured network datasets that contain sensitive identifiers, such as IP addresses, that can be used in multiple applications. For example, to create an AI or machine learning model for cyber security, e.g., to detect attacks, or for performance analysis, e.g., identify bottlenecks or predict performance. In addition, a market analysis that was conducted for potential applications of this technology identified a broader range of applications of our anonymization technology beyond the network sector that includes healthcare, banking, insurance, securities, finance (FISB), data brokering, cloud services, ad sales, and government.

97 MATHEMATICS AND COMPUTING↗

United States Multi-Sector Dynamics land use and land cover base maps to support Human-Earth System Modeling

Datasets are land use and land cover (LULC) rasterized base maps at 30-m resolution for the conterminous United States (CONUS) for the years 2008, 2011, 2016, and 2019. Separate base maps are provided where LULC classifications are thematically congruent with Community Land Model (CLM), Land Use Harmonization (LUH2), and Global Change Analysis Model (GCAM), and a detailed decomposition of all combined land classes into a Multisector Dynamics (MSD) LULC product. Base maps were developed using empirically derived satellite (National Land Cover Dataset, MODIS) and combined observation datasets (Crop Data Layer, Protected Areas Database) and represent the most up-to-date accurate information on LULC in the CONUS. The four datasets encompass four different landcover classification systems: MSD Layers - The raw landcover classes obtained from reclassifying NLCD and USDA Crop data layers into a respective landcover class GCAM Layers - The MSD classes mosaiced, reclassified, and combined into the respective GCAM landcover classes CLM Layers - Similar process to GCAM layers, but mosaiced, reclassified, and combined MSD layers to their respective PFT classes LUH2 Layers - Similar process to both GCAM and CLM Layers, but mosaiced, reclassified and combined the MSD layers to align with the respective states

Food↗

Data Security Defense: Modeling and Detection of Synchrophasor Data Spoofing Attack for Grid Edge

Data security and cyberattack have become critical issues in the distributed power system where adversaries can swap the source information of sensors or even spoof and alter measurements. However, the cyber security of the power system is challenged by the unpredictability and stealth of the spoofing attacks. Here, to protect the data security at the grid edge, this paper developed a synchrophasor data spoofing attack detection framework based on the time-frequency feature extraction techniques including the short-time Fourier transform (STFT) and object detection network for real-time synchrophasor data categorization and spoofing attack localization. The proposed approach outperforms earlier work in terms of spoofing attack detection and offers a vital localization function employing distributed synchrophasor sensors.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Analysis of Slow Spill Data for the Mu2e Experiment

The execution of the Mu2e experiment requires a stable, low-intensity proton beam from the Delivery Ring to produce clean data and protect equipment. This is done by performing a “slow extraction,” which is the gradual contraction of the stable region within the accelerator’s beam pipe. The Delivery Ring is currently unable to perform slow extraction with the stability required by Mu2e. To resolve this, the FAN-C team is training machine learning models with the purpose of replacing the Delivery Ring’s current PID controllers with AI-powered controllers. Training these models requires clean, processed data from slow spills. Over the course of this project, data from previous slow spills were processed and analyzed, and the clean data, graphs, and insights gained from the process were provided to the FAN-C team to assist them in their efforts.

Osborn, Thomas [Purdue U., West Lafayette]↗

SNM Radiation Signature Classification Using Different Semi-Supervised Machine Learning Models

The timely detection of special nuclear material (SNM) transfers between nuclear facilities is an important monitoring objective in nuclear nonproliferation. Persistent monitoring enabled by successful detection and characterization of radiological material movements could greatly enhance the nuclear nonproliferation mission in a range of applications. Supervised machine learning can be used to signal detections when material is present if a model is trained on sufficient volumes of labeled measurements. However, the nuclear monitoring data needed to train robust machine learning models can be costly to label since radiation spectra may require strict scrutiny for characterization. Therefore, this work investigates the application of semi-supervised learning to utilize both labeled and unlabeled data. As a demonstration experiment, radiation measurements from sodium iodide (NaI) detectors are provided by the Multi-Informatics for Nuclear Operating Scenarios (MINOS) venture at Oak Ridge National Laboratory (ORNL) as sample data. Anomalous measurements are identified using a method of statistical hypothesis testing. After background estimation, an energy-dependent spectroscopic analysis is used to characterize an anomaly based on its radiation signatures. In the absence of ground-truth information, a labeling heuristic provides data necessary for training and testing machine learning models. Supervised logistic regression serves as a baseline to compare three semi-supervised machine learning models: co-training, label propagation, and a convolutional neural network (CNN). In each case, the semi-supervised models outperform logistic regression, suggesting that unlabeled data can be valuable when training and demonstrating value in semi-supervised nonproliferation implementations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Data for Training and Testing Radiation Detection Algorithms in an Urban Environment

The US government routinely performs radiological response deployments to search for the presence of illicit nuclear materials (e.g., highly enriched uranium and weapons-grade plutonium) in a specified area. The deployments can be intelligence driven, in support of law enforcement, and for planned events such as WrestleMania, presidential inaugurations, or political conventions. In a typical deployment, radiation detection systems carried by human operators or mounted on vehicles move in a clearing pattern through the search area. Search teams rely on radiation detection algorithms running on these systems in real time to alert them to the presence of an illicit threat source. The detection and identification of sources is complicated by large variation of natural radiation background throughout a search area and the potential presence of localized non-threat sources such as patients undergoing treatment with medical isotopes. As a result, detection algorithms must be carefully balanced between missing real sources (false negatives) and reporting too many false alarms (false positives).The purpose of this data set is to spur innovations in detecting, identifying, and localizing nuclear materials inurban search missions.

07 ISOTOPE AND RADIATION SOURCES↗

Measuring Cities with Software-Defined Sensors

The Chicago Array of Things (AoT) project, funded by the US National Science Foundation, created an experimental, urban-scale measurement capability to support diverse scientific studies. Initially conceived as a traditional sensor network, collaborations with many science communities guided the project to design a system that is remotely programmable to implement Artificial Intelligence (AI) within the devices-at the “edge” of the network-as a means for measuring urban factors that heretofore had only been possible with human observers, such as human behavior including social interaction. The concept of “software-defined sensors” emerged from these design discussions, opening new possibilities, such as stronger privacy protections and autonomous, adaptive measurements triggered by events or conditions. We provide examples of current and planned social and behavioral science investigations uniquely enabled by software-defined sensors as part of the SAGE project, an expanded follow-on effort that includes AoT.

97 MATHEMATICS AND COMPUTING↗

Erratum to: Data Fusion to Support Integrated Nuclear Detonation Detection [Slides]

The original document (LA-UR-22-29547) contained minor equation errors that approximate correct equations that couple the multi-sensor, serial system detector thresholds for a seismic Rayleigh wave detector and an acoustic energy detector. Those errors appeared on slides 62-67. This erratum associates the following slides with the erroneous slides. The numbers of the erroneous slides are marked at the upper right in small text in orange. A result of those errors over-predict the performance of the two-sensor serial network, that is, the former document provides an optimistic estimate of system performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A secure measurement unit for an inspection system used in nuclear arms-control verification

Verification of arms control treaties may require information barriers to protect sensitive data acquired during the verification process. These information barriers are commonly implemented in software, but must store and operate on data that are vulnerable to attack and tampering. Here, we present a new Secure Measurement Unit based on a modified pipeline SAR ADC architecture. The idea is to move the information barrier back into the digitization process so that prohibited data is never created. A prototype 65nm CMOS Secure Measurement Unit (SMU) was tested using U-235, U-238, Co-60, Cs-137 and Am-241 sources. Here, the prototype accurately digitized allowable signals while being immune to side-channel attacks on signals within preset forbidden regions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cyber-Informed Engineering Research and Development Guide

This document provides guidance on incorporating Cyber Informed Engineering (CIE) principles into the research and development (R&D) of operational technology systems and tools, facilitating the creation and adoption of innovative technologies that are secure and resilient by design. As technological innovation and research are becoming pivotal for economic and national security, cybersecurity has emerged as a paramount concern across industries and sectors. The challenge of integrating robust cybersecurity measures is imperative to safeguard critical infrastructure, protect sensitive data, and preserve national security interests.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Side-channel Leakage Assessment Metrics: A Case Study of GIFT Block Ciphers

Determination of an adequate level of security and providing subsequent mechanisms to achieve it, is one of the most pressing problems regarding embedded computing devices. While there are some solutions available for resource-rich computer systems, direct application of these solutions to resource-constrained environments are often unfeasible. The fundamental problem for such resource-constrained systems is the fact that current cryptographic algorithms utilize significant energy consumption and storage overhead. Both the cryptographic algorithm and its physical implementation affect the resilience of a cryptosystem against side-channel attacks. A side-channel attack represents a process that exploits leakages in order to extract sensitive information such as the key. This paper focuses on Correlation Power Analysis (CPA) which is side-channel attack based on the power consumption leakage. In 2016 the U.S. Commerce Department’s National Institute of Standards and Technology (NIST) initiated the call for proposals of new cryptographic algorithms to strengthen the cryptographic defense of networked devices against cyberattacks and to protect the data created by those innumerable device. This work evaluates S-boxes used by NIST candidates PICCOLO, GIFT, and PRESENT, as well as several S-box variants that demonstrated sufficient weaknesses against classical cryptanalysis, for a quantitative comparison in terms of resiliency to CPA attack. Three well-known theoretical metrics are evaluated: transparency order (TO and RTO), nonlinearity, and signal-to-noise (SNR) ratio, aiming to characterize the resistance of these S-boxes against adversaries exploiting physical leakages. Experimental results from attacks on an 8- bit XMEGA were obtained via the ChipWhisperer platform and of all the S-boxes evaluated, GIFT64 with a PICCOLO S-box was found to be the most susceptible to CPA. Results showed that variations in TO and RTO were not sufficient to ensure practical CPA resistance and that among S-boxes with equal non-linearity there were no significant differences in the TO and SNR variants.

97 MATHEMATICS AND COMPUTING↗

Side-channel Leakage Assessment Metrics: A Case Study of GIFT Block Ciphers

Determination of an adequate level of security and providing subsequent mechanisms to achieve it, is one of the most pressing problems regarding embedded computing devices. While there are some solutions available for resource-rich computer systems, direct application of these solutions to resource-constrained environments are often unfeasible. The fundamental problem for such resource-constrained systems is the fact that current cryptographic algorithms utilize significant energy consumption and storage overhead. Both the cryptographic algorithm and its physical implementation affect the resilience of a cryptosystem against side-channel attacks. A side-channel attack represents a process that exploits leakages in order to extract sensitive information such as the key. This paper focuses on Correlation Power Analysis (CPA) which is side-channel attack based on the power consumption leakage. In 2016 the U.S. Commerce Department’s National Institute of Standards and Technology (NIST) initiated the call for proposals of new cryptographic algorithms to strengthen the cryptographic defense of networked devices against cyberattacks and to protect the data created by those innumerable device. This work evaluates S-boxes used by NIST candidates PICCOLO, GIFT, and PRESENT, as well as several S-box variants that demonstrated sufficient weaknesses against classical cryptanalysis, for a quantitative comparison in terms of resiliency to CPA attack. Three well-known theoretical metrics are evaluated: transparency order (TO and RTO), nonlinearity, and signal-to-noise (SNR) ratio, aiming to characterize the resistance of these S-boxes against adversaries exploiting physical leakages. Experimental results from attacks on an 8- bit XMEGA were obtained via the ChipWhisperer platform and of all the S-boxes evaluated, GIFT64 with a PICCOLO S-box was found to be the most susceptible to CPA. Results showed that variations in TO and RTO were not sufficient to ensure practical CPA resistance and that among S-boxes with equal non-linearity there were no significant differences in the TO and SNR variants.

97 MATHEMATICS AND COMPUTING↗