Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Anonymization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Pseudonymization at Scale: OLCF’s Summit Usage Data Case Study

The analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available.

Maheshwari, Ketan↗

Scalable and Secure Power Outage Data Reporting: A Hexagonal Geospatial Approach

Power outages disrupt critical infrastructure and cause billions of dollars in economic losses annually in the United States. Accurate and granular outage reporting is vital for effective restoration and mitigation. This paper examines the integration of the Hexagonal Hierarchical Geospatial Indexing System (H3) to enhance power outage reporting, leveraging its uniform grid structure, scalable resolutions, and support for privacy-preserving analysis. Using high-resolution LandScan Global population data and K-anonymization techniques, this work achieves a balance between data granularity and privacy. Results show that lower privacy thresholds (e.g., K-anonymity = 2) enable higher resolution, while stricter thresholds (e.g., >15 people per hex) reduce granularity, potentially affecting localized responses. State-and county-level resolution case studies demonstrate H3’s adaptability and the trade-offs between precision and privacy. The proposed H3-based framework offers a scalable and efficient solution for geospatial data integration within the energy sector, such as outage data, aiding utilities and regulators in improving resilience and response efforts, particularly in disaster-prone regions.

Ahmad, Nasir [ORNL] (ORCID:0000000150677368)↗

A Privacy-Preserving Cyber Threat Intelligence Sharing System

Cyber Threat Intelligence (CTI) is a key resource for developing defensive strategies against potential cyber adversaries. Entities typically access CTI through open-source platforms, national agencies, or specialized commercial services. However, the bi-directional exchange of CTI is hindered by organizational trust boundaries, which complicate the sharing processes between entities and CTI providers. Centralized CTI services benefit from receiving suspicious cyber observables such as IP addresses, domain names, and email addresses from various entities. The aggregation allows for the correlation of widespread adversarial activities to enhance the alert and response mechanisms across the network of involved parties. Despite these benefits, openly sharing such observables incurs potential legal, regulatory, and reputational risks for the disclosing entities.This paper introduces a system designed to facilitate the secure exchange of cyber observables across trust boundaries without compromising the anonymity of the sharing entities. Here, we propose an architecture that leverages common web protocols alongside zero-knowledge proofs to authenticate members while maintaining anonymity. Additionally, we outline a privacy model tailored for STIX (Structured Threat Information eXpression) cyber observables to minimize the risk of inadvertently disclosing private information. Through our threat models, we assess the privacy implications of our proposed system and demonstrate its potential to enhance collaborative cyber defense efforts without exposing entities to undue risk.

BBS+ Signatures↗

Deep Design Data Portal (D3P) v0.01

The Deep Design Data Portal (D3P) tool was developed to demonstrate how readily accessible data sources, such as building energy model reports for design and baseline energy performance data for projects, can provide the data required for reporting to an industry initiative (AIA 2030 commitment), as well as more detailed data that makes the industry dataset more valuable to all stakeholders, enabling project level analysis and analysis of BEM industry trends. D3P provides an easier and less time-consuming way for firms to auto-extract data from this data source, compared to the current reporting workflows of the firms. The BEM reports are the first of several data sources that D3P could integrate. D3P also provides the ability for firms to review, compare, and evaluate the performance of their projects to not only their portfolio, but also to the larger anonymized industry dataset created each time a project is added to D3P. The intent of D3P is to become part of a data-sharing ecosystem to assist creating large anonymized industry datasets that are accessible to industry.

Regnier, Cynthia [Lawrence Berkeley National Labor↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Differential Privacy in Grid Kitchen: Implementation & Software Documentation

Sharing of power grid feeder models faces significant challenges due to the potential risk of exposing sensitive operational information. Traditional anonymization techniques have shown notable limitations in other sensitive domains, as evidenced by documented re-identification attacks that combine supposedly anonymized datasets with auxiliary information, raising concerns that similar vulnerabilities could affect power grid data. Consequently, there is a pressing need for a more rigorous privacy protection strategy that not only delivers formal mathematical guarantees but also preserves the analytical value of the shared models. To address this challenge, we have enhanced the Grid Kitchen framework by implementing differential privacy mechanisms within the distribution model dehydration pipeline. This implementation carefully calibrates and applies noise to sensitive attributes in feeder models according to configurable privacy levels—low, moderate, and high—each offering different balances between data utility and privacy protection. Our approach uses established noise functions (Gaussian for continuous data and Discrete Laplace for integer values) with parameters carefully calibrated so that the impact of individual data points is effectively masked in the final output. The integration leverages our Noise Catalog, which we developed to categorize feeder model properties by component type, data type, and sensitivity. This catalog guides the application of appropriate noise functions and privacy parameters ($\varepsilon$ and $\delta$) to each attribute, ensuring consistent privacy protection across the model while maintaining its structural integrity and analytical usefulness. This implementation also includes evaluation tools that allow model owners to assess the impact of privacy-preserving transformations before sharing data with external parties. This report provides documentation for the differential privacy capabilities added to the Grid Kitchen project. It includes a primer on differential privacy concepts and their importance in modern data sharing, details the architecture of our implementation, explains the privacy modes and parameter configurations, and offers practical guidance on using the code for applying differential privacy to grid feeder models. Through examples and code snippets, we demonstrate the effective application of these privacy-enhancing technologies, enabling utility operators and researchers to confidently share grid data while protecting sensitive information.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FOA 1861 Data Curation Overview

This document describes the process executed to collect, examine, and consolidate Phasor Measurement Unit (PMU) data from multiple transmission operators into a common dataset. The consolidated PMU data set was further anonymized and distributed to the Department of Energy Funding Opportunity Announcement (FOA) 1861 Big Data Analysis of Synchrophasor Data awardees.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning

Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets () according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ -aware server aggregation to improve model utility by re-weighting client updates according to their declared privacy budgets. However, gradient updates in FL retain structural patterns induced by non-independent and identically-distributed (non-IID) data, and these additional signals exposed by -aware aggregation create new opportunities for inference by an honest-but-curious server. In this work, we first show that a server equipped with gradient denoising and surrogate modeling can mount a Privacy Inference Attack that infers distributional attributes of clients and links updates from the same client across training rounds, measured via surrogate inference accuracy and linkage success, under realistic knowledge constraints. The Shuffle-Model has been widely studied as a defense against such inference risks by anonymizing update sources, but it is fundamentally incompatible with HDP-FL -aware aggregation. To address this challenge, we propose IntraShuffler, a middleware defense framework designed for HDP-FL systems. IntraShuffler introduces a privacy-aware shuffling mechanism that groups clients into privacy-compatible buckets and performs parameter-level shuffling within each bucket to disrupt persistent gradient structure while preserving -aware aggregation. Experiments across four different datasets show that IntraShuffler reduces gradient recoverability by over 60% and decreases surrogate inference accuracy from 0.78 to 0.33 while maintaining comparable model utility across multiple FL aggregation rules.

Riya, Farhin Farhad [ORNL]↗

Using advanced data structures to enable responsive security monitoring

Write-optimized data structures (WODS), offer the potential to keep up with cyberstream event rates and give sub-second query response for key items like IP addresses. These data structures organize logs as the events are observed. To work in a real-world environment and not fill up the disk, WODS must efficiently expire older events. As the basis for our research into organizing security monitoring data, we implemented a tool, called Diventi, to index IP addresses in connection logs using RocksDB (a write-optimized LSM tree). In this work, we extended Diventi to automatically expire data as part of the data structures’ normal operations. We guarantee that Diventi always tracks the N most recent events and tracks no more than N + k events for a parameter k < N, while ensuring the index is opportunistically pruned. To test Diventi at scale in a controlled environment, we used anonymized traces of IP communications collected at SuperComputing 2019. We synthetically extended the 2.4 billion connection events to 100 billion events. We tested Diventi vs. Elasticsearch, a common log indexing tool. In our test environment, Elasticsearch saw an ingestion rate of at best 37,000 events/s while Diventi sustained ingestion rates greater than 171,000 events/s. Our query response times were as much as 100 times faster, typically answering queries in under 80 ms. Furthermore, we saw no noticeable degradation in Diventi from expiration. We have deployed Diventi for many months where it has performed well and supported new security analysis capabilities.

97 MATHEMATICS AND COMPUTING↗

Interlaboratory Reproducibility of Contour Method Data in a High Strength Aluminum Alloy

The contour method for residual stress measurement has seen significant development, but an experimental reproducibility study utilizing physical samples has not been published. A double-blind reproducibly study is reported, having scope beginning with EDM cutting and ending with residual stress calculation. A reinforced I-beam sample geometry is identified for its unique residual stress profile when extracted from residual stress bearing quenched aluminum bar (7050-T74). Contour measurements are prescribed on a midplane of symmetry with dimensions 24.0 mm by 50.0 mm. Fourteen identically prepared samples are fabricated from a single long bar with well characterized and uniform residual stress. Five samples throughout the bar are identified for planning measurements to validate sample uniformity and overall suitability of the residual stress field. The planning measurements employ a range of techniques: contour method, neutron diffraction, and hole-drilling. Eight samples are distributed to an international group of participants to execute their standard measurement practice. A double-blind process is followed to provide anonymity. Results are provided by eight participants: six being self-similar and two being quite different, the latter set aside as outliers. An average residual stress field is established from non-outlying results and the spatial distribution of reproducibility standard deviation is determined. The average stress field ranges from -60 to 70 MPa and the reproducibility standard deviation averages 8.1 MPa on the measurement plane. The average reproducibility standard deviation is about 3 × larger for points within 1.0 mm of plane boundaries (17.6 MPa) than for the remaining points (6.1 MPa). Reproducibility standard deviation (among different labs) for contour method residual stress measurement is found to be very similar to repeatability standard deviation (in a single lab) reported in prior work. The reproducibility observed here, for the entire measurement process, is also similar to that found in a prior reproducibility study limited to contour method data analysis.

36 MATERIALS SCIENCE↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Testing SOAR tools in use

Investigations within Security Operation Centers (SOCs) are tedious as they rely on manual efforts to query diverse data sources, overlay related logs, correlate the data into information, and then document results in a ticketing system. Security Orchestration, Automation, and Response (SOAR) tools are a relatively new technology that promise, with appropriate configuration, to collect, filter, and display needed diverse information; automate many of the common tasks that unnecessarily require SOC analysts’ time; facilitate SOC collaboration; and, in doing so, improve both efficiency and consistency of SOCs. There has been no prior research to test SOAR tools in practice; hence, understanding and evaluation of their effect is nascent and needed. Here, in this paper, we design and administer the first hands-on user study of SOAR tools, involving 24 participants and six commercial SOAR tools. Our contributions include the experimental design, itemizing six characteristics of SOAR tools, and a methodology for testing them. We describe configuration of a cyber range test environment, including network, user, and threat emulation; a full SOC tool suite; and creation of artifacts allowing multiple representative investigation scenarios to permit testing. We present the first research results on SOAR tools. Concisely, our findings are that: per-SOC SOAR configuration is extremely important; SOAR tools increase efficiency and reduce context switching, although with potentially decreased ticketing accuracy/completeness; user preference is slightly negatively correlated with their performance with the tool; internet dependence varies widely among SOAR tools; and balance of automation with assisting decision making is preferred by senior participants. We deliver a public user- and tool-anonymized and -obfuscated version of the data.

97 MATHEMATICS AND COMPUTING↗

Feature review of photovoltaic modeling software utilizing blind performance assessment

While confidence in photovoltaic (PV) modeling software has always been essential, the rapid pace of new PV plant developments makes accuracy and credibility more critical than ever. Independent assessments, particularly through blind modeling comparisons, are therefore necessary to ensure unbiased benchmarking across PV modeling software. Previous studies have been limited by a narrow range of models compared, anonymized results, or system size. This study presents results from the first-ever onymous blind modeling comparison, evaluated using both lab- and utility-scale fixed-tilt, monofacial, south-facing systems at sub-hourly time intervals. Seven commercially used PV software tools were compared: 3E SynaptiQ, PlantPredict, PVsyst, RatedPower, SAM, SolarFarmer, and Solargis Evaluate. Predictions were submitted directly by software representatives, providing unique insights into each software’s implementation and resulting prediction behavior. Notable features, including plane-of-array (POA) transposition model, module temperature model, shading model, and performance model were analyzed and compared. Four summary tables compile these features of the software, serving as a resource to help users understand the methodological differences and select the most suitable software for their applications. The software tools show deviations from mean error in annual yield up to 2.5 % in the lab-scale system, increasing to 6.0 % for the utility-scale system. These differences arise from a combination of user decisions and the inherent behavior of the software, indicating the need for continuous and rigorous validation of modeling methods using these software tools against complex, real-world systems.

14 SOLAR ENERGY↗

Spatial and Temporal Characterization of Activity in Public Space, 2019–2020

The data reported here characterize spatial and temporal variation in the ratio of short-to-long-duration visits in public places (i.e., points of interest) in the United States for each week between January 2019 and December 2020. The underlying data on anonymized and aggregated foot traffic to public places is curated by SafeGraph, a geospatial data provider. In this work, we report the estimated number and duration of “short” (i.e., <4 hours) and “long” (i.e., >4 hours) visits to public places at the US census block group level. Long visits are shown to be a good proxy for workers based on formal economic data. We propose that short visits are more likely to represent nonobligate activities: people visiting a public place for leisure, shopping, entertainment, or civic or cultural engagement. Our work constructs a ratio of short to long visits, which can be used to inform population estimates for nonworker use of public space. These data may be useful for understanding how people’s use of public space has changed during the COVID-19 pandemic and, more generally, for understanding activity patterns in public.

99 GENERAL AND MISCELLANEOUS↗

Crowd cluster data in the USA for analysis of human response to COVID-19 events and policies

We provide data on daily social contact intensity of clusters of people at different types of Points of Interest (POI) by zip code in Florida and California. This data is obtained by aggregating fine-scaled details of interactions of people at the spatial resolution of 10 m, which is then normalized as a social contact index. We also provide the distribution of cluster sizes and average time spent in a cluster by POI type. This data will help researchers perform fine-scaled, privacy-preserving analysis of human interaction patterns to understand the drivers of the COVID-19 epidemic spread and mitigation. Current mobility datasets either provide coarse-level metrics of social distancing, such as radius of gyration at the county or province level, or traffic at a finer scale, neither of which is a direct measure of contacts between people. We use anonymized, de-identified, and privacy-enhanced location-based services (LBS) data from opted-in cell phone apps, suitably reweighted to correct for geographic heterogeneities, and identify clusters of people at non-sensitive public areas to estimate fine-scaled contacts.

60 APPLIED LIFE SCIENCES↗

Effectiveness of Privacy Techniques in Smart Metering Systems

Smart grid technologies enable timely energy billing for residential homes. The ability to react to energy demands during peak hours allows energy providers to conserve power and operate efficiently. However, these data streams are also susceptible to privacy attacks within the energy company and from outside hackers. We implemented four different privacy models: k-anonymous, l-diversity, t-closeness, and ε-differential privacy. We demonstrate the models’ effectiveness using a real-world dataset composed of 15 different residential households with energy consumption data spanning over a year.

Peralta-Peterson, Martin↗

Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-the-Meter Solar

The ever-growing high penetration of ubiquitously distributed energy resources, especially behind-the-meter solar (BTM) generations, has significant impacts on nodal load (i.e., net injection) profiles and consequently caused imperative operational challenges to system operators such as regional transmission organizations (RTOs). Illustrated by real-world nodal data and examples at PJM Interconnection, this paper first discusses the application and necessity of effectively extracting daily nodal load profiles in a non-intrusive manner. More importantly, a novel bi-level architecture, including Factorization Machines (FM) learning procedure has been proposed to effectively disaggregate not only one node but every node in an RTO service territory. Specifically, FM leaning is adopted to capture the interconnections between related features to better utilize the correlation between buses in the same region and between a single bus and the zonal load. The proposed bi-level technique is numerically validated using real-world, minute-level, normalized, and anonymized nodal data at PJM service territory.

behind the meter solar, load disaggregation, load ↗

Advancing Our Understanding of System Availability through the PV Fleet Performance Data Initiative

The PV Fleet Performance Data Initiative partners with photovoltaic (PV) fleet owners to collect time-series data of PV production data and publishes aggregated anonymized results of system performance metrics. With an extensive dataset drawn from over 2,200 PV systems across the United States, comprising 8.5 GW and 24,000 separate inverter data channels, this initiative aims to ensure that systemic risks in the US PV fleet are detected. The current work explores system availability, revealing a pronounced dependence on time, especially within the initial 6 months of system performance. Following this start-up period, the average system availability stabilizes. Statistical analyses illustrate a median (P5O) monthly availability of 0.991 and a dependence on system size with a negative trend in availability with increasing system size. This finding indicates that larger systems experience lower availability compared to their smaller counterparts.

inverter availability↗