Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Anonymization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Protecting Customer Privacy Through Distributed Energy Resource Anonymization

Due to their stochastic nature, the increase of Renewable Energy Resources (RERs) as a primary source of energy for power grids creates challenges regarding the reliability and resilience of the system. In order to combat these obstacles, expansion of Distributed Energy Resources (DERs) and their participation in Demand Response (DR) programs is necessary. Widespread participation requires prioritizing customer privacy and addressing concerns that may arise regarding communication between DERs and the Grid Service Provider (GSP). This paper discusses the use of flow reservation resources to split the operating cycles of DER load profiles into unique phases. The splitting of phases increases anonymization of the DERs by making it more difficult to determine the individual characteristics of the device. We discuss an example of this using simulated DER load profile data and examine the resulting effectiveness by using a machine learning algorithm for classification, called Support Vector Machine (SVM).

Distributed Energy Resource, Anonymization, Renewa↗

Mitigate: An Adaptive Network Data Anonymization Tool Using Condensation-Based Differential Privacy

Modern network devices collect a large amount of data that can be analyzed to identify bottlenecks, anomalies, cyber-attacks, etc. Therefore, there is often a need to analyze such collections of network data quite often by an external expert or by the research community. However, these collections of data contain sensitive, proprietary information. In order for the network data to be shared, it must first be anonymized. The overall objective of this project is to develop an innovative privacy management tool to anonymize network data and achieve sufficient privacy, acceptable data utility, and efficient data analysis at the same time. No existing anonymization methods can achieve all of these at the same time. The core of this technology is a differential private clustering algorithm that provides strong privacy protection, preserves data properties important for subsequent analysis, and allows the party receiving the anonymized data to conduct analysis directly on anonymized data without the need of decryption or any extra processing. The research carried out was to design, implement and verify a solution to this problem by completing the following tasks: 1) developing the core technology; 2) developing a context based method that automatically recommends fields that must be anonymized; 3) conducted experiments showing superior results using our approach compared to existing tools, and 4) developed an intuitive but basic user interface. The research that was conducted generated novel algorithmic techniques that utilize state-of-the-art methods such as condensation, differential privacy preservation, clustering, automated tuning based on contextual awareness, and recommendation techniques to specify columns to users for anonymization leading to optimal privacy that allows research analysis on the dataset. Experiments were conducted to evaluate the efficacy of these novel algorithmic techniques by performing analysis on original non-anonymized datasets, then conducting analysis on the same yet anonymized datasets and comparing the results of the analyses. Overall, the anonymized analysis results were within 1% of the original results, verifying that the generated technology not only guarantees a high level of privacy but also enables research analysis as if it were conducted on the original dataset. Potential applications of this technology include anonymization of any type of structured network datasets that contain sensitive identifiers, such as IP addresses, that can be used in multiple applications. For example, to create an AI or machine learning model for cyber security, e.g., to detect attacks, or for performance analysis, e.g., identify bottlenecks or predict performance. In addition, a market analysis that was conducted for potential applications of this technology identified a broader range of applications of our anonymization technology beyond the network sector that includes healthcare, banking, insurance, securities, finance (FISB), data brokering, cloud services, ad sales, and government.

97 MATHEMATICS AND COMPUTING↗

Anonymization of Network Traces Data through Condensation-based Differential Privacy

Network traces are considered a primary source of information to researchers, who use them to investigate research problems such as identifying user behavior, analyzing network hierarchy, maintaining network security, classifying packet flows, and much more. However, most organizations are reluctant to share their data with a third party or the public due to privacy concerns. Therefore, data anonymization prior to sharing becomes a convenient solution to both organizations and researchers. Although several anonymization algorithms are available, few of them allow sufficient privacy (organization need), acceptable data utility (researcher need), and efficient data analysis at the same time. This article introduces a condensation-based differential privacy anonymization approach that achieves an improved tradeoff between privacy and utility compared to existing techniques and produces anonymized network trace data that can be shared publicly without lowering its utility value. Our solution also does not incur extra computation overhead for the data analyzer. A prototype system has been implemented, and experiments have shown that the proposed approach preserves privacy and allows data analysis without revealing the original data even when injection attacks are launched against it. When anonymized datasets are given as input to graph-based intrusion detection techniques, they yield almost identical intrusion detection rates as the original datasets with only a negligible impact.

97 MATHEMATICS AND COMPUTING↗

Geospatial and Information Substitution and Anonymization Tool (GISA)

The Geospatial and Information Substitution and Anonymization Tool (GISA) incorporates techniques for obfuscating identifiable information from point data or documents, while simultaneously maintaining chosen variables to enable future use and meaningful analysis. This approach promotes collaboration and data sharing while also reducing the risk of exposure to sensitive information. GISA can be used in a number of different ways, including the anonymization of point spatial data, batch replacement/removal of user-specified terms from file names and from within file content, and aid with the selection and redaction of images and terms based on recommendations using natural language processing. Version 1 of the tool, published here, has updated functionality and enhanced capabilities to the beta version published in 2023. Please see User Documentation for further information on capabilities, as well as a guide for how to download and use the tool. If there are any feedback you would like to provide for the tool, please reach out with your feedback to edxsupport@netl.doe.gov. Disclaimer: This project was funded by the United States Department of Energy, National Energy Technology Laboratory, in part, through a site support contract. Neither the United States Government nor any agency thereof, nor any of their employees, nor the support contractor, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. The Geospatial and Information Substitution and Anonymization Tool (GISA) was developed jointly through the U.S. DOE Office of Fossil Energy and Carbon Management’s EDX4CCS Project, in part, from the Bipartisan Infrastructure Law.

Bipartisan Infrastructure Law↗

Distributed Energy Resource Management Systems: Preserving Customer Privacy through K-Anonymity

The smart grid represents the next generation of electricity distribution systems that utilizes recent technological innovations. It uses digital communication between its components and entities to attain more automation, self-sufficiency, and reliability. One of the many concerns in smart grid digital communication discussions is the possibility of violating customers’ privacy. Violating customers’ privacy imposes a significant barrier as smart grid desirable attributes are tightly tied to customers’ participation. Employing privacy models can address concerns regarding information privacy in smart grid digital communication. In this work, we provide an approach to utilizing K-anonymity to ensure data within the system excludes Personally Identifiable Information. Results suggest that a dynamically generated generalization hierarchy minimizes information loss incurred by the anonymization process.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

WholeTraveler Anonymized Data Phase 1

Phase 1 of the WholeTraveler Study data collection consisted of an online-only survey. This survey captured data on three categories of observable variation in the population relevant to transportation decisions. First, the survey collected traditional demographic data such as age, gender, income, and education level. Second, it collected data across personality, psychological, and preference categories. This included: 1. The "Big Five" inventory personality traits: openness to new experience, conscientiousness, extroversion, agreeableness, and neuroticism; 2. Risk and time preferences; and 3. Environmental preferences. Third, the survey collected data on historical behavior patterns including: 1. Adoption of (as well as interest in) new technologies or innovations (e.g., smartphones, PEVs, solar panels, adaptive cruise control [ACC]); 2. Car ownership history and current car ownership status; 3. Recent mode use across different time scales (e.g., previous week, previous month, previous year); and 4. Timing of major life events such as starting a family as well as overall lifecycle trajectory patterns. Data from Phase 1 and Phase 2 are linked by a unique respondent identifier. Anonymized versions of the Phase 1 and Phase 2 data are both available on Livewire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

WholeTraveler Anonymized Data Phase 2

Phase 2 of the WholeTraveler study consisted of a global positioning system (GPS) data collection. Phase 2 started immediately after the completion of the Phase 1 survey for any respondent who opted into Phase 2. The raw locational data collected have been processed into identified "trips" and some of those trips into identified "trip chains." Data from Phase 1 and Phase 2 are linked by a unique respondent identifier. Anonymized versions of the Phase 1 and Phase 2 data are both available on Livewire.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

PCAP Anonymizer

Explore the source record for details and available documents.

42 ENGINEERING↗

K-anonymity applied to the energy grid of things distributed energy resource management system

Smart grid infrastructure relies on information exchange between multiple actors in order to ensure system reliability. These actors include but are not limited to smart loads, grid control, and energy management technologies. As information exchange between these actors is susceptible to cyber-attacks, security and privacy issues are indispensable to ensure a reliable and stable grid. This position paper proposes a privacypreserving, trust-augmented secure scheme for a smart grid implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

K-anonymization

New algorithms to de-identify geolocations. A unique feature of our algorithms is the ability to compare de-identified values that are similar but not the same

Bleeker, Amelia↗

Pathways to a Sustainable Aviation Ecosystem: Flight DNA: An Anonymized Aviation Data Tool and Repository

The National Renewable Energy Laboratory (NREL) has deep experience developing secure, national data repositories, which it augments with analysis, technology and market expertise, high-performance computing, and innovative data visualization. By adapting the architecture of existing mobility databases (i.e., Fleet DNA, the Transportation Secure Data Center, the National Fuel Cell Technology Evaluation Center), NREL can build a powerful aviation data clearinghouse - Flight DNA - that helps stakeholders navigate the web of pitfalls and possibilities generated by new aviation technologies.

aviation↗

DOE EV Data Collection - Vehicle Data

Vehicle data consist of electric vehicle performance data collected directly from the vehicle during standard operations. Data were collected using onboard data loggers that were either installed by the project team or preinstalled by the original equipment manufacturer. Data recorded by the data loggers were made accessible via an online web portal or an application programming interface. Different data loggers were used (HEM, ViriCiti, and Geotab), and the method for each vehicle is defined in the vehicle attributes file. Some systems collected data on a “trip-level” basis, in which each row of a table represents a single trip (the period between a key-on and key-off event), whereas other data were collected on a per-day basis, in which each row represents a single day of operation. Data were collected over a range of data collection periods, depending on the project. Data have been anonymized by removing information or decreasing information resolution as necessary so that fleets are not identifiable. Due to the wide range of vehicle types represented and variation in data collection, data parameters and frequencies differ between vehicles and fleets The **Performance Data Daily/Trip Data Dictionaries** contain definitions for each available parameter associated with a vehicle’s operations, aggregated at either a daily or trip level. The parameters available will vary from vehicle to vehicle, but every possible parameter will be defined. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. The **Vehicle Data** tables contain the data from each vehicle’s operations, aggregated at either a daily or trip level, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DOE EV Data Collection - Charging Data

Charging data are collected from one of three sources, each with varying levels of additional information. These sources, in approximate order from most to least additional information, are: • The electric vehicle supply equipment (charger) • Onboard the vehicle itself • From a utility submeter. Many chargers provide software that allows for the collection and reporting of charging session data. If unavailable, data may be recorded by the charging vehicle’s onboard systems. If neither of these options is available, data can be acquired from utility submeters that simply track the energy flowing to one or more chargers. Data collected directly from the electric vehicle supply equipment (EVSE) are typically the most accurate and highest frequency. However, it is not always possible to discern which exact vehicle is being charged during any one session. EVSE-side data can be identified where a single charger ID but a range of vehicle IDs are present (e.g., CH001, EV001-EV005). Data collected from the vehicle’s onboard systems usually does not provide information on which exact charger is being used. Vehicle-side data can be identified where a single Vehicle ID but a range of Charger IDs are present (e.g., EV001, CH001-CH005). Data collected from utility submeters provide no information on which specific vehicle is charging or which specific charger is in use. Submeter data can be identified where multiple Vehicle IDs and multiple Charger IDs are present, but only a single Fleet ID is present (e.g., EV001-EV005, CH001-CH005, Fleet01). The **Charge Data Daily/Session Dictionaries** contains definitions for each available parameter collected as part of an individual charging session, aggregated at either a daily or session level. The parameters available will vary between vehicles and chargers. The **Charger Attributes** table contains specific charger characteristics, coded to at least one anonymous Charger ID and linked to either a single or a range of Vehicle IDs. Vehicle ID can be used as a key between charging data and vehicle attribute tables. The **Charger Attributes Data Dictionary** contains definitions for each available parameter collected on the physical and operational characteristics of the charging hardware itself. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables, and in cases where charging data are supplied, links a vehicle with the charger(s) that supplied it power. The **Charging Data** tables contain the data from each charger’s operations, coded to at least one anonymous Charger ID and linked to either a single or a range of Vehicle IDs. Vehicle ID can be used as a key between charging data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

NLR HPC Eagle Jobs Data and Additional Energy Metrics

Overview: Anonymized job-level records from the Eagle high-performance computing (HPC) system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, resource utilization, CPU/GPU energy consumption, and efficiency metrics. Sensitive fields (user, account, job name) are replaced with cryptographic hashes. System & Timeframe: Eagle was a 2,000-node, 8-petaflop system operated at NLR from 2019–2024. Data covers the full operational lifetime of the system. Slurm data was processed nightly; timestamps are in Mountain Time. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.eagle.job-anon.zip — Core anonymized job records (Hive-partitioned Parquet) esif.hpc.eagle.job-anon-energy-metrics.zip — Same records with additional iLO and Ganglia energy metrics datacard.md — Full dataset documentation ~13.8 million rows, 62 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct through a pipeline: Eagle Jobs API → Redpanda → StreamSets → HPCMON API → PostgreSQL. Node-level power from iLO (HP Integrated Lights-Out); GPU power from Ganglia monitoring, joined to jobs via node lists and time ranges. Preprocessing: Anonymization of name, user, and account fields via cryptographic hashing Derived columns: queue_wait, cpu_eff, max_mem_eff Simplified job state mapping (e.g., "CANCELLED BY 12345" → "CANCELLED") QoS accounting rules (buy-in, standby, or Slurm QoS value) CPU energy estimated from TDP (200W, Intel Xeon Gold 6154, 18 cores) Timezone-aware columns (_tz) sourced from LEX accounting database to correctly handle DST transitions Key Variables: Scheduling: job_id, partition, state_simple, submit_time_tz, start_time_tz, end_time_tz, queue_waitResources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, node_energy_total_watt_hours (iLO), gpu0/1_energy_total_watt_hours (Ganglia) Partitions: bigmem, bigmem-8600, bigscratch, csc, dav, ddn, debug, gpu, haswell, long, mono, short, standard Job States: CANCELLED, COMPLETED, FAILED, NODE_FAIL, OUT_OF_MEMORY, PENDING, RUNNING, TIMEOUT QoS Levels: Unknown, normal, buy-in, debug, penalty, high, standby Important Notes: Non-_tz timestamp columns may be off by one hour across DST boundaries; use _tz columns for time difference calculations Energy fields are null for jobs without monitoring coverage Job step records and raw Slurm JSONB fields are excluded from this extract Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗