Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Anonymization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Predicting the Identities of su(met-2) and met-3 in Neurospora crassa by Genome Resequencing

A significant number of classical genetic Neurospora crassa biochemical mutants remain anonymous, unassociated with a physical genome locus. By utilizing short read next-generation sequencing methods, it is possible to sequence the genomes of mutant strains rapidly and economically for the purpose of identifying genes associated with mutant phenotypes. We have taken this approach to connect genes and mutations to “methionineless” phenotypes in N. crassa.

59 BASIC BIOLOGICAL SCIENCES↗

Hourly natural gas usage surrounding NIST NEB and NWB GHG monitoring stations in 2023

This dataset includes hourly natural gas usage surrounding two National Institute of Standards and Technology (NIST) greenhouse gas (GHG) monitoring stations, Northwest Baltimore (NWB; 39.3445°N, 76.6851°W) and Northeast Baltimore (NEB; 39.3154°N, 76.5830°W), for the year 2023 in units of therms. The gas usage was provided by the local gas distribution company and includes hourly data averaged across groups of 15 or more addresses to maintain anonymity. The hourly data was averaged spatially within a radius of 1 km from each monitoring station, including all groups containing data from any address within 1km of the monitoring station. This dataset only includes gas usage from service points equipped with advanced metering infrastructure (AMI). Between the two monitoring sites, this dataset includes usage from a total of 3649 service points, which we estimate is at least 44% of the service points in the domain.

Kenion, Helen C. R. [School of Environment and Sus↗

Dronebase Photovoltaic (PV) Fleet Imagery Quantitative Evaluation (CRADA CRD-22-22941 Final Report)

Combine the Dronebase aerial imagery with corresponding sites in the NLR Photovoltaic (PV) Fleets database. By combining these two data sources in an aggregated, anonymized fashion, we can perform the following analyses: quantifying power loss due to outages caused by stuck trackers, string outages, and shading/snow, validate site metadata, including tilt and azimuth, and correlate.

14 SOLAR ENERGY↗

Distributed PV permitting survey responses

In 2019, NREL, in partnership with the Solar Energy Industries Association, developed a survey for PV installers on their experiences with delays and cancelations in the adoption process, including the impacts of permitting and interconnection. This dataset includes the anonymized responses from 147 representatives from 136 solar install companies.

14 SOLAR ENERGY↗

Packetized energy management control systems and methods of using the same

Aspects of the present disclosure include anonymous, asynchronous, and randomized control schemes for distributed energy resources (DERs). Such control schemes may include packetized energy management (PEM) control schemes for managing DERs that may provide near-optimal tracking performance under imperfect information and consumer quality of service (QoS) constraints.

Frolik, Jeffrey↗

Systems and methods for randomized, packet-based power management of conditionally-controlled loads and bi-directional distributed energy storage systems

The present disclosure provides a distributed and anonymous approach to demand response of an electricity system. The approach conceptualizes energy consumption and production of distributed-energy resources (DERs) via discrete energy packets that are coordinated by a cyber computing entity that grants or denies energy packet requests from the DERs. The approach leverages a condition of a DER, which is particularly useful for (1) thermostatically-controlled loads, (2) non-thermostatic conditionally-controlled loads, and (3) bi-directional distributed energy storage systems. In a first aspect of the present approach, each DER independently requests the authority to switch on for a fixed amount of time (i.e., packet duration). The coordinator determines whether to grant or deny each request based electric grid and/or energy or power market conditions. In a second aspect, bi-directional DERs, such as distributed-energy storage systems (DESSs) are further able to request to supply energy to the grid.

Frolik, Jeff↗

Analysis of Automated Fault Detection and Diagnosis Records as an Indicator of HVAC Fault Prevalence: Methodology and Preliminary Results

Faults in commercial buildings can cause energy waste and other performance problems such as reduced occupant comfort, reduced equipment longevity, and increased noise. However, it is currently unknown how commonly faults occur in different equipment types. A method has been developed to estimate the prevalence of faults in air handling units, air terminal units, and rooftop units. This method includes two types of data. The first is data from several automated fault detection and diagnostics (AFDD) software technologies. This type of data provides a large sample that represents a wide range of building types, geographical locations, and equipment types. It includes fault diagnoses from thousands of buildings around the United States, as well as anonymized metadata describing the building and equipment characteristics. The number of fault records is in the order of 107. However, despite the size and richness of the data sample, this data contains some degree of inaccuracy, i.e., false positive and false negative findings. Therefore, the study includes a second type of data, coming from manual inspection of buildings that have had the same AFDD methods applied to them (from the commercial AFDD offerings). Since the field tests are conducted in buildings with AFDD-generated fault prevalence data, they can be combined with the larger sample size to provide insight into the potential biases or lower sensitivity of the AFDD data. Once a library of fault prevalence data is built, it will be studied to provide further insight into the drivers of fault prevalence, for example, whether prevalence is correlated with building type, geographical location (which is tied to climate and to utility rates), building size, etc. This paper describes the methods developed for this study and illustrates them with preliminary data. It discusses some of the challenges of harmonizing disparate outputs from multiple AFDD vendors, application of a unifying fault taxonomy, and fault prevalence metrics.

Ebrahimi Fakhar, Amir↗

A Privacy-Preserving Strategy for the Trust Layer of the Energy Grid of Things Distributed Energy Resource Management System

Emergent from the shadows of the traditional grid flaws, the Smart Grid (SG) idea was born and led by government mandates toward cleaner energy production. The SG represents the next generation of electricity distribution systems that subsume recent technological innovations. It uses digital communication between its components and entities to attain more automation, self-sufficiency, and reliability. Unfortunately, this relatively new concept is not flawless; the intrinsic reliance on increased digital communication spreads open attack paths for adversaries. Therefore, finding solutions that address information exchange vulnerabilities has become imperative. The Energy Grid of Things (EGoT) is Portland State University’s (PSU’s) implementation of a Distributed Energy Resource Management System (DERMS). The EGoT DERMS requires access to customers’ information to achieve operational objectives. The system’s access to customers’ information needs to be restricted such that it does not violate customers’ privacy. Applying privacy protection models such as K-anonymity to EGoT DERMS sub-components safeguards that privacy. This thesis work proposes a strategy to ensure communication in the EGoT DERMS is privacy-preserving and secure. Specifically, it provides an approach to applying the Mondrian Algorithm to ensure data within the system excludes Personally Identifiable Information (PII) and provides means for securing the communication according to industry standards (IEEE 2030.5). Results suggest that the generalization hierarchy derived for the EGoT DERMS exhibits an Identical Generalization Hierarchy structure. Guarantees of sameness manifested in the test feeder topology would not hold in real-world scenarios. This thesis work proposes a strategy to ensure communication in the EGoT DERMS is privacy-preserving and secure. Specifically, it provides an approach to applying the Mondrian Algorithm to ensure data within the system excludes Personally Identifiable Information (PII) and provides means for securing the communication according to industry standards (IEEE 2030.5). Results suggest that the generalization hierarchy derived for the EGoT DERMS exhibits an Identical Generalization Hierarchy structure. Guarantees of sameness manifested in the test feeder topology would not hold in real-world scenarios.

Alsiad, Mohammed↗

Automated vehicle occupancy detection

Described herein are systems and methods for detecting the number of occupants in a vehicle. The detecting may be performed using a camera and a processing device. The detecting may be anonymous and the image of the interior of the vehicle is not stored on the processing device.

Moniot, Matthew Louis↗

Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART): Data Format and Preparation Guidance, Version 2.0

The Joint Office of Energy and Transportation maintains the Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART), which provides a centralized hub for submitting electric vehicle (EV) charging infrastructure data directed by the Federal Highway Administration (23 CFR 680.112) EV-ChART will provide a streamlined data submission process and an integrated set of analytic tools, connect to other data sources, and empower data sharing and access across stakeholders, including the public. Any data shared publicly will be aggregated and anonymized to stay in accordance with 23 CFR 680. This EV-ChART Data Format and Preparation Guidance provides a comprehensive overview of the data reporting requirements as authorized under 23 CFR 680.112. The guidance is intended to be used alongside the EV-ChART Data Input Template, which defines the tabular data structure that these data submissions must follow.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

Improving the LandScan USA Non-Obligate Population Estimate (NOPE)

Where do people go when they have nowhere to be? Nonobligate activities are a significant part of our social and cultural lives, but there are no existing large scale data which characterize spatial variability in population allocation for these activities. As large scale population estimates have ever-finer resolutions, gaps in our ability to estimate this population segment have an increasingly large impact on high resolution population estimates. In this paper, we demonstrate an improved method for estimating the spatial allocation of the non-obligate population - people who are not at work, school, or in another residential institution. This method builds upon on anonymized and aggregate data on visits to public places, allocating the non-obligate population proportionally to worker population while accounting for the estimated ratio of visitors to workers in public places.

Brelsford, Christa↗

A multi-level load shape clustering and disaggregation approach to characterize patterns of energy consumption behavior

This study presents representative electrical load shapes, disaggregated to the end-use level, for over 5000 customer clusters across California’s residential, commercial, industrial and agricultural sectors. We developed a novel, multi-level load shape clustering approach for residential and commercial sectors leveraging interval meter data for over 350,000 California utility customers collected as a part of the Phase 4 California Demand Response (DR) Potential Study. The clustering approach allowed us to identify typical consumption patterns and categorize customers based on their daily load shape displayed throughout the year. For example, we were able to identify customers with particular energy technologies such as electric vehicles and rooftop solar, as well as building occupancy types such as restaurants, grocery stores and even unoccupied buildings, based solely on whole-building interval data. We then combined the load shape-based clusters with other customer information including building type, climate, geographical area, total consumption and low-income status, to create a set of customer clusters based on both demographics and usage patterns. Total cluster electricity demand was then disaggregated into a wide variety of end-uses using weather normalization and other publicly available end-use load shape datasets. The resulting disaggregated cluster load shapes will be released in anonymized form as part of the Phase 4 DR Potential Study. They will have wide-ranging applications in energy research and policy analysis, including estimation of energy efficiency (EE) and DR potential on the end-use level, time-dependent valuation of EE savings, building stock modeling, and developing customer targeting strategies for EE and DR programs.

Murthy, Samanvitha↗

PV Inverter Availability from the U.S. PV Fleet

In the PV Fleet Performance Data Initiative, we partner with photovoltaic (PV) fleet owners to collect time-series PV production data, and publish aggregated, anonymized results. An assessment of system availability is conducted on 1128 systems which passed our data quality checks, and include cumulative energy meter data. Overall inverter availability is low in the first 6 months of system performance before reaching steady-state by the end of the first year. System-level aggregated data shows a median (P50) system availability of 0.99, and a lower P90 value of 0.95. A dependence on system size is also identified, with better inverter availability results for smaller PV systems. Potential causes of this effect may include the selection of inverter itself: smaller inverters 6kW-250kW showed better average availability than inverters 300kW-5MW. The elimination of string combiner boxes and lower energy impact when one particular inverter goes off-line are potential benefits of a string inverter-based PV system architecture. DNV also analyzed availability data from over 1100 operating systems and found similar trends. DNV's P50 industry guidance on expected availability has been updated to reflect the data and the following observations: utility scale systems have lower availability than DG systems, availability is lower in first year compared to subsequent years, and that actual availability is lower than expected.

fleet↗

Leveraging Hydropower Multi-Sensor Data for Inference and Age-Informed Modeling

Increased demand of operational flexibility such as faster ramp up/down in generation, and more frequent start/stops are putting hydropower plants and their associated components in unprecedented stress. Consequently, these plants are at the high risk of extended and more frequent outage to accommodate unscheduled, and unexpected maintenance. Therefore, hydropower plants are in critical need of data driven and age-informed analysis for their regular and unscheduled operation. Yet not all hydropower plants are exhaustively equipped with sensors and/or measurement streams for their respective components – demanding solutions on how to detect, identify, and locate the cause of any event from the unobservable. Idaho National Laboratory (INL) analyzed the anonymized measurements and event records from the Hydropower Research Institute (HRI) to address this issue, as part of the Water Power Technologies Office (WPTO) funded one year multi-lab project. First, we investigated how time series of multiple sensor measurements can be leveraged to identify an event “root cause” as well as to develop an inference (i.e., estimate the unobservable) problem. INL also investigated how individual hydropower components’ reaction or response times vary across the pre-event, during event, and post-event conditions – enabling the hydropower dynamic models to be age-informed. Finally, the impact of clustering multi-sensor time series on short-term vibration prediction is analyzed. INL will present key findings from these analyses and recommend next steps for stakeholder adoption.

13 HYDRO ENERGY↗

Systems and methods for randomized energy draw or supply requests

The present disclosure can provide a distributed and anonymous approach to demand response of an electricity system. The approach can conceptualize energy consumption and production of distributed-energy resources (DERs) via discrete energy packets that are coordinated by a cyber computing entity that grants or denies energy packet requests from the DERs. The approach leverages a condition of a DER, which is particularly useful for (1) thermostatically-controlled loads, (2) non-thermostatic conditionally-controlled loads, and (3) bi-directional distributed energy storage systems, among others. In a first aspect of the present approach, each DER independently requests the authority to switch on for a fixed amount of time (i.e., packet duration). The coordinator determines whether to grant or deny each request based electric grid and/or energy or power market conditions. In a second aspect, bi-directional DERs, such as distributed-energy storage systems (DESSs) are further able to request to supply energy to the grid.

Frolik, Jeff↗

EVs@Scale Next-Gen Profiles - Fleet Utilization 2024

As part of the U.S. Department of Energy’s EVs@Scale initiative, the Next-Gen Profiles (NGP) project provides a comprehensive, data-driven analysis of electric vehicle (EV) and electric vehicle supply equipment (EVSE) operations across real-world fleet deployments. This paper presents findings from the NGP’s Fleet Utilization study, which investigates operational behavior and asset usage across seventeen EV fleets and two EVSE fleets, encompassing a wide range of vehicle types and use cases. Data collected from diverse sources—varying in format and temporal resolution—are first reformatted into a unified structure. From this harmonized dataset, a suite of rigorously defined performance metrics is calculated at an hourly cadence, enabling consistent cross-comparison of charging, routing, and other key operational behaviors. Amid rapidly increasing EV adoption and growing demands for energy-efficient fleet operations, the analysis reveals clear utilization trends—including diurnal and weekly activity cycles, differences in short versus long charging session dependencies, and route-specific energy usage patterns. These findings highlight the need for tailored infrastructure strategies and the deployment of advanced energy management systems, such as Distributed Energy Resource Management Systems (DERMS) and Site Energy Management Systems (SEMS), which can optimize charging schedules and mitigate peak loads. By leveraging anonymized, harmonized datasets and standardized metrics, this study offers critical insights into fleet behavior and performance, providing a foundation to improve operational efficiency, reduce costs, and enable the scalable deployment of electrified transportation.

Wells, Landon↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗