Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Privacy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Quantum machine learning with differential privacy

Abstract Quantum machine learning (QML) can complement the growing trend of using learned models for a myriad of classification tasks, from image recognition to natural speech processing. There exists the potential for a quantum advantage due to the intractability of quantum operations on a classical computer. Many datasets used in machine learning are crowd sourced or contain some private information, but to the best of our knowledge, no current QML models are equipped with privacy-preserving features. This raises concerns as it is paramount that models do not expose sensitive information. Thus, privacy-preserving algorithms need to be implemented with QML. One solution is to make the machine learning algorithm differentially private, meaning the effect of a single data point on the training dataset is minimized. Differentially private machine learning models have been investigated, but differential privacy has not been thoroughly studied in the context of QML. In this study, we develop a hybrid quantum-classical model that is trained to preserve privacy using differentially private optimization algorithm. This marks the first proof-of-principle demonstration of privacy-preserving QML. The experiments demonstrate that differentially private QML can protect user-sensitive information without signficiantly diminishing model accuracy. Although the quantum model is simulated and tested on a classical computer, it demonstrates potential to be efficiently implemented on near-term quantum devices [noisy intermediate-scale quantum (NISQ)]. The approach’s success is illustrated via the classification of spatially classed two-dimensional datasets and a binary MNIST classification. This implementation of privacy-preserving QML will ensure confidentiality and accurate learning on NISQ technology.

97 MATHEMATICS AND COMPUTING↗

Privacy-Preserving Federated Learning for Science: Challenges and Research Directions

This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific AI models, in particular, foundation models~(FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy-- an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.

Kim, Kibaek [Argonne National Laboratory (ANL)]↗

Investigating Users’ Privacy Concerns of Internet of Things (IoT) Smart Devices

Although the number of smart Internet of Things (IoT) devices has grown in recent years, the public's perception of how effectively these devices secure IoT data has been questioned. Many IoT users do not have a good level of confidence in the security or privacy procedures implemented within IoT smart devices for protecting personal IoT data. Moreover, determining the level of confidence end users have in their smart devices is becoming a major challenge. In this paper, we present a study that focuses on identifying privacy concerns IoT end users have when using IoT smart devices. We investigated multiple smart devices and conducted a survey to identify users’ privacy concerns. Furthermore, we identify five IoT privacy-preserving (IoTPP) control policies that we define and employ in comparing the privacy measures implemented by various popular smart devices. Results from our study show that the over 86% of participants are very or extremely concerned about the security and privacy of their personal data when using smart IoT devices such as Google Nest Hub or Amazon Alexa. In addition, our study shows that a significant number of IoT users may not be aware that their personal data is collected, stored or shared by IoT devices.

Joy, Daniel↗

Privacy-Preserving Artificial Intelligence on Edge Devices: A Homomorphic Encryption Approach

Recent advancements in privacy-preserving artificial intelligence (AI) have paved the way for enhanced privacy in computational processes. A standing challenge, however, is the robust privacy preservation in AI algorithms, especially when integrated into edge devices and Internet-of-Thing (IoT) infrastructures. Most prevailing solutions have adopted traditional encryption methods which, though secure, often introduce significant overhead and potential dips in accuracy. In this study, we put forth an innovative approach, utilizing the CKKS encryption scheme, aiming to harmoniously balance computational efficiency with stringent data privacy. By harnessing the capabilities of Full Homomorphic Encryption (FHE) under the CKKS scheme, we ensure the preservation of privacy, successfully curbing the inherent noise traditionally linked with accuracy reductions in similar encryption-oriented solutions. Through comprehensive experiments, our approach showcased its potential as a strong contender for privacy preservation, demonstrating commendable performance across all tests, affirming that FHE is indeed viable for devices with constrained computational power and energy resources.

Khan, Muhammad Jahanzeb↗

Exploring the Utility-Privacy Trade-Off: Impacts of Semantic and Visit Types Ambiguities on Human Mobility Simulation

Humans are in perpetual movement, constantly traversing buildings, cities, waters, oceans, and countries. Mobility stands out as a major driving force shaping our modern societies. Capturing and explaining human behavior in a world of eight billion distinct mobility agendas is a complex challenge. With the rise of interconnected devices and platforms, such as smartphones, wearables, and point-of-interest data, largescale behavioral data has become more accessible, enabling rich insights into mobility patterns. However, the widespread availability of such data introduces significant ethical challenges. Detailed mobility data can inadvertently reveal sensitive personal information, including individuals' locations, habits, social interactions, and even political or religious affiliations. Beyond privacy breaches, the ethical implications of uncovering and potentially manipulating underlying behavioral patterns demand attention. Striking a balance between the utility of mobility models and the protection of individual privacy is therefore paramount. This paper explores the utility-privacy trade-offs in human mobility modeling, focusing on the impacts of introducing semantic and visit type ambiguities. By systematically examining how these ambiguities affect the fidelity of simulated trajectories and privacy risks, we provide a framework for evaluating ethical and privacy-conscious modeling practices. Our findings emphasize the need for methods that safeguard privacy without undermining the usefulness of mobility models, contributing to the responsible advancement of mobility science in alignment with ethical standards and societal expectations.

Amichi, Licia [ORNL] (ORCID:0000000177631394)↗

Federated Learning and Differential Privacy: What might AI-Enhanced co-design of microelectronics learn?

Data is a valuable commodity, and it is often dispersed over multiple entities. Sharing data or models created from the data is not simple due to concerns regarding security, privacy, ownership, and model inversion. This limitation in sharing can hinder model training and development. Federated learning can enable data or model sharing across multiple entities that control local data without having to share or exchange the data themselves. Differential privacy is a conceptual framework that brings strong mathematical guarantee for privacy protection and helps provide a quantifiable privacy guarantee to any data or models shared. The concepts of federated learning and differential privacy are introduced along with possible connections. Lastly, some open discussion topics on how federated learning and differential privacy can tied to AI-Enhanced co-design of microelectronics are highlighted.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

PAS: Privacy Algorithms in Systems

Today we face an explosion of data generation, ranging from health monitoring to national security infrastructure systems. More and more systems are connected to the Internet that collects data at regular time intervals. These systems share data and use machine learning methods for intelligent decisions, which resulted in numerous real-world applications (e.g., autonomous vehicles, recommendation systems, and heart-rate monitoring) that have benefited from it. However, these approaches are prone to identity thief and other privacy related cyber-security attacks. So, how can data privacy be protected efficiently in these scenarios? More dedicated efforts are needed to propose the integration of privacy techniques into existing systems and develop more advanced privacy techniques to address the complex challenges of multi-system connectivity and data fusion. Therefore, we have introduced Privacy Algorithms in Systems (PAS) at CIKM which provides a venue to gather academic researchers and industry researchers/practitioners to present their research in an effort to advance the frontier of this critical direction of privacy algorithms in systems.

Kotevska, Olivera↗

A review of privacy in energy applications

As the distribution system continues to experience an increase in distributed energy resource (DER) and electric vehicle (EV) penetration, so does the need for new solutions that can help grid operators manage and leverage their capabilities. This will undoubtedly lead to new operational schemes and business opportunities that will transform the traditional consumer into a prosumer who will be more actively engaged in grid operations. Although the field is still under active development, many of the potential use cases presented in literature or industry are built upon edge computing, two-way communications, and other innovative computational constructs to attain their goals. However, at their core, many use cases assume a great level of data access to aid with the decision-making process, an assumption that may need to be revised to ensure fair and equitable operational processes are maintained. This may be particularly true as edge resources are predicted to participate in retail-side, many-to-many, or peer-to-peer markets and thus may lead to financial impacts if data access considerations are ignored. The need to revise data access mechanisms can be further justified by the introduction of new participants into the operational process, who do not have the same level of trust, nor the incentives to focus on energy delivery as their primary objective. At the same time, more consumers are becoming aware of their own data, and the potential impacts of its abuse. To help solution developers better understand these risks, this report has been developed to offer an initial introduction to the topic of privacy. This is achieved by 1) Highlighting the need for privacy-aware solutions; 2) Encouraging system designers to be inquisitive about the status quo; 3) Documenting the existing threat space; 4) Presenting and evaluating tools that may be helpful towards enabling better privacy postures; and 5) Making recommendations to encourage the adoption of better practices. From a technical perspective, the report focuses on evaluating two potential techniques by applying them to the Transactive Energy Space. Based on the obtained results, it can be established that differential privacy (DP) methods may have limited applicability when highly correlated, time-series data records need to be protected. However, DP may be a powerful tool when it is used to aggregate and analyze mid-size and large-size data sets in a more traditional statistical environment. The second tool under evaluation is threshold cryptography, which can guarantee complete secrecy (and thus privacy) but requires the establishment of key management procedures and dedicated communication channels for key coordination. Therefore, due to its increased computational overhead, the use of threshold cryptography must be weighted using a cost/benefit analysis on a per-application basis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

Analyzing Data Privacy for Edge Systems

Internet-of-Things (IoT)-based streaming applications are all around us. Currently, we are transitioning from IoT processing being performed on the cloud to the edge. While moving to the edge provides significant networking efficiency benefits, IoT edge computing creates significant data privacy concerns. We propose a methodology that can successfully privacy protect the continual data streams generated by sensors on the edge device. We implement local differential privacy on streaming data and incorporate Bayesian inference and Gaussian process to evaluate the privacy policy. We demonstrate our methodology on a real-world smart meter testbed and identify the optimal privacy protection settings.

Kotevska, Olivera↗

Implications of privacy needs and interpersonal distancing mechanisms for space station design

Isolation, confinement, and the characteristics of microgravity will accentuate the need for privacy in the proposed NASA space station, yet limit the mechanism available for achieving it. This study proposes a quantitative model for understanding privacy, interpersonal distancing, and performance, and discusses the practical implications for Space Station design. A review of the relevant literature provided the basis for a database, definitions of physical and psychological distancing, loneliness, and crowding, and a quantitative model of situational privacy. The model defines situational privacy (the match between environment and task), and focuses on interpersonal contact along visual, auditory, olfactory, and tactile dimensions. It involves summing across pairs of crew members, contact dimensions, and time, yet also permits separate analyses of subsets of crew members and contact dimensions. The study concludes that performance will benefit when the type and level of contact afforded by the environment align with that required by the task. The key to achieving this is to design a flexible, definable, and redefinable interior environment that provides occupants with a wide array of options to meet their needs for solitude, limited social interaction, and open group activity. The report presents 49 recommendations in five categories to promote a wide range of privacy options despite the space station's volumetric limitations.

Harrison, Albert A.↗

A Privacy-Preserving Strategy for the Trust Layer of the Energy Grid of Things Distributed Energy Resource Management System

Emergent from the shadows of the traditional grid flaws, the Smart Grid (SG) idea was born and led by government mandates toward cleaner energy production. The SG represents the next generation of electricity distribution systems that subsume recent technological innovations. It uses digital communication between its components and entities to attain more automation, self-sufficiency, and reliability. Unfortunately, this relatively new concept is not flawless; the intrinsic reliance on increased digital communication spreads open attack paths for adversaries. Therefore, finding solutions that address information exchange vulnerabilities has become imperative. The Energy Grid of Things (EGoT) is Portland State University’s (PSU’s) implementation of a Distributed Energy Resource Management System (DERMS). The EGoT DERMS requires access to customers’ information to achieve operational objectives. The system’s access to customers’ information needs to be restricted such that it does not violate customers’ privacy. Applying privacy protection models such as K-anonymity to EGoT DERMS sub-components safeguards that privacy. This thesis work proposes a strategy to ensure communication in the EGoT DERMS is privacy-preserving and secure. Specifically, it provides an approach to applying the Mondrian Algorithm to ensure data within the system excludes Personally Identifiable Information (PII) and provides means for securing the communication according to industry standards (IEEE 2030.5). Results suggest that the generalization hierarchy derived for the EGoT DERMS exhibits an Identical Generalization Hierarchy structure. Guarantees of sameness manifested in the test feeder topology would not hold in real-world scenarios. This thesis work proposes a strategy to ensure communication in the EGoT DERMS is privacy-preserving and secure. Specifically, it provides an approach to applying the Mondrian Algorithm to ensure data within the system excludes Personally Identifiable Information (PII) and provides means for securing the communication according to industry standards (IEEE 2030.5). Results suggest that the generalization hierarchy derived for the EGoT DERMS exhibits an Identical Generalization Hierarchy structure. Guarantees of sameness manifested in the test feeder topology would not hold in real-world scenarios.

Alsiad, Mohammed↗

Privacy Amplification for Episodic Training Methods

It has been shown that differential privacy bounds improve when subsampling within a randomized mechanism. Episodic training, utilized in many standard machine learning techniques, uses a multistage subsampling procedure which has not been previously analyzed for privacy bound amplification. In this paper, we focus on improving the calculation of privacy bounds in episodic training by thoroughly analyzing privacy amplification due to subsampling with a multi-stage subsampling procedure. The newly developed bound can be incorporated into existing privacy accounting methods.

Tombs, Vandy↗

Privacy by Design in Distributed Edge Systems: Innovating Secure Workflows for Smart Cities

The proliferation of distributed edge systems, such as those in smart cities, healthcare, and industrial IoT, offers unprecedented opportunities for data processing closer to its source, thereby reducing latency and enhancing efficiency. However, these systems also present significant privacy challenges due to the handling of sensitive data from multiple sources. This article explores the critical need for designing privacy-preserving workflows in distributed edge systems to ensure data security while maximizing the potential of edge computing. By examining the challenges, technological advancements, and potential of privacy-by-design approaches, we highlight the importance of integrating advanced privacy-preserving techniques like federated learning, differential privacy, homomorphic encryption, secure multi-party computation, and zero-knowledge proofs. These innovations are crucial for enhancing data security, regulatory compliance, and public trust in smart city applications, ultimately leading to safer and more efficient urban environments.

Kotevska, Olivera↗

Privacy-preserving federated learning: Application to behind-the-meter solar photovoltaic generation forecasting

Here, the growing usage of decentralized renewable energy sources has made accurate estimation of their aggregated generation crucial for maintaining grid flexibility and reliability. However, the majority of distributed photovoltaic (PV) systems are behind-the-meter (BTM) and invisible to utilities, leading to three challenges in obtaining an accurate forecast of their aggregated output. Firstly, traditional centralized prediction algorithms used in previous studies may not be appropriate due to privacy concerns. There is therefore a need for decentralized forecasting methods, such as federated learning (FL), to protect privacy. Secondly, there has been no comparison between localized, centralized, and decentralized forecasting methods for BTM PV production, and the trade-off between prediction accuracy and privacy has not been explored. Lastly, the computational time of data-driven prediction algorithms has not been examined. This article presents a FL power forecasting method for PVs, which uses federated learning as a decentralized collaborative modeling approach to train a single model on data from multiple BTM sites. The machine learning network used to design this FL-based BTM PV forecasting model is a multi-layered perceptron, which ensures privacy and security of the data. Comparing the suggested FL forecasting model to non-private centralized and entirely private localized models revealed that it has a high level of accuracy, with an RMSE that is 18.17% lower than localized models and 9.9% higher than centralized models.

14 SOLAR ENERGY↗

Distributed Energy Resource Management Systems: Preserving Customer Privacy through K-Anonymity

The smart grid represents the next generation of electricity distribution systems that utilizes recent technological innovations. It uses digital communication between its components and entities to attain more automation, self-sufficiency, and reliability. One of the many concerns in smart grid digital communication discussions is the possibility of violating customers’ privacy. Violating customers’ privacy imposes a significant barrier as smart grid desirable attributes are tightly tied to customers’ participation. Employing privacy models can address concerns regarding information privacy in smart grid digital communication. In this work, we provide an approach to utilizing K-anonymity to ensure data within the system excludes Personally Identifiable Information. Results suggest that a dynamically generated generalization hierarchy minimizes information loss incurred by the anonymization process.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Privacy-Preserving Robust Consensus for Distributed Microgrid Control Applications

Consensus-based distributed control has been proposed for coordinating distributed energy resources (DERs) in microgrids (MGs). As one key component, distributed average observers are used to estimate the average of a group of reference signals (e.g., voltage, current, or power). State-of-the-art distributed average observers could lead to loss of privacy due to information exchange on the communication channels. The DERs' reference signals, which contain private information, could be inferred by an eavesdropper. In this article, a privacy-preserving distributed average observer is proposed that is based on robust consensus and uses the state decomposition method to preserve privacy. Compared to the existing methods, the proposed observer does not require the knowledge of the reference signal's derivative and gives accurate and smooth estimation, and is thus applicable for MG distributed control applications. A detailed analysis regarding the convergence and privacy properties of the proposed observer is presented. Here, the proposed observer is implemented on hardware controllers and validated in the context of distributed MG control applications through hardware-in-the-loop (HIL) tests.

24 POWER TRANSMISSION AND DISTRIBUTION↗