Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Privacy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Catalyzing deep decarbonization with federated battery diagnosis and prognosis for better data management in energy storage systems

Industrial data analytics methods play a central role in improving energy storage performance and efficiency, impacting the future of electrified transportation and renewable electricity generation. However, significant challenges hinder the large-scale deployment of batteries. Conventional methods rely on centralized collection and processing of fleet-level data, leading to database size issues and privacy concerns due to potential data breaches. To enable scalable deployment of battery management systems, this article proposes a federated battery diagnosis and prognosis model, which distributes the processing of battery standard current-voltage-time-usage data in a privacy-preserving manner. Instead of transferring the raw data, this approach communicates only the locally processed parameters, thus reducing communication load and preserving data confidentiality. The federated model offers a paradigm shift in battery health management through privacy-preserving distributed methods for battery data processing and lifetime prediction, ensuring the reliable and sustainable deployment of lithium-ion batteries in a rapidly evolving world.

asset health management↗

Regression Analysis with the Directed Infusion of Data

Integrating artificial intelligence and machine learning tools into industry necessitates large-scale collaborative efforts that ensure the robust and accurate execution of downstream analytics such as time series prediction, uncertainty quantification, grid optimization, and condition monitoring. However, concerns related to data privacy pervade the nuclear industry due to the proprietary nature of its data and the possibility of data leakage. Legacy techniques such as encryption often require the explicit transmission of data to trustworthy parties, thereby inviting data leakage concerns. The ideal collaboration scenario avoids the explicit dissemination of data/code while maintaining experimental fidelity, which is currently accomplished using various techniques such as trusted execution environments, homomorphic encryption, differential privacy, and multimatrix masking. These techniques, however, often necessitate a trade-off between trust, efficiency, and utility. This article extends a previously proposed technique called the directed infusion of data (DIOD) that ensures data privacy, allows for scalable obfuscation, and combats the risk of data leakage without compromising utility. The experiments discussed in this article examine a regression-type scenario using DIOD with the goal of preserving the inferential link between two variables. Using the point-kinetics equations, regression experiments compare the performance of a model trained using the original data to that of a model trained using the obfuscated data, which produced identical results. Our claim is further strengthened by an information theoretic proof and experiment, which showed that the inferential content between variables remains the same after obfuscation, thereby avoiding the required communication of the proprietary data.

47 - OTHER INSTRUMENTATION↗

Dynamical Sketching for Enhanced Communication Efficiency in Federated Learning

Federated learning (FL) has revolutionized distributed machine learning by enabling collaborative model training without sharing local data. However, communication efficiency and privacy guarantees remain significant challenges. This paper introduces a dynamic sketching mechanism in FL, optimizing the trade-off between communication efficiency and model accuracy. By dynamically selecting the sketch matrix size, our approach adapts to the evolving characteristics of the data and the model, ensuring optimal performance across diverse scenarios. We leverage Bayesian optimization to systematically tune the sketch parameters, achieving an effective balance between resource efficiency and model performance. Experimental results on the MNIST dataset using a convolutional neural network (CNN) architecture validate the proposed method's efficiency and scalability. Our dynamic sketching approach significantly outperforms fixed-size sketching techniques, achieving higher compression ratios (up to 62x) and providing better privacy guarantees while maintaining high model accuracy. These findings highlight the robustness and versatility of our approach and make it a valuable solution for privacy-preserving, communication-efficient federated learning.

Afrose, Sharmin [ORNL]↗

Efficient Client Selection in Federated Learning

Federated Learning (FL) enables decentralized machine learning while preserving data privacy. This paper proposes a novel client selection framework that integrates differential privacy and fault tolerance. The adaptive client selection adjusts the number of clients based on performance and system constraints, with noise added to protect privacy. Evaluated on the UNSW-NB15 and ROAD datasets for network anomaly detection, the method improves accuracy by 7% and reduces training time by 25 % compared to baselines. Fault tolerance enhances robustness with minimal performance trade-offs.

Marfo, William [University of Texas at El Paso,Dep↗

OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC

Federated Learning (FL) is critical for edge and High Performance Computing (HPC) where data is not centralized and privacy is crucial. We present OmniFed, a modular framework designed around decoupling and clear separation of concerns for configuration, orchestration, communication, and training logic. Its architecture supports configuration-driven prototyping and code-level override-what-you-need customization. We also support different topologies, mixed communication protocols within a single deployment, and popular training algorithms. It also offers optional privacy mechanisms including Differential Privacy (DP), Homomorphic Encryption (HE), and Secure Aggregation (SA), as well as compression strategies. These capabilities are exposed through well-defined extension points, allowing users to customize topology and orchestration, learning logic, and privacy/compression plugins, all while preserving the integrity of the core system. We evaluate multiple models and algorithms to measure various performance metrics. By unifying topology configuration, mixed-protocol communication, and pluggable modules in one stack, OmniFed streamlines FL deployment across heterogeneous environments. Github repository is available at https://github.com/at-aaims/OmniFed.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Exploration of Factors That Influence Willingness to Consider Pooled Rideshare

Ridesharing has become an increasingly prevalent form of transportation. Although transportation network companies such as Uber and Lyft initially started as a personal rideshare service where individuals ride alone or with people they know, rideshare services have been expanded to pooled rideshare—a dynamic rideshare system where an individual rides with passengers they do not know. Despite the growth in rideshare services worldwide, the use of pooled rideshare in the U.S.A. is relatively low compared to other forms of transportation. A national U.S. survey (N = 5385) was conducted to investigate reasons why individuals are willing or unwilling to consider pooled rideshare. Exploratory and confirmatory factor analyses were performed, where the exploratory factor analysis suggests five factors, specifically,service experience,time/cost,traffic/environment,privacy, andsafety. Model fit indices of the confirmatory factor analysis verified that these five factors can represent the factors behind riders’ willingness to consider pooled rideshare. Furthermore, a binomial logistic regression was conducted to explore how the five factors influence riders’ willingness to consider pooled rideshare. The three factors that influence riders’ willingness to consider pooled rideshare wereservice experience(B = 1.05),traffic/environment(B = .38), andtime/cost(B = .26), while a lack ofprivacy(B = −1.46) can be a deterrent for pooled rideshare.Safetyis important for those who are both willing and unwilling to consider the use of pooled rideshare. Understanding these factors is important for the future of pooled rideshare services in the U.S.A.

Engineering↗

Frequency Oracle for Sensitive Data Monitoring

As data privacy issues grow, finding the best privacy preservation algorithm for each situation is increasingly essential. This research has focused on understanding the frequency oracles (FO) privacy preservation algorithms. FO conduct the frequency estimation of any value in the domain. The aim is to explore how each can be best used and recommend which one to use with which data type. We experimented with different data scenarios and federated learning settings. Results showed clear guidance on when to use a specific algorithm.

Sances, Richard↗

Indoor Occupant Counting by RF Backscattering

Building HVAC (heating, ventilation and air conditioning) consumes approximately 13% of all energy consumption in USA. Motion detectors, cameras and user programmable thermostats have been shown to be ineffective for HVAC controls to save energy, mostly due to the user concerns of comfort, reliability and privacy. A new HVAC control system based on real-time occupant counting that is fully automated, highly accurate, economically sensible and preserving privacy and aesthetics can thus bring forth a disruptive impact to this large energy sector. Our indoor occupant monitoring technology is based on the radio-frequency identification system (RFID), deployed in the room, not on the occupants. One reader with four antennas can be deployed on the ceiling or behind the ceiling panels for every thousand square feet in home, office and assisted living, with or without room partitions. The sticker-like passive tag, 10 cents each and maintenance-free, are profusely hidden on the wall or inside the furniture at arbitrary position, preserving privacy and aesthetics. The large number of tags can realize diverse observation points to accommodate arbitrary room layouts, which is impractical by other active units of camera, infrared, radar or lidar. With 20 tags, the system can reliably detect the number of occupants. For 100 tags, occupant posture and location can be known. The technology has been verified in the research labs and test buildings with very high accuracy. When the real-time occupant number can be accurately known without assuming devices on occupants or occupant motion, the building HVAC system can be automated to achieve building energy saving without sacrificing occupant comfort. According to our limited testing in a few types of building models and the simplified cost calculation, the RFID system has low overall cost in production, deployment, operation and maintenance. The signal processing algorithm based on machine learning requires very small number of training cases as most learning is transferrable for various layouts, and very low computational needs during operation, according to our testing in four different room sizes and layouts. In our preliminary estimate from HVAC saving alone, the RFID system can potentially pay for itself within 1.5 years, in addition to the other enhancement in building automation systems (BAS). Our commercialization strategy and business pitch deck focus on venturing this Cornell occupant monitoring technology into BAS and energy management markets. We have identified three broad BAS market segments of senior living, residential buildings, and office buildings. We have put together the minimum viable product characteristics for these identified segments, including analyses on total cost and competing technologies, as well as the fit and technical gaps for these market segments. We intend to bring the technology to market by licensing or partnering with existing BAS vendors. A list of potential collaborators and licensing partner candidates was assembled for different aspects of integrating our technology into a potential product that can be used with BAS.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

An Ethics-Based Review of Generative Artificial Intelligence: Assuring Responsible Use (Version 1.0)

The rapid expansion of generative artificial intelligence (GenAI) has generated excitement regarding its potential benefits and concern over its ethical implications. Governments, corporations, and standards organizations have described ethical principles to direct GenAI's development and use; however, practical guidance for implementing these principles is limited. Addressing this gap is critical, especially considering the array of risks associated with GenAI, such as legal liabilities, privacy concerns, security threats, and potential misuse. Robust policies and procedures are critical to support responsible deployment of GenAI. This report examines Pacific Northwest National Laboratory (PNNL)’s approach to promoting responsible GenAI use. Proposed initiatives include developing policies based on ethical principles, creating a governance process to review projects relative to those principles, and implementing onboarding processes for training staff. The governance framework described in this report adapts the structure and principles of Institutional Review Boards (IRBs), traditionally used in human subjects research, for GenAI ethical review, providing oversight. Ethical principles guiding responsible GenAI usage include transparency and accountability, privacy, fairness, safety, security, and validity and reliability. To operationalize these principles, we propose forming a GenAI Assurance Council (GAC) that mirrors the IRB's structure. The GAC will evaluate GenAI projects across privacy, accountability, transparency, safety, security, fairness, and validity dimensions. Complementing policy and governance is AI literacy training to support staff understanding of GenAI's ethical implications. An initial training effort for AI Incubator Chat—a GenAI tool deployed at PNNL—showed promising results, underscoring the importance of clear guidelines and user accountability. Collaborative efforts and the dissemination of best practices are also discussed. The proposed GAC model and AI literacy training provide a blueprint for establishing ethical GenAI use and governance, offering practical tools to bridge the gap between ethical principles and real-world applications. The responsible integration of GenAI at PNNL entails a multifaceted approach involving policy development, ethical governance, and AI literacy training. The positive initial feedback and collaborative opportunities position PNNL to lead by example in GenAI's responsible use, reflecting a proactive stance in addressing the ethical, legal, and societal challenges associated with this emerging technology. PNNL's systematic and ethical approach to GenAI offers a model for other institutions to emulate, promoting safe and responsible technological advancements in the AI domain.

97 MATHEMATICS AND COMPUTING↗

Exploring Object Detection and Image Classification Tasks for Niche Use Case in Naturalistic Driving Studies

Naturalistic driving studies consist of drivers using their personal vehicles and provide valuable real-world data, but privacy issues must be handled very carefully. Drivers sign a consent form when they elect to participate, but passengers do not for a variety of practical reasons. However, their privacy must still be protected. One large study includes a blurred image of the entire cabin which allows reviewers to find passengers in the vehicle; this protects the privacy but still allows a means of answering questions regarding the impact of passengers on driver behavior. A method for automatically counting the passengers would have scientific value for transportation researchers. We investigated different image analysis methods for automatically locating and counting the non-drivers including simple face detection and fine-tuned methods for image classification and a published object detection method. We also compared the image classification using convolutional neural network and vision transformer backbones. Our studies show the image classification method appears to work the best in terms of absolute performance, although we note the closed nature of our dataset and nature of the imagery makes the application somewhat niche and object detection methods also have advantages. We perform some analysis to support our conclusion.

Peruski, Ryan↗

Barriers and Benefits: Understanding Riders’ Views on Pooled Rideshare in the U.S.

This manuscript provides actionable recommendations to enhance user satisfaction and address existing barriers regarding pooled rideshare (PR) in the United States. Despite PR’s intended benefits, such as reduced traffic congestion and cost savings, its adoption remains limited. To identify these actionable items, a U.S. nationwide survey with 5385 participants explored transportation preferences, barriers, and motivators for PR use in the summer of 2021. First, two factor analyses were conducted. The first factor analysis identified the five factors associated with one’s willingness to consider PR (time/cost, traffic/environment, safety, privacy, and service experience). The second factor analysis revealed the four factors related to ways to optimize one’s PR experience (comfort/ease of use, convenience, vehicle technology/accessibility, and passenger safety). Privacy concerns, for instance, were found to reduce the likelihood of PR adoption by 77%, and convenience had the potential to increase it by 156%. A structural equation model evaluated the relationships among these nine key factors influencing PR usage to develop the Pooled Rideshare Acceptance Model (PRAM). The privacy, safety, trust service, and convenience factors each had a significant large effect (Cohen’s f 2 > 0.35) on the model. PRAM was extended using multigroup analyses to reveal the nuanced impact of 16 demographics, including gender, generation, rideshare experience, etc., highlighting the need for tailored strategies to improve PR acceptance through the Pooled Rideshare Acceptance Model Multigroup Analyses (PRAMMAs). Multiple workshops were held with diverse audiences to translate the team’s findings to date into 84 actionable recommendations, categorized across topical areas like safety, routing, driver and passenger selection, user education, etc. These findings are a foundation for a future study to determine which items resonate with different user groups. In the meantime, the actional items serve as a user-driven resource for policymakers, transportation network companies, and researchers, offering a roadmap to potential improvements to PR services to address existing concerns with the goal of increasing the usage of PR.

actionable recommendations↗

FedADMP: A Joint Anomaly Detection and Mobility Prediction Framework via Federated Learning

With the proliferation of mobile devices and smart cameras, detecting anomalies and predicting their mobility are critical for enhancing safety in ubiquitous computing systems. Due to data privacy regulations and limited communication bandwidth, it is infeasible to collect, transmit, and store all data from mobile devices at a central location. To overcome this challenge, we propose FedADMP, a federated learning based joint Anomaly Detection and Mobility Prediction framework. FedADMP adaptively splits the training process between the server and clients to reduce computation loads on clients. To protect the privacy of user data, clients in FedADMP upload only intermediate model parameters to the cloud server. We also develop a differential privacy method to prevent the cloud server and external attackers from inferring private information during the model upload procedure. Extensive experiments using real-world datasets show that FedADMP consistently outperforms existing methods.

97 MATHEMATICS AND COMPUTING↗

Understanding and Modeling Pooled Rideshare Acceptance: Influential Factors, Preferred User Experiences, and Implications

Ridesharing allows people to share a vehicle with others traveling in the same direction, which can reduce costs and traffic congestion. Pooled rideshare (PR) services, such as UberX Share and Lyft Shared, offer an economical and environmentally friendly alternative by matching passengers traveling similar routes. However, despite these benefits, PR adoption remains low due to concerns about safety, privacy, and convenience. This research explores the factors influencing PR adoption and provides recommendations to improve user acceptance. A nationwide survey of 5,385 participants across the U.S. was conducted to understand why people choose or avoid PR. The study identified five key factors influencing PR consideration: safety, service experience, privacy, traffic/environment, and time/cost. Additional research examined ways to optimize PR experiences by identifying four critical factors: comfort/ease of use, convenience, vehicle technology/accessibility, and passenger safety. To measure the impact of these factors, a statistical model called the Pooled Rideshare Acceptance Model (PRAM) was developed, providing insights into how each element influences PR adoption. Further analysis using the Pooled Rideshare Acceptance Model Multigroup Analyses (PRAMMA) revealed how demographic characteristics such as age, gender, income, and past rideshare experience shape PR perceptions. Some key findings from the multigroup analyses showed that younger users valued technological features and environmental benefits, while older users prioritized reliability and service transparency. Additionally, privacy concerns were more significant for female users, while convenience was critical for higher-income groups. These results emphasize that a 'onesize-fits-all' approach to PR service design is not effective, highlighting the need for tailored strategies to address different user segments. Further, workshops were conducted with researchers and students to translate the findings into real-world solutions. These workshops and 3 all the statistical analyses led to the development of 95 actionable recommendations. The recommendations focus on key areas such as safety, service reliability, user education, and accessibility, offering tangible improvements to PR services. The insights from this study provide valuable guidance for policymakers, transportation network companies (TNCs), and researchers aiming to make PR services safer, more accessible, and widely accepted. By addressing user concerns, PR can become a more viable transportation option, supporting sustainable urban mobility and reducing reliance on private vehicles. Additionally, these findings emphasize the importance of user-centric service design in encouraging broader PR adoption. Future research should explore evolving trends in PR preferences, technological advancements, and policy changes to ensure continued improvements. By implementing these recommendations, PR services can better align with user expectations, enhance trust in shared mobility, and contribute to a more efficient transportation ecosystem.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Differentially Private Synthesis and Sharing of Network Data Via Bayesian Exponential Random Graph Models

Abstract Network data often contain sensitive relational information. One approach to protecting sensitive information while offering flexibility for network analysis is to share synthesized networks based on the information in originally observed networks. We employ differential privacy (DP) and exponential random graph models (ERGMs) and propose the DP-ERGM method to synthesize network data. We apply DP-ERGM to two real-world networks. We then compare the utility of synthesized networks generated by DP-ERGM, the DyadWise Randomized Response (DWRR) approach, and the Synthesis through Conditional distribution of Edge given nodal Attribute (SCEA) approach. In general, the results suggest that DP-ERGM preserves the original information significantly better than two other approaches in network structural statistics and inference for ERGMs and latent space models. Furthermore, DP-ERGM satisfies node DP through modeling the global network structure with ERGM, a stronger notion of privacy than the edge DP under which DWRR and SCEA operate.

graph synthesis↗

OASIS: Offsetting Active Reconstruction Attacks in Federated Learning

Federated Learning (FL) has garnered significant attention for its potential to protect user privacy while enhancing model training efficiency. For that reason, FL has found its use in various domains, from health care to industrial engineering, especially where data cannot be easily exchanged due to sensitive information or privacy laws. However, recent research has demonstrated that FL protocols can be easily compromised by active reconstruction attacks executed by dishonest servers. These attacks involve the malicious modification of global model parameters, allowing the server to obtain a verbatim copy of users' private data by inverting their gradient updates. Tackling this class of attack remains a crucial challenge due to the strong threat model. In this paper, we propose a defense mechanism, namely OASIS, based on image augmentation that effectively counteracts active reconstruction attacks while preserving model performance. We first uncover the core principle of gradient inversion that enables these attacks and theoretically identify the main conditions by which the defense can be robust regardless of the attack strategies. We then construct our defense with image augmentation showing that it can undermine the attack principle. Comprehensive evaluations demonstrate the efficacy of the defense mechanism highlighting its feasibility as a solution.

deep neural networks↗

Securing Federated Learning Against Active Reconstruction Attacks

Federated Learning (FL) has amassed notable attention for its ability to preserve user privacy while emphasizing the retainment of model training efficiency. Due to this potential, FL has been integrated in many domains, such as healthcare, finance, law, and industrial engineering, where data cannot be easily exchanged due to sensitive information and strict privacy laws. However, current research has indicated that FL protocols are easily compromised by active data reconstruction attacks employed by actively dishonest servers. The malicious modification of global model parameters allows an actively dishonest server to obtain a direct copy of users’ private data via gradient inversion. Here, this class of attacks is highly underexplored and continues to be a major challenge due to the intense threat model. In this paper, we propose OASIS as a scalable and modality-agnostic defense based on data augmentation that counteracts active data reconstruction attacks while preserving model performance. To generalize our defense, we uncover the intuition behind gradient inversion that enables these attacks and theoretically establish the conditions by which the defense can be considered robust regardless of attack design. From this, we formulate our defense with data augmentation that illustrates its ability to undermine the attack principle. We evaluate OASIS on five real-world datasets–two image-based (ImageNet and CIFAR100) and three text-based (Wikitext, Stack Overflow, and Shakespeare)–which span diverse uses cases such as vision tasks and language modeling. Comprehensive evaluations on these datasets exhibit the efficacy of OASIS and highlight its feasibility as a solution.

97 MATHEMATICS AND COMPUTING↗

Comparative Study of Differentially Private Data Synthesis Methods

When sharing data among researchers or releasing data for public use, there is a risk of exposing sensitive information of individuals in the data set. Data synthesis is a statistical disclosure limitation technique for releasing synthetic data sets with pseudo individual records. Traditional data synthesis techniques often rely on strong assumptions of a data intruder’s behaviors and background knowledge to assess disclosure risk. Differential privacy (DP) formulates a theoretical approach for a strong and robust privacy guarantee in data release without having to model intruders’ behaviors. Efforts have been made aiming to incorporate the DP concept in the data synthesis process. Here, we examine current DIfferentially Private Data Synthesis (DIPS) techniques for releasing individual-level surrogate data for the original data, compare the techniques conceptually and evaluate the statistical utility and inferential properties of the synthetic data via each DIPS technique through extensive simulation studies. Our work sheds light on the practical feasibility and utility of the various DIPS approaches, and suggests future research directions for DIPS.

97 MATHEMATICS AND COMPUTING↗

Federated Quantum Machine Learning

Distributed training across several quantum computers could significantly improve the training time and if we could share the learned model, not the data, it could potentially improve the data privacy as the training would happen where the data is located. One of the potential schemes to achieve this property is the federated learning (FL), which consists of several clients or local nodes learning on their own data and a central node to aggregate the models collected from those local nodes. However, to the best of our knowledge, no work has been done in quantum machine learning (QML) in federation setting yet. In this work, we present the federated training on hybrid quantum-classical machine learning models although our framework could be generalized to pure quantum machine learning model. Specifically, we consider the quantum neural network (QNN) coupled with classical pre-trained convolutional model. Our distributed federated learning scheme demonstrated almost the same level of trained model accuracies and yet significantly faster distributed training. It demonstrates a promising future research direction for scaling and privacy aspects.

97 MATHEMATICS AND COMPUTING↗