Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Privacy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗

Efforts to enhance reproducibility in a human performance research project

Background: Ensuring the validity of results from funded programs is a critical concern for agencies that sponsor biological research. In recent years, the open science movement has sought to promote reproducibility by encouraging sharing not only of finished manuscripts but also of data and code supporting their findings. While these innovations have lent support to third-party efforts to replicate calculations underlying key results in the scientific literature, fields of inquiry where privacy considerations or other sensitivities preclude the broad distribution of raw data or analysis may require a more targeted approach to promote the quality of research output. Methods: We describe efforts oriented toward this goal that were implemented in one human performance research program, Measuring Biological Aptitude, organized by the Defense Advanced Research Project Agency's Biological Technologies Office. Our team implemented a four-pronged independent verification and validation (IV&V) strategy including 1) a centralized data storage and exchange platform, 2) quality assurance and quality control (QA/QC) of data collection, 3) test and evaluation of performer models, and 4) an archival software and data repository. Results: Our IV&V plan was carried out with assistance from both the funding agency and participating teams of researchers. QA/QC of data acquisition aided in process improvement and the flagging of experimental errors. Holdout validation set tests provided an independent gauge of model performance. Conclusions: In circumstances that do not support a fully open approach to scientific criticism, standing up independent teams to cross-check and validate the results generated by primary investigators can be an important tool to promote reproducibility of results.

59 BASIC BIOLOGICAL SCIENCES↗

Considerations for Introducing Artificial Intelligence into Nuclear Power Plants

Advanced computational tools and techniques such as artificial intelligence and machine learning (AI/ML) can transform the nuclear power industry. This is necessary given that the economic viability of the existing fleet is in jeopardy and its labor-centric approach to operations and maintenance. Currently, AI/ML research is being undertaken for reactor system design and analysis including fault and accident prognosis, nuclear risk analysis such as plant safety and security evaluation, and plant operations and maintenance including predictive maintenance. Applications include both existing and advanced reactor technologies with the aim of improving operational and business efficiencies. Most every aspect of the organization can benefit, from instrumentation and control, to work planning, to human-machine interactions and business management. AI/ML in nuclear can simplify complex problems and produce more effective decision-making. Nonetheless, careful consideration must be given to the implementation of an AI/ML initiative. The aims of this research are to 1) review barriers to AI/ML adoption within the nuclear power industry, and 2) suggest potential solutions. These barriers are organized along five distinct categories (Figure 1) that are interconnected. The first are historical barriers that track the industry’s development over the decades including worldwide nuclear events that shaped public perceptions. The resulting federal scrutiny and intense safety culture that emerged are discussed. Technical barriers to AI/ML adoption are considerable, and include data privacy concerns, data governance, and the current lack of AI/ML expert knowledge at the plants. The main business case barrier remains cost, but an absence of an industry-wide vision and wide-scale adoption also produces reluctance. Stakeholder readiness is reviewed with special attention given to regulatory readiness. The 5-year strategic plan for AI readiness recently published by the U.S. Nuclear Regulatory Commission is highlighted. Last, adoption barriers at the user level are addressed including the importance of user experience and explainable AI. The AI adoption barriers described here are inter-related and ideally should be addressed in a holistic fashion.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dataset 3: A National Dataset on Actionable Items in Improving Pooled Rideshare, 2025.

Dataset 3: A National Dataset on Actionable Items in Improving Pooled Rideshare.” 2025. Dataset Description: Pooled Rideshare Acceptance Survey - Phase 3 (2025, N = 8,296). This dataset represents the third and final phase of a national survey aimed at understanding user acceptance and preferences related to pooled rideshare (PR) services in the United States. Building on insights from earlier phases, this phase expands both the sample size and the depth of analysis to support policymaking, transportation planning, and service design for sustainable mobility systems. The Phase 3 survey was administered online to a nationally representative sample of 8,296 U.S. adults. The sample includes a wide range of demographics. The survey retained core questions from previous phases while introducing 77 detailed service features (actionable items) to evaluate potential improvements to PR offerings. Each feature was designed to assess whether a specific improvement, such as enhanced safety measures, real-time ride tracking, or user training would increase participants’ willingness to adopt PR services. In addition, behavioral predictors, current rideshare habits, environmental attitudes, and perceived barriers (e.g., safety, privacy, and comfort) were captured. - Phase_3_Final - The dataset contains rows corresponding to individual respondents and columns representing survey items, demographic characteristics, and response values. The data is available in both .CSV and .SAV formats. - Phase_3_Final_MapFile - Accompanying this dataset is a data dictionary explaining each variable, value range, and coding schema. An .XLSX format of the full survey instrument is included to support interpretation and reuse of the dataset.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FEDERATED LEARNING ON STOCHASTIC NEURAL NETWORKS

Federated learning is a machine learning paradigm that leverages edge computing on client devices to optimize models while maintaining user privacy by ensuring that local data remain on the device. However, since all data are collected by clients, federated learning is susceptible to latent noise in local datasets. Factors such as limited measurement capabilities or human errors may introduce inaccuracies in client data. To address this challenge, we propose the use of a stochastic neural network as the local model within the federated learning framework. Stochastic neural networks not only facilitate the estimation of the true underlying states of the data but also enable the quantification of latent noise. We refer to our federated learning approach, which incorporates stochastic neural networks as local models, as federated stochastic neural networks. In this work we will present numerical experiments demonstrating the performance and effectiveness of our method, particularly in handling nonindependent and identically distributed data.

97 MATHEMATICS AND COMPUTING↗

The GABLE Report: Garbled Autonomous Bots Leveraging Ethereum

Simple but mission-critical internet-based applications that require extremely high reliability and availability could potentially benefit from running on robust public programmable blockchain platforms such as Ethereum. Unfortunately, program code running on such blockchains is ordinarily publicly viewable, rendering these platforms unsuitable for applications requiring strict privacy of application code, data, and results. However, might it be possible to encode an application's business logic and data for these platforms in such a way that it becomes impossible for unauthorized parties to infer any meaningful information whatsoever about the semantics of the data, and the operations being performed on that data? In this report, we describe GABLE (Garbled Autonomous Bots Leveraging Ethereum), a system concept developed at Sandia that achieves this security goal in a limited, but still useful range of circumstances. GABLE, uses simple but effective algorithms to permit secure private execution of garbled state machines (and more efficient garbled circuits) on public computing resources. We give an example working implementation for garbled state machines, written using the Python and Solidity programming languages, and outline how our methods can be extended to support a more powerful garbled universal circuit model of computation. The capability embodied by the GABLE, system has significant potential applications, a few of which we discuss in this report.

97 MATHEMATICS AND COMPUTING↗

Elucidating and predicting the dynamic evolution of water and land systems due to natural and energy-related forcings

Focal Area(s): 3. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI; & 1. Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: Interactions between water, land, and energy systems are complex and occur on a variety of scales, ranging from local to basinal to regional. Accurately predicting the behavior of ground water and surface water systems for 5-10 years and beyond requires an understanding of the current system and the ability to model both the natural system at scale and human-induced forcings related to energy and other activities. Artificial intelligence and machine learning (AI/ML) combined with modern compilation and integration efforts for U.S. groundwater and surface water systems present potential solutions to bolstering detailed physics-based models of these systems. Big data tied with ML and physics-based modeling can drive breakthroughs in understanding the earth system, but research is often impeded by data access (e.g., privacy issues), quality, formats, gaps, multi-source, multi-scale, integration, and spatiotemporal challenges. Effective integration of real data and simulated (synthetic) data that fill gaps is critical. Overcoming these complex data and model integration challenges will enable a transformational approach to acquiring enhanced understanding of environmental systems.

54 ENVIRONMENTAL SCIENCES↗

Reimagining Codesign for Advanced Scientific Computing: Report for the ASCR Workshop on Reimagining Codesign

In March 2021, the U.S. Department of Energy’s Advanced Scientific Computing Research program convened the Workshop on Reimagining Codesign. The workshop, also known as ReCoDe, was organized around discussions on eight topic areas: (1) codesign for traditional high-performance computing workloads; (2) codesign of memory/storage systems; (3) codesign of machine learning, neuromorphic, quantum, and other non-von Neumann accelerators; (4) codesign for edge computing and processing at experimental instruments; (5) codesign for security and privacy; (6) hardware design tools and open-source hardware for high-productivity codesign; (7) tools, software stack, and programming languages for high-productivity codesign; and (8) quantitative tools and data collection for modeling and simulation for codesign. The panels identified four Priority Research Directions from these deliberations: (1) breakthrough computing capabilities with targeted heterogeneity and rapid design; (2) software and applications that embrace radical architecture diversity; (3) engineered security and integrity, from transistors to applications; and (4) design with data-rich processes.

97 MATHEMATICS AND COMPUTING↗

Sketching Algorithms in Distributed Systems

In this position paper, we discuss exciting recent advancements in sketching algorithms applied to distributed systems. That is, we look at randomized algorithms that simultaneously reduce the data dimensionality, offer potential privacy benefits, while maintaining verifiably high levels of algorithm accuracy and performance in multi-node computational setups. We look at next steps and discuss the applicability to real systems.

97 MATHEMATICS AND COMPUTING↗

Adapting Secure MultiParty Computation to Support Machine Learning in Radio Frequency Sensor Networks

In this project we developed and validated algorithms for privacy-preserving linear regression using a new variant of Secure Multiparty Computation (MPC) we call "Hybrid MPC" (hMPC). Our variant is intended to support low-power, unreliable networks of sensors with low-communication, fault-tolerant algorithms. In hMPC we do not share training data, even via secret sharing. Thus, agents are responsible for protecting their own local data. Only the machine learning (ML) model is protected with information-theoretic security guarantees against honest-but-curious agents. There are three primary advantages to this approach: (1) after setup, hMPC supports a communication-efficient matrix multiplication primitive, (2) organizations prevented by policy or technology from sharing any of their data can participate as agents in hMPC, and (3) large numbers of low-power agents can participate in hMPC. We have also created an open-source software library named "Cicada" to support hMPC applications with fault-tolerance. The fault-tolerance is important in our applications because the agents are vulnerable to failure or capture. We have demonstrated this capability at Sandia's Autonomy New Mexico laboratory through a simple machine-learning exercise with Raspberry Pi devices capturing and classifying images while flying on four drones.

42 ENGINEERING↗

Homomorphic Encryption for Machine Learning and Artificial Intelligence Applications

Third-party and expert analysis is a cost-effective solution for solving specialized problems or processing large datasets related to reactor structural health monitoring and nondestructive evaluation. However, when handling proprietary information, third-party and expert analysts pose a privacy risk. To address this challenge, Homomorphic Encryption (HE) permits arithmetic operations on encrypted data without exposing the underlying data. Implementations of Machine Learning (ML) and Artificial Intelligence (AI) algorithms using HE greatly enhances the capabilities of third-party analysts while maintaining a low security risk. This paper details current success in applying Principal Component Analysis (PCA) and Fully Connected Neural Networks (NN) using the Microsoft SEAL implementation of the popular CKKS Fully Homomorphic Encryption (FHE) algorithm. The MNIST Handwritten Dataset is analyzed as a proof-of-concept demonstration of the implementations.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Adaptive Neurocontrol for Grid-Following Inverters

The new generation of power systems involve a large number of autonomous operating units with the flow of data being restricted by privacy concerns or infrastructure hurdles. The systems are also time-varying so the control should adjust in a timely manner to avoid failures that usually have very negative economical implications. Adaptive neurocontrol which takes elements from adaptive control (great for time-varying problems) and model identification (could be in the form of neural network) by using local available data is a compelling tool. In this work, we summarize analytical results for adaptive neurocontrol, followed with numerical verification of how the method could work for future grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Intelligent Energy Optimizer for Residential Buildings

Demand-side management in the buildings is essential for meeting grid flexibility needs in a highly renewable energy scenario. Appliance load monitoring helps decision making for demand-side management by providing the information on operation status/power consumption from different appliances in the buildings. Nonintrusive load monitoring (NILM) is an attractive option for appliance load monitoring using because it has lower cost for sensors and helps mitigate privacy concerns. In this study, the team used an event detection technique followed by two different methods for event classification. The results from k-means clustering showed that the events from a single appliance are often distributed in multiple clusters. Thus, the unsupervised method of NILM using k-means clustering used in this study was not very suitable for load disaggregation. The results from NILM showed that the F1 score for event classification was 0.77 for a heat pump water heater and very low for other appliances using the rule-based classification.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Database to Enable Facial Analysis for Driving Studies (DEFADS)

Naturalistic Driving Studies (NDS) collect and utilize data on drivers in real-world environments in instrumented vehicles. A common problem with such studies is driver privacy. In this work we collected a dataset of 77 human subjects performing scripted driving-related activities. We used three camera systems for the collection, including two high-resolution webcam devices as well as a third system from an actual NDS (the Second Strategic Highway Research Project, or SHRP2). This report covers the data collection process and summarizes the dataset, which will be made publicly available to researchers under a data usage license.

97 MATHEMATICS AND COMPUTING↗

Blockchain for Fault-Tolerant Grid Operations

Radial topology and vast geographic coverage make distribution systems prone to widespread power outages upon the failure of a single (or multiple) upstream component. Fault-handling algorithms depend heavily on correct estimations of the system’s state to effectively isolate the affected area and reduce the number of affected customers while maintaining operational safety. The work described here leverages the core features of distributed, consensus-based decision-making processes and the immutability of blockchain, and demonstrates their value in improving fault-tolerant grid operations. In this work, blockchain was used to create a trusted data-sharing platform that enables independent actors to reconstruct the system state; this enables distributed resources to make intelligent decisions with limited knowledge. Although the process requires data sharing, its algorithms have been designed to limit the amount of private information that is exchanged, which helps preserve business-sensitive data and maintain customer privacy. In addition, by reducing the information that must be shared, the communication requirements are also reduced; (however, an in-depth analysis of the communication requirements is beyond the scope of this project). The proposed use cases are intended to represent a foundational basis for third parties to develop functional solutions that can eventually be deployed in the field. To further provide guidance, the envisioned use cases have incorporated design requirements that consider the blockchain characteristics and a need to limit information from surrounding resources, which preserve the assumption and the possibility that such resources could belong to different entities. This report presents a detailed design of the three use cases with the tools needed to enable the analysis being tested. The implemented gross error detection method can detect mismatches when the error exceeds 3.8 times the sensor’s rated accuracy. Detection of the circuit breaker state successfully identified the correct states across all simulation tests. A distribution-system power-flow solution in the simulator OpenDSS generally possesses a convergency tolerance of 0.01% on the voltage magnitude. The evaluation of possible reconnection using voltage magnitude—preserving the data ownership—has a voltage magnitude difference smaller than 0.001% from the OpenDSS result. The results preserving data ownership have a difference within the expected power flow tolerance with full knowledge of the system, which surpasses expectations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Domestic Extremism: Countering the Threat Posed to Critical Assets

Domestic extremism has been a growing concern in the United States in recent months, as illustrated in multiple bulletins from the Department of Homeland Security (DHS) warning law enforcement partners of the heightened threat. As concerns about these actors grows, it is important that facilities in the U.S. and internationally that protect critical assets, such as sensitive information, hazardous materials, or critical infrastructure, have effective methods in place to secure those assets. DE has challenged security systems through the threat of insider attack and violence, creating a new threat to be countered in the Office of Radiological Security’s radiological source security mission. In this effort, we used a literature review and focus group discussions with experts in critical asset security and extremism to understand the nature of the domestic extremist threat, to identify best practices in securing assets, recognize potential gaps in security measures to be corrected, and recommend actions and next steps. Twenty-two subject matter experts participated in a series of five focus group sessions. Questions focused on definitions of domestic extremism, potential changes in the threat, best practices in securing facilities, assets, and personnel, and any perceived gaps. Upon completion of the focus groups, notes were analyzed thematically to identify any recurring patterns in the results. In addition, a review of academic, industry, and government literature was conducted to understand the threat, describe the process of radicalization to extremism, and to identify empirically informed practices in prevention and response. Results of this project demonstrated that further work is needed to define domestic extremism in law, regulation, and policy, to help the U.S. develop a consistent response to the threat within organizations. This is especially important, as SMEs emphasized the need for early intervention in prevention efforts, noting that organizations need clear guidance on when and how to intervene. In addition, the need for social media monitoring was discussed, although challenges remain to do so with appropriate respect for privacy and civil liberties concerns.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Open Source Evaluation Framework for Solar Forecasting

The Solar Forecast Arbiter is an open-source evaluation framework for solar forecasting. The framework enables evaluations of solar irradiance, solar power, and net-load forecasts that are impartial, repeatable and auditable. The Solar Forecast Arbiter addresses stakeholder-informed use cases including evaluation of forecast skill, comparisons to reference data sets, private forecast trials, and evaluation of probabilistic forecast skill. The framework includes a data validation toolkit, reference data sources, data privacy protocols, and benchmark forecast capabilities for intra-hour and day ahead forecast horizons. Reports and metrics communicate the relative merits of the test and benchmark forecasts. The reports are created from standardized templates and include graphics for qualitatively evaluating deterministic and probabilistic forecasts and standard metrics for quantitatively evaluating forecasts. The Solar Forecast Arbiter is designed to support all solar forecasting stakeholders, including Solar Forecasting 2 Topic Area 2 and Topic Area 3 teams.

14 SOLAR ENERGY↗

Inventory of Public Key Cryptography in US Electric Vehicle Charging

Electric vehicles (EVs) and charging infrastructure are networked systems, which employ high-level communications in support of charging and grid service decisions. Public key cryptography (PKC) underlies much of the security and privacy protections of the information exchange. We are entering a new epoch where quantum computing threats must be seriously considered. A sufficiently large quantum computer, so named Cryptographically Relevant Quantum Computer (QRQC), will be able to perform the mathematical operations to efficiently attack the underpinnings of traditional PKC, thus jeopardizing the digital foundations for trust, communications security, and data security. Estimates suggest a QRQC can break public key encryption and digital signatures in the manner of tens to hundreds of hours, compared to traditional computing that would demand more than 10 18 years in a brute force-style attack. A consensus belief of quantum theorists, quantum experimenters, and cryptographers suggest that the quantum threat will be likely realized in the next twenty years. To address the threat, post-quantum cryptography, which is cryptosystems that are designed to be secure against both traditional and quantum computing threats, must be adopted. Migration from traditional PKC to quantum-resilient cryptography is a global undertaking and likely represents the largest transition in computing history. The nascent state of EV public key infrastructure, combined with limited adoption of the vehicle secure charging features, presents an opportunity to establish a preference for quantum-resistant cryptography as a step on the migration path. Delays will stunt the efforts as rapidly accelerating EVs sales and huge infrastructure investments will create large growing bases of long-lived vehicles and infrastructure. Migration preparations can commence while NIST continues the process to standardize post-quantum cryptography (PQC), which are quantum-resilient algorithms designed to be secure against traditional and quantum computing threats. The first step in preparing EV charging is to identify the presence of traditional public key cryptography algorithms and applications. With this objective in mind, this report is intended to advise the vehicle manufacturers, charging station manufacturers, charging station operators, charge network providers and other EV charging stakeholders with information on traditional PKC application and the potential risks when PKC becomes insecure. This report, the first in a series of reports discussing the topics existing at the confluence of post-quantum cryptography adoption and EV charging, identifies traditional public key applications employed and identifies potential consequences of leaving EV charging infrastructure vulnerable to quantum computing. The focus remains squarely on the of EV charging and infrastructure with respect to PKC and is believed by the authors to complement the NIST SP 1800-38 Migration to Post-Quantum Cryptography. While the report is centered on infrastructure, there are implications to vehicles.

33 ADVANCED PROPULSION SYSTEMS↗