Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data security”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Transportation Secure Data Center: Frequently Asked Questions for Data Owners/Contributors

The Transportation Secure Data Center is a centralized repository for detailed transportation data from travel and transit surveys and studies conducted across the nation. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. Hundreds of datasets from surveys and studies of household travel and transit passenger travel are archived in the TSDC, including surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Detailed data from travel surveys and studies are extremely valuable for research purposes. However, the fine-grained information they contain could potentially be misused to identify individual travelers, so access to these data should only be granted with safeguards in place to protect participant privacy. The TSDC was created to address this challenge and to relieve public agencies from the burden of archiving their data and responding to data requests.

33 ADVANCED PROPULSION SYSTEMS

TSDC: Transportation Secure Data Center: Real-World Data for Planning, Modeling, and Analysis

The Transportation Secure Data Center is a centralized repository for high-resolution transportation data from hundreds of travel and transit surveys and studies. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. It houses surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Meanwhile, the Livewire Data Platform empowers research, industry, and academic partners to easily and securely preserve, maintain, share, discover, and gain access to transportation and mobility data. Livewire accommodates a range of datasets, including behavioral, experimental, model, analytical, and raw data at the vehicle, traveler, and system levels. Datasets support mobility research and planning spanning urban science, connected and automated vehicles, fueling and charging infrastructure, mobility decision science, multimodal transportation, vehicle efficiency, and more.

33 ADVANCED PROPULSION SYSTEMS

Privacy by Design in Distributed Edge Systems: Innovating Secure Workflows for Smart Cities

The proliferation of distributed edge systems, such as those in smart cities, healthcare, and industrial IoT, offers unprecedented opportunities for data processing closer to its source, thereby reducing latency and enhancing efficiency. However, these systems also present significant privacy challenges due to the handling of sensitive data from multiple sources. This article explores the critical need for designing privacy-preserving workflows in distributed edge systems to ensure data security while maximizing the potential of edge computing. By examining the challenges, technological advancements, and potential of privacy-by-design approaches, we highlight the importance of integrating advanced privacy-preserving techniques like federated learning, differential privacy, homomorphic encryption, secure multi-party computation, and zero-knowledge proofs. These innovations are crucial for enhancing data security, regulatory compliance, and public trust in smart city applications, ultimately leading to safer and more efficient urban environments.

Kotevska, Olivera

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,

Secure and Privacy Aware Data Sharing Approach for Smart Electric Vehicles

The integration of smart electric vehicles (SEVs) into smart cities marks a significant step toward creating efficient, sustainable, and connected urban spaces. However, secure and private data sharing is a major challenge as SEVs connect with smart city systems. The interaction between SEVs and consumer electronic devices (CEDs) raises serious concerns about data security and privacy. Here, to address these challenges, this article presents how blockchain technology and federated learning (FL) can address these issues. The proposed approach provides a secure and privacy-aware framework for data exchange between SEVs and CEDs in smart cities. The experiment results demonstrate the effectiveness of the proposed framework for secure data sharing and maintaining system reliability in smart city environments. It also enables trust and promotes the widespread adoption of interconnected urban technologies.

Das, Debashis [Meharry Medical College, Nashville,

The CanBikeCO Full Pilot: Long-Term Results and Analysis From an E-Bike Program in Colorado, USA

Personal micromobility devices like bicycles, e-bikes, and scooters are low- or zero-energy alternatives to single-occupancy vehicles. However, a lack of data has led to a dearth of data-driven research on personally owned e-bike usage. We present longitudinal findings from the CanBikeCO program, focused on e-bike adoption and use across demographics, trip characteristics, and geographies in the state of Colorado. CanBikeCO recorded travel survey data from low-income individuals provided with personal e-bikes by the Colorado Energy Office in six communities across Colorado from July 2021 to December 2022. The data were collected using a custom instance of the National Renewable Energy Laboratory OpenPATH platform, which combines passive data collection with semantic information such as trip mode and purpose labels. To our knowledge, there are no prior travel survey data on personally owned e-bikes with this range and scope. Insights from this unique dataset include: (i) work trips were 17% more likely than average trips to be taken on an e-bike, (ii) e-bikes were most often reported to replace cars (34% of e-bike trips) and other personal micromobility devices (22%), and (iii) participants favored walking for trips less than 1 mile, e-bikes for trips of 1-3 miles, and e-bikes, cars, or shared rides for trips of 3-20 miles. The data used to generate these results have been made available in the Transportation Secure Data Center. We find e-bike use is appealing across age groups and may be related to characteristics of land use, urban form, occupation, income, and car ownership. We conclude for this population that the energy demand added by e-bike use (induced demand and replacing non-motorized modes) is outweighed by the reduction in energy demand from replacement of single-occupancy vehicle trips with e-bike trips. Our findings suggest considerable potential for energy savings from personal e-bike ownership.

29 ENERGY PLANNING, POLICY, AND ECONOMY

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION

Municipality of Anchorage Household Travel Survey

The 2002 Household Travel Survey for the municipality of Anchorage, Alaska, collected demographic, socioeconomic, and travel information about households and persons (age five and older) on an assigned, 24-hour travel period. The main objective of the study, conducted by NuStats, was to improve the transportation system. 2,035 households were recruited to participate in the study. Of these, 1,293 completed travel logs depicting detailed data on driving habits, such as purpose of the trips, time of the data, and mode of transportation—spanning from April 1, 2002, to May 17, 2002. This dataset is part of the Metropolitan Travel Survey Archive, which includes travel surveys from numerous public agencies across the United States and is archived by the Transportation Secure Data Center to ensure their continued public availability.

1Hz data

Scalable Data Center Capacity for DOE's AI Prototype: A Rapidly Available Gigawatt Data Center for DOE

The multilaboratory Gigawatt Data Center working group was commissioned to identify approaches to rapidly establish federal data centers with scalable capacities up to 1,000 MW. These state-of-the-art facilities will serve as hubs for interdisciplinary collaboration, industry partnerships, and transformative applications of artificial intelligence. The proposed strategic shift includes facilitating multilaboratory collaboration, prioritizing operational efficiency, expanding public–private partnerships, optimizing investments, ensuring long-term contractual flexibility, supporting open science and secure data enclaves, and exploiting high-speed national networks. Owing to their extensive experience and best practices, the US Department of Energy national laboratories are uniquely positioned to lead this initiative. We recommend conducting a feasibility analysis to rapidly identify the optimal sites for this initiative, and the effort will likely involve private industry for design, construction, financing, and operational integration. We also propose establishing multiple geographically diverse sites to ensure energy resilience, high operational reliability, and a diverse user base, thereby effectively addressing the nation’s critical needs.

42 ENGINEERING

2001 Atlanta Household Travel Survey

The 2001 Atlanta Household Travel Survey collected demographic, socioeconomic, and travel information on work and non-work travel behavior for a 48-hour travel period. The study was conducted on behalf of the Atlanta Regional Commission, and it is an essential element in the transportation planning and modeling efforts for the 13-county Atlanta region. The main objective of the study was to produce data that could be used to develop and calibrate travel demand models for use in travel forecasting, land use planning, and air quality planning to improve the transportation system. The second component in the survey was the deployment of an electronic travel diary with person-based GPS and an accelerometer to collect health and activity information. Travel data includes trip generation, trip distribution, and modal choice. The survey recruited a total of 12,184 households to participate in the study. Of these, 8,069 households (66%) completed travel. The 8,069 households, when weighted, represent 21,323 persons, 14,449 vehicles, and 126,127 places visited from April 2001 through April 2002. This dataset is part of the Metropolitan Travel Survey Archive, which includes travel surveys from numerous public agencies across the United States and is archived by the Transportation Secure Data Center to ensure their continued public availability.

1Hz data

5G Communications in Nuclear: Potential Use Cases and Security Considerations

As fifth-generation (5G) communications continues to revolutionize the future of wireless technology, there is growing demand to utilize its benefits for critical infrastructures such as nuclear power plants (NPPs). In regard to achieving full automation and control in the operation of existing and future nuclear reactors, the unique capabilities of 5G can bring several potential advantages over other wireless technologies. However, a deep investigation is needed for the availability and security of 5G communications under various NPP operational scenarios. This article examines how 5G security capabilities can be architecturally deployed in nuclear applications so as to replace existing communication infrastructures. We discuss the current use of all wireless technologies in NPPs with their key features. Consequently, we investigated several NPP use cases in which 5G offers potential advantages but entails specific security considerations. The present article covers the characteristics of 5G communications, general challenges to its application in nuclear, and the security gaps that need to be addressed. We also highlight certain 5G security-by-design features that can help addressing current stringent NPP requirements. In addition, we discuss some future research direction that can facilitate the implementation of 5G in a nuclear facility. The findings presented herein can help foster 5G deployment in NPPs, thus enabling secured data transmission, cost savings, and increased operational efficiency with enhanced reliability.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN

Securing 3D NAND Without Density Loss via In-Situ Encryption Using a Single Transistor XOR Cell

In this article, we push lightweight XOR-based in-situ encryption to extreme density by proposing a singletransistor XOR memory cell and applying it to 3D NAND, enabling secure data storage without density loss. Using a ferroelectric field-effect transistor (FeFET) as an example technology, we demonstrate that: i) a single-transistor memory can realize the XOR function by exploiting the ability to charge the source and drain separately and control current flow direction, eliminating the need for conventional encrypted cells that rely on complementary devices; ii) with a XOR-based cipher, encryption and decryption can be mapped to in-situ array operations, where ciphertext is stored as the threshold voltage (VTH) states of FeFETs in a NAND string, and decryption is achieved through read operations using key-dependent complementary source/drain bias; iii) the proposed technique is scalable to multi-level cell (MLC) storage by encrypting and decrypting data bit by bit; iv) using an integrated NAND FeFET array, we experimentally demonstrate encryption and decryption operations for both single-level cell (SLC) and MLC storage; v) systemlevel benchmarking shows that the proposed technique achieves 48× and 278× improvements in encryption and decryption throughput, respectively, compared to AES.

36 MATERIALS SCIENCE

Reverse Engineering of Medical Devices for Innovation and Advancement in Healthcare

• Medical technology is rapidly evolving and introducing new functionality and methodologies that suggest a higher risk for common vulnerability exposures (CVE) • Introduction of functionalities like Wi-Fi, Bluetooth, and internet connectivity require devices to be rigorously evaluated for vulnerabilities • Subsequently like many other fields cell-phone interconnectivity suggests a significantly higher level of risk to critical infrastructure and data security

Baldwin, David [Savannah River National Laboratory

Scalable and Secure Power Outage Data Reporting: A Hexagonal Geospatial Approach

Power outages disrupt critical infrastructure and cause billions of dollars in economic losses annually in the United States. Accurate and granular outage reporting is vital for effective restoration and mitigation. This paper examines the integration of the Hexagonal Hierarchical Geospatial Indexing System (H3) to enhance power outage reporting, leveraging its uniform grid structure, scalable resolutions, and support for privacy-preserving analysis. Using high-resolution LandScan Global population data and K-anonymization techniques, this work achieves a balance between data granularity and privacy. Results show that lower privacy thresholds (e.g., K-anonymity = 2) enable higher resolution, while stricter thresholds (e.g., >15 people per hex) reduce granularity, potentially affecting localized responses. State-and county-level resolution case studies demonstrate H3’s adaptability and the trade-offs between precision and privacy. The proposed H3-based framework offers a scalable and efficient solution for geospatial data integration within the energy sector, such as outage data, aiding utilities and regulators in improving resilience and response efforts, particularly in disaster-prone regions.

Ahmad, Nasir [ORNL] (ORCID:0000000150677368)