Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data protection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Fault Detection in the Solvent Extraction Process with Non-traditional Sensors

Solvent extraction is used to separate metals or other complexes into two different immiscible liquids and is an essential component of the PUREX Process. Improvements to the solvent extraction process can directly contribute to an organization’s ability to ensure the purity and recovery of special nuclear materials from spent nuclear fuel. Such improvements can benefit nuclear reprocessing efforts, increase fuel reutilization, and limit the concentration of actinides in nuclear waste repositories.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

MSIIP Poster

This poster is a brief summary of my project and my experience during my summer internship.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Emerging Technologies for Privacy Preservation in Energy Systems

This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.

Cali, Umit↗

Classification of animal sounds in a hyperdiverse rainforest using convolutional neural networks with data augmentation

To protect tropical forest biodiversity, we need to be able to detect it reliably, cheaply, and at scale. Automated detection of sound producing animals from passively recorded soundscapes via machine-learning approaches is a promising technique towards this goal, but it is constrained by the necessity of large training data sets. Using soundscapes from a tropical forest in Borneo and a Convolutional Neural Network model (CNN), we investigate i) the minimum viable training data set size for accurate prediction of call types (‘sonotypes’), and ii) the extent to which data augmentation and transfer learning can overcome the issue of small and imbalanced training data sets. We found that even relatively high sample sizes (>80 per sonotype) lead to mediocre accuracy, which however improved significantly with data augmentation and transfer learning, including at extremely small sample sizes (3 per sonotype), regardless of taxonomic group or call characteristics. Neither transfer learning nor data augmentation alone achieved high accuracy. Our results suggest that transfer learning and data augmentation could make the use of CNNs to classify species’ vocalizations feasible even for small soundscape-based projects with many rare species. Retraining our open-source model requires only basic programming skills which makes it possible for individual conservation initiatives to match their local context, in order to enable more evidence-informed management of biodiversity.

54 ENVIRONMENTAL SCIENCES↗

Defender Policy Evaluation and Resource Allocation against MITRE ATT&CK Data and Evaluations

Protecting against multi-step attacks of uncertain duration and timing forces defenders into an indefinite, always ongoing, resource-intensive response. To effectively allocate resources, a defender must be able to analyze multi-step attacks under assumption of constantly allocating resources against an uncertain stream of potentially undetected attacks. To achieve this goal, we present a novel methodology that applies a game-theoretic approach to the attack, attacker, and defender data derived from MITRE´s ATT&CK ® Framework. Time to complete attack steps is drawn from a probability distribution determined by attacker and defender strategies and capabilities. This constraints attack success parameters and enables comparing different defender resource allocation strategies. By approximating attacker-defender games as Markov processes, we represent the attacker-defender interaction, estimate the attack success parameters, determine the effects of attacker and defender strategies, and maximize opportunities for defender strategy improvements against an uncertain stream of attacks. This novel representation and analysis of multi-step attacks enables defender policy optimization and resource allocation, which we illustrate using the data from MITRE´ s APT3 ATT&CK ® Framework.

97 MATHEMATICS AND COMPUTING↗

Measurement of Local Differential Privacy Techniques for IoT-based Streaming Data

Various Internet of Things (IoT) devices generate complex, dynamically changed, and infinite data streams. Adversaries can cause harm if they can access the user’s sensitive raw streaming data. For this reason, protecting the privacy of the data streams is crucial. In this paper, we explore local differential privacy techniques for streaming data. We compare the techniques and report the advantages and limitations. We also present the effect on component (e.g., smoother, perturber) variations of distribution-based local differential privacy. We find that combining distribution-based noise during perturbation provides more flexibility to the interested entity.

Afrose, Sharmin↗

Specific Gamma-Ray Dose Constants with Current Emission Data

The specific gamma-ray dose constant represents the gamma effective dose rate due to a point source of unit activity of a given nuclide at 1 m. New tabulations of specific gamma-ray dose constants have been made using current gamma emission data from the SCALE 6.2.3 software package and International Commission on Radiological Protection Publication 107, combined with the effective dose per fluence conversion coefficients (antero-posterior orientation) of International Commission on Radiological Protection Publication 116. SCALE data cover 1,264 nuclides, and International Commission on Radiological Protection Publication 107 data include 1,192 nuclides, with only 777 nuclides in common between the two sets.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Analyzing Data Privacy for Edge Systems

Internet-of-Things (IoT)-based streaming applications are all around us. Currently, we are transitioning from IoT processing being performed on the cloud to the edge. While moving to the edge provides significant networking efficiency benefits, IoT edge computing creates significant data privacy concerns. We propose a methodology that can successfully privacy protect the continual data streams generated by sensors on the edge device. We implement local differential privacy on streaming data and incorporate Bayesian inference and Gaussian process to evaluate the privacy policy. We demonstrate our methodology on a real-world smart meter testbed and identify the optimal privacy protection settings.

Kotevska, Olivera↗

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Synchronized Waveforms – A Frontier of Data-Based Power System and Apparatus Monitoring, Protection, and Control

Voltage and current waveforms contain the most authentic and granular information on the behaviors of power systems. In recent years, it has become possible to synchronize waveform data measured from different locations. Thus large-scale coordinated analyses of multiple waveforms over a wide area are within our reach. This development could unleash a set of new concepts, strategies, and tools for monitoring, protecting, and controlling power systems and apparatuses. This paper presents an in-depth review and analysis of the advancements in synchronized waveform data, including measurement devices, data characteristics, use cases, and comparisons with synchrophasor data. Based on the findings, five strategies are proposed to discover and develop synchronized waveform based applications over multiple application areas. The paper also presents three complementary measurement platforms and two data screening algorithms for application implementation. It further discusses committee activities and standard developments useful to explore the full potential of the data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Development and Characterization of Cementitious Waste Forms for Immobilization of Granular Activated Carbon, Silver Mordenite, and HEPA Filter Media Solid Secondary Waste

At the Department of Energy’s Hanford site, over 53 million gallons of chemically complex and radioactive wastes have been stored in 177 underground tanks. The Hanford Tank Waste Treatment and Immobilization Plant (WTP) is under construction and is designed to treat and immobilize these wastes. During operations of WTP, solid secondary wastes (SSWs) will be generated as a result of waste treatment, vitrification, off-gas management, and supporting process activities. SSW treatment processes and resulting disposal pathways for the final disposition form of the SSW are needed to support direct feed low activity waste (DFLAW) operations and facilitate continued operation of WTP. The SSWs produced through WTP operations are expected to include used process equipment, contaminated tools and instruments, decontamination wastes, high-efficiency particulate air (HEPA) filters, carbon absorption beds (granular activated carbon, GAC), silver mordenite (AgM) and spent ion-exchange resins. These waste streams are planned to be immobilized in a cementitious waste form and disposed of either as stabilized/blended (non-debris) or encapsulated (debris) in a cementitious waste form. Accordingly, cementitious waste forms from these streams were included in the 2017 Integrated Disposal Facility (IDF) Performance Assessment (PA). The input data used to represent these SSW forms in the 2017 IDF PA involved many assumptions and associated uncertainties. This data limitation was due to the lack of material- and site-specific data available for representative SSW materials in cementitious matrices. To verify the assumed values used in the IDF PA and fill this limitation in available data, Washington River Protection Solutions, LLC (WRPS), has initiated a program targeted toward gathering site specific data relevant to Hanford SSW disposal. The work within this report is a continuation of this ongoing program.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Adapting Secure MultiParty Computation to Support Machine Learning in Radio Frequency Sensor Networks

In this project we developed and validated algorithms for privacy-preserving linear regression using a new variant of Secure Multiparty Computation (MPC) we call "Hybrid MPC" (hMPC). Our variant is intended to support low-power, unreliable networks of sensors with low-communication, fault-tolerant algorithms. In hMPC we do not share training data, even via secret sharing. Thus, agents are responsible for protecting their own local data. Only the machine learning (ML) model is protected with information-theoretic security guarantees against honest-but-curious agents. There are three primary advantages to this approach: (1) after setup, hMPC supports a communication-efficient matrix multiplication primitive, (2) organizations prevented by policy or technology from sharing any of their data can participate as agents in hMPC, and (3) large numbers of low-power agents can participate in hMPC. We have also created an open-source software library named "Cicada" to support hMPC applications with fault-tolerance. The fault-tolerance is important in our applications because the agents are vulnerable to failure or capture. We have demonstrated this capability at Sandia's Autonomy New Mexico laboratory through a simple machine-learning exercise with Raspberry Pi devices capturing and classifying images while flying on four drones.

42 ENGINEERING↗

Canopy reflectance spectroscopy, thermal images and digital photographs, San Lorenzo, Panama, 2020

This data package comprises reflectance spectra, thermal and visible images of top of canopy tree crowns at the San Lorenzo Protected Area, Panama. Data were collected using the Smithsonian Tropical Research Institute (STRI) canopy access crane, from positions approximately 4 m above the target tree canopy. Canopy reflectance spectra were collected using an HR-1024i full range spectroradiometer (350–2500 nm, SVC, Poughkeepsie, NY, USA) with a 14-degree lens. Canopy thermal images were captured using an 8640-S series USB Calibrated Thermal Camera (ICI International, Beaumont, TX, USA). Visible images were collected using an AW130 waterproof/shockproof camera (Nikon, Tokyo, Japan). Data were collected on three days, on 16 canopies from 10 tree species. On both January 30, 2020, and February 27, 2020, we collected spectra at three time points, in the morning, midday, and afternoon. On February 18, 2020, we collected data around midday only. Due to technical issues, digital photography was collected on 30 January only. This dataset includes unprocessed data only, including spectral data (.sig or .sed), thermal images (.i16), and photographs (.raw), a detailed method description (.pdf), species information (.csv) and ESS-DIVE file-level metadata (FLMD.csv).

54 ENVIRONMENTAL SCIENCES↗

Estimated capital costs of fish exclusion technologies for hydropower facilities

Hydropower is a reliable source of renewable energy, and its future expansion is likely to be in the form of either smaller new stream development (NSD) projects or powering existing non-powered dams. Thresholds for entrainment risk to fish and the requirements for fish exclusion at hydropower facilities often differ depending on the species involved, the characteristics of the facility, and the goals of stakeholders, but little quantitative information is present within the literature regarding the specific costs of fish exclusion measures. Cost data associated with protection, mitigation, and enhancement (PM&E) measures related to positive barrier screening were identified using keyword searches of an existing environmental mitigation cost data set and manual extraction from regulatory licensing documents available in the Federal Energy Regulatory Commission (FERC) eLibrary. This approach yielded a total of 50 p.m.&E mitigation measures with estimated capital construction costs pertaining to positive barrier screens and represented <10% of the 171 total FERC project dockets available in the data set. These data were highly skewed toward conventional relicensing projects, as <7% were associated with NSD projects. Results indicate highly variable costs are associated with fish screening, with flow-normalized costs one to two orders of magnitude higher for screening with the highest exclusion capability (≤0.09 in. spacing) compared with coarser screening (1–2 in.). These data provide an initial baseline for estimating exclusion costs for hydropower development and may help developers consider options for more fish-friendly generation technologies, though gaps remain relating to a lack of data, particularly for NSD projects.

13 HYDRO ENERGY↗

Update of Emission Factors of Greenhouse Gases and Criteria Air Pollutants, and Generation Efficiencies of the U.S. Electricity Generation Sector

The last decade has seen a steady evolution of the electricity generation sector. Fuels used for electricity generation have shifted from coal to cleaner energy sources such as natural gas and renewables including solar, wind, and other renewable sources. The share of U.S. electricity generated from coal decreased from 45% in 2010 to 24% in 2019, and is expected to decrease further to 13% by 2050. The conversion efficiency of electricity generation has also increased gradually for fuels such as natural gas due as less-efficient old generators are retired and more-efficient generators replace them. These changes in the electricity generation industry are likely to cause changes in the emissions from power generation units. Emission factors of greenhouse gases (GHG) including CO 2 , CH 4 , and N 2 O, and criteria air pollutants (CAPs) including CO, NO x , PM 10 , PM 2.5 , and SO x , from power plants are important parameters for estimating life-cycle emissions associated with vehicle electrification, energy systems, and the production of materials and chemicals. The electricity generation technologies and associated emission factors in the Greenhouse Gases, Regulated Emissions, and Energy Use in Technologies (GREET) model need to be updated to reflect recent developments in the electricity generation sector. The most recent update of the electricity generation emission factors in GREET adopted a mixed method. The emission factors of CH 4 , N 2 O, NO x , and SOx were estimated using a “topdown” approach by dividing the total emissions by the total net electricity generation, because emission data of these pollutants are readily available in the Emissions & Generation Resource Integrated Database (eGRID). For other CAPs such as CO, VOC, PM 10 , and PM 2.5 , emission data were not reported in eGRID. A “bottom-up” method was used to estimate the emission factors for these pollutants by considering generic uncontrolled emission factors and the pollutant removal efficiencies of emission control technologies adopted in the electricity generation sector. However, the uncontrolled emission factors and the emission removal efficiencies of various emission control technologies considered in the 2012 study came from the legacy AP-42 emission factors, and may not reflect the actual emission performances of the electricity generation sector of today. To leverage new data that recently became available, especially emission data measured from continuous emission monitoring systems (CEMS), we developed a new “top-down” approach to estimate efficiencies and GHG and CAP emission factors for electricity generation from combustion of individual fuel types by individual combustion technologies on the basis of power-generation data from U.S. Energy Information Administration’s (EIA’s) form EIA-923, and plant emission data from Environmental Protection Agency’s (EPA’s) Clean Air Markets Division (CAMD) dataset and National Emissions Inventory (NEI) dataset. Detailed discussion of the method and data used in this study can be found in Section 2.1. With this topdown approach, we aim to improve the estimates of energy efficiencies and emission factors for power plants using a more consistent methodology, and to update the emission factors, generation efficiencies, and generation technologies mixes in GREET to reflect recent technology advancements in the electricity generation sector.

20 FOSSIL-FUELED POWER PLANTS↗

Recent Experience with the CMS Data Management System

The CMS[1] experiment manages a large-scale data infrastructure, currently handling over 200 PB of disk and 500 PB of tape storage and transferring more than 1 PB of data per day on average between various WLCG[2] sites. Utilizing Rucio[3] for high-level data management, FTS[4] for data transfers, and a variety of storage and network technologies at the sites, CMS confronts inevitable challenges due to the system’s growing scale and evolving nature. Key challenges include managing transfer and storage failures, optimizing data distribution across different storages based on production and analysis needs, implementing necessary technology upgrades and migrations, and efficiently handling user requests. The data management team has established comprehensive monitoring to supervise this system and has successfully addressed many of these challenges. The team’s efforts aim to ensure data availability and protection, minimize failures and manual interventions, maximize transfer throughput and resource utilization, and provide reliable user support. This paper details the operational experience of CMS with its data management system in recent years, focusing on the encountered challenges, the effective strategies employed to overcome them and the ongoing challenges as we prepare for future demands.

Öztürk, Hasan [CERN]↗

Investigating Users’ Privacy Concerns of Internet of Things (IoT) Smart Devices

Although the number of smart Internet of Things (IoT) devices has grown in recent years, the public's perception of how effectively these devices secure IoT data has been questioned. Many IoT users do not have a good level of confidence in the security or privacy procedures implemented within IoT smart devices for protecting personal IoT data. Moreover, determining the level of confidence end users have in their smart devices is becoming a major challenge. In this paper, we present a study that focuses on identifying privacy concerns IoT end users have when using IoT smart devices. We investigated multiple smart devices and conducted a survey to identify users’ privacy concerns. Furthermore, we identify five IoT privacy-preserving (IoTPP) control policies that we define and employ in comparing the privacy measures implemented by various popular smart devices. Results from our study show that the over 86% of participants are very or extremely concerned about the security and privacy of their personal data when using smart IoT devices such as Google Nest Hub or Amazon Alexa. In addition, our study shows that a significant number of IoT users may not be aware that their personal data is collected, stored or shared by IoT devices.

Joy, Daniel↗

Heavy-Duty Vehicle Activity Updates for MOVES Using NREL Fleet DNA and CE-CERT Data

The U.S. Environmental Protection Agency's (EPA's) Motor Vehicle Emission Simulator (MOVES) is a publicly available tool used by researchers and policymakers to help understand motor vehicle emission sources at a national, county, and project level. Estimates of heavy-duty activity in the most recent version of the model at the time this work was conducted, MOVES2014, was identified as an area in need of improvement. The start activity in MOVES2014 is based on a limited and dated data set. In addition, MOVES2014 relies on drive cycles that represent on-network activity but do not account for idling activity that occurs on off-network roads, such as at a distribution center, while the truck is queuing or during loading and unloading. As a result, MOVES2014 may currently underestimate the number of starts and idle and soak time for heavy-duty trucks in real-world operation. The National Renewable Energy Laboratory (NREL) has previously leveraged its expansive Fleet DNA database of heavy-duty vehicles to idle and start activity for six of the nine heavy-duty vehicle source types of classes in the MOVES model. The data available in Fleet DNA from 416 conventional, diesel-powered vehicles provided activity estimates from more than 120,000 hours of operation throughout 14,682 vehicle days between October 2006 and January 2016. NREL calculated start fraction, starts per day, soak fraction, and idle fraction by hour of the day for each vehicle type, state, and vocation, and provided results in .CSV files that can be translated to MOVES table inputs. The idle and start activity from this initial analysis of Fleet DNA data was used to develop default idle and start data for heavy-duty vehicles in MOVES3. Satisfied with the results from the Fleet DNA data used for MOVES3, the EPA asked NREL to extend this start/soak/idle analysis using additional data from a larger number of vehicles for a potential future update to the MOVES model. Such a data set was achieved from a project led by the University of California at Riverside, College of Engineering, Center for Environmental Research & Technology (CE-CERT) and funded by California Air Resources Board. Specifically, this data set consists of 90 heavy-duty vehicles operated mainly in California, which can be separated into five of the nine heavy-duty vehicle classes in the MOVES model. In addition, the heavy-duty activity database collected by CE-CERT provided activity estimates from more than 44,000 hours of operation throughout 4,724 vehicle days between November 2014 and September 2016. This report details the analysis of the heavy-duty activity database collected from the University of California at Riverside by providing graphical analysis and context for the start, soak, and idle distributions. The comparison of the related results from both the Fleet DNA and CE-CERT data sets are documented as well.

33 ADVANCED PROPULSION SYSTEMS↗