Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Anonymization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Swarm autonomic agents with self-destruct capability

Systems, methods and apparatus are provided through which in some embodiments an autonomic entity manages a system by generating one or more stay alive signals based on the functioning status and operating state of the system. In some embodiments, an evolvable synthetic neural system is operably coupled to one or more evolvable synthetic neural systems in a hierarchy. The evolvable neural interface receives and generates heartbeat monitor signals and pulse monitor signals that are used to generate a stay alive signal that is used to manage the operations of the synthetic neural system. In another embodiment an asynchronous Alice signal (Autonomic license) requiring valid credentials of an anonymous autonomous agent is initiated. An unsatisfactory Alice exchange may lead to self-destruction of the anonymous autonomous agent for self-protection.

Hinchey, Michael G.↗

Important Changes to User Access at the NASA CDDIS

The Crustal Dynamics Data Information System (CDDIS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to make space geodesy and geodynamics related data and derived products available in a central archive, to maintain information about the archival of these data, to disseminate these data and information in a timely manner to a global scientific research community, and to provide user based tools for the exploration and use of the archive. Since its inception, the user community has utilized anonymous ftp for accessing and downloading files from the CDDIS archive. Although this protocol allows users to easily automate file downloads, many organizations, data systems, and users have already migrated from ftp or are actively pursuing a move away from the protocol due to problems from a system and security standpoint. Furthermore, U.S. Government agencies have become increasingly concerned about this legacy protocol and ensuring data integrity for the user community have begun recently to disallow the use of the ftp protocol. The CDDIS, operated by NASA GSFC, must therefore address these concerns and provide alternative methods for access to its archive for continued easy and automated download of its contents. This poster will discuss the upcoming changes at CDDIS and provide examples on transitioning from anonymous ftp.

Noll, Carey E.↗

FOA 1861 Data Curation Overview

This document describes the process executed to collect, examine, and consolidate Phasor Measurement Unit (PMU) data from multiple transmission operators into a common dataset. The consolidated PMU data set was further anonymized and distributed to the Department of Energy Funding Opportunity Announcement (FOA) 1861 Big Data Analysis of Synchrophasor Data awardees.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning

Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets () according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ -aware server aggregation to improve model utility by re-weighting client updates according to their declared privacy budgets. However, gradient updates in FL retain structural patterns induced by non-independent and identically-distributed (non-IID) data, and these additional signals exposed by -aware aggregation create new opportunities for inference by an honest-but-curious server. In this work, we first show that a server equipped with gradient denoising and surrogate modeling can mount a Privacy Inference Attack that infers distributional attributes of clients and links updates from the same client across training rounds, measured via surrogate inference accuracy and linkage success, under realistic knowledge constraints. The Shuffle-Model has been widely studied as a defense against such inference risks by anonymizing update sources, but it is fundamentally incompatible with HDP-FL -aware aggregation. To address this challenge, we propose IntraShuffler, a middleware defense framework designed for HDP-FL systems. IntraShuffler introduces a privacy-aware shuffling mechanism that groups clients into privacy-compatible buckets and performs parameter-level shuffling within each bucket to disrupt persistent gradient structure while preserving -aware aggregation. Experiments across four different datasets show that IntraShuffler reduces gradient recoverability by over 60% and decreases surrogate inference accuracy from 0.78 to 0.33 while maintaining comparable model utility across multiple FL aggregation rules.

Riya, Farhin Farhad [ORNL]↗

Using advanced data structures to enable responsive security monitoring

Write-optimized data structures (WODS), offer the potential to keep up with cyberstream event rates and give sub-second query response for key items like IP addresses. These data structures organize logs as the events are observed. To work in a real-world environment and not fill up the disk, WODS must efficiently expire older events. As the basis for our research into organizing security monitoring data, we implemented a tool, called Diventi, to index IP addresses in connection logs using RocksDB (a write-optimized LSM tree). In this work, we extended Diventi to automatically expire data as part of the data structures’ normal operations. We guarantee that Diventi always tracks the N most recent events and tracks no more than N + k events for a parameter k < N, while ensuring the index is opportunistically pruned. To test Diventi at scale in a controlled environment, we used anonymized traces of IP communications collected at SuperComputing 2019. We synthetically extended the 2.4 billion connection events to 100 billion events. We tested Diventi vs. Elasticsearch, a common log indexing tool. In our test environment, Elasticsearch saw an ingestion rate of at best 37,000 events/s while Diventi sustained ingestion rates greater than 171,000 events/s. Our query response times were as much as 100 times faster, typically answering queries in under 80 ms. Furthermore, we saw no noticeable degradation in Diventi from expiration. We have deployed Diventi for many months where it has performed well and supported new security analysis capabilities.

97 MATHEMATICS AND COMPUTING↗

Interlaboratory Reproducibility of Contour Method Data in a High Strength Aluminum Alloy

The contour method for residual stress measurement has seen significant development, but an experimental reproducibility study utilizing physical samples has not been published. A double-blind reproducibly study is reported, having scope beginning with EDM cutting and ending with residual stress calculation. A reinforced I-beam sample geometry is identified for its unique residual stress profile when extracted from residual stress bearing quenched aluminum bar (7050-T74). Contour measurements are prescribed on a midplane of symmetry with dimensions 24.0 mm by 50.0 mm. Fourteen identically prepared samples are fabricated from a single long bar with well characterized and uniform residual stress. Five samples throughout the bar are identified for planning measurements to validate sample uniformity and overall suitability of the residual stress field. The planning measurements employ a range of techniques: contour method, neutron diffraction, and hole-drilling. Eight samples are distributed to an international group of participants to execute their standard measurement practice. A double-blind process is followed to provide anonymity. Results are provided by eight participants: six being self-similar and two being quite different, the latter set aside as outliers. An average residual stress field is established from non-outlying results and the spatial distribution of reproducibility standard deviation is determined. The average stress field ranges from -60 to 70 MPa and the reproducibility standard deviation averages 8.1 MPa on the measurement plane. The average reproducibility standard deviation is about 3 × larger for points within 1.0 mm of plane boundaries (17.6 MPa) than for the remaining points (6.1 MPa). Reproducibility standard deviation (among different labs) for contour method residual stress measurement is found to be very similar to repeatability standard deviation (in a single lab) reported in prior work. The reproducibility observed here, for the entire measurement process, is also similar to that found in a prior reproducibility study limited to contour method data analysis.

36 MATERIALS SCIENCE↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Testing SOAR tools in use

Investigations within Security Operation Centers (SOCs) are tedious as they rely on manual efforts to query diverse data sources, overlay related logs, correlate the data into information, and then document results in a ticketing system. Security Orchestration, Automation, and Response (SOAR) tools are a relatively new technology that promise, with appropriate configuration, to collect, filter, and display needed diverse information; automate many of the common tasks that unnecessarily require SOC analysts’ time; facilitate SOC collaboration; and, in doing so, improve both efficiency and consistency of SOCs. There has been no prior research to test SOAR tools in practice; hence, understanding and evaluation of their effect is nascent and needed. Here, in this paper, we design and administer the first hands-on user study of SOAR tools, involving 24 participants and six commercial SOAR tools. Our contributions include the experimental design, itemizing six characteristics of SOAR tools, and a methodology for testing them. We describe configuration of a cyber range test environment, including network, user, and threat emulation; a full SOC tool suite; and creation of artifacts allowing multiple representative investigation scenarios to permit testing. We present the first research results on SOAR tools. Concisely, our findings are that: per-SOC SOAR configuration is extremely important; SOAR tools increase efficiency and reduce context switching, although with potentially decreased ticketing accuracy/completeness; user preference is slightly negatively correlated with their performance with the tool; internet dependence varies widely among SOAR tools; and balance of automation with assisting decision making is preferred by senior participants. We deliver a public user- and tool-anonymized and -obfuscated version of the data.

97 MATHEMATICS AND COMPUTING↗

Feature review of photovoltaic modeling software utilizing blind performance assessment

While confidence in photovoltaic (PV) modeling software has always been essential, the rapid pace of new PV plant developments makes accuracy and credibility more critical than ever. Independent assessments, particularly through blind modeling comparisons, are therefore necessary to ensure unbiased benchmarking across PV modeling software. Previous studies have been limited by a narrow range of models compared, anonymized results, or system size. This study presents results from the first-ever onymous blind modeling comparison, evaluated using both lab- and utility-scale fixed-tilt, monofacial, south-facing systems at sub-hourly time intervals. Seven commercially used PV software tools were compared: 3E SynaptiQ, PlantPredict, PVsyst, RatedPower, SAM, SolarFarmer, and Solargis Evaluate. Predictions were submitted directly by software representatives, providing unique insights into each software’s implementation and resulting prediction behavior. Notable features, including plane-of-array (POA) transposition model, module temperature model, shading model, and performance model were analyzed and compared. Four summary tables compile these features of the software, serving as a resource to help users understand the methodological differences and select the most suitable software for their applications. The software tools show deviations from mean error in annual yield up to 2.5 % in the lab-scale system, increasing to 6.0 % for the utility-scale system. These differences arise from a combination of user decisions and the inherent behavior of the software, indicating the need for continuous and rigorous validation of modeling methods using these software tools against complex, real-world systems.

14 SOLAR ENERGY↗

Spatial and Temporal Characterization of Activity in Public Space, 2019–2020

The data reported here characterize spatial and temporal variation in the ratio of short-to-long-duration visits in public places (i.e., points of interest) in the United States for each week between January 2019 and December 2020. The underlying data on anonymized and aggregated foot traffic to public places is curated by SafeGraph, a geospatial data provider. In this work, we report the estimated number and duration of “short” (i.e., <4 hours) and “long” (i.e., >4 hours) visits to public places at the US census block group level. Long visits are shown to be a good proxy for workers based on formal economic data. We propose that short visits are more likely to represent nonobligate activities: people visiting a public place for leisure, shopping, entertainment, or civic or cultural engagement. Our work constructs a ratio of short to long visits, which can be used to inform population estimates for nonworker use of public space. These data may be useful for understanding how people’s use of public space has changed during the COVID-19 pandemic and, more generally, for understanding activity patterns in public.

99 GENERAL AND MISCELLANEOUS↗

Crowd cluster data in the USA for analysis of human response to COVID-19 events and policies

We provide data on daily social contact intensity of clusters of people at different types of Points of Interest (POI) by zip code in Florida and California. This data is obtained by aggregating fine-scaled details of interactions of people at the spatial resolution of 10 m, which is then normalized as a social contact index. We also provide the distribution of cluster sizes and average time spent in a cluster by POI type. This data will help researchers perform fine-scaled, privacy-preserving analysis of human interaction patterns to understand the drivers of the COVID-19 epidemic spread and mitigation. Current mobility datasets either provide coarse-level metrics of social distancing, such as radius of gyration at the county or province level, or traffic at a finer scale, neither of which is a direct measure of contacts between people. We use anonymized, de-identified, and privacy-enhanced location-based services (LBS) data from opted-in cell phone apps, suitably reweighted to correct for geographic heterogeneities, and identify clusters of people at non-sensitive public areas to estimate fine-scaled contacts.

60 APPLIED LIFE SCIENCES↗

Effectiveness of Privacy Techniques in Smart Metering Systems

Smart grid technologies enable timely energy billing for residential homes. The ability to react to energy demands during peak hours allows energy providers to conserve power and operate efficiently. However, these data streams are also susceptible to privacy attacks within the energy company and from outside hackers. We implemented four different privacy models: k-anonymous, l-diversity, t-closeness, and ε-differential privacy. We demonstrate the models’ effectiveness using a real-world dataset composed of 15 different residential households with energy consumption data spanning over a year.

Peralta-Peterson, Martin↗

Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-the-Meter Solar

The ever-growing high penetration of ubiquitously distributed energy resources, especially behind-the-meter solar (BTM) generations, has significant impacts on nodal load (i.e., net injection) profiles and consequently caused imperative operational challenges to system operators such as regional transmission organizations (RTOs). Illustrated by real-world nodal data and examples at PJM Interconnection, this paper first discusses the application and necessity of effectively extracting daily nodal load profiles in a non-intrusive manner. More importantly, a novel bi-level architecture, including Factorization Machines (FM) learning procedure has been proposed to effectively disaggregate not only one node but every node in an RTO service territory. Specifically, FM leaning is adopted to capture the interconnections between related features to better utilize the correlation between buses in the same region and between a single bus and the zonal load. The proposed bi-level technique is numerically validated using real-world, minute-level, normalized, and anonymized nodal data at PJM service territory.

behind the meter solar, load disaggregation, load ↗

Advancing Our Understanding of System Availability through the PV Fleet Performance Data Initiative

The PV Fleet Performance Data Initiative partners with photovoltaic (PV) fleet owners to collect time-series data of PV production data and publishes aggregated anonymized results of system performance metrics. With an extensive dataset drawn from over 2,200 PV systems across the United States, comprising 8.5 GW and 24,000 separate inverter data channels, this initiative aims to ensure that systemic risks in the US PV fleet are detected. The current work explores system availability, revealing a pronounced dependence on time, especially within the initial 6 months of system performance. Following this start-up period, the average system availability stabilizes. Statistical analyses illustrate a median (P5O) monthly availability of 0.991 and a dependence on system size with a negative trend in availability with increasing system size. This finding indicates that larger systems experience lower availability compared to their smaller counterparts.

inverter availability↗

Deep Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-The-Meter Solar

The ever-growing integration of distributed energy resources (DERs), especially behind-the-meter (BTM) solar generations, poses imperative operational challenges to system operators such as regional transmission organizations (RTOs). It is important for RTOs to effectively and accurately extract actual load profiles at the transmission level for a single node with significant BTM solar injection. This paper first illustrates the necessity of disaggregating the daily actual load profile of a single node. Furthermore, by segmenting nodes with selected timeseries features, nodes with significant BTM solar generation are identified. Lastly, a bi-level framework is proposed, comprising reference node disaggregation and DeepFM nodal disaggregation, aimed at disaggregating the nodal load profiles from which system operators require more information. By adopting a hybrid Deep Factorization Machine (DeepFM) model, the model achieve accurate results by extracting both linear and nonlinear relations between nodes in the same region and the zonal load and nodal load profile. To overcome the lack of ground truth, this paper segments the load profile into daytime, nighttime, and zero-crossing points and utilizes the latter two for evaluation purposes. The proposed disaggregation procedure is validated using real world, minute-level, normalized, and anonymized nodal data in the PJM service territory.

42 ENGINEERING↗

Contrasting student and staff perceptions of preclinical‐to‐clinical transition at a Chilean dental school

Abstract Introduction Dental education is a challenging and demanding field of study as students are expected to acquire various competencies to fulfil their professional requirements after graduation. The objective of this study was to investigate and compare dental students' and clinical staff instructors' perceptions of the preclinical‐to‐clinical transition training at a Dental School in Santiago, Chile. Material and Methods Two questionnaires containing 11 quantitative and one qualitative item were developed to assess our year three, four and five ( n = 244) dental undergraduate students' challenges when they begin treating patients, and clinical staff ( n = 78) perceptions of the preparedness to treat patients of the same students. Both questionnaires were voluntarily and anonymously implemented eight weeks after the beginning of the 2019 academic year. Responses were analysed using a Chi‐squared test for each quantitative question, while qualitative comments were studied to form themes and dimensions. RESULTS A total of 234 (96%) students and 60 (77%) instructors completed their respective questionnaire. There were considerable variations between students in the different years of the programme, as well as between students and staff members. Students and instructors felt the former had enough knowledge to treat patients though it was difficult for them to apply it in clinical practice. Again, both believed they could communicate with patients, but third year students asked for more training on this. Regarding practical skills, fourth‐ and fifth‐year students felt prepared but not third year students, who preferred to work in pairs with senior students, a preference that was shared by the instructors. All student groups asked clinical staff to provide more frequent, constructive and consistent feedback and felt that the difference between simulation and clinical environments and the amount of clinical work to fulfil clinical requirements made them feel stressed. Another mentioned stressor was students' low self‐confidence when working with patients. Among the requested improvements, students requested better training on how the dental clinic works to save time. Conclusions Preclinical‐to‐clinical transition training presents several challenges. Some of the problems highlighted by both students and clinical staff members persisted with the transition after three, four and even five years of training, which needs to be addressed.

Tricio, Jorge↗

Broadband, 920-nm mirror thin film damage competition

This year’s competition proposed to survey the state-of-the-art broadband, near-IR multilayer dielectric (MLD) mirrors designed for ultra-short, pulsed laser applications. The requirements for the coatings were a minimum reflection of 99.5% at 45-degree incidence angle for S-polarization from 830 nm to 1010 nm and group delay dispersion (GDD) < ± 50 fs 2 . The participants in this effort selected the coating materials, coating design, and deposition method. Samples were damage tested at a single testing facility to enable direct comparison among the participants using a 25 ± 5 fs OPCPA laser system operating at 5 Hz. A double blind test assured sample and submitter anonymity. The damage performance results, sample rankings, details of the deposition processes, coating materials and substrate cleaning methods are shared here. We found that multilayer coatings using tantala and/or hafnia as high index materials were top performers within several coating deposition groups. Specifically, dense coatings by ion-beam sputtering (IBS), magnetron sputtering (MS), and electron-beam ion assisted deposition (e-beam IAD) exhibited highest damage initiation onset (LIDT) while e-beam coatings were low performers. In addition, damage growth onset (LDGT) was also examined and the results are reported here for all samples as this performance metric plays an important role in establishing the safe operational conditions for larger aperture, ultrashort pulsed lasers. As a result, not all coating samples in the survey met the GDD requirements stated above and associated measurements are discussed in the context of the present and past competitions focused on similar broadband, near-IR MLD coatings.

42 ENGINEERING↗

Using AI tools to analyze periodic phenomena

Here, in this report, we describe how students can use natural language to prompt ChatGPT to conduct the analysis of complex periodic acceleration data collected using their smartphones. Students may use this approach to characterize human physiological tremor frequency. Students can choose to use their own data, or, for privacy purposes, use an anonymized set of data provided to them.

Klay, Jennifer L. [California Polytechnic State Un↗