Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Anonymization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

PV Fleet Performance Data Initiative Program and Methodology

The US Department of Energy’s PV Fleet Performance Data Initiative has been launched in order to collect and evaluate production data across multiple PV fleet partners. Performance statistics are anonymized, aggregated and shared to represent a snapshot of the US commercial and utility-scale fleet. Production data have been collected from over 1500 systems representing more than 1.3 GWdc capacity. Preliminary analysis indicates median performance loss rates are in line with previous publications of system degradation, on the order of –0.6%/yr to –0.9%/yr (preliminary numbers subject to change). These values are higher than module-only degradation rates which are often used in pro-forma estimates of project performance and economics, potentially exposing owner/operators to increased risk if systems under-perform over time.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Advancing Our Understanding of System Availability through the PV Fleet Performance Data Initiative

The PV Fleet Performance Data Initiative partners with photovoltaic (PV) fleet owners to collect time-series data of PV production data and publishes aggregated anonymized results of system performance metrics. With an extensive dataset drawn from over 2,200 PV systems across the United States, comprising 8.5 GW and 24,000 separate inverter data channels, this initiative aims to ensure that systemic risks in the US PV fleet are detected. The current work explores system availability, revealing a pronounced dependence on time, especially within the initial 6 months of system performance. Following this start-up period, the average system availability stabilizes. Statistical analyses illustrate a median (P5O) monthly availability of 0.991 and a dependence on system size with a negative trend in availability with increasing system size. This finding indicates that larger systems experience lower availability compared to their smaller counterparts.

inverter availability↗

Deep Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-The-Meter Solar

The ever-growing integration of distributed energy resources (DERs), especially behind-the-meter (BTM) solar generations, poses imperative operational challenges to system operators such as regional transmission organizations (RTOs). It is important for RTOs to effectively and accurately extract actual load profiles at the transmission level for a single node with significant BTM solar injection. This paper first illustrates the necessity of disaggregating the daily actual load profile of a single node. Furthermore, by segmenting nodes with selected timeseries features, nodes with significant BTM solar generation are identified. Lastly, a bi-level framework is proposed, comprising reference node disaggregation and DeepFM nodal disaggregation, aimed at disaggregating the nodal load profiles from which system operators require more information. By adopting a hybrid Deep Factorization Machine (DeepFM) model, the model achieve accurate results by extracting both linear and nonlinear relations between nodes in the same region and the zonal load and nodal load profile. To overcome the lack of ground truth, this paper segments the load profile into daytime, nighttime, and zero-crossing points and utilizes the latter two for evaluation purposes. The proposed disaggregation procedure is validated using real world, minute-level, normalized, and anonymized nodal data in the PJM service territory.

42 ENGINEERING↗

Contrasting student and staff perceptions of preclinical‐to‐clinical transition at a Chilean dental school

Abstract Introduction Dental education is a challenging and demanding field of study as students are expected to acquire various competencies to fulfil their professional requirements after graduation. The objective of this study was to investigate and compare dental students' and clinical staff instructors' perceptions of the preclinical‐to‐clinical transition training at a Dental School in Santiago, Chile. Material and Methods Two questionnaires containing 11 quantitative and one qualitative item were developed to assess our year three, four and five ( n = 244) dental undergraduate students' challenges when they begin treating patients, and clinical staff ( n = 78) perceptions of the preparedness to treat patients of the same students. Both questionnaires were voluntarily and anonymously implemented eight weeks after the beginning of the 2019 academic year. Responses were analysed using a Chi‐squared test for each quantitative question, while qualitative comments were studied to form themes and dimensions. RESULTS A total of 234 (96%) students and 60 (77%) instructors completed their respective questionnaire. There were considerable variations between students in the different years of the programme, as well as between students and staff members. Students and instructors felt the former had enough knowledge to treat patients though it was difficult for them to apply it in clinical practice. Again, both believed they could communicate with patients, but third year students asked for more training on this. Regarding practical skills, fourth‐ and fifth‐year students felt prepared but not third year students, who preferred to work in pairs with senior students, a preference that was shared by the instructors. All student groups asked clinical staff to provide more frequent, constructive and consistent feedback and felt that the difference between simulation and clinical environments and the amount of clinical work to fulfil clinical requirements made them feel stressed. Another mentioned stressor was students' low self‐confidence when working with patients. Among the requested improvements, students requested better training on how the dental clinic works to save time. Conclusions Preclinical‐to‐clinical transition training presents several challenges. Some of the problems highlighted by both students and clinical staff members persisted with the transition after three, four and even five years of training, which needs to be addressed.

Tricio, Jorge↗

Broadband, 920-nm mirror thin film damage competition

This year’s competition proposed to survey the state-of-the-art broadband, near-IR multilayer dielectric (MLD) mirrors designed for ultra-short, pulsed laser applications. The requirements for the coatings were a minimum reflection of 99.5% at 45-degree incidence angle for S-polarization from 830 nm to 1010 nm and group delay dispersion (GDD) < ± 50 fs 2 . The participants in this effort selected the coating materials, coating design, and deposition method. Samples were damage tested at a single testing facility to enable direct comparison among the participants using a 25 ± 5 fs OPCPA laser system operating at 5 Hz. A double blind test assured sample and submitter anonymity. The damage performance results, sample rankings, details of the deposition processes, coating materials and substrate cleaning methods are shared here. We found that multilayer coatings using tantala and/or hafnia as high index materials were top performers within several coating deposition groups. Specifically, dense coatings by ion-beam sputtering (IBS), magnetron sputtering (MS), and electron-beam ion assisted deposition (e-beam IAD) exhibited highest damage initiation onset (LIDT) while e-beam coatings were low performers. In addition, damage growth onset (LDGT) was also examined and the results are reported here for all samples as this performance metric plays an important role in establishing the safe operational conditions for larger aperture, ultrashort pulsed lasers. As a result, not all coating samples in the survey met the GDD requirements stated above and associated measurements are discussed in the context of the present and past competitions focused on similar broadband, near-IR MLD coatings.

42 ENGINEERING↗

Using AI tools to analyze periodic phenomena

Here, in this report, we describe how students can use natural language to prompt ChatGPT to conduct the analysis of complex periodic acceleration data collected using their smartphones. Students may use this approach to characterize human physiological tremor frequency. Students can choose to use their own data, or, for privacy purposes, use an anonymized set of data provided to them.

Klay, Jennifer L. [California Polytechnic State Un↗

Safe and Private Forward-trading Platform for Transactive Microgrids

Power grids are evolving at an unprecedented pace due to the rapid growth of distributed energy resources (DER) in communities. These resources are very different from traditional power sources, as they are located closer to loads and thus can significantly reduce transmission losses and carbon emissions. However, their intermittent and variable nature often results in spikes in the overall demand on distribution system operators (DSO). To manage these challenges, there has been a surge of interest in building decentralized control schemes, where a pool of DERs combined with energy storage devices can exchange energy locally to smooth fluctuations in net demand. Building a decentralized market for transactive microgrids is challenging, because even though a decentralized system provides resilience, it also must satisfy requirements such as privacy, efficiency, safety, and security, which are often in conflict with each other. As such, existing implementations of decentralized markets often focus on resilience and safety but compromise on privacy. In this article, we describe our platform, called TRANSAX, which enables participants to trade in an energy futures market, which improves efficiency by finding feasible matches for energy trades, enabling DSOs to plan their energy needs better. TRANSAX provides privacy to participants by anonymizing their trading activity using a distributed mixing service, while also enforcing constraints that limit trading activity based on safety requirements, such as keeping planned energy flow below line capacity. We show that TRANSAX can satisfy the seemingly conflicting requirements of efficiency, safety, and privacy. We also provide an analysis of how much trading efficiency is lost. Trading efficiency is improved through the problem formulation, which accounts for temporal flexibility, and system efficiency is improved using a hybrid-solver architecture. Lastly, we describe a testbed to run experiments and demonstrate its performance using simulation results.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

Location Generalizer

This software produces general location descriptions from time-series location data collected from mobile devices that record global positioning system (GPS) coordinates over time. The purpose of this software is to convert detailed location history data, which is considered personally identifiable information (PII), into non-traceable, anonymized, generic information that is useful to researchers but does not contain PII. The software is specifically designed for use with trigger-based data that describe the parked locations and dwell times of automobiles.

Smart, John↗

Building Performance Database API (BPD API) v2.1

The Building Performance Database (BPD) is the largest publicly-available source of measured energy performance data for buildings in the United States. It contains information about the building's energy use, location, and physical and operational characteristics. The BPD can be used by building owners, operators, architects and engineers to compare a building's energy efficiency against customized peer groups, identify energy efficiency opportunities, and set energy efficiency targets. It can also be used by energy efficiency program implementers and policymakers to analyze energy efficiency features and trends in the building stock. The BPD compiles data from various data sources, converts it into a standard format, cleanses and quality checks the data, and provides users with access to the data in a way that maintains anonymity for data providers. This software is the database and the Application Programming Interface (API). Users can utilize the BPD's data to develop their own applications using the API. Version 2.1 included a major update for multiple years of data and refactoring of code for faster queries.

Mathew, Paul↗

EyeON

EyeON: Eye on Operational technology Software Supply Chain attacks have risen drastically over the past few years, none more well-known and impactful than the SolarWinds compromise. Criminal organizations inserted an attack vector into a specific version of the source code, giving themselves an air of credibility. Once news broke on SolarWinds, identifying compromised sites was very difficult, even knowing the culprit update. Software Bills of Materials (SBOM) have been touted as the solution to reclaiming control of your software supply chain. Deployment of SBOMs has been slow, however, due to conflicting standards, opaque storage requirements, and vendor adoption. Additionally, the path from obtaining an SBOM and securing your supply chain is unclear; how can an SBOM library be leveraged to provide insight to your attack surface? The EyeON tool, sponsored by Department of Energy Cybersecurity, Energy Security, and Emergency Response (DoE CESER), aims to address these gaps by providing an encapsulated solution to tracking which updates have been installed in an enterprise, and alerting system administrators to vulnerabilities as they become known. Similar to a virus scanner, EyeON is a command line tool to parse either a single file, nested directory structure, or filesystem. It collects data such as signature (hashes), version information, VirusTotal tags, compiler, compilation date, and code signing information. Users will anonymously submit scan data periodically to DoE CESER, who will then compile a database of known software products employed by Critical Infrastructure and broadcast alerts based on discovered flaws as they arise.

Tenzing, Wangmo↗

Differentially Private Adaptive Noise Injection (DP-ANI) v1.0

Location data is collected from users continuously to understand their mobility patterns. Releasing the user trajectories may compromise user privacy. Therefore, the general practice is to release aggregated location datasets. However, private information may still be inferred from an aggregated version of location trajectories. Differential privacy (DP) protects the query output against inference attacks regardless of background knowledge. This software implements a differential privacy-based privacy model that protects the user's origins and destinations from being inferred from aggregated mobility datasets. This is achieved by injecting Planar Laplace noise to the user origin and destination GPS points. The noisy GPS points are then transformed into a link representation using a link-matching algorithm. Finally, the link trajectories form an aggregated mobility network. The injected noise level is selected using the Sparse Vector Mechanism. This DP selection mechanism considers the link density of the location and the functional category of the localized links. Compared to the different baseline models, including a k-anonymity method, our differential privacy-based aggregation model offers query responses that are close to the raw data in terms of aggregate statistics at both the network and trajectory-levels with maximum 9% deviation from the baseline in terms of network length.

Peisert, Sean [Lawrence Berkeley National Laborato↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Summit Darshan Archival Dataset

Summit Darshan Archival Dataset contains 2021 Summit Darshan log data for 25 applications and is grouped into science domains. The dataset is processed, and all the propriety fields are anonymized. The resultant data is converted into a tabular structure and saved in parquet file format. In this notebook, we demonstrate how to access the data. Data Organization: The data is organized into two directories: Darshan total (`darshan_total`): List all the high levels generated by the `darshan-parser --total` command on `.darshan` files. There is one parquet file for each application. Note: `uid` and `exe` field are masked Darshan detail (`darshan_detail`): This data contains detailed job level log information extracted by command `darshan-parser` on the raw `.darshan` files. The data is sorted by directory hierarchy in the order of `year/month/day (2021/12/07)`. For instance, to get the data for a `job_id` 3819766 of application `App11`, which was executed on `2021-12-07`can be accessed as follows. Note:`uid` and `filename` fields are masked

97 MATHEMATICS AND COMPUTING↗

Disproportionate impacts of COVID-19 in a large US city

COVID-19 has disproportionately impacted individuals depending on where they live and work, and based on their race, ethnicity, and socioeconomic status. Studies have documented catastrophic disparities at critical points throughout the pandemic, but have not yet systematically tracked their severity through time. Using anonymized hospitalization data from March 11, 2020 to June 1, 2021 and fine-grain infection hospitalization rates, we estimate the time-varying burden of COVID-19 by age group and ZIP code in Austin, Texas. During this 15-month period, we estimate an overall 23.7% (95% CrI: 22.5–24.8%) infection rate and 29.4% (95% CrI: 28.0–31.0%) case reporting rate. Individuals over 65 were less likely to be infected than younger age groups (11.2% [95% CrI: 10.3–12.0%] vs 25.1% [95% CrI: 23.7–26.4%]), but more likely to be hospitalized (1,965 per 100,000 vs 376 per 100,000) and have their infections reported (53% [95% CrI: 49–57%] vs 28% [95% CrI: 27–30%]). We used a mixed effect poisson regression model to estimate disparities in infection and reporting rates as a function of social vulnerability. We compared ZIP codes ranking in the 75th percentile of vulnerability to those in the 25th percentile, and found that the more vulnerable communities had 2.5 (95% CrI: 2.0–3.0) times the infection rate and only 70% (95% CrI: 60%-82%) the reporting rate compared to the less vulnerable communities. Inequality persisted but declined significantly over the 15-month study period. Our results suggest that further public health efforts are needed to mitigate local COVID-19 disparities and that the CDC’s social vulnerability index may serve as a reliable predictor of risk on a local scale when surveillance data are limited.

60 APPLIED LIFE SCIENCES↗

Rural EVSE Planning and Analysis

The dataset includes detailed anonymized public charging station usage from several rural stations on the ChargePoint and Shell Recharge Solutions (formerly Greenlots) networks situated in and around Athens, Ohio, a rural Appalachian community. Both Level 2 and DC fast charging stations are represented. Historical data in the set date back to 2019; additional data will be uploaded semiannually until the project's completion in 2023. Each charging session recorded includes information on date and time, location, charging station level, session duration, energy delivered, and fuel savings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EV Profile Capture

NextGen Profiles' EV profile capture efforts aimed to explore the variance in performance and evaluate how different operational conditions influence production EV charging behavior. Data were collected at a frequency of 10 Hz from both the EV and EVSE during each charge session. These charge session parameters were then entered into a time-series database for further analysis. The data were gathered under different operational conditions to examine the effects of various factors such as battery state of charge, battery temperature, vehicle condition, smart charge management, and EVSE limitations. The EV profile capture dataset includes extensive high-power charging data from 16 different EVs—comprising light-, medium-, and heavy-duty vehicles—along with EVSE from various suppliers. To protect confidentiality, the EV and EVSE metadata are anonymized, and the publicly released datasets are aggregated to 0.1-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EVSE Characterization

NextGen Profiles' EVSE characterization efforts explored performance variability in production EVSE through the use of EV emulation equipment and assessed how different operational conditions influence charging behavior. Data were collected at a frequency of 10 Hz from both the EV emulator and EVSE during each charge session and stored in a time-series database for further analysis. As part of the NextGen Profiles project, characterization of high-power EVSE was performed on both conductive and wireless charging infrastructure; however, only conductive charging data are currently included in this repository. This EVSE characterization was performed over a range of DC output currents and voltages, covering both nominal and off-nominal test conditions. This EVSE characterization dataset includes high-power charging data from two types of 350-kW-capable EVSE using liquid-cooled Combined Charging System-1 (CCS1, North American version) cables and connectors. To protect confidentiality, all EVSE metadata are anonymized, and the publicly released datasets are metered at 10-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗