Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Anonymization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-the-Meter Solar

The ever-growing high penetration of ubiquitously distributed energy resources, especially behind-the-meter solar (BTM) generations, has significant impacts on nodal load (i.e., net injection) profiles and consequently caused imperative operational challenges to system operators such as regional transmission organizations (RTOs). Illustrated by real-world nodal data and examples at PJM Interconnection, this paper first discusses the application and necessity of effectively extracting daily nodal load profiles in a non-intrusive manner. More importantly, a novel bi-level architecture, including Factorization Machines (FM) learning procedure has been proposed to effectively disaggregate not only one node but every node in an RTO service territory. Specifically, FM leaning is adopted to capture the interconnections between related features to better utilize the correlation between buses in the same region and between a single bus and the zonal load. The proposed bi-level technique is numerically validated using real-world, minute-level, normalized, and anonymized nodal data at PJM service territory.

behind the meter solar, load disaggregation, load ↗

Advancing Our Understanding of System Availability through the PV Fleet Performance Data Initiative

The PV Fleet Performance Data Initiative partners with photovoltaic (PV) fleet owners to collect time-series data of PV production data and publishes aggregated anonymized results of system performance metrics. With an extensive dataset drawn from over 2,200 PV systems across the United States, comprising 8.5 GW and 24,000 separate inverter data channels, this initiative aims to ensure that systemic risks in the US PV fleet are detected. The current work explores system availability, revealing a pronounced dependence on time, especially within the initial 6 months of system performance. Following this start-up period, the average system availability stabilizes. Statistical analyses illustrate a median (P5O) monthly availability of 0.991 and a dependence on system size with a negative trend in availability with increasing system size. This finding indicates that larger systems experience lower availability compared to their smaller counterparts.

inverter availability↗

Deep Factorization Machine Learning for Disaggregation of Transmission Load Profiles with High Penetration of Behind-The-Meter Solar

The ever-growing integration of distributed energy resources (DERs), especially behind-the-meter (BTM) solar generations, poses imperative operational challenges to system operators such as regional transmission organizations (RTOs). It is important for RTOs to effectively and accurately extract actual load profiles at the transmission level for a single node with significant BTM solar injection. This paper first illustrates the necessity of disaggregating the daily actual load profile of a single node. Furthermore, by segmenting nodes with selected timeseries features, nodes with significant BTM solar generation are identified. Lastly, a bi-level framework is proposed, comprising reference node disaggregation and DeepFM nodal disaggregation, aimed at disaggregating the nodal load profiles from which system operators require more information. By adopting a hybrid Deep Factorization Machine (DeepFM) model, the model achieve accurate results by extracting both linear and nonlinear relations between nodes in the same region and the zonal load and nodal load profile. To overcome the lack of ground truth, this paper segments the load profile into daytime, nighttime, and zero-crossing points and utilizes the latter two for evaluation purposes. The proposed disaggregation procedure is validated using real world, minute-level, normalized, and anonymized nodal data in the PJM service territory.

42 ENGINEERING↗

Contrasting student and staff perceptions of preclinical‐to‐clinical transition at a Chilean dental school

Abstract Introduction Dental education is a challenging and demanding field of study as students are expected to acquire various competencies to fulfil their professional requirements after graduation. The objective of this study was to investigate and compare dental students' and clinical staff instructors' perceptions of the preclinical‐to‐clinical transition training at a Dental School in Santiago, Chile. Material and Methods Two questionnaires containing 11 quantitative and one qualitative item were developed to assess our year three, four and five ( n = 244) dental undergraduate students' challenges when they begin treating patients, and clinical staff ( n = 78) perceptions of the preparedness to treat patients of the same students. Both questionnaires were voluntarily and anonymously implemented eight weeks after the beginning of the 2019 academic year. Responses were analysed using a Chi‐squared test for each quantitative question, while qualitative comments were studied to form themes and dimensions. RESULTS A total of 234 (96%) students and 60 (77%) instructors completed their respective questionnaire. There were considerable variations between students in the different years of the programme, as well as between students and staff members. Students and instructors felt the former had enough knowledge to treat patients though it was difficult for them to apply it in clinical practice. Again, both believed they could communicate with patients, but third year students asked for more training on this. Regarding practical skills, fourth‐ and fifth‐year students felt prepared but not third year students, who preferred to work in pairs with senior students, a preference that was shared by the instructors. All student groups asked clinical staff to provide more frequent, constructive and consistent feedback and felt that the difference between simulation and clinical environments and the amount of clinical work to fulfil clinical requirements made them feel stressed. Another mentioned stressor was students' low self‐confidence when working with patients. Among the requested improvements, students requested better training on how the dental clinic works to save time. Conclusions Preclinical‐to‐clinical transition training presents several challenges. Some of the problems highlighted by both students and clinical staff members persisted with the transition after three, four and even five years of training, which needs to be addressed.

Tricio, Jorge↗

Broadband, 920-nm mirror thin film damage competition

This year’s competition proposed to survey the state-of-the-art broadband, near-IR multilayer dielectric (MLD) mirrors designed for ultra-short, pulsed laser applications. The requirements for the coatings were a minimum reflection of 99.5% at 45-degree incidence angle for S-polarization from 830 nm to 1010 nm and group delay dispersion (GDD) < ± 50 fs 2 . The participants in this effort selected the coating materials, coating design, and deposition method. Samples were damage tested at a single testing facility to enable direct comparison among the participants using a 25 ± 5 fs OPCPA laser system operating at 5 Hz. A double blind test assured sample and submitter anonymity. The damage performance results, sample rankings, details of the deposition processes, coating materials and substrate cleaning methods are shared here. We found that multilayer coatings using tantala and/or hafnia as high index materials were top performers within several coating deposition groups. Specifically, dense coatings by ion-beam sputtering (IBS), magnetron sputtering (MS), and electron-beam ion assisted deposition (e-beam IAD) exhibited highest damage initiation onset (LIDT) while e-beam coatings were low performers. In addition, damage growth onset (LDGT) was also examined and the results are reported here for all samples as this performance metric plays an important role in establishing the safe operational conditions for larger aperture, ultrashort pulsed lasers. As a result, not all coating samples in the survey met the GDD requirements stated above and associated measurements are discussed in the context of the present and past competitions focused on similar broadband, near-IR MLD coatings.

42 ENGINEERING↗

Using AI tools to analyze periodic phenomena

Here, in this report, we describe how students can use natural language to prompt ChatGPT to conduct the analysis of complex periodic acceleration data collected using their smartphones. Students may use this approach to characterize human physiological tremor frequency. Students can choose to use their own data, or, for privacy purposes, use an anonymized set of data provided to them.

Klay, Jennifer L. [California Polytechnic State Un↗

Building Performance Database API (BPD API) v2.1

The Building Performance Database (BPD) is the largest publicly-available source of measured energy performance data for buildings in the United States. It contains information about the building's energy use, location, and physical and operational characteristics. The BPD can be used by building owners, operators, architects and engineers to compare a building's energy efficiency against customized peer groups, identify energy efficiency opportunities, and set energy efficiency targets. It can also be used by energy efficiency program implementers and policymakers to analyze energy efficiency features and trends in the building stock. The BPD compiles data from various data sources, converts it into a standard format, cleanses and quality checks the data, and provides users with access to the data in a way that maintains anonymity for data providers. This software is the database and the Application Programming Interface (API). Users can utilize the BPD's data to develop their own applications using the API. Version 2.1 included a major update for multiple years of data and refactoring of code for faster queries.

Mathew, Paul↗

EyeON

EyeON: Eye on Operational technology Software Supply Chain attacks have risen drastically over the past few years, none more well-known and impactful than the SolarWinds compromise. Criminal organizations inserted an attack vector into a specific version of the source code, giving themselves an air of credibility. Once news broke on SolarWinds, identifying compromised sites was very difficult, even knowing the culprit update. Software Bills of Materials (SBOM) have been touted as the solution to reclaiming control of your software supply chain. Deployment of SBOMs has been slow, however, due to conflicting standards, opaque storage requirements, and vendor adoption. Additionally, the path from obtaining an SBOM and securing your supply chain is unclear; how can an SBOM library be leveraged to provide insight to your attack surface? The EyeON tool, sponsored by Department of Energy Cybersecurity, Energy Security, and Emergency Response (DoE CESER), aims to address these gaps by providing an encapsulated solution to tracking which updates have been installed in an enterprise, and alerting system administrators to vulnerabilities as they become known. Similar to a virus scanner, EyeON is a command line tool to parse either a single file, nested directory structure, or filesystem. It collects data such as signature (hashes), version information, VirusTotal tags, compiler, compilation date, and code signing information. Users will anonymously submit scan data periodically to DoE CESER, who will then compile a database of known software products employed by Critical Infrastructure and broadcast alerts based on discovered flaws as they arise.

Tenzing, Wangmo↗

Differentially Private Adaptive Noise Injection (DP-ANI) v1.0

Location data is collected from users continuously to understand their mobility patterns. Releasing the user trajectories may compromise user privacy. Therefore, the general practice is to release aggregated location datasets. However, private information may still be inferred from an aggregated version of location trajectories. Differential privacy (DP) protects the query output against inference attacks regardless of background knowledge. This software implements a differential privacy-based privacy model that protects the user's origins and destinations from being inferred from aggregated mobility datasets. This is achieved by injecting Planar Laplace noise to the user origin and destination GPS points. The noisy GPS points are then transformed into a link representation using a link-matching algorithm. Finally, the link trajectories form an aggregated mobility network. The injected noise level is selected using the Sparse Vector Mechanism. This DP selection mechanism considers the link density of the location and the functional category of the localized links. Compared to the different baseline models, including a k-anonymity method, our differential privacy-based aggregation model offers query responses that are close to the raw data in terms of aggregate statistics at both the network and trajectory-levels with maximum 9% deviation from the baseline in terms of network length.

Peisert, Sean [Lawrence Berkeley National Laborato↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Summit Darshan Archival Dataset

Summit Darshan Archival Dataset contains 2021 Summit Darshan log data for 25 applications and is grouped into science domains. The dataset is processed, and all the propriety fields are anonymized. The resultant data is converted into a tabular structure and saved in parquet file format. In this notebook, we demonstrate how to access the data. Data Organization: The data is organized into two directories: Darshan total (`darshan_total`): List all the high levels generated by the `darshan-parser --total` command on `.darshan` files. There is one parquet file for each application. Note: `uid` and `exe` field are masked Darshan detail (`darshan_detail`): This data contains detailed job level log information extracted by command `darshan-parser` on the raw `.darshan` files. The data is sorted by directory hierarchy in the order of `year/month/day (2021/12/07)`. For instance, to get the data for a `job_id` 3819766 of application `App11`, which was executed on `2021-12-07`can be accessed as follows. Note:`uid` and `filename` fields are masked

97 MATHEMATICS AND COMPUTING↗

Disproportionate impacts of COVID-19 in a large US city

COVID-19 has disproportionately impacted individuals depending on where they live and work, and based on their race, ethnicity, and socioeconomic status. Studies have documented catastrophic disparities at critical points throughout the pandemic, but have not yet systematically tracked their severity through time. Using anonymized hospitalization data from March 11, 2020 to June 1, 2021 and fine-grain infection hospitalization rates, we estimate the time-varying burden of COVID-19 by age group and ZIP code in Austin, Texas. During this 15-month period, we estimate an overall 23.7% (95% CrI: 22.5–24.8%) infection rate and 29.4% (95% CrI: 28.0–31.0%) case reporting rate. Individuals over 65 were less likely to be infected than younger age groups (11.2% [95% CrI: 10.3–12.0%] vs 25.1% [95% CrI: 23.7–26.4%]), but more likely to be hospitalized (1,965 per 100,000 vs 376 per 100,000) and have their infections reported (53% [95% CrI: 49–57%] vs 28% [95% CrI: 27–30%]). We used a mixed effect poisson regression model to estimate disparities in infection and reporting rates as a function of social vulnerability. We compared ZIP codes ranking in the 75th percentile of vulnerability to those in the 25th percentile, and found that the more vulnerable communities had 2.5 (95% CrI: 2.0–3.0) times the infection rate and only 70% (95% CrI: 60%-82%) the reporting rate compared to the less vulnerable communities. Inequality persisted but declined significantly over the 15-month study period. Our results suggest that further public health efforts are needed to mitigate local COVID-19 disparities and that the CDC’s social vulnerability index may serve as a reliable predictor of risk on a local scale when surveillance data are limited.

60 APPLIED LIFE SCIENCES↗

Rural EVSE Planning and Analysis

The dataset includes detailed anonymized public charging station usage from several rural stations on the ChargePoint and Shell Recharge Solutions (formerly Greenlots) networks situated in and around Athens, Ohio, a rural Appalachian community. Both Level 2 and DC fast charging stations are represented. Historical data in the set date back to 2019; additional data will be uploaded semiannually until the project's completion in 2023. Each charging session recorded includes information on date and time, location, charging station level, session duration, energy delivered, and fuel savings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EV Profile Capture

NextGen Profiles' EV profile capture efforts aimed to explore the variance in performance and evaluate how different operational conditions influence production EV charging behavior. Data were collected at a frequency of 10 Hz from both the EV and EVSE during each charge session. These charge session parameters were then entered into a time-series database for further analysis. The data were gathered under different operational conditions to examine the effects of various factors such as battery state of charge, battery temperature, vehicle condition, smart charge management, and EVSE limitations. The EV profile capture dataset includes extensive high-power charging data from 16 different EVs—comprising light-, medium-, and heavy-duty vehicles—along with EVSE from various suppliers. To protect confidentiality, the EV and EVSE metadata are anonymized, and the publicly released datasets are aggregated to 0.1-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EVSE Characterization

NextGen Profiles' EVSE characterization efforts explored performance variability in production EVSE through the use of EV emulation equipment and assessed how different operational conditions influence charging behavior. Data were collected at a frequency of 10 Hz from both the EV emulator and EVSE during each charge session and stored in a time-series database for further analysis. As part of the NextGen Profiles project, characterization of high-power EVSE was performed on both conductive and wireless charging infrastructure; however, only conductive charging data are currently included in this repository. This EVSE characterization was performed over a range of DC output currents and voltages, covering both nominal and off-nominal test conditions. This EVSE characterization dataset includes high-power charging data from two types of 350-kW-capable EVSE using liquid-cooled Combined Charging System-1 (CCS1, North American version) cables and connectors. To protect confidentiality, all EVSE metadata are anonymized, and the publicly released datasets are metered at 10-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fleet Utilization

A key goal of NextGen Profiles' fleet utilization study was to conduct a comprehensive, strategic, and standardized assessment of the operational behavior and utilization patterns across EV and EVSE production-ready fleets. These data-driven insights were intended to inform current fleet management strategies and support future infrastructure planning, ensuring the effective adoption and adaptation of the growing EV fleet market. The study applied a series of metrics defined in NextGen Profiles to evaluate diverse fleet operations across various use cases, emphasizing trends in charging, routing, and other critical behaviors. The fleet utilization dataset includes these three sets of metrics from 17 EV fleets, each consisting of a wide range of vehicle types and operational categories, as well as two EVSE fleets. Data were collected from a variety of sources and reformatted into a unified structure before metric computation, ensuring consistency and comparability across all fleets. To protect confidentiality, all fleet metadata are anonymized, and the publicly released metric datasets are aggregated to an hourly cadence.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Chelation Modeling of a Plutonium-238 Inhalation Incident Treated with Delayed DTPA

This work describes an analysis, using a previously established chelation model, of the bioassay data collected from a worker who received delayed chelation therapy following a plutonium-238 inhalation. The details of the case have already been described in two publications. The individual was treated with Ca-DTPA via multiple intravenous injections and then nebulizations beginning several months after the intake and continuing for four years. The exact date and circumstances of the intake are unknown. However, interviews with the worker suggested that the intake occurred via inhalation of a soluble plutonium compound. The worker provided daily urine and fecal bioassay samples throughout the chelation treatment protocol, including samples collected before, during, and after the administration of Ca-DTPA. Unlike the previous two publications presenting this case, the current analysis explicitly models the combined biokinetics of the plutonium-DTPA chelate. Further, using the previously established chelation model, it was possible to fit the data through optimizing only the intake (day and magnitude), solubility, and absorbed fraction of nebulized Ca-DTPA. This work supports the hypothesis that the efficacy of the delayed chelation treatment observed in this case results mainly from chelation of cell-internalized plutonium by Ca-DTPA (intracellular chelation). It also demonstrates the validity of the previously established chelation model. As the bioassay data were modified to ensure data anonymization, the calculation of the “true” committed effective dose was not possible. However, the treatment-induced dose inhibition (in percentage) was calculated.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

AmeriFlux US-EA5 Uvalde Ranch Mesquite Woodland

This is the AmeriFlux version of the carbon flux data for the site US-EA5 Uvalde Ranch Mesquite Woodland. Site Description - This tower was located on a private ranch located approximately 25 km northwest of Uvalde, TX. The tower was installed on a trailer and situated amongst primarily mesquite trees. The lat/long provided here is approximate as the landowner wishes to maintain anonymity.

McKinney, Tyson↗