Leveraging Battery Archive to Evaluate Cycle Life Prediction Methods Across a Variety of Datasets
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Abstract. Atmospheric observations in remote locations offer a possibility of exploring trace gas and particle concentrations in pristine environments. However, data from remote areas are often contaminated by pollution from local sources. Detecting this contamination is thus a central and frequently encountered issue. Consequently, many different methods exist today to identify local contamination in atmospheric composition measurement time series, but no single method has been widely accepted. In this study, we present a new method to identify primary pollution in remote atmospheric datasets, e.g., from ship campaigns or stations with a low background signal compared to the contaminated signal. The pollution detection algorithm (PDA) identifies and flags periods of polluted data in five steps. The first and most important step identifies polluted periods based on the derivative (time derivative) of a concentration over time. If this derivative exceeds a given threshold, data are flagged as polluted. Further pollution identification steps are a simple concentration threshold filter, a neighboring points filter (optional), a median, and a sparse data filter (optional). The PDA only relies on the target dataset itself and is independent of ancillary datasets such as meteorological variables. All parameters of each step are adjustable so that the PDA can be “tuned” to be more or less stringent (e.g., flag more or fewer data points as contaminated). The PDA was developed and tested with a particle number concentration dataset collected during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition in the central Arctic. Using strict settings, we identified 62 % of the data as influenced by local contamination. Using a second independent particle number concentration dataset also collected during MOSAiC, we evaluated the performance of the PDA against the same dataset cleaned by visual inspection. The two methods agreed in 94 % of the cases. Additionally, the PDA was successfully applied to a trace gas dataset (CO2), also collected during MOSAiC, and to another particle number concentration dataset, collected at the high-altitude background station Jungfraujoch, Switzerland. Thus, the PDA proves to be a useful and flexible tool to identify periods affected by local contamination in atmospheric composition datasets without the need for ancillary measurements. It is best applied to data representing primary pollution. The user-friendly and open-access code enables reproducible application to a wide suite of different datasets. It is available at https://doi.org/10.5281/zenodo.5761101 (Beck et al., 2021).
Electricity generation in the United States entails significant water usage and greenhouse gas emissions. However, accurately estimating these impacts is complex due to the intricate nature of the electric grid and the dynamic electricity mix. Existing methods to estimate the environmental consequences of electricity use often generalize across large regions, neglecting spatial and temporal variations in water usage and emissions. Consequently, electric grid dynamics, such as temporal fluctuations in renewable energy resources, are often overlooked in efforts to mitigate environmental impacts. The U.S. Department of Energy (DOE) has initiated the development of resilient energyshed management systems, requiring detailed information on the local electricity mix and its environmental impacts. This study supports DOE's goal by incorporating geographic and temporal variations in the electricity mix of the local electric grid to better understand the environmental impacts of electricity end users. We offer hourly estimates of the U.S. electricity mix, detailing fuel types, water withdrawal intensity, and water consumption intensity for each grid balancing authority through our publicly accessible tool, the Water Integrated Mapping of Power and Carbon Tracker (Water IMPACT). While our primary focus is on evaluating water intensity factors, our dataset and programming scripts for historical and real-time analysis also include evaluations of carbon dioxide (equivalence) intensity within the same modeling framework. This integrated approach offers a comprehensive understanding of the environmental footprint associated with electricity generation and use, enabling informed decision-making to effectively reduce Scope 2 water usage and emissions.
Abstract. Gaussian process (GP) regression is a flexible modeling technique used to predict outputs and to capture uncertainty in the predictions. However, the GP regression process becomes computationally intensive when the training spatial dataset has a large number of observations. To address this challenge, we introduce a scalable GP algorithm, termed MuyGPs, which incorporates nearest-neighbor and leave-one-out cross-validation during training. This approach enables the evaluation of large spatial datasets with state-of-the-art accuracy and speed in certain spatial problems. Despite these advantages, conventional quadratic loss functions used in the MuyGPs optimization, such as root mean squared error (RMSE), are highly influenced by outliers. We explore the behavior of MuyGPs in cases involving outlying observations and, subsequently, develop a robust approach to handle and mitigate their impact. Specifically, we introduce a novel leave-one-out loss function based on the pseudo-Huber function (LOOPH) that effectively accounts for outliers in large spatial datasets within the MuyGPs framework. Our simulation study shows that the LOOPH loss method maintains accuracy despite outlying observations, establishing MuyGPs as a powerful tool for mitigating unusual observation impacts in the large data regime. In the analysis of US ozone data, MuyGPs provides accurate predictions and uncertainty quantification, demonstrating its utility in managing data anomalies. Through these efforts, we advance the understanding of GP regression in spatial contexts.
Developing methods to understand and control defect formation in nanomaterials offers a promising route for materials discovery. Monolayer MX 2 phases represent a particularly compelling case for defect engineering of nanomaterials due to the large variability in their physical properties as different defects are introduced into their structure. However, effective identification and quantification of defects remain a challenge even as high-throughput scanning transmission electron microscopy methods improve. This study highlights the benefits of employing first principles calculations to produce digital twins for training deep learning segmentation models for defect identification in monolayer MX 2 phases. Around 600 defect structures were obtained using density functional theory calculations, with each monolayer MX 2 structure being subjected to multislice simulations for the purpose of generating the digital twins. Several deep learning segmentation architectures were trained on this dataset, and their performances evaluated under a variety of conditions such as recognizing defects in the presence of unidentified impurities, beam damage, grain boundaries, and with reduced image quality from low electron doses. Further, this digital twin approach allows benchmarking different deep learning architectures on a theory dataset, which enables the study of defect classification under a broad array of finely controlled conditions. It thus opens the door to resolving the underpinning physical reasons for model shortcomings and potentially chart paths forward for automated discovery of materials defect phases in experiments.
Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.
The SATELLITE_EAGLES_PNNL NetCDF dataset contains a suite of satellite- and reanalysis-derived atmospheric and surface parameters on a regular latitude–longitude grid. The dataset includes core geophysical fields such as land fraction, aerosol optical depth at multiple wavelengths (465, 550, 667, and 865 nm), sea surface temperature, estimated inversion strength, and various thermodynamic and dynamic quantities (e.g., relative humidity, vertical velocity, boundary-layer height, and surface fluxes) from both MERRA and ERA reanalysis products, provided as daily-mean and instantaneous values. A major component of the dataset consists of MODIS-retrieved cloud microphysical properties, including cloud droplet number concentration, cloud effective radius, optical thickness, and liquid water path, provided for three compositing regimes (“All,” “Q06,” and “G18”). Corresponding cloud-top parameters—temperature, height, and pressure—along with total and domain-mean cloud fraction fields are also included. The file further integrates additional satellite data from AMSR-E (for cloud water, rain water, and surface precipitation retrievals) and CERES (for top-of-atmosphere radiative fluxes, cloud fractions, and albedo). This dataset is designed to evaluate aerosol–cloud interactions in warm clouds, emphasizing the use of MODIS for deriving cloud droplet number concentration and liquid water path statistics. The complementary satellite and reanalysis fields are co-located and time-matched to the same instantaneous MODIS observations, enabling consistent comparisons between cloud properties, aerosol loading, and large-scale meteorological conditions. The dataset is recently featured in Christensen et al. (2025), Machine Learning Reveals Strong Grid-Scale Dependence in the Satellite Nd–LWP Relationship, EGUsphere [preprint], https://doi.org/10.5194/egusphere-2025-3850, 2025.
The SATELLITE_EAGLES_PNNL NetCDF dataset contains a suite of satellite- and reanalysis-derived atmospheric and surface parameters on a regular latitude–longitude grid. The dataset includes core geophysical fields such as land fraction, aerosol optical depth at multiple wavelengths (465, 550, 667, and 865 nm), sea surface temperature, estimated inversion strength, and various thermodynamic and dynamic quantities (e.g., relative humidity, vertical velocity, boundary-layer height, and surface fluxes) from both MERRA and ERA reanalysis products, provided as daily-mean and instantaneous values. A major component of the dataset consists of MODIS-retrieved cloud microphysical properties, including cloud droplet number concentration, cloud effective radius, optical thickness, and liquid water path, provided for three compositing regimes (“All,” “Q06,” and “G18”). Corresponding cloud-top parameters—temperature, height, and pressure—along with total and domain-mean cloud fraction fields are also included. The file further integrates additional satellite data from AMSR-E (for cloud water, rain water, and surface precipitation retrievals) and CERES (for top-of-atmosphere radiative fluxes, cloud fractions, and albedo). This dataset is designed to evaluate aerosol–cloud interactions in warm clouds, emphasizing the use of MODIS for deriving cloud droplet number concentration and liquid water path statistics. The complementary satellite and reanalysis fields are co-located and time-matched to the same instantaneous MODIS observations, enabling consistent comparisons between cloud properties, aerosol loading, and large-scale meteorological conditions. The dataset is recently featured in Christensen et al. (2025), Machine Learning Reveals Strong Grid-Scale Dependence in the Satellite Nd–LWP Relationship, EGUsphere [preprint], https://doi.org/10.5194/egusphere-2025-3850, 2025.
As the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD).
A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.
Bias in medical image segmentation can lead to unequal performance across demographic subgroups, raising concerns about fairness and reliability in clinical AI systems. While deep learning models have achieved high segmentation accuracy, ensuring equitable performance across race and gender remains a significant challenge, particularly in privacy-sensitive healthcare environments. This study investigates fairness-aware medical image segmentation for hip and knee radiographs using deep learning models evaluated in both centralized and Federated Learning (FL) settings. We introduce Curriculum Learning (CL) strategies and Progressive Loss (PL) functions to regulate sample difficulty during training. In addition, we propose two novel fairness-oriented federated learning algorithms, Federated Intersection over Union (FedIoU) and Federated Intersection over Union with Outlier Analysis (FedIoUoutlier). Experiments are conducted using multiple segmentation backbones and simulated multi-site data partitions derived from the Osteoarthritis Initiative dataset. Model performance is evaluated using Intersection over Union (IoU), IoU standard deviation, Skewed Error Ratio (SER), and Min-Max Disparity across race and gender subgroups. Statistical significance was verified using paired t-tests to compare per-sample IoU performance against baseline configurations. Across both hip and knee segmentation tasks, curriculum learning and progressive loss strategies consistently improved segmentation accuracy and reduced demographic performance disparities in centralized training. In federated settings, fairness-aware aggregation further enhanced performance. Notably, FedIoUoutlier combined with balanced curriculum learning and tiered progressive loss achieved the highest mean IoU while yielding the lowest SER and Min-Max Disparity, indicating improved fairness without sacrificing accuracy. In several configurations, federated models matched or exceeded the performance of optimized centralized models, with statistically significant improvements in per-sample IoU over baseline configurations. The results demonstrate that structured training strategies and fairness-aware federated aggregation can jointly improve accuracy, stability, and demographic fairness in medical image segmentation. By integrating curriculum learning, progressive loss, and novel FL algorithms, this work provides a practical pathway toward equitable and privacy-preserving AI systems for medical imaging.
Conventional differential privacy (DP) assumes record independence, limiting effectiveness on real-world datasets with temporal, spatial, or structural correlations. These dependencies undermine privacy guarantees and degrade utility in domains like healthcare, IoT, and smart city analytics. We propose Maximum Information Correlated Differential Privacy (MIC-DP), a novel framework that dynamically calibrates noise based on statistical dependencies. MIC-DP uses the Maximum Information Coefficient (MIC) to capture both linear and nonlinear correlations without explicit modeling, enabling adaptive sensitivity adjustment and improved privacy–utility trade-offs. Evaluations on healthcare (MIMIC), demographic (ACI), and synthetic datasets show that MIC-DP reduces mean absolute error (MAE) by up to 5.2% under strict privacy budgets (ϵ≤1), with aggregate utility improvements reaching 18% across datasets and evaluation metrics. MIC-DP provides formal (ϵ,δ)-privacy guarantees, scales efficiently with feature count, and supports deployment in moderate-scale, privacy-sensitive applications. Its tunable performance and runtime efficiency make MIC-DP suitable for privacy-sensitive applications where low-latency analytics and strong privacy guarantees must coexist. These results demonstrate MIC-DP’s effectiveness as a correlation-aware solution for practical DP.
Here, we develop a ytopt autotuning framework that leverages Bayesian optimization to explore the parameter space search and compare four different supervised learning methods within Bayesian optimization and evaluate their effectiveness. We select six of the most complex PolyBench benchmarks and apply the newly developed LLVM Clang/Polly loop optimization pragmas to the benchmarks to optimize them. We then use the autotuning framework to optimize the pragma parameters to improve their performance. The experimental results show that our autotuning approach outperforms the other compiling methods to provide the smallest execution time for the benchmarks syr2k, 3mm, heat-3d, lu, and covariance with two large datasets in 200 code evaluations for effectively searching the parameter spaces with up to 170,368 different configurations. We find that the Floyd-Warshall benchmark did not benefit from autotuning. To cope with this issue, we provide some compiler option solutions to improve the performance. Then we present loop autotuning without a user's knowledge using a simple mctree autotuning framework to further improve the performance of the Floyd-Warshall benchmark. We also extend the ytopt autotuning framework to tune a deep learning application.
Modern network devices collect a large amount of data that can be analyzed to identify bottlenecks, anomalies, cyber-attacks, etc. Therefore, there is often a need to analyze such collections of network data quite often by an external expert or by the research community. However, these collections of data contain sensitive, proprietary information. In order for the network data to be shared, it must first be anonymized. The overall objective of this project is to develop an innovative privacy management tool to anonymize network data and achieve sufficient privacy, acceptable data utility, and efficient data analysis at the same time. No existing anonymization methods can achieve all of these at the same time. The core of this technology is a differential private clustering algorithm that provides strong privacy protection, preserves data properties important for subsequent analysis, and allows the party receiving the anonymized data to conduct analysis directly on anonymized data without the need of decryption or any extra processing. The research carried out was to design, implement and verify a solution to this problem by completing the following tasks: 1) developing the core technology; 2) developing a context based method that automatically recommends fields that must be anonymized; 3) conducted experiments showing superior results using our approach compared to existing tools, and 4) developed an intuitive but basic user interface. The research that was conducted generated novel algorithmic techniques that utilize state-of-the-art methods such as condensation, differential privacy preservation, clustering, automated tuning based on contextual awareness, and recommendation techniques to specify columns to users for anonymization leading to optimal privacy that allows research analysis on the dataset. Experiments were conducted to evaluate the efficacy of these novel algorithmic techniques by performing analysis on original non-anonymized datasets, then conducting analysis on the same yet anonymized datasets and comparing the results of the analyses. Overall, the anonymized analysis results were within 1% of the original results, verifying that the generated technology not only guarantees a high level of privacy but also enables research analysis as if it were conducted on the original dataset. Potential applications of this technology include anonymization of any type of structured network datasets that contain sensitive identifiers, such as IP addresses, that can be used in multiple applications. For example, to create an AI or machine learning model for cyber security, e.g., to detect attacks, or for performance analysis, e.g., identify bottlenecks or predict performance. In addition, a market analysis that was conducted for potential applications of this technology identified a broader range of applications of our anonymization technology beyond the network sector that includes healthcare, banking, insurance, securities, finance (FISB), data brokering, cloud services, ad sales, and government.
The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.
Federated Learning (FL) has amassed notable attention for its ability to preserve user privacy while emphasizing the retainment of model training efficiency. Due to this potential, FL has been integrated in many domains, such as healthcare, finance, law, and industrial engineering, where data cannot be easily exchanged due to sensitive information and strict privacy laws. However, current research has indicated that FL protocols are easily compromised by active data reconstruction attacks employed by actively dishonest servers. The malicious modification of global model parameters allows an actively dishonest server to obtain a direct copy of users’ private data via gradient inversion. Here, this class of attacks is highly underexplored and continues to be a major challenge due to the intense threat model. In this paper, we propose OASIS as a scalable and modality-agnostic defense based on data augmentation that counteracts active data reconstruction attacks while preserving model performance. To generalize our defense, we uncover the intuition behind gradient inversion that enables these attacks and theoretically establish the conditions by which the defense can be considered robust regardless of attack design. From this, we formulate our defense with data augmentation that illustrates its ability to undermine the attack principle. We evaluate OASIS on five real-world datasets–two image-based (ImageNet and CIFAR100) and three text-based (Wikitext, Stack Overflow, and Shakespeare)–which span diverse uses cases such as vision tasks and language modeling. Comprehensive evaluations on these datasets exhibit the efficacy of OASIS and highlight its feasibility as a solution.
In this paper, a framework is outlined to generate realistic artificial data (RAD) as a tool for comparing different models developed for safety analysis. The primary focus of transportation safety analysis is on identifying and quantifying the influence of factors contributing to traffic crash occurrence and its consequences. The current framework of comparing model structures using only observed data has limitations. With observed data, it is not possible to know how well the models mimic the true relationship between the dependent and independent variables. Further, real datasets do not allow researchers to evaluate the model performance for different levels of complexity of the dataset. RAD offers an innovative framework to address these limitations. Hence, we propose a RAD generation framework embedded with heterogeneous causal structures that generates crash data by considering crash occurrence as a trip level event impacted by trip level factors, demographics, roadway and vehicle attributes. Within our RAD generator we employ three specific modules: (a) disaggregate trip information generation, (b) crash data generation and (c) crash data aggregation. For disaggregate trip information generation, we employ a daily activity-travel realization for an urban region generated from an established activity-based model for the Chicago region. We use this data of more than 2 million daily trips to generate a subset of trips with crash data. For trips with crashes crash location, crash type, driver/vehicle characteristics, and crash severity. The daily RAD generation process is repeated for generating crash records at yearly or multi-year resolution. In conclusion, the crash databases generated can be employed to compare frequency models, severity models, crash type and various other dimensions by facility type - possibly establishing a universal benchmarking system for alternative model frameworks in safety literature.
Abstract Land surface models such as the Community Land Model Version 5 (CLM5) are essential tools for simulating the behavior of the terrestrial system. Despite the extensive application of CLM5, limited attention has been paid to the underlying uncertainties associated with its hydrological parameters and how these uncertainties affect water resource applications. To address this long-standing issue, we use five meteorological datasets to conduct a comprehensive hydrological parameter uncertainty characterization of CLM5 over the hydroclimatic gradients of the conterminous United States. Key datasets produced from the uncertainty characterization experiment include: a benchmark dataset of CLM5 default hydrological performance, parameter sensitivities for 28 hydrological metrics, and large-ensemble outputs for CLM5 hydrological predictions. The presented datasets will assist CLM5 calibration and support broad applications, such as evaluating drought and flood vulnerabilities. The datasets can be used to identify the hydroclimatological conditions under which parametric uncertainties demonstrate substantial effects on hydrological predictions and clarify where further investigations are needed to understand how hydrological prediction uncertainties interact with other Earth system processes.