Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Extended experimental inferential structure determination method in determining the structural ensembles of disordered protein states

Proteins with intrinsic or unfolded state disorder comprise a new frontier in structural biology, requiring the characterization of diverse and dynamic structural ensembles. Here we introduce a comprehensive Bayesian framework, the Extended Experimental Inferential Structure Determination (X-EISD) method, which calculates the maximum log-likelihood of a disordered protein ensemble. X-EISD accounts for the uncertainties of a range of experimental data and back-calculation models from structures, including NMR chemical shifts, J-couplings, Nuclear Overhauser Effects (NOEs), paramagnetic relaxation enhancements (PREs), residual dipolar couplings (RDCs), hydrodynamic radii (R h ), single molecule fluorescence Förster resonance energy transfer (smFRET) and small angle X-ray scattering (SAXS). We apply X-EISD to the joint optimization against experimental data for the unfolded drkN SH3 domain and find that combining a local data type, such as chemical shifts or J-couplings, paired with long-ranged restraints such as NOEs, PREs or smFRET, yields structural ensembles in good agreement with all other data types if combined with representative IDP conformers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis of a Computational Framework for Bayesian Inverse Problems: Ensemble Kalman Updates and MAP Estimators under Mesh Refinement

This paper analyzes a popular computational framework to solve infinite-dimensional Bayesian inverse problems, discretizing the prior and the forward model in a finite-dimensional weighted inner product space. We demonstrate the benefit of working on a weighted space by establishing operator-norm bounds for finite element and graph-based discretizations of Matérn-type priors and deconvolution forward models. For linear-Gaussian inverse problems, we develop a general theory to characterize the error in the approximation to the posterior. We also embed the computational framework into ensemble Kalman methods and MAP estimators for nonlinear inverse problems. Furthermore, our operator-norm bounds for prior discretizations guarantee the scalability and accuracy of these algorithms under mesh refinement.

Bayesian inverse problem↗

Quantifying uncertainty for deep learning based forecasting and flow-reconstruction using neural architecture search ensembles

Classical problems in computational physics such as data-driven forecasting and signal reconstruction from sparse sensors have recently seen an explosion in deep neural network (DNN) based algorithmic approaches. However, most DNN models do not provide uncertainty estimates, which are crucial for establishing the trustworthiness of these techniques in downstream decision making tasks and scenarios. In recent years, ensemble-based methods have achieved significant success for the uncertainty quantification in DNNs on a number of benchmark problems. However, their performance on real-world applications remains under-explored. In this work, we present an automated approach to DNN discovery and demonstrate how this may also be utilized for ensemble-based uncertainty quantification. Specifically, we propose the use of a scalable neural and hyperparameter architecture search for discovering an ensemble of DNN models for complex dynamical systems. We highlight how the proposed method not only discovers high-performing neural network ensembles for our tasks, but also quantifies uncertainty seamlessly. This is achieved by using genetic algorithms and Bayesian optimization for sampling the search space of neural network architectures and hyperparameters. Subsequently, a model selection approach is used to identify candidate models for an ensemble set construction. Afterwards, a variance decomposition approach is used to estimate the uncertainty of the predictions from the ensemble. We demonstrate the feasibility of this framework for two tasks — forecasting from historical data and flow reconstruction from sparse sensors for the sea-surface temperature. In conclusion, we demonstrate superior performance from the ensemble in contrast with individual high-performing models and other benchmarks.

Deep ensembles↗

Protein folding from heterogeneous unfolded state revealed by time-resolved X-ray solution scattering

One of the most challenging tasks in biological science is to understand how a protein folds. In theoretical studies, the hypothesis adopting a funnel-like free-energy landscape has been recognized as a prominent scheme for explaining protein folding in views of both internal energy and conformational heterogeneity of a protein. Despite numerous experimental efforts, however, comprehensively studying protein folding with respect to its global conformational changes in conjunction with the heterogeneity has been elusive. Here we investigate the redox-coupled folding dynamics of equine heart cytochrome c (cyt-c) induced by external electron injection by using time-resolved X-ray solution scattering. A systematic kinetic analysis unveils a kinetic model for its folding with a stretched exponential behavior during the transition toward the folded state. With the aid of the ensemble optimization method combined with molecular dynamics simulations, we found that during the folding the heterogeneously populated ensemble of the unfolded state is converted to a narrowly populated ensemble of folded conformations. These observations obtained from the kinetic and the structural analyses of X-ray scattering data reveal that the folding dynamics of cyt-c accompanies many parallel pathways associated with the heterogeneously populated ensemble of unfolded conformations, resulting in the stretched exponential kinetics at room temperature. This finding provides direct evidence with a view to microscopic protein conformations that the cyt-c folding initiates from a highly heterogeneous unfolded state, passes through still diverse intermediate structures, and reaches structural homogeneity by arriving at the folded state.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Class imbalance in out-of-distribution datasets: Improving the robustness of the TextCNN for the classification of rare cancer types

In the last decade, the widespread adoption of electronic health record documentation has created huge opportunities for information mining. Natural language processing (NLP) techniques using machine and deep learning are becoming increasingly widespread for information extraction tasks from unstructured clinical notes. Disparities in performance when deploying machine learning models in the real world have recently received considerable attention. In the clinical NLP domain, the robustness of convolutional neural networks (CNNs) for classifying cancer pathology reports under natural distribution shifts remains understudied. In this research, we aim to quantify and improve the performance of the CNN for text classification on out-of-distribution (OOD) datasets resulting from the natural evolution of clinical text in pathology reports. We identified class imbalance due to different prevalence of cancer types as one of the sources of performance drop and analyzed the impact of previous methods for addressing class imbalance when deploying models in real-world domains. Our results show that our novel class-specialized ensemble technique outperforms other methods for the classification of rare cancer types in terms of macro F1 scores. We also found that traditional ensemble methods perform better in top classes, leading to higher micro F1 scores. Based on our findings, we formulate a series of recommendations for other ML practitioners on how to build robust models with extremely imbalanced datasets in biomedical NLP applications.

60 APPLIED LIFE SCIENCES↗

Large Scale Study of Ligand–Protein Relative Binding Free Energy Calculations: Actionable Predictions from Statistically Robust Protocols

The accurate and reliable prediction of protein–ligand binding affinities can play a central role in the drug discovery process as well as in personalized medicine. Of considerable importance during lead optimization are the alchemical free energy methods that furnish an estimation of relative binding free energies (RBFE) of similar molecules. Recent advances in these methods have increased their speed, accuracy, and precision. This is evident from the increasing number of retrospective as well as prospective studies employing them. However, such methods still have limited applicability in real-world scenarios due to a number of important yet unresolved issues. Here, we report the findings from a large data set comprising over 500 ligand transformations spanning over 300 ligands binding to a diverse set of 14 different protein targets which furnish statistically robust results on the accuracy, precision, and reproducibility of RBFE calculations. We use ensemble-based methods which are the only way to provide reliable uncertainty quantification given that the underlying molecular dynamics is chaotic. These are implemented using TIES (Thermodynamic Integration with Enhanced Sampling). Results achieve chemical accuracy in all cases. Ensemble simulations also furnish information on the statistical distributions of the free energy calculations which exhibit non-normal behavior. We find that the “enhanced sampling” method known as replica exchange with solute tempering degrades RBFE predictions. We also report definitively on numerous associated alchemical factors including the choice of ligand charge method, flexibility in ligand structure, and the size of the alchemical region including the number of atoms involved in transforming one ligand into another. Our findings provide a key set of recommendations that should be adopted for the reliable application of RBFE methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Solar Uncertainty Management and Mitigation for Exceptional Reliability in Grid Operations (SUMMER-GO): Project Final Report

The Solar Uncertainty Management and Mitigation for Exceptional Reliability in Grid Operation (SUMMER-GO) project was recently completed through a collaboration among the National Renewable Energy Laboratory, Maxar, the Electric Reliability Council of Texas (ERCOT), the University of Texas at Dallas, the University of California Berkeley, and the University of Colorado Boulder. The project made significant advances in probabilistic solar power forecasting, both through the development of Bayesian model averaging methods for ensemble forecasting and in bringing these and other advancements into practice with Maxar's delivery of operational forecasts to ERCOT. In addition to creating more reliable solar power forecasts, the project developed methods for their utilization in power system operations. These include the development of risk-aware unit commitment and economic dispatch algorithms and methods to reformulate probabilistic forecasts to be used in these power system operational models. Dynamic power system reserve methods were also developed, which have been shown in silico to create economic savings and reliability improvements on an ERCOT-like system as well as financial savings in the ERCOT system through more granular consideration of the uncertainty associated with solar power forecasts. Finally, a situational awareness tool to help grid operators better understand solar power forecast uncertainty in daily operations was developed and extensively vetted.

14 SOLAR ENERGY↗

Uncertainty quantification of mass models using ensemble Bayesian model averaging

Developments in the description of the masses of atomic nuclei have led to various nuclear mass models that provide predictions for masses across the whole chart of nuclides. These mass models play an important role in understanding the synthesis of heavy elements in the rapid neutron capture ( r ) process. However, it is still a challenging task to estimate the size of uncertainty associated with the predictions of each mass model. In this work, a method called ensemble Bayesian model averaging (EBMA) is introduced to quantify the uncertainty of one-neutron separation energies (S 1 n ) which are directly relevant in the calculations of r -process observables. Here, this Bayesian method provides a natural way to perform model averaging, selection, and uncertainty quantification, by combining the mass models as a mixture of normal distributions whose parameters are optimized against the experimental data, employing the Markov chain Monte Carlo method using the no-u-turn sampler. The EBMA model optimized with all the experimental S 1 n from the AME2003 nuclides are shown to provide reliable uncertainty estimates when tested with the new data in the AME2020.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ensemble variational Fokker-Planck methods for data assimilation

Particle flow filters solve Bayesian inference problems by smoothly transforming a set of particles into samples from the posterior distribution. Particles move in state space under the flow of an McKean-Vlasov-Itˆo process. This work introduces the Variational Fokker-Planck (VFP) framework for data assimilation, a general approach that includes previously known particle flow filters as special cases. The McKean-Vlasov-Itˆo process that transforms particles is defined via an optimal drift that depends on the selected diffusion term. It is established that the underlying probability density - sampled by the ensemble of particles - converges to the Bayesian posterior probability density. For a finite number of particles the optimal drift contains a regularization term that nudges particles toward becoming independent random variables. Based on this analysis, we derive computationally-feasible approximate regularization approaches that penalize the mutual information between pairs of particles, and avoid particle collapse. Moreover, the diffusion plays a role akin to a particle rejuvenation approach that aims to alleviate particle collapse. The VFP framework is very flexible. Different assumptions on prior and intermediate probability distributions can be used to implement the optimal drift, and localization and covariance shrinkage can be applied to alleviate the curse of dimensionality. A robust implicit-explicit method is discussed for the efficient integration of stiff McKean- Vlasov-Itˆo processes. Here, the effectiveness of the VFP framework is demonstrated on three progressively more challenging test problems, namely the Lorenz ’63, Lorenz ’96 and the quasi-geostrophic equations.

97 MATHEMATICS AND COMPUTING↗

Estimating Watershed Subsurface Permeability From Stream Discharge Data Using Deep Neural Networks

Subsurface permeability is a key parameter in watershed models that controls the contribution from the subsurface flow to stream flows. Since the permeability is difficult and expensive to measure directly at the spatial extent and resolution required by fully distributed watershed models, estimation through inverse modeling has had a long history in subsurface hydrology. The wide availability of stream surface flow data, compared to groundwater monitoring data, provides a new data source to infer soil and geologic properties using integrated surface and subsurface hydrologic models. As most of the existing methods have shown difficulty in dealing with highly nonlinear inverse problems, we explore the use of deep neural networks for inversion owing to their successes in mapping complex, highly nonlinear relationships. We train various deep neural network (DNN) models with different architectures to predict subsurface permeability from stream discharge hydrograph at the watershed outlet. The training data are obtained from ensemble simulations of hydrographs corresponding to an permeability ensemble using a fully-distributed, integrated surface-subsurface hydrologic model. The trained model is then applied to estimate the permeability of the real watershed using its observed hydrograph at the outlet. Our study demonstrates that the permeabilities of the soil and geologic facies that make significant contributions to the outlet discharge can be more accurately estimated from the discharge data. Their estimations are also more robust with observation errors. Compared to the traditional ensemble smoother method, DNNs show stronger performance in capturing the nonlinear relationship between permeability and stream hydrograph to accurately estimate permeability. Our study sheds new light on the value of the emerging deep learning methods in assisting integrated watershed modeling by improving parameter estimation, which will eventually reduce the uncertainty in predictive watershed models.

54 ENVIRONMENTAL SCIENCES↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Ensemble Simulations on Leadership Computing Systems

Scientific productivity can be enhanced through workflow management tools, relieving large High Performance Computing (HPC) system users from the tedious tasks of scheduling and designing the complex computational execution of scientific applications. This paper presents a study on the usage of ensemble workflow tools to accelerate science using the Summit and Frontier supercomputing systems. The research aims to connect science domain simulations using Oak Ridge Leadership Computing Facility (OLCF) supercomputing platforms with ensemble workflow methods in order to accelerate HPC-enabled discovery and boost scientific impact. We present the coupling, porting and optimization of Radical-Cybertools on three applications: Chroma, NAMD and LAMMPS. The tools augment traditional HPC monolithic runs with a pilot scheduler. Lessons-learned are discussed for physics, biology and materials science applications. We discuss intrinsic limitations of coupling and porting ensemble workflow tools to applications that run on large HPC systems. The origins of technical challenges and their solutions developed during the implementation process are discussed. Data management strategies, OLCF’s policies for ensembles, and natively supported workflow tools are also summarized.

Georgiadou, Antigoni [ORNL] (ORCID:000000020977631↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

An ensemble data assimilation modeling system for operational outdoor microalgae growth forecasting

Microalgae have received increasing attention as a potential feedstock for biofuel or biobased products. Forecasting the microalgae growth is beneficial for managers in planning pond operations and harvesting decisions. This study proposed a biomass forecasting system comprised of the Huesemann Algae Biomass Growth Model (BGM), the Modular Aquatic Simulation System in Two Dimensions (MASS2), ensemble data assimilation (DA), and numerical weather prediction Global Ensemble Forecast System (GEFS) ensemble meteorological forecasts. The novelty of this study is to seek the use of ensemble DA to improve both BGM and MASS2 model initial conditions with the assimilation of biomass and water temperature measurements and consequently improve short-term biomass forecasting skills. This study introduces the theory behind the proposed integrated biomass forecasting system, with an application undertaken in pseudo-real-time in three outdoor ponds cultured with Chlorella sorokiniana in Delhi, California, United States. Results from all three case studies demonstrate that the biomass forecasting system improved the short-term (i.e., 7-day) biomass forecasting skills by about 60% on average, comparing to forecasts without using the ensemble DA method. Given the satisfactory performances achieved in this study, it is probable that the integrated BGM-MASS2-DA forecasting system can be used operationally to inform managers in making pond operation and harvesting planning decisions.

59 BASIC BIOLOGICAL SCIENCES↗

Uncertainty quantification of a physics-informed model based on sparse identification of a Thermal Energy Distribution System

Integrated energy systems (IES)s are crucial for enhancing the economy and efficiency of power generation sources (e.g., nuclear energy) necessary to unleash American energy dominance. These systems can be integrated with thermal energy storage (TES) and intermittent renewable energies to optimize overall energy use, peak-load regulation, and demand-side responses. However, the stabilization of energy generation, transport, and utilization introduces operational complexities that exceed the challenges of managing each sub-component individually. Currently, though IESs rely on human operators for efficiency and stability, reducing human error risk and enhancing performance through automation is highly desirable. Recent advances at Idaho National Laboratory have demonstrated successful control of the Thermal Energy Distributed System (TEDS). However, the automatic control system depends on a deterministic Sparse Identification of Nonlinear Dynamics with Control (SINDyC) model, which are trained based on simulation data from physics-based simulations. Because of uncertainties in physics-based simulation, SINDyC model results in large discrepancies against experimental data and cannot be reliably used in automatic control. In this paper, we present an innovative approach to address these discrepancies by quantifying uncertainties and developing a more robust model. We first generated trajectories by using first-principles physics codes to encapsulate the experiment. Next, we trained thousands of models by randomly sampling these trajectories. We then collapsed all those models into one probabilistic SINDyC by fitting a multivariate Gaussian distribution onto the resulting coefficient’s distribution. Despite its simplicity, our approach successfully produced 95% confidence intervals that captured the experimental trajectories. It even did so with a higher probability and better U-pooling score across six of the seven relevant quantities of interest (QoIs), as compared to other classical approaches. In conclusion, ongoing research is focusing on generating new experimental trajectories to validate this approach, and on employing Bayesian calibration to refine parametric uncertainties and guide future model development efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Ensemble Monte Carlo calculations with five novel moves

We introduce five novel types of Monte Carlo (MC) moves that brings the number of moves of ensemble MC calculations from three to eight. So far such calculations have relied on affine invariant stretch moves that were originally introduced by Christen (2007), walk moves by Goodman and Weare (2010) and quadratic moves by Militzer (2023). Ensemble MC methods have been very popular because they harness information about the fitness landscape from a population of walkers rather than relying on expert knowledge. Here we modified the affine method and employed a simplex of points to set the stretch direction. We adopt the simplex concept to quadratic moves. We also generalize quadratic moves to arbitrary order. Finally, we introduce directed moves that employ the values of the probability density while all other types of moves rely solely on the location of the walkers. We apply all algorithms to the Rosenbrock density in 2 and 20 dimensions and to the ring potential in 12 and 24 dimensions. We evaluate their efficiency by comparing error bars, autocorrelation time, travel time, and the level of cohesion that measures whether any walkers were left behind. Our code is open source.

97 MATHEMATICS AND COMPUTING↗

The sensitive surface chemistry of Co-free, Ni-rich layered oxides: identifying experimental conditions that influence characterization results

Recent studies have suggested that Co-free, Ni-rich layered cathodes (e.g., doped LiNiO2) can provide promising battery performance for practical applications. However, these layered cathodes suffer from significant surface instability during various stages of the sample history, which generates inherent challenges for achieving stable battery performance and obtaining statistically representative characterization results. To reliably report the surface chemistry of these materials, delicate controls of stepwise sample preparation are required. In this study, we aim to reveal how the surface chemistry of LiNiO2 based materials changes with various environments, including human exhalation, sample storage, sample preparation, electrochemistry cycling, and surface doping. Our results demonstrate that the surface of these materials is highly reactive and prone to alter at various stages of sample handling and characterization. The sensitive surface could impact the interpretation of the surface chemical and structural information, including surface carbonate formation, transition metal reduction and dissolution, and surface reconstruction. Importantly, the heterogeneity of the surface degradation calls for a consolidation of nanoscale, high-resolution characterization, and ensemble-averaged methods in order to improve statistical representation. Furthermore, the doping chemistry can effectively mitigate the surface degradation and improve overall battery performance due to the enhanced surface oxygen retention. Our study highlights the necessity of strict measurements through complementary characterizations at multiple length scales to eliminate unintentional biased conclusions.

surface chemistry, Co-free Ni-rich cathodes, Istab↗