Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Machine Learning for Mapping Multipactor Susceptibility in RF Systems: Capabilities and Generalization Constraints

Multipactor is a surface-driven electron avalanche phenomenon that degrades the performance and reliability of radio-frequency (RF) systems in particle accelerator and vacuum electronics applications. Multipactor behavior in a given device structure is conventionally assessed through susceptibility charts, which provide a parameter-space characterization of the instability. In this work, we assess the capabilities of machine-learning (ML) models to learn and predict such susceptibility charts and analyze the constraints governing their generalization across materials. Using a simulation-derived dataset spanning six distinct secondary-electron-yield material profiles in a canonical two-surface planar geometry, we train supervised regression models and artificial neural networks to predict the time-averaged electron growth rate, δavg, across the relevant parameter space. Model performance is evaluated using metrics that explicitly probe the structure of susceptibility charts, including Intersection over Union, Structural Similarity Index, and correlation analysis. Tree-based ensemble models outperform neural-network models in reconstructing susceptibility regions and in generalizing across material domains. Principal-component analysis reveals disjoint material feature distributions, indicating that the piecewise mode structure of multipactor susceptibility is difficult to represent with a single global model and that generalization is constrained by data coverage rather than by model complexity. An exhaustive reduced-coverage study further shows that sparse material-space coverage can yield mean performance in the same general range but producing large variability in the susceptibility-region overlap. These results clarify the capabilities of ML-based surrogate models for parameter-space characterization of multipactor discharge. They also provide guidance for their appropriate use in RF system design.

43 PARTICLE ACCELERATORS↗

Combining PCA and nonlinear fitting of peak models to re-evaluate C 1s XPS spectrum of cellulose

Cellulose is an example of a material that responds to XPS by the creation of new chemistry not present in the as-received sample. While improvements in instrumentation may be seen in general as beneficial to surface science, recent studies have shown that the consequences for some materials are detrimental. Here, in this work, these problems are illustrated through an analysis of cellulose spectra obtained during a degradation study. C 1s spectra are decomposed into two well-formed component curves that are open to chemical interpretation. In particular, a component-curve representative of pure cellulose is obtained as well as a second component curve that implies cellulose is degraded through the creation of carbon chemistry involving C—O, C$=$O and O—C$=$O. Since cellulose is a crystalline material, formed through the alignment of molecules under the influence of hydrogen bonds, the analysis and findings presented in this paper are relevant to any material analyzed by XPS whose properties are dependent on hydrogen bonds. The analysis techniques are based on an informed vectorial approach, which extracts directly from data spectral shapes that are used to monitor sample degradation via linear least squares optimization. Related mathematics of Principal Component Analysis and linear analysis are presented.

36 MATERIALS SCIENCE↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗

Enhancing physicochemical, bioactive, and nutritional properties of sweet potatoes: Ultrasonic contact drying with slot jet nozzles compared to hot-air drying and freeze drying

Sweet potatoes are a rich source of nutrients and bioactive compounds, but their quality can be impacted by the drying process. This study investigates the impact of slot jet reattachment (SJR) nozzle and ultrasound (US) combined drying (SJR + US) on sweet potato quality, compared to freeze-drying (FD), SJR drying, and hot air drying (HAD). SJR + US drying at 50 °C closely resembled FD in enhancing quality attributes and outperformed HAD and SJR in key areas such as rehydration, shrinkage ratios, and nutritional composition. Notably, SJR + US at 50 °C produced the highest total starch (36.84 g/100 g), total dietary fiber (8.48 g/100 g), total phenolic content (158.19 mg GAE/100 g), total flavonoid content (119.08 mg QE/g), DPPH antioxidant activity (6.44 μmol TE/g), β-carotene (31.98 mg/100 g), and vitamin C (5.27 mg/100 g). It also exhibited higher glass transition temperatures (Tg: 14.49 °C), indicating better stability at room temperature. The hardness values for SJR + US samples were similar to FD, while HAD samples had the highest hardness. SJR + US at 50 °C resulted in the lowest total color changes (ΔE), indicating minimal impact on appearance. Additionally, FTIR analysis revealed that peaks in specific spectral regions indicated superior preservation of bioactive compounds in SJR + US samples compared to other methods, which was also confirmed by principal component analysis (PCA) and heatmap visualization. Overall, these findings suggest that SJR + US is an effective alternative to conventional drying techniques, significantly improving the quality of dried sweet potatoes.

Color↗

Lipid droplet-associated proteins in alcohol-associated fatty liver disease: A proteomic approach

The earliest manifestation of alcohol-associated liver disease (ALD) is steatosis characterized by deposition of fat in specialized organelles called lipid droplets (LDs). While alcohol administration causes a rise in LD numbers in the hepatocytes, little is known regarding their characteristics that allow their accumulation and size to increase. The aim of the present study is to gain insights into underlying pathophysiological mechanisms by investigating the ethanol-induced changes in hepatic LD proteome as a function of LD size. Adult male Wistar rats (180–200 g BW) were fed with ethanol liquid diet for 6 weeks. At sacrifice, large-, medium-, and small-sized hepatic LD subpopulations (LD1, LD2, and LD3, respectively) were isolated and subjected to morphological and proteomic analyses. Morphological analysis of LD1-LD3 fractions of ethanol-fed rats clearly demonstrated that LD1 contained larger LDs compared with LD2 and LD3 fractions. Our preliminary results from principal component analysis showed that the proteome of different-sized hepatic LD fractions was distinctly different. Proteomic data analysis identified over 2000 proteins in each LD fraction with significant alterations in protein abundance among the three LD fractions. Among the altered proteins, several were related to fat metabolism, including synthesis, incorporation of fatty acid, and lipolysis. Ingenuity pathway analysis revealed increased fatty acid synthesis, fatty acid incorporation, LD fusion, and reduced lipolysis in LD1 compared to LD3. Overall, the proteomic findings indicate that the increased level of protein that facilitates fusion of LDs combined with an increased association of negative regulators of lipolysis dictates the generation of large-sized LDs during the development of alcohol-associated hepatic steatosis. Several significantly altered proteins were identified in different-sized LDs isolated from livers of ethanol-fed rats. Ethanol-induced increases in specific proteins that hinder LD lipid metabolism led to the accumulation and persistence of large-sized LDs in the liver.

60 APPLIED LIFE SCIENCES↗

Correlations between Texture Profile Analysis and Sensory Evaluation of Cured Largemouth Bass Meat (Micropterus salmoides)

Texture is an important factor in evaluating the quality of aquatic products. To evaluate the texture properties of cured large mouth bass, edible sodium chloride (0, 1.0, 2.0, and 3.0%) was smeared to the bass meat. Texture profile analysis (TPA) and sensory evaluation were performed to evaluate the quality of the cured samples, and the correlations of the indexes in the two methods were analyzed. Two principal components were obtained from the TPA indexes and sensory evaluation indexes, and the cumulative variance contribution rates were 73.87% and 72.99%, respectively. Results from the principal component analysis showed that the main indicators that affected the TPA were gumminess and springiness, while those that affected sensory evaluation were chewiness and adhesiveness. The TPA index and sensory evaluation could be effectively improved when the sodium chloride added to the bass meat was 1%. In the correlation analysis, sensory springiness was negatively correlated with TPA hardness ( P < 0.05 , r = −0.553) but positively correlated with TPA chewiness ( P < 0.05 , r = 0.596). After stepwise regression analysis, the prediction equation between the sensory springiness and TPA hardness was obtained as SSp=5.770−0.002Ha. These results provide a basis for predicting the quality of large mouth bass cured products.

Li, Meijin↗

Machine Learning Model Geotiffs - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains geotiffs, supporting shapefiles and readmes for the inputs and output models of algorithms explored in the Nevada Geothermal Machine Learning project, meant to accompany the final report. Layers include: Artificial Neural Network (ANN), Extreme Learning Machine (ELM), Bayesian Neural Network (BNN), Principal Component Analysis (PCA/PCAk), Non-negative Matrix Factorization (NMF/NMFk), input rasters of feature sets, and positive/negative training sites. See readme .txt files and final report for additional metadata. A submission linking the full codebase for generating machine learning output models is available under "related resources" on this page.

15 GEOTHERMAL ENERGY↗

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Genetic variation in nitrogen‐use efficiency and its associated traits in dryland winter wheat ( Triticum aestivum L.) cultivars released from the 1940s to the 2010s in Shaanxi Province, China

Abstract BACKGROUND Improving the nitrogen‐use efficiency (NUE) of wheat can help mitigate the problems of poor soil fertility under dryland conditions. We conducted field experiments using three nitrogen (N) fertilization levels (0, 120, and 180 kg ha −1 ) applied to eight dryland wheat cultivars to assess NUE and its associated traits. RESULTS The grain yield significantly increased with the improvement in variety, mainly as a result of a substantial increase in 1000‐grain weight and harvest index. Modern wheat varieties have stabilized at an optimal plant height and exhibited improved performance in terms of NUE, partial N productivity, N harvest index, and grain protein content compared to older varieties. The NUE of wheat gradually increased with variety replacement. The net photosynthesis rate of the flag leaves in the filling stage improved with the year of cultivar release; Increasing soil–plant analysis development (SPAD) values of flag leaves in the flowering and filling stages were observed over time, with the flag leaves of modern varieties showing a high chlorophyll content in the filling stage. Additionally, the principal component analysis showed that the SPAD value, grain number per unit area, transpiration rate, leaf area, and grain protein content positively contributed to the clustering of the N180 and modern cultivars (from the 2000s to 2010s). CONCLUSION Overall, high levels of N application did not significantly improve the NUE of wheat. However, modern wheat varieties can optimize N distribution, increase flag leaf photosynthetic capacity, and improve photosynthesis ability, thus enhancing NUE to achieve high yields under a suitable level of N supply. © 2022 Society of Chemical Industry.

Lian, Huida↗

Finding Hidden Patterns in High Resolution Wind Flow Model Simulations

Wind flow data is critical in terms of investment decisions and policy making. High resolution data from wind flow model simulations serve as a supplement to the limited resource of original wind flow data collection. Given the large size of data, finding hidden patterns in wind flow model simulations are critical for reducing the dimensionality of the analysis. In this work, we first perform dimension reduction with two autoencoder models: the CNN-based autoencoder (CNN-AE) [1], and hierarchical autoencoder (HIER-AE) [2], and compare their performance with the Principal Component Analysis (PCA). We then investigate the super-resolution of the wind flow data. By training a Generative Adversarial Network (GAN) with 300 epochs, we obtained a trained model with 2× resolution enhancement. We compare the results of GAN with Convolutional Neural Network (CNN), and GAN results show finer structure as expected in the data field images. Also, the kinetic energy spectra comparisons show that GAN outperforms CNN in terms of reproducing the physical properties for high wavenumbers and is critical for analysis where high-wavenumber kinetics play an important role.

97 MATHEMATICS AND COMPUTING↗

Uncertainty quantification for Multiphase-CFD simulations of bubbly flows: a machine learning-based Bayesian approach supported by high-resolution experiments

In this paper, we developed a machine learning-based Bayesian approach to inversely quantify and reduce the uncertainties of multiphase computational fluid dynamics (MCFD) simulations for bubbly flows. The proposed approach is supported by high-resolution two-phase flow measurements, including those by double-sensor conductivity probes, high-speed imaging, and particle image velocimetry. Local distributions of key physical quantities of interest (QoIs), including the void fraction and phasic velocities, are obtained to support the Bayesian inference. In the process, the epistemic uncertainties of the closure relations are inversely quantified while the aleatory uncertainties from stochastic fluctuations of the system are evaluated based on experimental uncertainty analysis. The combined uncertainties are then propagated through the MCFD solver to obtain uncertainties of the QoIs, based on which probability-boxes are constructed for validation. The proposed approach relies on three machine learning methods: feedforward neural networks and principal component analysis for surrogate modeling, and Gaussian processes for model form uncertainty modeling. The whole process is implemented within the framework of an open-source deep learning library PyTorch with graphics processing unit (GPU) acceleration, thus ensuring the efficiency of the computation. The results demonstrate that with the support of high-resolution data, the uncertainties of MCFD simulations can be significantly reduced. The proposed approach has the potential for other applications that involve numerical models with empirical parameters.

42 ENGINEERING↗

Topological Data Analysis for Particulate Gels

Soft gels, formed via the self-assembly of particulate materials, exhibit intricate multiscale structures that provide them with flexibility and resilience when subjected to external stresses. Here, this work combines particle simulations and topological data analysis (TDA) to characterize the complex multiscale structure of soft gels. Our TDA analysis focuses on the use of the Euler characteristic, which is an interpretable and computationally scalable topological descriptor that is combined with filtration operations to obtain information on the geometric (local) and topological (global) structure of soft gels. We reduce the topological information obtained with TDA using principal component analysis (PCA) and show that this provides an informative low-dimensional representation of the gel structure. We use the proposed computational framework to investigate the influence of gel preparation (e.g., quench rate, volume fraction) on soft gel structure and to explore dynamic deformations that emerge under oscillatory shear in various response regimes (linear, nonlinear, and flow). Our analysis provides evidence of the existence of hierarchical structures in soft gels, which are not easily identifiable otherwise. Moreover, our analysis reveals direct correlations between topological changes of the gel structure under deformation and mechanical phenomena distinctive of gel materials, such as stiffening and yielding. In summary, we show that TDA facilitates the mathematical representation, quantification, and analysis of soft gel structures, extending traditional network analysis methods to capture both local and global organization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

XRF-MAPS

x-ray fluorescence (XRF) imaging typically involves the creation and analysis of 3D data sets, where at each scan position a full energy dispersive x-ray spectrum is recorded. This allows one to later process the data in a variety of different approaches, e.g., by spectral region of interest (ROI) summation with or without background subtraction, principal component analysis, or fitting. XRF-Maps is a C++ open source software package that implements these functions to provide a tool set for the analysis of XRF data sets.

GLOWACKI, ARTHURT↗

Modeling the Spectral Diversity of Quasars in the Sixteenth Data Release from the Sloan Digital Sky Survey

Abstract We present a new approach to capturing the broad diversity of emission-line and continuum properties in quasar spectra. We identify populations of spectrally similar quasars through pixel-level clustering on 12,968 high signal-to-noise ratio (S/N) spectra from the Sloan Digital Sky Survey (SDSS) in the redshift range of 1.57 < z < 2.4. Our clustering analysis finds 396 quasar spectra that are not assigned to any population, 15 misclassified spectra, and 6 quasars with incorrect redshifts. We compress the quasar populations into a library of 684 high-S/N composite spectra, anchored in redshift space by the Mg ii emission line. Principal component analysis on the library results in an eigenspectrum basis spanning 1067–4007 Å. We model independent samples of SDSS quasar spectra with the eigenbasis, allowing for a free redshift parameter. Our models achieve a median reduced χ 2 on non–broad absorption line quasar spectra that is reduced by 8.5% relative to models using the eigenspectra from the SDSS spectroscopic pipeline. A significant contribution to the relative improvement is from the ability to reconstruct the range of emission-line variation. The redshift estimates from our model are consistent with the Mg ii emission-line redshift with an average offset that displays 51.4% less redshift-dependent variation relative to the SDSS eigenspectra. Our method for developing quasar spectra models can improve automated classification and predict the intrinsic spectrum in regions affected by intervening absorbers such as Ly α , C iv , and Mg ii , thus benefiting studies of large-scale structure.

79 ASTRONOMY AND ASTROPHYSICS↗

Chemistry imaging and distribution analysis of rare earth elements in coal using LIBS and LA-ICP-MS instruments

Currently, demand for rare earth elements (REEs) increased significantly. Coal is actively evaluated as potential economic sources for extraction of REEs. Here, in this work, laser-induced breakdown spectroscopy (LIBS) was evaluated for rapid estimation of REEs content and their distribution in the natural coal samples. The results were compared with similar laser ablation–inductively coupled plasma–mass spectrometry (LA-ICP-MS) measurements. Thirteen coal samples (nine standard samples and five natural samples) were used in this study. Powder samples were pressed into pellets while coal chunks were directly ablated for data recording. Pellets of the powder standard samples were used to optimize the data acquisition system and then data recorded with this optimized system was used to identify the proper data acquisition and analysis models. After establishing the proper data acquisition system and analysis model using the standard samples, natural coal samples in powder form and their chunks were utilized to record LIBS and LA-ICP-MS spectra. Multivariate calibration models were developed using four of the natural samples, which were evaluated by predicting the REE content in the fifth sample. Principal component analysis was performed on the LIBS data obtained from the natural samples and it classified all the samples with high accuracy. Two-dimensional (2D) elemental mapping on coal chunk samples was also performed using both LIBS and LA-ICP-MS to study the distribution of REEs in the samples. The resulting elemental images and their correlations can be used to infer mineral distributions.

01 COAL, LIGNITE, AND PEAT↗

Using real-time data analysis to conduct next-generation synchrotron fatigue studies

Next-generation experimental techniques, like high energy X-ray diffraction microscopy (HEDM), usher in new opportunities to collect the grain-scale data necessary for understanding the evolving processes that drive fatigue failure. In this study, we present a framework for monitoring the evolution of a deforming polycrystal, in real-time, by applying principal component analysis (PCA) to raw X-ray diffraction image data. We applied this framework to inform in-situ HEDM measurements of a cyclically loaded Inconel-718 superalloy. Further, we discovered correlations between PCA of the diffraction data and the physical processes in the polycrystal. Lastly, we discuss extending this framework in future HEDM fatigue studies.

36 MATERIALS SCIENCE↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗