Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical feature extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Topological Approach for Motion Track Discrimination

Detecting small targets at range is difficult because there is not enough spatial information present in an image sub-region containing the target to use correlation-based methods to differentiate it from dynamic confusers present in the scene. Moreover, this lack of spatial information also disqualifies the use of most state-of-the-art deep learning image-based classifiers. Here, we use characteristics of target tracks extracted from video sequences as data from which to derive distinguishing topological features that help robustly differentiate targets of interest from confusers. In particular, we calculate persistent homology from time-delayed embeddings of dynamic statistics calculated from motion tracks extracted from a wide field-of-view video stream. In short, we use topological methods to extract features related to target motion dynamics that are useful for classification and disambiguation and show that small targets can be detected at range with high probability.

Emerson, Tegan H.↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Statistical White-Line Analysis in High-Throughput TXM-XANES for Chemical State Quantification

The transmission X-ray microscopy (TXM) based X-ray absorption near-edge structure (XANES) technique provides three-dimensional mapping of element-specific chemical states at nanometer-scale spatial resolution and micrometer-scale fields of view. However, compared to conventional volume-averaged XANES (VA-XANES) measurements, the inherently small voxel size in TXM-XANES leads to a lower signal-to-noise ratio, making full-spectrum analysis computationally demanding and less robust. Here, we present the structural and compositional conditions for a statistical white-line analysis framework under which chemical state information can be directly extracted from the white-line peak position in voxel spectra without the need for voxel-wise background subtraction or normalization, under well-defined structural and compositional conditions. The method is validated on layered oxide cathode materials, where low-order polynomial fitting accurately reproduces white-line features, and the extracted energy distributions correlate strongly with VA-XANES results. This statistical approach enables high-throughput, dose-efficient, and noise-robust chemical state quantification in TXM-XANES, offering broad applicability to functional materials requiring nanoscale oxidation-state mapping.

TXM↗

Chain trajectories, domain shapes, and terminal boundaries in block copolymers

The packing geometry of macromolecules in complex mesophases is of key importance to self-organization in synthetic and biological soft materials. While approximate or heuristic models rely on often-untested assumptions about how flexible molecules “fit in” to distinct locations of complex assemblies, physical assemblies derive from ensembles of fluctuating conformations, obscuring the connection between mesophase geometry and the underlying arrangements. Here, we present an approach to extract and analyze features of molecular packing in diblock block copolymer (BCP) melts, a prototypical soft matter system, based on the statistical description of chain conformations in self-consistent field (SCF) theory. We show how average BCP chain trajectories in ordered morphologies can be analyzed from the SCF-derived orientational order parameter of chain segments. We use these extracted trajectories to analyze the features of local packing geometry, including chain bending and tilt, as well as the terminal boundaries that delineate distinct domains in ordered BCP morphologies. In conclusion, we illustrate this analysis by focusing on measurable features of packing frustration in 2D (columnar) and 3D (spherical and bicontinuous) morphologies, notably establishing an explicit link between chain conformations in complex morphologies and their medial geometry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sensor selection and tool wear prediction with data‐driven models for precision machining

Abstract Estimation of tool wear in precision machining is vital in the traditional subtractive machining industry to reduce processing cost, improve manufacturing efficiency and product quality. In this vein, fusion of time and frequency‐domain features of commonly sensed signals can provide an early indication of tool wear and improve its prediction accuracy for prognostics and health management. This paper presents a data‐driven methodology and a complete tool chain for the inference of precision machining tool wear from fused machine measurements, such as cutting force, power, audio and vibration signals, and quantify the usefulness of each measurement. Indicators of tool wear are extracted from time‐domain signal statistics, frequency‐domain analysis, and time‐frequency domain analysis. Correlation coefficients between the extracted features (indicators) and the tool wear are used to select the most informative features. Principal Component Analysis and Partial Least‐Squares are used to reduce the dimensionality of the feature space. Regression models, including linear regression, support vector regression, Decision tree regression, neural network regression and Gaussian process regression, are used to predict the tool wear using data from a Haas milling machine performing spiral boss face milling. The performance of the regression models based on subsets of sensors validates the preliminary estimates about the saliency of the sensors. The experimental results show that the proposed methods can predict the machine tool wear precisely, with readily available sensor measurements. Neural network and Gaussian process regression were able to achieve good estimates of tool wear at different machine operating conditions. The most informative signal in predicting tool wear was shown to be the vibration signal. Time‐frequency domain features were the most informative features among the combination of features of three domains. In addition, using partial least squares components extracted from the original features of signals led to higher prediction accuracy.

Han, Seulki↗

Decoding the Mechanisms of Phase Transitions from In Situ Microscopy Observations

Abstract Analysis of the temperature‐ and stimulus‐dependent imaging data toward elucidation of the physical transformations is an ubiquitous problem in multiple fields. Here, temperature‐induced phase transition in BaTiO 3 is explored using the machine learning analysis of domain morphologies visualized via variable‐temperature scanning transmission electron microscopy (STEM) imaging data. This approach is based on the multivariate statistical analysis of the time or temperature dependence of the statistical descriptors of the system, derived in turn from the categorical classification of observed domain structures or projection on the continuous parameter space of the feature extraction‐dimensionality reduction transform. The proposed workflow offers a powerful tool for the exploration of the dynamic data based on the statistics of image representation as a function of the external control variable to visualize the transformation pathways during phase transitions and chemical reactions. This can include the mesoscopic STEM data as demonstrated here, but also optical, chemical imaging, etc., data. It can further be extended to the higher dimensional spaces, for example, analysis of the combinatorial libraries of materials compositions.

Valleti, Sai Mani Prudhvi↗

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

In Situ Investigation of Upper Airway Occlusion in Sleep Disordered Breathing Using Ultrasonic Transducer Arrays

This work presents a novel application of ultrasound for the real-time, non-invasive investigation of occlusion of the upper airway during events of obstructive sleep apnea/hypopnea syndrome. It is hypothesized that ultrasonic pulses applied to the neck during apneic events produce spectral and temporal features that can detect apnea occurrence. Theoretical models of ultrasound propagation and an in vitro test were conducted to test this hypothesis in both transmission and reflection modes. Complete specifications and technical details of the system design and fabrication, which is mounted on each subject’s neck, are presented, including the methodology. Nine patients (seven male and two female, mean age of 42 years, with a range of 25 to 56 years, and body mass index 37.6 ± 6.6 kg/m2) were recruited for a full night study, which included simultaneous nocturnal polysomnography for the validation of the results. Nine temporal features and four spectral features were extracted from the envelope of the received pulse waveform. These were used to compute 26 metrics to quantify the changes in the ultrasonic waveforms between normal breathing and apneic events. The statistical analysis of the collected ultrasonic data showed that at least two or more of the proposed features could detect apneic events in all subjects. The findings establish the feasibility of the proposed method as a cost-effective and non-invasive OSAHS screening tool.

60 APPLIED LIFE SCIENCES↗

Automated, high-accuracy classification of textured microstructures using a convolutional neural network

Crystallographic texture is an important descriptor of material properties but requires time-intensive electron backscatter diffraction (EBSD) for identifying grain orientations. While some metrics such as grain size or grain aspect ratio can distinguish textured microstructures from untextured microstructures after significant grain growth, such morphological differences are not always visually observable. This paper explores the use of deep learning to classify experimentally measured textured microstructures without knowledge of crystallographic orientation. A deep convolutional neural network is used to extract high-order morphological features from binary images to distinguish textured microstructures from untextured microstructures. The convolutional neural network results are compared with a statistical Kolmogorov–Smirnov tests with traditional morphological metrics for describing microstructures. Results show that the convolutional neural network achieves a significantly improved classification accuracy, particularly at early stages of grain growth, highlighting the capability of deep learning to identify the subtle morphological patterns resulting from texture. The results demonstrate the potential of a convolutional neural network as a tool for reliable and automated microstructure classification with minimal preprocessing.

36 MATERIALS SCIENCE↗

Monitoring and flaw detection during wire-based directed energy deposition using in-situ acoustic sensing and wavelet graph signal analysis

The goal of this work is to detect flaw formation in the wire-based directed energy deposition (W-DED) process using in-situ sensor data. The W-DED studied in this work is analogous to metal inert gas electric arc welding. The adoption of W-DED in industry is limited because the process is susceptible to stochastic and environmental disturbances that cause instabilities in the electric arc, eventually leading to flaw formation, such as porosity and suboptimal geometric integrity. Moreover, due to the large size of W-DED parts, it is difficult to detect flaws post-process using non-destructive techniques, such as X-ray computed tomography. Accordingly, the objective of this work is to detect flaw formation in W-DED parts using data acquired from an acoustic (sound) sensor installed near the electric arc. To realize this objective, we develop and apply a novel wavelet integrated graph theory approach. The approach extracts a single feature called graph Laplacian Fiedler number from the noise-contaminated acoustic sensor data, which is subsequently tracked in a statistical control chart. Using this approach, the onset of various types of flaws are detected with a false alarm rate less-than 2%. This work demonstrates the potential of using advanced data analytics for in-situ monitoring of W-DED.

42 ENGINEERING↗

Rolling Root Mean Square Based Multimodal Anomaly Detection for Real Time Monitoring of Smart Grid

Reliable real-time monitoring is valuable for maintaining the operational integrity of modern electrical smart grids. Deployment of heterogeneous sensing technologies in substations has enabled high-resolution, multichannel waveform monitoring, but also introduces challenges for anomaly detection due to noise, baseline drift, and modality-dependent signal characteristics. In this work, we present a computationally efficient unsupervised method for multimodal event detection based on Rolling Root Mean Square based Event Detection (RRMSED). The method is developed using in-house, field deployed sensors collecting data at a utility substation. The sensing system comprises voltage and current sensors, triaxial accelerometers, and magnetometers, collectively capturing electrical, vibrational, and magnetic waveform measurements at high temporal resolution. RRMSED operates by extracting rolling RMS energy features and their first-order temporal differences from consecutive waveform segments for each channel and then applying channel-specific statistical thresholds learned from historical data. A persistence-based exceedance logic is employed to robustly identify transient events while suppressing impulsive noise, and to provide precise temporal localization with high resolution. The framework is designed for continuous server-side operation and can be deployed in real time without requiring complex models. Experiments on simulated waveform data with known ground truth demonstrate low false positive (FP) and false negative (FN) rates. Application to real substation data shows RRMSED to identify events that are not captured by conventional monitoring indicators including fast transient detection algorithm currently deployed in the system. These results indicate that rolling RMS based features provide an effective and practical basis for real-time multimodal event detection in smart-grid substations.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

In Situ Transmission Electron Microscopy: Signal processing challenges and examples

Transmission electron microscopy (TEM) is a powerful tool for imaging material structure and characterizing material chemistry. Recent advances in data collection technology for TEM have enabled high-volume and high-resolution data collection at a microsecond frame rate. Here, taking advantage of these advances in data collection rates requires the development and application of data processing tools, including image analysis, feature extraction, and streaming data processing techniques. In this article, we highlight a few areas in materials science that have benefited from combining signal processing and statistical analysis with data collection capabilities in TEM and present a future outlook on opportunities of integrating signal processing with automated TEM data analysis.

36 MATERIALS SCIENCE↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

State of the art astronomical simulations have provided datasets which enabled the training of novel deep learning techniques for constraining cosmological parameters. However, differences in subgrid physics implementation and numerical approximations among simulation suites lead to differences in simulated datasets, which pose a hard challenge when trying to generalize across diverse data domains and ultimately when applying models to observational data. Recent work reveals deep learning algorithms are able to extract more information from complex cosmological simulations than summary statistics like power spectra. We introduce Domain Adaptive Graph Neural Networks (DA-GNNs), trained on CAMELS data, inspired by CosmoGraphNet (Villanueva-Domingo et al 2023). By utilizing GNNs, we can capitalize on their capacity to capture both astrophysical and topological features of galaxy distributions. Mixing these capabilities with domain adaptation techniques such as Maximum Mean Discrepancy (MMD), which enable extraction of domain-invariant features, our framework demonstrates enhanced accuracy and robustness. We present experimental results, including the alignment of distributions across domains through data visualization. These findings suggest that DA-GNNs are an efficient way of extracting domain independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Comprehensive Neural Posterior Estimation for Galaxy-Galaxy Strong Lensing

We present a deep learning model based on neural posterior estimation (NPE) for comprehensive extraction of astrophysical parameters from galaxy-scale strong gravitational lenses. The unprecedentedly large amount of galaxy-scale strong lenses expected in future cosmological surveys (${\cal O}(10^5)$) promises to enable valuable statistical constraints in various studies ranging from galaxy formation to the nature of dark matter, but it also poses a significant challenge for traditional modelling pipelines. To this end, our automated model includes several new, state-of-the-art features and approaches leveraging the framework of simulation-based inference (SBI). We infer a total of 20 parameters describing the mass and light profiles of both lens and source galaxies, using simulated raw multi-band data modelled under noise and observing conditions expected by the Legacy Survey of Space and Time (LSST), with its summary statistics generated by a residual network. We examine the efficacy of multi-band data in extracting nearly 20 model parameters simultaneous from strong lensing images including lens light. Finally, We perform a comprehensive set of diagnostics for SBI models, evaluating the model's prediction accuracy, stability, and uncertainty quantification.

Zhao, Roy J. [Chicago U., KICP]↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

From Points to Planes: A Workflow for Converting Three‐Dimensional Point Cloud Data Into Discrete Fracture Network Flow and Transport Models

We present the Point cLoud Algorithm for NEtwork Extraction of Discrete Fracture Networks (PLANE-DFN), a point cloud–based algorithm for automatic fracture network extraction designed to support discrete fracture network (DFN) modeling workflows. PLANE-DFN segments three-dimensional fracture planes from raw point cloud data using RANdom SAmple Consensus coupled with statistical outlier removal and density-based clustering to isolate individual fracture features. Each candidate plane is constrained against site-specific structural constraints based on strike and dip. After segmentation, each fracture is converted into a 2-D convex polygon suitable for meshing and simulation. The PLANE-DFN algorithm is validated by comparing geometric and flow and transport data against data from dfnWorks simulations with ensembles of plane-fit networks. We find that the flow and transport in plane-fit networks are comparable to dfnWorks-generated networks when realistic network geometry is maintained. The PLANE-DFN algorithm provides an automated and streamlined workflow to transform point clouds of data into DFN network geometry.

54 ENVIRONMENTAL SCIENCES↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗