Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Domain knowledge”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

CI-MOR Final Report: Analysis and Validation of Critical Infrastructure Models using Model Order Reduction

This report summarizes the research and capabilities developed as part of the project “Analysis and Validation of Critical Infrastructure Models using Model Order Reduction” (CI-MOR) LDRD project. CI-MOR research enables the solution of large, complex optimization models that naturally arise in national security challenges involving critical infrastructures. Specifically, CI-MOR researchers developed methods to (1) rigorously approximate complex, nonlinear optimization formulations, (2) identify alternative near-optimal solutions, (3) accelerate optimization workflows used for complex applications, and (4) rigorously integrate domain knowledge in stochastic-process models. This report provides an overview of the research done in CI-MOR, and we describe application exemplars used to illustrate CI-MOR capabilities. Furthermore, we describe the software developed by CI-MOR that researchers can leverage to analyze new applications.

97 MATHEMATICS AND COMPUTING↗

Celeritas Midterm SciDAC Report

Celeritas is a new Monte Carlo (MC) code that helps satisfy the increasing demand for high energy physics (HEP) detector simulation, using Graphics Processing Unit (GPU) hardware on high performance computing (HPC) systems to model Large Hadron Collider (LHC) experiments and beyond. This report details the project’s progress midway through its SciDAC funding period, highlighting the first complete implementation of standard electromagnetic (EM) physics on GPUs, initial results for performance and scalability on Leadership Computing Facilities (LCFs), and preliminary integration into the CMS and ATLAS experiments. By integrating HEP domain knowledge with expertise in MC transport, Celeritas has catalyzed a shift in the HEP community’s perception of GPU platforms as the future for HPC simulations.

97 MATHEMATICS AND COMPUTING↗

A Data-Driven Approach to Real-World Degradation of Backsheets

The objectives of this project are as follows: • The population behavior of fielded modules in various conditions of use • Predictions of materials in specific climatic zones • Understanding of a module’s local environment in the field on its degradation It aims to understand how and what backsheet materials of photovoltaic modules degrade in the real world field. • Field Survey Protocol This project is started from the protocol, because all the data, information, domain knowledge are from experience of the real world field surveys, which is based on the protocol. The Protocol is explored from the experience of the field survey observations. With the increasing of the field surveys, it is refined for three versions, which are Task 1.0 (Section 3.1), Task 6.0 (Section 3.6), Task 10.0 (Section 3.10) respectively. It includes a document and a training video, which is able to direct other teams to follow the same procedure with the sites surveyed during this project. The documents include the detailed information but is not limited to the terminology definition, instruments SOP, preparation items for the surveys, form for the data collections, the order of the information collection. The final version of the protocol can be found at Appendix A, see Section 5. Additional, It can also be found at Open Science Framework (OSF), see Section 3.18 for detail. • Written Waiver Request Before staring the field surveys, the request of the waiver for international surveys is completed, because the limitation of climate zone in the United States, see Appendix B in Section 6 for the request documents. Unfortunately, only 1 international site from Taiwan, China can be finished, due to the COVID-19. • Field Survey According to the protocol we built in Section 3.1, 3.6 and 3.10, 41 sites have been surveyed across seven different climate zones (Cfa, Csa, Csb, BSk, Dfa, Dfb, Am). A variety of materials, including Polyethylene Naphthalate (PEN), Polyethylene Terephthalate (PET), Polyvinyl Fluoride (PVF), Polyvinylidene Fluoride (PVDF), Acrylic PVDF, Fluoroethylene Vinyl Ether (FEVE), and Glass, were identified. These sites are located in various states including California, South Carolina, New Mexico, Maryland, Ohio, Tennessee, Florida, Massachusetts, Illinois, Minnesota, Oregon, Colorado, and Taiwan, Republic of China. The ages of the sites ranged from 2 - 38 years in service and the field size varied from 1 MW - 25 MW. All requirements for the modeling have been satisfied. Some observations like ’Edge Effect’ for the rows and Junction box heating will also be a useful knowledge to build the model. Section 3.7 provides detailed information on the sites visited during this reporting period.

14 SOLAR ENERGY↗

An Advanced Machine Learning and Artificial Intelligence System for Demonstrating Radiation Regulatory Compliance in DOE Accelerator Facilities

In this Phase II proposal, Applied Research LLC (ARLLC), Thomas Jefferson National Accelerator Facility (Jefferson Lab), and Old Dominion University (ODU) propose the combination of domain knowledge (beam characteristics, fixed structural shielding, earthen burden (the soil and foliage added to the dome of the experimental halls as additional shielding), etc.), machine learning (ML) and/or artificial intelligence (AI) to correlate a variety of multi-modal onsite signals and the radiation fields seen in accessible areas of the accelerator site and the site boundary. The ML/AI will consider the complex influence of environmental parameters affecting the radon contribution of the measurements, focusing on actual data obtained from Jefferson Lab. In Phase I, the coded beam and location data were fed into a deep learning model to predict doses at several designated locations in Jefferson Lab’s facility. Moreover, a dense radiation map was generated using only a sparse collection of the samples in a facility. In Phase II, we will develop a software prototype containing a radiation prediction algorithm, dense radiation map algorithms, and background noise prediction algorithms, with actual data used to evaluate the prototype. This work will provide a framework for evaluation of radiation measurement results around the site based on learned responses. In addition, the proposed approach allows more granular mapping of radiation levels. Better understanding and communication of these levels is related to the overall approach in keeping doses to personnel ALARA.

43 PARTICLE ACCELERATORS↗

Hierarchical Bayesian Modeling for Cosmology: Can NPE reliably replace MCMC?

Hierarchical neural posterior estimation has its place Hierarchical Bayesian Modeling (HBM) combined with MCMC algorithms has been shown to provide more robust and accurate inference for real-world phenomena in which nature takes a nested form. However, MCMC-based inference can be computationally expensive, and its performance often suffers for complex posterior geometries. These costs are especially pertinent for HBM. Studies have recently demonstrated the potential for a flexible, expressive, and amortized hierarchical neural posterior estimator (HNPE) built on Normalizing Flows. These studies have mostly been performed on simple datasets, or they focus on a single parameter from each level of the hierarchy. A systematic study analyzing how both hierarchical methods compare for more complex and realistic datasets is necessary before applying HNPE for scientific measurements. Here, we re-explore the theory behind HNPE and conduct comparative numerical experiments of HNPE and MCMC-based HBM methods on real and synthetic data, including strong gravitational lensing simulations. In particular, we use a suite of diagnostics to show trade-offs in terms of accuracy, precision, time to train or sample, reproducibility, and the need for expert domain knowledge. Especially for higher dimensional and complex posteriors, HNPE is expected to drastically improve on time for inference, accuracy, and precision with an upfront training time cost.

Hur, Rachel [Chicago U.] (ORCID:000900089890445X)↗

Zentropy Theory for Transformative Functionalities of Magnetic and Superconducting Materials

The proposed research developed the zentropy theory through applications to complex magnetic materials and superconductors under the hypothesis that the emergent properties of complex magnetic materials and superconductors can be predicted by statistical mechanics of ergodic microstates with their partition functions computed from DFT-predicted free energies. The key objective is to develop approaches to systematically determine the types and number of microstates and the supercell size in DFT-based calculations through convergency of macroscopic functionalities, with the incorporation of our mixed-space approach accounting for the interactions between periodic supercells. In addition to use scientific intuitions to guide the design of important microstates, the key innovation of the proposed research is to integrate the domain knowledge and the material-property-descriptor database (MPDD) with 4 million microstates, which is supported by our deep neural network machine learning models (SIPFENN: structure-informed prediction of formation energy using neural networks) and integrated with our high throughput DFT Tool Kit (DFTTK). For complex magnetic materials, one of the objectives is to develop approaches to calculate short-range ordering from the statistical distribution of each microstate. For superconductors, the divergency of quasiparticle effective mass at a quantum critical point will be investigated, and the superconducting and non-superconducting microstates will be delineated through analysis of electronic band structure, density of states, charge density, and Fermi surface.

36 MATERIALS SCIENCE↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Statistical analysis and degradation pathway modeling of photovoltaic minimodules with varied packaging strategies

Degradation pathway models constructed using network structural equation modeling (netSEM) are used to study degradation modes and pathways active in photovoltaic (PV) system variants in exposure conditions of high humidity and temperature. This data-driven modeling technique enables the exploration of simultaneous pairwise and multiple regression relationships between variables in which several degradation modes are active in specific variants and exposure conditions. Durable and degrading variants are identified from the netSEM degradation mechanisms and pathways, along with potential ways to mitigate these pathways. A combination of domain knowledge and netSEM modeling shows that corrosion is the primary cause of the power loss in these glass/backsheet PV minimodules. We show successful implementation of netSEM to elucidate the relationships between variables in PV systems and predict a specific service lifetime. The results from pairwise relationships and multiple regression show consistency. This work presents a greater opportunity to be expanded to other materials systems.

electrical measurements↗

A Review of Physics-Informed Machine Learning in Fluid Mechanics

Physics-informed machine-learning (PIML) enables the integration of domain knowledge with machine learning (ML) algorithms, which results in higher data efficiency and more stable predictions. This provides opportunities for augmenting—and even replacing—high-fidelity numerical simulations of complex turbulent flows, which are often expensive due to the requirement of high temporal and spatial resolution. In this review, we (i) provide an introduction and historical perspective of ML methods, in particular neural networks (NN), (ii) examine existing PIML applications to fluid mechanics problems, especially in complex high Reynolds number flows, (iii) demonstrate the utility of PIML techniques through a case study, and (iv) discuss the challenges and opportunities of developing PIML for fluid mechanics.

42 ENGINEERING↗

Application of Systems Engineering Principles and Techniques in Biological Big Data Analytics: A Review

In the past few decades, we have witnessed tremendous advancements in biology, life sciences and healthcare. These advancements are due in no small part to the big data made available by various high-throughput technologies, the ever-advancing computing power, and the algorithmic advancements in machine learning. Specifically, big data analytics such as statistical and machine learning has become an essential tool in these rapidly developing fields. As a result, the subject has drawn increased attention and many review papers have been published in just the past few years on the subject. Different from all existing reviews, this work focuses on the application of systems, engineering principles and techniques in addressing some of the common challenges in big data analytics for biological, biomedical and healthcare applications. Specifically, this review focuses on the following three key areas in biological big data analytics where systems engineering principles and techniques have been playing important roles: the principle of parsimony in addressing overfitting, the dynamic analysis of biological data, and the role of domain knowledge in biological data analytics.

dynamic analysis↗

Machine Reading at Scale: A Search Engine for Scientific and Academic Research

The Internet, much like our universe, is ever-expanding. Information, in the most varied formats, is continuously added to the point of information overload. Consequently, the ability to navigate this ocean of data is crucial in our day-to-day lives, with familiar tools such as search engines carving a path through this unknown. In the research world, articles on a myriad of topics with distinct complexity levels are published daily, requiring specialized tools to facilitate the access and assessment of the information within. Recent endeavors in artificial intelligence, and in natural language processing in particular, can be seen as potential solutions for breaking information overload and provide enhanced search mechanisms by means of advanced algorithms. As the advent of transformer-based language models contributed to a more comprehensive analysis of both text-encoded intents and true document semantic meaning, there is simultaneously a need for additional computational resources. Information retrieval methods can act as low-complexity, yet reliable, filters to feed heavier algorithms, thus reducing computational requirements substantially. In this work, a new search engine is proposed, addressing machine reading at scale in the context of scientific and academic research. It combines state-of-the-art algorithms for information retrieval and reading comprehension tasks to extract meaningful answers from a corpus of scientific documents. The solution is then tested on two current and relevant topics, cybersecurity and energy, proving that the system is able to perform under distinct knowledge domains while achieving competent performance.

Sousa, Norberto (ORCID:0000000329194817)↗

Gearbox bearing crack growth prognostics and uncertainty quantification with physics-informed machine learning

This paper introduces the extreme theory of functional connections (X-TFC), a physics-informed machine learning algorithm, and tailors it to estimate the remaining useful life (RUL) of wind turbine gearbox bearings experiencing fatigue crack growth. Unlike purely data-driven methods, X-TFC embeds a physics model, based on Head's theory in this work, into its training objective. The core of X-TFC is a random-projection single-layer neural network trained via an extreme learning machine, which requires only limited damage progression data and solves for output weights with a least-squares optimization algorithm. A composite loss function balances the network's fit to observed degradation data against the residuals of the governing crack growth differential equation, ensuring the learned damage trajectory remains physically plausible. When applied to a vibration-based health-index (HI) dataset measured during the growth of a crack on the inner ring of a high-speed bearing in a wind turbine gearbox (Bechhoefer and Dubé, 2020), X-TFC achieves near-zero prediction bias. Even when trained on only the first 10 %–20 % of the damage progression data, with sufficient physics weighting its predictions remain monotonic and smooth, delivering high prognosability and trendability. To quantify the epistemic uncertainty, we employ a Monte Carlo ensemble of independently initialized X-TFC models trained on noise-perturbed data, which yields confidence intervals around each RUL estimate and captures both model-parameter and epistemic uncertainty. In addition to a vibration-based HI, we demonstrate that the proposed framework can be directly applied to a supervisory control and data acquisition (SCADA) data-based HI (Eftekhari Milani et al., 2026) measured during similar wind turbine gearbox bearing crack faults, preserving its accuracy and interpretability. This extension shows the versatility of our approach, which is applicable to bearings of multiple gearbox manufacturers, models, and ratings using only SCADA data. By integrating domain knowledge with machine learning, X-TFC offers a rapid, reliable tool for crack prognostics. Its adaptability to other bearing failure modes, such as pitch bearing ring cracks, positions X-TFC as a powerful enabler of data-driven, physics-informed asset management in the wind energy sector and beyond.

17 WIND ENERGY↗

Decreasing wind speed extrapolation error via domain-specific feature extraction and selection

Abstract. Model uncertainty is a significant challenge in the wind energy industry and can lead to mischaracterization of millions of dollars' worth of wind resources. Machine learning methods, notably deep artificial neural networks (ANNs), are capable of modeling turbulent and chaotic systems and offer a promising tool to produce high-accuracy wind speed forecasts and extrapolations. This paper uses data collected by profiling Doppler lidars over three field campaigns to investigate the efficacy of using ANNs for wind speed vertical extrapolation in a variety of terrains, and it quantifies the role of domain knowledge in ANN extrapolation accuracy. A series of 11 meteorological parameters (features) are used as ANN inputs, and the resulting output accuracy is compared with that of both standard log-law and power-law extrapolations. It is found that extracted nondimensional inputs, namely turbulence intensity, current wind speed, and previous wind speed, are the features that most reliably improve the ANN's accuracy, providing up to a 65 % and 52 % increase in extrapolation accuracy over log-law and power-law predictions, respectively. The volume of input data is also deemed important for achieving robust results. One test case is analyzed in depth using dimensional and nondimensional features, showing that the feature nondimensionalization drastically improves network accuracy and robustness for sparsely sampled atmospheric cases.

17 WIND ENERGY↗

Graph Neural Networks for Particle Reconstruction in High Energy Physics detectors

Pattern recognition problems in high energy physics are notably different from traditional machine learning applications in computer vision. Reconstruction algorithms identify and measure the kinematic properties of particles produced in high energy collisions and recorded with complex detector systems. Two critical applications are the reconstruction of charged particle trajectories in tracking detectors and the reconstruction of particle showers in calorimeters. These two problems have unique challenges and characteristics, but both have high dimensionality, high degree of sparsity, and complex geometric layouts. Graph Neural Networks (GNNs) are a relatively new class of deep learning architectures which can deal with such data effectively, allowing scientists to incorporate domain knowledge in a graph structure and learn powerful representations leveraging that structure to identify patterns of interest. In this work we demonstrate the applicability of GNNs to these two diverse particle reconstruction problems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The CYBER security – Competency Health and Maturity Progression (CYBER-CHAMP) model: Extending the National Initiative for Cybersecurity Education (NICE) Framework Across Organizational Security

Problem Statement: There is a pervasive talent deficit in the cybersecurity industry that prevents employers from being able to fill their open positions efficiently. A holistic approach to security is required to ensure organizations have adequate prevention and response capabilities in case of a cyberattack. Specifically, industrial control systems (ICS’s) and their operational technology (OT) components have become a constant target for cyberattacks. Research Questions: It is proposed that the NICE Framework should be extended in the following areas: 1) Include guidance regarding the job roles and competencies for both IT and OT professionals. 2) Offer step-by-step solutions, based on the work role mappings from the NICE Framework, to increase cybersecurity through employee training and education. 3) Provide a streamlined, lifecycle approach to building a cybersecurity program. Contribution: The CYBER security – Competency Health and Maturity Progression (CYBER-CHAMP©) model provides a customized solution for businesses to understand their education gaps in organizational security and target areas for improvement. Rationale: The Framework for Improving Critical Infrastructure Cybersecurity v1.1 addresses ICS but does not offer a measurement of cybersecurity maturity or clear methods to ascertain an organization’s current risk profile. In Phases 1 and 5 of the model, measurements are provided to help an organization build their current and target risk profiles. The NICE framework provides a structure for planning an IT cybersecurity workforce, but the OT aspects of cybersecurity are only briefly discussed. The model uses Phases 2-3 to examine the competencies of an organization’s workforce, which includes both IT and OT roles. Current frameworks do not offer next steps to increase an organization’s cybersecurity. During Phase 4, employees’ roles are mapped to training, education, and/or certifications from common vendors. Investigative Approach: The model provides measurements and metrics for both an organization’s status and continual improvement. This improvement methodology includes guidance for creating an overall strategic plan for security improvement via products designed to increase an organization’s operational readiness through workforce competency health. Lessons Learned: Depending on who was participating, there were contradicting answers given in Phase 1 due to different security cultures in the organization. This revelation has influenced the steps listed in the User’s Guide, where Phase 1’s first recommended step is to assemble a team that champions the facilitation and implementation of the model in the organization. During Phase 2, the discovery was made that organizations may be missing roles that are necessary to perform critical cybersecurity functions. By understanding the functional roles and competencies needed, they can contract or hire cybersecurity help to fill these gaps. Implications: Using the model, organizations can discuss quantitative measures for improvement as a business case for advancing their security program. Future research can validate and extend the present theory and model to a variety of environments. It is of interest to investigate additional security roles and knowledge domains that are used to build standardized cybersecurity curriculum.

97 MATHEMATICS AND COMPUTING↗

Snowmass Computational Frontier: Topical Group Report on Quantum Computing

Quantum computing will play a pivotal role in the High Energy Physics (HEP) science program over the early parts of the 21$^{st}$ Century, both as a major expansion of our capabilities across the Computational Frontier, and in synthesis with quantum sensing and quantum networks. This report outlines how Quantum Information Science (QIS) and HEP are deeply intertwined endeavors that benefit enormously from a strong engagement together. Quantum computers do not represent a detour for HEP, rather they are set to become an integral part of our discovery toolkit. Problems ranging from simulating quantum field theories, to fully leveraging the most sensitive sensor suites for new particle searches, and even data analysis will run into limiting bottlenecks if constrained to our current computing paradigms. Easy access to quantum computers is needed to build a deeper understanding of these opportunities. In turn, HEP brings crucial expertise to the national quantum ecosystem in quantum domain knowledge, superconducting technology, cryogenic and fast microelectronics, and massive-scale project management. The role of quantum technologies across the entire economy is expected to grow rapidly over the next decade, so it is important to establish the role of HEP in the efforts surrounding QIS. Fully delivering on the promise of quantum technologies in the HEP science program requires robust support. It is important to both invest in the co-design opportunities afforded by the broader quantum computing ecosystem and leverage HEP strengths with the goal of designing quantum computers tailored to HEP science.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗