A Comparison of Online Model-Based Anomaly Detection Methods for a Lithium-Ion Battery Cell
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
In May 2025, nearly 700 participants from all around the world coalesced at 10 regional nodes and a few satellite nodes to take part in a global hackathon of kilometer-scale (horizontal grid spacing < 10 km) regional and global Earth system models. Exciting science is emerging from these efforts, ranging across novel model analysis, new ways of integrating with satellite data, and emulation with machine learning. New technologies were trialed that enable the community to work in new and complementary ways to democratize access to global information at a local scale from a set of the world’s highest-resolution climate models. The hackathon demonstrated how exascale data can be organized to be accessible to anyone. Fundamentally, the community could apply these techniques and technologies to move toward more participative models for coproduction and delivery of diverse sources of climate information for climate scientists and citizens alike.
Alkali-silica reaction (ASR) causes concrete degradation, leading to cracking, rebar corrosion, and reduced structural integrity, which raises safety concerns. Ultrasonic nondestructive evaluation (NDE) effectively assesses concrete properties and monitors ASR progression. However, its deployment and analysis require specialized expertise and subjective interpretation. As computational power increases, artificial intelligence (AI) and machine learning (ML) algorithms are increasingly being used to automate NDE data analysis across various industries for AI-assisted automation. Regulatory agencies are adapting to this technological shift, prompting a need to evaluate current ML technologies’ capabilities and limitations in assessing concrete material properties and damage. This report presents a comparative analysis of four ML regression models for predicting concrete material damage induced by ASR expansion using long-term ultrasonic data monitoring. The models investigated include linear regression (LR), support vector regression (SVR), shallow neural networks (NN), and deep neural networks (DNN). LR, SVR, and shallow NN models use features extracted from ultrasonic signals, whereas the DNN model processes time-domain ultrasonic signals and frequency spectra directly. The study systematically compared the models’ performance from various perspectives, including model input, prediction performance, and generalization ability. The findings indicate significant variability in model performance, with some ML algorithms achieving very high or very low prediction accuracy depending on the preprocessing and feature engineering (extraction and selection) applied. Key insights include the observation that shallow ML models (LR, SVR, and shallow NNs) require meticulous preprocessing and feature extraction to achieve high accuracy. In contrast, the DNN model, although it bypasses the need for feature engineering, necessitates extensive preprocessing to mitigate noise and computational demands. The SVR model emerged as the top performer among the shallow models, and the DNN model exhibited superior performance on specific datasets but struggled with generalization across specimens from different batches. Additionally, the SVR model is sensitive to temperature variations, whereas the DNN model is robust in this regard. Using recurrent neural networks is recommended for future ASR expansion prediction studies. Recurrent neural networks’ inherent ability to capture temporal dependencies and long-term patterns makes them well suited for analyzing sequential ultrasonic monitoring data. Overall, the results and conclusions of this study could provide insights into the capabilities and effectiveness of ML when applied to ultrasonic NDE data and help identify best practices for using ML for ultrasonic NDE of concrete material properties.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Objective performance and economic modeling of solar thermal plants is of keen interest to many EPRI funders. One solar thermal technology that is not currently available to model in any non-vendor, non-proprietary tool is linear Fresnel. This technology has garnered enough interest from EPRI funders to merit investing in a tool to objectively model its performance. Early in 2010, EPRI performed a comparison of modeling solar thermal power plants using the IPSEPRO, CNRS and Solar Advisor (SAM) tools. After completing this effort, EPRI decided to adopt NREL’s Solar Advisor Model as its default modeling tool based in part on user friendliness, flexibility, number of technologies covered, integrated financial model and ease of running sensitivities. Furthermore, it was recognized that NREL continues to invest considerable time and resources into improving capabilities and functionality of the model.
The ALICE Collaboration reports the measurement of semi-inclusive distributions of charged-particle jets recoiling from a high transverse momentum (high p T ) hadron trigger in proton-proton and central Pb-Pb collisions at $\sqrt{s_{NN}}$ = 5.02 TeV. A data-driven statistical method is used to mitigate the large uncorrelated background in central Pb-Pb collisions. Recoil jet distributions are reported for jet resolution parameter R = 0.2, 0.4, and 0.5 in the range 7 < p T.jet < 140 GeV=c and trigger-recoil jet azimuthal separation π/2 < Δφ < π. The measurements exhibit a marked medium-induced jet yield enhancement at low p T and at large azimuthal deviation from Δφ ~ π. The enhancement is characterized by its dependence on Δφ, which has a slope that differs from zero by 4.7σ. Comparisons to model calculations incorporating different formulations of jet quenching are reported. These comparisons indicate that the observed yield enhancement arises from the response of the QGP medium to jet propagation.
The frequency and intensity of summer heat waves in East Asia have increased sharply in recent decades, significantly impacting public health and the economy. The Arctic-Siberian Plain (ASP) teleconnection pattern has been identified as a key driver, with ASP warming amplifying atmospheric circulation patterns conducive to extreme temperatures. This study evaluates the ability of Coupled Model Inter-comparison Project phase 6 models to simulate the ASP pattern across interannual variability (IAV) and intra-seasonal variability (ISV) timescales using the Common Basis Function method. The multi-model mean shows statistically significant pattern correlations with ERA5 reanalysis, with correlation coefficients of 0.90 and 0.99 for IAV and ISV, respectively. While the ASP pattern is generally well captured, models exhibit substantial inter-model diversity in the intensity and position of anticyclonic anomalies over the ASP and East Asia. Models with ASP pattern variability similar to reanalysis better reproduce extreme East Asian temperatures, whereas those over- or underestimating ASP variability exhibit lower skill. These performance differences are related to differences in simulating key variables associated with the development of the ASP pattern. Our findings highlight the role of the ASP pattern in modulating extreme heat events, as models with improved ASP simulations align more closely with observed temperature extremes. Refining ASP representations in models could enhance seasonal heat wave predictions, improving climate adaptation strategies.
Electric field waveforms of light carry rich information about dynamical events on a broad range of timescales. The insight that can be reached from their analysis, however, depends on the accuracy of retrieval from noisy data. In this article, we present a novel approach for waveform retrieval based on supervised deep learning. We demonstrate the performance of our model by comparison with conventional denoising approaches, including wavelet transform and Wiener filtering. The model leverages the enhanced precision obtained from the nonlinearity of deep learning. The results open a path toward an improved understanding of physical and chemical phenomena in field-resolved spectroscopy.
Research on plasma generated by electrical discharge in gas-filled capillaries plays an important role in advancing laser wakefield accelerators. We present a comprehensive study of the temporal and spatial distribution of plasma density within a short, 1 cm long, square-shaped capillary filled with hydrogen gas. Two transverse capillary sizes, 500 μm and 300 μm, were investigated achieving peak plasma densities of 0.9 × 10 18 cm −3 and 2.5 × 10 18 cm −3 , respectively, at the capillary center. In addition, we explore how these distributions depend on discharge parameters, specifically discharge current and gas flow through the capillary. Plasma density was determined by analyzing the Stark broadening of the hydrogen Balmer-alpha line. The results obtained were compared with the theoretical models and simulations. The comparison between the model predictions and the experimental data at the transient ionization stage of the discharge reveals a discrepancy of a factor of ∼1.3–2.1, depending on the capillary size, which is thoroughly discussed.
This paper introduces STRIPE (Simulated Transport of RF Impurity Production and Emission), an advanced modeling framework developed to analyze material erosion and the global transport of eroded impurities originating from radio-frequency (RF) antenna structures in magnetic confinement fusion devices. STRIPE integrates multiple physics modules: SolEdge3x for scrape-off-layer plasma profiles, COMSOL for 3D RF rectified sheath potentials, RustBCA for erosion yields and surface interactions, and global impurity transport for 3D ion energy-angle distributions and impurity transport. The framework is applied to an ion cyclotron RF-heated L-mode discharge (#57877) in the WEST tokamak, where it predicts a thirty-fold increase in gross tungsten erosion at antenna limiters during the transition from ohmic to ICRH operation. Additionally, under ICRH conditions, a tenfold enhancement in erosion is observed when comparing RF sheath effects to purely thermal sheath conditions. High-charge-state oxygen ions ($\mathrm{O}$ 6+ and above) are identified as the dominant contributors to tungsten sputtering. To validate the model, a synthetic diagnostic tool based on inverse photon efficiency (S/XB coefficients) from the ColRadPy collisional-radiative model enables direct comparison with spectroscopic measurements. Model predictions using a plasma composition of 1% oxygen and 99% deuterium show good agreement with observed W − I (400.9 nm) emission for discharge #57877, supporting the accuracy of the STRIPE framework. This study focuses specifically on gross erosion calculations to demonstrate STRIPE’s capabilities. Future extensions of this work will incorporate net erosion, re-deposition, self-sputtering effects, and whole-device modeling of sputtered tungsten impurity transport. STRIPE is also being applied to other RF-heated linear and toroidal devices, offering valuable insights for antenna design, impurity control, and performance optimization in next-generation fusion reactors.
Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.
Abstract In Vadas et al. (2024, https://doi.org/10.1029/2024ja032521 ), we modeled the atmospheric gravity waves (GWs) during 11–14 January 2016 using the HIAMCM, and found that the polar vortex jet generates medium to large‐scale, higher‐order GWs in the thermosphere. In this paper, we model the traveling ionospheric disturbances (TIDs) generated by these GWs using the HIAMCM‐SAMI3 and compare with ionospheric observations from ground‐based Global Navigation Satellite System (GNSS) receivers, Incoherent Scatter Radars (ISR) and the Super Dual Auroral Radar Network (SuperDARN). We find that medium to large‐scale TIDs are generated worldwide by the higher‐order GWs from this event. Many of the TIDs over Europe and Asia have concentric ring/arc‐like structure, and most of those over North/South America have planar wave structure and occur during the daytime. Those over North/South America propagate southward and are generated by higher‐order GWs from Europe/Asia which propagate over the Arctic. These latter TIDs can be misidentified as arising from geomagnetic forcing. We find that the higher‐order GWs that propagate to Africa and Brazil from Europe may aid in the formation of equatorial plasma bubbles (EPBs) there. We find that the simulated GWs, TIDs and EPBs agree with EISCAT, PFISR, GNSS, and SuperDARN measurements. We find that the higher‐order GWs are concentrated at N at 200 km, in agreement with GOCE and CHAMP data. Thus the polar vortex jet is important for generating TIDs in the northern winter ionosphere via multi‐step vertical coupling through GWs.
An approximate and easily applied analytical model was developed for heat transfer calculations of heat exchangers consisting of multiple rows and columns of heat transfer fluid flow channels with semi-elliptical cross sections. Heat exchangers of this type are being developed by using ceramic material and additive manufacturing for high temperature and pressure-concentrating solar electric power plants. Calculations using the model require only the geometrical dimensions and flow conditions of the heat exchanger. Comparisons of modeling predictions, with both simulation results and experimental data, were conducted to verify the viability of the model. The results showed good agreement where almost all the modeling predictions were within 20% of the simulation results or the experimental data. Finally, the proposed modeling approach is more generally applicable to heat transfer analysis of heat exchangers with similar flow channel configurations to those considered in this study.
There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).
Ever-increasing wind turbine size has challenged predictive capabilities on several fronts. Here, to address part of the blade structural modeling uncertainty, a systematic model fidelity comparison study was conducted on commonly used finite elements. pyNuMAD was utilized to create beam, shell, and solid models of a 100 m long blade undergoing large static deflections. The solid model avoided the use of layered-solid elements by resolving core and facesheet layers. An unprecedented model with 73.7 million elements revealed insights that have never been possible from prior experimental and numerical studies. As compared to the solid element model, the tip deflection from the shell and beam model was found to be about 2% and 4.3% too low, respectively. The twist from the beam model was found to be about 5.6% too high, while the twist from shell model was 24% too low, though improvement was demonstrated with mesh refinement. The beam model adhesive stresses were more accurate than the shell model. Out-of-plane stresses were of great significance near geometric and material discontinuities, and neither the shell nor beam model captured these effects well. Failure predictions from beam, shell, or layered-solid models are unlikely to be reliable at trailing edges, adhesives, ply-drops, spar-cap boundaries.
Effective policymaking to achieve net zero greenhouse gas emissions demands an understanding of the complex drivers of, and barriers to, consumer adoption behavior via behaviorally realistic energy system models. Existing models tend to oversimplify by focusing on homogenized financial factors while neglecting consumer heterogeneity and non-monetary influences. This study develops and applies a comprehensive framework for evaluating the behavioral realism of consumer adoption models, informed by the adoption literature. It introduces a typology for factors influencing low-carbon technology adoption decisions: monetary and non-monetary factors relating to household characteristics, psychology, technological attributes, and contextual conditions. Next, reviews of the consumer adoption and decision-making literature identify the most influential adoption factor categories for distributed solar photovoltaics, electric vehicles, and air-source heat pumps. Finally, the extent to which a selection of energy system models accounts for these adoption factors is assessed. Existing models predominantly emphasize the economic aspects of technology, which are generally identified as the most important factors. Where the models fall short — in considering moderately important factor categories — sector-specific and agent-based models can offer more behaviorally realistic insights. This study sheds light on which types of factors are most important for consumer adoption decisions and investigates how well current models rise to the challenge of behavioral realism. The end-to-end analysis presented enables internally consistent comparisons across models and energy technologies. This research advances timely conversations on consumer adoption. It could inform more behaviorally realistic energy system modeling, and thereby more effective decarbonization policymaking.