SEARCH · Engineering Papers
Results for “Data analysis”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Statistical data analysis of x-ray spectroscopy data enabled by neural network accelerated Bayesian inference
Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.
An R shiny graphical user interface for highprecision mass spectrometric data analysis
• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility
What Is Data Analysis?
A quick guide to understand your data and use it to tell compelling stories. Data analysis helps develop insights for research projects, planning interventions, or systematic information gathering. This guide highlights important aspects of the data analysis process.
BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions
Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.
High Crimes and Misdemeanors in Data Analysis
Invited talk going over some of the various data analysis problems that have occurred over the years including the most extreme example of falsifying data.
Adapt: A Weather Radar Data Analysis and Nowcasting Platform for Informed Adaptive Scanning
SF-26-021 Adapt is a data processing platform for real-time data analysis, short term prediction of targets convective cells and tracking for archived data. It provides tools for downloading, processing, segmenting, projecting, analyzing, and visualizing storm cell data from weather radar. The pipeline includes cell detection, motion estimation using optical flow, cell property extraction, and persistence to NetCDF and SQLite/Parquet for guiding adaptive scanning.
Topological Data Analysis for Particulate Gels
Soft gels, formed via the self-assembly of particulate materials, exhibit intricate multiscale structures that provide them with flexibility and resilience when subjected to external stresses. Here, this work combines particle simulations and topological data analysis (TDA) to characterize the complex multiscale structure of soft gels. Our TDA analysis focuses on the use of the Euler characteristic, which is an interpretable and computationally scalable topological descriptor that is combined with filtration operations to obtain information on the geometric (local) and topological (global) structure of soft gels. We reduce the topological information obtained with TDA using principal component analysis (PCA) and show that this provides an informative low-dimensional representation of the gel structure. We use the proposed computational framework to investigate the influence of gel preparation (e.g., quench rate, volume fraction) on soft gel structure and to explore dynamic deformations that emerge under oscillatory shear in various response regimes (linear, nonlinear, and flow). Our analysis provides evidence of the existence of hierarchical structures in soft gels, which are not easily identifiable otherwise. Moreover, our analysis reveals direct correlations between topological changes of the gel structure under deformation and mechanical phenomena distinctive of gel materials, such as stiffening and yielding. In summary, we show that TDA facilitates the mathematical representation, quantification, and analysis of soft gel structures, extending traditional network analysis methods to capture both local and global organization.
Active multi-mode data analysis to improve fault diagnosis in AHUs
Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.
HDG-1 Fiber Bragg grating data analysis
The main goal of the High dose graphite 1 Advanced test reactor experiment was to study nuclear grade graphite at high fluences. Additional supplementary optical fiber instrumentation was added to this long duration experiment for instrumentation development purposes. The supplementary instrumentation consisted of two pure silica core, fluorine doped cladding optical fibers each etched with 9 fiber Bragg gratings, one fiber being heat treated for 9 hours at 750 C and 16 hours at 750 C, the other being heat treated for 24 hours at 550 C and 48 hours at 650 C. Fiber Bragg gratings are known to have issues of measurement drift when in high temperature and high radiation environments like what is encountered in the Advanced test reactor. At the culmination of this experiment, the optical fibers saw ~1.3E21 n/cm2 total fluence, which is at the highest fluences that fiber Bragg gratings have been studied to date. Reported here is the analysis of this data including radiation induced shift and changes in sensitivity.
Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)
This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.
MapsTorch : automatic differentiation for X-ray fluorescence data analysis
X-ray fluorescence (XRF) is a popular spectroscopy technique for elemental analysis. Spectrum fitting and parameter tuning are at the core of XRF analysis and are conventionally manually intensive, especially for synchrotron experiments involving large amounts of diverse samples. This work introduces the automatic differentiation (AD) technique to XRF and an open-source package called MapsTorch. By transforming an analytical model of the XRF spectrum into a differentiable computation graph with AD, MapsTorch enables robust optimization of parameters and elemental intensities. We evaluate MapsTorch by conducting computational experiments on a large number of historical synchrotron XRF datasets and compare its performance with the currently practiced fitting tool NLopt. The results show that MapsTorch consistently achieves high-quality fits and often leads to better fitting quality than NLopt, particularly in tasks such as initial spectrum fitting and elemental intensity refinement. The robust performance of MapsTorch paves the way for developing automated and high-throughput XRF data analysis workflows to handle the increasing data volumes expected from next-generation synchrotron facilities.
Using feature importance as an exploratory data analysis tool on Earth system models
Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.
Evaluation of Machine Learning Models for Automated Data Analysis in In-Service Nuclear Power Plant Inspections
The commercial nuclear power industry is facing a potential shortage of certified nondestructive evaluation (NDE) analysts to meet future in-service inspection demands. Automated data analysis (ADA) currently supports human inspectors in tasks such as eddy current evaluations for steam generator examinations. Machine learning (ML) systems are nearing the capability to pass performance demonstration tests for ultrasonic testing (UT) inspections of reactor pressure vessel upper head penetrations in nuclear power plants (NPPs). Current research and development is focused on assisted analysis (AA) of ADA versus fully automated examinations. This presentation will cover assessment of ML flaw detection on dissimilar metal weld (DMW) piping joints.
Open-Source Data Analysis Tool for Spectral Small-Angle X-ray Scattering Using Spectroscopic Photon-Counting Detector
Spectral small-angle X-ray scattering (sSAXS) is a powerful technique for material characterization from thicker samples by capturing elastic X-ray scattering data in angle- and energy-dispersive modes at small angles. This approach is enabled by the use of a 2D spectroscopic photon-counting detector that provides energy and position information of scattered photons when a sample is irradiated by a polychromatic X-ray beam. Here, we describe an open-source tool with a graphical interface for analyzing sSAXS data obtained from a 2D spectroscopic photon-counting detector with a large number of energy bins. The tool takes system geometry parameters and raw detector data to output 1D scattering patterns and a 2D spatially-resolved scattering map in the energy range of interest. We validated these features using data from samples of caffeine powder with well-known scattering peaks. This open-source tool will facilitate sSAXS data analysis for various material characterization applications.
Doppler Backscattering Data Analysis and Integrated Modeling with OMFIT
One Modeling Framework for Integrated Tasks (OMFIT) is a widely used software tool in the magnetic fusion research community. OMFIT provides magnetic fusion energy researchers with a framework for the development of special-purpose physics modules. This paper describes an OMFIT physics module pertaining to the Doppler Backscattering (DBS) fusion plasma diagnostic. DBS measures density fluctuations and flow velocity through plasma scattering of electromagnetic waves. The OMFIT DBS module was developed to analyze experimental DBS data and facilitate modeling of DBS systems installed on multiple tokamak devices. The OMFIT DBS module is designed to support several analysis workflows: detailed analysis of experimental data, experimental planning, and theory-based synthetic diagnostic modeling. The DBS module uses integrated modeling by leveraging other OMFIT physics modules to perform tasks related to DBS, e.g. ray/beam–tracing simulations, edge-localized mode–synchronized data analysis, magnetic equilibrium reconstruction, and fitting kinetic profile data. Furthermore, this paper describes several supported workflows and serves a reference for the OMFIT DBS module.
Data for 3-Hydroxypropionic Acid Recovery from Fermentation Broth through Novel Downstream Processing: Technoeconomic Analysis
This study develops and validates a simplified, fully solvent-free downstream processing (DSP) strategy for high-purity recovery of 3-hydroxypropionic acid (3-HP) from real fermentation broth containing 62.3 g/L of 3-HP. Optimized activated carbon treatment achieved 98% color removal, while Amberlite IRA-67 was operated at pH 4.5 and 30 °C to minimize product loss. This is the first integrated demonstration of a fully solvent-free DSP enabling recovery of bio-based 3-HP as both a solid sodium salt and a concentrated aqueous solution, supported by techno-economic analysis. At lab scale, the process achieved 77.3% recovery of sodium 3-HP with 83.2% (w/w) purity and produced a 30% (w/v) aqueous solution. Techno-economic analysis yielded minimum selling prices of $0.551/kg for the solution and $0.892/kg for the salt, both below target thresholds for cost-competitive bio-acrylic acid production. Overall, these results demonstrate an efficient, scalable, and economically viable industrial pathway for 3-HP recovery.
In situ Synchrotron X‐ray Metrology Boosted by Automated Data Analysis for Real‐time Monitoring of Cathode Calcination
Abstract Synchrotron X‐ray‐based in situ metrology is advantageous for monitoring the synthesis of battery materials, offering high throughput, high spatial and temporal resolution, and chemical sensitivity. However, the rapid generation of massive data poses a challenge to on‐site, on‐the‐fly analysis needed for real‐time process monitoring. Here, a weighted lagged cross‐correlation (WLCC) similarity approach is presented for automated data analysis, which merges with in situ synchrotron X‐ray diffraction metrology to monitor the calcination process of the archetypal nickel‐based cathode, LiNiO 2 . The WLCC approach, incorporating variables that account for peak shifts and width changes associated with structural transformations, enables rapid extraction of phase progression within 10 seconds from tens of diffraction patterns. Details are captured, from initial precursors to intermediates and the final layered LiNiO 2 , providing information for agile on‐site adjustments during experiments and complementing post hoc diffraction analysis by offering insights into early‐stage phase nucleation and growth. Expanding this data‐powered platform paves the way for real time calcination process monitoring and control, which is pivotal to quality control in battery cathode manufacturing.