Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH

What Is Data Analysis?

A quick guide to understand your data and use it to tell compelling stories. Data analysis helps develop insights for research projects, planning interventions, or systematic information gathering. This guide highlights important aspects of the data analysis process.

29 ENERGY PLANNING, POLICY, AND ECONOMY

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE

Data Analysis for GOPEX Image Frames

This article describes the data analysis based on the image frames received at the Solid State Imaging (SSI) camera of the Galileo Optical Experiment (GOPEX) demonstration conducted between December 9 and 16, 1992. Laser uplink was successfully established between the ground and the Galileo spacecraft during its second Earth-gravity-assist phase in December 1992. SSI camera frames were acquired which contained images of detected laser pulses transmitted from the Table Mountain Facility (TMF), Wrightwood, California, and the Starfire Optical Range (SOR), Albuquerque, New Mex/co. Laser pulse data were processed using standard image-processing techniques at the Multimission Image Processing Laboratory (MIPL) for preliminary pulse identification and to produce public re/ease images. Subsequent image analysis corrected for background noise to measure received pulse intensities. Data were plotted to obtain histograms on a dmly basis and were then compared with theoretical results derived from applicable weak-turbulence and strong-turbulence considerations. This article describes processing steps and compares the theories with the experimented results. Quantitative agreement was found in both turbulence regimes, and better agreement would have been found, given more received laser pulses. Future experiments should consider methods to reliably measure low-intensity pulses, and through experimented planning to geometrically locate pulse positions with greater certainty.

B M Levine

Adapt: A Weather Radar Data Analysis and Nowcasting Platform for Informed Adaptive Scanning

SF-26-021 Adapt is a data processing platform for real-time data analysis, short term prediction of targets convective cells and tracking for archived data. It provides tools for downloading, processing, segmenting, projecting, analyzing, and visualizing storm cell data from weather radar. The pipeline includes cell detection, motion estimation using optical flow, cell property extraction, and persistence to NetCDF and SQLite/Parquet for guiding adaptive scanning.

Raut, Bhupendra Ashokrao [Argonne National Laborat

Active multi-mode data analysis to improve fault diagnosis in AHUs

Faults in heating, ventilation and air conditioning systems can lead to increased energy consumption, occupant comfort issues, and reduced equipment lifetime. Commercial fault detection and diagnosis (FDD) tools has been increasingly deployed in U.S. commercial buildings. While they are helping to achieve energy efficiency and operational reliability, there remain gaps in their fault diagnostic capabilities. The diagnostic results often contain multiple distinct candidate root causes (CRCs) or offer no insight into CRCs. This study developed a novel active rule-based multi-mode data analysis method to enhance diagnostic resolution by applying proven rule sets and additional new rules to data from multiple known operational modes. The proposed method was demonstrated using enhanced air handling unit performance assessment rule sets and validated with the simulated data of two air handling units. New metrics, namely, reduced number of CRCs and improvement ratio, were developed to quantify the improvement of fault diagnostic resolution. The validation results showed that the proposed method effectively reduced the number of CRCs in contrast to analyzing data solely for a single mode of operation. It achieved a median improvement ratio of 80% in 19 test cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

HDG-1 Fiber Bragg grating data analysis

The main goal of the High dose graphite 1 Advanced test reactor experiment was to study nuclear grade graphite at high fluences. Additional supplementary optical fiber instrumentation was added to this long duration experiment for instrumentation development purposes. The supplementary instrumentation consisted of two pure silica core, fluorine doped cladding optical fibers each etched with 9 fiber Bragg gratings, one fiber being heat treated for 9 hours at 750 C and 16 hours at 750 C, the other being heat treated for 24 hours at 550 C and 48 hours at 650 C. Fiber Bragg gratings are known to have issues of measurement drift when in high temperature and high radiation environments like what is encountered in the Advanced test reactor. At the culmination of this experiment, the optical fibers saw ~1.3E21 n/cm2 total fluence, which is at the highest fluences that fiber Bragg gratings have been studied to date. Reported here is the analysis of this data including radiation induced shift and changes in sensitivity.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING

MapsTorch : automatic differentiation for X-ray fluorescence data analysis

X-ray fluorescence (XRF) is a popular spectroscopy technique for elemental analysis. Spectrum fitting and parameter tuning are at the core of XRF analysis and are conventionally manually intensive, especially for synchrotron experiments involving large amounts of diverse samples. This work introduces the automatic differentiation (AD) technique to XRF and an open-source package called MapsTorch. By transforming an analytical model of the XRF spectrum into a differentiable computation graph with AD, MapsTorch enables robust optimization of parameters and elemental intensities. We evaluate MapsTorch by conducting computational experiments on a large number of historical synchrotron XRF datasets and compare its performance with the currently practiced fitting tool NLopt. The results show that MapsTorch consistently achieves high-quality fits and often leads to better fitting quality than NLopt, particularly in tasks such as initial spectrum fitting and elemental intensity refinement. The robust performance of MapsTorch paves the way for developing automated and high-throughput XRF data analysis workflows to handle the increasing data volumes expected from next-generation synchrotron facilities.

X-ray fluorescence

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)

Evaluation of Machine Learning Models for Automated Data Analysis in In-Service Nuclear Power Plant Inspections

The commercial nuclear power industry is facing a potential shortage of certified nondestructive evaluation (NDE) analysts to meet future in-service inspection demands. Automated data analysis (ADA) currently supports human inspectors in tasks such as eddy current evaluations for steam generator examinations. Machine learning (ML) systems are nearing the capability to pass performance demonstration tests for ultrasonic testing (UT) inspections of reactor pressure vessel upper head penetrations in nuclear power plants (NPPs). Current research and development is focused on assisted analysis (AA) of ADA versus fully automated examinations. This presentation will cover assessment of ML flaw detection on dissimilar metal weld (DMW) piping joints.

36 MATERIALS SCIENCE

Test Data Analysis of the Thermodynamic Vent System-Augmented Top Spray Injector Liquid Nitrogen Transfer Experiments

Traditionally, a cryogenic tank must be pre-chilled to some “target” temperature before the main vent valve can be closed to attempt a non-vented fill (NVF) of cryogenic liquid propellant. This methodology is particularly attractive for performing in-space transfer of cryogens due to the unknown location of the liquid/vapor interface in microgravity and the high likelihood of venting liquid if the vent valve is opened during transfer. This paper presents in-depth test data analysis of a Thermodynamic Vent System (TVS) augmented injector used for cryogenic tank chilldown and fill experiments of a thin-walled Titanium tank. Eight tests were conducted using liquid nitrogen across a range of inlet conditions and boundary conditions, and three different chilldown/fill methods. For four of the tests, the injector sprays liquid into the tank as normal, but also uses a TVS heat exchanger to cool the metallic injector itself as well as the main incoming liquid stream. Results show that using the TVS augmented injector simplifies transfer operation via enhanced condensation at the injector surface at the cost of sacrificing only a small amount of propellant.

Thermodynamic Vent System

Doppler Backscattering Data Analysis and Integrated Modeling with OMFIT

One Modeling Framework for Integrated Tasks (OMFIT) is a widely used software tool in the magnetic fusion research community. OMFIT provides magnetic fusion energy researchers with a framework for the development of special-purpose physics modules. This paper describes an OMFIT physics module pertaining to the Doppler Backscattering (DBS) fusion plasma diagnostic. DBS measures density fluctuations and flow velocity through plasma scattering of electromagnetic waves. The OMFIT DBS module was developed to analyze experimental DBS data and facilitate modeling of DBS systems installed on multiple tokamak devices. The OMFIT DBS module is designed to support several analysis workflows: detailed analysis of experimental data, experimental planning, and theory-based synthetic diagnostic modeling. The DBS module uses integrated modeling by leveraging other OMFIT physics modules to perform tasks related to DBS, e.g. ray/beam–tracing simulations, edge-localized mode–synchronized data analysis, magnetic equilibrium reconstruction, and fitting kinetic profile data. Furthermore, this paper describes several supported workflows and serves a reference for the OMFIT DBS module.

Doppler backscattering

Data for 3-Hydroxypropionic Acid Recovery from Fermentation Broth through Novel Downstream Processing: Technoeconomic Analysis

This study develops and validates a simplified, fully solvent-free downstream processing (DSP) strategy for high-purity recovery of 3-hydroxypropionic acid (3-HP) from real fermentation broth containing 62.3 g/L of 3-HP. Optimized activated carbon treatment achieved 98% color removal, while Amberlite IRA-67 was operated at pH 4.5 and 30 °C to minimize product loss. This is the first integrated demonstration of a fully solvent-free DSP enabling recovery of bio-based 3-HP as both a solid sodium salt and a concentrated aqueous solution, supported by techno-economic analysis. At lab scale, the process achieved 77.3% recovery of sodium 3-HP with 83.2% (w/w) purity and produced a 30% (w/v) aqueous solution. Techno-economic analysis yielded minimum selling prices of $0.551/kg for the solution and $0.892/kg for the salt, both below target thresholds for cost-competitive bio-acrylic acid production. Overall, these results demonstrate an efficient, scalable, and economically viable industrial pathway for 3-HP recovery.

Bioproducts

Brazilian CBP - Technoeconomic analysis data

This data is related to the paper entitled "Techno-economic analysis of sugarcane bagasse and straw conversion into cellulosic ethanol via consolidated bioprocessing". That features the evaluation of sugarcane bagasse and straw conversion to ethanol at stand-alone facilities generating electricity from residues. The following scenarios were evaluated: Conventional, featuring hydrothermal pretreatment, fungal cellulase, and yeast fermentation (current commercial standard); Mid-term consolidated bioprocessing (CBP), relying on bagasse solubilization without pretreatment or cotreatment; and Mature CBP, incorporating cotreatment but no pretreatment and considering significant technological advance of the CBP. Available here are the spreadsheets used for Material and Energy balance calculation, Capital and Operational costs estimation and Cash flow analysis. Also available are the description and python code used for Monte Carlo analysis of the ethanol and capital investment variations. This data can be used as a source to implement other techno-economic analysis in the biorefinary context.

09 BIOMASS FUELS

Hydrology Copilot: A Cloud-Native Ai System for Hydrological Data Analysis

The emergence of AI-driven Earth observation systems promises to broaden access to petabyte-scale geospatial data beyond domain specialists. However, translating this vision into operational scientific infrastructure requires addressing fundamental challenges in data virtualization, code transparency, and domain-specific reasoning. We present Hydrology Copilot, a cloud-native AI framework for natural-language-driven analysis of Earth observation data. To demonstrate operational capabilities at scale, we implement the system using NASA's North American Land Data Assimilation System version 3 (NLDAS-3), which provides surface meteorological forcing and land-surface model output across North and Central America at 1-km resolution, from which drought diagnostics are derived. The system integrates five core contributions: (1) scalable data virtualization using Kerchunk-based cloud optimized access, achieving a 1.5 to 4.6 times improvement in I/O latency across benchmark queries spanning regional single-day extractions (4.6 times speedup) to continental monthly aggregations (1.5 times speedup); (2) transparent code generation through Microsoft Azure AI Foundry agents that expose executable Python workflows for scientific verification; (3) persistent conversational memory enabling multi-turn analytical discourse across sessions; (4) intelligent query validation that enforces dataset boundaries and resolves ambiguous requests before execution; and (5) a multi-agent architecture coordinating query parsing, code generation, and visualization. We evaluate the system through drought-monitoring workflows, demonstrating reliable code generation, accurate results validated against reference computations and the operational U.S. Drought Monitor, and efficient operation across increasingly complex tasks. By bridging natural-language interfaces with rigorous hydrological analysis, Hydrology Copilot advances beyond proof-of-concept demonstrations to provide a deployable framework for operational Earth science applications.

Data virtualization

Establishing Data Analysis Pipeline for Bulk ATAC-Seq Datasets

We developed an analysis pipeline for transposase-accessible chromatin sequencing (ATAC-Seq) data derived from bulk samples, which brings together publicly available R packages in addition to command-line tools designed for analysis of bulk ATAC-Seq data and can be run on any computer running a Linux-like operating system such as Ubuntu or Apple OSX.

97 MATHEMATICS AND COMPUTING

AGC-2 Graphite Preirradiation Data Analysis Report

This report describes the specimen loading order and documents all preirradiation examination material property measurement data for graphite specimens contained within the Second Advanced Graphite Capsule (AGC 2) irradiation capsule. The AGC 2 capsule is the second in six planned irradiation capsules comprising the Advanced Graphite Creep (AGC) test series. The AGC test series is used to irradiate graphite specimens in order to garner quantitative data necessary for predicting the irradiation behavior and operating performance of new nuclear grade graphites. This testing will ascertain the in service behavior of the graphite for pebble bed and prismatic very high temperature reactor designs. Similar to the First Advanced Graphite Capsule (AGC 1) preirradiation examination report, material property tests were conducted on specimens from 18 nuclear grade graphite types. However, AGC 2 tested an increased number of specimens (i.e., 512) prior to loading them into the AGC 2 irradiation assembly. All AGC 2 specimen testing was conducted at Idaho National Laboratory from July 2009 to August 2010. This report also details the specimen loading methodology for graphite specimens inside the AGC 2 irradiation capsule. The AGC 2 capsule design requires “matched pair” creep specimens that have similar dose levels above and below the neutron flux profile mid plane. This provides similar specimens with and without an applied load. Analysis in this document utilizes the neutron flux profile calculated for the AGC 2 capsule design, the capsule dimensions, and the size (i.e., length) of the selected graphite specimens to create a stacking order that produces “matched pairs” of graphite specimens above and below the AGC 2 capsule elevation mid point, thus providing specimens with similar neutron dose levels.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS