Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Time-series prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network

Mixed Delay/Nondelay Embeddings Based Neuromorphic Computing with Patterned Nanomagnet Arrays

Patterned nanomagnet arrays (PNAs) have been shown to exhibit a strong geometrically frustrated dipole interaction. Some PNAs have also shown emergent domain wall dynamics. Previous works have demonstrated methods to physically probe these magnetization dynamics of PNAs to realize neuromorphic reservoir systems that exhibit chaotic dynamical behavior and high-dimensional nonlinearity. These PNA reservoir systems from prior works leverage echo state properties and linear/nonlinear short-term memory of component reservoir nodes to map and preserve the dynamical information of the input time-series data into nondelay spatial embeddings. Such mappings enable these PNA reservoir systems to imitate and predict/forecast the input time series data. However, these prior PNA reservoir systems are based solely on the nondelay spatial embeddings obtained at component reservoir nodes. As a result, they require a massive number of component reservoir nodes, or a very large spatial embedding (i.e., high-dimensional spatial embedding) per reservoir node, or both, to achieve acceptable imitation and prediction accuracy. These requirements reduce the practical feasibility of such PNA reservoir systems. To address this shortcoming, we present a mixed delay/nondelay embeddings-based PNA reservoir system. Our system uses a single PNA reservoir node with the ability to obtain a mixture of delay/nondelay embeddings of the dynamical information of the time-series data applied at the input of a single PNA reservoir node. Our analysis shows that when these mixed delay/nondelay embeddings are used to train a perceptron at the output layer, our reservoir system outperforms existing PNA-based reservoir systems for the imitation of NARMA 2, NARMA 5, NARMA 7, and NARMA 10 time series data, and for the short-term and long-term prediction of the Mackey Glass time series data.

Ti, Changpeng

Do slow orbital periodicities appear in the record of earth's magnetic reversals?

Time-series spectral analysis has been performed on the dates of geomagnetic reversals of the last 20 Myr BP and earlier. Possible evidence is found from the presence of high spectral peaks for two very long periodicities, 0.4 Myr and 1.3 Myr, that may be associated with slow variations of the earth's orbital eccentricity as predicted by Berger. However, statistical significance tests and a number of other arguments do not confirm the two detections.

Stothers, Richard B.

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.

36 MATERIALS SCIENCE

Evaluation of data driven low-rank matrix factorization for accelerated solutions of the Vlasov equation

Low-rank methods have shown success in accelerating simulations of a collisionless plasma described by the Vlasov equation, but still rely on computationally costly linear algebra every time step. We propose a data-driven factorization method using artificial neural networks, specifically with convolutional layer architecture, that trains on existing simulation data. At inference time, the model outputs a low-rank decomposition of the distribution field of the charged particles, and we demonstrate that this step is faster than the standard linear algebra technique. Numerical experiments show that the method achieves comparable reconstruction accuracy for interpolation tasks, generalizing to unseen test data in a manner beyond just memorizing training data; patterns in factorization also inherently followed the same numerical trend as those within algebraic methods (e.g., truncated singular-value decomposition). However, when training on the first 70% of a time-series data and testing on the remaining 30%, the method fails to meaningfully extrapolate. Despite this limiting result, the technique may have benefits for simulations in a statistical steady-state or otherwise showing temporal stability. These results suggest that while the model offers a computationally efficient alternative for datasets with temporal stability, its current formulation is best suited for interpolation rather than for predicting future states in time-evolving systems. This study thus lays the groundwork for further refinement of neural network-based approaches to low-rank matrix factorization in high-dimensional plasma simulations.

97 MATHEMATICS AND COMPUTING

Regional climate change predictions from the Goddard Institute for Space Studies high resolution GCM

A new diagnostic tool is developed for examining relationships between the synoptic scale circulation and regional temperature distributions in GCMs. The 4 x 5 deg GISS GCM is shown to produce accurate simulations of the variance in the synoptic scale sea level pressure distribution over the U.S. An analysis of the observational data set from the National Meteorological Center (NMC) also shows a strong relationship between the synoptic circulation and grid point temperatures. This relationship is demonstrated by deriving transfer functions between a time-series of circulation parameters and temperatures at individual grid points. The circulation parameters are derived using rotated principal components analysis, and the temperature transfer functions are based on multivariate polynomial regression models. The application of these transfer functions to the GCM circulation indicates that there is considerable spatial bias present in the GCM temperature distributions. The transfer functions are also used to indicate the possible changes in U.S. regional temperatures that could result from differences in synoptic scale circulation between a 1XCO2 and a 2xCO2 climate, using a doubled CO2 version of the same GISS GCM.

Crane, Robert G.

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method or recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification mode to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed.

Merrill, W.

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method of recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification made to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed. Previously announced in STAR as N82-20339

Merrill, W.

A deep generative model for deciphering cellular dynamics and in silico drug discovery in complex diseases

Human diseases are characterized by intricate cellular dynamics. Single-cell transcriptomics provides critical insights, yet a persistent gap remains in computational tools for detailed disease progression analysis and targeted in silico drug interventions. Here we introduce UNAGI, a deep generative neural network tailored to analyse time-series single-cell transcriptomic data. This tool captures the complex cellular dynamics underlying disease progression, enhancing drug perturbation modelling and screening. When applied to a dataset from patients with idiopathic pulmonary fibrosis, UNAGI learns disease-informed cell embeddings that sharpen our understanding of disease progression, leading to the identification of potential therapeutic drug candidates. Validation using proteomics reveals the accuracy of UNAGI’s cellular dynamics analysis, and the use of the fibrotic cocktail-treated human precision-cut lung slices confirms UNAGI’s predictions that nifedipine, an antihypertensive drug, may have anti-fibrotic effects on human tissues. UNAGI’s versatility extends to other diseases, including COVID, demonstrating adaptability and confirming its broader applicability in decoding complex cellular dynamics beyond idiopathic pulmonary fibrosis, amplifying its use in the quest for therapeutic solutions across diverse pathological landscapes.

Neural Network

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Primary production estimates from recordings of solar-stimulated fluorescence in the equatorial Pacific at 150 deg W

Biological, optical, and hydrographical data were collected on the WEC88 cruise along 150 deg W and during a 6-day time-series station on the equator during February/March 1988. This area was characterized by a subsurface chlorophyll maximum (SCM), located at 50-70 m depth at the equator and descending down to 120-125 m at the north and south end of the transect. Highest primary production rates were near-surface and confined to the equatorial region and stations between 7 deg and 11 deg N. To determine the relationship between solar-stimulated fluorescence (centered at 683 nm wavelength) and primary production, a production-fluorescence model based on phytoplankton physiology and marine optics is described. Results of model calculations predict that there is a linear relation between production and fluorescence. A comparison between morning and midday measurements of the production-fluorescence relation showed that there was some difference between the two, whereas evening measurements, on the other hand, were distinctly different from the morning/midday ones. This seems to suggest that diurnal variations contribute significantly to variability in the quantum yield of photochemical processes. The ratio of the quantum yield of photosynthesis to the quantum yield of fluorescence ranged between 0.24 and 0.44 molC/Ein for all stations. The highest value for this ratio occurred at the equatorial stations, indicating that latitudinal variability could have an effect on the production-fluorescence relation.

Stegmann, P. M.

The brighter-fatter effect in the JWST MIRI Si:As IBC detectors: I. Observations, impact on science, and modeling

Context.The Mid-Infrared Instrument (MIRI) on board theJames WebbSpace Telescope (JWST) uses three Si:As impurity band conduction (IBC) detector arrays. The output voltage level of each MIRI detector pixel is digitally recorded by sampling up the ramp. For uniform or low-contrast illumination, the pixel ramps become nonlinear in a predictable way, but in areas of high contrast, the nonlinearity curve becomes much more complex. The origin of the effect is poorly understood and currently not calibrated out of the data. Aims.We provide observational evidence of the brighter-fatter effect (BFE) in MIRI conventional and high-contrast coronagraphic imaging, low-resolution spectroscopy, and medium-resolution spectroscopy data, and we investigate the physical mechanism that gives rise to the effect on the MIRI detector pixel raw voltage integration ramps. Methods.We used public data from the JWST/MIRI commissioning and Cycle 1 phase. We also developed a numerical electrostatic model of the MIRI detectors using a modified version of the publicPoisson_CCDcode. Results.We find that the physical mechanism behind the BFE manifesting in MIRI data is fundamentally different to that of charge-coupled devices and photodiode arrays such as the Hawaii-XRG near-infrared detectors used by the NIRISS, NIRCam, and NIRSpec instruments on board JWST. Observationally, the BFE makes the JWST MIRI data yield 10–25% larger point sources and spectral line profiles as a function of the relative level of de-biasing of neighboring detector pixels. This broadening impacts the MIRI absolute flux calibration, time-series observations of faint companions, and point spread function modeling and subtraction. We also find that the intra-pixel 2D profile of the shrinking Si:As IBC detector depletion region directly impacts the accuracy of the pixel ramp nonlinearity calibration model.

Astronomy & Astrophysics

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University

Hydroboost

HydroBoost is the most realistic revenue optimization tool for the hybridization of hydropower and battery energy storage systems to date. The innovative representation of how operators actually schedule hydropower in practice results in more realistic predictions of revenue and operations. Unlike other optimization tools, HydroBoost generates forecast energy prices with uncertainty to use in the optimization. This allows HydroBoost to give users a range of potential revenue with an upper bound using the perfect foresight pricing and a lower bound using a naive persistence forecast model. Additional forecast can be generated and used in the optimization, such as additive models, random forest, and neural networks to give further insight into potential revenue. HydroBoost has been designed to be applicable for both run-of-river and reservoir storage sites. The primary focus is on the day-ahead market and requires year-long data with an hour time-step. All time-series input and constraints are contained in an Excel worksheet for convince. The user will run the forecasting generation first with a Python script to give the optimization model the necessary requirements. Next the optimization is ran using Julia and results are generated and stored into a directory as csv files. HydroBoost includes an additional module to generate figures based on the results of the optimization simulation. The results help analyze the results and users to draw insights into how the hydro and battery systems are operated and the revenue each is producing. Additionally, the difference between the perfect foresight model and models that include forecast can easily be inspected.

Phillips, TylerB. [Idaho National Laboratory (INL)

Groundwater table elevation and temperature from 2015 to 2024 at the Lower Montane site in the East River Watershed, Colorado.

This groundwater level elevation and temperature data package is aimed at improving the predictive understanding of hydro-biogeochemical processes at the lower montane site in the East River Watershed, Colorado. The dataset is obtained using pressure transducers placed in shallow wells in the floodplain. This dataset contains data from wells with Location ID's ER-DOW (alias DO1West), ER-DOE (alias DO2East), ER-MBA1 (alias M1Bend1), ER-MBA2 (alias M1Bend2), ER-UPW (alias UP1West), ER-UPM (alias UP2), ER-UPE (alias UP3East). Another dataset contains the data from wells with Location ID's ER-CPA1 to ER-CPA6. Each file contains the water level elevation and the water temperature. Water level elevation has been obtained using the barometric pressure from the pressure transducer (Hobos sensor) in the well, barometric pressure from a sensor in air located at the same site (lower montane), depth from top-of-casing (TOC) to sensor measurement point, and TOC elevation. Data have been checked with a few measurements of water table depths. A real-time kinematic (RTK) global positioning system (GPS) has been used to survey the TOC (data in file Well_Location.csv). The water level elevation is given in UTM13N Geoid2012AB. While depth to water level is not present in the data files, it can be easily calculated with the TOC and distance to ground provided in the GPS coordinate file. The dataset quality is discussed in Collection/Analysis section of the methods. Time-series of measurements were initially added to the archive for the period 2015 to 2019, and later updated with time-series until 2024 (end of data collection). The dataset contains 8 *.csv data files, and 3 *.csv metadata files. Feel free to contact the author with any questions or collaboration interests. The publication year was updated from "2020" to "2025" to reflect the revised version of this dataset.

54 ENVIRONMENTAL SCIENCES

Groundwater table elevation and temperature from 2015 to 2024 across Meander C at the Lower Montane site in the East River Watershed, Colorado.

This groundwater level elevation and temperature data package is aimed at improving the predictive understanding of hydro-biogeochemical processes at the lower montane site in the East River Watershed, Colorado. The dataset is obtained using pressure transducers placed in shallow wells in the floodplain. This dataset contains data from wells ER-CPA1 to ER-CPA6. Another dataset contains the data from wells at nearby Locations. Each file contains the water level elevation and the water temperature. Water level elevation has been obtained using the barometric pressure from the pressure transducer (Hobos sensor) in the well, barometric pressure from a sensor in air located at the same site (lower montane), depth from top-of-casing (TOC) to sensor measurement point, and TOC elevation. Data have been checked with a few measurements of water table depths. A real-time kinematic (RTK) global positioning system (GPS) has been used to survey the TOC (data in file Well_Location.csv). The water level elevation is given in UTM13N Geoid2012AB. While depth to water level is not present in the data files, it can be easily calculated with the TOC and distance to ground provided in the GPS coordinate file. The dataset quality is discussed in Collection/Analysis section of the methods. Time-series of measurements were initially added to the archive for the period 2015 to 2019, and later updated with time-series until 2024 (end of data collection). The dataset contains 7 *.csv data files, and 3 *.csv metadata files. Feel free to contact the author with any questions or collaboration interests. The publication year was updated from "2020" to "2025" to reflect the revised version of this dataset.

54 ENVIRONMENTAL SCIENCES

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]