Learning Flexible Time-windowed Granger Causality Integrating Heterogeneous Interventional Time Series Data
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This study investigates vegetation dynamics in boreal forests of Interior Alaska, focusing on topography, fire history, and climate influences. The study area includes Bonanza Creek Experimental Forest (BCEF) and surrounding region, categorized by topography (upland, floodplain, lowland) and fire history. Using Mann–Kendall trend and Theil–Sen slope analyses on Landsat-derived spectral metrics: Normalized Difference Vegetation Index (NDVI) and Normalized Burn Ratio (NBR), we observed a shift from browning to greening trends, particularly in historically burned areas. The photosynthetic activity in burned upland converged with unburned areas ~30 years post-fire, coincident with a shift towards deciduous dominance during post-fire succession. Normalized Difference Moisture Index (NDMI) trends revealed a significant increase in vegetation moisture content across all topographies. We introduce Effective Seasonal Precipitation Index (ESPI), which combines prior-year annual precipitation with current-year spring snow depth. Its positive correlation with NDMI highlights its potential for monitoring vegetation moisture dynamics at the landscape scale. Furthermore, by correlating dendrochronology-based climate indices, we found strong correlation between NDMI and normalized Supplemental Precipitation Index (nSPI), across topographies. Overall, this research provides critical insights into how climate and fire influence interior boreal vegetation, highlighting the effects of increased precipitation, and topography on shaping differential vegetation responses across the landscape.
Controlling radiation doses at potential radioactive facilities is critical to ensuring the safety of both personnel and the public. At the Thomas Jefferson National Accelerator Facility (JLab), multiple sensors are deployed around the three experimental halls to monitor key parameters, including single-beam current, energy levels, current leakage, and radiation values during accelerator operations. In this study, we developed a Multi-task Transformer model, MTL_TX, to accurately estimate radiation doses at sensor locations based on historical data, with the aim of enhancing safety in accelerator facilities and surrounding public areas. To improve estimation accuracy, we integrated two innovative components into the proposed model: hierarchical feature embedding (HFE) and multi-level decomposition attention (MDA). Additionally, the multi-task learning (MTL) framework effectively leverages correlations among multiple sensors, enabling individual estimations for each sensor. MTL_TX achieved outstanding results on data collected in 2018, with an MSE of 0.1464, an RMSE of 0.2353, and an R 2 score of 0.8584. Furthermore, when trained on 2018 data, MTL_TX exhibited excellent generalization capability to unseen datasets from 2016 to 2019, achieving an MSE of 0.1407, an RMSE of 0.2263, and an R 2 score of 0.8831. These results demonstrate a significant improvement over existing state-of-the-art models.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.
Patterned nanomagnet arrays (PNAs) have been shown to exhibit a strong geometrically frustrated dipole interaction. Some PNAs have also shown emergent domain wall dynamics. Previous works have demonstrated methods to physically probe these magnetization dynamics of PNAs to realize neuromorphic reservoir systems that exhibit chaotic dynamical behavior and high-dimensional nonlinearity. These PNA reservoir systems from prior works leverage echo state properties and linear/nonlinear short-term memory of component reservoir nodes to map and preserve the dynamical information of the input time-series data into nondelay spatial embeddings. Such mappings enable these PNA reservoir systems to imitate and predict/forecast the input time series data. However, these prior PNA reservoir systems are based solely on the nondelay spatial embeddings obtained at component reservoir nodes. As a result, they require a massive number of component reservoir nodes, or a very large spatial embedding (i.e., high-dimensional spatial embedding) per reservoir node, or both, to achieve acceptable imitation and prediction accuracy. These requirements reduce the practical feasibility of such PNA reservoir systems. To address this shortcoming, we present a mixed delay/nondelay embeddings-based PNA reservoir system. Our system uses a single PNA reservoir node with the ability to obtain a mixture of delay/nondelay embeddings of the dynamical information of the time-series data applied at the input of a single PNA reservoir node. Our analysis shows that when these mixed delay/nondelay embeddings are used to train a perceptron at the output layer, our reservoir system outperforms existing PNA-based reservoir systems for the imitation of NARMA 2, NARMA 5, NARMA 7, and NARMA 10 time series data, and for the short-term and long-term prediction of the Mackey Glass time series data.
These data contain the hourly plant-level power time series from PLUSWIND. The time series is derived from the HRRR model and corrected for density and losses. The PLUSWIND dataset hosted on the Wind Data Hub here (https://a2e.energy.gov/project/pluswind) contains additional versions with meteorological corrections. Data are provided for years 2018 - 2021. The data hosted at this location include conversion from capacity factor output (normalized output) to total hourly production (units of MW).
In this research, we examine the relationship between aerial IR defect analysis and photovoltaic (PV) performance data for twelve utility- and commercial-scale solar sites in the United States. To do this, we fuse the site diagram geoJSON's, aerial infrared thermography (aIRT) defect analyses, and associated inverter time series, allowing for a direct comparison between site defects and time series data. Defect analyses were provided by Zeitview, under its Solar Insights platform. Following the data fusion process, we look at the relationship between system performance and aIRT defects. We investigate the relationship between degradation and hotspot defects, as well as the relationship between AC power data and offline strings and misaligned modules. In general, system degradation was not affected by long-term or balance-of-system (BoS) defects as they occurred infrequently in the data set. However, for one system, a near statistically significant relationship (p-value=0.057) was found when comparing the degradation of inverter blocks with several multi-hotspot defects to all other inverter blocks without this particular defect. There was strong alignment when comparing short-term recoverable module defects such as stuck trackers and offline strings to time series data. In general, we found that when an inverter block has more than 80% of modules flagged for one of these defects, its AC power time data is flat-lined and the inverter block is not producing.
Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.
The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset is part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with other energy technologies. This dataset contains inputs and outputs from simulations of a floating marine hydrokinetic turbine over approximately half a tidal cycle (~6.6 hours). Inflow conditions were derived from field measurements in Alaska’s Cook Inlet and represent a tidal environment in which the current speed ramps from near 0 m/s to a peak of 3 m/s and back. The original acoustic doppler current profiler dataset is publicly available on the Marine and Hydrokinetic Data Repository. In a full tidal cycle, the flow reverses and the rotor would reorient; this reversal was not modeled. In the Cook Inlet campaign , turbulence intensity was similar in both directions. Two inflow cases are included. In the first case, labeled “raw” in the files, the measured current time series was used directly in the InflowWind module of OpenFAST. Speed and direction were applied as a function of time and elevation, uniformly in the horizontal direction. With full spatial coherence, this approach captures high turbulent variability and results in pronounced power fluctuations, so it is considered a conservative, near-worst-case representation of loading. In the second case, labeled “average” in the files, a 30-minute moving average was applied to extract the slowly varying mean speed. The residual fluctuations about this mean were used to generate spatially varying, full-field turbulence inputs with TurbSim, giving a more physically realistic representation of the inflow across the rotor disk. Two random realizations were used to produce distinct inflow conditions for two OpenFAST simulations representing a two-turbine array. The same turbulence intensity is applied across the full time series, producing larger fluctuations at the start and end, where the mean speed is low. The second case is the more appropriate framework for performance and power assessment but overpredicts turbulence at lower flow speeds and underpredicts it at higher speeds. As the floating platform moves and the rotor changes its x-position, Taylor’s frozen turbulence hypothesis used by InflowWind assumes a constant rather than a time-varying mean velocity, introducing some inaccuracy in the velocity plane sampling. The turbine modeled is the 500-kW Reference Model 1, a horizontal-axis two-bladed hydrokinetic turbine on a four-column floating semisubmersible substructure . Simulations were performed using OpenFAST v4.1 with the Reference Open Source Controller (ROSCO) v2.10. All input files required to reproduce the simulations are included. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel . This unit supports up to 2.5 MW, but NLR has only a single 1.25-MW stack. The datasets report hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. The system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operating current of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The simulated tidal turbine time series data was translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz. Each zip file represents a single tidal electrolysis experiment and is named: {technology}_{inflow method}_{number of 500 kW tidal turbines connected} For instance, “tidal-500kW-RM1_average_2.zip” is a 6-hour experiment using the 500-kW tidal reference model, scaled by 2x (1-MW) to better match the electrolyzer maximum of 1.25MW, fed with the 30-minute moving average current case. Each zip folder contains the following files: A .csv file of raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. A .csv file combines all tidal profiles as "combined_tidal_experiments.csv." A separate experiment, “characterization_200.zip,” shows the MC250 electrolyzer steady-state response with 30-minute load steps over 5 hours and is accessible with this entry.
Here, we present a method that uses protein levels to predict times series of metabolite concentrations. Understanding this type of pathway dynamics is important in order to predict the behavior of the pathway and, more pragmatically, to be able to design biological systems (such as strains bioengineered to produce chemical products) reliably. Typically, for this purpose, kinetic models consisting of differential equations based on the Michaelis-Menten dynamics have been used in the past. However, these methods can rarely produce good fits to measured data time series. Possibly, this happens because the kinetic constants are unknown or are different from the ones measured in vivo, or perhaps because Michaelis-Menten dynamics is not a satisfactory description. In order to improve the predictive nature of these kinetic models we have eliminated the Michaelis-Menten description of pathway dynamics and we have substituted it by algorithms that automatically learn these dynamics from previously obtained metabolomics and proteomics data using machine learning approaches. Specifically, kinetic deep learning uses deep learning to map proteomics time series to metabolite concentration time series, instead of learning the first metabolite derivative and integrating in (as in the first version of kinetic learning). This approach is shown to provide good to excellent results with a data set specifically collected for this purpose.
The dataset contains temperature measurements from distributed temperature profiling (DTP) systems (Dafflon et al., 2022; Wielandt et al., 2022; Wang et al., 2024a; Fiolleau et al., 2024) deployed vertically above the ground surface at a large number of locations from 2021 to 2024. The research is designed to improve understanding of the local heterogeneity in snow depth and snow thermal insulation dynamics, as well as their interactions in a discontinuous permafrost region (Wang et al., 2025). The DTP systems were deployed at 96 locations in a watershed along the Nome-Teller road at mile marker 27 (T27) and at 54 locations on a hillslope along the Kougarok road at mile marker 64 (K64) in the Seward Peninsula, Alaska. The probe location information is stored in Probe_locations_T27.csv and Probe_locations_K64.csv. Temperature measurements were recorded at 15-minute intervals using high-precision digital sensors (accuracy: ±0.1°C, resolution: 0.0078°C). The temperature probes, either 1.4 m or 1.6 m long, contain sensors spaced every 5 cm or 10 cm along their length. The temperature data are stored in compressed files following the format: DTP_snow_air_temperature_(site)_(start)_(end).zip, where site is either T27 or K64, and start and end represent the time series period. Within each ZIP file, individual CSV files are named by probe ID and contain temperature records at different heights above the ground surface.This dataset also includes derived snow depth time series over three snow seasons, estimated from temperature measurements. Snow depth was estimated by identifying the consecutive sensor pair that exhibited the largest drop in high-frequency temperature fluctuations (detailed in the methods). These data are stored in: Snow_depths_flags_(site)_(start)_(end).csv, which includes snow depth time series and corresponding quality flags (defined in the methods) from different probes. Additionally, the dataset includes derived metrics and supporting measurements at selected locations over two snow seasons, contributing to the manuscript of Wang et al., 2025. These locations were chosen based on the availability of high-quality snow depth time series during both seasons. The additional data include: (1) Air temperature proxies measured from the top sensors on the pole when they were not buried by snow, stored in Air_temperature_proxies_(site)_(start)_(end).csv (2) Ground interface temperature, recorded at 3 cm above the ground, stored in Ground_interface_temperature_(site)_(start)_(end).csv (3) Site characteristics, including vegetation height, elevation, and the topographic position index (TPI) within a 50 m radius, stored in Selected_probe_locations_gps_vegheight_tpi_elevation_(site).csv. These metrics were derived from 1 m resolution summer LiDAR-based digital elevation models and digital surface models from Singhania et al., 2023, DOI:10.5440/1832016. Metadata files include data descriptions (_dd.csv) for tabular data. All included files are listed and described in xxxx_flmd.csv.This dataset is an updated version of a previous archive (Wang et al., 2024b, DOI: 10.15485/2475020), incorporating multiple seasons and improved snow depth estimation. Please note that due to large amount of information present in this dataset, many specificities associated with the acquisition of snow temperature, air temperature proxy and estimation of snow depth, and the future archiving of additional datasets on the soil temperature, thaw depth and soil characteristics at these locations, the author would welcome being contacted by people planning to use this dataset.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).
This presentation uses a simulation by the author in 1991 and newer developements in 2025 to illustrate strategies to address problems that arise when steady state assumption is applied in time series simulation: 1) high resolution time series data; 2) distirbution functiion; 3) machine learning. The presentation does not report new findings (previously published material is cited).
This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.
Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.
The computer code assumes that time series representing the physical signals of the vehicle have been extracted from the CAN bus. The main input of the computer code is the multivariate time series representation of the signals in the CAN bus. The computer code cluster these time series using agglomerative hierarchical clustering from benign and attack datasets. Based on this, it generates probability distributions from the similarity of the obtained clusters based in each scenario---benign and attack---using the CluSim method (https://github.com/Hoosier-Clusters/clusim). Finally, it compares how a new data collection compares with the previous distribution to provide and probability score for an intrusion.