Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Temporal data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Artificial Intelligence and Machine Learning Applications in Modern Power Systems

Machine learning (ML) and artificial intelligence (AI) algorithms offer valuable tools for the analysis and interpretation of large datasets. These tools have the capability to uncover insights that may not be readily apparent within these datasets. In recent years, the integration of ML and AI has become increasingly prevalent in various applications within the power system domain. One of the earliest instances of machine learning in power systems can be traced back to demand forecasting, where artificial neural networks were employed for short-term load forecasting. In contemporary power systems, an abundance of high-resolution geospatial and temporal data is generated at various time intervals, ranging from sub-seconds (Phasor Measurement Units or PMUs) to seconds (Supervisory Control and Data Acquisition or SCADA), minutes (Process Information or PI), and extending to days, months, and years. These datasets contain valuable information concerning system reliability and performance. This information holds the potential to offer critical insights into system operations, as well as solutions for predicting and mitigating contingencies to prevent cascading outages. Despite the immense power of machine learning tools, system operators, planners, and utilities often exhibit hesitancy in fully embracing AI-enabled system operations and planning. This cautious approach persists, even as numerous diverse applications of machine learning continue to emerge in the realm of power systems. In this chapter, our focus will delve deep into ML and AI applications tailored for power systems. These applications aim to furnish system operators with enhanced situational awareness and augment their decision-making capabilities, especially during challenging operating conditions. Specific areas of interest encompass root cause analyses of electricity market datasets and the strategic selection of representative samples from vast power system databases for training ML/AI models. Finally, the chapter will conclude with a short discussion on the future of ML/AI in power systems and possible directions that the industry is moving towards.

power system applications, machine learning (ML), ↗

A semantics-driven framework to enable demand flexibility control applications in real buildings

Decarbonising and digitalising the energy sector requires scalable and interoperable Demand Flexibility (DF) applications. Semantic models are promising technologies for achieving these goals, but existing studies focused on DF applications exhibit limitations. These include dependence on bespoke ontologies, lack of computational methods to generate semantic models, ineffective temporal data management and absence of platforms that use these models to easily develop, configure and deploy controls in real buildings. This paper introduces a semantics-driven framework to enable DF control applications in real buildings. The framework supports the generation of semantic models that adhere to Brick and SAREF while using metadata from Building Information Models (BIM) and Building Automation Systems (BAS). The work also introduces a web platform that leverages these models and an actor and microservices architecture to streamline the development, configuration and deployment of DF controls. The paper demonstrates the framework through a case study, illustrating its ability to integrate diverse data sources, execute DF actuation in a real building, and promote modularity for easy reuse, extension, and customisation of applications. The paper also discusses the alignment between Brick and SAREF, the value of leveraging BIM data sources, and the framework's benefits over existing approaches, demonstrating a 75% reduction in effort for developing, configuring, and deploying building controls.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Material Fracturing and Failure Simulation Datasets

Fracturing is a fundamental physics phenomena with broad relevance across multiple domains, ranging from infrastructure integrity, aerospace durability, reservoir production, and seismic events. We present a diverse dataset of simulated fracture evolution and material failure generated from two numerical solvers: the phase-field method and the combined finite-discrete element method (FDEM). These solvers differ in formulation, physical fidelity, and computational efficiency. The dataset includes five materials: PBX, anisotropic shale, tungsten, aluminum, and steel. For each, phase-field simulations span 400,000 cases: 200,000 under uniaxial tension and 200,000 under biaxial tension. The computationally expensive FDEM simulations include 90,000 split evenly among PBX, shale, and tungsten under uniaxial loading. All simulations begin with randomized initial fracture patterns. Each entry includes temporal data capturing fracture propagation dynamics. This comprehensive dataset is designed to support the development of foundational or surrogate machine learning approaches for predicting material failure. While no such models are introduced here, the dataset lays a robust foundation for advancing future research and innovation in these areas.

36 MATERIALS SCIENCE↗

Acoustic Tomography of the Atmosphere: A Large-Eddy Simulation Sensitivity Study

Accurate measurement of atmospheric turbulent fluctuations is critical for understanding environmental dynamics and improving models in applications such as wind energy. Advanced remote sensing technologies are essential for capturing instantaneous velocity and temperature fluctuations. Acoustic tomography (AT) offers a promising approach that utilizes sound travel times between an array of transducers to reconstruct turbulence fields. This study presents a systematic evaluation of the time-dependent stochastic inversion (TDSI) algorithm for AT using synthetic travel-time measurements derived from large-eddy simulation (LES) fields under both neutral and convective atmospheric boundary-layer conditions. Unlike prior work that relied on field observations or idealized fields, the LES framework provides a ground-truth atmospheric state, enabling quantitative assessment of TDSI retrieval reliability, sensitivity to travel-time measurement noise, and dependence on covariance model parameters and temporal data integration. A detailed sensitivity analysis was conducted to determine the best-fit model parameters, identify the tolerance thresholds for parameter mismatch, and establish a maximum spatial resolution. The TDSI algorithm successfully reconstructed large-scale velocity and temperature fluctuations with root mean square errors ( RMSE s) below 0.35 m/s and 0.12 K, respectively. Spectral analysis established a maximum spatial resolution of approximately 1.4 m, and reconstructions remained robust for travel-time measurement uncertainties up to 0.002 s. These findings provide critical insights into the operational limits of TDSI and inform future applications of AT for atmospheric turbulence characterization and system design.

17 WIND ENERGY↗

Dynamic Temporal Graph Sequence Data for Resilience-Oriented Distribution Network Reconfiguration

This dataset comprises temporal dynamic graph sequences generated from power grid simulations focused on grid reconfiguration to enhance resilience. The simulations model failure propagation under varying conditions, with nodes assigned distinct failure probabilities. For each time step, the dataset captures the evolution of node states (functional or failed) and features critical to grid operations, such as pv_output, load_profile, load_dispatch, dg_output, loss, and voltage. Node types include sources, normal loads, and nodes with specific equipment like PVs, micro turbines, or shunt capacitors. The dataset is structured to support the training of dynamic graph neural networks, facilitating research on node feature prediction and edge dynamics under failure scenarios. Three distinct configurations are included, providing a robust foundation for modeling power grid resilience.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Persistent global greening over the last four decades using novel long-term vegetation index data with enhanced temporal consistency

Advanced Very High-Resolution Radiometer (AVHRR) satellite observations have provided the longest global daily records from 1980s, but the remaining temporal inconsistency in vegetation index datasets has hindered reliable assessment of vegetation greenness trends. To tackle this, we generated novel global long-term Normalized Difference Vegetation Index (NDVI) and Near-Infrared Reflectance of vegetation (NIRv) datasets derived from AVHRR and Moderate Resolution Imaging Spectroradiometer (MODIS). We addressed residual temporal inconsistency through three-step post processing including cross-sensor calibration among AVHRR sensors, orbital drifting correction for AVHRR sensors, and machine learning-based harmonization between AVHRR and MODIS. After applying each processing step, we confirmed the enhanced temporal consistency in terms of detrended anomaly, trend and interannual variability of NDVI and NIRv at calibration sites. Our refined NDVI and NIRv datasets showed a persistent global greening trend over the last four decades (NDVI: 0.0008 yr -1 ; NIRv: 0.0003 yr -1 ), contrasting with those without the three processing steps that showed rapid greening trends before 2000 (NDVI: 0.0017 yr -1 ; NIRv: 0.0008 yr -1 ) and weakened greening trends after 2000 (NDVI: 0.0004 yr -1 ; NIRv: 0.0001 yr -1 ). These findings highlight the importance of minimizing temporal inconsistency in long-term vegetation index datasets, which can support more reliable trend analysis in global vegetation response to climate changes.

54 ENVIRONMENTAL SCIENCES↗

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

TRAILS Output Files

Overview This data repository contains ZIP files that store compressed versions of the output of running the WaterPaths utility planning and management tool in the DU Re-Evaluation mode (to download the tool, please see this GitHub repository). The tool was used to simulate the six-utility North Carolina Research Triangle problem. Details on the contents of each ZIP file can be seen below. Data details Temporal range: Weekly data for 2,344 weeks from 2015 to 2060 (45 years). Spatial range: Six water utilities in the North Carolina Research Triangle region (0: Chapel Hil/OWASA, 1: Durham, 2: Cary, 3: Raleigh, 4: Pittsboro, and 5: Chatham) File types: CSV and OUT Different solutions available The solution numbers correspond to the different pathway strategies (henceforth referred to as "solutions") discussed in paper's main and supporting text (abstract and link to the paper here). They are as follows: Sol92: The Durham-focused pathway strategy Sol132: The Raleigh-focused pathway strategy Sol140: The regionally-robust pathway strategy Objectives files These files can be accessed by unzipping solXX_objectives_pathways.zip that contains 1,000 Objectives_RDMXX_solsXX_to_XX.csv files. Each CSV file will consist of a row representing all the objective values for that specific solution, while every six columns represents the reliability, restriction frequency, infrastructure net present value ($ mil), peak financial cost, worst-case cost, and unit cost ($ per MG; in that order) for each of the six utilities. There will be 1,000 such files, denoting the performance of the six utilities across the 1,000 deeply uncertain states of the world (DU SOWs). Pathway files These files can be accessed by unzipping solXX_objectives_pathways.zip that contains 1,000 Pathways_sXX_RDMXX.out file. Each OUT corresponds to the set of infrastructure being triggered in a specific DU SOW, and each file will have the name file will consist of four tab-delimited columns that are described as follows: Realization: The realization in which an infrastructure options being triggered utility: The utility currently triggering infrastructure week: The week in which a specific infrastructure option is being triggered infra.: The infrastructure option being triggered If the OUT file contains only the header line, no infrastructure was triggered for that specific DU SOW. Policies files These files can be obtained by unzipping Policies.zip. Each of the 1,000 CSV files within the unzipped folder will contain weekly water use restriction policies for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: 0rest_m: restriction multiplier for utility 0 (values between 0 and 1) 1rest_m: restriction multiplier for utility 1 (values between 0 and 1) 2rest_m: restriction multiplier for utility 2 (values between 0 and 1) 3rest_m: restriction multiplier for utility 3 (values between 0 and 1) 4rest_m: restriction multiplier for utility 4 (values between 0 and 1) 5rest_m: restriction multiplier for utility 5 (values between 0 and 1) 0transf: transfer volume for utility 0 (in MGD) 1transf: transfer volume for utility 1 (in MGD) 2transf: transfer volume for utility 2 (in MGD) 3transf: transfer volume for utility 3 (in MGD) 4transf: transfer volume for utility 4 (in MGD) 5transf: transfer volume for utility 5 (in MGD) Water Sources files These files can be obtained by unzipping WaterSources_subset.zip. Each of the 100 CSV files within the unzipped folder will contain weekly state variables at each water source for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: Xvolume: available water volume from source X (in MGD) Xs_area: surface area of source X (in ACF) Xdemand: demand drawn from a water source from source X (in MGD) Xup_spill: upstream spillage from source X (in MGD) Xww_inflow: wastewater inflow from source X (in MGD) Xcatch_inflow: upstream catchment inflow to source X (in MGD) Xevap: evaporation multiplier for source X (values between 0 and 1) Xds_spill: downstream spillage from source X (in MGD) X_Y_alloc_cap: the allocated capacity from source X to utility Y (values between 0 and 1) X_Y_alloc_dem: the allocated demand from source X to utility Y (values between 0 and 1) Xtrmt_alloc_Y: the allocated treatment capacity from source X to utility Y (values between 0 and 1) Utilities files These files can be obtained by unzipping Utilities_subset.zip. Each of the 100 CSV files within the unzipped folder will contain weekly state variables at each utility for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: Xst_vol: total available storage volume of utility X (in MG) Xcapacity: total storage capacity of utility X (in MG) Xnet_inf: : net inflow for all storage infrastructure for utility X (in MGD) Xst_rof: short term ROF for utility X (values between 0 and 1) Xst_stor_rof: short-term storage ROF for utility X (values between 0 and 1) Xst_trmt_rof: short-term treatment ROF for utility X (values between 0 and 1) Xlt_rof: long-term ROF for utility X (values between 0 and 1) Xlt_stor_rof: long-term storage ROF for utility X (values between 0 and 1) Xlt_trmt_rof: long-term treatment ROF for utility X (values between 0 and 1) Xrest_demand: restricted demand for utility X (in MGD) Xunrest_demand: unrestricted demand for utility X (in MGD) Xunfulf_demand: unfulfilled demand for utility X (in MGD) Xwastewater: wastewater return for utility X (in MGD) Xtreat_capacity: total treatment capacity for utility X (in MG) Xcont_fund: reserve (contingency) fund balance for utility X Xins_pout: insurance payout for utility X (% annual volumetric revenue) Xins_price: insurance price for utility X (% annual volumetric revenue) Xinfra_npv: infrastructure net present value for utility ($mil) Xst_vol: total available storage volume of utility X (in MG) Xdebt_serv: debt service for utility X (usually once per year if the infrastructure is triggered; % annual volumetric revenue) Xstor_vol: total stored volume (in MGD) Xobs_ann_dem: observed annual demand for utility X (in MGD) Xproj_dem: projected annual demand for utility X (in MGD) Xpv_debt_serv: present value of debt service payments for utility X (% annual volumetric revenue) Xgross_rev: gross revenue for utility X ($mil) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

Artificial Intelligence↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

Estimating soybean yields from high-temporal-resolution multi-source data using deep learning

Accurate and timely crop yield prediction is crucial for ensuring food security and maintaining stable agricultural markets. In recent years, there has been a surge in interest in leveraging high-temporal-resolution, multi-source data for effective crop growth monitoring and yield estimation. A notable challenge arises from the difficulty in capturing the intricate interactions between variables across different time steps within these high-temporal-resolution time series datasets. This complexity hinders the reliable extraction of yield information from voluminous and often noisy datasets, especially during periods of extreme weather events. Here, in this study, we propose an Attention and Graph Isomorphism Network-enhanced Bi-directional Long Short-Term Memory network (AGB-LSTM) for estimating county-level soybean yield in the United States. This model integrates a diverse set of remote sensing data, including Near-Infrared Reflectance of Vegetation (NIRv), Sun-Induced chlorophyll Fluorescence (SIF), and Gross Primary Productivity (GPP), along with environmental covariates. The AGB-LSTM effectively leverages information related to crop yield from high-temporal-resolution time series data (5-days), achieving an accuracy of R²= 0.67 and rRMSE = 14.46%. This approach significantly outperforms traditional machine learning methods such as Random Forest (RF) (R²= 0.52, rRMSE = 17.36%) and Bi-LSTM (R²= 0.58, rRMSE = 16.17%). Sensitivity experiments with different time steps and ranges demonstrated that our model could accurately and stably predict yields 1 to 2 months before harvest. Moreover, data with a finer temporal resolution consistently improved prediction performance, resulting in an approximately 20% increase in and an approximately 20% decrease in rRMSE compared to using monthly composites. We also evaluated the robustness of the model under extreme climate events and observed strong performance (R²= 0.50, rRMSE = 21.32%). Finally, yield mapping for major soybean-producing regions in North America in 2023 revealed spatial patterns that closely matched USDA yield reports. Our findings suggest that the AGB-LSTM model is a promising and effective method for estimating yield and has notable potential for global crop yield forecasting.

Deep learning↗

Data Quality Monitoring for the Hadron Calorimeters Using Transfer Learning for Anomaly Detection

The proliferation of sensors brings an immense volume of spatio-temporal (ST) data in many domains, including monitoring, diagnostics, and prognostics applications. Data curation is a time-consuming process for a large volume of data, making it challenging and expensive to deploy data analytics platforms in new environments. Transfer learning (TL) mechanisms promise to mitigate data sparsity and model complexity by utilizing pre-trained models for a new task. Despite the triumph of TL in fields like computer vision and natural language processing, efforts on complex ST models for anomaly detection (AD) applications are limited. In this study, we present the potential of TL within the context of high-dimensional ST AD with a hybrid autoencoder architecture, incorporating convolutional, graph, and recurrent neural networks. Motivated by the need for improved model accuracy and robustness, particularly in scenarios with limited training data on systems with thousands of sensors, this research investigates the transferability of models trained on different sections of the Hadron Calorimeter of the Compact Muon Solenoid experiment at CERN. The key contributions of the study include exploring TL’s potential and limitations within the context of encoder and decoder networks, revealing insights into model initialization and training configurations that enhance performance while substantially reducing trainable parameters and mitigating data contamination effects.

47 OTHER INSTRUMENTATION↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 2 Sensor Data v2-1

This is the version v2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Please see v2-1 TEMPEST L2 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

Multimodal super-resolution: discovering hidden physics and its application to fusion plasmas

Understanding complex physical systems often requires integrating data from multiple diagnostics, each with limited resolution or coverage. We present a machine learning framework that reconstructs synthetic high-temporal-resolution data for a target diagnostic using information from other diagnostics, without direct target measurements during the inference. This multimodal super-resolution technique improves diagnostic robustness and enables monitoring even in case of measurement failures or degradation. Applied to fusion plasmas, our method targets edge-localized modes (ELMs), which can damage plasma-facing materials. By reconstructing super-resolution Thomson Scattering data from complementary diagnostics, we uncover fine-scale plasma dynamics and validate the role of resonant magnetic perturbations (RMPs) in ELM suppression through magnetic island formation. The approach provides new observation supporting the plasma profile flattening due to these islands. Our results demonstrate the framework’s ability to generate high-fidelity synthetic diagnostics, offering a powerful tool for ELM control development in future reactors like ITER. The approach is broadly transferable to other domains facing sparse, incomplete, or degraded diagnostic data, opening new avenues for discovery.

Jalalvand, Azarakhsh [Princeton Univ., NJ (United ↗

Data-Driven Modeling and Correction of Vehicle Dynamics

We develop a data-driven framework for learning and correcting nonautonomous vehicle dynamics. Physics-based vehicle models are often simplified for tractability and therefore exhibit inherent model-form uncertainty, motivating the need for data-driven correction. Moreover, nonautonomous dynamics are governed by time-dependent control inputs, which pose challenges in learning predictive models directly from temporal snapshot data. To address these, we reformulate the vehicle dynamics via a local parameterization of the time-dependent inputs, yielding a modified system composed ofa sequence of local parametric dynamical systems. Here, we approximate these parametric systems using two complementary approaches. First, we employ the dimension reduction and interpolation in parameter space (DRIPS) methodology to construct efficient linear surrogate models, equipped with lifted observable spaces and manifold-based operator interpolation. This enables data-efficient learning of vehicle models whose dynamics admit accurate linear representations in the lifted spaces. Second, for more strongly nonlinear systems, we employ flow map learning (FML), a deep neural network (DNN) approach that approximates the parametric evolution map without requiring special treatment of nonlinearities. We further extend FML with a transfer-learning-based model correction procedure, enabling the correction of misspecified prior models using only a sparse set of high-fidelity or experimental measurements, without assuming a prescribed form for the correction term. Through a suite of numerical experiments on unicycle, simplified bicycle, and slip-based bicycle models, we demonstrate that DRIPS offers robust and highly data-efficient learning of nonautonomous vehicle dynamics, while FML provides expressive nonlinear modeling and effective correction of model-form errors under severe data scarcity.

data-driven modeling↗

Daily, 30 m Resolution NDSI Data for the East River Watershed, CO for 2000-2020

This dataset contains daily Normalized Difference Snow Index (NDSI) values at 30 m spatial resolution for the East River watershed in Colorado, USA. The temporal range of these data includes water years 2001-2020. These data were created using the Spatial and Temporal Adaptive Reflectance Fusion Model (STARFM). This model fuses low spatial and high temporal resolution data from MODIS (500 m, daily) with high spatial and low temporal resolution data from Landsat (30 m, 16 days) to create a 30m synthetic daily snow product. This product allows for the analysis of historical snow covered area trends in the East River Watershed at fine spatiotemporal resolutions where it was not available previously. This research was performed as a part of the Department of Energy’s Subsurface Biogeochemical Research Program with the primary intent of better understanding the timing and spatial patterns of water delivery to the Critical Zone in mountain watersheds. Each .zip file contains one "water year" of data (October 1 - September 30; i.e., water year 2010 starts October 1, 2010 and ends September 30, 2011). Each zip file contains the following: STARFM daily Normalized Difference Snow Index (NDSI) fusion data files in GeoTiff format with one layer for each day between Landsat data acquisition dates (i.e., for dates of Landsat acquisition, the Landsat image is included for that date). The study area is located in an area of Landsat path overlap, so Landsat dates acquisitions are every 7-9 days. Landsat NDSI files containing the high spatial (30m), low temporal (7-9 days due to Landsat path overlap) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Dates for which no Landsat data were obtained are included as NoData layers. MODIS NDSI files containing the high temporal (daily), low spatial (500m) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Please note the MODIS data were resampled to 30m pixels for input into the STARFM model. The data have a scale factor of 10,000 and a no data value of -32767. The projection of all datasets is WGS 84 (EPSG: 4326), which has a latitude/longitude based degree resolution of 0.0002694946 X 0.0002694946, and approximates to the 30 m spatial resolution mentioned above. The Layer Index files in .csv format. They contain information for each layer in the above GeoTiff files regarding the corresponding date for each layer, the fraction of pixels in the image that contain valid data (missing data is due to either cloud cover or poor data quality; these values are not percent snow cover). Dates of Landsat overpass are indicated in these files. If no Landsat data were able to be obtained due to cloud cover or lack of Landsat Tier 1 data available on Google Earth Engine, this is also noted.

EARTH SCIENCE > CRYOSPHERE > SNOW/ICE↗

CERF: IM3 Projected Western US Power Plant Locations

Overview The Capacity Expansion Regional Feasibility (CERF) model is an open-source geospatial python package that provides new power plant locations at a 1km resolution. The model ingests U.S. state or regional-scale electricity system capacity expansion plans, such as those produced by the Global Change Analysis Model (GCAM-USA), and identifies feasible, site-specific locations for individual new power plants (renewable and non-renewable). CERF combines high-resolution geospatial suitability analyses with an economic algorithm that selects individual plant siting locations based on grid interconnection costs and the locational marginal value of new generation. The model incorporates a wide range of dynamic constraints and opportunities, such as protected lands, population density, existing infrastructure, and water availability. This dataset provides CERF power plant siting results for IM3 Phase 2 simulations across eight different scenarios for the Western US through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 CERF siting results in this dataset correspond to capacity expansion plans in the GCAM-USA IM3 Phase 2 simulation data and are available for each of the above scenarios. Data Details Temporal Range: 2015-2055 in 5-year timesteps. Note that 2015 is the experiment base year and 2020 and beyond represent model simulation years. Spatial Range: Plant locations are provided for the eleven states in the Western US including Arizona, California, Colorado, Idaho, Montana, New Mexico, Nevada, Oregon, Utah, Washington, and Wyoming. Spatial Resolution: 1 km-squared, provided in x and y coordinates Geospatial Projection: Albers Equal Area Conic (ESRI:102003) File Type: csv The dataset contains subdirectories for each of the eight scenarios described in the overview. Each scenario folder contains two subfolders with the following information: 1. Power Plant Data This directory contains a single .csv file of power plant locations for both pre-existing (non-CERF sited plants in operation in 2015) and new (CERF-sited) power plants across the temporal range along with additional CERF model output parameters for CERF-sited plants. Plant with a siting year earlier than 2020 correspond to facilities that are operational leading into the first timestep CERF simulation. For a more detailed description of CERF model output parameters, see the CERF model documentation. Note that the cerf_plant_id parameter is unique within each scenario file but not across scenario files. Parameter Descriptions scenario - Name of scenario cerf_plant_id - Unique siting identifier cerf_sited - If True, indicates that plant was sited by CERF model. If False, indicates pre-existing facility region_name - Name of region (state) tech_id - Technology ID tech_name - Full generation technology name inclusive of cooling type (if applicable) and additional characteristics tech_simple - Simplified generation technology type unit_size_mw - Power plant unit size (MW) xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) index - Index position in the flattend 2D array buffer_in_km - Exclusion buffer around site (km) sited_year - Year of siting retirement_year - Year of retirement lmp_zone - Locational marginal price (LMP) zone ID locational_marginal_price_usd_per_mwh - Locational marginal price ($/MWh) generation_mwh_per_year - Generation output (MWh/yr) operating_cost_usd_per_year - Cost of plant operations ($/yr) net_operational_value - Net operational value based on LMP and and operating costs ($/yr) interconnection_cost - Cost of interconnection for transmission & gas pipeline (if applicable) net_locational_cost -- Difference of interconnection cost and operating value ($/yr) capacity_factor_fraction - Capacity factor (fraction) carbon_capture_rate_fraction - Carbon capture rate (fraction) fuel_co2_content_tons_per_btu - Fuel CO2 content (tons/Btu) fuel_price_usd_per_mmbtu - Fuel price ($/MMBtu) fuel_price_esc_rate_fraction - Fuel price escalation rate (fraction) heat_rate_btu_per_kWh - Heat rate (Btu/kWh) lifetime_yrs - Technology lifetime for annuity (years) operational_life_yrs - Operational lifetime for retirement (years) variable_om_usd_per_mwh - Variable operation and maintenance costs of yearly capacity use ($/MWh) variable_om_esc_rate_fraction - Variable operation and maintenance costs escalation rate (fraction) carbon_tax_usd_per_ton - Carbon tax ($/ton) carbon_tax_esc_rate_fraction - Carbon tax escalation rate (fraction) 2. Storage Data This directory contains information on new and pre-existing energy storage facilities operational in each timestep along with various storage operational parameters. The 2015 timestep provides pre-existing energy storage data and corresponds with facilities that are operational leading into the first model simulation timestep. Note that coordinates in the storage files correspond to the interconnection point on the grid (substation location), not individual energy storage locations. Energy storage is added in a cumulative process at each given interconnection point. That is, each individual file provides the total operational storage capacity interconnected to the specified substation for the given timestep, inclusive of previously installed storage at that location and new storage installed in that timestep at that location. Parameters scenario - Name of scenario timestep - Simulation timestep name - Unique storage identifier s_typ - Type of energy storage technology (battery or pumped storage hydro) s_node - Node ID of interconnecting substation xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) charge_rate - Maximum charge rate (power capacity) of storage system (MW) discharge_rate - Maximum discharge rate (power capacity) of storage system (MW) duration - Duration of storage system (hours) max_SoC - Allowed maximum state of charge (energy capacity) of storage system (MWh) min_SoC -Allowed minimum state of charge (energy capacity) of storage system (MWh) charge_eff - Efficiency of charge (fraction between 0 and 1) discharge_eff - Efficiency of discharge (fraction between 0 and 1) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

CERF↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗