Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

TEAMER: Triton Systems Oscillating Water Column Modeling Data and Report

This dataset provides the output of six Wave Energy Converter Simulator (WEC-Sim) simulations and accompanying documentation for the modeling of Triton Systems' oscillating water column (OWC) system at tank scale (validated using available data for tuning the model, Tests 1-2) and deployment scale (for which no validation data is available, Tests 4-6). Included are the output data in a MATLAB file structure, a comprehensive report on the modeling and design of the Triton OWC system, and a link to the WEC-Sim GitHub page. This work was supported by funding from TEAMER RFTS 5 (Request for Technical Support).

16 TIDAL AND WAVE POWER↗

Data-model files associated with the manuscript "The Effects of Spatial and Temporal Resolution of Gridded Meteorological Forcing on Watershed Hydrological Responses" (Shuai et al., 2022 HESS)

This data package contains the model inputs and outputs used in "The Effects of Spatial and Temporal Resolution of Gridded Meteorological Forcing on Watershed Hydrological Responses" (Shuai et al., 2022 HESS). The data.zip file contains the data used to drive the model simulations. The model.zip file contains the XML input file for ATS. The notebook.zip file contains the Jupyter notebooks for pre- and post- processing model results. The figures.zip file contains the raw figures associated with the manuscript.Meteorological forcing plays a critical role in accurately simulating the watershed hydrological cycle. With the advancement of high-performance computing and the development of integrated watershed models, simulating the watershed hydrological cycle at high temporal (hourly to daily) and spatial resolution (10s of meters) has become efficient and computationally affordable. These hyperresolution watershed models require high resolution of meteorological forcing as model input to ensure the fidelity and accuracy of simulated responses. In this study, we utilized the Advanced Terrestrial Simulator (ATS), an integrated watershed model, to simulate surface and subsurface flow and land surface processes using unstructured meshes at the Coal Creek Watershed near Crested Butte (Colorado). We compared simulated watershed hydrologic responses including streamflow, and distributed variables such as evapotranspiration, snow water equivalent (SWE), and groundwater table driven by three publicly available, gridded meteorological forcings (GMFs) -- Daily Surface Weather and Climatological Summaries (Daymet), Parameter-elevation Regressions on Independent Slopes Model (PRISM), and North American Land Data Assimilation System (NLDAS). By comparing various spatial resolutions (ranging from 400 m to 4 km) of PRISM, the simulated streamflow only becomes marginally worse when spatial resolution of meteorological forcing is coarsened to 4 km (or 30% of the watershed area). However, the 4 km resolution has much worse performance than finer resolution in spatially distributed variables such as SWE. Using temporally disaggregated PRISM, we compared models forced by different temporal resolutions (hourly to daily), sub-daily resolution preserves the dynamic watershed responses (e.g., diurnal fluctuation of streamflow) that are absent in results forced by daily resolution. Conversely, the simulated streamflow shows better performance using daily resolution compared to that using sub-daily resolution. Our findings suggest that the choice of GMF and its spatiotemporal resolution depends on the quantity of interest and its spatial and temporal scale, which may have important implications on model calibration and watershed management decisions.

54 ENVIRONMENTAL SCIENCES↗

Data for "Examining Organic Acid Production Potential and Growth-Coupled Strategies in Issatchenkia orientalis Using Constraint-Based Modeling"

Growth-coupling product formation can facilitate strain stability by aligning industrial objectives with biological fitness. Organic acids make up many building block chemicals that can be produced from sugars obtainable from renewable biomass. Issatchenkia orientalis is a yeast strain tolerant to acidic conditions and is thus a promising host for industrial production of organic acids. Here, we use constraint-based methods to assess the potential of computationally designing growth-coupled production strains for I. orientalis that produce 22 different organic acids under aerobic or microaerobic conditions. We explore native and engineered pathways using glucose or xylose as the carbon substrates as proxy constituents of hydrolyzed biomass. We identified growth-coupled production strategies for 37 of the substrate-product pairs, with 15 pairs achieving production for any growth rate. We systematically assess the strain design solutions and categorize the underlying principles involved.

Bioproducts↗

Second-generation downscaled earth system model data using generative machine learning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. As with the first Sup3rCC version, the data still include temperature, wind speed and direction at multiple heights, pressure, three components of downwelling solar radiation, and relative humidity—all at 4-kilometer (km) hourly resolution over the contiguous United States. This is a 25x spatial enhancement and 24x temporal enhancement of the source 100-km daily-average ESM data. This extension of the Sup3rCC dataset includes data from six ESMs from two shared socioeconomic pathways (SSPs) totaling 400 years of data with multiple future projections of changing meteorological conditions. The scenario selection was based on a structured evaluation of historical ESM skill and comprehensive representation of possible trajectories of future climate change in temperature, humidity, precipitation, solar irradiance, and near-surface wind speeds. The inclusion of multiple future projections is intended to enable users to assess key drivers of un 36 certainty and variability. All data are double-bias corrected, resulting in a product that can be used out-of-the-box for energy system analysis with minimal historical bias. The potential applications of Sup3rCC data extend to various topics in renewable energy resource assessment, energy systems modeling, and grid resilience studies. High-resolution future meteorological projections are critical for evaluating the effects of changing meteorological conditions on renewable energy generation, energy demand, and for optimizing energy storage and grid infrastructure. The 4-km hourly resolution of the downscaled data enables understanding of spatial and temporal variability at the scales necessary for energy system operational planning. In addition, the dataset can support risk assessments by providing detailed information on possible future extreme weather events and long-term meteorological variability at scales relevant to energy infrastructure. By offering an enhanced representation of possible future meteorological conditions, the second-generation Sup3rCC dataset enables more precise modeling of energy resilience and adaptation strategies in response to changing meteorological conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Landmark-embedded Gaussian process with applications for functional data modeling

In practice, we often need to infer the value of a target variable from functional observation data. A challenge in this task is that the relationship between the functional data and the target variable is very complex: the target variable not only influences the shape but also the location of the functional data. In addition, due to the uncertainties in the environment, the relationship is probabilistic, that is, for a given fixed target variable value, we still see variations in the shape and location of the functional data. To address this challenge, we present a landmark-embedded Gaussian process model that describes the relationship between the functional data and the target variable. A unique feature of the model is that landmark information is embedded in the Gaussian process model so that both the shape and location information of the functional data are considered simultaneously in a unified manner. Gibbs-Metropolis-Hasting algorithm is used for model parameters estimation and target variable inference. The performance of the proposed framework is evaluated by extensive numerical studies and a case study of nano-sensor calibration.

42 ENGINEERING↗

Data for Impact of Vertical and Seasonal Variation in Leaf Traits on Simulating Soybean Canopy Photosynthesis via 1D and 3D Modeling

Accurate modeling of photosynthesis is crucial for predicting crop productivity and quantifying the carbon cycle in agroecosystems. Leaf traits are essential inputs for modeling canopy photosynthesis. Yet, many existing models still use fixed plant functional type (PTF)-based values to parameterize leaf traits under a big-leaf or two-big-leaf assumption, neglecting their vertical profiles and seasonal changes. This simplification may introduce significant uncertainties in estimating gross primary productivity (GPP). In this study, we simulated soybean GPP and tested the effects of vertical and seasonal variation in three key leaf photosynthetic traits: the maximum carboxylation rate at 25 °C (Vcmax25), leaf chlorophyll content (LCC), and leaf mass per area (LMA) in the 1D-SCOPE and 3D-Helios models. Weekly field measurements were conducted during the growing season of 2024 to support the simulation. We designed ten leaf trait parameterization schemes by incorporating different combinations of vertical profiles and seasonal changes, while assuming homogeneous canopy architecture in both models. Our results revealed that Vcmax25 vertical and seasonal variation had the strongest influence on simulated GPP in both 1D and 3D models, while LCC and LMA effects were minimal. Particularly, the scheme with an empirically parameterized Vcmax25 profile achieved comparable performance to the scheme with the measured Vcmax25 profile. Both 1D-SCOPE and 3D-Helios accurately modeled GPP (SCOPE: R2 = 0.87, Bias = 0.55 µmol m⁻² s⁻¹; Helios: R2 = 0.9, Bias = 0.22 µmol m⁻² s⁻¹) under the most complex scheme, and their responses to vertical and seasonal variation in leaf traits were consistent, demonstrating the robustness of our findings. Based on our findings, we propose a scalable framework for parameterizing leaf traits to improve GPP simulations. This study contributes to improving the representation of leaf trait dynamics in canopy-level photosynthesis models, potentially enhancing our ability to predict crop productivity and understand agroecosystem carbon dynamics.

Photosynthesis↗

YEASTRACT+: a portal for the exploitation of global transcription regulation and metabolic model data in yeast biotechnology and pathogenesis

Abstract YEASTRACT+ (http://yeastract-plus.org/) is a tool for the analysis, prediction and modelling of transcription regulatory data at the gene and genomic levels in yeasts. It incorporates three integrated databases: YEASTRACT (http://yeastract-plus.org/yeastract/), PathoYeastract (http://yeastract-plus.org/pathoyeastract/) and NCYeastract (http://yeastract-plus.org/ncyeastract/), focused on Saccharomyces cerevisiae, pathogenic yeasts of the Candida genus, and non-conventional yeasts of biotechnological relevance. In this release, YEASTRACT+ offers upgraded information on transcription regulation for the ten previously incorporated yeast species, while extending the database to another pathogenic yeast, Candida auris. Since the last release of YEASTRACT+ (January 2020), a fourth database has been integrated. CommunityYeastract (http://yeastract-plus.org/community/) offers a platform for the creation, use, and future update of YEASTRACT-like databases for any yeast of the users’ choice. CommunityYeastract currently provides information for two Saccharomyces boulardii strains, Rhodotorula toruloides NP11 oleaginous yeast, and Schizosaccharomyces pombe 972h-. In addition, YEASTRACT+ portal currently gathers 304 547 documented regulatory associations between transcription factors (TF) and target genes and 480 DNA binding sites, considering 2771 TFs from 11 yeast species. A new set of tools, currently implemented for S. cerevisiae and C. albicans, is further offered, combining regulatory information with genome-scale metabolic models to provide predictions on the most promising transcription factors to be exploited in cell factory optimisation or to be used as novel drug targets. The expansion of these new tools to the remaining YEASTRACT+ species is ongoing.

Teixeira, Miguel Cacho (ORCID:0000000256766174)↗

Modeling Data Movement Performance on Heterogeneous Architectures

The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.

97 MATHEMATICS AND COMPUTING↗

Block Island Noise Modeling Data

Noise propagation near Block Island was simulated to assess environmental impacts of impact pile driving during wind turbine construction. Computational models complement field-recorded acoustic data, providing insights into sound attenuation, spectral variability, and propagation dynamics. The dataset includes: 1. Propagation Models: Simulated underwater sound fields documenting sound pressure and directional variability across frequency bands and distances. 2. Spectral Analysis (LTSA): Long-term averages and processed outputs calculating acoustic intensity over time. 3. Visualization Files: Graphs, 2D/3D simulation results, and reference calculations used in sound modeling.

17 WIND ENERGY↗

Northern Pacific Turbulence Intensity Model Data

The dataset is a subset of simulations carried out for the northern Pacific region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the lidar buoys deployed off the coast of California (Humbolt and Morro Bay) are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing. Simulations are performed from November to December in 2020 and May through June in 2021. Model outputs every 10 minutes, permitting comparison with corresponding lidar observations.

17 WIND ENERGY↗

Gulf of Mexico Turbulence Intensity Model Data

The dataset is a subset of simulations carried out for the Gulf of Mexico region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing. Simulations are performed from January through June in 2020 and output every 10 minutes, which permits comparison with corresponding lidar observations.

17 WIND ENERGY↗

Gulf of Mexico Turbulence Intensity Model Data

The dataset is a subset of simulations carried out for the Gulf of Mexico region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from the NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing. Simulations are performed from January through June in 2020 and output every 10 minutes, which permits comparison with corresponding lidar observations.

17 WIND ENERGY↗

Mid-Atlantic Turbulence Intensity Model Data

The dataset is a subset of simulations carried out for the mid-Atlantic region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Air-Sea Interaction Tower (ASIT) of Woods Hole Oceanographic Institution’s Martha's Vineyard Coastal Observatory are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing. Simulations are performed from February through June in 2020 and output every 10 minutes, which permits comparison with corresponding lidar observations.

17 WIND ENERGY↗

Model America: Data and Models for every U.S. Building

The 5-year goal of the “Model America” concept was to generate a model of every building in the United States. This data repository delivers on that goal with "Model America v1". Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,715,609 buildings detected in the United States. Of this number, 122,146,671 (97.2%) buildings resulted in a successful generation and simulation of a building energy model. This dataset includes the full 125 million buildings. Future updates may include additional buildings, data improvements, or other algorithmic model enhancements in "Model America v2". This dataset contains OSM and IDF zip files for every U.S. county. Each zip file contains the generated buildings from that county. The .csv input data contains the following data fields: 1. ID - the Unique Building Identifier (UBID), generated using the Pacific Northwest National Laboratory (PNNL) BuildingID framework 2. Centroid - building center location in latitude/longitude (from Footprint2D) 3. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 4. State_abbr - state name 5. Area - estimate of total conditioned floor area (ft2) 6. Area2D - footprint area (ft2) 7. Height - building height (ft) 8. NumFloors - number of floors (above-grade) 9. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 10. CZ - ASHRAE Climate Zone designation 11. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 12. Standard - building vintage This data is made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy’s (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). Update (September 23, 2025): We corrected the ID field in all state-level.csv input files to ensure one-to-one consistency with the corresponding .osm and .idf output files. The schema and file structure are unchanged; only the values in the ID column were modified. No files were added or removed, and the .zip bundles (containing .osm / .idf) are unchanged. The corrected .csv inputs were re-extracted in March 2025 from the original data generated ~ 2021 (Theta supercomputer runs), and published here to align input IDs with model outputs. Update (September 6, 2026): The Model America dataset was updated to replace the previous building ID field with the Unique Building Identifier (UBID), using the Pacific Northwest National Laboratory (PNNL) BuildingID framework. UBIDs provide standardized, location-based identifiers for individual building footprints and improve interoperability with other building and geospatial datasets. The data files containing the previous building identifiers were updated to include UBIDs. This update standardizes building identification; the underlying Model America building characteristics and energy simulation results were not recomputed as part of this update.

54 ENVIRONMENTAL SCIENCES↗

A Primer on Dose-Response Data Modeling in Radiation Therapy

An overview of common approaches used to assess a dose response for radiation therapy–associated endpoints is presented, using lung toxicity data sets analyzed as a part of the High Dose per Fraction, Hypofractionated Treatment Effects in the Clinic effort as an example. Each component presented (eg, data-driven analysis, dose-response analysis, and calculating uncertainties on model prediction) is addressed using established approaches. Specifically, the maximum likelihood method was used to calculate best parameter values of the commonly used logistic model, the profile-likelihood to calculate confidence intervals on model parameters, and the likelihood ratio to determine whether the observed data fit is statistically significant. The bootstrap method was used to calculate confidence intervals for model predictions. Correlated behavior of model parameters and implication for interpreting dose response are discussed.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Harmonic Modeling, Data Generation and Analysis of Power Electronics-Interfaced Residential Loads

The share of electronics-based residential load is expected to rise as devices such as variable frequency drives (VFDs), electric vehicle chargers, and inverter-based distributed energy resources (DERs), e.g., photovoltaic (PV) systems become more common. These loads may introduce significant harmonics into power networks that need to be closely studied in order to perform accurate load modeling and forecasting. However, it can be difficult to obtain harmonic-rich voltage and current data - necessary for identifying accurate load models - for residential electrical loads. Recognizing this need, we identify and model a set of electronics-based end-use loads and DERs in an electromagnetic transients program (EMTP) tool for a residence with a single- phase split-phase supply. Further, a procedure is developed to model harmonic interactions between end-use loads connected to the same non-ideal supply voltage in a residential setting. Finally, a harmonic-rich dataset produced via the proposed procedure is utilized to identify frequency coupling matrix (FCM) based load model. Numerical results demonstrate the accuracy of the model, and explore model identifiability with limited data points.

harmonics, power quality, load modeling↗

Model America - data and models of every U.S. building

The 5-year goal of the 'Model America' concept was to generate a model of every building in the United States. This data repository delivers on that goal. Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,714,640 buildings detected in the United States and this dataset contains 122,930,327 (97.8%) buildings which resulted in a successful simulation. Future, annual updates have been proposed that may include additional buildings, data improvements, or other algorithmic enhancements. This dataset of 122.9 million buildings includes: Models (state_county.zip) - OpenStudio (v3.1.0) and EnergyPlus (v9.4) building energy models. Please note that the download requires the free Globus Connect Personal (https://www.globus.org/globus-connect-personal); Each model has approximately 3,000 building input descriptors that can be extracted. Please see the EnergyPlus(v9.4) 2,784-page Input/Output Reference Guide (https://energyplus.net/sites/all/modules/custom/nrel_custom/pdfs/pdfs_v9.4.0/InputOutputReference.pdf) for everything that can be retrieved or simulated from these models. These models were derived from the following metadata, which is not included in this dataset: 1. ID - unique building ID 2. County - county name 3. State - state name 4. CZ - ASHRAE Climate Zone designation 5. Clim_Zone - text label of climate zone 6. est_year - estimated year of construction 7. est_commercial - estimated building type (0=residential, 1=commercial) 8. Centroid - building center location in latitude/longitude (from Footprint2D) 9. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 10. Height - building height (meters) 11. Area2D - footprint area (ft2) 12. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 13. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 14. NumFloors - number of floors (above-grade) 15. Area - estimate of total conditioned floor area (ft2) 16. Standard - building vintage. These models are made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy's (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). This research used resources of the Argonne Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC02-06CH11357. Please cite as: New, Joshua R., Adams, Mark, Bass, Brett, Berres, Anne, and Clinton, Nicholas (2021). 'Model America - data and models of every U.S. building. [Data set].' Constellation, doi.ccs.ornl.gov/ui/doi/339, April 14, 2021

24 POWER TRANSMISSION AND DISTRIBUTION↗

Modeling Data Flows with Network Calculus in Cyber-Physical Systems: Enabling Feature Analysis for Anomaly Detection Applications

The electric grid is becoming increasingly cyber-physical with the addition of smart technologies, new communication interfaces, and automated grid-support functions. Because of this, it is no longer sufficient to only study the physical system dynamics, but the cyber system must also be monitored as well to examine cyber-physical interactions and effects on the overall system. To address this gap for both operational and security needs, cyber-physical situational awareness is needed to monitor the system to detect any faults or malicious activity. Techniques and models to understand the physical system (the power system operation) exist, but methods to study the cyber system are needed, which can assist in understanding how the network traffic and changes to network conditions affect applications such as data analysis, intrusion detection systems (IDS), and anomaly detection. In this paper, we examine and develop models of data flows in communication networks of cyber-physical systems (CPSs) and explore how network calculus can be utilized to develop those models for CPSs, with a focus on anomaly and intrusion detection. This provides a foundation for methods to examine how changes to behavior in the CPS can be modeled and for investigating cyber effects in CPSs in anomaly detection applications.

97 MATHEMATICS AND COMPUTING↗