Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Northern Pacific Turbulence Intensity Model Data in Observational Space

The dataset archives model-simulated turbulence intensity and meteorological profiles and timeseries at the lidar buoys sites off the coast of California (Humboldt and Morro Bay). The simulated data are interpolated in time and/or space according to observed quantities. The simulations were carried out for the north Pacific region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing.

17 WIND ENERGY↗

Gulf of Mexico Turbulence Intensity Model Data in Observational Space

The dataset archives model-simulated turbulence intensity and meteorological profiles and timeseries at the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars. The simulated data are interpolated in time and/or space according to observed quantities. The simulations were carried out for the Gulf of Mexico region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing.

17 WIND ENERGY↗

Data-model files associated with the manuscript "Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California"

This package contains the data, simulation setups, notebooks and figures used in “Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California” (Xu et al., 2025). In this study, we selected Elkhorn Slough, a tidal estuary, in California, to investigate the impact of wetland restoration and sea level rise on coastal hydrology using the process-based coastal hydrologic model, Advanced Terrestrial Simulator (ATS), informed by site-specific data. We designed a novel modeling workflow for incorporating wetland restoration features into land cover and soil properties for the model parameterization. The validation results demonstrate a strong agreement between modeled and observed data. We studied the characteristics of coastal watershed hydrology, then focused on the surface water dynamics at two wetland sites within Elkhorn Slough, a reference site and a restored site. Our simulation results indicate that the restored site successfully maintains surface elevation, resulting in reduced surface inundation. We also examined the impact of wetland restoration under expected sea level rise over the next few decades. The low-lying Yampah Marsh, the reference site, is likely to be inundated due to future sea level rise when highest tides arrive; while a higher percentage of Hester Marsh, the restored site, would retain marsh vegetation in coming decades, regardless of tidal conditions. Our study provides important information for examining the outcome of restoration practices that include surface elevation in tidal wetlands under climate changes.Several files can be found from this data package.1. README.md: This file describes the title, journal, co-authors, abstract, repository structure and model version.2. Simulation_Setups.zip: The file contains the model configuration files (XML format) for ATS. 3. Notebooks.zip: The file contains the Jupyter notebooks for generating the pre- and post-restoration meshes and the meshes of future scenarios. 4. Figures.zip: The file contains the figures used in the manuscript.5. Data.zip: The file contains the data used to drive the model simulations, including watershed and wetlands boundaries, mesh files and references to additional datasets (e.g., meteorological forcing, tidal dataset, DEMs, land cover, soil properties). Also, it contains water level observations at the restored wetland.

54 ENVIRONMENTAL SCIENCES↗

The future of Earth system prediction: Advances in model-data fusion

Predictions of the Earth system, such as weather forecasts and climate projections, require models informed by observations at many levels. Some methods for integrating models and observations are very systematic and comprehensive (e.g., data assimilation), and some are single purpose and customized (e.g., for model validation). We review current methods and best practices for integrating models and observations. We highlight how future developments can enable advanced heterogeneous observation networks and models to improve predictions of the Earth system (including atmosphere, land surface, oceans, cryosphere, and chemistry) across scales from weather to climate. As the community pushes to develop the next generation of models and data systems, there is a need to take a more holistic, integrated, and coordinated approach to models, observations, and their uncertainties to maximize the benefit for Earth system prediction and impacts on society.

54 ENVIRONMENTAL SCIENCES↗

Aerial and Processed Model Data Representing As-built Conditions in Coastal Port Arthur, Texas in May 2025

This dataset was collected by the Co-Design Team of the Southeast Texas Urban Integrated Field Lab, a research initiative led by the University of Texas at Austin and funded by the U.S. Department of Energy. The broader project focuses on developing climate-resilient design solutions for the Beaumont–Port Arthur region, with more information available at www.setx-uifl.org. Our team conducted aerial surveys of the Port Arthur coastal neighborhood in May 2025, before the start of construction scheduled for Summer 2026. These pre-construction datasets are designed to facilitate comparative analyses, including pre- and post-construction assessments and simulated inundation scenario evaluations. Aerial images were captured using DroneDeploy autonomous flight systems, with imagery processed through the DroneDeploy engine. All original aerial photographs are provided in JPG format and organized in zipped folders by area. The processed data package includes: 3D surface models Orthomosaics Geospatial and topographic mappings Point clouds For guidance on file contents, structure, and recommended usage, please refer to the included README file.

2D mapping↗

Model data for Flood Frequency Analysis using Stochastic Storm Transposition and an Integrated Surface-Subsurface Hydrological Model

This archived provides scripts and input files used for the implementation of a novel approach to conduct process-based Flood Frequency Analysis using a Stochastic Storm Transposition (SST) and an Integrated Surface-Subsurface Hydrological Model (ISSHM). As a proof-of-concept, this study uses the ISSHM, Advanced Terrestrial Simulator (Amanzi-ATS) model, and the SST model, RainyDay, to conduct flood frequency analysis by simulating the flood response to 5,000 annual synthetic storm events in a ~2000 km2 Southeast Texas watershed.The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include TXT, CSV, DAT, SBATCH, SHP, TIF, NetCDF, and HDF5 files, which can be read through Python scripts. The input files for the ATS model and RainyDay model have .XML and .SST extensions, respectively, and can be edited in any commonly used text editors.This archive contains:* Scripts and data files essential for generating the ATS model input. It uses the Watershed Workflow package to produce both mesh and ATS input files. * Jupyter notebooks designated for the ATS model evaluation, covering both long-term simulations and 40 rainfall-runoff events.* Input files required to simulate SST storm events using RainyDay.

54 ENVIRONMENTAL SCIENCES↗

Downscaled Earth System Model Data for Resilient Energy System Planning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. In this presentation, we explore the output characteristics of the dataset and various validation analyses. We also present and discuss plans for the integration of this data into power system planning models using a decision-making under deep uncertainty (DMDU) methodology.

97 MATHEMATICS AND COMPUTING↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Revolutionizing Materials Design: The Intersection of Quantum Mechanics and Data Modeling

The field of materials design is currently experiencing a notable evolution, driven by the convergence of sophisticated computational methodologies based on first principles and data-driven modeling approaches. I will review our recent endeavors employing AI/ML to expedite first-principles simulations and mitigate traditional methods' temporal and spatial limitations. Central to our efforts is developing and utilizing ML interatomic potentials (MLPs) across a diverse spectrum of materials. We show that MLPs serve as invaluable tools for navigating the complexities of the simulations, such as understanding the behavior of MgO at extreme environments of ~1 terapascal and temperatures >10,000 Kelvin. Moreover, we show that MLPs can provide precise details of the intricate dynamics governing the oxidation processes of binary alloy systems due to the competition between surface segregation and reconstruction tendencies. In summation, advancements in MLPs open the door to fresh possibilities in material modeling and, ultimately, discovery.

Saidi, Wissam↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

TEAMER: Triton Systems Oscillating Water Column Modeling Data and Report

This dataset provides the output of six Wave Energy Converter Simulator (WEC-Sim) simulations and accompanying documentation for the modeling of Triton Systems' oscillating water column (OWC) system at tank scale (validated using available data for tuning the model, Tests 1-2) and deployment scale (for which no validation data is available, Tests 4-6). Included are the output data in a MATLAB file structure, a comprehensive report on the modeling and design of the Triton OWC system, and a link to the WEC-Sim GitHub page. This work was supported by funding from TEAMER RFTS 5 (Request for Technical Support).

16 TIDAL AND WAVE POWER↗

Data-model files associated with the manuscript "The Effects of Spatial and Temporal Resolution of Gridded Meteorological Forcing on Watershed Hydrological Responses" (Shuai et al., 2022 HESS)

This data package contains the model inputs and outputs used in "The Effects of Spatial and Temporal Resolution of Gridded Meteorological Forcing on Watershed Hydrological Responses" (Shuai et al., 2022 HESS). The data.zip file contains the data used to drive the model simulations. The model.zip file contains the XML input file for ATS. The notebook.zip file contains the Jupyter notebooks for pre- and post- processing model results. The figures.zip file contains the raw figures associated with the manuscript.Meteorological forcing plays a critical role in accurately simulating the watershed hydrological cycle. With the advancement of high-performance computing and the development of integrated watershed models, simulating the watershed hydrological cycle at high temporal (hourly to daily) and spatial resolution (10s of meters) has become efficient and computationally affordable. These hyperresolution watershed models require high resolution of meteorological forcing as model input to ensure the fidelity and accuracy of simulated responses. In this study, we utilized the Advanced Terrestrial Simulator (ATS), an integrated watershed model, to simulate surface and subsurface flow and land surface processes using unstructured meshes at the Coal Creek Watershed near Crested Butte (Colorado). We compared simulated watershed hydrologic responses including streamflow, and distributed variables such as evapotranspiration, snow water equivalent (SWE), and groundwater table driven by three publicly available, gridded meteorological forcings (GMFs) -- Daily Surface Weather and Climatological Summaries (Daymet), Parameter-elevation Regressions on Independent Slopes Model (PRISM), and North American Land Data Assimilation System (NLDAS). By comparing various spatial resolutions (ranging from 400 m to 4 km) of PRISM, the simulated streamflow only becomes marginally worse when spatial resolution of meteorological forcing is coarsened to 4 km (or 30% of the watershed area). However, the 4 km resolution has much worse performance than finer resolution in spatially distributed variables such as SWE. Using temporally disaggregated PRISM, we compared models forced by different temporal resolutions (hourly to daily), sub-daily resolution preserves the dynamic watershed responses (e.g., diurnal fluctuation of streamflow) that are absent in results forced by daily resolution. Conversely, the simulated streamflow shows better performance using daily resolution compared to that using sub-daily resolution. Our findings suggest that the choice of GMF and its spatiotemporal resolution depends on the quantity of interest and its spatial and temporal scale, which may have important implications on model calibration and watershed management decisions.

54 ENVIRONMENTAL SCIENCES↗

Data for "Examining Organic Acid Production Potential and Growth-Coupled Strategies in Issatchenkia orientalis Using Constraint-Based Modeling"

Growth-coupling product formation can facilitate strain stability by aligning industrial objectives with biological fitness. Organic acids make up many building block chemicals that can be produced from sugars obtainable from renewable biomass. Issatchenkia orientalis is a yeast strain tolerant to acidic conditions and is thus a promising host for industrial production of organic acids. Here, we use constraint-based methods to assess the potential of computationally designing growth-coupled production strains for I. orientalis that produce 22 different organic acids under aerobic or microaerobic conditions. We explore native and engineered pathways using glucose or xylose as the carbon substrates as proxy constituents of hydrolyzed biomass. We identified growth-coupled production strategies for 37 of the substrate-product pairs, with 15 pairs achieving production for any growth rate. We systematically assess the strain design solutions and categorize the underlying principles involved.

Bioproducts↗

Second-generation downscaled earth system model data using generative machine learning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. As with the first Sup3rCC version, the data still include temperature, wind speed and direction at multiple heights, pressure, three components of downwelling solar radiation, and relative humidity—all at 4-kilometer (km) hourly resolution over the contiguous United States. This is a 25x spatial enhancement and 24x temporal enhancement of the source 100-km daily-average ESM data. This extension of the Sup3rCC dataset includes data from six ESMs from two shared socioeconomic pathways (SSPs) totaling 400 years of data with multiple future projections of changing meteorological conditions. The scenario selection was based on a structured evaluation of historical ESM skill and comprehensive representation of possible trajectories of future climate change in temperature, humidity, precipitation, solar irradiance, and near-surface wind speeds. The inclusion of multiple future projections is intended to enable users to assess key drivers of un 36 certainty and variability. All data are double-bias corrected, resulting in a product that can be used out-of-the-box for energy system analysis with minimal historical bias. The potential applications of Sup3rCC data extend to various topics in renewable energy resource assessment, energy systems modeling, and grid resilience studies. High-resolution future meteorological projections are critical for evaluating the effects of changing meteorological conditions on renewable energy generation, energy demand, and for optimizing energy storage and grid infrastructure. The 4-km hourly resolution of the downscaled data enables understanding of spatial and temporal variability at the scales necessary for energy system operational planning. In addition, the dataset can support risk assessments by providing detailed information on possible future extreme weather events and long-term meteorological variability at scales relevant to energy infrastructure. By offering an enhanced representation of possible future meteorological conditions, the second-generation Sup3rCC dataset enables more precise modeling of energy resilience and adaptation strategies in response to changing meteorological conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Landmark-embedded Gaussian process with applications for functional data modeling

In practice, we often need to infer the value of a target variable from functional observation data. A challenge in this task is that the relationship between the functional data and the target variable is very complex: the target variable not only influences the shape but also the location of the functional data. In addition, due to the uncertainties in the environment, the relationship is probabilistic, that is, for a given fixed target variable value, we still see variations in the shape and location of the functional data. To address this challenge, we present a landmark-embedded Gaussian process model that describes the relationship between the functional data and the target variable. A unique feature of the model is that landmark information is embedded in the Gaussian process model so that both the shape and location information of the functional data are considered simultaneously in a unified manner. Gibbs-Metropolis-Hasting algorithm is used for model parameters estimation and target variable inference. The performance of the proposed framework is evaluated by extensive numerical studies and a case study of nano-sensor calibration.

42 ENGINEERING↗

Data for Impact of Vertical and Seasonal Variation in Leaf Traits on Simulating Soybean Canopy Photosynthesis via 1D and 3D Modeling

Accurate modeling of photosynthesis is crucial for predicting crop productivity and quantifying the carbon cycle in agroecosystems. Leaf traits are essential inputs for modeling canopy photosynthesis. Yet, many existing models still use fixed plant functional type (PTF)-based values to parameterize leaf traits under a big-leaf or two-big-leaf assumption, neglecting their vertical profiles and seasonal changes. This simplification may introduce significant uncertainties in estimating gross primary productivity (GPP). In this study, we simulated soybean GPP and tested the effects of vertical and seasonal variation in three key leaf photosynthetic traits: the maximum carboxylation rate at 25 °C (Vcmax25), leaf chlorophyll content (LCC), and leaf mass per area (LMA) in the 1D-SCOPE and 3D-Helios models. Weekly field measurements were conducted during the growing season of 2024 to support the simulation. We designed ten leaf trait parameterization schemes by incorporating different combinations of vertical profiles and seasonal changes, while assuming homogeneous canopy architecture in both models. Our results revealed that Vcmax25 vertical and seasonal variation had the strongest influence on simulated GPP in both 1D and 3D models, while LCC and LMA effects were minimal. Particularly, the scheme with an empirically parameterized Vcmax25 profile achieved comparable performance to the scheme with the measured Vcmax25 profile. Both 1D-SCOPE and 3D-Helios accurately modeled GPP (SCOPE: R2 = 0.87, Bias = 0.55 µmol m⁻² s⁻¹; Helios: R2 = 0.9, Bias = 0.22 µmol m⁻² s⁻¹) under the most complex scheme, and their responses to vertical and seasonal variation in leaf traits were consistent, demonstrating the robustness of our findings. Based on our findings, we propose a scalable framework for parameterizing leaf traits to improve GPP simulations. This study contributes to improving the representation of leaf trait dynamics in canopy-level photosynthesis models, potentially enhancing our ability to predict crop productivity and understand agroecosystem carbon dynamics.

Photosynthesis↗

YEASTRACT+: a portal for the exploitation of global transcription regulation and metabolic model data in yeast biotechnology and pathogenesis

Abstract YEASTRACT+ (http://yeastract-plus.org/) is a tool for the analysis, prediction and modelling of transcription regulatory data at the gene and genomic levels in yeasts. It incorporates three integrated databases: YEASTRACT (http://yeastract-plus.org/yeastract/), PathoYeastract (http://yeastract-plus.org/pathoyeastract/) and NCYeastract (http://yeastract-plus.org/ncyeastract/), focused on Saccharomyces cerevisiae, pathogenic yeasts of the Candida genus, and non-conventional yeasts of biotechnological relevance. In this release, YEASTRACT+ offers upgraded information on transcription regulation for the ten previously incorporated yeast species, while extending the database to another pathogenic yeast, Candida auris. Since the last release of YEASTRACT+ (January 2020), a fourth database has been integrated. CommunityYeastract (http://yeastract-plus.org/community/) offers a platform for the creation, use, and future update of YEASTRACT-like databases for any yeast of the users’ choice. CommunityYeastract currently provides information for two Saccharomyces boulardii strains, Rhodotorula toruloides NP11 oleaginous yeast, and Schizosaccharomyces pombe 972h-. In addition, YEASTRACT+ portal currently gathers 304 547 documented regulatory associations between transcription factors (TF) and target genes and 480 DNA binding sites, considering 2771 TFs from 11 yeast species. A new set of tools, currently implemented for S. cerevisiae and C. albicans, is further offered, combining regulatory information with genome-scale metabolic models to provide predictions on the most promising transcription factors to be exploited in cell factory optimisation or to be used as novel drug targets. The expansion of these new tools to the remaining YEASTRACT+ species is ongoing.

Teixeira, Miguel Cacho (ORCID:0000000256766174)↗