Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “knowledge representation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data-driven modeling and control of dynamical systems using Koopman and Perron-Frobenius operators

This dissertation studies the data-driven modeling and control problem of nonlinear systems by exploiting the linear operator theoretic framework involving Koopman and Perro-Frobenius operator. A systematic linear-operator based controller design procedure has been established, which can be used to solve a variety of nonlinear control problems, including feedback stabilization using control Lyapunov functions, optimal quadratic regulation using Koopman eigenfunctions and convex optimization formulation of optimal control problem using P-F and Koopman operator approximation. As the core of data-driven modeling, we first propose a new algorithm for the finite-dimensional approximation of the linear transfer Koopman and Perron-Frobenius operator from time-series data. We argue that the existing approach for the finite-dimensional approximation of these transfer operators such as Dynamic Mode Decomposition (DMD) and Extended Dynamic Mode Decomposition (EDMD) do not capture two important properties of these operators, namely positivity and Markov property. The algorithm we propose preserves these two properties. We call the proposed algorithm as naturally structured DMD (NSDMD) since it retains the inherent properties of these operators. Naturally structured DMD algorithm leads to a better approximation of the steady-state dynamics of the system regarding computing Koopman and Perron- Frobenius operator eigenfunctions and eigenvalues. However, preserving positivity property is critical for capturing the real transient dynamics of the system. This positivity property of the transfer operators and it's finite-dimensional approximation play an important role for controller and estimator design of nonlinear systems. To solve the feedback stabilization problem for nonlinear control systems, we tried to take advantage of the Koopman operator framework. The Koopman operator approach provides a linear representation for a nonlinear dynamical system and a bilinear representation for a nonlinear control system. The problem of feedback stabilization of a nonlinear control system is then transformed to the stabilization of a bilinear control system. We propose a control Lyapunov function (CLF)-based approach for the design of stabilizing feedback controllers for the bilinear system. The search for finding a CLF for the bilinear control system is formulated as a convex optimization problem. This leads to a schematic procedure for designing CLF-based stabilizing feedback controllers for the bilinear system and hence the original nonlinear system. Another advantage of the proposed controller design approach outlined in this dissertation is that it does not require explicit knowledge of system dynamics. In particular, the bilinear representation of a nonlinear control system in the Koopman eigenfunction space can be obtained from time-series data. Next, we study the optimal quadratic regulation problem for nonlinear systems. The linear operator theoretic framework involving the Koopman operator is used to lift the dynamics of nonlinear control system to an infinite-dimensional bilinear system. The optimal quadratic regulation problem for nonlinear system is formulated in terms of the finite-dimensional approximation of the bilinear system. A convex optimization-based approach is proposed for solving the quadratic regulator problem for bilinear system. We applied a variety of examples and compared the simulation results between our framework and conventional LQR control using linearized model. For more general optimal control problems, we provide a density-function based convex formulation for the optimal control problem of the nonlinear system. The convex formulation relies on the duality result in the stability theory of a dynamical system involving density function and Perron-Frobenius operator. The optimal control problem is formulated as an infinite-dimensional convex optimization program. The finite-dimensional approximation of the optimization problem relies on the recent advances made in the data-driven computation of the Koopman operator, which is dual to the Perron-Frobenius operator. Simulation results are presented to demonstrate the application of the developed framework.

Huang, Bowen↗

COVID19 Disease Map, a computational knowledge repository of virus–host interaction mechanisms

We need to effectively combine the knowledge from surging literature with complex datasets to propose mechanistic models of SARS-CoV-2 infection, improving data interpretation and predicting key targets of intervention. Here, we describe a large-scale community effort to build an open access, interoperable and computable repository of COVID-19 molecular mechanisms. The COVID-19 Disease Map (C19DMap) is a graphical, interactive representation of disease-relevant molecular mechanisms linking many knowledge sources. Notably, it is a computational resource for graph-based analyses and disease modelling. To this end, we established a framework of tools, platforms and guidelines necessary for a multifaceted community of biocurators, domain experts, bioinformaticians and computational biologists. The diagrams of the C19DMap, curated from the literature, are integrated with relevant interaction and text mining databases. We demonstrate the application of network analysis and modelling approaches by concrete examples to highlight new testable hypotheses. This framework helps to find signatures of SARS-CoV-2 predisposition, treatment response or prioritisation of drug candidates. Such an approach may help deal with new waves of COVID-19 or similar pandemics in the long-term perspective.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling dust mineralogical composition: sensitivity to soil mineralogy atlases and their expected climate impacts

Soil dust aerosols are a key component of the climate system, as they interact with short- and long-wave radiation, alter cloud formation processes, affect atmospheric chemistry and play a role in biogeochemical cycles by providing nutrient inputs such as iron and phosphorus. The influence of dust on these processes depends on its physicochemical properties, which, far from being homogeneous, are shaped by its regionally varying mineral composition. The relative amount of minerals in dust depends on the source region and shows a large geographical variability. However, many state-of-the-art Earth system models (ESMs), upon which climate analyses and projections rely, still consider dust mineralogy to be invariant. The explicit representation of minerals in ESMs is more hindered by our limited knowledge of the global soil composition along with the resulting size-resolved airborne mineralogy than by computational constraints. In this work we introduce an explicit mineralogy representation within the state-of-the-art Multiscale Online Nonhydrostatic AtmospheRe CHemistry (MONARCH) model. We review and compare two existing soil mineralogy datasets, which remain a source of uncertainty for dust mineralogy modeling and provide an evaluation of multiannual simulations against available mineralogy observations. Soil mineralogy datasets are based on measurements performed after wet sieving, which breaks the aggregates found in the parent soil. Our model predicts the emitted particle size distribution (PSD) in terms of its constituent minerals based on brittle fragmentation theory (BFT), which reconstructs the emitted mineral aggregates destroyed by wet sieving. Our simulations broadly reproduce the most abundant mineral fractions independently of the soil composition data used. Feldspars and calcite are highly sensitive to the soil mineralogy map, mainly due to the different assumptions made in each soil dataset to extrapolate a handful of soil measurements to arid and semi-arid regions worldwide. For the least abundant or more difficult-to-determine minerals, such as iron oxides, uncertainties in soil mineralogy yield differences in annual mean aerosol mass fractions of up to ~ 100 %. Although BFT restores coarse aggregates including phyllosilicates that usually break during soil analysis, we still identify an overestimation of coarse quartz mass fractions (above 2 µm in diameter). In a dedicated experiment, we estimate the fraction of dust with undetermined composition as given by a soil map, which makes up ~ 10 % of the emitted dust mass at the global scale and can be regionally larger. Changes in the underlying soil mineralogy impact our estimates of climate-relevant variables, particularly affecting the regional variability of the single-scattering albedo at solar wavelengths or the total iron deposited over oceans. All in all, this assessment represents a baseline for future model experiments including new mineralogical maps constrained by high-quality spaceborne hyperspectral measurements, such as those arising from the NASA Earth Surface Mineral Dust Source Investigation (EMIT) mission.

54 ENVIRONMENTAL SCIENCES↗

Qualitative trend analysis based on a mixed-integer representation

Shape constrained spline fitting is a useful method to impose prior knowledge onto flexible semi-parametric models during parameter estimation. Most typically, the function shape is imposed through order restrictions on the regression coefficients. The intended shape is considered known or selected based on heuristic rules. In this study, we present a method to estimate the optimal set of order restrictions to segment a univariate data series into episodes with distinct shapes. This is also known as the qualitative trend analysis (QTA) problem. The obtained solution uses a trade-off between lack-of-fit and model complexity. Further, our practical implementation takes inspiration from the generalized order restricted information criterion (GORIC) for inequality-constrained model selection. From this, one learns (a) that QTA can be formulated as a mixed-integer quadratic program (MIQP) and (b) that the newly proposed mixed order restricted information criterion (MORIC) enables optimal segmentation. This is illustrated through didactic case studies.

42 ENGINEERING↗

Soil Carbon Saturation: What Do We Really Know?

Managing soils to increase organic carbon storage presents a potential opportunity to mitigate and adapt to global change challenges, while providing numerous co-benefits and ecosystem services. However, soils differ widely in their potential for carbon sequestration, and knowledge of biophysical limits to carbon accumulation may aid in informing priority regions. Consequently, there is great interest in assessing whether soils exhibit a maximum capacity for storing organic carbon, particularly within organo–mineral associations given the finite nature of reactive minerals in a soil. While the concept of soil carbon saturation has existed for over 25 years, recent studies have argued for and against its importance. Here, we summarize the conceptual understanding of soil carbon saturation at both micro- and macro-scales, define key terminology, and address common concerns and misconceptions. We review methods used to quantify soil carbon saturation, highlighting the theory and potential caveats of each approach. Critically, we explore the utility of the principles of soil carbon saturation for informing carbon accumulation, vulnerability to loss, and representations in process-based models. We highlight key knowledge gaps and propose next steps for furthering our mechanistic understanding of soil carbon saturation and its implications for soil management.

Environmental sciences↗

Identifying intragenic functional modules of genomic variations associated with cancer phenotypes by learning representation of association networks

Background Genome-wide Association Studies (GWAS) aims to uncover the link between genomic variation and phenotype. They have been actively applied in cancer biology to investigate associations between variations and cancer phenotypes, such as susceptibility to certain types of cancer and predisposed responsiveness to specific treatments. Since GWAS primarily focuses on finding associations between individual genomic variations and cancer phenotypes, there are limitations in understanding the mechanisms by which cancer phenotypes are cooperatively affected by more than one genomic variation. Results This paper proposes a network representation learning approach to learn associations among genomic variations using a prostate cancer cohort. The learned associations are encoded into representations that can be used to identify functional modules of genomic variations within genes associated with early- and late-onset prostate cancer. The proposed method was applied to a prostate cancer cohort provided by the Veterans Administration’s Million Veteran Program to identify candidates for functional modules associated with early-onset prostate cancer. The cohort included 33,159 prostate cancer patients, 3181 early-onset patients, and 29,978 late-onset patients. The reproducibility of the proposed approach clearly showed that the proposed approach can improve the model performance in terms of robustness. Conclusions To our knowledge, this is the first attempt to use a network representation learning approach to learn associations among genomic variations within genes. Associations learned in this way can lead to an understanding of the underlying mechanisms of how genomic variations cooperatively affect each cancer phenotype. This method can reveal unknown knowledge in the field of cancer biology and can be utilized to design more advanced cancer-targeted therapies.

60 APPLIED LIFE SCIENCES↗

NGEE Arctic Integrated Modeling (IM2): Improved subgrid hillslope hydrologic connectivity

This data product represents the integration of new code capability for arctic tundra hillslope hydrologic processes into the Energy Exascale Earth System Model (E3SM), through the E3SM Land Model (ELM) component. This code integration is the result of collaborative effort between the NGEE Arctic project and the E3SM project. The current ELM represents water movement primarily through vertical processes, such as precipitation, canopy interception, evaporation, infiltration, and soil water movement. Lateral water movement—such as surface runoff, subsurface flow, and river transport—plays a significant role in the hydrological cycle, especially in regions with varied topography. While E3SM includes a runoff routing component representing water transport in the river network, the lateral transport of water at the subgrid scale within the land model has previously not been taken into account. With the recent development of topographic units within the ELM subgrid data structure, there is an opportunity to simulate hillslope hydrologic connectivity by introducing water transport along topographic gradients. We expect that more realistic representation of hillslope hydrologic processes will lead to improved predictions of both soil water content and river network flows. Lateral transport of water at and near the surface is represented as a sub-grid process in this new code development. Water is tracked as it moves from higher to lower elevations within a gridcell. This capability uses the nested hierarchical sub-grid scheme within ELM to connect water fluxes from sub-grid elements with higher elevation to those with lower elevation. This data record consists of a single document (pdf format) that describes the theoretical basis for the hillslope hydrology processes added to ELM, and describes the modifications made to the ELM code. The Methods section of this metadata record includes a link to the public E3SM code repository where the exact code modifications as integrated in E3SM can be accessed. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Thornton, Peter E [ORNL] (ORCID:0000000247595158)↗

Accounting for herbaceous communities in process‐based models will advance our understanding of “grassy” ecosystems

Abstract Grassland and other herbaceous communities cover significant portions of Earth's terrestrial surface and provide many critical services, such as carbon sequestration, wildlife habitat, and food production. Forecasts of global change impacts on these services will require predictive tools, such as process‐based dynamic vegetation models. Yet, model representation of herbaceous communities and ecosystems lags substantially behind that of tree communities and forests. The limited representation of herbaceous communities within models arises from two important knowledge gaps: first, our empirical understanding of the principles governing herbaceous vegetation dynamics is either incomplete or does not provide mechanistic information necessary to drive herbaceous community processes with models; second, current model structure and parameterization of grass and other herbaceous plant functional types limits the ability of models to predict outcomes of competition and growth for herbaceous vegetation. In this review, we provide direction for addressing these gaps by: (1) presenting a brief history of how vegetation dynamics have been developed and incorporated into earth system models, (2) reporting on a model simulation activity to evaluate current model capability to represent herbaceous vegetation dynamics and ecosystem function, and (3) detailing several ecological properties and phenomena that should be a focus for both empiricists and modelers to improve representation of herbaceous vegetation in models. Together, empiricists and modelers can improve representation of herbaceous ecosystem processes within models. In so doing, we will greatly enhance our ability to forecast future states of the earth system, which is of high importance given the rapid rate of environmental change on our planet.

59 BASIC BIOLOGICAL SCIENCES↗

Post-fire soil respiration in late growing season (2023 and 2024), Kougarok Fire Complex, Seward Peninsula, Alaska

Field soil respiration data collected in 2023 and 2024 from burned and unburned tussock tundra sites in the Kougarok Fire Complex, near Nome, on the Seward Peninsula of Alaska. Specifically, we measured soil properties and late-growing season CO2 fluxes in patches of unique plant functional types (forbs, shrubs, and graminoids) across two years in tundra recovering from repeated wildfires over the decade. The goal was to identify the main drivers of soil respiration in Arctic tundra underlain by discontinuous permafrost that is recovering from two recent, repeated wildfires that differed in fire age and number of times burned, thereby resulting in different levels of vegetation and subsurface property changes (i.e., successional trajectories). There are five files in *.csv format with one data file and four data description files including data, dictionary, methods, terminology, and file-level metadata. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), is a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic Phase 3 project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Santos, Fernanda [ORNL] (ORCID:0000000191555623)↗

NGEE Arctic Integrated Modeling (IM3): Improved snow-vegetation interaction

This data product represents the integration of new code capability for arctic tundra snow-vegetation-terrain interactions into the Energy Exascale Earth System Model (E3SM), through the E3SM Land Model (ELM) component. This code integration is the result of collaborative effort between the NGEE Arctic project and the E3SM project. The NGEE Arctic project developed a total of six Integrated Modeling (IM) modules informed by observations and experiments. New ELM capability represented by this data product (IM3) falls into three categories: 1) Downscaling from gridcell to topographic unit level when working through the existing coupler bypass code. 2) Four new parameters (taper, stocking, bendresist, and vegshape) have been added to ELM to allow for flexible definition of snow-vegetation interactions. 3) Vegshape and bendresist parameters are used to calculate the fraction of leaf area and/or stem area buried by snow for a given snow depth. This data record consists of a single document (pdf format) that describes the theoretical basis for the snow-vegetation-terrain interactions added to ELM, and describes the modifications made to the ELM code. The Methods section of this metadata record includes a link to the public E3SM code repository where the exact code modifications as integrated in E3SM can be accessed. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Thornton, Peter E [ORNL] (ORCID:0000000247595158)↗

Multi-level optimization with the koopman operator for data-driven, domain-aware, and dynamic system security

Cyber-Physical Systems (CPSs) like the power grid are critically important but also increasingly vulnerable; ensuring reliable system operation in the face of disruptions is becoming more and more challenging. Multi-Level Optimization (MLO) is a powerful way to model adversarial interactions, which naturally makes it applicable to studying CPS security. However, MLO typically does not address underlying system dynamics, and incorporating nonlinear dynamics is generally infeasible. In this paper, we show how to combine MLO with the Koopman Operator (KO) to remedy this. The KO maps nonlinear dynamics to a lifted space in which those dynamics are linear, thus making it ideal for use with MLO. Moreover, the structure of the KO also provides convenient ways to incorporate domain knowledge into the data-driven process of learning the KO representation of a given system. Here we then demonstrate the use of MLO-KO on a small example problem taken from the power grid domain, discuss the scalability and computational cost of MLO-KO, and identify future research directions for this work.

42 ENGINEERING↗

QLiG: Query Like a Graph For Subgraph Matching

A graph is a natural and flexible modeling approach to represent entities and relationships between them in real-world. A Knowledge Graphs (KG) is a specialized graph with formal and structured representation of facts, relationships, annotated with semantic descriptions. Subgraph matching is one of the fundamental graph problems to identify relationships, interactions and activities of interest within a large graph. A query specification is a collection of abstract components, operations, and constraints to express a pattern. The specification can be implemented in different ways based on underlying data model. Various graph query specifications have been developed over the years and have led to the development of different open-sourced and vendor-specific query languages. Such specification are modeled as an extension of relational algebra used to develop relational query languages such as SQL. Such relational concepts do not inherently support graph queries. There is a need to represent graph queries in terms on graph-based components to expedite query construction by non-database experts. We present a graph-based query approach QLiG (pronounced cleeg), to perform subgraph matching in Labeled Property Graph. We present the query specifications, salient features, and a use case to show functional examples.

Purohit, Sumit↗

TriGORank: A Gene Ontology Enriched Learning-to-Rank Framework for Trigenic Fitness Prediction

Machine learning (ML) has been gaining interest in the metabolic engineering community as a means to automate prediction tasks. In this work, we introduce and study the task of using ML to recommend high-fitness triplet mutants as candidates for wet-lab experiments. We first utilize individual fitness and digenic fitness scores as features and train machine learning models that produce a ranked list, from high to low fitness scores, for triplet gene mutants of S. cerevisiae. Then, we incorporate prior metabolic knowledge from an existing gene ontology, by designing a novel graph representation and deducing features that can capture gene similarity and gene interactions. Lastly, experimental results show that our proposed gene ontology enriched model, termed TriGORank, improves both performance and explainability.

Labhishetty, Sahiti↗

Foundations of automatic feature extraction at LHC–point clouds and graphs

Abstract Deep learning algorithms will play a key role in the upcoming runs of the Large Hadron Collider (LHC), helping bolster various fronts ranging from fast and accurate detector simulations to physics analysis probing possible deviations from the Standard Model. The game-changing feature of these new algorithms is the ability to extract relevant information from high-dimensional input spaces, often regarded as “replacing the expert” in designing physics-intuitive variables. While this may seem true at first glance, it is far from reality. Existing research shows that physics-inspired feature extractors have many advantages beyond improving the qualitative understanding of the extracted features. In this review, we systematically explore automatic feature extraction from a phenomenological viewpoint and the motivation for physics-inspired architectures. We also discuss how prior knowledge from physics results in the naturalness of the point cloud representation and discuss graph-based applications to LHC phenomenology.

Bhardwaj, Akanksha↗

Landscaper v1

Understanding the inner workings of machine learning models through their loss landscapes offers crucial insights into model properties, optimization dynamics, and generalizability. However, accessing these insights has traditionally required specialized mathematical expertise, limiting broader adoption. Landscaper is an open-source Python package designed to bridge this gap. Landscaper seamlessly integrates a suite of multi-dimensional loss landscape analyses with cutting-edge topological data analysis (TDA) methods. This powerful combination makes both fundamental loss landscape analysis and advanced TDA techniques accessible to the broader scientific ML community, without requiring deep pre-existing mathematical knowledge. Landscaper offers three key functionalities: * Construction: Builds detailed loss landscape representations through versatile low and high-dimensional sampling techniques. * Quantification: Applies advanced metrics, including a novel topological data analysis (TDA) based smoothness metric, enabling new perspectives on model behavior. * Visualization: Offers intuitive tools to visualize and interpret loss landscapes, providing actionable insights beyond traditional performance metrics.

Weber, Gunther [Lawrence Berkeley National Laborat↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Near-Surface Hydrology and Soil Properties Drive Heterogeneity in Permafrost Distribution, Vegetation Dynamics, and Carbon Cycling in a Sub-Arctic Watershed: Modeling Archive

This Modeling Archive is in support of a NGEE-Arctic publication: Shirley et al. (2022) “Near-Surface Hydrology and Soil Properties Drive Heterogeneity in Permafrost Distribution, Vegetation Dynamics, and Carbon Cycling in a Sub-Arctic Watershed". [DOI].The dataset contains outputs from the global sensitivity analysis (GSA) of the “ecosys” model as reported in Shirley et al. (2022). The study showed that discontinuous permafrost environments are characterized by complex feedback loops and strong spatial heterogeneity which is created by variability in near-surface hydrology and soil properties. Additionally, the study demonstrated that missing representation of sub-grid heterogeneity in terrestrial ecosystem models can lead to biased estimates of the high-latitude carbon budget. Included in this dataset are the factor values for each run in the GSA and the model outputs used in this study. Included are two *.csv data files and one *.pdf.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Multi-resolution Arctic Shrub Cover Dataset Derived from UAS and Airborne SfM and LiDAR (2013-2025)

We synthesized 177 unoccupied aerial system flights and 77 airborne flights across the Arctic and created a multi-resolution benchmark data of low-to-tall shrub fractional cover leveraging Structure-from-Motion and Light Detection and Ranging. The resulting dataset covered a total of 1899 km2 across Alaska, Western Canada, Sweden, and Siberian Arctic, including key sites from the Oro Arctic to the High Arctic. The dataset is organized into 6 primary data collection directories (“Abisko,” “AWI,” “ERE,” “Fairbanks,” “NGEE,” “Toolik”), each containing site and flight subdirectories. Flight directories include shrub cover rasters (*.tifs) at 1 m, 5 m, and 30 m resolution, the canopy height model at 1 m resolution (*.tifs), and a bounding box *.kml file. For the AWI, Abisko, NGEE, and Fairbanks collections, we also include the GCC raster at 1 m resolution (*.tif). Files are organized by Collection > Site > Flight Name > Data Files. Flight rasters are in the local UTM zone and the .kml files are in the geographic coordinate system EPSG 4326. We also include a .csv file that details the source datasets for every flight. The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research. Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.

canopy height model↗